Does AI Really Make You More Productive? The Evidence Just Got Shakier
- Lille My

- 11 minutes ago
- 5 min read

Image via Wikimedia Commons (public domain).
Welcome to this week's scan of how AI is reshaping research — curated from Nature, Science, PNAS and the leading business journals.
When "AI made me more productive" is really just how you defined the treatment
One of the most quoted claims about AI in science is that researchers who start using large language models publish substantially more. The difficulty is that nobody observes adoption directly, so studies infer it — typically from the first paper that a detector flags as LLM-assisted. In a new PNAS letter, Scientific production in the era of large language models: Outcome-triggered treatment timing and spurious event-study dynamics, the authors show that this inference rule is the whole problem. Because high-output months are mechanically more likely to contain a paper that gets detected, treatment timing becomes entangled with the outcome it is supposed to explain — and the event study produces a clean, convincing post-treatment jump even when no causal effect exists.
The demonstration is unusually direct. Using reconstructed arXiv data, the authors run the same design with random treatment assignment, with neutral (non-AI) keyword triggers, with inverted treatment, and over pre-ChatGPT placebo periods — and every variant reproduces the same shape of post-treatment gains. Simulations with a true effect of exactly zero reproduce the published pattern too. Their conclusion is blunt: first-detection timing alone manufactures the evidence. The critique targets Kusumegi et al., *Science* 390, 1240–1243 (2025), but it generalises to any design where the moment of "adoption" is read off the output stream itself. That covers a large and growing share of the AI-and-productivity literature, in scientometrics and in firm-level studies alike.
It lands in the same week as a useful reality check on the other end of the pipeline. Nature reports on a preprint in which an agentic system was set loose on the concepts of two computer-science papers; it developed them, but the original authors were unimpressed. Princeton's Sayash Kapoor, a co-author, told Nature: "I don't think full automation of open-ended research is on the horizon right now" (AI isn't ready to research itself). Taken together, the two stories describe a field where the ceiling on autonomous AI research is lower than advertised — and where the evidence that AI accelerates *human* research is more fragile than the headline numbers suggest.
Why it matters: If you are writing a paper, a grant, or a policy memo that leans on "AI raises research productivity by X%," check how adoption was measured before you cite the number. More broadly, this is a textbook case of an identification problem that arrives disguised as a data problem: when your treatment indicator is derived from the same process that generates your outcome, staggered event-study plots will look beautiful and mean nothing. The same trap is waiting in every study that dates AI adoption from AI-detected output — including the ones your own field is about to publish.
More from this week
Platform AI policies change creator behaviour — even when nobody is forced to use AI
Using two natural experiments on leading Chinese visual-arts platforms — Lofter launching an AI image generator, and Graffiti Kingdom prohibiting AI-generated artwork — the authors find creator activity fell after the pro-AI move and rose after the anti-AI move. Multiple tests point to stance signalling rather than tool features or enforcement: creators read the policy as a statement about the platform's commitment to human work and the future competitive environment. Effects were largest among higher-popularity, multi-homing and AI-averse creators, with replacement risk, perceived low quality and copyright infringement the dominant stated concerns.
Generative AI reduced price discrimination in housing valuations
Housing price discrimination — comparable homes valued lower in minority-dominant neighbourhoods — has long been attributed to human bias, and earlier work suggested conventional AI models, even fairness-targeted ones, do not fix it. This study compares AI-generated with human-generated housing prices across 284,749 U.S. properties and finds generative AI can actually alleviate the gap, then probes the mechanisms behind that counter-intuitive result. It is a rare case of a generative model reducing rather than laundering an existing bias, and the mechanism analysis is what makes it citable.
Simulating a bank run with synthetic depositors
The author builds a representative population of synthetic depositor agents with assigned demographic attributes, exposes them to a viral panic post, randomises bank communication interventions, and validates the model's responses against human benchmarks before generating systematic message variants. Estimated withdrawal propensities are then fed into a contagion model that propagates withdrawals across a proximity network. Direct, personalised communications with strong reassurances and explicit survival clauses substantially reduced withdrawal intent — an inexpensive way to pre-test crisis messaging that would be impossible to run on real depositors.
When AI accuracy parity meets strategic self-disclosure
This analytical paper examines how consumers' strategic decisions to disclose their identity interact with a firm's AI recommendation and AI-learning investment choices under an accuracy-parity policy. The setup matters because parity requirements are usually analysed as a pure constraint on the algorithm, ignoring that consumers respond to them by changing what they reveal — which in turn changes the data the algorithm learns from. Relevant to anyone modelling fairness regulation, personalisation or consumer privacy.
Stop reading p > 0.05 as "no effect"
Tomek, Caldwell and Eisner tackle one of the most persistent misinterpretations in the literature: treating a statistically non-significant result as evidence of no difference. A non-significant p-value is consistent with two very different worlds — no meaningful difference, or a real difference that the sample was too small or too noisy to detect. The paper sets out why and how to test for practical equivalence instead. Directly useful for anyone reporting null results, including the increasingly common "AI made no difference" findings.
Governing AI agents by profile, not by rule
Atoosa Kasirzadeh and Iason Gabriel argue in Nature for governing AI systems through agentic profiles — characterisations of what an AI agent can do, on whose behalf and with what autonomy — rather than through blanket rules aimed at models. For researchers deploying agents in data collection, coding or analysis, this framing is a useful vocabulary for describing exactly how much independent action a system was given.
That's this week. Forward it to the colleague who is about to cite an AI-productivity number — and explore AI tools for your own research at gaiforresearch.com.




Comments