AI Is Auditing the Scientific Record — and Finding 75-Year-Old Mistakes
- Lille My

- 1 hour ago
- 4 min read

Image via Wikimedia Commons (public domain).
Welcome to this week's scan of how AI is reshaping research — curated from Nature, Science and the leading business and social science journals.
AI is auditing the scientific record — and the record is losing
For three years the debate about AI and scientific integrity has run in one direction: what happens when machines pollute the literature with fabricated text and hallucinated citations. This week Nature reports on the traffic running the other way. A theoretical chemist at Zhejiang Lab in Hangzhou was using an AI system to predict boiling points when it kept producing values that clashed with long-accepted entries in a 75-year-old reference database. He went back to the original literature by hand — and found that the reference data, not the model, were wrong. In other cases the same approach surfaced a typo in a published paper and incorrect century-old boiling-point measurements. The technology, Nature notes, is proving unusually adept at finding faults in old papers and reference works. The caveat matters as much as the finding: these tools "make mistakes like humans do," so every flag still needs manual verification before it means anything.
That capability lands in the middle of an argument about where AI belongs in the quality-control pipeline. In *Management Science*, a team argues in "Fighting Fire with Fire" that AI-accelerated authorship has driven submission volumes past what volunteer reviewer pools can absorb, degrading turnaround times and decision accuracy — and that journals therefore have little choice but to bring large language models into the review workflow, starting with a concrete, bounded LLM-assisted pipeline. A commentary in the same issue pushes back on the framing rather than the tools: don't build a better fax machine. Bolting AI onto the existing architecture of submission, review and publication preserves a system designed for a scarcity of manuscripts and a scarcity of readers, when the interesting opportunity is to rethink how knowledge gets evaluated and disseminated at all. Meanwhile *Science* makes a parallel argument about how we measure AI itself, calling for successors to the Turing test built around diversity, partnership and systemicity rather than imitation of a single human.
Why it matters: Post-publication error detection is the first job where AI's weaknesses barely matter. A model that flags a suspect value at scale is useful even at modest precision, because a human verifies every hit and the cost of a false positive is one afternoon in the archives. That is a very different economics from AI-assisted peer review, where a wrong call quietly kills or passes a paper. Expect the first durable AI role in scholarly quality control to be *after* publication, not before it — and expect some uncomfortable weeks when a tool points at a reference table your own field has been citing for decades. If your work depends on canonical constants, benchmark datasets or hand-curated tables, this is the year to ask when anyone last checked them.
More from this week
AI agents are spotting decades-old errors in the literature
A chemist at Zhejiang Lab found that his AI system's 'wrong' boiling-point predictions were actually right, and the 75-year-old reference data they contradicted were in error. Related cases turned up a typo in a published paper and incorrect century-old measurements. Nature stresses that AI fact-checkers remain unreliable arbiters on their own — they make human-like mistakes — so flagged items need manual verification before anything is corrected.
The case for infusing AI into peer review
Writing in Management Science, the authors argue that AI adoption by authors has accelerated article production and submission rates, straining reviewer capacity and hurting journal metrics like turnaround time and decision accuracy. Their response is a concrete workflow that deploys large language models inside the review process as a first step rather than a wholesale replacement of human judgement. It is one of the first serious operational proposals from a top journal, not just a position statement.
Don't build a better fax machine
This companion commentary accepts that AI can ease reviewer shortages and raise productivity, but warns that such applications largely preserve the existing architecture of scientific publishing. Drawing on operations management and innovation research, it argues AI instead creates an opening to reimagine more fundamental aspects of how scholarly knowledge is evaluated and disseminated. The disagreement is not about whether AI works, but about what problem the field should be solving.
The next Turing tests
Imitation of an individual human was always a narrow target, and it has become a misleading one. This Science piece argues that new approaches to evaluating AI progress should be organised around diversity of capability, human–AI partnership, and systemic effects rather than a single pass/fail conversation. For researchers who evaluate AI tools as part of their own methods sections, it is a useful reframing of what 'good performance' should even mean.
Generative AI as a causal prediction tool for untested content
Generative AI makes new content cheap to produce, which paradoxically makes selection harder: managers face an exploded option set with little chance to test any of it. This Journal of Marketing Research paper develops a causal prediction framework for evaluating novel unstructured treatments — text, images, creative variants — that have never been fielded. The underlying problem generalises well beyond marketing to any researcher who wants effect estimates for stimuli they have not yet run.
LLMs may make us 'do more, less well'
Reported in Nature, a modelling study predicts that LLM adoption speeds up every phase of the research process, but that the acceleration shows up as more papers rather than more polished ones: researchers jump to the next project instead of refining the current one. Co-author Carl Bergstrom is explicit that the tools are not the root cause — the result reflects an incentive system that already rewards quantity over quality, and 'LLMs hold up a mirror to problems that we already have.'
That's this week. Forward it to a colleague who still trusts every number in the handbook — and explore more at gaiforresearch.com.




Comments