57% of Papers Show AI's Fingerprints — But Can We Really Count Them?
- Lille My
- 6 hours ago
- 5 min read

Image: "British Museum Reading Room Panorama Feb 2006" by Diliff (CC BY 2.5) via Wikimedia Commons, recoloured.
Welcome to this week's scan of how AI is reshaping research — curated from Nature, Science, PNAS and the leading business and social science journals.
How much of the literature is AI-written? PNAS is now arguing about the ruler, not the reading
In May, Kyle Siler published what is probably the largest audit of AI's footprint in the published literature: an analysis of the full texts of 7.3 million journal articles from 2020–2025 across Elsevier, Frontiers, MDPI and PLoS. Building a corpus of 228 'focal words' whose frequency jumped sharply after 2022 in ways consistent with LLM output, the study estimates that by 2025 roughly 57% of published articles showed evidence of LLM influence — up from 12% in 2023. The more interesting result for social scientists is not the headline number but the stratification: difference-in-differences models find LLM-associated language varies sharply by region, institutional rank, publisher, discipline and journal tier, with economic development and distance from English as a first language among the strongest regional predictors. Lower-ranked institutions show higher rates than elite universities; young for-profit publishers show elevated rates against their competitors. Adoption, in Siler's framing, is pervasive but socially stratified.
This week PNAS published the pushback. In a letter titled "Lexical change is not a calibrated measure of LLM prevalence or its determinants", Chad M. Topaz and Utsav Bahl challenge exactly the inferential step the headline number depends on: that a rise in the frequency of certain words can be converted into a prevalence rate for LLM use, or used to identify who is adopting the tools. Siler's reply reframes the contribution as the interpretation of a post-2022 lexical shift in academic prose rather than a calibrated instrument reading. The disagreement matters because word-frequency methods have quietly become the field's default: Kobak and colleagues used excess-vocabulary analysis across more than 15 million biomedical abstracts to put a lower bound of 13.5% on LLM-processed 2024 abstracts, and a Nature Human Behaviour study applied a related estimation approach to scientific papers. The estimates differ by a factor of four, which is itself the story.
Why it matters: If you work with text as data — and increasingly, most social scientists and marketing researchers do — this is a live methodological warning about your own pipeline. Lexical proxies are cheap, scalable and intuitive, which is precisely why they get treated as measurements rather than as correlates in need of validation. The same critique applies well beyond AI detection: to sentiment dictionaries, to topic-model labels, to any construct inferred from word counts and then regressed on demographics or institutional rank. Before your next text-as-data paper, ask what a calibration study for your measure would look like — because reviewers are about to start asking.
More from this week
Most organizational coordination may be unnecessary — and AI agents helped prove it
Harang Ju borrows a precise criterion from distributed systems theory: coordination is genuinely required only when a task specification is nonmonotonic, meaning later information can invalidate earlier conclusions. He shows Thompson's classic taxonomy of interdependence maps onto that criterion, then applies the resulting decision rule to 65 APQC workflows and 13,417 O*NET tasks scored with a calibrated LLM, plus multiagent AI simulations. Under his decompositions, 74% of workflows and 42% of tasks are monotonic — implying that 24 to 57% of coordination spending is unnecessary for correctness. A rare case of multiagent AI serving as a measurement instrument for classic organization theory.
Machine learning from just 14 examples
Sparse training data is the standard reason 'we can't use ML here.' Shan and colleagues combine genome-informed data augmentation with contrastive learning in protein language space, starting from only 14 known phenazine-modifying sequences, in an approach they call ML-CITO. It led to the discovery of PTC, an enzyme catalysing a phenazine modification long presumed to occur only through nonenzymatic chemistry, confirmed by simulation and biochemical characterisation. The generalisable lesson is the recipe: use domain structure to manufacture informative training signal rather than waiting for a big dataset to appear.
Stable Diffusion as an instrument for cultural analysis
Rather than using generative AI to make images, Kim and co-authors use Stable Diffusion to read them, extracting two kinds of latent information from 500 years of Western painting: formal aspects such as colour, and contextual aspects such as subject matter. Contextual information aligns far more strongly with conventional artistic periods, styles and individual artists than formal elements do. The team then validates that geometry by infusing prospective contexts into historical artworks and synthesising paintings consistent with target-period style. For anyone doing computational analysis of cultural products, it is a template for treating a generative model's embedding space as measurement apparatus.
What the photos in annual reports are actually telling investors
Ben-Rephael, Ronen, Ronen and Zhou use machine-learning algorithms to assess the informativeness of photos (not graphs or charts) in firms' annual reports, developing a measure of 'content reinforcement' — how far the information investors extract from images complements the textual narrative. Firms use more images when asset growth, business complexity and less readable disclosures make information processing costly, and they increase image use after an exogenous decline in analyst coverage. Greater visual prevalence and stronger text reinforcement are associated with more accurate, less dispersed analyst forecasts. A clean example of multimodal ML opening a disclosure channel that text-only research has ignored.
Modularity doesn't replace coordination — it depends which kind
Amrit Tiwana draws on Simon's theory of the artifact to disentangle platform modularity into three facets: decoupling, interfaces, and 'design rules' that circumscribe what apps may do internally. Coordination complements the separability-fostering facets but substitutes for the integration-fostering ones, and design rules — ubiquitous in platforms yet theoretically overlooked — generate equivocality that coordination must resolve. The paper extends the argument explicitly to agentic AI and hybrid platform frontiers, where the question of how much coordination an ecosystem owner should buy is about to get much sharper.
Practical: AI graphical abstracts, and the image mistakes to avoid
Nature ran a pair of practical guides this week: one walking through how to produce a graphical abstract with AI tools in minutes, and a companion piece offering three tips for avoiding AI image mistakes in science. Worth reading together rather than separately — the speed gain in the first is only safe once you have internalised the failure modes in the second, given how quickly journals have moved to police AI-generated figures.
That's this week. Forward it to a colleague who's about to run a text-as-data study — and explore AI tools for your own research at gaiforresearch.com.
