Almost Every Biomedical Paper Now Shows Signs of AI — And the Skill Has Shifted
- Lille My

- 12 minutes ago
- 4 min read

Image via Wikimedia Commons (public domain).
Welcome to this week's scan of how AI is reshaping research — curated from Nature, Science and the leading management journals.
AI is in almost every paper. The open question is who's directing whom.
The headline number is hard to look away from. A preprint posted to arXiv on 12 August, covered by Nature, estimates that by December 2025 close to 90% of papers in the PubMed Central repository showed signs of LLM-assisted writing or editing — 77% across the whole of 2025, and 52% for 2024. The method extends the 'excess vocabulary' approach the same group published in Science Advances: rather than trying to detect AI directly, it tracks a set of non-content words whose frequency jumped abruptly after ChatGPT's release. Usage is uneven across a paper — highest in Discussion and Abstract sections, lowest in Methods — which is roughly where you would expect a writing aid, rather than a research aid, to leave its traces. The authors are explicit about the limits: English-language papers only, and PubMed Central is not the whole of science.
Treat the precise figure with care and the direction with none. Word-frequency detection cannot distinguish a researcher who asked a model to polish a non-native-English draft from one who generated a discussion section wholesale, and earlier estimates using narrower methods put 2024 usage at 12–16% — far below this study's 31% for abstracts and 52% for full texts. What the trend does establish is that AI-assisted writing has stopped being an edge case and become the default condition of the literature you read and review. If you are still framing disclosure policy around whether authors used AI, the interesting question has already moved on to how, and where in the manuscript.
Two management papers published this month suggest what that 'how' should look like. In California Management Review, Teck-Hua Ho and Catherine Yeung propose a principal worker model in which each knowledge worker acts as both task architect — decomposing work and delegating subtasks to AI agents — and agent orchestrator, optimising and synthesising agent outputs into a coherent deliverable. Writing in MIT Sloan Management Review, Jennifer Sloan and Vern Glaser make a parallel argument: stop prompting AI and start directing it, by defining an agent's context (what it can access), capabilities (what it can do), and orientation (what it attends to) — and by running multiple agents with different orientations over the same data to compare approaches.
Why it matters: The 90% figure and the orchestration papers are describing the same shift from opposite ends. AI is already woven through scholarly output, largely as an invisible writing assistant — the least accountable and least interesting use of it. The research-design move is to make the delegation explicit: decide which subtasks an agent owns, give it a defined orientation, keep the interpretive authority yourself, and write down what you did. That is both better methodology and, increasingly, what disclosure statements will need to say.
More from this week
Explainable AI has a cognitive cost
Explainable AI is usually justified as a fix for over- and under-reliance on algorithmic advice, but the empirical evidence is mixed. Boyaci, Canyakmaz and de Véricourt develop an analytical model of human–AI collaboration that captures the distinct features of human and machine intelligence and, crucially, accounts for the decision-maker's cognitive effort and fatigue. The framework offers a theoretical explanation for why adding explanations does not reliably improve joint performance — useful for anyone designing AI-assisted decision studies.
Most AI-patent datasets are measuring the wrong thing
Innovation research increasingly leans on patent data to measure AI activity, but the methods used to decide which patents count as 'AI' have rarely been validated. Wu, Min, Ding and Wang assess patent-class approaches and the USPTO's Artificial Intelligence Patent Dataset against an independent, human-expert-annotated ground truth, document substantial performance limitations, and introduce a scalable framework that classifies better. If your identification strategy rests on an off-the-shelf AI-patent flag, this is worth reading before your next revision.
The AI assistant market isn't winner-take-all
Cohen, Hage-Youssef, McCarthy and Sokol use consumer mobile data for the AI assistant category to challenge the winner-take-all framing of AI competition. They identify three coexisting and growing strategies — scale (ChatGPT), adjacency (Gemini) and premium specialisation (Claude) — and find that major product launches reward the launcher without measurably harming rivals. They argue these archetypes are not AI-specific but recur whenever a disruptive technology opens a market before competitive positions harden.
Predicting which brand slogans stick
Ada Aka and John McCoy build predictive models of how well a brand slogan is remembered and how strongly it cues its brand. They identify candidate memorability factors through a systematic literature review, then operationalise them using natural language processing, human ratings and constructs derived from text embeddings, and evaluate predictive accuracy across models. It's a clean template for how to turn a scattered qualitative literature into machine-measurable constructs.
Nature reviews the safety and security of clinical LLMs
Clusmann, Freyer, Ostermann, Ferber and colleagues survey the safety and security landscape for large language models deployed in healthcare settings. With clinical LLM pilots multiplying faster than the evidence base behind them, a consolidated account of failure modes and safeguards in a general-science venue is a useful reference point — including for social scientists studying AI adoption in high-stakes professional settings.
A reinforcement-learning gym for fluid dynamics
Lagemann, Mokbel, Gondrum, Rüttgers and co-authors introduce HydroGym, a reinforcement learning platform for fluid dynamics. Fluids resist conventional control because their dynamics are high-dimensional, nonlinear and multiscale; RL has made striking progress in domains such as protein folding and games, which had shared benchmarks and standardised environments that flow control lacked. HydroGym supplies them — a reminder that in many fields the bottleneck to AI progress is infrastructure, not model capability.
That's this week. Forward it to a colleague who's still copy-pasting into ChatGPT — and explore AI tools for your own research at gaiforresearch.com.




Comments