AI in Research: The Agent Swarm Gets a Lab Bench

Image: "Andrew+ robot" by Pocar19 (CC BY-SA 4.0) via Wikimedia Commons, recoloured.
Welcome to this week's scan of how AI is reshaping research — curated from Nature, Science, PNAS and the leading journals.
From chatbot to lab crew: AI agents start running experiments
For two years the 'AI scientist' has mostly lived on a screen. This week it moved closer to the bench. As Nature reports, Anthropic has launched a biology wet lab where human scientists and AI agents design and run experiments together. Its first announced find came from roughly 950 AI agents that spent more than 21 hours exploring databases of billions of proteins in a self-directed way — discussing interim results and choosing next steps among themselves. Asked to look for proteins that might partner with reverse transcriptases, the agents noticed a short DNA sequence repeating near such a gene in a giant virus, then found similar repeat arrays in other viral genomes — patterns reminiscent of bacterial CRISPR systems. The caveats are real: the preprint is not yet peer reviewed, what the sequences do is unknown, and there is no known DNA-cutting enzyme partnered with them. The company's own head of life sciences calls it "a promising lead" — the hard part, characterising function, is still human, physical work.
The plumbing for that physical work is changing too. A second Nature story describes the Model Hardware Standard (MHS), a software framework from a collaboration between Anthropic and HHMI's Janelia Research Campus that lets instruments from different vendors 'talk' to each other — and lets an AI agent orchestrate them. A Carnegie Mellon team testing it reports setting up an experiment in hours, a process that can "easily take months". Meanwhile, Science flags a stranger side-effect of letting agents collaborate at scale: swarms of chatting agents tend to coin their own 'weird terms' — a natural but, the piece argues, disturbing by-product.
Why it matters: The unit of AI-assisted research is shifting from one person prompting one model to many agents searching, deciding and even operating equipment with limited human checkpoints. That raises throughput, but it also moves the researcher's job toward auditing: reading agent logs that may drift into private jargon, and deciding which of hundreds of machine-flagged 'weird things' deserve scarce lab time. For social scientists, the same pattern is coming to data collection and analysis pipelines — which makes provenance, transparency and human verification the skills to invest in now.
More from this week
Human–AI teams underperform — because experiments rarely let humans learn
A recent meta-analysis concluded that human–AI combinations, on average, do not beat the better of the two acting alone. Julian Berger, Ralph Hertwig, Ralf Kurvers and colleagues reanalysed all 74 studies in it and found most designs omitted features that support human learning, such as outcome feedback. Studies with feedback showed tentatively higher synergy; AI explanations paired with feedback were associated with positive synergy, while explanations without feedback were associated with negative synergy. The authors note causal tests are still needed — a clear design lesson for anyone running human–AI experiments.
AI cannot replace human participants, PNAS authors argue
In a new PNAS piece, Jim Hoover, Alan D. J. Cooke, Geoff Tomaino and Su Hyun Lee argue — as their title states plainly — that AI cannot replace human participants in behavioral research. It adds a high-profile voice to the synthetic-respondent debate we covered in August, and is worth reading before you pilot 'silicon samples' in a consumer or survey study.
p-values for machine-learning features
Kay Giesecke, Enguerrand Horel and Chartsiri Jirachotkulthorn introduce AICO (Add-In COvariates), which tests whether each feature contributes to a model's predictive performance by masking its information and measuring the change. It yields exact, finite-sample feature p-values and confidence intervals without retraining, surrogate models or distributional assumptions, and is demonstrated on credit scoring and mortgage-behaviour prediction. For business and finance researchers using ML, it offers a statistically principled answer to 'which variables matter?'.
OpenAI claims a Millennium Prize solution — physicists weigh in
OpenAI has claimed that its most advanced model solved a Millennium Prize Problem by showing that the Navier–Stokes equations for fluid flow can produce singularities. Nature's feature explores what that would mean for physics — experts note the equations were already known to fail in some regimes, such as rarefied gases — and the wider debate it has opened about AI's mathematical abilities, whether models absorb others' unpublished work, and who deserves credit.
Drug firms' private data supercharge AI protein models
Public databases lack enough examples of how proteins and drugs interact. Nature reports on an AI system trained on more than 20,000 protein structures contributed by pharmaceutical companies that outperforms AlphaFold-like models trained only on public data. It is a telling case for any field: when the decisive training data are proprietary, who gets to build — and scrutinise — the best models?
That's this week. Forward it to a colleague who's still copy-pasting into ChatGPT — and explore AI tools for your own research at gaiforresearch.com.





Comments