top of page

AI in Research: When Papers Become Agents

13 hours ago
4 min read

Image via Wikimedia Commons (public domain).

Welcome to this week's scan of how AI is reshaping research — curated from Nature, Science and the leading business journals.

From static papers to working agents

Anyone who has tried to reuse a published method knows the pain: find the code, fix the environment, decode the supplementary materials. In Nature, Jiacheng Miao, Joe R. Davis, Jonathan K. Pritchard and colleagues introduce Paper2Agent, an automated framework that converts a research paper into an AI agent acting as a "virtual corresponding author." Multiple agents analyse the paper and its codebase to build a model context protocol (MCP) server, then generate and run tests to make it robust. Connected to a chat agent such as Claude Code, the resulting 'paper agent' can answer scientific queries in natural language while invoking the paper's own tools and workflows. In case studies built on AlphaGenome, Scanpy and TISSUE, the authors report that the agents reproduced the original papers' results and handled novel user queries — and several paper agents collaborated to prioritise a causal gene for psoriasis (Nature news coverage).

The same week, Science published the Virtual Biotech — an organisation of AI agents modelled on a drug-development company, with divisions for target discovery, safety assessment, modality selection and clinical development. In one demonstration, over 37,000 agents annotated outcomes from 55,984 trials, finding that drugs targeting cell-type-specific genes were 48% more likely to reach market with 32% fewer adverse events. The system also proposed a lung-cancer therapeutic strategy and inferred possible failure mechanisms of a terminated ulcerative colitis trial. The authors frame it explicitly as a *human-guided* multi-agent system.

Why it matters: Together these papers point to a shift in what a research output is — from a document you read to a system you query, and from a single analyst to an organisation of agents coding evidence at scale. For social scientists and business researchers, the Paper2Agent logic transfers directly: a replication package that can explain and re-run itself would lower the barrier to reuse and make reproducibility checks far cheaper. The Virtual Biotech's trial-annotation exercise is, in effect, large-scale content analysis by agents — a template for systematic reviews and archival coding. The open questions are the familiar ones: who validates the agents' coding, and how do journals credit and review a paper that talks back?

More from this week

A generalist AI for abdominal CT boosts radiologists

RADAR is a vision-language model trained on more than 400,000 contrast-enhanced abdominal CT examinations and 15 million anatomy-wise image-text pairs, learning directly from clinical reports without manual annotation. Across internal and external multi-centre evaluations it performed strongly on 18 anatomical structures and 146 imaging findings. In a reader study, RADAR assistance increased the diagnostic sensitivity of 26 radiologists by roughly 10% — a useful data point for anyone studying human-AI collaboration in expert work.

Deep learning meets structural econometrics

Yicheng Song and Tianshu Sun propose the Sequential Search Transformer (SST), a deep structural econometric model that embeds sequential search theory into an end-to-end trainable deep learning architecture, with a formal identification analysis. Applied to clickstream data from a U.S. e-commerce site, SST outperformed state-of-the-art deep learning and structural models in predicting searches and purchases, and policy experiments showed gains for product recommendations and new-product promotion. A clear model for marketing researchers who want predictive power without giving up interpretability.

Why generative AI firms might skip the free tier

An analytical model in Information Systems Research examines how 'AI adaptive learning' — user interactions refining a model during fine-tuning — changes freemium strategy. The authors find that highly effective adaptive learning can paradoxically hurt profitability by making the free version too appealing, that firms may optimally offer no free version even when service costs are minimal, and that a free tier does not necessarily raise consumer surplus, because better AI can let firms raise premium prices.

Facial recognition cut airport departure delays

Exploiting the first terminal-wide implementation of facial recognition at a U.S. airport, this study finds roughly a 16% reduction in departure delays and a 6% reduction in arrival delays, with no increase in early departures or arrivals. Gains were smaller for flights to Asia and Africa — destinations with a higher share of non-Caucasian passengers — and larger for bigger aircraft, a reminder that AI's operational benefits can be distributed unequally.

Making algorithmic advice people will actually follow

Chan, Delage and Lin tackle a classic algorithm-aversion problem: 'optimal' recommendations get ignored when they clash with how people judge a situation. Their method learns decision-makers' evaluations from past choices, builds an uncertainty set around the inferred parameters and uses robust optimisation, with statistical and performance guarantees. In a simulated Toronto food-delivery study, it substantially improved estimated courier adherence while keeping delivery times comparable.

That's this week. Forward it to a colleague whose replication package still lives in a zip file — and explore AI tools for your own research at gaiforresearch.com.

Comments


This initiative is supported by the following organizations:

  • Twitter
  • LinkedIn
  • YouTube
logo_edited.png
bottom of page