top of page

One Dataset, Many AI Analysts — and Any Answer You Want

Welcome to this week's scan of how AI is reshaping research — curated from Nature, Science, PNAS and the leading business journals.

When defensible analyses are cheap, evidence gets abundant — and untrustworthy

For decades, the practical brake on p-hacking was labour. Many-analyst studies — independent teams testing the same hypothesis on the same data and regularly reaching conflicting conclusions — proved the multiverse of defensible analyses was real, but running one required costly human coordination. This week Martin Bertran, Riccardo Fogliato and Zhiwei Steven Wu removed the brake. Writing in PNAS, they built fully autonomous AI analysts on large language models, each independently executing a complete analysis pipeline on a fixed dataset and hypothesis, with a separate AI auditor screening every run for methodological validity. Across three datasets, the AI analysts reproduced the analytic dispersion observed in human many-analyst studies: substantial spread in effect sizes, *P*-values and conclusions, traceable to identifiable choices in preprocessing, model specification and inference.

The finding that should stop you is the next one. The outcomes are steerable. Reassigning the analyst persona or swapping the underlying LLM shifts the distribution of results — and it does so *even among runs the auditor judged methodologically sound*. This is not a story about AI making mistakes that better prompting will fix. It is a story about a large space of genuinely defensible analyses that now costs almost nothing to search. As the authors put it, when defensible analyses are cheap to generate, evidence becomes abundant and vulnerable to selective reporting. The economics of motivated reasoning have quietly inverted: the expensive thing used to be finding the analysis you wanted, and the cheap thing was honesty.

Why it matters: The same capability also supplies the fix, and the authors are specific about it. Treating analyst results as a *distribution* rather than a point estimate makes analytic uncertainty visible instead of hiding it in a single reported path. Deploying AI analysts against a published specification reveals how much disagreement stems from design choices the authors left underspecified — which is a genuinely new diagnostic for reviewers and replicators. Their conclusion is a transparency norm worth adopting before your field's editors impose one: if an AI generated your analysis, report the multiverse, not the run you liked. Pair this with last week's Nature paper showing LLMs can forecast experimental results before data collection, and the shape of the problem is clear — AI is getting good at telling you what you'll find, and even better at telling you what you'd like to find.

More from this week

Generative AI made Kenya's weakest entrepreneurs worse off

In a field experiment randomising access to a GPT-4-powered business assistant among Kenyan entrepreneurs, the authors could not reject the null of no average treatment effect on revenues or profits. The distribution told the real story: the effect for entrepreneurs who were low performing at baseline was over 0.20 standard deviations lower than for initial high performers, with subsample analyses showing low performers did nearly 10% worse while high performers may have benefited by over 15%. Strikingly, this was not driven by differences in the questions asked or the advice the AI gave, but by which advice entrepreneurs chose to implement. A caution for anyone reading average effects in AI adoption research.

AI eats the data it needs — and the case for taxing it

This Review of Financial Studies paper models a data-AI feedback loop in which data quality affects AI productivity, which influences adoption, which in turn shapes the composition and quality of future data. Calibrated to evidence on synthetic training loops, it predicts hump-shaped labour dynamics — short-term displacement that partially reverses as data quality deteriorates. A Grossman–Stiglitz-style externality emerges: AI adopters free-ride on the human-generated actions supplying the novel information AI itself relies on. The policy conclusion is counterintuitive — in a competitive market AI should be taxed to correct the externality, but a concentrated AI industry overcorrects, making a subsidy optimal.

AI-assisted creativity raises the floor and lowers the ceiling

Across four studies spanning short-story writing, circular-economy solutions, humour and collaborative storytelling, the authors found AI consistently limited the diversity of ideas across groups. The result is a paradox: higher average quality, but fewer breakthrough outliers. For research teams the implication mirrors the 'monoculture' worry in science — individual outputs improve while the collective search narrows, and the outliers are precisely what long-run innovation runs on. The authors' prescription is workflow design that harnesses AI without flattening originality.

Which 'AI scientist' actually suits your lab?

Ewen Callaway surveys the department's worth of general-purpose 'AI scientist' platforms now competing for lab budgets, including Anthropic's Claude Science — unveiled on 30 June with biology research in mind — alongside Co-Scientist and Biomni. Benchling's Ashu Singhal estimates fewer than 20% of labs have fully embedded AI scientists into their research. The framing is practical: hypothesis-generating systems fit the earliest stages of a project, while task-specific tools handle work like genomic data analysis later on. One anecdote does most of the persuading — Stanford's Euan Ashley led the first clinical analysis of a human genome in 2010 with a team of 31 scientists over nine months; he recently asked Claude to analyse his own genome to the same standard, and it took 30 minutes.

Algorithmic status inequality is not a data-cleaning problem

This paper introduces 'algorithmic status inequality' — enduring disparities in social position, influence and resource access reinforced by AI systems — as a lens on technological stratification in organisations. The integrative model shows how computational beliefs (cultural assumptions embedded in algorithmic design) interact with computational inequalities (disparities in technical capability) to produce persistent status hierarchies through self-reinforcing feedback loops, illustrated in recruitment, healthcare and legal assistance. The argument cuts against both counterpoints in the debate: neither purely technical approaches such as diversifying training data nor purely market-based approaches such as relying on competitive dynamics are sufficient.

What 400,000 investor queries reveal about how people really use AI

Blankespoor, Croom and Grant pair archival evidence from over 400,000 investor queries to a major brokerage's GenAI chatbot with a survey of more than 2,000 retail investors. Investors most often use GenAI to interpret and contextualise financial information and market movements, alongside stock screening and streamlining complex research, with usage shifting over time from screening and high-level company assessments toward detailed monitoring and interpretation of firm news. Nearly half of surveyed investors report using GenAI primarily to accelerate and simplify information processing, and more sophisticated retail investors lead adoption while deploying it for more complex tasks. A rare look at real usage logs rather than lab tasks.

That's this week. Forward it to a colleague who is about to let an agent run their analysis — and explore AI tools for your own research at gaiforresearch.com.

Comments


This initiative is supported by the following organizations:

  • Twitter
  • LinkedIn
  • YouTube
logo_edited.png
bottom of page