top of page

Can LLMs Stand In for Human Subjects?

Image: "Empty seats" by ProjectManhattan (CC BY-SA 3.0) via Wikimedia Commons, recoloured.

Welcome to this week's scan of how AI is reshaping research — curated from Nature, Science, PNAS and the leading business journals.

LLMs as synthetic subjects: cheap pilots, or a shortcut too far?

The most consequential methods debate in the social sciences right now is whether a language model can stand in for a human participant. In Strategic Management Journal, Matteo Tranchero, Cecil-Francis Brenninkmeijer, Arul Murugan and Abhishek Nagaraj offer one of the most concrete answers yet for strategy research: a framework for designing and running simulated experiments with LLM-powered agents as synthetic subjects. Applying it to the classic exploration–exploitation dilemma, they report that LLM-based experiments reproduce patterns observed among human participants — which is exactly the result that makes the approach both attractive and hazardous.

The authors position the method as a tool for rapid, low-cost prototyping of human experiments and for generating novel hypotheses — not as a replacement for human data. That framing matters. A synthetic pilot that costs a few dollars and runs overnight can kill a bad design before you spend real money on a panel, or sweep a parameter space no IRB-approved study could afford. But reproducing a known behavioural pattern is a weak test: a model trained on decades of published social science should recover findings that appear in its training data. Matching the textbook result is table stakes, not validation — the interesting question is whether these agents get the *novel* cases right, and that is precisely where you have no human benchmark to check against.

Why it matters: If you run experiments — in marketing, management, IS or psychology — this is now a live design choice rather than a curiosity. The defensible use is upstream: pilot your stimuli, prune your conditions, generate hypotheses, then collect human data for anything you intend to claim. The indefensible use is downstream: reporting agent behaviour as evidence about people. Reviewers are already sharpening this distinction, and a paper that blurs it will get caught. Read the framework, borrow the prototyping workflow, and keep humans in the confirmatory studies.

More from this week

PNAS asks whether research stays a human enterprise

Michael Hochberg and Peter H. Thrall use a PNAS perspective to put the question bluntly: AI will reorganize science, and it is an open matter whether research remains a human enterprise. It is a useful companion to the empirical work on AI adoption, because it shifts the frame from individual productivity to the structure of the scientific enterprise itself — who sets the questions, who gets credit, and what the human role becomes.

Security cameras become an 'offline clickstream'

Physical retailers have never had the real-time behavioural visibility that clickstream data gives e-commerce. In Information Systems Research, Rubing Li, Wen Wang, Kaiquan Xu and Anindya Ghose build a video analytics framework that converts ordinary in-store security footage into a structured behavioural record — using person re-identification, trajectory reconstruction, pose estimation and vision-language models to predict in-store purchase. For consumer researchers it is both a rich new data source and a live privacy question worth thinking through before the data lands in your lab.

Medical AI has a measurement problem

Writing in Nature, Arjun K. Manrai argues that the field evaluating medical AI is measuring the wrong things. The critique generalises well beyond medicine: any discipline that adopts leaderboard-style benchmarks inherits the gap between scoring well on a curated test set and being useful in the messy setting the tool is meant to serve. Worth reading if you are designing evaluations for AI tools in your own field.

AI-generated adult content and offline crime

Siddharth Bhattacharya, Yun Young Hur and Gal Oestreicher-Singer report in MIT Sloan Management Review that an analysis of detailed crime data from Japan's National Police Agency found statistically significant increases in rape during periods of peak engagement with material tagged as AI-generated and adults-only. The authors frame this as a looming reputational and legal risk for firms providing such tools and venues. As with any observational design on sensitive outcomes the identification deserves close reading — but the question it raises about generative AI's downstream social effects is one business scholars cannot leave alone.

An AI 'Raygun' resizes proteins

Nature reports on Raygun, an AI system that can shrink or enlarge proteins — a capability its developers say opens the door to easier protein editing. It is a reminder that while the social sciences debate synthetic subjects, generative models in the life sciences are producing designs that get tested at a bench rather than argued about in a seminar.

Interpretable treatment policies, learned from observational data

Nathanael Jo, Sina Aghaei, Andrés Gómez and Phebe Vayanos present a mixed-integer-optimization method in Management Science for learning optimal prescriptive trees — interpretable treatment-assignment policies in binary-tree form — from observational rather than randomized data. For fields where an intervention has to be explained to a regulator, a clinician or a manager, an interpretable policy of moderate depth is often worth more than a marginally better black box.

That's this week. Forward it to a colleague who's about to run a pilot study — and explore AI tools for your own research at gaiforresearch.com.

Comments


This initiative is supported by the following organizations:

  • Twitter
  • LinkedIn
  • YouTube
logo_edited.png
bottom of page