Randomized trial assesses LLM personas' predictive abilities and social influence in human simulation, indicating their strengths and limitations.
Large language models are increasingly proposed as simulators of individual humans, yet the proposals are rarely tested against strong statistical baselines under matched information. This work runs three frozen, pre-registered challenges. On individual survey-response prediction, a compact amortized person-embedding model (Mirror) reaches .623 accuracy, above the best statistical baseline (.589) and above interview-grounded LLM personas (.553 to .591). On group dynamics, LLM agents reproduce social-influence revision better than a fitted DeGroot model, but cannot reliably create partisan bias from an assigned identity at small scale, and both LLM families over-conform. The value of LLMs in human simulation lies in question understanding and influence propagation, not person-level prediction.
No takes yet. Share an insight, caveat, or question.
Umar Aslam (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: