LLM social-simulation agents pass standard fidelity checks yet fail as evidence; we define mechanism non-identification, give a reproducible trigger and trace-level diagnostic, and propose a six-item disclosure checklist.
Abstract
This position paper argues that LLM social simulations cannot substitute for human-subject evidence without an identification strategy. Recent work shows high predictive fidelity --- GPT-4 simulations correlate strongly with human treatment effects across hundreds of contrasts, including a post-training-cutoff subset --- while critical work shows synthetic respondents fail regression, prompt-sensitivity, and temporal-stability tests. These findings are not contradictory: prediction asks whether outputs match observed outcomes, while substitution asks whether the simulation identifies the social data-generating process. We formalize the problem as observational equivalence among three mechanisms: training-prior retrieval, prompt-induced role compliance, and genuine interactional emergence. The same outcome distribution can be rationalized by multiple combinations of these mechanisms, so predictive fit alone cannot identify emergence. We audit the principal LLM social-simulation literature through this lens, concede the strongest predictive-fidelity result, and show why it cannot license replacing respondents or experiments. NeurIPS should require an identification standard: simulations may generate hypotheses or forecasts, but they become evidence only when their identifying assumptions are explicit, testable, and stress-tested. Positioned for FAGEN, the paper documents an under-recognized agentic failure mode: multi-agent LLM systems can pass standard final-score fidelity checks (Hewitt et al. r=0.85 across 476 effects) and remain reliable on terminal-success benchmarks while systematically failing as evidence about a target social mechanism, because no current trace-level diagnostic separates training-prior retrieval, prompt-induced role compliance, and interactional emergence. We give an operational definition (mechanism non-identification), a reproducible audit trigger applicable to four pre-existing agent papers, and a six-item disclosure checklist that surfaces the failure at the trace layer rather than the final-score layer.