The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study
Abstract
Large language models (LLMs) are increasingly used to simulate synthetic users in applications in behavioral science, economics, policymaking, and tech, such as conducting large-scale simulated user surveys or evaluating assistive agents. However, because LLMs are trained on observational data, interventions in experiments with synthetic users can induce unintended shifts in latent user attributes, causing the simulated population itself to differ across treatment conditions and confounding effect estimates. We formalize this phenomenon as selection bias due to user drift and show how these intervention-dependent shifts can inflate or attenuate observed treatment effect estimates. To diagnose user drift, we propose using negative control outcomes—attributes that should remain invariant under intervention—to identify distribution shifts across intervention conditions, providing evidence of confounding. To mitigate drift, we study adjusting users' persona specifications by eliciting additional confounders, finding that targeted, setting-relevant confounders can substantially reduce bias across survey-style and multi-turn agent evaluations.