Causal Effects with Unobserved Unit Types in Interacting Human–AI Systems
Abstract
Randomized controlled trials (RCTs) are the gold standard for establishing causal effects and informing evidence-based policy. As AI agents become pervasive on digital platforms, policymakers face a new challenge: measuring the causal impact of interventions on humans when the population is a mixture of humans and AI agents whose identities are unknown. Without such tools, governance decisions risk relying on misleading population-level estimates that conflate human and AI responses. We introduce a causal inference framework for experimentation in mixed human--AI populations. Each unit is associated with a prior probability of being human, and outcomes evolve through dynamic interactions on an unobserved network. We derive state evolution equations showing that aggregate outcomes depend on unit types only through their average composition, enabling consistent estimation of human-specific treatment effects from randomized data. We validate on an LLM-driven social platform simulation and demonstrate that standard estimators miss the true human effect entirely, while our method recovers it accurately, providing a principled methodology for evidence-based regulatory decisions in human--AI ecosystems.