From Business Metrics to Behavioral Personas: Controllable User Simulation for Pre-Deployment Agent Testing
Zhenyu Zhang ⋅ Yuan Ling ⋅ Dingyang Chen ⋅ Xinyang Shen
Abstract
How do you systematically stress-test a conversational agent before deploying it to real customers? We present a framework that generates controllable, behaviorally diverse user personas grounded in real business data. All seller data used in this work is synthetic, generated to preserve the statistical properties of real operational profiles while protecting seller privacy; the method is equally applicable to real data for simulation grounding. Starting from operational metrics, we produce structured personas with 7 tunable personality dimensions and 10 adversarial archetypes targeting 5 failure modes of LLM-based agents. Personality parameters produce large behavioral effects (41\% of dimension--feature pairs show Cohen's $d{>}0.8$) but exhibit frequently non-linear responses---a finding with implications for simulator calibration. Used as a quality gate for a seller enrollment agent, the framework revealed that vanilla and chain-of-thought agents fail on 61\% and 50\% of adversarial personas, respectively, while state-tracking ("believed-state") agents achieve 75--77\% goal rates---a ${\sim}$37pp architectural advantage (vanilla vs.\ believed-state) constant across model tiers. Per-failure-mode diagnostics directly informed the selection of the believed-state variant.
Chat is not available.
Successful Page Load