Behavioral World Models as Missing Infrastructure for Responsible Generative Audio
Abstract
Generative audio systems are evaluated almost exclusively on signal-level quality metrics: naturalness, intelligibility, and speaker similarity. These metrics measure whether audio sounds good; they do not measure what it does to people who interact with it over time. A voice companion that sounds natural but systematically reduces engagement, inflates emotional arousal, or accelerates churn is a deployment failure no signal-level metric detects. We argue that behavioral world models, which track and predict user psychological state across multi-turn, multi-session interactions, are the missing evaluation and deployment infrastructure for responsible generative audio. We (i) formalize what such a model must provide, (ii) sketch what a concrete instantiation entails and outline preliminary evidence that one is achievable with current methods, and (iii) propose three behavioral outcome metrics (Engagement Trajectory AUC, Emotional Arousal Calibration Error, and Counterfactual Retention Gain) that can accompany acoustic quality metrics as standard practice.