Chorus: When Synthetic Audiences Should Decide or Abstain
Abstract
Agentic systems are currently integrated into many real-world applications. Persona-based workflows are also involved to those systems. Chorus is built to support decisions with synthetically generated personas. We propose a selective decision support system with synthetically generated personas which reduces wrong decisive recommendations by abstaining on uncertain cases. We tested Chorus using Gemini 2.5 Flash-Lite on 96 outcome-labeled cases, mainly public A/B cases with a small number of human-panel labels. With 64 personas, the forced directional prediction was correct for 75/96 (78.1%). An interval-style synthetic-vote diagnostic covered 67/96 (69.8%), was correct on 56/67 covered cases (83.6%), and reduced wrong decisive recommendations from 21/96 (21.9%) to 11/96 (11.5%), while abstaining on 29 cases (30.2%). These results do not provide calibrated population inference, causal lift estimates, or distribution-free risk control. Instead, they show that selective evaluation metrics expose weaknesses missed by winner accuracy alone. Chorus is best viewed as an empirical decision-support system for abstention, disagreement inspection, and validation routing, not as a replacement for A/B testing or human panels.