From Debate to Decision: Distribution-Free Act-or-Defer Control for Multi-Agent LLM Pipelines
Mengdie Flora Wang ⋅ Haochen Xie ⋅ Guanghui Wang ⋅ Aijing Gao ⋅ Guang Yang ⋅ Ziyuan Li ⋅ Qucy W Qiu ⋅ Fangwei Han ⋅ Hengzhi Qiu ⋅ Yajing Huang ⋅ Bing Zhu ⋅ Jae Oh Woo
Abstract
Agentic LLM pipelines repeatedly face act-or-defer decisions---whether to answer, route, continue, or escalate---yet deployed heuristics (top-1 confidence, majority vote, consensus stopping) carry no finite-sample risk certificate. We study how to attach a distribution-free validity layer to a black-box agentic pipeline, using multi-agent debate as a case study where naive consensus stopping can confidently commit wrong answers. Our premise is that the right object to calibrate in an agentic system is the composed pipeline output, not the individual agents. Conformal Social Choice aggregates verbalized probabilities from $N$ heterogeneous agents via a linear opinion pool and applies split conformal prediction post-hoc on the aggregated belief, producing sets $C(x)$ with $\Pr[y \in C(x)] \geq 1-\alpha$ under exchangeability alone---no per-agent calibration. A hierarchical action policy maps singletons to autonomous action, multi-element sets to escalation, and empty sets to anomaly review. The contribution is a statistical systems framework, not a new conformal theorem: it identifies what to calibrate (the composed output), where to intervene (after aggregation, before action), and how to convert set size into deployment behavior. Across a full domain-by-round stress test on eight MMLU-Pro domains with three heterogeneous agents, the layer maintains near-nominal coverage while set size adapts to task ambiguity and singleton decisions increase as debate progresses; at $\alpha{=}0.05$ it intercepts 81.9\% of wrong-consensus cases while introducing only 2 new wrong singletons out of 4{,}158 (240:1 ratio), and preliminary robustness checks on GPQA and LogiQA show similar behavior without retuning.
Chat is not available.
Successful Page Load