Bayesian Truth Serum for LLMs?
Shuo L Liu
Abstract
Bayesian Truth Serum (Prelec, 2004) asks each voter to predict how the other voters will answer. For strategic human voters, the rule rewards truthful reports at equilibrium. We test whether this signal still helps when the voters are language models. The original assumptions fit poorly. The models are not strategic agents, and they largely share pretraining data. On 871 code review items judged by four language models, every BTS variant we test is less accurate than plain majority vote. The gap is about one point. A paired McNemar test finds the gap statistically significant, and it remains after Benjamini and Hochberg correction across the full grid (BTS$_E$ at $p{=}0.031$, BTS$_C$ at $p{=}0.0039$, both BTS$_I$ and the Surprisingly Popular rule at $p{=}0.0078$). A 150 item probe over MMLU, TruthfulQA, and ARC Challenge shows the same pattern. Calibration tells a different story. Every BTS variant lowers expected calibration error by 21 to 32 percent relative to majority vote. The Surprisingly Popular rule of Prelec, McCoy and Seung (2017) has the lowest mean miscalibration under split conformal coverage ($0.0212$ versus $0.0390$ for the full BTS score). A sequential anytime valid test is uninformative on the hardest tier because the panel is saturated. We therefore withdraw the claim that the team beats the best individual reviewer pending a harder benchmark. A persuasion audit holds at 27 of 27 sessions (one sided 95 percent lower bound $0.895$). The failure mode is structural. Shared pretraining breaks the conditional independence assumption. Language models are not strategic agents. Their stated confidences are decoder samples, not coherent posteriors. The same shrinkage that hurts accuracy also dampens overconfident tails, which calibration metrics reward. Here BTS is best viewed as a calibration estimator inside a modern uncertainty layer, not as a replacement for plain voting.
Chat is not available.
Successful Page Load