RoPoLL: Robust Panel of LLM Judges
Anish Acharya ⋅ Ka-woon Pan ⋅ Brian Verkhovsky
Abstract
The LLM Jury: a panel of LLM-as-Judge models that collectively scores language-model outputs, has quickly become a practical alternative to single-judge evaluation, yet in-depth studies of its statistical behavior remain scarce. In this work we initiate a theoretical investigation of jury aggregation. We formalize the LLM Jury and characterize judge failures e.g., mode collapse, sycophancy, safety refusals, as Byzantine faults under the Huber contamination model, showing that standard polling~\citep{verga2024replacing} admits unbounded bias from even a single corrupted judge. We propose \textsc{RoPoLL} (\textbf{Ro}bust \textbf{P}anel \textbf{o}f \textbf{LL}M-As-Judge), a robust LLM Jury via robust mean estimation. We establish finite-sample error bounds for \textsc{RoPoLL} under contamination, showing it tolerates up to half the jury being adversarial while converging at rate $O(\sigma\sqrt{d/N})$. We validate the theory with large-scale experiments: 13 judges spanning 4B--675B parameters, three benchmarks with human and GPT-4 ground truth, and systematic adversarial injection at contamination rates up to 50\%. On biased heavy-tailed and cross-dimensional attacks, \textsc{RoPoLL} reduces RMSE by one to three orders of magnitude over the arithmetic mean; a three-judge \textsc{RoPoLL} jury at 38\,B total parameters outperforms the best single 675\,B judge under 30\% corruption, demonstrating that robust aggregation of cheap models is strictly more parameter-efficient than scaling a single judge.
Chat is not available.
Successful Page Load