$\text{P}^4$Bench: Contamination-Proof, Publicly Verifiable, and Privacy-Preserving LLM Evaluation via Zero-Knowledge Proofs
Enhan Zhao ⋅ Yuanrui Zhang ⋅ Zhang Zhang ⋅ Wei Wu ⋅ Di He
Abstract
Mathematical benchmarks are central to evaluating the reasoning abilities of large language models (LLMs), but data contamination undermines their reliability when benchmark questions and answers appear in training data. The most common way to handle contamination is to keep benchmarks private, hiding both test problems and answers from model providers and the public. Although this reduces direct test-set leakage, it sacrifices transparency: the community must trust a third-party benchmark organizer to validate the problems, run the evaluation, verify the answers, and report scores, rather than being able to verify the results independently. To address this trade-off, we propose $\text{P}^4$Bench, a contamination-proof, publicly verifiable, and privacy-preserving evaluation framework based on zero-knowledge proofs (ZKPs). The key idea is to publish benchmark problems while keeping answers private: model providers submit cryptographic proofs that their answers are correct, and anyone can verify these proofs without learning the answers. This simultaneously enables public benchmark release, answer privacy, and independent verification. We instantiate $\text{P}^4$Bench on Boolean satisfiability (SAT) and discrete logarithm problem (DLP) tasks, encoding answer verification as zk-SNARK circuits. Experiments with frontier LLMs show that the framework can effectively distinguish model capabilities while keeping proof generation and verification practical.
Chat is not available.
Successful Page Load