Conformal Risk Control for AI-Assisted Mathematical Derivations: A Workflow with Finite-Sample Correctness Guarantees
Siddharth Karuturi ⋅ Kaustubh Bukkapatnam ⋅ Laksh Patel
Abstract
Large language models (LLMs) are increasingly used by researchers to assist with mathematical derivations, yet current practice offers no formal correctness guarantees. Researchers either verify every LLM-generated step manually—a significant time burden—or accept outputs heuristically, risking error propagation. We present \textsc{CRC-Derive}, a calibrated acceptance workflow grounded in conformal risk control (CRC) that provides finite-sample, distribution-free guarantees on the false acceptance rate (FAR) of LLM-generated proof steps. Given a small domain-specific calibration set of annotated steps (a calibration set of $n \ge 185$ suffices), the workflow selects a data-driven threshold $\hat{\tau}(\alpha)$ such that the expected fraction of incorrect steps passed to the researcher is at most $\alpha$. We prove six theorems covering: step-level FAR control, end-to-end proof composition via Bonferroni correction, a decision-theoretic optimal risk level, online adaptation for non-exchangeable sequential proofs, calibration sample complexity, and an efficiency ordering by score calibration quality. A human study with 10 ML researchers across 100 derivation tasks shows a $41\%$ reduction in time-to-correct-proof ($p = 4.8 \times 10^{-29}$, paired $t$-test) and a $77\%$ reduction in propagated errors ($p = 5 \times 10^{-4}$) relative to unassisted manual review. All code, prompts, and calibration data are released publicly.
Chat is not available.
Successful Page Load