Holm-Bonferroni Verification of Math Solvers: A Hessian-Mismatch Diagnostic for Benchmark-Level Failure Modes
Jinze Yu ⋅ Yanping Deng
Abstract
Verifying math solvers and AI-driven optimization systems across a benchmark suite is a multiple-comparison problem with correlated test statistics, yet benchmark practice routinely reports uncorrected $p$-values or none at all. We propose a verification protocol pairing Holm-Bonferroni multiple-comparison correction with a problem-specific structural diagnostic that attributes per-dataset rejections to algorithmic mechanism. We instantiate the protocol on a foundational math solver --- $\ell_{2}$-regularized logistic regression --- where a fast closed-form solver ($R1{+}10$, single Cholesky) is compared against warm-start L-BFGS across 11 datasets. The Holm-Bonferroni-corrected paired $t$-test rejects the iso-accuracy null on exactly 2 of 11 datasets (Adult, $p<10^{-3}$; Covertype, $p=0.001$). The rejected datasets are precisely those with smallest Hessian-mismatch index $\mu_{*} := \lambda_{\min}(\mathbf{M}^{-1/2}\mathbf{H}(\boldsymbol{\theta}^{*})\mathbf{M}^{-1/2})$, a generalized Rayleigh quotient that controls the solver's asymptotic linear convergence rate. The pre-iteration estimate $\hat{\mu}_{5}$ fails as a decision rule (FP rate $100\%$ at $\tau=0.05$): $\hat{\mu}_{5}$ values cluster in $[0.08, 0.21]$ regardless of true $\mu_{*}$. The protocol generalizes to verification of any math AI system whose error has a spectral characterization: replace $\mu_{*}$ with the problem-specific convergence diagnostic. We compare four correction procedures (Bonferroni, Holm-Bonferroni, Benjamini-Hochberg, no correction) and four test statistics (paired $t$, Wilcoxon, sign-test, bootstrap CI), all consistent on the two rejections.
Chat is not available.
Successful Page Load