Multiple-Comparison Testing for Cross-Validation Path Optimization Benchmarks: A Hessian-Mismatch Diagnostic for Holm--Bonferroni Failure Modes
Jinze Yu ⋅ Yanping Deng
Abstract
Comparing optimization solvers on K-fold cross-validation accuracy across D datasets is a D-comparison problem with correlated test statistics, yet the standard practice reports uncorrected p-values or no p-values at all. We instantiate a Holm--Bonferroni paired t-test protocol on a benchmark of 11 datasets comparing a fast closed-form / fixed-preconditioner logistic-regression solver ($R1{+}10$) to warm-start L-BFGS. Two of 11 datasets reject the iso-accuracy null hypothesis (Adult, $p<10^{-3}$; Covertype, $p=0.001$). We then show, as a post-hoc diagnostic, that the rejections are precisely the datasets with small Hessian-mismatch index $\mu_{*} := \lambda_{\min}(\mathbf{M}^{-1/2}\mathbf{H}(\boldsymbol{\theta}^{*})\mathbf{M}^{-1/2})$, where $\mathbf{M} = 2\lambda\mathbf{I}+\tfrac{1}{4}\mathbf{X}^{\top}\mathbf{X}$ is the fixed preconditioner and $\mathbf{H}(\boldsymbol{\theta}^{*})$ is the optimum-Hessian. We further test whether the progressive estimate $\hat{\mu}_{k}$ from early iterates is a valid pre-iteration decision rule and reject this hypothesis: across 10 datasets $\hat{\mu}_{5}$ values cluster in $[0.08, 0.21]$ regardless of $\mu_{*}$ (false-positive rate 100\% at the $\hat{\mu}_{5} \leq 0.05$ threshold). The work illustrates how multiple-comparison correction interacts with a structural diagnostic in optimization benchmarking.
Chat is not available.
Successful Page Load