Too Sharp, Too Sure: When Calibration Follows Curvature
Abstract
Modern neural networks can achieve high accuracy while remaining poorly calibrated, producing confidence estimates that do not match empirical correctness. Yet calibration is often treated as a post-hoc attribute. We take a different perspective: we study calibration as a \emph{training-time} phenomenon on small vision tasks, and ask how it co-evolves with loss-landscape geometry. We identify a tight coupling between calibration, curvature, and margins during training of deep networks under multiple gradient-based methods. Empirically, Expected Calibration Error (ECE) closely tracks curvature-based sharpness throughout optimization. Mathematically, we show that both ECE and Gauss--Newton curvature are controlled, up to problem-specific constants, by the same margin-dependent exponential tail functional along the trajectory. Causal tests via interventions that target sharpness versus directional curvature confirm the mechanism, with directional interventions yielding more reliable in-sample calibration gains.