Bayesian model selection and misspecification testing in imaging inverse problems only from noisy and partial measurements
Abstract
Modern imaging techniques heavily rely on Bayesian statistical models to address difficult image reconstruction and restoration tasks. This paper addresses the objective evaluation of such models in settings where ground truth is unavailable, with a focus on model selection and misspecification diagnosis. Existing unsupervised model evaluation methods are often unsuitable for computational imaging due to their high computational cost and incompatibility with modern image priors defined implicitly via machine learning models. We herein propose a general methodology for unsupervised model selection and misspecification detection in Bayesian imaging sciences, based on a novel combination of Bayesian cross-validation and data fission, a randomized measurement splitting technique. The approach is compatible with any Bayesian imaging sampler, including diffusion and plug-and-play samplers. We demonstrate the methodology through experiments involving various scoring rules and types of model misspecification, where we achieve excellent selection and detection accuracy with a low computational cost.
Lay Summary
Scientists in medical imaging, astronomy, and computational photography routinely reconstruct hidden images from incomplete or noisy measurements — for example, turning raw MRI signals into a brain scan. These reconstructions depend on mathematical models that encode assumptions about what plausible images look like, and a poorly chosen model can subtly distort the result in ways that mislead downstream decisions. Traditionally, verifying the model requires a ground-truth image for comparison, which simply isn't available in real-world settings. We propose a way to evaluate these models from a single noisy measurement, with no ground truth required. Our approach borrows the idea of cross-validation from statistics: we artificially split the noise inside the measurement into two complementary parts, producing two synthetic "half-views" of the unknown image. We then ask whether the model, after being shown one half, can accurately predict the other. A well-suited model predicts faithfully; a misspecified one fails in a measurable way. The method works with any modern image-reconstruction technique, including those built from deep learning. We demonstrate it on brain MRI reconstruction and image deblurring, where it reliably flags misspecified models — a practical safety check for scientific and medical settings where ground truth is unavailable.