Flow Matching Calibration for Simulation-Based Inference under Model Misspecification
Abstract
Simulation-based inference (SBI) is transforming experimental sciences by enabling parameter estimation in complex non-linear models from simulated data. A persistent challenge, however, is model misspecification. In a Bayesian setting, targeting posterior distributions, errors may arise from the simulator, the noise or prior modelling. These model components are only approximations of reality, and severe mismatches can yield biased or overconfident posteriors. We address this issue by introducing Flow Matching Corrected Posterior Estimation (FMCPE), a framework that leverages the flow matching paradigm to refine simulation-trained posterior estimators using a small set of calibration samples. Our approach proceeds in two stages: first, a posterior approximator is trained on abundant simulated data; second, flow matching transports its predictions toward the true posterior supported by calibration observations. We rely on the later to guide the correction, without requiring explicit knowledge of the misspecification form or of which model components are affected. This design enables FMCPE to combine the scalability of SBI with robustness to distributional shift. Across synthetic benchmarks and real-world datasets, we show that our proposal consistently mitigates the effects of misspecification, delivering improved inference accuracy and uncertainty quantification compared to standard SBI baselines, while remaining computationally efficient.
Lay Summary
Scientists often want to determine a hidden parameter of a physical system from what they can observe — for example, inferring a property of a distant star from its light. This already challenging problem can be even harder when the model linking the parameter to the observation is only an approximation of the real process, which is almost always the case. We developed a method that combines two sources of information: large amounts of simulated data, which are cheap and easy to generate from a possibly imperfect model, and a small number of real-world observations, obtained through controlled experiments or costly measurements. Using recent advances in generative AI, we built a two-stage approach: first we train a model on the abundant simulations, then we use the real observations to correct its predictions toward the truth. Crucially, this correction works without requiring any knowledge about how or where the model is wrong. We believe this can be of interest in many experimental sciences where trustworthy data is hard to obtain but generating approximate data is easy