Frequentist Consistency of Prior-Data Fitted Networks for Causal Inference
Abstract
Foundation models based on prior-data fitted networks (PFNs) have shown strong empirical performance in causal inference by framing the task as an in-context learning problem. However, it is unclear whether PFN-based causal estimators provide uncertainty quantification that is consistent with classical frequentist estimators. In this work, we address this gap by analyzing the frequentist consistency of PFN-based estimators for the average treatment effect (ATE). (1) We show that existing PFNs, when interpreted as Bayesian ATE estimators, can exhibit prior-induced confounding bias: the prior is not asymptotically overwritten by data, which, in turn, prevents frequentist consistency. (2) As a remedy, we suggest employing a calibration procedure based on a one-step posterior correction (OSPC). We show that the OSPC helps to restore frequentist consistency and can yield a semi-parametric Bernstein-von Mises theorem for calibrated PFNs (i.e., both the calibrated PFN-based estimators and the classical semi-parametric efficient estimators converge in distribution with growing data size). (3) Finally, we implement OSPC through tailoring martingale posteriors on top of the PFNs. In this way, we are able to recover functional nuisance posteriors from PFNs, required by the OSPC. In multiple (semi-)synthetic experiments, PFNs calibrated with our martingale posterior OSPC produce ATE uncertainty that (i) asymptotically matches frequentist uncertainty and (ii) is well calibrated in finite samples in comparison to other Bayesian ATE estimators.
Lay Summary
Many decisions in medicine, public policy, and business depend on estimating what would happen if we changed an action, such as giving a treatment or introducing a policy. New AI systems called prior-data fitted networks can make these estimates quickly because they learn from many simulated datasets before seeing a real one. However, we found that these systems can be overconfident in causal settings: the simulated worlds used for training often contain too little confounding, meaning too few cases where treatment choices and outcomes are linked through shared background factors. As a result, the AI may underestimate how uncertain it should be about the true effect of an intervention. We propose a calibration method that adjusts the AI’s estimate using a correction inspired by classical causal inference. This correction makes the AI’s uncertainty behave more like uncertainty from well-established statistical estimators, while still keeping the speed and flexibility of prior-data fitted networks. To make the correction practical, we also show how to recover the needed uncertainty about hidden modeling components from the network’s predictions. Across several synthetic and semi-synthetic benchmarks, the corrected method provides more reliable uncertainty than naïvely using the networks as causal estimators. Our work shows that foundation models for tabular data can be useful for causal inference, but only when their uncertainty is carefully calibrated.