Causal Foundation Models Perform Better without Post-treatment Variables
Abstract
Causal Foundation Models (CFMs) amortize Bayesian causal inference by pretraining on synthetic datasets, enabling zero-shot conditional average treatment effect (CATE) estimation. This paper studies the structural bias induced when post-treatment covariates are included in the inference-time query set of CFMs. The bias decomposes into three terms, and the CATE estimation error of two representative CFMs, Do-PFN and CausalPFN, is consistent with the corresponding theoretical bounds under mediator and collider conditioning. In an Oracle-Exclude experiment, removing post-treatment covariates from the query set reduces estimation error by approximately 25% for Do-PFN and up to 72% for CausalPFN without retraining. As a practical alternative to oracle exclusion, Treatment-Centric Local Discovery (TC-LD) filters query covariates before inference. It recovers 85.5%-92.6% of the Oracle headroom on the synthetic benchmark and detects all synthetically injected mediators in semi-synthetic experiments on IHDP and ACIC 2016.