Implicit Reward Alignment For Training Causally-Coherent Tabular Data Generators
Abstract
Foundation models for structured data are increasingly being used as queryable generators for scenario planning and counterfactual analysis, requiring models that are statistically realistic, causally coherent, and capable of conditional querying from partial inputs. Yet existing approaches to tabular data generation either optimize solely for distributional fidelity, or impose causality through explicit structural assumptions, which are untenable in a real-world setting. We argue that counterfactual queryability is a key missing reliability axis, for tabular FMs and introduce Causal Reward Aligned Fine-Tuning (CRAFT), a reinforcement learning framework where language models are trained to produce causally-consistent samples only through implicit reward signals. Across multiple semi-synthetic and benchmark settings, we compare against 15 baseline models spanning LLM-, diffusion-, VAE-, and GAN-based generators, using multiple metrics along both realism and causal coherence dimensions, and show that reward-based alignment improves counterfactual accuracy and treatment-effect estimation, even in settings where baseline models achieve comparable distributional realism. More broadly, these findings suggest that intervention-consistent generative behavior can emerge from alignment objectives alone, without strong architectural priors or explicit encoding of the causal graph.