\plurel-to-\rdbpfn: Schema-Guided Synthetic Relational Pretraining
Mohammad S Abolhasani ⋅ Viswanath Ganapathy
Abstract
Relational Foundation Models (RFMs) require large-scale synthetic relational databases for pretraining, but existing approaches tightly couple data generation with the model training pipeline. We study whether \plurel{}, a general-purpose synthetic relational database generator, can serve as an external data source for \rdbpfn{}, a relational in-context learner originally pretrained on its own specialized synthetic prior using ${\sim}1.8$M tasks across two stages. We build a conversion pipeline that maps \plurel{}-generated databases—including externally constructed binary prediction tasks—into the \rdbpfn{} training format and evaluate three curriculum strategies: \sgf{} (real-world schema then fully synthetic), \fs{} (diverse synthetic schemas throughout), and \sgl{} (fully synthetic then real-world schema). Using only ${\sim}5{,}500$ relational databases (${\sim}33$K tasks)—roughly $55\times$ fewer tasks than the original protocol—and no single-table warm-up, our best curriculum (\sgf{}) achieves $0.6346$ average ROC-AUC across 19 real benchmark tasks at 1024-shot context, recovering 87.6\% of the published \rdbpfn{} performance ($0.7245$). At 64-shot context, the gap narrows to 93.8\% ($0.6116$ vs.\ $0.6517$). Our results demonstrate that external synthetic generators can provide useful pretraining signals for RFMs when combined with appropriate curriculum design and that exposure to a real-world schema early in training is substantially more effective than late-stage schema adaptation.
Chat is not available.
Successful Page Load