Mitigating Label Shift in Tabular In-Context Learning via Test-Time Posterior Adjustment
Abstract
TabPFN has recently gained attention as a foundation model for tabular datasets, achieving strong performance by leveraging in-context learning on synthetic data. However, we find that TabPFN is vulnerable to label shift, often overfitting to the majority class in the training dataset. To address this limitation, we propose DistPFN, the first test-time posterior adjustment method designed for tabular foundation models. DistPFN rescales predicted class probabilities by downweighting the influence of the training prior (i.e., the class distribution of the context) and emphasizing the contribution of the model’s predicted posterior, without architectural modification or additional training. We further introduce DistPFN-T, which incorporates temperature scaling to adaptively control the adjustment strength based on the discrepancy between prior and posterior. We evaluate our methods on over 250 OpenML datasets, demonstrating substantial improvements for various TabPFN-based models in classification tasks under label shift, while maintaining strong performance in standard settings without label shift. Code is available at this repository: https://github.com/seunghan96/DistPFN.
Lay Summary
Tabular foundation models have recently shown strong performance on structured datasets such as medical records, finance, and scientific data. However, we find that these models can become biased when the class distribution in the training examples differs from that of the test data, often favoring majority classes and making unreliable predictions. To address this issue, we propose a simple test-time adjustment method called DistPFN that recalibrates predicted probabilities without requiring additional training or changes to the model architecture. Our method improves robustness under distribution shifts while preserving the original strengths of the foundation model. Extensive experiments across diverse datasets demonstrate that DistPFN consistently improves prediction reliability and classification performance.