When Simpler ICL Outperforms Pretrained Tabular Foundation Models for RNA Editing
Ran Eisenberg ⋅ Efraim Rahamim ⋅ Erez Levanon ⋅ Ofir Lindenbaum
Abstract
Pretrained tabular In-Context Learning (ICL) models promise to transfer to new structured-data tasks, but biomedical tables often differ sharply from their benchmark regimes. We study RNA editing prediction from single-cell gene expression profiles, where the model must predict each cell's editing index. Evaluation against tabular foundation models shows they achieve lower performance and are slower than a task-trained alternative. We investigate whether ICL can be simplified: instead of deploying a large pretrained ICL model directly, we train an explicit retrieval-based ICL adapter with attention-based multiple-instance learning (MIL) over genes and a gated correction from similar labeled training cells. This task-trained adapter achieves the best overall rank correlation, outperforming both tabular foundation models and task-trained baselines on most tissues, while requiring only ${\approx}1.4$ minutes of inference per fold, a $45{\times}$ speedup. Its additive prediction form separates the query-cell gene score from the retrieved-cell context correction, providing gene-level and support-set explanations without post-hoc attributions. For some shifted biomedical regression tables, simpler domain-structured ICL can be stronger, faster, and more interpretable than direct pretrained tabular ICL.
Chat is not available.
Successful Page Load