Position: Multi-Omics Foundation Models Need a Modality Identifiability Standard, Not Just Aggregate Accuracy
Abstract
Multi-omics foundation models that integrate DNA, RNA, protein, and other biological modalities are increasingly used for phenotype prediction, drug response, and disease modeling (Cui et al. 2024; Theodoris et al. 2023). The standard reporting practice evaluates these models by aggregate prediction accuracy on held-out cells or patients. We argue this practice obscures a structural identifiability failure: when modalities partially measure the same underlying biological signal (a near-universal property of multi-omics data, where DNA variants, RNA expression, and protein abundance all reflect overlapping regulatory mechanisms), the contribution of each modality to phenotype prediction is not identifiable from aggregate data alone, regardless of cohort size. We support the position with a controlled simulation: with 3 modalities, 5 features each, and 500 samples, the per-modality posterior std on coefficients collapses from 0.005 (full-rank, identifiable) to 0.817 (rank-deficient, essentially the prior) as the shared-signal dimension grows from 0 to 5; the design rank drops from 15 to 5. Critically, even in the unidentifiable case, the sum of modality contributions is recovered exactly, which is exactly what the aggregate accuracy metric measures, while the per-modality contributions remain at the prior. Claims like "the protein modality drives prediction" or "adding DNA improves performance" may not be licensed by the design even when aggregate accuracy is high. We propose a four-item modality identifiability disclosure standard for FM4LS papers: modality-design rank reporting, per-modality posterior contraction quantification, shared-signal estimation, and ablation-vs-identification distinction.