Position: Saturation in Single-Cell Foundation Model Benchmarks Signals Identifiability Failure, Not Solved Capability
Abstract
Single-cell foundation models (scFMs) including scGPT (Cui et al. 2024), Geneformer (Theodoris et al. 2023), and scFoundation now report cell-type classification accuracies above 85-95% on standard benchmarks (Hou et al. 2026). The community routinely interprets clustering of top scFMs near these accuracies as evidence that "the benchmark is solved" or that "differences between scFMs are negligible." We argue both interpretations are wrong. Saturation is, structurally, an identifiability-failure signal: as observed pass-rates compress against the upper bound, the Fisher information on the latent capability of each scFM collapses exponentially, and the data become uninformative about pairwise capability contrasts. We support the claim with a simulation modeling 13 scFMs on 200 benchmark cells: as mean pass-rate rises from 0.50 to 0.95, the probability that the empirically-best scFM is also the truly-best scFM drops from 0.76 to 0.35, and Fisher information at phat=0.95 is only 18% of its maximum at phat=0.5 (theta=3 vs theta=0). Bayesian posterior probabilities Pr(thetatop > thetasecond | y) degrade similarly. We propose a four-item identifiability-aware scFM benchmarking standard: pass-rate distribution disclosure, Fisher information reporting, posterior pairwise comparison, and saturation-triggered redesign protocol.