The Perception–Physics Paradox: Probing Scientific Alignment with TC-Bench
Abstract
While Vision Foundation Models (VFMs) excel at predictive tasks on satellite imagery, their performance can arise from visual correlations rather than underlying structural invariants, making certain perception-based out-of-distribution accuracy a poor proxy for scientific utility. As a result, models may look correct without reasoning correctly—a discrepancy we term the Perception–Physics Paradox. To address this gap, we introduce Scientific Alignment as an implicit objective for representation learning in scientific domains. We study a principled, testable aspect of scientific alignment through Structural Isomorphism, which requires latent representations to uniquely identify physical systems up to a linear reparameterization. This perspective induces a hierarchy of necessary conditions and yields a systematic probing protocol for physical and causal interpretability. To operationalize this framework, we release TC-Bench, a foundational global dataset and automated construction pipeline for tropical cyclone research, and show that current VFMs rely on visual shortcuts that collapse in extreme regimes, indicating that scientific alignment does not arise as a natural byproduct of visual scaling alone.
Lay Summary
Artificial intelligence models are increasingly used to analyze satellite images for scientific problems, including weather and climate. However, a model that appears to make accurate predictions from images may still rely on visual shortcuts rather than understanding the underlying physical system. In this paper, we study this issue for tropical cyclones, where small physical differences can matter greatly for storm intensity, even when satellite images look very similar. We introduce the Perception–Physics Paradox: AI models can look correct without reasoning correctly. To test this, we build TC-Bench, a global and reproducible benchmark dataset for tropical cyclone research, and evaluate several widely used vision foundation models on satellite images of storms. We find that these models often perform reasonably well on average, but their internal representations break down for intense cyclones, precisely where accurate physical understanding matters most. Our results suggest that scaling up visual models alone is not enough for scientific use. Future models should be evaluated not only by prediction accuracy, but also by whether they preserve meaningful physical structure.