Towards Benchmarking Time Series Foundation Models on Native Scientific Data
Lewis ODonnell ⋅ Arinbjörn Kolbeinsson ⋅ Benedikt Kolbeinsson ⋅ Marc Deisenroth
Abstract
Time series foundation models are evaluated on benchmarks that pre-align observations to a uniform grid, discarding the irregular, multi-rate structure present in real scientific data. We show empirically that preprocessing-induced degradation is large, architecture-specific, and invisible under conventional benchmarks. We propose a cross-domain benchmark spanning nuclear fusion, healthcare, and climate, stored in a unified observation set format that preserves native timestamps and mixed modalities without grid alignment and define four evaluation tasks and a hierarchical metric system to assess model capabilities.
Chat is not available.
Successful Page Load