StarEmbed: Benchmarking Time Series Foundation Models on Astronomical Observations of Variable Stars
Weijian Li ⋅ Hong-Yu Chen ⋅ Nabeel Rehemtulla ⋅ Ved Shah ⋅ Dongho Kim ⋅ Dennis Wu ⋅ Qinjie Lin ⋅ Adam Miller ⋅ Han Liu
Abstract
Current time series foundation model (TSFM) training corpora largely omit data with certain complexities like irregular temporal sampling. Astronomical time series of stellar fluxes (``light curves'') are available in immense quantities and exhibit irregular sampling, multiple variates, and heteroskedasticity. We introduce $\texttt{StarEmbed}$, the first public benchmark for light curves comprised of real observations of $\sim$40,000 stars across seven classes and evaluations in clustering, classification, and out-of-distribution (OOD) source detection. We benchmark TSFMs with differing architecture and training strategies as well as domain-specific transformers. Our results demonstrate that the $\texttt{Chronos}$ family, despite being pre-trained on regularly sampled non-astronomical data, yields state-of-the-art (SOTA) performance in light curve clustering and OOD detection. While no TSFM strictly surpasses the classification performance of the long-established domain baseline, they do demonstrate excellent generalization abilities. $\texttt{StarEmbed}$ marks a step toward universal light curve embeddings and improved TSFM performance on challenging data.
Lay Summary
Modern observatories collect brightness measurements for billions of stars, creating far more time-varying data than can be efficiently analyzed with traditional, hand-designed astronomy pipelines. These data are also especially challenging for AI because telescope observations are irregular: measurements may be separated by minutes, days, or months, and the data are noisy or incomplete. This makes astronomy a challenging test case for time-series foundation models, which are AI models designed to understand many kinds of time-varying data. We introduce $\texttt{StarEmbed}$, the first public benchmark built from real observations of about 40,000 variable stars. These stars change brightness in repeating patterns, and our benchmark tests whether general-purpose time-series AI models can group similar stars, classify star types, and identify unusual stars that do not fit known categories. Our results show that some general-purpose time-series models, even though they were not trained on astronomy data, can learn useful patterns from stars’ changing brightness over time. They perform especially well at finding unusual stars, while traditional astronomy features still remain strongest for some classification tasks. $\texttt{StarEmbed}$ provides a shared testbed for improving AI models on irregular scientific time series and will help astronomers choose the best AI model for analyzing the massive datasets produced by modern observatories.
Successful Page Load