Recall Residualisation: Decontaminating Foundation-Model Evaluation on Public Time-Series Benchmarks
Abstract
Foundation-model evaluations on public time-series benchmarks risk conflating forecasting skill with parametric recall: factor returns, macroeconomic indicators, and climate indices are widely mirrored in pretraining data, so a model’s apparent “skill” on such a benchmark may reflect retrieval of labels rather than predictive ability. Standard evaluation cannot distinguish the two. We propose Recall Residualization: regress any LLM-derived signal on the model’s own direct value-recall of the benchmark, and report the residual signal–truth correlation as decontaminated skill. A worst-case orthogonal decomposition gives an upper bound; OLS gives a co-located linear point estimate. In a 20-cell study across four benchmarks and five frontier LLMs, apparent skill up to |ρ| = 0.93 collapses to |ρ| ≤ 0.17 after residualization, with LeakShare ≥ 0.97 in 19/20 cells. Dedicated time-series foundation models forecast the same months at or near AR(1)-level performance under their native numeric-history interface, indicating that the contamination channel is specific to date- and label-conditioned text queries. Code here: https://anonymous.4open.science/r/recall-residualization-1D49