Nonlinear RNNs as a Compute Shortcut for Time Series Foundation Models
Levente Zolyomi ⋅ David Stap ⋅ Sebastian Böck ⋅ Günter Klambauer ⋅ Sepp Hochreiter
Abstract
Time series foundation models have converged on Transformer and linear-RNN backbones that scale efficiently but do not provide an explicit nonlinear recurrent state-update mechanism, despite nonstationarity and long-range dependencies being central to time series data. Pure nonlinear recurrent models track state but scale poorly, suggesting an expressivity-efficiency trade-off. We show the trade-off is avoidable: adding even a single nonlinear sLSTM layer to three popular scalable backbones (Transformer, Gated DeltaNet, and mLSTM) consistently improves forecasting accuracy on GIFT-Eval across five parameter scales (1M--80M). The gain is largest at small scale, reaching roughly 4--5\% for Transformer and Gated DeltaNet at 1M, and shrinks to under 1\% for most 80M comparisons. Yet at matched wall-clock training budget, the 10M hybrid outperforms a same-sized baseline trained with $\geq 4\times$ more compute, and a single sLSTM placed early in the stack recovers most of the benefit gained from adding more state tracking layers. Nonlinear state tracking therefore functions as a compute shortcut: a cheap inductive bias that substitutes for capacity and training time.
Chat is not available.
Successful Page Load