Signature-Kernel Based Evaluation Metrics for Robust Probabilistic and Tail-Event Forecasting
Benjamin Redhead ⋅ Thomas Lee ⋅ Peng Gu ⋅ Víctor Elvira ⋅ Amos Storkey
Abstract
Standard sample-based metrics like Continuous Ranked Probability Score (CRPS) and Quantile Loss (QL) are frequently used in the evaluation of probabilistic time-series forecasting. However, these metrics fail to capture multivariate correlations or assess the ability to capture distribution tails. These global metrics are insensitive to errors on high-utility events in distributional tails due to over-representation of the distribution's body. To accurately measure distributional fidelity on tail-regions and regions of high utility while still having a proper scoring rule, we introduce a family of censored signature kernel metrics. Our proposed metrics concentrate evaluation of forecasts to a focus region, representing high-utility or distributional tails, by collapsing the body of the distribution to a single pivot. Our benchmarks on time-series foundation models (TSFMs) reveal that while a clear ranking can be formed for distribution capture, there is no clear winner for tasks like systemic load prediction. This indicates that models with a strong ability to capture an overall distribution do not produce forecasts with high downstream utility. To encourage the use of the proposed metrics we open-source the efficient signature kernel (ESK) library. This library facilitates batch computation, with custom triton kernels, achieving a speed-up of up to 3.58$\times$ compared to implementations via the popular SigKernel library.
Chat is not available.
Successful Page Load