ConTSG-Bench: A Unified Benchmark for Conditional Time Series Generation
Abstract
Conditional time series generation plays a critical role in addressing data scarcity and enabling causal analysis in real-world applications. Despite its increasing importance, the field lacks a standardized and systematic benchmarking framework for evaluating generative models across diverse conditions. To address this gap, we introduce the Conditional Time Series Generation Benchmark (ConTSG-Bench). ConTSG-Bench comprises a suite of large-scale, well-aligned datasets spanning diverse conditioning modalities and levels of semantic abstraction, enabling systematic evaluation of representative generation methods across these dimensions with a comprehensive suite of metrics for generation fidelity and condition adherence. Both the quantitative benchmarking and in-depth analyses of conditional generation behaviors have revealed the traits and limitations of the current approaches, highlighting critical challenges and promising research directions, particularly with respect to precise structural controllability and downstream task utility under complex conditions.
Lay Summary
Many important decisions rely on data collected over time, such as heart signals, weather records, traffic flows, and energy use. In many settings, there is too little real data, or the data is hard to share, so researchers build AI systems that create artificial examples. These examples are useful only when they look realistic and follow a user’s request, such as “make the signal rise steadily” or “show a pattern linked to this weather condition.” Fair comparison has been difficult, because studies often use different datasets, requests, and tests. We introduce ConTSG-Bench, a shared testbed that brings together multiple datasets and several ways of giving requests, including simple labels, structured descriptions, and everyday language. It checks whether generated data looks real, follows the request, handles detailed local instructions, works on new combinations of requests, and helps train other models. Our experiments show that language-based requests can be powerful, but current systems still struggle with precise details and unfamiliar combinations. By releasing common data, tests, and code, ConTSG-Bench can help researchers build more reliable tools for creating time-based data in areas such as health, weather, traffic, and energy.