FactoryNet: A Large-Scale Dataset toward Industrial Time-Series Foundation Models
Abstract
We introduce the first universal pretraining corpus for industrial time-series data: FactoryNet. FactoryNet contains 51 million datapoints across 23,000 end-to-end task executions (13,300 real, 9,800 synthetic) on six embodiments. These are unified by a shared schema that enables robust zero-shot cross-embodiment transfer and highly parameter-efficient anomaly detection. We introduce a novel schema: Setpoint, Effort, Feedback, Context (S-E-F-C), which underlies the entire pipeline. This schema maps any actuated system into a common representational frame. The corpus spans 27 annotated anomaly types, alongside healthy baselines and counterfactual pairs across robotic manipulation and machining domains. Cross-embodiment transfer experiments yield positive results: under bias-aware metrics, our model demonstrates fair cross-embodiment transfer capabilities on the evaluated source-target pair. Additionally, 24 schema-aligned signals achieve competitive anomaly detection performance compared to high-dimensional baselines. We release FactoryNet as a growing, multi-embodiment dataset to drive progress toward industrial foundation models.