Grounded Time-Series Anomaly Detection with Compact Multimodal Models and Preference Optimization
Abstract
We study time-series anomaly detection (TSAD) with compact multimodal models that jointly perform anomaly interval localization and generate verifiable natural-language explanations. We introduce a staged V0--V3 framework that progressively increases input richness and output structure, enabling a 3B vision-language model to match GPT-5 on out-of-distribution benchmarks while producing citation-grounded explanations. A key limitation is that post-SFT explanations exhibit structured grounding errors that invalidate factual correctness under strict verification. We therefore adopt DPOP, a positive-anchoring extension, to preserve the likelihood of well-formed outputs during optimization, raising the fully-truthful explanation rate from 72.4\% to 92.6\% while maintaining localization performance. Our results show that explicit anchoring makes grounded preference optimization practical for reliable structured reasoning with compact multimodal models under strict output constraints.