Formally Exploring Visual Anomaly Detection Evaluation Metrics
Abstract
Inaccurate Visual Anomaly Detection (VAD) can lead to critical failures in safety-sensitive domains, including autonomous navigation and industrial surveillance. With the increasing abundance and rapid proliferation of VAD algorithms, their reliable evaluation has become increasingly important and challenging. Commonly used evaluation metrics often fail to capture practically relevant aspects of model behavior, yielding inconsistent or misleading assessments by overlooking errors such as redundant detections and the spatial distribution of false positives. In this paper, we formalize the requirements for VAD evaluation by introducing a set of axiomatic, verifiable properties that an evaluation metric should satisfy. Through a systematic analysis of state-of-the-art evaluation methods, we show that none satisfies all proposed properties. To address this gap, we introduce SAAM-ALARM, a novel evaluation metric that satisfies these properties. Our results show that SAAM-ALARM provides a more nuanced and theoretically sound assessment, establishing a stronger standard for performance benchmarking in VAD.
Lay Summary
Visual Anomaly Detection systems are used in safety-concerned settings such as self-driving vehicles and industrial monitoring to detect unusual or potentially dangerous events. While many methods have been developed for this task, it is still difficult to fairly evaluate and compare them, because commonly used evaluation methods can sometimes give confusing or misleading results. In this work, we study the limitations of current evaluation practices and show that they often fail to capture important types of mistakes made by these systems. To better understand what a good evaluation should look like, we define a set of simple, clear requirements that evaluation methods should satisfy. We find that existing approaches do not fully meet these requirements. To address this, we introduce a new evaluation method called SAAM-ALARM, which is designed to provide a more reliable way of measuring performance. Our results show that it gives a clearer and more consistent picture of how well anomaly detection systems actually work.