Learning to Watch: Active Video Anomaly Understanding via Interleaved Policy Optimization
Abstract
Lay Summary
Long videos often contain only a few moments that explain what is really happening. In safety-related videos, for example, a person may seem to be standing normally until a brief frame reveals a fall, a fight, or another unusual event. Many AI systems make their decision after looking at a fixed set of snapshots. If those snapshots miss the important moment, the system has little chance to correct itself. Our work builds an AI video reviewer, Anom-pi, that can decide when it needs to look again. It can zoom in on selected moments, check what happened just before or after them, and inspect suspicious parts of the video more carefully. It also learns not to keep looking forever: once it has enough evidence, it stops and gives a clear answer. Across several public video collections, this active reviewing strategy helps the system make better decisions in difficult cases while avoiding unnecessary extra viewing.