Position: Efficient Dual-View AI-Generated Video Detection Needs a Visual-Language Dual View
Abstract
High-fidelity AI-generated video (AIGC-V) challenges the use of video as reliable evidence in multimodal systems. We argue that efficient AIGC-V detection requires a vision-language dual view because authenticity evidence appears at different levels of abstraction and cost: from local artifacts and temporal coherence to cross-modal alignment, provenance, and world-level plausibility. Rather than treating efficiency as a smaller-model objective alone, we frame it as evidence-aware verification: identifying the level of evidence that is both relevant to the suspected manipulation and sufficient for a trustworthy authenticity judgment. To ground this position, we reinterpret AIGC-V detection as Factual Fidelity Verification and organize 230 works through a four-layer Vision-Language Dual-View taxonomy. The visual view covers intrinsic cue analysis and spatiotemporal consistency, while the language view covers cross-modal consistency and world-level reasoning. Across these layers, existing methods reveal a trade-off between evidence richness, computational cost, and human interpretability. This framing connects AIGC-V detection to resource-aware video retrieval, reasoning, and question answering without recasting detection itself as open-ended QA. We close with a research agenda for robust, cost-aware, and human-auditable video verification.