Can We Build a Monolithic Model for Fake Image Detection? SICA: Semantic-Induced Constrained Adaptation for Unified-Yet-Discriminative Artifact Feature Space Reconstruction
Bo Du ⋅ Xiaochen Ma ⋅ Xuekang Zhu ⋅ Zhe Yang ⋅ Chaoqun Niu ⋅ Jian liu ⋅ Ji-Zhe Zhou
Abstract
Fake Image Detection (FID), aiming at unified detection across four image forensic subdomains, is critical in real-world forensic scenarios. Compared with ensemble approaches, monolithic FID models are theoretically more promising, but to date, consistently yield inferior performance in practice. In this work, we identify the intrinsic distinctness of artifacts across subdomains—a critical barrier we term the "Ji-Zhe phenomenon". Driven by this phenomenon, we diagnose the cause of this underperformance for the first time: the collapse of the artifact feature space. The core challenge for developing a practical monolithic FID model thus boils down to the "unified-yet-discriminative" reconstruction of the artifact feature space. To address this paradoxical challenge, we hypothesize that high-level semantics can serve as a structural prior for the reconstruction, and further propose Semantic-Induced Constrained Adaptation (SICA), the first monolithic FID paradigm. Extensive experiments on our $ \textit{OpenMMSec} $ dataset demonstrate that SICA outperforms 15 state-of-the-art methods and reconstructs the target unified-yet-discriminative artifact feature space in a near-orthogonal manner, thus firmly validating our hypothesis. The code and dataset will be made publicly available.
Lay Summary
Detecting fake images is crucial, but building a single AI tool to catch every type of forgery, from deepfakes to AI-generated art, is difficult because the visual clues for each are completely different. We discovered that forcing an AI to learn all these unrelated clues at once causes its learning system to collapse. To solve this, we developed a new method that uses the image's overall context (like whether it is a face or a document) to organize and safely guide the AI in spotting specific fake traces. Tested on our massive new dataset of over 330,000 images, this single, unified model outperformed 15 top-tier tools, bringing us closer to a reliable, all-in-one fake image detector.
Successful Page Load