Where Concept Erasure Should Occur: Concept–Layer Alignment in Text-to-Video Diffusion Models
Abstract
Text-to-video diffusion transformers encode semantic information unevenly across model depth, which constrains effective concept erasure. We identify a representational bottleneck, termed concept–layer topological alignment, under which target concepts exhibit higher separability at certain representational depths. Outside these depths, concept and non-target signals remain strongly entangled, limiting the effectiveness of depth-specific erasure. This observation reframes concept erasure as the problem of identifying representational depths where concept–non-target separation naturally emerges. Motivated by this structural constraint, we introduce CLEAR, a separability-driven optimization framework for concept erasure that explicitly enforces concept–layer alignment. CLEAR operationalizes this principle by formulating layer selection as an optimization problem over concept–non-target separability, rather than relying on layer-agnostic or heuristic choices. To enable this, we introduce a separability-aware objective that favors layers exhibiting stronger concept–non-target separation. Experiments on large-scale text-to-video diffusion models demonstrate that enforcing concept--layer alignment leads to more precise concept suppression while preserving overall generative quality.
Lay Summary
Modern text-to-video AI systems can generate highly realistic videos, but they may also produce harmful, copyrighted, or sensitive content. Existing methods for removing such concepts often apply changes at fixed locations inside the model, which can lead to incomplete removal or unintended damage to overall video quality. In this work, we show that different concepts are represented at different depths inside the model, and we introduce a method called CLEAR that automatically identifies where concept removal should occur. Our approach removes unwanted concepts more effectively while better preserving video quality and semantic consistency, helping make AI video generation systems safer and more controllable.