Rethinking Forgery Attacks on Semantic Watermarks in Black-Box Settings: A Geometric Distortion Perspective
Abstract
Recent studies have shown that semantic watermarks, which embed information into the initial noise of latent diffusion models (LDMs), are vulnerable to black-box forgery attacks. However, existing methods primarily rely on empirical evidence and lack a rigorous theoretical understanding of the conditions under which such attacks succeed or fail. To bridge this gap, we rethink the nature of such attacks through the lens of rate-distortion in the latent space. Our analysis identifies an irreducible distortion floor due to structural mismatches between proxy and target models, which fundamentally limits the fidelity of forged watermarks. We further characterize this distortion as structured geometric deviations on the latent manifold, in the form of global drift and local deformation rather than stochastic noise. Leveraging these insights, we propose a scheme-agnostic detection method that distinguishes forged samples before watermark verification. Extensive experiments demonstrate the effectiveness of our method across diverse black-box scenarios, while preserving robustness to common distortions.
Lay Summary
How can we tell whether an AI-generated image really came from the system claimed by its watermark? Watermarks are meant to leave hidden signs inside generated images, helping people trace where an image came from. But recent attacks show that someone can take a watermarked image, run it through another image-generation model, and create a new image that still appears to carry the original watermark. This could falsely blame an image service for content it did not actually produce. We study why these attacks are hard to carry out perfectly. Our key finding is that when an attacker uses a different model, the image’s hidden generation record changes in subtle but structured ways. These changes are like fingerprints left by the forgery process. Based on this observation, we design a method that checks whether a watermarked image looks genuine or forged before applying the usual watermark test. Our method can be added to existing watermarking systems without redesigning them, and experiments show that it works across different image-generation models, watermarking methods, and attack settings.