Round-Trip Latent Geometry in Diffusion VAEs Enables Covert Channels
Catherine Ge-Wang ⋅ Tushar Nagar ⋅ Joy Z. Yang
Abstract
Diffusion VAEs define learned latent interfaces through which small perturbations may survive decoding and re-encoding. We which latent positions and directions are preserved by the composite map $\mathrm{Enc}\circ\mathrm{Dec}$, and how this preservation enables covert image channels. Using training-free signed latent perturbations as probes, we measure spatial carrier stability, content dependence, directional gain, cross-VAE transfer, and monitor detectability across datasets and VAE checkpoints. We find that perturbation survival is highly non-uniform across latent positions and image content, and that a local directional-gain estimate better explains carrier reliability than a direction-agnostic stability heuristic. Across CIFAR-10, Caltech101, and a 1{,}000-image ImageNet-family subset, and across 3 VAE architectures, our perturbations are reliably recoverable with $>97\%$ bit accuracy at $\epsilon=2.0$, and the channel survives realistic image transformations at higher perturbation strengths. Our results suggest that covert communication in multimodal agent settings is mediated by interpretable structure in VAE round-trip geometry, and that representation-level monitoring is necessary for detecting such channels.
Chat is not available.
Successful Page Load