Contrastive Diffusion Alignment: Learning Structured Latents for Controllable Generation
Abstract
Diffusion models excel at generation, but their latent spaces are high dimensional and not explicitly organized for interpretation or control. We introduce ConDA (Contrastive Diffusion Alignment), a plug-and-play geometry layer that applies contrastive learning to pretrained diffusion latents using auxiliary variables (e.g., time, stimulation parameters, facial action units). ConDA learns a low-dimensional embedding whose directions align with underlying dynamical factors, consistent with recent contrastive learning results on structured and disentangled representations. In this embedding, simple nonlinear trajectories support smooth interpolation, extrapolation, and counterfactual editing while rendering remains in the original diffusion space. ConDA separates editing and rendering by lifting embedding trajectories back to diffusion latents with a neighborhood-preserving kNN decoder and is robust across inversion solvers. Across fluid dynamics, neural calcium imaging, therapeutic neurostimulation, facial expression dynamics, and monkey motor cortex activity, ConDA yields more interpretable and controllable latent structure than linear traversals and conditioning-based baselines, indicating that diffusion latents encode dynamics-relevant structure that can be exploited by an explicit contrastive geometry layer.
Lay Summary
AI models can now create realistic images and scientific data, but it is often hard to understand or control exactly how they make changes. This is a major challenge in science and medicine, where researchers want to study how complex systems evolve, respond to treatments, or change under different conditions. We introduce ConDA, a method that helps make these AI models easier to guide. It uses simple information, such as time, experimental settings, facial movements, or brain stimulation levels, to organize the model’s internal knowledge in a more meaningful way. This makes it easier to create smooth changes, predict future states, and explore “what-if” scenarios. We test ConDA on examples from fluid flow, brain activity, facial expressions, and movement planning. Our results show that AI models can capture important patterns in complex systems, and ConDA helps researchers uncover and control those patterns more clearly.