Blind Audio Restoration using Contrastive Diffusion Guidance
Abstract
We study blind audio restoration, where recorded signals are degraded by unknown and potentially composite corruptions. This setting is intrinsically ill-posed: many clean audio signals may be compatible with the same degraded observation. We address this ambiguity with a diffusion-based posterior sampler that generates restorations that are compatible with the observed recording. However, unlike standard generative inverse problems, the absence of a known forward model makes the likelihood score difficult to compute. Rather than estimating the degradation operator or relying on a partial approximation, we learn a contrastive embedding space in which measurement consistency can be evaluated directly. We show that the resulting embedding-space guidance provides a practical surrogate for the true likelihood score, enabling the diffusion sampler to move toward the desired posterior distribution. Experiments on historical piano recordings suggest that AudioCoGuide can effectively address fully blind audio inverse problems through contrastive guidance.