Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution
Abstract
Generative priors in Image Super-Resolution (SR) often compromise faithful restoration, we attribute this limitation to a fundamental spectral misalignment between isotropic objectives and the intrinsic natural image manifold. While Direct Preference Optimization offers a path to alignment, its reliance on spectrally flat Gaussian noise fails to distinguish authentic high-frequency details from hallucinations. To bridge this geometric gap, we propose ASASR, a theoretically grounded framework that recasts the generative flow into a Sobolev-induced Riemannian geometry by explicitly coloring the noise transition kernel to mirror natural spectral decay. Driving this geometric alignment, we integrate a parametric adversary grounded in the Riesz Representation Theorem, which synthesizes targeted negative samples equivalent to worst-case Sobolev gradients to direct optimization along the tangent space of plausible structural failures. Extensive evaluations demonstrate that ASASR outperforms leading generative baselines, particularly in preserving spectral consistency and structural fidelity, offering a robust solution that effectively mitigates artifacts.
Lay Summary
Image super-resolution aims to recover clear high-resolution images from low-resolution inputs. Although recent generative methods can create sharp and realistic-looking results, they often introduce hallucinated details that are not faithful to the original image. We identify this issue as a spectral mismatch: common training objectives and Gaussian noise treat image frequencies too uniformly, while natural images follow structured frequency patterns. To address this, we propose ASASR, a framework that aligns the generative process with the natural image spectrum. Instead of using spectrally flat noise, ASASR colors the noise transition to reflect natural spectral decay, encouraging more faithful restoration of edges, textures, and structures. We also introduce an adversarial mechanism that generates targeted negative samples, helping the model recognize and avoid plausible but incorrect details. Experiments show that ASASR improves spectral consistency and structural fidelity over leading generative super-resolution methods, producing sharper images with fewer artifacts and hallucinations.