Stable Spectral Copula Alignment for Robust Multimodal Learning
Abstract
Multimodal alignment can fail under deployment shift because standard objectives entangle cross-modal dependence with marginal-sensitive geometry. Stable Spectral Copula Alignment (SSCA) provides a deployment protocol for copula-stable dependence under approximately coordinate-wise monotone marginal distortions, together with auditable, label-free diagnostics for monitoring and mitigation. SSCA combines (i) clipped soft-rank Gaussianization that suppresses marginal effects while tracking tie and approximation errors, (ii) dependence-weighted sliced Wasserstein hub coupling for globally coherent multiway alignment with cycle auditing, and (iii) diagonal-stabilized block-spectral learning with eigengap-normalized Davis-Kahan diagnostics, yielding an actionable subspace-risk inequality. A calibrated gate maps diagnostic proxies to a reliability signal with a measurable false-alarm/miss trade-off, enabling stability-mode updates, budgeted remediation, and conservative no-update fallback for out-of-scope drift. Evaluations on MOSEI/MELD, MSCOCO, and CC3M-500K show improved performance under perturbation and substantially reduced degradation under controlled monotone distortions, raw-pipeline drifts, and frozen-feature retrieval stress tests.
Lay Summary
Many AI systems need to combine text, images, sound, and video. In real use, these inputs often change after deployment. Images may be compressed, sensors may use different settings, software may be updated, or some signals may be missing. These changes can make a system less reliable, even when the useful connection between the inputs is still present. This work introduces a way to make such systems more stable. It reduces sensitivity to harmless changes in scale or format, keeps different input types matched more consistently, and checks whether the current data is still safe to learn from. When the check suggests that the data has moved outside the reliable range, the system stops risky updates and falls back to a safer setting. Tests on emotion recognition and image-text search show smaller performance drops under both controlled changes and more realistic deployment changes. The method also provides warning signals that make deployed systems easier to monitor.