INFER: Learning Implicit Neural Frequency Response Fields for Confined Acoustic Environments
Abstract
Neural acoustic fields often model time-domain impulse responses, which struggle to capture the frequency-selective wave behaviors that dominate confined, resonant environments. To address this, we propose INFER (Implicit Neural Frequency Response fields), a framework that directly learns continuous, complex-valued frequency response fields. Unlike prior time-domain methods, our frequency-first approach enables three key innovations: (1) end-to-end learning of frequency-specific attenuation and phase delay in 3D space; (2) a physics-based Kramers–Kronig consistency constraint that causally regularizes attenuation and phase delay; and (3) perceptual and hardware-aware spectral supervision that prioritizes critical auditory bands. We evaluate INFER across diverse settings, ranging from standard room-scale benchmarks (MeshRIR, RAF) to challenging, highly reverberant environments like real car cabins. Our approach significantly outperforms time- and hybrid-domain baselines, reducing average magnitude and phase reconstruction errors by over 39\% and 51\%, respectively, demonstrating state-of-the-art accuracy in modeling complex acoustic spaces.
Lay Summary
The way sound behaves inside a small, enclosed space like a car cabin is surprisingly complicated. The seats, windows, dashboard, and trim each absorb and reflect sound differently, producing strong echoes and resonances. To deliver clean, accurately positioned audio (whether a song, a phone call, or a directional safety alert warning a driver of a hazard on the right) engineers need a precise acoustic map of the cabin. Today such maps are built by hand-tuning, painstaking in-car measurements, or expensive simulations that quickly become inaccurate when seats recline or windows roll down. We built a system called INFER that learns this acoustic map directly from a modest set of measurements. Unlike earlier learning-based methods that first reconstruct the sound waveform and then convert it, INFER works directly in terms of how each frequency, or pitch, travels through the space. It links how sound fades to how it shifts in time, keeping the learned map physically realistic. On both standard rooms and real car cabins, INFER reconstructs sound more accurately than prior methods, advancing high-fidelity in-vehicle audio and directional safety alerts.