PolarDepth: Monocular Transparent Object Depth from Polar-Physics Priors
Abstract
Depth estimation for transparent objects remains a fundamental challenge, as RGB-based cues often fail in regions affected by refraction and light transmission. Polarization provides physically grounded information related to surface orientation and material properties, offering reliable geometric cues even in the absence of texture. In this work, we introduce PolarDepth, a monocular framework that incorporates both RGB and polarization inputs, including the degree and angle of linear polarization (DoLP and AoLP), to estimate dense depth and localize transparent regions. PolarDepth injects polarization-derived physical priors by estimating the refractive index, zenith angle, and azimuth angle from polarization measurements and embedding them into an implicit geometric representation that constrains depth inference in ambiguous transparent regions. To support model development and evaluation, we introduce PTOD, a dataset with synchronized RGB, polarization, and depth data and manually annotated transparent region masks. Experimental results demonstrate that PolarDepth achieves state-of-the-art performance in transparent object depth estimation. The findings highlight the effectiveness of embedding polarization-derived physical priors into learned representations for robust perception in complex visual environments.
Lay Summary
Standard AI systems often struggle to "see" transparent objects like glass bottles or windows because they rely on color and texture, which glass lacks. Instead of detecting the object, the AI "sees through" it to the background, causing robots to drop glassware or autonomous systems to miss clear obstacles. Our research introduces PolarDepth, a system that reveals these invisible surfaces by analyzing "polarization"—a hidden property of light that describes how light waves twist as they bounce off a surface. Just as polarized sunglasses filter out glare to reveal the world more clearly, PolarDepth uses the physics of light to calculate the exact shape and distance of transparent surfaces. We combine standard camera images with these physical light patterns to create a reliable 3D map of transparent objects, even in messy or low-contrast environments. To support future innovation, we have also released a massive new dataset of glass objects captured with specialized light sensors. This work paves the way for smarter robotic assistants in homes and laboratories and ensures that future autonomous technologies can navigate our transparent world with newfound safety and clarity.