Cross-Tactile Sensor Representation Learning
Abstract
Visuo-tactile sensors have been widely adopted in robotic manipulation. However, inherent heterogeneity in sensor designs hinders the learning of unified tactile representations in cross-sensor scenarios. Existing methods that focus on reconstruction or task-specific supervision often fail to capture the common information between different tactile sensors, particularly in the presence of substantial sensor variations, resulting in limited generalization to unseen sensors. To address this, we propose Cross-Tactile Sensor Representation Learning (CTSRL), a unified framework for sensor-agnostic tactile representation learning. CTSRL introduces a Cross-Sensor Modulator (CSM) to eliminate sensor-specific biases and adopts a two-stage learning paradigm: (1) leveraging aligned synthetic data for cross-sensor self-supervised learning to extract shared latent representations across sensor domains; and (2) integrating real-world multimodal tactile data to bridge the sim-to-real semantic gap through cross-modal alignment, thereby enriching representations with fine-grained semantic attributes. Experimental results show that our method demonstrates strong multi-sensor generalization, significantly improving sensor-agnostic representation learning.
Lay Summary
Robots need a sense of touch to manipulate objects, but different touch sensors produce distinct signals—similar to how different cameras capture images. This diversity makes it difficult to learn a universal understanding of touch that works for any sensor. Existing methods that reconstruct sensor readings or are designed for specific tasks often fail to capture what is common across sensors, especially when sensor designs vary a lot, leading to poor generalization to new sensors. We propose a framework that learns sensor-agnostic touch representations. It uses an adaptation module to remove sensor-specific differences and follows a two-stage process. First, we use aligned synthetic data from multiple sensors to learn shared patterns without manual labels. Second, we align these patterns with real-world data that combines touch and vision, closing the gap between simulation and reality and adding semantic detail. Experiments show our method generalizes well to unseen sensors, significantly improving universal touch perception.