Fourier Features Let Agents Learn High Precision Policies with Imitation Learning
Abstract
High-precision robotic manipulation requires fine-grained spatial reasoning that is often difficult to achieve with RGB-only policies due to depth ambiguity and perspective scale issues. Policies that leverage 3D information directly, such as those based on point clouds, offer a stronger geometric prior over purely image-based ones, yet their performance remains highly task-dependent. We hypothesize that this discrepancy may be due to the spectral bias of neural networks towards learning low frequency functions, which especially affects architectures conditioned on slow-moving Cartesian features. We thus propose to map point clouds from Cartesian space into high-dimensional Fourier space, effectively equipping the point cloud encoder with direct access to high-frequency features. We experimentally validate the use of Fourier features on challenging manipulation tasks from the RoboCasa and ManiSkill3 benchmarks and on a real robot setup. Despite their simplicity, we find that Fourier features provide significant benefits across diverse encoder architectures and benchmarks and are robust across hyperparameters. Our results indicate that Fourier features let policies leverage geometric details more effectively than Cartesian features, showing their potential as a general-purpose tool for point cloud-based imitation learning. We provide source code and videos on our project page: https://fourier-il.github.io/fourier-il.
Lay Summary
Although neural networks can learn anything in theory, in practice they are bad at learning high frequency functions, i.e. when the output needs to change very fast. High frequency functions are essential in robotic control, because a robot may have to change its movements sharply even if its environment has not changed much - think of a lining up a key with a lock. This problem has a well-known solution: Fourier features, a technique that transforms the inputs themselves into something rapidly changing, so the network does not need to. Even though this method is well-known, it is almost never used when training robots to see with point clouds. Point clouds are collections of points in 3D space, and even though they are better than camera images for understanding the 3D world, they force the network to learn a high frequency function. We apply Fourier features to robot learning on point clouds and find that it is not only effective, but much more robust than one might expect. In fact, they do not require much tuning and can be integrated into almost any neural network architecture. We argue that this technique should be used in essentially every neural network trained on point clouds.