WorldComp2D: Spatio-semantic Representations of Object Identity and Location from Local Views
Abstract
Lay Summary
For AI to interact with the world, it must understand both "what" an object is and "where" it is located. Most current AI models try to learn this by processing massive amounts of data in a way that is computationally heavy, making them slow and difficult to run on basic devices like standard laptops or smartphones. We developed WorldComp2D, a new, lightweight framework that changes how AI organizes its internal knowledge. Instead of a messy black box of information, WorldComp2D explicitly organizes data like a map, grouping information based on an object's identity and its physical closeness to other things. We tested this by teaching the AI to find facial features, like the corners of eyes or the tip of a nose. Our approach allows the AI to perform complex spatial tasks using up to four times less memory and half the processing power of current state-of-the-art models. This proves that we can make AI much more efficient without sacrificing accuracy. By making these models lighter and faster, we can bring sophisticated spatial intelligence to everyday devices, from better mobile apps to smarter, more responsive robotics.