The Data Manifold under the Microscope
Abstract
Lay Summary
Modern AI systems can recognize images and make predictions impressively well, but our mathematical understanding of why they work still lags behind. Many explanations assume that real data has a simpler hidden shape, for example because images of the same object change smoothly as the object moves, rotates, or changes size. This idea is hard to test because clean mathematical examples are often too simple, while real image collections are too messy to measure precisely. We introduce a benchmark that creates controlled image collections where this hidden shape can be measured much more accurately. We build these collections by starting from simple object datasets and varying factors such as position, size, rotation, and viewpoint in a dense and organized way. This allows researchers to measure how large, curved, and folded the hidden shape is, and to check whether existing mathematical predictions match observed behavior. We use the benchmark to study two cases, how current error bounds scale with data size, and how a network changes data geometry inside its layers. Our work provides a practical testing ground for checking theories of deep learning and for developing better tools to measure the geometry of data.