How High is ‘High’? Rethinking the Roles of Dimensionality in Topological Data Analysis and Manifold Learning
Abstract
Lay Summary
In machine learning, working with data that has many features (known as high-dimensional data) is often seen as an obstacle where the hidden structure in the data either becomes obscured by the high dimensions, or is assumed to only be recoverable when the number of features vastly exceeds the number of samples. The former issue arises when additional features are not informative, i.e. they are mostly noise. This paper re-thinks the latter assumption, in the situation where each dimension does include some information about the hidden structure. We introduce a mathematical framework proving that if data contains enough meaningful variability, we can accurately recover its hidden structure without requiring the number of features to outnumber the number of samples. We applied our theory to groundbreaking neuroscience data tracking brain activity in rats, providing evidence that the brain activity produces a faithful representation of 2D physical space.