Sharp Concentration Bounds for Bundle-Valued Statistics on Manifolds
Swagatam Das ⋅ Vaclav Snasel
Abstract
Many geometric statistics and manifold learning pipelines routinely produce observations---such as tangent vectors or local frames---whose natural home is a varying family of fibers attached to different points of a base manifold, rather than a single shared vector space. Forming empirical averages requires transporting these observations to a common reference fiber, introducing curvature- and holonomy-driven effects absent from classical concentration theory. We develop a non-asymptotic concentration theory for such transported empirical means, deriving finite-sample, dimension-free Hoeffding- and Bernstein-type bounds via sharp Hilbert-space inequalities. When shortest paths to the reference point are non-unique, transport becomes path-dependent and introduces a deterministic holonomy bias; we isolate and quantify this bias through bundle curvature and loop geometry, with sharp closed-form formulas for the tangent bundle of a round sphere. The resulting bias--variance decomposition separates the stochastic fluctuation decaying at the classical $n^{-1/2}$ rate in sample size $n$, from a curvature-driven error floor that no amount of additional data can eliminate; minimax lower bounds confirm both terms are unavoidable. We further establish a robust median-of-means estimator achieving optimal rates under heavy tails, and a central limit theorem in the reference fiber. Controlled experiments on the sphere validate all theoretical predictions.
Lay Summary
When machine learning models process data that lives on curved surfaces---such as wind directions on a globe, brain fiber orientations in medical scans, or features in graph neural networks---simply averaging that data the usual way can introduce hidden errors driven by the curvature of the underlying space, not by lack of data. This paper provides the first rigorous mathematical guarantees that precisely separate how much error comes from having too few samples (which shrinks as you collect more data) from how much is an irreducible geometric distortion that no amount of additional data can fix. These results give practitioners concrete, computable error bounds for a wide class of modern geometric machine learning pipelines.
Successful Page Load