Joint Learning in the Gaussian Single Index Model
Abstract
Lay Summary
Many modern learning problems involve discovering a low-dimensional structure hidden inside high-dimensional data. In this work, we study a fundamental setting where the prediction only depends on an unknown one-dimensional direction together with an unknown nonlinear function. Learning both objects simultaneously leads to a highly non-convex optimization problem, which is generally difficult to analyze theoretically. We show that a natural learning procedure can nevertheless reliably recover both the hidden direction and the nonlinear relationship from Gaussian data. Our analysis provides precise guarantees on the speed of convergence and reveals that learning can succeed even when the initial estimate is initially pointing in the “wrong” direction. We also show how the method can be implemented efficiently in practice using kernel-based techniques. Overall, our work provides new theoretical and practical insights into how low-dimensional structure can be learned from high-dimensional data.