Signal Strength Estimation in Logistic Regression Using Data Splitting
Abstract
Logistic regression is widely used in applications; however, when the dimension scales with the sample size, theory reveals that the asymptotic behavior of common M-estimators depends on nonzero bias and variance factors, which are functions of the signal strength. To leverage the theory to design valid statistical methodologies, it is essential to obtain accurate estimates of the signal strength. In this work, we utilize a data-splitting strategy to efficiently estimate the signal strength. To alleviate issues caused by separable data, we analyze the exact asymptotics of an M-estimator with a data-driven, non-decomposable regularizer that adapts to the true covariance structure. We justify the validity of our method through both theoretical analysis and numerical experiments.
Lay Summary
This paper studies logistic regression in high-dimensional settings, where the number of variables is comparable to the number of observations. In such regimes, standard methods can be biased or even fail to exist, and key theoretical results depend on an unknown quantity called signal strength, which measures how strongly features influence outcomes. However, estimating this signal strength is challenging because the true model parameters cannot be accurately recovered. The authors propose a simple and principled solution based on data splitting. They divide the data into two parts, fit the model separately, and use the difference and similarity between the two estimates to infer underlying signal and noise levels. By leveraging theoretical relationships, they then recover the signal strength. The method is provably consistent, works even when standard estimators fail, and does not require knowing the data covariance structure. This provides a practical and reliable tool for inference in modern high-dimensional problems.