Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
Gradient descent is one of the most widely used iterative algorithms in modern statistical learning. However, its precise algorithmic dynamics in high-dimensional settings remain only partially understood, which has limited its broader potential for statistical inference applications. This paper provides a precise, nonasymptotic joint distributional characterization of gradient descent iterates and their debiased statistics in a broad class of empirical risk minimization problems, in the so-called mean-field regime where the sample size is proportional to the signal dimension. Our nonasymptotic state evolution theory holds for both general nonconvex loss functions and non-Gaussian data, and reveals the central role of two Onsager correction matrices that precisely characterize the nontrivial dependence among all gradient descent iterates in the mean-field regime. Leveraging the joint state evolution characterization, we show that the gradient descent iterate retrieves approximate normality after a debiasing correction via a linear combination of observable loss derivative directions from all past iterates. Crucially, the debiasing coefficients are directly linked to the Onsager correction matrices, which can be estimated in a fully data-driven manner via the proposed gradient descent inference algorithm. This leads to a new algorithmic statistical inference framework based on debiased gradient descent, which (i) applies to a broad class of models with both convex and nonconvex losses, (ii) remains valid at each iteration without requiring algorithmic convergence and (iii) exhibits a certain robustness to possible model misspecification. As a by-product, our framework also provides algorithmic estimates of the generalization error at each iteration. We demonstrate our theory and inference methods in the canonical single-index regression model and a generalized logistic regression model, where the natural loss functions may exhibit arbitrarily nonconvex landscapes. Our analysis further shows that, in linear regression with squared loss, the proposed debiased gradient descent iterate eventually coincides with the debiased convex regularized estimator in a mean-field distributional sense, and the quality of statistical inference for the unknown signal aligns exactly with the generalization error achieved along the algorithmic trajectory.more » « lessFree, publicly-accessible full text available June 1, 2027
-
Abstract Understanding the stochastic behavior of random projections of geometric sets constitutes a fundamental problem in high dimension probability that finds wide applications in diverse fields. This paper provides a kinematic description for the behavior of Gaussian random projections of closed convex cones, in analogy to that of randomly rotated cones studied in Amelunxen et al. (2014, Inf. Inference J. IMA, 3, 224–294). Formally, let $$K$$ be a closed convex cone in $$\mathbb{R}^{n}$$, and $$G\in \mathbb{R}^{m\times n}$$ be a Gaussian matrix with i.i.d. $$\mathscr{N}(0,1)$$ entries. We show that $$GK\equiv \{G\mu : \mu \in K\}$$ behaves like a randomly rotated cone in $$\mathbb{R}^{m}$$ with statistical dimension $$\min \{\delta (K),m\}$$, in the following kinematic sense: for any fixed closed convex cone $$L$$ in $$\mathbb{R}^{m}$$,$$$ \begin{align*} &\delta(L)+\delta(K)\ll m\, \Rightarrow\, L\cap GK = \{0\} \hbox{ with high probability},\nonumber\\ &\delta(L)+\delta(K)\gg m\, \Rightarrow\, L\cap GK \neq \{0\} \hbox{ with high probability}. \end{align*} $$ Similar kinematic descriptions are obtained for Gaussian random pre-images, and certain Gaussian random projections of general closed convex sets. The practical utility and broad applicability of the prescribed approximate kinematic formulae are demonstrated in a number of distinct problems arising from statistical learning, mathematical programming and asymptotic geometric analysis. In particular, we prove (i) new phase transitions of the existence of cone-constrained maximum likelihood estimators in logistic regression, (ii) new phase transitions of the cost optimum of deterministic conic programs with random constraints and (iii) a local version of the Gaussian Dvoretzky-Milman theorem that describes almost deterministic, low-dimensional behaviours of subspace sections of randomly projected convex sets. The proofs of our results exploit the full strength of comparison inequalities for Gaussian processes. Compared to the conic integral geometry method in Amelunxen et al. (2014, Inf. Inference J. IMA, 3, 224–294), our method has the advantage of circumventing the rigid requirement of exact kinematic formulae that are typically unavailable for random projections and general closed convex sets.more » « lessFree, publicly-accessible full text available February 19, 2027
-
The Ridgeless minimum $$\ell_2$$-norm interpolator in overparametrized linear regression has attracted considerable attention in recent years in both machine learning and statistics communities. While it seems to defy conventional wisdom that overfitting leads to poor prediction, recent theoretical research on its $$\ell_2$$-type risks reveals that its norm minimizing property induces an `implicit regularization' that helps prediction in spite of interpolation. This paper takes a further step that aims at understanding its precise stochastic behavior as a statistical estimator. Specifically, we characterize the distribution of the Ridgeless interpolator in high dimensions, in terms of a Ridge estimator in an associated Gaussian sequence model with positive regularization, which provides a precise quantification of the prescribed implicit regularization in the most general distributional sense. Our distributional characterizations hold for general non-Gaussian random designs and extend uniformly to positively regularized Ridge estimators. As a direct application, we obtain a complete characterization for a general class of weighted $$\ell_q$$ risks of the Ridge(less) estimators that are previously only known for $q=2$ by random matrix methods. These weighted $$\ell_q$$ risks not only include the standard prediction and estimation errors, but also include the non-standard covariate shift settings. Our uniform characterizations further reveal a surprising feature of the commonly used generalized and $$k$$-fold cross-validation schemes: tuning the estimated $$\ell_2$$ prediction risk by these methods alone lead to simultaneous optimal $$\ell_2$$ in-sample, prediction and estimation risks, as well as the optimal length of debiased confidence intervals.more » « lessFree, publicly-accessible full text available January 1, 2027
-
Abstract Distance covariance is a popular dependence measure for two random vectors $$X$$ and $$Y$$ of possibly different dimensions and types. Recent years have witnessed concentrated efforts in the literature to understand the distributional properties of the sample distance covariance in a high-dimensional setting, with an exclusive emphasis on the null case that $$X$$ and $$Y$$ are independent. This paper derives the first non-null central limit theorem for the sample distance covariance, and the more general sample (Hilbert–Schmidt) kernel distance covariance in high dimensions, in the distributional class of $(X,Y)$ with a separable covariance structure. The new non-null central limit theorem yields an asymptotically exact first-order power formula for the widely used generalized kernel distance correlation test of independence between $$X$$ and $$Y$$. The power formula in particular unveils an interesting universality phenomenon: the power of the generalized kernel distance correlation test is completely determined by $$n\cdot \operatorname{dCor}^{2}(X,Y)/\sqrt{2}$$ in the high-dimensional limit, regardless of a wide range of choices of the kernels and bandwidth parameters. Furthermore, this separation rate is also shown to be optimal in a minimax sense. The key step in the proof of the non-null central limit theorem is a precise expansion of the mean and variance of the sample distance covariance in high dimensions, which shows, among other things, that the non-null Gaussian approximation of the sample distance covariance involves a rather subtle interplay between the dimension-to-sample ratio and the dependence between $$X$$ and $$Y$$.more » « less
An official website of the United States government
