NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Lazy Estimation of Variable Importance for Large Neural Networks

Yue Gao, Abby Stevens (January 2022, Proceedings of the 39th International Conference on Machine Learning)

As opaque predictive models increasingly impact many areas of modern life, interest in quantifying the importance of a given input variable for making a specific prediction has grown. Recently, there has been a proliferation of model-agnostic methods to measure variable importance (VI) that analyze the difference in predictive power between a full model trained on all variables and a reduced model that excludes the variable(s) of interest. A bottleneck common to these methods is the estimation of the reduced model for each variable (or subset of variables), which is an expensive process that often does not come with theoretical guarantees. In this work, we propose a fast and flexible method for approximating the reduced model with important inferential guarantees. We replace the need for fully retraining a wide neural network by a linearization initialized at the full model parameters. By adding a ridge-like penalty to make the problem convex, we prove that when the ridge penalty parameter is sufficiently large, our method estimates the variable importance measure with an error rate of O(1/n) where n is the number of training samples. We also show that our estimator is asymptotically normal, enabling us to provide confidence bounds for the VI estimates. We demonstrate through simulations that our method is fast and accurate under several data-generating regimes, and we demonstrate its real-world applicability on a seasonal climate forecasting example.
more » « less
Full Text Available
ISLET: Fast and Optimal Low-Rank Tensor Regression via Importance Sketching

https://doi.org/10.1137

Zhang, A. R; Luo, Y; Raskutti, G; Yuan, M. (June 2020, SIAM journal on mathematics of data science)

In this paper, we develop a novel procedure for low-rank tensor regression, namely Importance Sketching Low-rank Estimation for Tensors (ISLET). The central idea behind ISLET is importance sketching, i.e., carefully designed sketches based on both the responses and low-dimensional structure of the parameter of interest. We show that the proposed method is sharply minimax optimal in terms of the mean-squared error under low-rank Tucker assumptions and under the randomized Gaussian ensemble design. In addition, if a tensor is low-rank with group sparsity, our procedure also achieves minimax optimality. Further, we show through numerical study that ISLET achieves comparable or better mean-squared error performance to existing state-of-the-art methods while having substantial storage and run-time advantages including capabilities for parallel and distributed computing. In particular, our procedure performs reliable estimation with tensors of dimension $p = O(10^8)$ and is 1 or 2 orders of magnitude faster than baseline methods.
more » « less
Full Text Available
Stochastic Gradient Descent in Correlated Settings: A Study on Gaussian Processes

Chen, H; Zheng, L; al Kontar, R; Raskutti, G. (January 2020, Advances in neural information processing systems)
null (Ed.)
Stochastic gradient descent (SGD) and its variants have established themselves as the go-to algorithms for large-scale machine learning problems with independent samples due to their generalization performance and intrinsic computational advantage. However, the fact that the stochastic gradient is a biased estimator of the full gradient with correlated samples has led to the lack of theoretical understanding of how SGD behaves under correlated settings and hindered its use in such cases. In this paper, we focus on the Gaussian process (GP) and take a step forward towards breaking the barrier by proving minibatch SGD converges to a critical point of the full loss function, and recovers model hyperparameters with rate O(1/K) up to a statistical error term depending on the minibatch size. Numerical studies on both simulated and real datasets demonstrate that minibatch SGD has better generalization over state-of-the-art GP methods while reducing the computational burden and opening a new, previously unexplored, data size regime for GPs.
more » « less
Full Text Available
The bias of isotonic regression

https://doi.org/10.1214/20-EJS1677

Dai, R.; Song, H.; Barber, R. F.; Raskutti, G. (January 2020, Electronic journal of statistics)

We study the bias of the isotonic regression estimator. While there is extensive work characterizing the mean squared error of the iso- tonic regression estimator, relatively little is known about the bias. In this paper, we provide a sharp characterization, proving that the bias scales as O(n−β/3) up to log factors, where 1 ≤ β ≤ 2 is the exponent correspond- ing to H ̈older smoothness of the underlying mean. Importantly, this result only requires a strictly monotone mean and that the noise distribution has subexponential tails, without relying on symmetric noise or other restrictive assumptions.
more » « less
Full Text Available
PUlasso: High-Dimensional Variable Selection With Presence-Only Data

https://doi.org/10.1080/01621459.2018.1546587

Song, Hyebin; Raskutti, Garvesh (December 2018, Journal of the American Statistical Association)

Full Text Available

Search for: All records