NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Provable Multi-Task Representation Learning by Two-Layer ReLU Neural Networks

Collins, L; Hassani, H; Soltanolkotabi, M; Mokhtari, A; Shakkottai, S (June 2024, National Library of Medicine catalog)

An increasingly popular machine learning paradigm is to pretrain a neural network (NN) on many tasks offline, then adapt it to downstream tasks, often by re-training only the last linear layer of the network. This approach yields strong downstream performance in a variety of contexts, demonstrating that multitask pretraining leads to effective feature learning. Although several recent theoretical studies have shown that shallow NNs learn meaningful features when either (i) they are trained on a single task or (ii) they are linear, very little is known about the closer-to-practice case of nonlinear NNs trained on multiple tasks. In this work, we present the first results proving that feature learning occurs during training with a nonlinear model on multiple tasks. Our key insight is that multi-task pretraining induces a pseudo-contrastive loss that favors representations that align points that typically have the same label across tasks. Using this observation, we show that when the tasks are binary classification tasks with labels depending on the projection of the data onto an 𝑟-dimensional subspace within the 𝑑 ≫𝑟-dimensional input space, a simple gradient-based multitask learning algorithm on a two-layer ReLU NN recovers this projection, allowing for generalization to downstream tasks with sample and neuron complexity independent of 𝑑. In contrast, we show that with high probability over the draw of a single task, training on this single task cannot guarantee to learn all 𝑟 ground-truth features.
more » « less
Full Text Available
On the Role of Attention in Prompt-tuning

Oymak, S.; Rawat, A. S.; Soltanolkotabi, M.; & Thrampoulidis, C. (January 2023, International Conference on Machine Learning)

Full Text Available
Outlier-Robust Sparse Estimation via Non-Convex Optimization

Cheng, Y.; Diakonikolas, I.; Ge, R.; Gupta, S.; Kane, D.; Soltanolkotabi, M. (December 2022, Advances in Neural Information Processing Systems (NeurIPS))

Full Text Available
Understanding Over-parameterization in Generative Adversarial Networks

Balaji, Y; Sajedi, M; Kalibhat, N; Ding, M; Stöger, D; Soltanolkotabi, M; Feizi, S. (January 2021, International Conference on Learning Representations (ICLR))

Full Text Available
Lagrange Coded Computing: Optimal Design for Resiliency, Security and Privacy

Yu, Q; Li, S; Raviv, N; Mousavi_Kalan, M; Soltanolkotabi, M; Avestimehr, S (January 2019, Proceedings of Machine Learning Research)

Full Text Available

Search for: All records