HyLo: a hybrid low-rank natural gradient descent method

Mu, Baorun; Soori, Saeed; Can, Bugra; Gurbuzbalaban, Mert; Dehnavi, Maryam Mehri

Citation Details

This work presents a Hybrid Low-Rank Natural Gradient Descent method, called HyLo, that accelerates the training time of deep neural networks. Natural gradient descent (NGD) requires computing the inverse of the Fisher information matrix (FIM), which is typically expensive at large-scale. Kronecker factorization methods such as KFAC attempt to improve NGD's running time by approximating the FIM with Kronecker factors. However, the size of Kronecker factors increases quadratically as the model size grows. Instead, in HyLo, we use the Sherman-Morrison-Woodbury variant of NGD (SNGD) and propose a reformulation of SNGD to resolve its scalability issues. HyLo uses a computationally-efficient low-rank factorization to achieve superior timing for Fisher inverses. We evaluate HyLo on large models including ResNet-50, U-Net, and ResNet-32 on up to 64 GPUs. HyLo converges 1.4×-2.1× faster than the state-of-the-art distributed implementation of KFAC and reduces the computation and communication time up to 350× and 10.7× on ResNet-50. more »

Award ID(s):: 1814888 2053485

NSF-PAR ID:: 10399586

Author(s) / Creator(s):: Mu, Baorun; Soori, Saeed; Can, Bugra; Gurbuzbalaban, Mert; Dehnavi, Maryam Mehri

Date Published:: 2022-01-01

Journal Name:: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis

Volume:: 47

Page Range / eLocation ID:: 1-16

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this