Automating Dependence-Aware Parallelization of Machine Learning Training on Distributed Shared Memory

Wei, Jinliang; Gibson, Garth A.; Gibbons, Phillip B.; Xing, Eric P.

doi:10.1145/3302424.3303954

Citation Details

Automating Dependence-Aware Parallelization of Machine Learning Training on Distributed Shared Memory

Machine learning (ML) training is commonly parallelized using data parallelism. A fundamental limitation of data parallelism is that conflicting (concurrent) parameter accesses during ML training usually diminishes or even negates the benefits provided by additional parallel compute resources. Although it is possible to avoid conflicting parameter accesses by carefully scheduling the computation, existing systems rely on programmer manual parallelization and it remains a question when such parallelization is possible. We present Orion, a system that automatically parallelizes serial imperative ML programs on distributed shared memory. The core of Orion is a static dependence analysis mechanism that determines when dependence-preserving parallelization is effective and maps a loop computation to an optimized distributed computation schedule. Our evaluation shows that for a number of ML applications, Orion can parallelize a serial program while preserving critical dependences and thus achieve a significantly faster convergence rate than data-parallel programs and a matching convergence rate and comparable computation throughput to state-of-the-art manual parallelizations including model-parallel programs. more »

Award ID(s):: 1725663

NSF-PAR ID:: 10136255

Author(s) / Creator(s):: Wei, Jinliang; Gibson, Garth A.; Gibbons, Phillip B.; Xing, Eric P.

Date Published:: 2019-03-01

Journal Name:: EuroSys '19: Proceedings of the Fourteenth EuroSys Conference

Volume:: 14

Page Range / eLocation ID:: 1 to 17

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
https://doi.org/10.1145/3302424.3303954

More Like this