A Tight Convergence Analysis for Stochastic Gradient Descent with Delayed Updates

Arjevani, Yossi; Shamir, Ohad; Srebro, Nathan

Citation Details

We provide tight finite-time convergence bounds for gradient descent and stochastic gradient descent on quadratic functions, when the gradients are delayed and reflect iterates from τ rounds ago. First, we show that without stochastic noise, delays strongly affect the attainable optimization error: In fact, the error can be as bad as non-delayed gradient descent ran on only 1/τ of the gradients. In sharp contrast, we quantify how stochastic noise makes the effect of delays negligible, improving on previous work which only showed this phenomenon asymptotically or for much smaller delays. Also, in the context of distributed optimization, the results indicate that the performance of gradient descent with delays is competitive with synchronous approaches such as mini-batching. Our results are based on a novel technique for analyzing convergence of optimization algorithms using generating functions. more »

Award ID(s):: 1718970

PAR ID:: 10085929

Author(s) / Creator(s):: Arjevani, Yossi; Shamir, Ohad; Srebro, Nathan

Date Published:: 2018-06-26

Journal Name:: arXiv.org

ISSN:: 2331-8422

Page Range / eLocation ID:: arXiv:1806.10188

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Journal Article:
The DOI is not currently available.

More Like this