Large-scale machine learning training, in particular distributed stochastic descent, needs to be robust to inherent system variability such as straggling and random communication delays. This work considers a training framework where each worker node is allowed to perform model updates and the resulting models are averaged periodically. We the true speed of error convergence with respect to wall-clock time(instead of the number of iterations), and analyze how it is affected by the of averaging. The main contribution is the design of AdaComm, an communication strategy that starts with infrequent averaging to save delay and improve convergence speed, and then increases the frequency in order to achieve a low error floor. Rigorous on training deep neural networks show that AdaComm can take 3\× less time than fully synchronous SGD, and still reach the same final loss.
No takes yet. Share an insight, caveat, or question.
Wang et al. (2018) studied this question.