The implementation of a vast majority of machine learning (ML) algorithms down to solving a numerical optimization problem. In this context, Gradient Descent (SGD) methods have long proven to provide good, both in terms of convergence and accuracy. Recently, several approaches have been proposed in order to scale SGD to solve large ML problems. At their core, most of these approaches are following a-reduce scheme. This paper presents a novel parallel updating algorithm for, which utilizes the asynchronous single-sided communication paradigm. to existing methods, Asynchronous Parallel Stochastic Gradient Descent(ASGD) provides faster (or at least equal) convergence, close to linear scaling stable accuracy.
No takes yet. Share an insight, caveat, or question.
Keuper et al. (2015) studied this question.