The increasing size of machine learning models, especially deep neural network models, can improve the model generalization capability. However, large models require more training data and more computing resources (such as GPU clusters) to train. In distributed training, the communication overhead of exchanging gradients or models among workers becomes a potential system bottleneck that limits the system scalability. Recently, many research works aim to reduce communication time of two types of distributed deep learning architectures, centralized and decentralized.
No takes yet. Share an insight, caveat, or question.
Tang et al. (2020) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: