The past few years have witnessed growth in the computational requirements training deep convolutional neural networks. Current approaches parallelize onto multiple devices by applying a single parallelization strategy(e.g., data or model parallelism) to all layers in a network. Although easy to about, these approaches result in suboptimal runtime performance in-scale distributed training, since different layers in a network may different parallelization strategies. In this paper, we propose-wise parallelism that allows each layer in a network to use an individual strategy. We jointly optimize how each layer is parallelized by a graph search problem. Our evaluation shows that layer-wise outperforms state-of-the-art approaches by increasing training, reducing communication costs, achieving better scalability to GPUs, while maintaining original network accuracy.
No takes yet. Share an insight, caveat, or question.
Jia et al. (2018) studied this question.