Abstract
This paper proposes DROPSIGNSGD, a communication-efficient and network-fault tolerant algorithm for training deep neural networks in a distributed and synchronous fashion. In DROPSIGNSGD, all numerical elements communicated between machines are either 1 or −1, represented by only one bit. More importantly, DROPSIGNSGD does not decline the benchmark accuracy on the ImageNet dataset when compared with the traditional distributed stochastic gradient descent algorithm, owing to a little trick in memorizing unused gradients. Experimental results are supported by a mathematical proof showing that DROPSIGNSGD converges under standard assumptions.
Author supplied keywords
Cite
CITATION STYLE
Phong, L. T., & Phuong, T. T. (2020). Distributed signsGD with improved accuracy and network-fault tolerance. IEEE Access, 8, 191839–191849. https://doi.org/10.1109/ACCESS.2020.3032637
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.