ACCL: Architecting Highly Scalable Distributed Training Systems With Highly Efficient Collective Communication Library
作者:Jianbo Dong, Shaochuang Wang, Fei Feng, Zheng Cao, Heng Pan, Lingbo Tang, Pengcheng Li, Hao Li, Qianyuan Ran, Yiqun Guo, Shanyuan Gao, Xin Long, Jie Zhang, Yong Li, Zhisheng Xia, Liuyihan Song, Yingya Zhang, Pan Pan, Guohui Wang, Xiaowei Jiang · 发表于:IEEE Micro · 年份:2021 · DOI:10.1109/mm.2021.3091475 · 被引用次数:34 · 研究领域:Advanced Memory and Neural Computing、Ferroelectric and Negative Capacitance Devices、Parallel Computing and Optimization Techniques
Distributed systems have been widely adopted for deep neural networks model training. However, the scalability of distributed training systems is largely bounded by the communication cost. We design a highly efficient collective communication library, namely Alibaba Collective Communication Library (ACCL), to build distributed training systems with linear scalability. ACCL provides optimized algorithms to fully make use of heterogeneous interconnects simultaneously. And the experimental results show significant performance improvement.