UB-Mesh: A Hierarchically Localized nD-FullMesh Data Center Network Architecture
作者:Heng Liao, Bingyang Liu, Xianping Chen, Zhigang Guo, Chuanning Cheng, Jianbing Wang, Xiangyu Chen, Peng Dong, Rui Meng, Wenjie Liu, Zhe Zhou, Ziyang Zhang, Yan‐Zhi Gai, Cunle Qian, Yi Xiong, Zhongjun Cheng, Jing Xia, Yongsheng Ma, Xi Chen, Wenhua Du, Shizhong Xiao, C. C. Li, Yong Qin, Li Xiong, Yu Zhou, Chen Lv, Lei Chen, Buyun Wang, Pei-Chen Wu, Junen Gao, Xingyue Li, Jian He, Shizhuan Yan, Bill McColl · 发表于:IEEE Micro · 年份:2025 · DOI:10.1109/mm.2025.3592688 · 被引用次数:14 · 研究领域:Cloud Computing and Resource Management、Graph Theory and Algorithms、Distributed and Parallel Computing Systems
The scaling of Large-scale Language Models (LLMs) demands unprecedented computational power and bandwidth. We present UB-Mesh, an innovative AI datacenter network architecture that enhances scalability, performance, and cost-efficiency through a hierarchical nD-FullMesh topology. Unlike traditional symmetrical designs, UB-Meshoptimizes LLM training by prioritizing localized data movement and minimizing switch usage. The architecture features UB-Mesh-Pod, a physical implementation of 4D-FullMesh using custom hardware including NPUs, CPUs, Low/High-Radix Switches (LRS/HRS), and NICs, interconnected via our Unified Bus (UB) technology for dynamic resource allocation. For network optimization, we introduce All-Path-Routing (APR) to efficiently manage data traffic. Combined with topology-aware performance tuning and robust reliability mechanisms like 64 + 1 backup, UB-Meshachieves 2.04× better cost-efficiency and 7.2% higher availability than Clos networks. These innovations address the critical challenges of building practical, high-performance AI infrastructure at scale.