A Transformer-Block-Wise Collaborative Training Mechanism with Hybrid Parallelism Over Heterogeneous Networks
作者:Jiewei Chen, Jingrong Wang, Shaoyong Guo, Jiakai Hao, Xuesong Qiu, Zehui Xiong · 年份:2025 · DOI:10.1109/wcnc61545.2025.10978332 · 被引用次数:1 · 研究领域:Advanced Graph Neural Networks、Brain Tumor Detection and Classification、Neural Networks and Applications
With the rise of AI-Generated Content (AIGC) services in wireless networks, efficient and high-quality distributed training of Large Language Models (LLMs) has become essential for enabling the large-scale application of next generation AI technologies. However, the extensive parameters of LLMs impose significant demands on memory, computing power and communication resources in heterogeneous networks. To efficiently utilize the dispersed network resources, this paper presents a First-Pipeline- Then-Federated Learning (FPTFL) approach with a hybrid parallel scheduling strategy to facilitate the training of Transformer-based LLMs. We propose a block-wise splitting mechanism to partition the Transformer's encoder into distinct segments, which are deployed cross individual devices. The encoder parameters and intermediate smashed data are uploaded to the edge server, where the whole model is updated through federated aggregation. Particularly, we develop a fine-grained computation-efficient method based on pipeline parallelism, enabling the segments to cooperatively train the entire encoder. An optimization problem is formulated to determine the LLM segments and the number of micro-batches under network resource constraints, with the goal of minimizing the total latency of LLM training services. Simulation results demonstrate that our approach enables Transformer-based model training on resource-constrained devices, preserves model performance, and reduces waiting time.