Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism

作者:Zizhao Mo, Jianxiong Liao, Huanle Xu, Zhi Zhou, Chengzhong Xu · 年份:2025 · DOI:10.1145/3712285.3759784 · 被引用次数:2 · 研究领域:Parallel Computing and Optimization Techniques、Advanced Data Storage Technologies、Cloud Computing and Resource Management

The significant resource demands in LLM serving prompts production clusters to fully utilize heterogeneous hardware by partitioning LLM models across a mix of high-end and low-end GPUs. However, existing parallelization approaches often struggle to scale efficiently in heterogeneous environments due to their coarse-grained and static parallelization strategies.