Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU Clusters

作者:Wenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye, Chengzhong Xu · 年份:2025 · DOI:10.1145/3689031.3696074 · 被引用次数:10 · 研究领域:Advanced Neural Network Applications、Machine Learning and Data Classification、Domain Adaptation and Few-Shot Learning

Deep learning (DL) inference services are widely recognized as crucial workloads in large-scale cloud clusters. However, due to the stringent latency requirements, cloud providers often over-provision GPU resources, resulting in underutilization of the available GPU potential. Although co-locating tasks on the same device can enhance utilization, ensuring Service Level Objectives (SLOs) guarantees for multiplexing highly dynamic inference services becomes extremely challenging due to significant resource interference.