A Survey on Machine Learning-Based HPC I/O Analysis and Optimization
作者:Peng Jj, Lihua Yang, Huijun Wu, Wenzhe Zhang, Zhenwei Wu, Wei Zhang, Jiaxin Li, Yiqin Dai, Yong Dong · 发表于:IEEE Transactions on Parallel and Distributed Systems · 年份:2025 · DOI:10.1109/tpds.2025.3639682 · 被引用次数:1 · 研究领域:Advanced Data Storage Technologies、Distributed and Parallel Computing Systems、Parallel Computing and Optimization Techniques
The soaring computing power of HPC systems supports numerous large-scale applications, which generate massive data volumes and diverse I/O patterns, leading to severe I/O bottlenecks. Analyzing and optimizing HPC I/O is therefore critical. However, traditional approaches are typically customized and lack the adaptability required to cope with dynamic changes in HPC environments. To address the challenge, Machine Learning (ML) has been increasingly adopted to automate and enhance I/O analysis and optimization. Given sufficient I/O traces from HPC systems, ML can learn underlying I/O behaviors, extract actionable insights, and dynamically adapt to evolving workloads to improve performance. In this survey, we propose a novel taxonomy that aligns HPC I/O problems with learning tasks to systematically review existing studies. Through this taxonomy, we synthesize key findings on research distribution, data preparation, and model selection. Finally, we discuss several directions to advance the effective integration of ML in HPC I/O systems.