Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Performability Analysis for Large-Scale Multi-State Computing Systems: Methodologies, Advances, and Future Directions

作者:Yuchang Mo, Yuan Fan, Chunyu Miao, Mirlan Chynybaev, Faer Gui, Rengui Zhang, Jianyong Hu, Jinbin Mu, Akylbek Chymyrov · 发表于:ICCK Transactions on Systems Safety and Reliability · 年份:2025 · DOI:10.62762/tssr.2025.527003 · 被引用次数:3 · 研究领域:Reliability and Maintenance Optimization、Software Reliability and Analysis Research、Software System Performance and Reliability

Large-scale computing systems, such as cloud data centers, grid infrastructures, and high-performance computing clusters, are the backbone of modern information technology ecosystems. These systems typically consist of numerous heterogeneous, multi-state computing nodes that exhibit varying performance levels due to component failures, degradation, or dynamic resource allocation. Performability analysis, which integrates both system reliability and performance evaluations to quantify the probability of the system operating at a specified performance level, is critical for ensuring the efficient, reliable, and cost-effective operation of these complex systems. This paper presents a comprehensive review of recent advancements in performability analysis for large-scale multi-state computing systems over the past decade. It classifies existing research into three core methodological categories: binary decision diagram (BDD)-based approaches, multi-valued decision diagram (MDD)-based approaches, and comparative benchmarking with traditional methods (e.g., continuous-time Markov chains (CTMC), universal generating function (UGF)). For each category, the paper details key methodologies, algorithmic innovations, and practical applications. Additionally, the promising future directions are proposed to address emerging challenges, such as handling dynamic system behaviors, integrating real-time data, and optimizing resource allocation for performability. This review provides a valuable...