Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Determining the optimal number of clusters by Enhanced Gap Statistic in K-mean algorithm

作者:Iliyas Karim Khan, Hanita Daud, Nooraini Binti Zainuddin, Rajalingam Sokkalingam, Muhammad Farooq, Muzammil Elahi Baig, Gohar Ayub, Mudasar Zafar · 发表于:Egyptian Informatics Journal · 年份:2024 · DOI:10.1016/j.eij.2024.100504 · 被引用次数:45 · 研究领域:Advanced Clustering Algorithms Research、Face and Expression Recognition、Bayesian Methods and Mixture Models

Unsupervised learning, particularly K-means clustering, seeks to partition data into clusters with distinct intra-class cohesion and inter-class disparity. However, the arbitrary selection of clusters in K-means introduces challenges, leading to trial and error in determining the Optimal Number of Clusters (ONC). To address this, various methodologies have been devised, among which the Gap Statistic is prominent. Gap Statistic reliance on expected values for reference data selection poses limitations, especially in scenarios involving diverse scale, noise, and overlapping data. To tackle these challenges, this study introduces Enhanced Gap Statistic (EGS), which standardizes reference data using an exponential distribution within the Gap Statistic framework, integrating an adjustment factor for a more dependable estimation of the ONC. Application of EGS to K-means clustering facilitates accurate ONC determination. For comparison purposes, EGS is benchmarked against traditional Gap Statistic and other established methods used for ONC selection in K-means, evaluating accuracy and efficiency across datasets with varying characteristics. The results demonstrate EGS superior accuracy and efficiency, affirming its effectiveness in diverse data environments.