Development and validation of a machine learning model for cardiovascular disease risk prediction in type 2 diabetes patients
作者:Chunming Xu, Fachao Shi, Wenlong Ding, Cunming Fang, Caoyang Fang, Caoyang Fang, Caoyang Fang · 发表于:Scientific Reports · 年份:2025 · DOI:10.1038/s41598-025-18443-7 · 被引用次数:16 · 研究领域:Diabetes Management and Research、Artificial Intelligence in Healthcare、Diabetes, Cardiovascular Risks, and Lipoproteins
Patients with type 2 diabetes mellitus (T2DM) have a significantly higher risk of cardiovascular disease (CVD) compared to the general population. Accurately predicting this risk is crucial for developing personalized treatment plans and public health interventions. This study aims to develop and validate a model for predicting CVD risk in T2DM patients using the Boruta feature selection algorithm and machine learning methods. We analyzed data from the National Health and Nutrition Examination Survey (NHANES) from 1999 to 2018. Six machine learning (ML) models, including Multilayer Perceptron (MLP), Light Gradient Boosting Machine (LightGBM), Decision Tree (DT), Extreme Gradient Boosting (XGBoost), Logistic Regression (LR), and k-Nearest Neighbors (KNN), were employed for model development and validation. Boruta was used for optimal feature selection. The performance of the machine learning models was comprehensively evaluated using ROC curves, accuracy, and other related metrics. Shapley Additive Explanation (SHAP) analysis was conducted for visual interpretation, and the Shinyapps.io platform was utilized to deploy the best-performing models as web-based applications. A total of 4,015 T2DM patients were included, among which 999 (24.9%) had CVD. Model evaluation revealed significant overfitting with the KNN algorithm, which showed perfect discrimination in the training set but performed poorly in the test set (AUC = 0.64). In contrast, XGBoost demonstrated more consistent p...