Is explainability the missing link in artificial intelligence-based diabetes prediction?
作者:F. Zaier, M. Zribi, H. Aounallah-Skhiri · 发表于:European Journal of Public Health · 年份:2025 · DOI:10.1093/eurpub/ckaf161.1294 · 被引用次数:1
Abstract Background With the rising use of artificial intelligence (AI) in healthcare, the need for model transparency and interpretability is increasingly emphasized. This study aimed to identify key risk factors for diabetes and compare the predictive performance of an interpretable logistic regression (LR) model with advanced machine learning (ML) algorithms using explainable AI (XAI) tools. Methods Data were obtained from the 2016 Tunisian Health Examination Survey. LR was used to assess diabetes risk factors through adjusted odds ratios (aOR) and 95% confidence intervals (CI). ML models included Decision Tree (DT), Gradient Boosting (GB), Artificial Neural Network (ANN), and Random Forest (RF). Model performance was assessed via accuracy, recall, F1-score, and area under the curve (AUC). For interpretability, we used Local Interpretable Model-agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP) to visualize feature importance. Results A total of 8,894 adults were included. The diabetes prevalence was 18.9% [17.1-21.2]. Significant predictors in LR included dyslipidemia (aOR=2.98 [2.88-3.30]), low socioeconomic status (aOR=1.27 [1.16-1.38]), urban residence (aOR=1.19 [1.02-1.35]), and higher BMI (aOR=1.19 [1.11-1.29]). LR achieved an accuracy of 82.8%, recall of 97.2%, F1-score of 90.1%, and AUC of 77.8%. Among ML models, DT performed the worst (AUC=59.6%). GB and RF outperformed other models in AUC (79.7% and 77.6%, respectively), accuracy (83.3% and 83....