Quality safety and disparity of an AI chatbot in managing chronic diseases: simulated patient experiments
作者:Yafei Si, Yurun Meng, Xi Chen, Ruopeng An, Limin Mao, Bingqin Li, Hazel Bateman, Han Zhang, Hongbin Fan, Jiaqi Zu, Shaoqing Gong, Zhongliang Zhou, Yudong Miao, Xiaojing Fan, Gang Chen · 发表于:npj Digital Medicine · 年份:2025 · DOI:10.1038/s41746-025-01956-w · 被引用次数:5 · 研究领域:Healthcare Systems and Reforms、Health Systems, Economic Evaluations, Quality of Life、Chronic Disease Management Strategies
The rapid development of AI solutions reveals opportunities to address the underdiagnosis and poor management of chronic conditions in developing settings. Using the method of simulated patients and experimental designs, we evaluate the quality, safety, and disparity of medical consultation with ERNIE Bot in China among 384 patient-AI trials. ERNIE Bot reached a diagnostic accuracy of 77.3%, correct drug prescriptions of 94.3%, but prescribed high rates of unnecessary medical tests (91.9%) and unnecessary medications (57.8%). Disparities were observed based on patient age and household economic status, with older and wealthier patients receiving more intensive care. Under standardized conditions, ERNIE Bot, ChatGPT, and DeepSeek demonstrated higher diagnostic accuracy but a greater tendency toward overprescription than human physicians. The results suggest the great potential of ERNIE Bot in empowering quality, accessibility, and affordability of healthcare provision in developing contexts, but also highlight critical risks related to safety and amplification of sociodemographic disparities.