Scholay

学术搜索 · AI 审稿 · LaTeX 协作

AI-Driven Large Language Models in Health Consultations for HIV Patients

作者:Chunyan Zhao, Chang Song, Tong Yang, Ai-Chun Huang, Hang-Biao Qiang, Chun-Ming Gong, Jingsong Chen, Qingdong Zhu · 发表于:Journal of Multidisciplinary Healthcare · 年份:2025 · DOI:10.2147/jmdh.s533621 · 被引用次数:3 · 研究领域:Artificial Intelligence in Healthcare and Education、HIV/AIDS Research and Interventions、Mobile Health and mHealth Applications

Purpose: This study endeavors to conduct a comprehensive assessment on the performance of large language models (LLMs) in health consultation for individuals living with HIV, delve into their applicability across a diverse array of dimensions, and provide evidence-based support for clinical deployment. Patients and Methods: A 23-question multi-dimensional HIV-specific question bank was developed, covering fundamental knowledge, diagnosis, treatment, prognosis, and case analysis. Four advanced LLMs-ChatGPT-4o, Copilot, Gemini, and Claude-were tested using a multi-dimensional evaluation system assessing medical accuracy, comprehensiveness, understandability, reliability, and humanistic care (which encompasses elements such as individual needs attention, emotional support, and ethical considerations). A five-point Likert scale was employed, with three experts independently scoring. Statistical metrics (mean, standard deviation, standard error) were calculated, followed by consistency analysis, difference analysis, and post-hoc testing. Results: Claude obtained the most outstanding performance with regard to information comprehensiveness (mean score 4.333), understandability (mean score 3.797), and humanistic care (mean score 2.855); Copilot demonstrated proficiency in diagnostic questions (mean score 3.880); Gemini illustrated exceptional performance in case analysis (mean score 4.111). Based on the post-hoc analysis, Claude outperformed other models in thoroughness and humanist...