Scholay

学术搜索 · AI 审稿 · LaTeX 协作

A clinician-based comparative study of large language models in answering medical questions: the case of asthma

作者:Yong Yin, Mei Zeng, Hansong Wang, Haibo Yang, Chuxiong Zhou, Fen Jiang, Shang‐Jung Wu, Tingyue Huang, Shuahua Yuan, Jilei Lin, Mingyu Tang, Jiande Chen, Bin Dong, Jiajun Yuan, Dan Xie · 发表于:Frontiers in Pediatrics · 年份:2025 · DOI:10.3389/fped.2025.1461026 · 被引用次数:10 · 研究领域:Artificial Intelligence in Healthcare and Education、Genomics and Rare Diseases、Machine Learning in Healthcare

Objective: This study aims to evaluate and compare the performance of four major large language models (GPT-3.5, GPT-4.0, YouChat, and Perplexity) in answering 32 common asthma-related questions. Materials and methods: Seventy-five clinicians from various tertiary hospitals participated in this study. Each clinician was tasked with evaluating the responses generated by the four large language models (LLMs) to 32 common clinical questions related to pediatric asthma. Based on predefined criteria, participants subjectively assessed the accuracy, correctness, completeness, and practicality of the LLMs' answers. The participants provided precise scores to determine the performance of each language model in answering pediatric asthma-related questions. Results: GPT-4.0 performed the best across all dimensions, while YouChat performed the worst in all dimensions. Both GPT-3.5 and GPT-4.0 outperformed the other two models, but there was no significant difference in performance between GPT-3.5 and GPT-4.0 or between YouChat and Perplexity. Conclusion: GPT and other large language models can answer medical questions with a certain degree of completeness and accuracy. However, clinical physicians should critically assess internet information, distinguishing between true and false data, and should not blindly accept the outputs of these models. With advancements in key technologies, LLMs may one day become a safe option for doctors seeking information.