Evaluating large language models in pediatric fever management: a two-layer study
作者:Guijun Yang, Hejun Jiang, Shuhua Yuan, Mingyu Tang, Jing Zhang, Jilei Lin, Jiande Chen, Jiajun Yuan, Liebin Zhao, Yonggao Yin · 发表于:Frontiers in Digital Health · 年份:2025 · DOI:10.3389/fdgth.2025.1610671 · 被引用次数:2 · 研究领域:Thermal Regulation in Medicine、Health Literacy and Information Accessibility、Pediatric Pain Management Techniques
Background Pediatric fever is a prevalent concern, often causing parental anxiety and frequent medical consultations. While large language models (LLMs) such as ChatGPT, Perplexity, and YouChat show promise in enhancing medical communication and education, their efficacy in addressing complex pediatric fever-related questions remains underexplored, particularly from the perspectives of medical professionals and patients’ relatives. Objective This study aimed to explore the differences and similarities among four common large language models (ChatGPT3.5, ChatGPT4.0, YouChat, and Perplexity) in answering thirty pediatric fever-related questions and to examine how doctors and pediatric patients’ relatives evaluate the LLM-generated answers based on predefined criteria. Methods The study selected thirty fever-related pediatric questions answered by the four models. Twenty doctors rated these responses across four dimensions. To conduct the survey among pediatric patients’ relatives, we eliminated certain responses that we deemed to pose safety risks or be misleading. Based on the doctors’ questionnaire, the thirty questions were divided into six groups, each evaluated by twenty pediatric relatives. The Tukey post-hoc test was used to check for significant differences. Some of pediatric relatives was revisited for deeper insights into the results. Results In the doctors’ questionnaire, ChatGPT3.5 and ChatGPT4.0 outperformed YouChat and Perplexity in all dimensions, with no signifi...