Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Comparative performance evaluation of large language models in answering esophageal cancer-related questions: a multi-model assessment study

作者:Zhanwen He, Lilan Zhao, Genglin Li, J. Wang, Songyu Cai, Pengjie Tu, Jingbo Chen, Jia‐Jia Wu, Juan Zhang, Ruiqi Chen, Yangyun Huang, Xiaojie Pan, Wenshu Chen · 发表于:Frontiers in Digital Health · 年份:2025 · DOI:10.3389/fdgth.2025.1670510 · 被引用次数:6 · 研究领域:Artificial Intelligence in Healthcare and Education、Esophageal Cancer Research and Treatment、Radiomics and Machine Learning in Medical Imaging

Background: Esophageal cancer has high incidence and mortality rates, leading to increased public demand for accurate information. However, the reliability of online medical information is often questionable. This study systematically compared the accuracy, completeness, and comprehensibility of mainstream large language models (LLMs) in answering esophageal cancer-related questions. Methods: In total, 65 questions covering fundamental knowledge, preoperative preparation, surgical treatment, and postoperative management were selected. Each model, namely, ChatGPT 5, Claude Sonnet 4.0, DeepSeek-R1, Gemini 2.5 Pro, and Grok-4, was queried independently using standardized prompts. Five senior clinical experts, including three thoracic surgeons, one radiologist, and one medical oncologist, evaluated the responses using a five-point Likert scale. A retesting mechanism was applied for the low-scoring responses, and intraclass correlation coefficients were used to assess the rating consistency. The statistical analyses were conducted using the Friedman test, the Wilcoxon signed-rank test, and the Bonferroni correction. Results: All the models performed well, with average scores exceeding 4.0. However, the following significant differences emerged: Gemini excelled in accuracy, while ChatGPT led in completeness, particularly in surgical and postoperative contexts. Minor differences appeared in fundamental knowledge, but notable disparities were found in complex areas. Retesting showed ...