Performance of large language model in cross-specialty medical scenarios
作者:Zhen Cui, Wuzheng Liu, Xuan Tian, Conglei You, Xiangyu Meng, HuiJuan ZHANG, Kangzi Gong, Xu Wang, Jun Wu · 发表于:Journal of Translational Medicine · 年份:2025 · DOI:10.1186/s12967-025-07577-x · 被引用次数:2 · 研究领域:Artificial Intelligence in Healthcare and Education、Genomics and Rare Diseases、Machine Learning in Healthcare
BACKGROUND: Large language models (LLMs) demonstrate transformative potential in healthcare, yet their diagnostic and therapeutic accuracy across medical specialties remains inadequately characterized. METHODS: This study aimed to compare diagnostic and therapeutic capabilities of GPT-4o, GPT-3.5-Turbo, Claude-3-Sonnet across 12 medical specialties using standardized clinical vignettes. 50 PubMed-derived clinical cases between 2007 and 2024 were assessed. Two board-certified physicians independently evaluated LLMs outputs, with a senior clinician adjudicating discrepancies. All LLMs received identical text-based case descriptions with or without images, generating free-text diagnostic and therapeutic recommendations for blinded, randomized evaluation. RESULTS: Among the three evaluated LLMs, GPT-4o demonstrated superior diagnostic accuracy (median 10; IQR, 7.5–10), outperforming Claude-3-Sonnet (median 8; IQR, 2.8–10; P = .02) and GPT-3.5-Turbo (median 4; IQR, 1–9.3; P < .0001). A narrow IQR and minimal variation (SD = 2.9; range = 5.0) reflected high consistency in diagnostic outputs across diverse medical fields. For therapeutic recommendations, GPT-4o (median 10, IQR 0–10) outperformed GPT-3.5-Turbo (median 0, IQR 0–6.3; P = .0005) but showed no significant advantage over Claude-3-Sonnet (median 5, IQR 0–10; P = .45). CONCLUSION: This study demonstrates that advanced LLMs, particularly GPT-4o, have significant potential to support clinical diagnostics, showing high accurac...