Comparative evaluation of large language models for hip fracture-related patient questions: DeepSeek-V3-FW, Gemini 2.0 Flash, and ChatGPT-4.5
作者:Yejin Zhang, Tao Huang, Chaoran Liu, Anna N Miller, Mingliang Yang, Ian A Harris, Takeshi Sawaguchi, Theodore Miclau, Maoyi Tian, Chun Sing Chui, Ning Zhang, Wing Hoi Cheung, Ronald Man Yeung Wong · 发表于:Digital Health · 年份:2026 · DOI:10.1177/20552076251412989 · 被引用次数:5 · 研究领域:Artificial Intelligence in Healthcare and Education、Hip and Femur Fractures、Topic Modeling
Background: Large language models (LLMs) are increasingly used in healthcare for patient education and clinical decision support. However, systematic benchmarking in real-world clinical contexts remains limited, particularly for high-risk conditions such as hip fractures. Objective: To evaluate and compare the performance of three state-of-the-art LLMs-DeepSeek-V3-FW, Gemini 2.0 Flash, and ChatGPT-4.5-in answering standardized patient questions on hip fracture management. Methods: Thirty standardized questions covering general knowledge, diagnosis, treatment, and rehabilitation were developed by three specialists in orthopedics and traumatology. Each LLM generated responses independently. Three experienced orthopedic surgeons assessed accuracy (4-point scale) and comprehensiveness (5-point scale). Statistical analyses included Kruskal-Wallis and chi-squared tests. Results: All models demonstrated high reliability, with 96.7% of responses rated "Good" or "Excellent" and none rated "Poor." Mean accuracy scores were comparable across models, and comprehensiveness averaged 4.8/5. DeepSeek-V3-FW tended to provide longer, structured answers and performed best in general knowledge, while Gemini 2.0 Flash excelled in diagnosis and rehabilitation and produced the most concise responses. ChatGPT-4.5 offered shorter, conversational answers with similar accuracy and detail. Conclusions: The three LLMs showed strong capabilities in delivering accurate and comprehensive information on hip ...