Evaluating Generative AI Large Language Models for Urticaria Management: A Comparative Analysis of DeepSeek‐R1 and ChatGPT‐4o
作者:Mengyao Yang, Jingchen Liang, Luyue Zhang, Hongshan Liu, Ying Chen, Yawen Wang, Cancan Qi, Han Ma, Ziyun Gao, Xinyue Zhang, Xinwu Niu, Xiaopeng Wang, Jianwen Ren, Jingyi Yuan, Weihui Zeng, Zhao Wang · 发表于:Clinical and Translational Allergy · 年份:2025 · DOI:10.1002/clt2.70113 · 被引用次数:2 · 研究领域:Urticaria and Related Conditions、Drug-Induced Adverse Reactions、Dermatology and Skin Diseases
INTRODUCTION: Urticaria is a prevalent condition affecting a significant portion of the global population. Both dermatologists and patients require access to up-to-date and accurate information. Traditional search engines often fall short in meeting these needs. Despite the growing reliance on AI for medical inquiries, the accuracy and quality of AI-generated remain understudied. This study aims to evaluate and compare the performance of two widely used AI models, ChatGPT-4o and DeepSeek-R1, in addressing urticaria-related queries. METHODS: An e-Delphi procedure was employed to generate and refine a set of urticaria-related questions, as well as to develop an evaluation framework for AI-generated responses. ChatGPT-4o and DeepSeek-R1 were then prompted with the finalized questions, and their responses were recorded. A single-blind comparative assessment was conducted among 67 participants (29 dermatologists and 38 non-dermatologists). The responses from both AI models were assessed across simplicity, accuracy, professionalism, clinical feasibility, comprehensibility, and completeness. RESULTS: DeepSeek-R1 outperformed ChatGPT-4o in most metrics. Dermatologists rated DeepSeek significantly higher in simplicity (p < 0.001), accuracy (p < 0.001), completeness (p = 0.001), professionalism (p < 0.001), and clinical feasibility (p < 0.001). Non-dermatologists found DeepSeek's responses more concise (p < 0.001) and comprehensible (p < 0.001). Both models showed comparable integratio...