Evaluation of the Quality and Reliability of ChatGPT‐4's Responses on Allergen Immunotherapy Using Validated Instruments for Health Information Quality Assessment
作者:Iván Chérrez-Ojeda, Torsten Zuberbier, Gabriela Rodas‐Valero, Jorge Sanchez, Michael Rudenko, Stephanie Dramburg, Pascal Demoly, Davide Caimmi, René Maximiliano Gómez, German D. Ramón, Ghada E. Fouda, Kim R. Quimby, Herberto José Chong‐Neto, Oscar Calderón Llosa, José Ignacio Larco, Olga Patricia Monge-Ortega, Marco Faytong‐Haro, Oliver Pfaar, Jean Bousquet, Karla Robles‐Velasco · 发表于:Clinical and Translational Allergy · 年份:2025 · DOI:10.1002/clt2.70130 · 被引用次数:4 · 研究领域:Artificial Intelligence in Healthcare and Education、Misinformation and Its Impacts、Immune responses and vaccinations
BACKGROUND: Chat Generative Pre-Trained Transformer 4 (ChatGPT-4) represents an advancing large language model (LLM) with potential applications in medical education and patient care. While Allergen Immunotherapy (AIT) can change the course of allergic diseases, it can also bring uncertainty to patients, who turn to readily available resources such as ChatGPT-4 to address these doubts. This study aimed to use validated tools to evaluate the information provided by ChatGPT-4 regarding AIT in terms of quality, reliability, and readability. METHODS: In accordance with EAACI clinical guidelines about AIT, 24 questions were selected and introduced in ChatGPT-4. Independent reviewers evaluated ChatGPT-4 responses using three validated tools: the DISCERN instrument (quality), JAMA Benchmark criteria (reliability), and Flesch-Kincaid Readability Tests (readability). Descriptive statistics summarized findings across categories. RESULTS: ChatGPT-4 responses were generally rated as "fair quality" on DISCERN, with strengths in classification/formulations and special populations. Notably, the tool provided good-quality responses on the preventive effects of AIT in children and premedication to reduce adverse reactions. However, JAMA Benchmark scores consistently indicated "insufficient information" (median = 0-1), primarily due to absent authorship, attribution, disclosure, and currency. Readability analyses revealed a college graduate-level requirement, with most responses classified as ...