Ophthalmological Question Answering and Reasoning Using OpenAI o1 vs Other Large Language Models
作者:Sahana Srinivasan, X. C. Ai, Minjie Zou, Ke Zou, Hyunjae Kim, Thaddaeus Wai Soon Lo, Krithi Pushpanathan, Gabriel Dawei Yang, Jocelyn Hui Lin Goh, Yiming Kong, Anran Li, Maxwell Singer, Kai Pu Jin, Fares Antaki, David Ziyou Chen, Dianbo Liu, Ron Afshari Adelman, Qingyu Chen, Yih Chung Tham · 发表于:JAMA Ophthalmology · 年份:2025 · DOI:10.1001/jamaophthalmol.2025.2413 · 被引用次数:19 · 研究领域:Artificial Intelligence in Healthcare and Education、Genomics and Rare Diseases、Multimodal Machine Learning Applications
Importance: OpenAI's recent large language model (LLM) o1 has dedicated reasoning capabilities, but it remains untested in specialized medical fields like ophthalmology. Evaluating o1 in ophthalmology is crucial to determine whether its general reasoning can meet specialized needs or if domain-specific LLMs are warranted. Objective: To assess the performance and reasoning ability of OpenAI's o1 compared with other LLMs on ophthalmological questions. Design, Setting, and Participants: In September through October 2024, the LLMs o1, GPT-4o (OpenAI), GPT-4 (OpenAI), GPT-3.5 (OpenAI), Llama 3-8B (Meta), and Gemini 1.5 Pro (Google) were evaluated on 6990 standardized ophthalmology questions from the Medical Multiple-Choice Question Answering (MedMCQA) dataset. The study did not analyze human participants. Main Outcomes and Measures: Models were evaluated on performance (accuracy and macro F1 score) and reasoning abilities (text-generation metrics: Recall-Oriented Understudy for Gisting Evaluation [ROUGE-L], BERTScore, BARTScore, AlignScore, and Metric for Evaluation of Translation With Explicit Ordering [METEOR]). Mean scores are reported for o1, while mean differences (Δ) from o1's scores are reported for other models. Expert qualitative evaluation of o1 and GPT-4o responses assessed usefulness, organization, and comprehensibility using 5-point Likert scales. Results: The LLM o1 achieved the highest accuracy (mean, 0.877; 95% CI, 0.870 to 0.885) and macro F1 score (mean, 0.877; 9...