Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Evaluating Artificial Intelligence Chatbots in Oral and Maxillofacial Surgery Board Exams: Performance and Potential.

作者:R. Mahmoud, A. Shuster, S. Kleinman, S. Arbel, C. Ianculovici, O. Peleg · 发表于:Journal of oral and maxillofacial surgery · 年份:2024 · DOI:10.1016/j.joms.2024.11.007 · 被引用次数:26 · 研究领域:Medicine

BACKGROUND While artificial intelligence has significantly impacted medicine, the application of large language models (LLMs) in oral and maxillofacial surgery (OMS) remains underexplored. PURPOSE This study aimed to measure and compare the accuracy of 4 leading LLMs on OMS board examination questions and to identify specific areas for improvement. STUDY DESIGN, SETTING, AND SAMPLE An in-silico cross-sectional study was conducted to evaluate 4 artificial intelligence chatbots on 714 OMS board examination questions. PREDICTOR VARIABLE The predictor variable was the LLM used - LLM 1 (Generative Pretrained Transformer 4o [GPT-4o], OpenAI, San Francisco, CA), LLM 2 (Generative Pretrained Transformer 3.5 [GPT-3.5], OpenAI, San Francisco, CA), LLM 3 (Gemini, Google, Mountain View, CA), and LLM 4 (Copilot, Microsoft, Redmond, WA). MAIN OUTCOME VARIABLES The primary outcome variable was accuracy, defined as the percentage of correct answers provided by each LLM. Secondary outcomes included the LLMs' ability to correct errors on subsequent attempts and their performance across 11 specific OMS subject domains: Medicine and Anesthesia, Dentoalveolar and Implant Surgery, Maxillofacial Trauma, Maxillofacial Infections, Maxillofacial Pathology, Salivary Glands, Oncology, Maxillofacial Reconstruction, Temporomandibular Joint Anatomy and Pathology, Craniofacial and Clefts, and Orthognathic Surgery. COVARIATES No additional covariates were considered. ANALYSES Statistical analysis...