Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis
作者:Shahid Akhtar Akhund, Sadia Qazi, Muhammad Atif Mazhar, Aftab Ahmed Shaikh, Eshal Atif, Adeen Shaikh, Hassan Shaibah, Mohd Khirulnizam Musa, Sabiya N.Qazi, Mateen A. Khan, Akef Obeidat · 发表于:Medical Education Online · 年份:2026 · DOI:10.1080/10872981.2026.2684837 · 研究领域:Artificial Intelligence in Healthcare and Education、Clinical Reasoning and Diagnostic Skills、Innovations in Medical Education
Background Medical education faces increasing demand for scalable grading solutions. Large language models have been proposed as automated graders for high-stakes assessments, but evidence for their reliability in multimodal practical examinations remains limited.Methods This retrospective inter-rater reliability study compared three LLMs: ChatGPT-4o, Gemini 2.5 Flash, and Claude 3.5 Haiku with original human reference scores on 19 integrated anatomy, histology, and physiology objective structured practical examination items completed by 309 pre-medical students. Human reference scores were assigned during live summative grading by two content experts who divided items, with each item marked by one expert across all students; no human inter-rater reliability estimate was available. Items required simultaneous image interpretation and short-answer responses. Models were evaluated using zero-shot prompting reflecting deployment-realistic conditions. Agreement was assessed using Spearman's ρ, Cohen's κ, intraclass correlation coefficients, and Bland-Altman analysis. Item-level gap analysis compared student success rates with LLM performance across items.Results Rank-order correlations were strong across all models (ρ = 0.784–0.921), but categorical agreement diverged substantially. Pass/fail agreement ranged from 52.1% (ChatGPT; κ = 0.067, slight) to 84.1% (Claude; κ = 0.680, substantial). Bland-Altman analysis showed inconsistent systematic bias: Claude over-scored by +2.5 poin...