Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Benchmarking GPT-5 in radiation oncology: measurable gains, but persistent need for expert oversight

作者:Ugur Dinç, Jibak Sarkar, Philipp Schubert, Sabine Semrau, Thomas Weißmann, Andre Karius, Johann Brand, Bernd‐Niklas Axer, Ahmed Mohamed Gomaa, Pluvio Stephan, Ishita Sheth, Sogand Beirami, Annette Schwarz, Udo S. Gaipl, Benjamin Frey, Christoph Bert, Stefanie Corradini, Rainer Fietkau, Florian Putz · 发表于:Frontiers in Oncology · 年份:2025 · DOI:10.3389/fonc.2025.1695468 · 被引用次数:3 · 研究领域:Radiomics and Machine Learning in Medical Imaging、Management of metastatic bone disease、Medical Imaging Techniques and Applications

Introduction Large language models (LLM) have shown great potential in clinical decision support and medical education. GPT-5 is a novel LLM system that has been specifically marketed towards oncology use. This study comprehensively benchmarks GPT-5 for the field of radiation oncology. Methods Performance was assessed using two complementary benchmarks: (i) the American College of Radiology Radiation Oncology In-Training Examination (TXIT, 2021), comprising 300 multiple-choice items, and (ii) a curated set of 60 authentic radiation oncologic vignettes representing diverse disease sites and treatment indications. For the vignette evaluation, GPT-5 was instructed to generate structured therapeutic plans and concise two-line summaries. Four board-certified radiation oncologists independently rated outputs for correctness, comprehensiveness, and hallucinations. Inter-rater reliability was quantified using Fleiss’ κ . GPT-5–14 results were compared to published GPT-3.5 and GPT-4 baselines. Results On the TXIT benchmark, GPT-5 achieved a mean accuracy of 92.8%, outperforming GPT-4 (78.8%) and GPT-3.5 (62.1%). Domain-specific gains were most pronounced in dose specification and diagnosis. In the vignette evaluation, GPT-5’s treatment recommendations were rated highly for correctness (mean 3.24/4, 95% CI: 3.11–3.38) and comprehensiveness (3.59/4, 95% CI: 3.49–3.69). Hallucinations were rare, flagged in 10.0% of all individual reviewer assessments (24 of 240), and no patient case reac...