Scholay

学术搜索 · AI 审稿 · LaTeX 协作

RadCLARE: an automated clinical language engine for detecting semantic errors in radiology reports

作者:Feng Pan, Jie Lou, Yusheng Guo, Du Wang, Zhonghua Wang, Qianqian Fan, Hao Wang, Chuansheng Zheng, Lian Yang · 发表于:European Radiology Experimental · 年份:2025 · DOI:10.1186/s41747-025-00659-x · 被引用次数:1 · 研究领域:Radiology practices and education、Artificial Intelligence in Healthcare and Education、Topic Modeling

BACKGROUND: Errors in radiology reports can result in inappropriate/harmful decisions. We investigated whether large language models can reduce the error rate. MATERIALS AND METHODS: We developed the radiology-specific clinical language anomaly recognition engine (RadCLARE) network, an automated engine based on the bidirectional encoder representations from transformers (BERT)-base model, designed to detect semantic errors in Chinese radiology reports and trained using 1.4 million reports, including 615,920 digital radiography, 560,310 computed tomography reports, and 223,480 magnetic resonance reports. One thousand reports were randomly selected for expert manual annotation. Inter-reader agreement for error detection and classification was assessed using Cohen κ and Gwet AC1. The RadCLARE's detection was compared against the expert references. Changes in error rates before (baseline test dataset, BTD) and after (experimental test dataset, ETD) RadCLARE implementation were analyzed. Finally, radiologists were invited to complete questionnaires to evaluate satisfaction and rate the system across five dimensions. RESULTS: Among the 1,000 reports, a total of 506 errors were identified as the reference standard. Inter-reader agreement was substantial for error detection (κ = 0.77) and excellent for error classification (Gwet AC1 = 0.94). RadCLARE successfully detected 437/506 errors, with 87.3% accuracy, 88.3% precision, 86.4% recall, and 87.4% F1-score. The BTD comprised 571,264...