Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Revolution or risk?—Assessing the potential and challenges of GPT-4V in radiologic image interpretation

作者:Marc Huppertz, Robert Malte Siepmann, David B. Topp, Omid Nikoubashman, Can Yüksel, Christiane Katharina Kuhl, Daniel Truhn, Sven Nebelung · 发表于:European Radiology · 年份:2024 · DOI:10.1007/s00330-024-11115-6 · 被引用次数:52 · 研究领域:Artificial Intelligence in Healthcare and Education、Explainable Artificial Intelligence (XAI)、Radiology practices and education

OBJECTIVES: ChatGPT-4 Vision (GPT-4V) is a state-of-the-art multimodal large language model (LLM) that may be queried using images. We aimed to evaluate the tool's diagnostic performance when autonomously assessing clinical imaging studies. MATERIALS AND METHODS: A total of 206 imaging studies (i.e., radiography (n = 60), CT (n = 60), MRI (n = 60), and angiography (n = 26)) with unequivocal findings and established reference diagnoses from the radiologic practice of a large university hospital were accessed. Readings were performed uncontextualized, with only the image provided, and contextualized, with additional clinical and demographic information. Responses were assessed along multiple diagnostic dimensions and analyzed using appropriate statistical tests. RESULTS: With its pronounced propensity to favor context over image information, the tool's diagnostic accuracy improved from 8.3% (uncontextualized) to 29.1% (contextualized, first diagnosis correct) and 63.6% (contextualized, correct diagnosis among differential diagnoses) (p ≤ 0.001, Cochran's Q test). Diagnostic accuracy declined by up to 30% when 20 images were re-read after 30 and 90 days and seemed unrelated to the tool's self-reported confidence (Spearman's ρ = 0.117 (p = 0.776)). While the described imaging findings matched the suggested diagnoses in 92.7%, indicating valid diagnostic reasoning, the tool fabricated 258 imaging findings in 412 responses and misidentified imaging modalities or anatomic regions in...