Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Evaluating ChatGPT’s diagnostic potential for pathology images

作者:Liya Ding, Lei Fan, Miao Shen, Yawen Wang, Kaiqin Sheng, Zijuan Zou, Huimin An, Zhinong Jiang · 发表于:Frontiers in Medicine · 年份:2025 · DOI:10.3389/fmed.2024.1507203 · 被引用次数:22 · 研究领域:Artificial Intelligence in Healthcare and Education、Radiomics and Machine Learning in Medical Imaging、Radiology practices and education

Background: Chat Generative Pretrained Transformer (ChatGPT) is a type of large language model (LLM) developed by OpenAI, known for its extensive knowledge base and interactive capabilities. These attributes make it a valuable tool in the medical field, particularly for tasks such as answering medical questions, drafting clinical notes, and optimizing the generation of radiology reports. However, keeping accuracy in medical contexts is the biggest challenge to employing GPT-4 in a clinical setting. This study aims to investigate the accuracy of GPT-4, which can process both text and image inputs, in generating diagnoses from pathological images. Methods: This study analyzed 44 histopathological images from 16 organs and 100 colorectal biopsy photomicrographs. The initial evaluation was conducted using the standard GPT-4 model in January 2024, with a subsequent re-evaluation performed in July 2024. The diagnostic accuracy of GPT-4 was assessed by comparing its outputs to a reference standard using statistical measures. Additionally, four pathologists independently reviewed the same images to compare their diagnoses with the model's outputs. Both scanned and photographed images were tested to evaluate GPT-4's generalization ability across different image types. Results: GPT-4 achieved an overall accuracy of 0.64 in identifying tumor imaging and tissue origins. For colon polyp classification, accuracy varied from 0.57 to 0.75 in different subtypes. The model achieved 0.88 accura...