Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Evaluating AI in medicine: a comparative analysis of expert and ChatGPT responses to colorectal cancer questions

作者:Wen Peng, Yifei Feng, Yao Cui, Sheng Zhang, Han Zhuo, Tianzhu Qiu, Yi Zhang, Junwei Tang, Yanhong Gu, Yueming Sun · 发表于:Scientific Reports · 年份:2024 · DOI:10.1038/s41598-024-52853-3 · 被引用次数:39 · 研究领域:Artificial Intelligence in Healthcare and Education、Radiomics and Machine Learning in Medical Imaging、Meta-analysis and systematic reviews

Colorectal cancer (CRC) is a global health challenge, and patient education plays a crucial role in its early detection and treatment. Despite progress in AI technology, as exemplified by transformer-like models such as ChatGPT, there remains a lack of in-depth understanding of their efficacy for medical purposes. We aimed to assess the proficiency of ChatGPT in the field of popular science, specifically in answering questions related to CRC diagnosis and treatment, using the book "Colorectal Cancer: Your Questions Answered" as a reference. In general, 131 valid questions from the book were manually input into ChatGPT. Responses were evaluated by clinical physicians in the relevant fields based on comprehensiveness and accuracy of information, and scores were standardized for comparison. Not surprisingly, ChatGPT showed high reproducibility in its responses, with high uniformity in comprehensiveness, accuracy, and final scores. However, the mean scores of ChatGPT's responses were significantly lower than the benchmarks, indicating it has not reached an expert level of competence in CRC. While it could provide accurate information, it lacked in comprehensiveness. Notably, ChatGPT performed well in domains of radiation therapy, interventional therapy, stoma care, venous care, and pain control, almost rivaling the benchmarks, but fell short in basic information, surgery, and internal medicine domains. While ChatGPT demonstrated promise in specific domains, its general efficiency...