Scholay

学术搜索 · AI 审稿 · LaTeX 协作

The performance evaluation of artificial intelligence ERNIE bot in Chinese National Medical Licensing Examination

作者:Leiyun Huang, Jinghan Hu, Qingjin Cai, Guangjie Fu, Zhenglin Bai, Yongzhen Liu, Ji Zheng, Zengdong Meng · 发表于:Postgraduate Medical Journal · 年份:2024 · DOI:10.1093/postmj/qgae062 · 被引用次数:17 · 研究领域:Artificial Intelligence in Healthcare and Education

Clinical medical students are often tasked with acquiring a vast amount of theoretical knowledge and clinical experience across various subdisciplines during their academic journey. With limited time and energy, it is impractical for students to complete clinical rotations in every department, making it crucial to prioritize theoretical studies. By focusing on theoretical learning, students can indirectly gain clinical experience and develop a comprehensive understanding of different disciplines. Evaluating students’ performance in theoretical studies is a key component in assessing their progress. As a result, some researchers [1–4] are exploring the use of artificial intelligence in the National Medical Licensing Examination to assess its potential applications in medical research, education, and beyond. The study by Richard C. Armitage et al. [1] highlighted the impressive performance of the large language model (LLM) GPT-4 in the Membership of the Royal College of General Practitioners examination in an English context. In a similar vein, U. Hin Lai et al. [2] demonstrated ChatGPT’s strong performance in the United Kingdom Medical Licensing Assessment. However, a comparison of these two studies [3, 4] revealed that while GPT-4 excelled in the UK exams, it did not perform well in the Chinese National Medical Licensing Examination. ERNIE Bot is a LLM developed by Baidu, a Chinese company. It is trained based on a vast amount of Chinese-language content, and its version 3.5 ...