Assessing the possibility of using large language models in ocular surface diseases
作者:Qian Ling, Yan-Mei Zeng, Hong Qi, Xian-Zhe Qian, Jin‐Yu Hu, Chong-Gang Pei, Hong Wei, Jie Zou, Cheng Chen, Xiaoyu Wang, Xu Chen, Zhenkai Wu, Yi Shao · 发表于:International Journal of Ophthalmology · 年份:2024 · DOI:10.18240/ijo.2025.01.01 · 被引用次数:18 · 研究领域:Artificial Intelligence in Healthcare and Education、AI in cancer detection、Clinical Reasoning and Diagnostic Skills
AIM: To assess the possibility of using different large language models (LLMs) in ocular surface diseases by selecting five different LLMS to test their accuracy in answering specialized questions related to ocular surface diseases: ChatGPT-4, ChatGPT-3.5, Claude 2, PaLM2, and SenseNova. METHODS: A group of experienced ophthalmology professors were asked to develop a 100-question single-choice question on ocular surface diseases designed to assess the performance of LLMs and human participants in answering ophthalmology specialty exam questions. The exam includes questions on the following topics: keratitis disease (20 questions), keratoconus, keratomalaciac, corneal dystrophy, corneal degeneration, erosive corneal ulcers, and corneal lesions associated with systemic diseases (20 questions), conjunctivitis disease (20 questions), trachoma, pterygoid and conjunctival tumor diseases (20 questions), and dry eye disease (20 questions). Then the total score of each LLMs and compared their mean score, mean correlation, variance, and confidence were calculated. RESULTS: GPT-4 exhibited the highest performance in terms of LLMs. Comparing the average scores of the LLMs group with the four human groups, chief physician, attending physician, regular trainee, and graduate student, it was found that except for ChatGPT-4, the total score of the rest of the LLMs is lower than that of the graduate student group, which had the lowest score in the human group. Both ChatGPT-4 and PaLM2 were mor...