Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Clinical utility and readability of large language models’ responses to headache-related inquiries: A comparative evaluation with stratified application recommendations based on Chinese online health communities

作者:Qiansu Yang, Zhao Dong, Yinglin Yang, J Liu, Le Cai, Nan Bai, Tianlin Wang · 发表于:Intelligent Pharmacy · 年份:2026 · DOI:10.1016/j.ipha.2026.05.003 · 研究领域:Health Literacy and Information Accessibility、Artificial Intelligence in Healthcare and Education、Electronic Health Records Systems

Headache disorders affect a large proportion of the global population, and Large Language Models (LLMs) show potential in the provision of online health information services, with their response accuracy and readability requiring verification. This cross-sectional study used an expert panel of five headache specialists and three clinical neuroscience pharmacists to evaluate the quality and readability of LLM(represented by GPT-5-main) generated responses to 53 headache-related public inquiries paired with responses from physicians and pharmacists from Chinese online health communities (OHC)-Haodf.com (Jan 2020 to Dec 2024) via six evaluation dimensions and the AlphaReadabilityChinese tool, with stratified analyses by inquiry difficulty (Level 1: simple; Level 2: complex). The results showed that LLM responses were significantly longer than responses from physicians and pharmacists ( P <0.001), with a 27.4% mean accuracy for authorship distinction; LLM produced fewer incorrect contents (PR=0.60, 95% CI=0.50-0.71) and lower harm risk (PR=0.85, 95% CI=0.80-0.91), along with better alignment to medical consensus (PR=1.63, 95% CI=1.40-1.90), relevance (PR=2.32, 95% CI=1.79-3.00) and completeness (PR=2.95, 95% CI=2.31-3.75) than responses from physicians and pharmacists. Stratified analysis indicated that LLM achieved non-inferior health information quality to physicians and pharmacists for Level 1 inquiries (all P >0.05), and we recommend that Level 2 inquiries be subject to human...