Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Rx-LLM: a benchmarking suite to evaluate safe large language model performance for medication-related tasks

作者:Xingmeng Zhao, Kaitlin Blotske, Moriah Cargile, Adeleine Tilley, Brian Murray, Yanjun Gao, Kelli Henry, Susan Smith, Erin F. Barreto, Seth R. Bauer, Sunghwan Sohn, Tianming Liu, Tell Bennett, Mitch Cohen, Andrea Sikora · 发表于:medRxiv · 年份:2025 · DOI:10.64898/2025.12.01.25341004 · 被引用次数:4 · 研究领域:Topic Modeling、Artificial Intelligence in Healthcare and Education、Machine Learning in Healthcare

Background: For large language models (LLMs) to reach their potential as information technology tools that make medication use safer, clinically relevant benchmarks capable of automated grading and designed specifically to measure the performance of LLMs for medication tasks are required. The purpose of this study was to design a suite of benchmarking tests reflective of Comprehensive Medication Management (CMM; the standard of care for medication optimization) and quantify the baseline performance of the latest LLMs. Methods: We established six benchmarks representing critical stages of the CMM process: drug formulation matching, drug order (sig) generation, drug route matching, drug-drug interaction identification, renal dose identification, and drug-indication matching. For each benchmark, we curated a clinician-annotated dataset comprising 250 standardized input-output pairs including both inpatient and outpatient medications. We evaluated the clinical knowledge retrieval capabilities of three LLMs: GPT-4o-mini, MedGemma-27B, and LLaMA3-70B. We employed a zero-shot prompting strategy, excluding in-context examples, to assess the models' internal clinical knowledge rather than their few-shot learning potential. To check reliability, each model was run three times using a temperature of 0.7 (a mid-range value of an LLM setting controlling text generation randomness). Performance was assessed using task-specific evaluation metrics including precision (positive predictive val...