A Dataset for Evaluating Large Language Models on Chinese National Medical Licensing Examinations
作者:Hui Zong, Bairong Shen · 发表于:Zenodo (CERN European Organization for Nuclear Research) · 年份:2026 · DOI:10.5281/zenodo.18951465 · 被引用次数:1 · 研究领域:Computer science、Information retrieval、Natural language processing、Data science、Artificial intelligence、World Wide Web、Medical education
The dataset integrates question-answer pairs from three sources, including PubMed, GitHub, and MedExamLLM. CNMLEQA comprises two subsets: CNMLEQA-10k (9,890 questions) and CNMLEQA-3k (2,949 questions), each consisting of multiple-choice questions with five options and one correct answer. Questions are annotated with key dimensions including: (1) question type (knowledge-based or case-based), (2) auxiliary metadata such as examination year, 3) clinical scenario information across five dimensions: disease or diagnosis, surgery, medication, laboratory examination, and symptom or sign. Annotation was conducted by clinical experts.