Scholay

学术搜索 · AI 审稿 · LaTeX 协作

A hybrid framework with large language models for rare disease phenotyping

作者:Jinge Wu, Hang Dong, Zexi Li, Haowei Wang, R Li, Arijit Patra, Chengliang Dai, Waqar Ali, Phil Scordis, Honghan Wu · 发表于:BMC Medical Informatics and Decision Making · 年份:2024 · DOI:10.1186/s12911-024-02698-7 · 被引用次数:29 · 研究领域:Genomics and Rare Diseases、Biomedical Text Mining and Ontologies、Topic Modeling

PURPOSE: Rare diseases pose significant challenges in diagnosis and treatment due to their low prevalence and heterogeneous clinical presentations. Unstructured clinical notes contain valuable information for identifying rare diseases, but manual curation is time-consuming and prone to subjectivity. This study aims to develop a hybrid approach combining dictionary-based natural language processing (NLP) tools with large language models (LLMs) to improve rare disease identification from unstructured clinical reports. METHODS: We propose a novel hybrid framework that integrates the Orphanet Rare Disease Ontology (ORDO) and the Unified Medical Language System (UMLS) to create a comprehensive rare disease vocabulary. SemEHR, a dictionary-based NLP tool, is employed to extract rare disease mentions from clinical notes. To refine the results and improve accuracy, we leverage various LLMs, including LLaMA3, Phi3-mini, and domain-specific models like OpenBioLLM and BioMistral. Different prompting strategies, such as zero-shot, few-shot, and knowledge-augmented generation, are explored to optimize the LLMs' performance. RESULTS: The proposed hybrid approach demonstrates superior performance compared to traditional NLP systems and standalone LLMs. LLaMA3 and Phi3-mini achieve the highest F1 scores in rare disease identification. Few-shot prompting with 1-3 examples yields the best results, while knowledge-augmented generation shows limited improvement. Notably, the approach uncovers a ...