Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Improving large language models for miRNA information extraction via prompt engineering

作者:Rongrong Wu, Hui Zong, Erman Wu, Jiakun Li, Yiqing Zhou, Chi Zhang, Yingbo Zhang, Jiao Wang, Tong Tang, Bairong Shen · 发表于:Computer Methods and Programs in Biomedicine · 年份:2025 · DOI:10.1016/j.cmpb.2025.109033 · 被引用次数:4 · 研究领域:Biomedical Text Mining and Ontologies、Topic Modeling、Genomics and Rare Diseases

OBJECTIVE: Large language models (LLMs) demonstrate significant potential in biomedical knowledge discovery, yet their performance in extracting fine-grained biological information, such as miRNA, remains insufficiently explored. Accurate extraction of miRNA-related information is essential for understanding disease mechanisms and identifying biomarkers. This study aims to comprehensively evaluate the capabilities of LLMs in miRNA information extraction through diverse prompt learning strategies. METHODS: Three high-quality miRNA information extraction datasets were constructed to support the benchmarking and training of generative LLMs, specifically Re-Tex, Re-miR and miR-Cancer. These datasets encompass three types of entities: miRNAs, genes, and diseases, along with their relationships. The accuracy and reliability of three LLMs, including GPT-4o, Gemini, and Claude, were evaluated and compared with traditional models. Different prompt engineering strategies were implemented to enhance the LLMs' performance, including baseline prompts, 5-shot Chain of Thought prompts, and generated knowledge prompts. RESULTS: The combination of optimized prompt strategies significantly improved overall entity extraction performance across both trained and untrained datasets. Generated knowledge prompting achieved the highest performance, with maximum F1 scores of 76.6 % for entity extraction and 54.8 % for relationship extraction. Comparative analysis indicated GPT-4o exhibited superior pe...