Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Fine-tuning medical language models for enhanced long-contextual understanding and domain expertise

作者:Qimin Yang, Jiexin Chen, Yue Sun, Yapeng Wang, Tao Tan · 发表于:Quantitative Imaging in Medicine and Surgery · 年份:2025 · DOI:10.21037/qims-2024-2655 · 被引用次数:17 · 研究领域:Artificial Intelligence in Healthcare and Education、Machine Learning in Healthcare、Topic Modeling

Background: Since the emergence of large language models (LLMs), a large number of applications have emerged in vertical fields, such as medical, legal, and subject education. By fine-tuning the pre-trained base model, professional knowledge can be well parameterized into model capabilities, enabling it to have better performance in specific fields. However, we observed that although the fine-tuned model has improved domain-specific knowledge, the performance of medical LLMs (Med-LLMs) in long-context understanding has declined significantly due to the large amount of knowledge-intensive fine-tuning, especially compared with the general language model with similar parameters. This study aims to investigate the problem of the decline in performance of Med-LLMs in long-context understanding. Methods: We designed a series of experiments to conduct open-book professional knowledge tests related to the medical field on models using different fine-tuning methods to evaluate their long-context understanding capabilities in the medical field. These experiments included benchmarks of general language models, benchmarks of medical language models, tests that adjusted the ratio and amount of general data and professional data during fine-tuning, and experimental data to determine the best data composition to optimize professional models and achieve a balance between long-context performance and specific domain knowledge. Results: Our experimental framework evaluated 5 general-purpose LL...