Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Clinician expertise and prompt engineering enhance cancer information extraction in electronic health records by small language models

作者:Federica Corso, V. Peppoloni, Laura Mazzeo, giuseppe leone, Luana Passos, V. Miskovic, Justin Armanini, A. Ferrarin, Isabella C. Wiest, Fabian Wolf, Giulia Montelatici, Rebecca Romanò, Ambrosini Paolo, Tommaso Capoccia, Stefano Natangelo, Simone Rota, Paola Andena, Marta De Ponti, Alessandra Russo, Giulia Stasi, Leonardo Provenzano, Andrea Spagnoletti, Marco Meazza Prina, Chiara Cavalli, Claudia Giani, Roberta Serino, Michele Borracino, C. Bonalume, Rosa Maria Di Mauro, C. Agosta, Andra Diana Dumitrascu, Giorgia Di Liberti, Giulia Corrao, Teresa Beninato, Monica Ganzinelli, Mario Occhipinti, Marta Brambilla, Claudia Proto, Jakob Nikolas Kather, Alessandra Pedrocchi, Filippo de Braud, Giuseppe Lo Russo, Paolo Baili, Arsela Prelaj · 发表于:Communications Medicine · 年份:2026 · DOI:10.1038/s43856-026-01790-5 · 研究领域:Topic Modeling、Electronic Health Records Systems、Machine Learning in Healthcare

Real-world data (RWD) in unstructured electronic health records (EHRs) is crucial for understanding complex diseases like cancer, but extracting structured information is challenging due to linguistic variability, semantic complexity, and privacy concerns. This study evaluates the performance of four small, locally deployable language models for information extraction from Italian EHRs. We examine three prompting strategies (zero-shot, few-shot, and annotated few-shot) across English and Italian, involving clinicians with varying expertise to assess the impact of prompt design on accuracy. We evaluate the performance of four open-source small language models (SLMs) for clinical information extraction from Italian electronic health records (EHRs) in the APOLLO 11 trial on non-small cell lung cancer (NSCLC). The extraction protocol involves four steps: problem definition, data preprocessing, Large Language Model (LLM)-based information extraction, and output evaluation. We show that general-purpose models (e.g., LLaMA 3.1 8B) outperform biomedical models in most tasks, particularly in extracting binary features. Multiclass variables such as TNM (Tumor, Node, Metastasis) staging, PD-L1 (Programmed death-ligand 1), and ECOG-PS (Eastern Cooperative Oncology Group-Performance Status) are more difficult due to implicit language and lack of standardization. Few-shot prompting and native-language inputs significantly improve performance and reduced hallucinations. Clinical expertise e...