Scholay

学术搜索 · AI 审稿 · LaTeX 协作

The “David Vs Goliath” Study: Application of Large Language Models (LLM) for Automatic Medical Information Retrieval from Multiple Data Sources to Accelerate Clinical and Translational Research in Hematology

作者:Mattia Delleani, S. D'amico, Elisabetta Sauta, G. Asti, Elena Zazzetti, Alessia Campagna, L. Lanino, G. Maggioni, M. C. Grondelli, Alessandro Forcina Barrero, Pierandrea Morandini, M. Ubezio, G. Todisco, Antonio Russo, C. Tentori, Alessandro Buizza, Arturo Bonometti, Cesare Lancellotti, Luca di Tommaso, Daoud Rahal, Marilena Bicchieri, V. Savevski, Armando Santoro, V. Santini, Francesc Solé, U. Platzbecker, Pierre Fenaux, M. Díez-Campelo, R. Komrokji, G. Garcia-Manero, T. Haferlach, S. Kordasti, A. Zeidan, G. Castellani, M. D. Della Porta · 发表于:Blood · 年份:2024 · DOI:10.1182/blood-2024-205621 · 被引用次数:3

Background. The innovation process in hematology requires access to a large amount of healthcare data. However, 97% of patient data produced by hospitals remains unused (Source: Deloitte, Health Data, 2023), primarily due to privacy limitations, lack of data harmonization from different sources, and the unstructured and dispersed nature of the information. Large Language Models (LLM) are computational models capable of performing general-purpose language generation and other natural language processing tasks. These models acquire these abilities by learning statistical relationships from vast amounts of text through a computationally intensive self-supervised and semi-supervised training process. LLMs have been increasingly utilized in healthcare to enhance diagnostics, streamline patient interactions, and improve overall clinical workflows. In this project, we analyze the potential of Artificial Intelligence (AI) solutions based on LLM for data retrieval, extraction and generation to create standardized datasets to accelerate clinical and translational research in blood diseases in hematology. Aims. The “David vs Goliath” study was conducted by Synthema EU consortium with the following aims to: 1) develop AI solution leveraging LLM for information retrieval, extraction and generation of research-ready datasets from multiple medical sources; 2) evaluate clinical and statistical fidelity of AI-retrieved dataset through a specific Validation Framework (VF); 3) validate the re...