Artificial intelligence can match domain experts in evidence extraction and critical appraisal of microbial oncogenesis research publications
作者:Kaela Kokkas, Hairong Wang, Richard Klein, Nazir A. Ismail, Natalie Irwin, Mohammad Z. Moonsamy, Kubendran Naidoo, Jeremy Nel, Ekene E. Nweke, Raveen Parboosing, Emmanuel K. Sekyi, Rebecca T. van Dorsten, Bruce A. Bassett, Robert F. Breiman · 发表于:Frontiers in Cellular and Infection Microbiology · 年份:2026 · DOI:10.3389/fcimb.2026.1876326 · 研究领域:Artificial Intelligence in Healthcare and Education、AI in cancer detection、Biomedical Text Mining and Ontologies
Background: Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying and confirming novel microbial oncogenicity could yield strategies and tools that will reduce disease burdens. However, relevant evidence may be dispersed across a vast biomedical literature that is infeasible for humans to comprehensively synthesize. Large Language Models (LLMs) may enable scalable, expert-level systematic evidence synthesis to identify high priority microbe-cancer pairs; however, such capabilities have not yet been demonstrated. Methods: Domain experts were recruited to create a human-validated test dataset to benchmark the performance of LLMs (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, and GPT-5 Nano) on 24 original research papers using Mouse Mammary Tumor Virus-Like Virus and breast cancer as a case study. We devised a structured template for evidence extraction and appraisal of papers, consisting of multiple choice, Likert-scale, multi-select, and free-text question types (77 question items across 24 papers). Agreement between (1) experts, and (2) experts and each LLM, was determined per question instance using novel scoring metrics. LLMs were assessed by comparing inter-expert and expert-LLM agreement score distributions to determine whether LLMs behaved as additional experts by either increasing or maintaining inter-expert agreement. Free-text responses were further evaluated qualitatively. Results: Across all question types, LLM responses aligned closely wit...