Host transcriptomics and machine learning for secondary bacterial infections in patients with COVID-19: a prospective, observational cohort study
作者:Meagan Carney, Tiana M Pelaia, Tracy Chew, Sally Teoh, Amy Phu, Karan Kim, Ya Wang, Jonathan R. Iredell, Yoann Zerbib, Anthony S. McLean, Klaus Schughart, Benjamin Tang, Maryam Shojaei, Kirsty R. Short, Meagan Carney, Tiana M Pelaia, Tracy Chew, Sally Teoh, Amy Phu, Karan Kim, Ya Wang, Jonathan R. Iredell, Gabriella Cirmena, Alberto Ballestrero, Allan W. Cripps, Amanda J. Cox, Andrea De Maria, Arutha Kulasinghe, Carl G. Feng, Damien Chaussabel, Darawan Rinchai, Davide Bedognetti, Gabriele Zoppoli, Gunawan Gunawan, Irani Thevarajan, Jennifer Audsley, John‐Sebastian Eden, Marcela Kralovcova, Marek Nalos, Marko Radic, Martin Matějovič, Michele Bedognetti, Miroslav Průcha, Mohammed Toufiq, Narasaraju Teluguakula, Nicholas P. West, Paolo Cremonesi, Philip N Britton, Ricardo Garcia Branco, Rostyslav Bilyy, Stephen Macdonald, Thomas Karvunidis, Tim N Kwan, Velma Herwanto, Win Sen Kuan, Yoann Zerbib, Anthony S. McLean, Klaus Schughart, Benjamin Tang, Maryam Shojaei, Kirsty R. Short · 发表于:The Lancet Microbe · 年份:2024 · DOI:10.1016/s2666-5247(23)00363-4 · 被引用次数:14 · 研究领域:Antibiotic Use and Resistance、Gut microbiota and health、Pneumonia and Respiratory Infections
BACKGROUND: Viral respiratory tract infections are frequently complicated by secondary bacterial infections. This study aimed to use machine learning to predict the risk of bacterial superinfection in SARS-CoV-2-positive individuals. METHODS: In this prospective, multicentre, observational cohort study done in nine centres in six countries (Australia, Indonesia, Singapore, Italy, Czechia, and France) blood samples and RNA sequencing were used to develop a robust model of predicting secondary bacterial infections in the respiratory tract of patients with COVID-19. Eligible participants were older than 18 years, had known or suspected COVID-19, and symptoms of a recent respiratory infection. A control cohort of participants without COVID-19 who were older than 18 years and with no infection symptoms was also recruited from one Australian centre. In the pre-analysis phase, data were filtered to include only individuals with complete blood transcriptomics and patient data (ie, age, sex, location, and WHO severity score at the time of sample collection). The dataset was then divided randomly (4:1) into a training set (80%) and a test set (20%). Gene expression data in the training set and control cohort were used for differential expression analysis. Differentially expressed genes, along with WHO severity score, location, age, and sex, were used for feature selection with least absolute shrinkage and selection operator (LASSO) in the training set. For LASSO analysis, samples were ...