Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Machine learning based survival prediction of DLBCL: Using multimodal data.

作者:Fahad Ahmed, Zachary AK Frosch, Rashmi Khanal, Lalit Sehgal, Shazia Nakhoda, Marcus Messmer, Mariusz Wasik, Nicholas Mackrides, Yibin Yang, Reza Nejati · 发表于:Blood · 年份:2025 · DOI:10.1182/blood-2025-7068 · 研究领域:Lymphoma Diagnosis and Treatment、Radiomics and Machine Learning in Medical Imaging、AI in cancer detection

Abstract Background: Diffuse-large B-cell lymphoma (DLBCL) is a heterogeneous disease with outcomes influenced by various clinical and molecular predictors. Our aim for this study was to develop a machine learning based prediction of overall survival using clinical, laboratory, and gene expression data. Methods: We previously analyzed publicly available normalized expression data from NCBI-GEO (GSE181063), encompassing 1,311 DLBCL patients; where we extracted 41 genes that were associated with overall survival using statistical modeling. Here we developed a binary gene matrix that was generated from median expression of all the study population and anything with higher than median expression was considered (high expression) for the differential gene expression of 41 genes and clinical (Age, Gender, first line of treatment, curative intent, ECOG, B symptoms, Stage and IPI score), Imaging (number of lymph nodes), and lab data (LDH, CBC findings) was added to it. Outcomes included time-points (>6 months, >1, >3, >5 and >10 year(s)) and were binarized (alive = 1 and dead = 0). SMOTE (Synthetic Minority Over-sampling Technique) was used because of the imbalanced nature of biological and clinical data. Nine different machine learning models (MLMs) were developed, and the best MLMs were ranked and presented here. MLMs included: RandomForest (RF), Gradient Boosting (GB), XGBoost (XGB), AdaBoost (AB), Logistic Regression (LR), Naïve Bayes (NB), Multi...