Integrating multi-layer perceptron and random forest in an ensemble framework for improved genomic prediction accuracy and SHAP-derived interpretability of residual feed intake in cattle
作者:Edwin Ong Jun Kiat, Mark Mooney, Faisal I. Rezwan, Hui Wang, Masoud Shirali · 发表于:Figshare · 年份:2026 · DOI:10.6084/m9.figshare.c.8641292 · 研究领域:Artificial intelligence、Computer science、Machine learning、Mathematics、Statistics、Biology、Data mining
Abstract Background Feed efficiency (FE) is recognized as a vital component of sustainable dairy production, with residual feed intake (RFI) serving as a key metabolic indicator of FE independent of production levels. However, the genetic improvement of this complex trait is limited by the inability of conventional genomic Best Linear Unbiased Prediction (gBLUP) model to capture complex, non-linear genetic architectures and epistatic interactions. To address these limitations, this study aims to compare the predictive performance of machine learning (ML) approaches, specifically Random Forest (RF) and Multi-Layer Perceptron (MLP) models, against standard gBLUP using genomic data from 220 UK Holstein cows genotyped with the BovineSNP50 v3 BeadChip with 47,446 quality-controlled single nucleotide polymorphisms (SNPs), phenotyped for RFI from 1996–2023. SHapley Additive exPlanations (SHAP) were applied to interpret SNP feature importance from the ML models, and an ensemble framework was implemented to leverage the complementary strengths of RF and MLP. Results While the gBLUP model exhibited moderate predictive performance, the RF model demonstrated greater stability and accuracy compared to gBLUP, and the MLP showed higher variance across random states. The ensemble framework achieved the highest coefficient of determination (R2 = 0.39) and lowest root mean squared error (RMSE = 0.086). SHAP interpretability analysis revealed distinct genomic architectures between the models. T...