Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Effects of data transformation and model selection on feature importance in microbiome classification data

作者:Zuzanna Karwowska, Oliver Aasmets, Mait Metspalu, Metspalu A, Lili Milani, Tõnu Esko, Tomasz Kościółek, Elin Org · 发表于:Microbiome · 年份:2025 · DOI:10.1186/s40168-024-01996-6 · 被引用次数:38 · 研究领域:Gut microbiota and health、Machine Learning in Bioinformatics、Gene expression and cancer classification

BACKGROUND: Accurate classification of host phenotypes from microbiome data is crucial for advancing microbiome-based therapies, with machine learning offering effective solutions. However, the complexity of the gut microbiome, data sparsity, compositionality, and population-specificity present significant challenges. Microbiome data transformations can alleviate some of the aforementioned challenges, but their usage in machine learning tasks has largely been unexplored. RESULTS: Our analysis of over 8500 samples from 24 shotgun metagenomic datasets showed that it is possible to classify healthy and diseased individuals using microbiome data with minimal dependence on the choice of algorithm or transformation. Presence-absence transformations performed comparably to abundance-based transformations, and only a small subset of predictors is necessary for accurate classification. However, while different transformations resulted in comparable classification performance, the most important features varied significantly, which highlights the need to reevaluate machine learning-based biomarker detection. CONCLUSIONS: Microbiome data transformations can significantly influence feature selection but have a limited effect on classification accuracy. Our findings suggest that while classification is robust across different transformations, the variation in feature selection necessitates caution when using machine learning for biomarker identification. This research provides valuable in...