Spurious model comparisons are widespread in biomedical artificial intelligence
作者:Tianchu Zeng, Hui Li, Shaoshi Zhang, Yan Quan Tan, Fang Tian, Csaba Orbán, Lijun An, Wanyu Che, Jingwen Cheng, Joanna Su Xian Chong, Niousha Dehestani, Zijian Dong, Xin Li, Zhizhou Li, Mervyn Jun Rui Lim, Yi Lin, Qinrui Ling, Zijie Ling, Xi Zhi Low, Sina Mansour L., Kwun Kei Ng, Thuan Tinh Nguyen, Leon Qi Rong Ooi, Shreya Pande, Xing Qian, Jingxuan Ruan, Z WANG, Yapei Xie, Chen Zhang, Yichi Zhang, K Patil, Linden Parkes, Elvisha Dhamala, Sidhant Chopra, Andrew Zalesky, Avram Holmes, S Eickhoff, Juan Helen Zhou, Olivier Renaud, Nico Dosenbach, Konrad P. Körding, Danilo Bzdok, Thomas Nichols, B T Thomas Yeo · 发表于:bioRxiv (Cold Spring Harbor Laboratory) · 年份:2026 · DOI:10.64898/2026.05.17.724301 · 被引用次数:1 · 研究领域:Machine Learning in Materials Science、Computational Drug Discovery Methods、Cell Image Analysis Techniques
Machine learning is accelerating biomedical research. Cross-validation is widely used to compare predictive performance - not only to benchmark algorithms, but also to inform scientific applications, such as ranking biomarkers. However, prediction performance estimates across cross-validation folds are not independent. Standard tests for comparing prediction performance (e.g., paired t-test) assume independence and can therefore inflate false positive rates. In a PRISMA-guided meta-analysis of 210 studies (impact factor ≥15, 1 June 2020 - 1 June 2025), we find that 97% ignored fold dependence when comparing prediction performance. This problem is ubiquitous across scientific fields and unaffected by impact factor, rigor-promoting policies, or open science practices. Simulations across 420 scenarios spanning four diverse datasets show that ignoring fold dependence leads to invalid false positive control in most settings. Repeated cross-validation further compounds this problem, with false positive rates rising toward 100% as the number of repetitions grows. Existing fold-dependence-aware tests rely on strong assumptions because the variance of fold-level statistics and the between-fold correlation cannot be disentangled under standard cross-validation. We therefore propose the SHARP (Split-HAlf RePeated) test, a simple modification to standard cross-validation that enables direct estimation of variance and correlation. Benchmarked against 12 tests, SHARP provides the best over...