Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Cross validation for model selection: A review with examples from ecology

作者:Luke A. Yates, Zach Aandahl, Shane A. Richards, Barry W. Brook · 发表于:Ecological Monographs · 年份:2022 · DOI:10.1002/ecm.1557 · 被引用次数:457 · 研究领域:Species Distribution and Climate Change、Data Analysis with R、Soil Geostatistics and Mapping

Abstract Specifying, assessing, and selecting among candidate statistical models is fundamental to ecological research. Commonly used approaches to model selection are based on predictive scores and include information criteria such as Akaike's information criterion, and cross validation. Based on data splitting, cross validation is particularly versatile because it can be used even when it is not possible to derive a likelihood (e.g., many forms of machine learning) or count parameters precisely (e.g., mixed‐effects models). However, much of the literature on cross validation is technical and spread across statistical journals, making it difficult for ecological analysts to assess and choose among the wide range of options. Here we provide a comprehensive, accessible review that explains important—but often overlooked—technical aspects of cross validation for model selection, such as: bias correction, estimation uncertainty, choice of scores, and selection rules to mitigate overfitting. We synthesize the relevant statistical advances to make recommendations for the choice of cross‐validation technique and we present two ecological case studies to illustrate their application. In most instances, we recommend using exact or approximate leave‐one‐out cross validation to minimize bias, or otherwise k ‐fold with bias correction if k < 10. To mitigate overfitting when using cross validation, we recommend calibrated selection via our recently introduced modified one‐standard‐error ...