Imputation of missing data in life‐history trait datasets: which approach performs the best?
作者:Caterina Penone, Ana D. Davidson, Kevin T. Shoemaker, Moreno Di Marco, Carlo Rondinini, Thomas M. Brooks, Bruce E. Young, Catherine H. Graham, Gabriel C. Costa · 发表于:Methods in Ecology and Evolution · 年份:2014 · DOI:10.1111/2041-210x.12232 · 被引用次数:476 · 研究领域:Wildlife Ecology and Conservation、Evolution and Paleontology Studies、Genetic and phenotypic traits in livestock
Summary Despite efforts in data collection, missing values are commonplace in life‐history trait databases. Because these values typically are not missing randomly, the common practice of removing missing data not only reduces sample size, but also introduces bias that can lead to incorrect conclusions. Imputing missing values is a potential solution to this problem. Here, we evaluate the performance of four approaches for estimating missing values in trait databases (K‐nearest neighbour (kNN), multivariate imputation by chained equations (mice), missForest and Phylopars), and test whether imputed datasets retain underlying allometric relationships among traits. Starting with a nearly complete trait dataset on the mammalian order Carnivora (using four traits), we artificially removed values so that the percent of missing values ranged from 10% to 80%. Using the original values as a reference, we assessed imputation performance using normalized root mean squared error. We also evaluated whether including phylogenetic information improved imputation performance inkNN, mice, and missForest (it is a required input in Phylopars). Finally, we evaluated the extent to which the allometric relationship between two traits (body mass and longevity) was conserved for imputed datasets by looking at the difference (bias) between the slope of the original and the imputed datasets or datasets with missing values removed. Three of the tested approaches (mice, missForest and Phylopars), result...