Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Flexible Imputation of Missing Data

作者:Hakan Demirtaş · 发表于:Journal of Statistical Software · 年份:2018 · DOI:10.18637/jss.v085.b04 · 被引用次数:726 · 研究领域:Machine Learning and Data Classification、Machine Learning in Healthcare

Missingness is a commonly occurring phenomenon in many applications.Determining a suitable analytical approach in the absence of complete observations is a major focus of scientific inquiry due to the extra sophistication that arises through missing data.Incompleteness generally complicates the statistical analysis in terms of reduced statistical power, biased parameter estimates, and degraded confidence intervals, and thereby may lead to false inferences.Developments in computational statistics have produced flexible missing-data procedures with a sound statistical basis.One of these procedures involves multiple imputation (MI), which is a stochastic simulation technique in which the missing values are replaced by m > 1 simulated versions.Subsequently, each of the simulated complete data sets is analyzed by standard methods, and the results are combined into a single inferential statement that formally incorporates missing-data uncertainty to the modeling process.MI has gained widespread acceptance and popularity in the last few decades.It has some well-accepted advantages: First, MI allows researchers to use conventional models and software; an imputed data set may be analyzed by literally any method that would be appropriate if the data were complete.As computing environments and statistical models grow increasingly complex, the value of using familiar methods and software becomes more pronounced.Second, there are still many classes of problems for which no direct maximum ...