Variable Importance Plots—An Introduction to the vip Package
作者:M. Greenwell Brandon, C. Boehmke Bradley · 发表于:The R Journal · 年份:2020 · DOI:10.32614/rj-2020-013 · 被引用次数:353 · 研究领域:Engineering and Materials Science Studies
In the era of "big data", it is becoming more of a challenge to not only build state-of-the-art predictive models, but also gain an understanding of what's really going on in the data.For example, it is often of interest to know which, if any, of the predictors in a fitted model are relatively influential on the predicted outcome.Some modern algorithms-like random forests (RFs) and gradient boosted decision trees (GBMs)-have a natural way of quantifying the importance or relative influence of each feature.Other algorithms-like naive Bayes classifiers and support vector machines-are not capable of doing so and model-agnostic approaches are generally used to measure each predictor's importance.Enter vip, an R package for constructing variable importance scores/plots for many types of supervised learning algorithms using model-specific and novel model-agnostic approaches.We'll also discuss a novel way to display both feature importance and feature effects together using sparklines, a very small line chart conveying the general shape or variation in some feature that can be directly embedded in text or tables.1 Although "interpretability" is difficult to formally define in the context of ML, we follow Doshi-Velez and Kim (2017) and describe "interpretable" as the ". . .ability to explain or to present in understandable terms to a human."2 In this context "importance" can be defined in a number of different ways.In general, we can describe it as the extent to which a feature has a...