{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T12:31:10Z","timestamp":1786537870331,"version":"3.56.0"},"reference-count":46,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2021,10,29]],"date-time":"2021-10-29T00:00:00Z","timestamp":1635465600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,10,29]],"date-time":"2021-10-29T00:00:00Z","timestamp":1635465600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100000086","name":"Directorate for Mathematical and Physical Sciences","doi-asserted-by":"publisher","award":["1712554"],"award-info":[{"award-number":["1712554"]}],"id":[{"id":"10.13039\/100000086","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000086","name":"Directorate for Mathematical and Physical Sciences","doi-asserted-by":"publisher","award":["1712041"],"award-info":[{"award-number":["1712041"]}],"id":[{"id":"10.13039\/100000086","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000086","name":"Directorate for Mathematical and Physical Sciences","doi-asserted-by":"publisher","award":["2015400"],"award-info":[{"award-number":["2015400"]}],"id":[{"id":"10.13039\/100000086","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Stat Comput"],"published-print":{"date-parts":[[2021,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper reviews and advocates against the use of permute-and-predict (PaP) methods for interpreting black box functions. Methods such as the variable importance measures proposed for random forests, partial dependence plots, and individual conditional expectation plots remain popular because they are both model-agnostic and depend only on the pre-trained model output, making them computationally efficient and widely available in software. However, numerous studies have found that these tools can produce diagnostics that are highly misleading, particularly when there is strong dependence among features. The purpose of our work here is to (i) review this growing body of literature, (ii) provide further demonstrations of these drawbacks along with a detailed explanation as to why they occur, and (iii) advocate for alternative measures that involve additional modeling. In particular, we describe how breaking dependencies between features in hold-out data places undue emphasis on sparse regions of the feature space by forcing the original model to extrapolate to regions where there is little to no data. We explore these effects across various model setups and find support for previous claims in the literature that PaP metrics can vastly over-emphasize correlated features in both variable importance measures and partial dependence plots. As an alternative, we discuss and recommend more direct approaches that involve measuring the change in model performance after muting the effects of the features under investigation.\n<\/jats:p>","DOI":"10.1007\/s11222-021-10057-z","type":"journal-article","created":{"date-parts":[[2021,10,29]],"date-time":"2021-10-29T07:03:07Z","timestamp":1635490987000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":189,"title":["Unrestricted permutation forces extrapolation: variable importance requires at least one more model, or there is no free variable importance"],"prefix":"10.1007","volume":"31","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2648-1167","authenticated-orcid":false,"given":"Giles","family":"Hooker","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lucas","family":"Mentch","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Siyu","family":"Zhou","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,10,29]]},"reference":[{"issue":"4","key":"10057_CR1","doi-asserted-by":"publisher","first-page":"2249","DOI":"10.1016\/j.csda.2007.08.015","volume":"52","author":"KJ Archer","year":"2008","unstructured":"Archer, K.J., Kimes, R.V.: Empirical characterization of random forest variable importance measures. Comput. Stat. Data Anal. 52(4), 2249\u20132260 (2008)","journal-title":"Comput. Stat. Data Anal."},{"issue":"5","key":"10057_CR2","doi-asserted-by":"publisher","first-page":"2055","DOI":"10.1214\/15-AOS1337","volume":"43","author":"RF Barber","year":"2015","unstructured":"Barber, R.F., Cand\u00e8s, E.J., et al.: Controlling the false discovery rate via knockoffs. Ann. Stat. 43(5), 2055\u20132085 (2015)","journal-title":"Ann. Stat."},{"key":"10057_CR3","unstructured":"B\u00e9nard, C., Da Veiga, S., Scornet, E.: Mda for random forests: inconsistency, and a practical solution via the sobol-mda. arXiv preprintarXiv:2102.13347 (2021)"},{"issue":"1","key":"10057_CR4","doi-asserted-by":"publisher","first-page":"175","DOI":"10.1111\/rssb.12340","volume":"82","author":"TB Berrett","year":"2020","unstructured":"Berrett, T.B., Wang, Y., Barber, R.F., Samworth, R.J.: The conditional permutation test for independence while controlling for confounders. J. R. Stat. Soc. Ser. B Stat. Methodol. 82(1), 175\u2013197 (2020)","journal-title":"J. R. Stat. Soc. Ser. B Stat. Methodol."},{"key":"10057_CR5","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman, L.: Random Forests. Mach. Learn. 45, 5\u201332 (2001)","journal-title":"Mach. Learn."},{"key":"10057_CR6","doi-asserted-by":"crossref","unstructured":"Candes, E., Fan, Y., Janson, L., Lv, J.: Panning for gold: \u201cmodel-x\u201d knockoffs for high dimensional controlled variable selection. J. R. Stat. Soc. Ser. B Stat. Methodol. 80(3), 551\u2013577 (2018)","DOI":"10.1111\/rssb.12265"},{"key":"10057_CR7","doi-asserted-by":"publisher","first-page":"2420","DOI":"10.1214\/12-EJS749","volume":"6","author":"G Chastaing","year":"2012","unstructured":"Chastaing, G., Gamboa, F., Prieur, C., et al.: Generalized Hoeffding\u2013Sobol decomposition for dependent variables-application to sensitivity analysis. Electron. J. Stat. 6, 2420\u20132448 (2012)","journal-title":"Electron. J. Stat."},{"key":"10057_CR8","unstructured":"Coleman, T., Peng, W., Mentch, L.: Scalable and efficient hypothesis testing with random forests. arXiv preprintarXiv:1904.07830 (2019)"},{"issue":"1","key":"10057_CR9","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1186\/1471-2105-7-3","volume":"7","author":"R D\u00edaz-Uriarte","year":"2006","unstructured":"D\u00edaz-Uriarte, R., De Andres, S.A.: Gene selection and classification of microarray data using random forest. BMC Bioinform. 7(1), 3 (2006)","journal-title":"BMC Bioinform."},{"key":"10057_CR10","first-page":"66","volume":"1\u201315","author":"H Fanaee-T","year":"2013","unstructured":"Fanaee-T, H., Gama, J.: Event labeling combining ensemble detectors and background knowledge. Prog. Artif. Intell. 1\u201315, 66 (2013)","journal-title":"Prog. Artif. Intell."},{"key":"10057_CR11","unstructured":"Fisher, A., Rudin, C., Dominici, F.: All models are wrong, but many are useful: Learning a variable\u2019s importance by studying an entire class of prediction models simultaneously. J. Mach. Learn. Res. 20(177), 1\u201381 (2019)"},{"key":"10057_CR12","first-page":"66","volume":"1189\u20131232","author":"JH Friedman","year":"2001","unstructured":"Friedman, J.H.: Greedy function approximation: a gradient boosting machine. Ann. Stat. 1189\u20131232, 66 (2001)","journal-title":"Ann. Stat."},{"issue":"1","key":"10057_CR13","doi-asserted-by":"publisher","first-page":"44","DOI":"10.1080\/10618600.2014.907095","volume":"24","author":"A Goldstein","year":"2015","unstructured":"Goldstein, A., Kapelner, A., Bleich, J., Pitkin, E.: Peeking inside the black box: visualizing statistical learning with plots of individual conditional expectation. J. Comput. Graph. Stat. 24(1), 44\u201365 (2015)","journal-title":"J. Comput. Graph. Stat."},{"key":"10057_CR14","doi-asserted-by":"publisher","first-page":"15","DOI":"10.1016\/j.csda.2015.04.002","volume":"90","author":"B Gregorutti","year":"2015","unstructured":"Gregorutti, B., Michel, B., Saint-Pierre, P.: Grouped variable importance with random forests and application to multiple functional data analysis. Comput. Stat. Data Anal. 90, 15\u201335 (2015)","journal-title":"Comput. Stat. Data Anal."},{"issue":"3","key":"10057_CR15","doi-asserted-by":"publisher","first-page":"66","DOI":"10.1198\/106186007X237892","volume":"16","author":"G Hooker","year":"2007","unstructured":"Hooker, G.: Generalized functional Anova diagnostics for high-dimensional functions of dependent variables. J. Comput. Graph. Stat. 16(3), 66 (2007)","journal-title":"J. Comput. Graph. Stat."},{"issue":"4","key":"10057_CR16","doi-asserted-by":"publisher","first-page":"558","DOI":"10.1002\/sim.7803","volume":"38","author":"H Ishwaran","year":"2019","unstructured":"Ishwaran, H., Lu, M.: Standard errors and confidence intervals for variable importance in random forest regression, classification, and survival. Stat. Med. 38(4), 558\u2013582 (2019)","journal-title":"Stat. Med."},{"key":"10057_CR17","unstructured":"Lehmann, E.L., Romano, J.P.: Testing Statistical Hypotheses. Springer (2006)"},{"key":"10057_CR18","doi-asserted-by":"crossref","unstructured":"Lei, J., G\u2019Sell, M., Rinaldo, A., Tibshirani, R.J., Wasserman, L.: Distribution-free predictive inference for regression. J. Am. Stat. Assoc. 113(523), 1094\u20131111 (2018)","DOI":"10.1080\/01621459.2017.1307116"},{"key":"10057_CR19","unstructured":"Li, X., Wang, Y., Basu, S., Kumbier, K., Yu, B.: A debiased mdi feature importance measure for random forests. arXiv preprint arXiv:1906.10845 (2019)"},{"key":"10057_CR20","unstructured":"Liu, Y., Zheng, C.: Auto-encoding knockoff generator for fdr controlled variable selection. arXiv preprint arXiv:1809.10765 (2018)"},{"key":"10057_CR21","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1080\/03610926.2020.1764042","volume":"66","author":"M Loecher","year":"2020","unstructured":"Loecher, M.: Unbiased variable importance for random forests. Commun. Stat. Theory Methods 66, 1\u201313 (2020)","journal-title":"Commun. Stat. Theory Methods"},{"key":"10057_CR22","doi-asserted-by":"crossref","unstructured":"Lou, Y., Caruana, R., Gehrke, J., Hooker, G.: Accurate intelligible models with pairwise interactions. In: Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 623\u2013631. ACM (2013)","DOI":"10.1145\/2487575.2487579"},{"key":"10057_CR23","unstructured":"Lundberg, S.M., Lee, S.-I.: A unified approach to interpreting model predictions. In: Proceedings of the 31st International Conference on Neural Information Processing Systems, pp. 4768\u20134777 (2017)"},{"issue":"1","key":"10057_CR24","first-page":"841","volume":"17","author":"L Mentch","year":"2016","unstructured":"Mentch, L., Hooker, G.: Quantifying uncertainty in random forests via confidence intervals and hypothesis tests. J. Mach. Learn. Res. 17(1), 841\u2013881 (2016)","journal-title":"J. Mach. Learn. Res."},{"issue":"3","key":"10057_CR25","doi-asserted-by":"publisher","first-page":"589","DOI":"10.1080\/10618600.2016.1256817","volume":"26","author":"L Mentch","year":"2017","unstructured":"Mentch, L., Hooker, G.: Formal hypothesis tests for additive structure in random forests. J. Comput. Graph. Stat. 26(3), 589\u2013597 (2017)","journal-title":"J. Comput. Graph. Stat."},{"key":"10057_CR26","unstructured":"Mentch, L., Zhou, S.: Getting better from worse: augmented bagging and a cautionary tale of variable importance. arXiv preprint arXiv:2003.03629 (2020)"},{"key":"10057_CR27","unstructured":"Nelsen, R.B.: An Introduction to Copulas. Springer (2007)"},{"issue":"1","key":"10057_CR28","doi-asserted-by":"publisher","first-page":"110","DOI":"10.1186\/1471-2105-11-110","volume":"11","author":"KK Nicodemus","year":"2010","unstructured":"Nicodemus, K.K., Malley, J.D., Strobl, C., Ziegler, A.: The behaviour of random forest permutation-based variable importance measures under predictor correlation. BMC Bioinform. 11(1), 110 (2010)","journal-title":"BMC Bioinform."},{"key":"10057_CR29","doi-asserted-by":"crossref","unstructured":"Owen, A.B.: Sobol\u2019indices and shapley value. SIAM\/ASA J. Uncertain. Quantif. 2(1), 245\u2013251 (2014)","DOI":"10.1137\/130936233"},{"key":"10057_CR30","doi-asserted-by":"crossref","unstructured":"Ribeiro, M.T., Singh, S., Guestrin, C.: Why should i trust you? Explaining the predictions of any classifier. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 1135\u20131144. ACM (2016)","DOI":"10.1145\/2939672.2939778"},{"key":"10057_CR31","unstructured":"Roosen, C.B.: Visualization and Exploration of High-dimensional Functions Using the Functional ANOVA Decomposition. Ph. D. thesis, Stanford University (1995)"},{"issue":"5","key":"10057_CR32","doi-asserted-by":"publisher","first-page":"206","DOI":"10.1038\/s42256-019-0048-x","volume":"1","author":"C Rudin","year":"2019","unstructured":"Rudin, C.: Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat. Mach. Intell. 1(5), 206\u2013215 (2019)","journal-title":"Nat. Mach. Intell."},{"key":"10057_CR33","unstructured":"Simonyan, K., Vedaldi, A., Zisserman, A.: Deep inside convolutional networks: visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034 (2013)"},{"key":"10057_CR34","doi-asserted-by":"crossref","unstructured":"Slack, D., Hilgard, S., Jia, E., Singh, S., Lakkaraju, H.: Fooling lime and shap: adversarial attacks on post hoc explanation methods. In: Proceedings of the AAAI\/ACM Conference on AI, Ethics, and Society, pp. 180\u2013186 (2020)","DOI":"10.1145\/3375627.3375830"},{"key":"10057_CR35","first-page":"407","volume":"1","author":"IM Sobol","year":"1993","unstructured":"Sobol, I.M.: Sensitivity analysis for non-linear mathematical models. Math. Model. Comput. Exp. 1, 407\u2013414 (1993)","journal-title":"Math. Model. Comput. Exp."},{"issue":"1","key":"10057_CR36","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1186\/1471-2105-8-25","volume":"8","author":"C Strobl","year":"2007","unstructured":"Strobl, C., Boulesteix, A.-L., Zeileis, A., Hothorn, T.: Bias in random forest variable importance measures: illustrations, sources and a solution. BMC Bioinform. 8(1), 25 (2007)","journal-title":"BMC Bioinform."},{"issue":"1","key":"10057_CR37","doi-asserted-by":"publisher","first-page":"307","DOI":"10.1186\/1471-2105-9-307","volume":"9","author":"C Strobl","year":"2008","unstructured":"Strobl, C., Boulesteix, A.-L., Kneib, T., Augustin, T., Zeileis, A.: Conditional variable importance for random forests. BMC Bioinform. 9(1), 307 (2008)","journal-title":"BMC Bioinform."},{"key":"10057_CR38","unstructured":"Tan, S., Caruana, R., Hooker, G., Koch, P., Gordo, A.: Learning global additive explanations for neural nets using model distillation. arXiv preprint arXiv:1801.08640 (2018)"},{"key":"10057_CR39","doi-asserted-by":"crossref","unstructured":"Tan, S., Caruana, R., Hooker, G., Lou, Y.: Distill-and-compare: auditing black-box models using transparent model distillation. In: Proceedings of the 2018 AAAI\/ACM Conference on AI, Ethics, and Society, pp. 303\u2013310 (2018)","DOI":"10.1145\/3278721.3278725"},{"issue":"14","key":"10057_CR40","doi-asserted-by":"publisher","first-page":"1986","DOI":"10.1093\/bioinformatics\/btr300","volume":"27","author":"L Tolo\u015fi","year":"2011","unstructured":"Tolo\u015fi, L., Lengauer, T.: Classification with correlated features: unreliability of feature ranking and solutions. Bioinformatics 27(14), 1986\u20131994 (2011)","journal-title":"Bioinformatics"},{"key":"10057_CR41","first-page":"1341","volume":"10","author":"E Tuv","year":"2009","unstructured":"Tuv, E., Borisov, A., Runger, G., Torkkola, K.: Feature selection with ensembles, artificial variables, and redundancy elimination. J. Mach. Learn. Res. 10, 1341\u20131366 (2009)","journal-title":"J. Mach. Learn. Res."},{"key":"10057_CR42","first-page":"841","volume":"31","author":"S Wachter","year":"2017","unstructured":"Wachter, S., Mittelstadt, B., Russell, C.: Counterfactual explanations without opening the black box: automated decisions and the gdpr. Harv. J. Law Technol. 31, 841 (2017)","journal-title":"Harv. J. Law Technol."},{"key":"10057_CR43","doi-asserted-by":"crossref","unstructured":"Williamson, B.D., Gilbert, P.B., Simon, N.R., Carone, M.: A unified approach for inference on algorithm-agnostic variable importance. arXiv preprint arXiv:2004.03683 (2020)","DOI":"10.1080\/01621459.2021.2003200"},{"key":"10057_CR44","doi-asserted-by":"crossref","unstructured":"Wood, S.N.: Generalized Additive Models: An Introduction with R. Chapman and Hall\/CRC (2006)","DOI":"10.1201\/9781420010404"},{"issue":"477","key":"10057_CR45","doi-asserted-by":"publisher","first-page":"235","DOI":"10.1198\/016214506000000843","volume":"102","author":"Y Wu","year":"2007","unstructured":"Wu, Y., Boos, D.D., Stefanski, L.A.: Controlling variable selection by the addition of pseudovariables. J. Am. Stat. Assoc. 102(477), 235\u2013243 (2007)","journal-title":"J. Am. Stat. Assoc."},{"key":"10057_CR46","unstructured":"Zhou, Z., Hooker, G.: Unbiased measurement of feature importance in tree-based methods. arXiv preprint arXiv:1903.05179 (2019)"}],"container-title":["Statistics and Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11222-021-10057-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11222-021-10057-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11222-021-10057-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,11,12]],"date-time":"2021-11-12T09:28:24Z","timestamp":1636709304000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11222-021-10057-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,29]]},"references-count":46,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2021,11]]}},"alternative-id":["10057"],"URL":"https:\/\/doi.org\/10.1007\/s11222-021-10057-z","relation":{},"ISSN":["0960-3174","1573-1375"],"issn-type":[{"value":"0960-3174","type":"print"},{"value":"1573-1375","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,10,29]]},"assertion":[{"value":"10 February 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 October 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 October 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"82"}}