{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,23]],"date-time":"2026-03-23T17:14:02Z","timestamp":1774286042102,"version":"3.50.1"},"reference-count":37,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2024,8,29]],"date-time":"2024-08-29T00:00:00Z","timestamp":1724889600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,8,29]],"date-time":"2024-08-29T00:00:00Z","timestamp":1724889600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Data Min Knowl Disc"],"published-print":{"date-parts":[[2024,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In order to trust the predictions of a machine learning algorithm, it is necessary to understand the factors that contribute to those predictions. In the case of probabilistic and uncertainty-aware models, it is necessary to understand not only the reasons for the predictions themselves, but also the reasons for the model\u2019s level of confidence in those predictions. In this paper, we show how existing methods in explainability can be extended to uncertainty-aware models and how such extensions can be used to understand the sources of uncertainty in a model\u2019s predictive distribution. In particular, by adapting permutation feature importance, partial dependence plots, and individual conditional expectation plots, we demonstrate that novel insights into model behaviour may be obtained and that these methods can be used to measure the impact of features on both the entropy of the predictive distribution and the log-likelihood of the ground truth labels under that distribution. With experiments using both synthetic and real-world data, we demonstrate the utility of these approaches to understand both the sources of uncertainty and their impact on model performance.<\/jats:p>","DOI":"10.1007\/s10618-024-01070-7","type":"journal-article","created":{"date-parts":[[2024,8,29]],"date-time":"2024-08-29T13:02:20Z","timestamp":1724936540000},"page":"4184-4216","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":29,"title":["Model-agnostic variable importance for predictive uncertainty: an entropy-based approach"],"prefix":"10.1007","volume":"38","author":[{"given":"Danny","family":"Wood","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9689-543X","authenticated-orcid":false,"given":"Theodore","family":"Papamarkou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Matt","family":"Benatan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Richard","family":"Allmendinger","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,8,29]]},"reference":[{"key":"1070_CR1","unstructured":"Antoran J, Bhatt U, Adel T, et al (2021) Getting a CLUE: a method for explaining uncertainty estimates. In: International conference on learning representations"},{"key":"1070_CR2","unstructured":"Blundell C, Cornebise J, Kavukcuoglu K, et al (2015) Weight uncertainty in neural networks. In: International conference on machine learning"},{"key":"1070_CR3","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman L (2001) Random forests. Mach Learn 45:5\u201332. https:\/\/doi.org\/10.1023\/A:1010933404324","journal-title":"Mach Learn"},{"key":"1070_CR4","doi-asserted-by":"publisher","unstructured":"Casalicchio G, Molnar C, Bischl B (2018) Visualizing the feature importance for black box models. In: Machine learning and knowledge discovery in databases: European conference, ECML PKDD, Springer, pp 655\u2013670. https:\/\/doi.org\/10.1007\/978-3-030-10925-7_40","DOI":"10.1007\/978-3-030-10925-7_40"},{"key":"1070_CR5","unstructured":"Chai LR (2018) Uncertainty estimation in Bayesian neural networks and links to interpretability. Master\u2019s thesis, University of Cambridge"},{"key":"1070_CR6","unstructured":"Chau SL, Muandet K, Sejdinovic D (2024) Explaining the uncertain: stochastic Shapley values for gaussian process models. Adv Neural Inf Process Syst 36"},{"key":"1070_CR7","doi-asserted-by":"publisher","unstructured":"Chen H, Covert IC, Lundberg SM, et al (2023) Algorithms to estimate Shapley value feature attributions. Nat Mach Intell pp 1\u201312. https:\/\/doi.org\/10.1038\/s42256-023-00657-x","DOI":"10.1038\/s42256-023-00657-x"},{"issue":"1","key":"1070_CR8","first-page":"9477","volume":"22","author":"IC Covert","year":"2021","unstructured":"Covert IC, Lundberg S, Lee SI (2021) Explaining by removing: a unified framework for model explanation. J Mach Learn Res 22(1):9477\u20139566","journal-title":"J Mach Learn Res"},{"key":"1070_CR9","unstructured":"Depeweg S, Hern\u00e1ndez-Lobato JM, Udluft S, et al (2017) Sensitivity analysis for predictive uncertainty in Bayesian neural networks. arXiv preprint arXiv:1712.03605"},{"key":"1070_CR10","unstructured":"Depeweg S, Hernandez-Lobato JM, Doshi-Velez F, et al (2018) Decomposition of uncertainty in Bayesian deep learning for efficient and risk-sensitive learning. In: International conference on machine learning"},{"key":"1070_CR11","doi-asserted-by":"publisher","unstructured":"Friedman JH (2001) Greedy function approximation: a gradient boosting machine. Ann Stat pp 1189\u20131232. https:\/\/doi.org\/10.1214\/aos\/1013203451","DOI":"10.1214\/aos\/1013203451"},{"key":"1070_CR12","unstructured":"Gal Y, Ghahramani Z (2016) Dropout as a Bayesian approximation: representing model uncertainty in deep learning. In: International conference on machine learning, pp 1050\u20131059"},{"key":"1070_CR13","unstructured":"Gardner JR, Pleiss G, Bindel D, et al (2018) GPyTorch: Blackbox matrix-matrix Gaussian process inference with GPU acceleration. In: Advances in neural information processing systems"},{"issue":"1","key":"1070_CR14","doi-asserted-by":"publisher","first-page":"44","DOI":"10.1080\/10618600.2014.907095","volume":"24","author":"A Goldstein","year":"2015","unstructured":"Goldstein A, Kapelner A, Bleich J et al (2015) Peeking inside the black box: visualizing statistical learning with plots of individual conditional expectation. J Comput Gr Stat 24(1):44\u201365. https:\/\/doi.org\/10.1080\/10618600.2014.907095","journal-title":"J Comput Gr Stat"},{"key":"1070_CR15","unstructured":"Guo C, Pleiss G, Sun Y, et al (2017) On calibration of modern neural networks. In: International conference on machine learning, pp 1321\u20131330"},{"key":"1070_CR16","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s11222-021-10057-z","volume":"31","author":"G Hooker","year":"2021","unstructured":"Hooker G, Mentch L, Zhou S (2021) Unrestricted permutation forces extrapolation: variable importance requires at least one more model, or there is no free variable importance. Stat Comput 31:1\u201316. https:\/\/doi.org\/10.1007\/s11222-021-10057-z","journal-title":"Stat Comput"},{"key":"1070_CR17","doi-asserted-by":"publisher","DOI":"10.5772\/intechopen.93310","author":"L Kelly","year":"2020","unstructured":"Kelly L, Sachan S, Ni L et al (2020) Explainable artificial intelligence for digital forensics: opportunities, challenges and a drug testing case study. Digital Forensic Sci. https:\/\/doi.org\/10.5772\/intechopen.93310","journal-title":"Digital Forensic Sci"},{"key":"1070_CR18","unstructured":"Liu J, Paisley J, Kioumourtzoglou MA, et al (2019) Accurate uncertainty estimation and decomposition in ensemble learning. In: Advances in neural information processing systems"},{"key":"1070_CR19","unstructured":"Lundberg SM, Lee SI (2017) A unified approach to interpreting model predictions. In: Advances in neural information processing systems"},{"issue":"1","key":"1070_CR20","doi-asserted-by":"publisher","first-page":"56","DOI":"10.1038\/s42256-019-0138-9","volume":"2","author":"SM Lundberg","year":"2020","unstructured":"Lundberg SM, Erion G, Chen H et al (2020) From local explanations to global understanding with explainable AI for trees. Nat Mach Intell 2(1):56\u201367. https:\/\/doi.org\/10.1038\/s42256-019-0138-9","journal-title":"Nat Mach Intell"},{"key":"1070_CR21","unstructured":"Mease D, Wyner A (2008) Evidence contrary to the statistical view of boosting. J Mach Learn Res 9(2)"},{"key":"1070_CR22","unstructured":"Molnar C (2022) Interpretable machine learning, 2nd edn. Independently Published"},{"key":"1070_CR23","doi-asserted-by":"publisher","DOI":"10.1007\/s10618-022-00901-9","author":"C Molnar","year":"2023","unstructured":"Molnar C, K\u00f6nig G, Bischl B et al (2023) Model-agnostic feature importance and effects with dependent features: a conditional subgroup approach. Data Min Knowl Discov. https:\/\/doi.org\/10.1007\/s10618-022-00901-9","journal-title":"Data Min Knowl Discov"},{"key":"1070_CR24","unstructured":"Moosbauer J, Herbinger J, Casalicchio G, et al (2021) Explaining hyperparameter optimization via partial dependence plots. In: Advances in neural information processing systems"},{"key":"1070_CR25","doi-asserted-by":"publisher","unstructured":"Mukhoti J, Kirsch A, van Amersfoort J, et al (2023) Deep deterministic uncertainty: a new simple baseline. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, https:\/\/doi.org\/10.1109\/CVPR52729.2023.02336","DOI":"10.1109\/CVPR52729.2023.02336"},{"key":"1070_CR26","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4612-0745-0","volume-title":"Bayesian learning for neural networks","author":"RM Neal","year":"2012","unstructured":"Neal RM (2012) Bayesian learning for neural networks, vol 118. Springer Science & Business Media, Berlin. https:\/\/doi.org\/10.1007\/978-1-4612-0745-0"},{"key":"1070_CR27","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa F, Varoquaux G, Gramfort A et al (2011) Scikit-learn: machine learning in Python. J Mach Learn Res 12:2825\u20132830","journal-title":"J Mach Learn Res"},{"key":"1070_CR28","doi-asserted-by":"publisher","unstructured":"Ribeiro MT, Singh S, Guestrin C (2016) \u2018Why should I trust you?\u2019 Explaining the predictions of any classifier. In: ACM SIGKDD international conference on knowledge discovery and data mining, pp 1135\u20131144. https:\/\/doi.org\/10.1145\/2939672.2939778","DOI":"10.1145\/2939672.2939778"},{"key":"1070_CR29","doi-asserted-by":"publisher","unstructured":"Shaker MH, H\u00fcllermeier E (2020) Aleatoric and epistemic uncertainty with random forests. In: Advances in intelligent data analysis XVIII: 18th international symposium on intelligent data analysis, https:\/\/doi.org\/10.1007\/978-3-030-44584-3_35","DOI":"10.1007\/978-3-030-44584-3_35"},{"key":"1070_CR30","doi-asserted-by":"publisher","unstructured":"Slack D, Hilgard S, Jia E, et al (2020) Fooling LIME and SHAP: adversarial attacks on post hoc explanation methods. In: AAAI\/ACM conference on AI, ethics, and society, https:\/\/doi.org\/10.1145\/3375627.3375830","DOI":"10.1145\/3375627.3375830"},{"key":"1070_CR31","unstructured":"Smith JW, Everhart JE, Dickson W, et al (1988) Using the ADAP learning algorithm to forecast the onset of diabetes mellitus. In: Annual symposium on computer application in medical care, p 261"},{"key":"1070_CR32","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/1471-2105-9-307","volume":"9","author":"C Strobl","year":"2008","unstructured":"Strobl C, Boulesteix AL, Kneib T et al (2008) Conditional variable importance for random forests. BMC Bioinform 9:1\u201311. https:\/\/doi.org\/10.1186\/1471-2105-9-307","journal-title":"BMC Bioinform"},{"key":"1070_CR33","unstructured":"Watson DS, O\u2019Hara J, Tax N, et al (2023) Explaining predictive uncertainty with information theoretic Shapley values. arXiv preprint arXiv:2306.05724"},{"key":"1070_CR34","doi-asserted-by":"publisher","unstructured":"Williams CK, Rasmussen CE (2006) Gaussian processes for machine learning. 3, MIT Press Cambridge, MA, https:\/\/doi.org\/10.7551\/mitpress\/3206.001.0001","DOI":"10.7551\/mitpress\/3206.001.0001"},{"key":"1070_CR35","unstructured":"Wimmer L, Sale Y, Hofman P, et al (2023) Quantifying aleatoric and epistemic uncertainty in machine learning: are conditional entropy and mutual information appropriate measures? In: Uncertainty in artificial intelligence"},{"key":"1070_CR36","doi-asserted-by":"publisher","unstructured":"Yeh IC (2007) Concrete compressive strength. UCI machine learning repository, https:\/\/doi.org\/10.24432\/C5PK67","DOI":"10.24432\/C5PK67"},{"key":"1070_CR37","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2022.108418","volume":"243","author":"X Zhang","year":"2022","unstructured":"Zhang X, Chan FT, Mahadevan S (2022) Explainable machine learning in image classification models: an uncertainty quantification perspective. Knowl-Based Syst 243:108418. https:\/\/doi.org\/10.1016\/j.knosys.2022.108418","journal-title":"Knowl-Based Syst"}],"container-title":["Data Mining and Knowledge Discovery"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10618-024-01070-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10618-024-01070-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10618-024-01070-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,28]],"date-time":"2024-10-28T09:14:08Z","timestamp":1730106848000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10618-024-01070-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,29]]},"references-count":37,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2024,11]]}},"alternative-id":["1070"],"URL":"https:\/\/doi.org\/10.1007\/s10618-024-01070-7","relation":{},"ISSN":["1384-5810","1573-756X"],"issn-type":[{"value":"1384-5810","type":"print"},{"value":"1573-756X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,8,29]]},"assertion":[{"value":"19 October 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 August 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 August 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}