{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T11:53:18Z","timestamp":1780401198097,"version":"3.54.1"},"reference-count":55,"publisher":"Springer Science and Business Media LLC","issue":"9","license":[{"start":{"date-parts":[[2025,8,5]],"date-time":"2025-08-05T00:00:00Z","timestamp":1754352000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,8,5]],"date-time":"2025-08-05T00:00:00Z","timestamp":1754352000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Flanders AI Research Program"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2025,9]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>The accurate representation of epistemic uncertainty is a challenging yet essential task in machine learning. A widely used representation corresponds to convex sets of probabilistic predictors, also known as credal sets. One popular way of constructing these credal sets is via ensembling or specialized supervised learning methods, where the epistemic uncertainty can be quantified through measures such as the set size or the disagreement among members. In principle, these sets should contain the true data-generating distribution. As a necessary condition for this validity, we adopt the strongest notion of calibration as a proxy. Concretely, we propose a novel statistical test to determine whether there is a convex combination of the set\u2019s predictions that is calibrated in distribution. In contrast to previous methods, our framework allows the convex combination to be instance-dependent, recognizing that different ensemble members may be better calibrated in different regions of the input space. Moreover, we learn this combination via proper scoring rules, which inherently optimize for calibration. Building on differentiable, kernel-based estimators of calibration errors, we introduce a nonparametric testing procedure and demonstrate the benefits of capturing instance-level variability on synthetic and real-world experiments.<\/jats:p>","DOI":"10.1007\/s10994-025-06844-8","type":"journal-article","created":{"date-parts":[[2025,8,5]],"date-time":"2025-08-05T18:41:53Z","timestamp":1754419313000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["A calibration test for evaluating set-based epistemic uncertainty representations"],"prefix":"10.1007","volume":"114","author":[{"given":"Mira","family":"J\u00fcrgens","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thomas","family":"Mortier","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eyke","family":"H\u00fcllermeier","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Viktor","family":"Bengs","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Willem","family":"Waegeman","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,8,5]]},"reference":[{"key":"6844_CR1","unstructured":"Abe, T., Buchanan, E. K., Pleiss, G., Zemel, R., & Cunningham, J. P. (2022). Deep ensembles work, but are they necessary? In Proceedings of NeurIPS, 35th advances in neural information processing systems."},{"key":"6844_CR2","unstructured":"Bengs, V., H\u00fcllermeier, E., & Waegeman, W. (2022). Pitfalls of epistemic uncertainty quantification through loss minimisation. In Proceedings of NeurIPS, 35th advances in neural information processing systems."},{"key":"6844_CR3","unstructured":"Bengs, V., H\u00fcllermeier, E., & Waegeman, W. (2023). On second-order scoring rules for epistemic uncertainty quantification. In Proceedings of ICML, 40th international conference on machine learning."},{"issue":"1","key":"6844_CR4","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1175\/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2","volume":"78","author":"GW Brier","year":"1950","unstructured":"Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1.","journal-title":"Monthly Weather Review"},{"key":"6844_CR5","doi-asserted-by":"publisher","first-page":"1512","DOI":"10.1002\/qj.456","volume":"135","author":"J Br\u00f6cker","year":"2009","unstructured":"Br\u00f6cker, J. (2009). Reliability, sufficiency, and the decomposition of proper scores. Quarterly Journal of the Royal Meteorological Society, 135, 1512\u20131519.","journal-title":"Quarterly Journal of the Royal Meteorological Society"},{"key":"6844_CR6","doi-asserted-by":"publisher","first-page":"651","DOI":"10.1175\/WAF993.1","volume":"22","author":"J Br\u00f6cker","year":"2007","unstructured":"Br\u00f6cker, J., & Smith, L. A. (2007). Increasing the reliability of reliability diagrams. Weather and Forecasting, 22, 651\u2013661.","journal-title":"Weather and Forecasting"},{"key":"6844_CR7","unstructured":"Caprio, M., Dutta, S., Jang, K. J., Lin, V., Ivanov, R., Sokolsky, O., & Lee, I. (2024). Credal Bayesian deep learning. Transactions on Machine Learning Research."},{"key":"6844_CR8","unstructured":"Chau, S. L., Schrab, A., Gretton, A., Sejdinovic, D., & Muandet, K. (2024). Credal two-sample tests of epistemic ignorance. arXiv preprint arXiv:2410.12921."},{"issue":"2","key":"6844_CR9","doi-asserted-by":"publisher","first-page":"199","DOI":"10.1016\/S0004-3702(00)00029-1","volume":"120","author":"FG Cozman","year":"2000","unstructured":"Cozman, F. G. (2000). Credal networks. Artificial Intelligence, 120(2), 199\u2013233.","journal-title":"Artificial Intelligence"},{"key":"6844_CR10","unstructured":"D\u00fcmbgen, L. (2017). Empirische Prozesse. Institut f\u00fcr Mathematische Statistik und Versicherungsmathematik der Universit\u00e4t Bern, Bern, Switzerland. https:\/\/books.google.de\/books?id=tmFMygEACAAJ."},{"key":"6844_CR11","unstructured":"Gal, Y., & Ghahramani, Z. (2016). Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of ICML, 32nd international conference on machine learning."},{"key":"6844_CR12","volume-title":"Bayesian data analysis","author":"A Gelman","year":"2004","unstructured":"Gelman, A., Carlin, J. B., Stern, H. S., & Rubin, D. B. (2004). Bayesian data analysis (2nd ed.). Boca Raton: Chapman and Hall.","edition":"2"},{"key":"6844_CR13","doi-asserted-by":"publisher","first-page":"359","DOI":"10.1198\/016214506000001437","volume":"102","author":"T Gneiting","year":"2007","unstructured":"Gneiting, T., & Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102, 359\u2013378.","journal-title":"Journal of the American Statistical Association"},{"key":"6844_CR14","unstructured":"Gruber, S., & Buettner, F. (2022). Better uncertainty calibration via proper scores for classification and beyond. In Proceedings of NeurIPS, 35th advances in neural information processing systems."},{"key":"6844_CR15","unstructured":"Gruber, C., Schenk, P. O., Schierholz, M., Kreuter, F., & Kauermann, G. (2023). Sources of uncertainty in machine learning\u2014a statisticians\u2019 view. arXiv preprint arXiv:2305.16703."},{"key":"6844_CR16","unstructured":"Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of ICML, 34th international conference on machine learning."},{"key":"6844_CR17","unstructured":"Gupta, K., Rahimi, A., Ajanthan, T., Mensink, T., Sminchisescu, C., & Hartley, R. (2020). Calibration of neural networks using splines. In Proceedings of ICML, international conference on learning representations."},{"key":"6844_CR18","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770\u2013778).","DOI":"10.1109\/CVPR.2016.90"},{"key":"6844_CR19","doi-asserted-by":"publisher","first-page":"965","DOI":"10.1002\/(SICI)1097-0258(19970515)16:9<965::AID-SIM509>3.0.CO;2-O","volume":"16","author":"DW Hosmer","year":"1997","unstructured":"Hosmer, D. W., Hosmer, T., Le Cessie, S., & Lemeshow, S. (1997). A comparison of goodness-of-fit tests for the logistic regression model. Statistics in Medicine, 16, 965\u2013980.","journal-title":"Statistics in Medicine"},{"key":"6844_CR20","unstructured":"H\u00fcllermeier, E., Destercke, S., & Shaker, M. H. (2022). Quantification of credal uncertainty in machine learning: A critical analysis and empirical comparison. In Proceedings of UAI, 34th conference on uncertainty in artificial intelligence."},{"key":"6844_CR21","doi-asserted-by":"publisher","first-page":"457","DOI":"10.1007\/s10994-021-05946-3","volume":"110","author":"E H\u00fcllermeier","year":"2021","unstructured":"H\u00fcllermeier, E., & Waegeman, W. (2021). Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods. Machine Learning, 110, 457\u2013506.","journal-title":"Machine Learning"},{"key":"6844_CR22","unstructured":"Javanmardi, A., Stutz, D., & H\u00fcllermeier, E. (2024). Conformalized credal set predictors. In Proceedings of NeurIPS, 37th advances in neural information processing systems."},{"key":"6844_CR23","unstructured":"Juergens, M., Meinert, N., Bengs, V., H\u00fcllermeier, E., & Waegeman, W. (2024). Is epistemic uncertainty faithfully represented by evidential deep learning methods? In Proceedings of ICML, 41st international conference on machine learning."},{"key":"6844_CR24","unstructured":"Kendall, A., & Gal, Y. (2017). What uncertainties do we need in bayesian deep learning for computer vision? In Proceedings of NeurIPS, 30th advances in neural information processing systems."},{"key":"6844_CR25","unstructured":"Kopetzki, A.-K., Charpentier, B., Z\u00fcgner, D., Giri, S., & G\u00fcnnemann, S. (2021). Evaluating robustness of predictive uncertainty estimation: Are Dirichlet-based models reliable? In Proceedings of ICML, 38th international conference on machine learning."},{"key":"6844_CR26","unstructured":"Krizhevsky, A., & Hinton, G. (2009). Learning multiple layers of features from tiny images."},{"key":"6844_CR27","doi-asserted-by":"crossref","unstructured":"Kull, M., & Flach, P. (2015). Novel decompositions of proper scoring rules for classification: Score adjustment as precursor to calibration. In Machine learning and knowledge discovery in databases.","DOI":"10.1007\/978-3-319-23528-8_5"},{"key":"6844_CR28","unstructured":"Kumar, A., Sarawagi, S., & Jain, U. (2018). Trainable calibration measures for neural networks from kernel mean embeddings. In Proceedings of ICML, 35th international conference on machine learning."},{"key":"6844_CR29","unstructured":"Lakshminarayanan, B., Pritzel, A., & Blundell, C. (2017). Simple and scalable predictive uncertainty estimation using deep ensembles. In Proceedings of NeurIPS, 30th advances in neural information processing systems."},{"key":"6844_CR30","unstructured":"L\u00f6hr, T., Hofman, P., Mohr, F., & H\u00fcllermeier, E. (2025). Credal prediction based on relative likelihood. arXiv preprint arXiv:2505.22332."},{"key":"6844_CR31","unstructured":"Luo, R., Bhatnagar, A., Bai, Y., Zhao, S., Wang, H., Xiong, C., Savarese, S., Ermon, S., Schmerling, E., & Pavone, M. (2022). Local calibration: Metrics and recalibration. In Proceedings of UAI, uncertainty in artificial intelligence."},{"key":"6844_CR32","unstructured":"Marx, C., Zalouk, S., & Ermon, S. (2024). Calibration by distribution matching: Trainable kernel calibration metrics. In Proceedings of NeurIPS, 36th advances in neural information processing systems."},{"key":"6844_CR33","doi-asserted-by":"crossref","unstructured":"Meinert, N., Gawlikowski, J., & Lavin, A. (2023). The unreasonable effectiveness of deep evidential regression. In Proceedings of AAAI, 37th proceedings of the AAAI conference on artificial intelligence.","DOI":"10.1609\/aaai.v37i8.26096"},{"key":"6844_CR34","unstructured":"Mortier, T., Bengs, V., H\u00fcllermeier, E., Luca, S., & Waegeman, W. (2023). On the calibration of probabilistic classifier sets. In Proceedings of the 26th international conference on artificial intelligence and statistics."},{"key":"6844_CR35","doi-asserted-by":"publisher","first-page":"429","DOI":"10.2307\/1428011","volume":"29","author":"A M\u00fcller","year":"1997","unstructured":"M\u00fcller, A. (1997). Integral probability metrics and their generating classes of functions. Advances in Applied Probability, 29, 429\u2013443.","journal-title":"Advances in Applied Probability"},{"issue":"4","key":"6844_CR36","doi-asserted-by":"publisher","first-page":"595","DOI":"10.1175\/1520-0450(1973)012<0595:ANVPOT>2.0.CO;2","volume":"12","author":"AH Murphy","year":"1973","unstructured":"Murphy, A. H. (1973). A new vector partition of the probability score. Journal of Applied Meteorology (1962\u20131982), 12(4), 595\u2013600.","journal-title":"Journal of Applied Meteorology (1962\u20131982)"},{"key":"6844_CR37","unstructured":"Naeini, M. P., Cooper, G. F., & Hauskrecht, M. (2015). Obtaining well calibrated probabilities using Bayesian binning. In Proceedings of AAAI, proceedings of the twenty-ninth AAAI conference on artificial intelligence."},{"key":"6844_CR38","doi-asserted-by":"publisher","first-page":"89","DOI":"10.1007\/s10994-021-06003-9","volume":"111","author":"V-L Nguyen","year":"2022","unstructured":"Nguyen, V.-L., Shaker, M. H., & H\u00fcllermeier, E. (2022). How to measure uncertainty in uncertainty sampling for active learning. Machine Learning, 111, 89\u2013122.","journal-title":"Machine Learning"},{"key":"6844_CR39","doi-asserted-by":"crossref","unstructured":"Niculescu-Mizil, A., & Caruana, R. (2005). Predicting good probabilities with supervised learning. In Proceedings of ICML, 22nd international conference on machine learning.","DOI":"10.1145\/1102351.1102430"},{"key":"6844_CR40","unstructured":"Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J., Lakshminarayanan, B., & Snoek, J. (2019). Can you trust your model\u2019s uncertainty? Evaluating predictive uncertainty under dataset shift. In Proceedings of NeurIPS, 31st advances in neural information processing systems."},{"key":"6844_CR41","unstructured":"Popordanoska, T., Gregor Gruber, S., Tiulpin, A., Buettner, F., & Blaschko, M. B. (2024). Consistent and asymptotically unbiased estimation of proper calibration errors. In Proceedings of the 27th international conference on artificial intelligence and statistics."},{"key":"6844_CR42","unstructured":"Popordanoska, T., Sayer, R., & Blaschko, M. (2022). A consistent and differentiable lp canonical calibration error estimator."},{"key":"6844_CR43","unstructured":"Rahimi, A., Shaban, A., Cheng, C.-A., Hartley, R., & Boots, B. (2020). Intra order-preserving functions for calibration of multi-class neural networks. In Proceedings of NeurIPS, 33rd advances in neural information processing systems."},{"key":"6844_CR44","unstructured":"Sale, Y., Caprio, M., & H\u00fcllermeier, E. (2023). Is the volume of a credal set a good measure for epistemic uncertainty? In Proceedings of UAI, 39th conference on uncertainty in artificial intelligence."},{"key":"6844_CR45","doi-asserted-by":"publisher","first-page":"16","DOI":"10.1016\/j.ins.2013.07.030","volume":"255","author":"R Senge","year":"2014","unstructured":"Senge, R., B\u00f6sner, S., Dembczy\u0144ski, K., Haasenritter, J., Hirsch, O., Donner-Banzhoff, N., & H\u00fcllermeier, E. (2014). Reliable classification: Learning classifiers that distinguish aleatoric and epistemic uncertainty. Information Sciences, 255, 16\u201329.","journal-title":"Information Sciences"},{"key":"6844_CR46","unstructured":"Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition."},{"key":"6844_CR47","unstructured":"Ulmer, D., Hardmeier, C., & Frellsen, J. (2023). Prior and posterior networks: A survey on evidential deep learning methods for uncertainty estimation. Transactions on Machine Learning Research."},{"key":"6844_CR48","volume-title":"Asymptotic statistics","author":"AW Vaart","year":"2000","unstructured":"Vaart, A. W. (2000). Asymptotic statistics. Cambridge: Cambridge University Press."},{"key":"6844_CR49","unstructured":"Vaicenavicius, J., Widmann, D., Andersson, C., Lindsten, F., Roll, J., & Sch\u00f6n, T. (2019). Evaluating model calibration in classification. In Proceedings of the twenty-second international conference on artificial intelligence and statistics."},{"key":"6844_CR50","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4899-3472-7","volume-title":"Statistical reasoning with imprecise probabilities","author":"P Walley","year":"1991","unstructured":"Walley, P. (1991). Statistical reasoning with imprecise probabilities. London: Chapman and Hall."},{"key":"6844_CR51","doi-asserted-by":"crossref","unstructured":"Wang, K., Shariatmadar, K., Manchingal, S. K., Cuzzolin, F., Moens, D., & Hallez, H. (2025). Creinns: credal-set interval neural networks for uncertainty estimation in classification tasks. In Neural networks.","DOI":"10.1016\/j.neunet.2025.107198"},{"key":"6844_CR52","unstructured":"Widmann, D., Lindsten, F., & Zachariah, D. (2019). Calibration tests in multi-class classification: A unifying framework. In Proceedings of NeurIPS, 32nd advances in neural information processing systems"},{"key":"6844_CR53","doi-asserted-by":"publisher","first-page":"241","DOI":"10.1016\/S0893-6080(05)80023-1","volume":"5","author":"DH Wolpert","year":"1992","unstructured":"Wolpert, D. H. (1992). Stacked generalization. Neural Networks, 5, 241\u2013259.","journal-title":"Neural Networks"},{"key":"6844_CR54","unstructured":"Zadrozny, B., & Elkan, C. (2001). Obtaining calibrated probability estimates from decision trees and naive Bayesian classifiers. In Proceedings of ICML, 18th international conference on machine learning."},{"key":"6844_CR55","doi-asserted-by":"crossref","unstructured":"Zadrozny, B., & Elkan, C. (2002). Transforming classifier scores into accurate multiclass probability estimates. In Proceedings of ACM SIGKDD, 8th international conference on knowledge discovery and data mining.","DOI":"10.1145\/775107.775151"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06844-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-025-06844-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06844-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,8]],"date-time":"2025-09-08T19:07:39Z","timestamp":1757358459000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-025-06844-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,5]]},"references-count":55,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2025,9]]}},"alternative-id":["6844"],"URL":"https:\/\/doi.org\/10.1007\/s10994-025-06844-8","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,5]]},"assertion":[{"value":"9 June 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 July 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 July 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 August 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"202"}}