{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,11]],"date-time":"2026-05-11T15:09:15Z","timestamp":1778512155100,"version":"3.51.4"},"reference-count":29,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2025,2,5]],"date-time":"2025-02-05T00:00:00Z","timestamp":1738713600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100006769","name":"Russian Science Foundation","doi-asserted-by":"publisher","award":["24-11-00272"],"award-info":[{"award-number":["24-11-00272"]}],"id":[{"id":"10.13039\/501100006769","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>The interpretability requirement is one of the largest obstacles when deploying machine learning models in various practical fields. Methods of eXplainable Artificial Intelligence (XAI) address those issues. However, the growing number of different solutions in this field creates a demand to assess the quality of explanations and compare them. In recent years, several attempts have been made to consolidate scattered XAI quality assessment methods into a single benchmark. Those attempts usually suffered from a focus on feature importance only, a lack of customization, and the absence of an evaluation framework. In this work, the eXplainable Artificial Intelligence Benchmark (XAIB) is proposed. Compared to existing benchmarks, XAIB is more universal, extensible, and has a complete evaluation ontology in the form of the Co-12 Framework. Due to its special modular design, it is easy to add new datasets, models, explainers, and quality metrics. Furthermore, an additional abstraction layer built with an inversion of control principle makes them easier to use. The benchmark will contribute to artificial intelligence research by providing a platform for evaluation experiments and, at the same time, will contribute to engineering by providing a way to compare explainers using custom datasets and machine learning models, which brings evaluation closer to practice.<\/jats:p>","DOI":"10.3390\/a18020085","type":"journal-article","created":{"date-parts":[[2025,2,5]],"date-time":"2025-02-05T10:09:52Z","timestamp":1738750192000},"page":"85","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Open and Extensible Benchmark for Explainable Artificial Intelligence Methods"],"prefix":"10.3390","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0125-960X","authenticated-orcid":false,"given":"Ilia","family":"Moiseev","sequence":"first","affiliation":[{"name":"Faculty of Digital Transformations, ITMO University, Saint Petersburg 197101, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9138-0736","authenticated-orcid":false,"given":"Ksenia","family":"Balabaeva","sequence":"additional","affiliation":[{"name":"Faculty of Digital Transformations, ITMO University, Saint Petersburg 197101, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8828-4615","authenticated-orcid":false,"given":"Sergey","family":"Kovalchuk","sequence":"additional","affiliation":[{"name":"Faculty of Digital Transformations, ITMO University, Saint Petersburg 197101, Russia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,2,5]]},"reference":[{"key":"ref_1","unstructured":"Bryce Goodman, S.F. (2016). European union regulations on algorithmic decision-making and a \u201cright to explanation\u201d. arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Markus, A.F., Kors, J.A., and Rijnbeek, P.R. (2021). The role of explainability in creating trustworthy artificial intelligence for health care: A comprehensive survey of the terminology, design choices, and evaluation strategies. J. Biomed. Inform., 113.","DOI":"10.1016\/j.jbi.2020.103655"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Abdullah, T.A., Zahid, M.S.M., and Ali, W. (2021). A review of interpretable ML in healthcare: Taxonomy, applications, challenges, and future directions. Symmetry, 13.","DOI":"10.3390\/sym13122439"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Molnar, C., Casalicchio, G., and Bischl, B. (2021). Interpretable machine learning\u2014A brief history, state-of-the-art and challenges. Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Springer.","DOI":"10.1007\/978-3-030-65965-3_28"},{"key":"ref_5","unstructured":"Doshi-Velez, F., and Kim, B. (2017). Towards a rigorous science of interpretable machine learning. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1145\/3236386.3241340","article-title":"The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery","volume":"16","author":"Lipton","year":"2018","journal-title":"Queue"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"110273","DOI":"10.1016\/j.knosys.2023.110273","article-title":"Explainable AI (XAI): A systematic meta-survey of current challenges and future opportunities","volume":"263","author":"Saeed","year":"2023","journal-title":"Knowl.-Based Syst."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3583558","article-title":"From anecdotal evidence to quantitative evaluation methods: A systematic review on evaluating explainable ai","volume":"55","author":"Nauta","year":"2023","journal-title":"ACM Comput. Surv."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1719","DOI":"10.1007\/s10618-023-00933-9","article-title":"Benchmarking and survey of explanation methods for black box models","volume":"37","author":"Bodria","year":"2023","journal-title":"Data Min. Knowl. Discov."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"141","DOI":"10.1016\/j.dss.2010.12.003","article-title":"An empirical evaluation of the comprehensibility of decision table, tree and rule based predictive models","volume":"51","author":"Huysmans","year":"2011","journal-title":"Decis. Support Syst."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Kulesza, T., Stumpf, S., Burnett, M., Yang, S., Kwan, I., and Wong, W.K. (2013, January 15\u201319). Too much, too little, or just right? Ways explanations impact end users\u2019 mental models. Proceedings of the 2013 IEEE Symposium on Visual Languages and Human Centric Computing, San Jose, CA, USA.","DOI":"10.1109\/VLHCC.2013.6645235"},{"key":"ref_12","first-page":"9525","article-title":"Sanity checks for saliency maps","volume":"31","author":"Adebayo","year":"2018","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Hase, P., and Bansal, M. (2020). Evaluating explainable AI: Which algorithmic explanations help users predict model behavior?. arXiv.","DOI":"10.18653\/v1\/2020.acl-main.491"},{"key":"ref_14","unstructured":"Zhang, H., Chen, J., Xue, H., and Zhang, Q. (2019). Towards a unified evaluation of explanation methods without ground truth. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Bhatt, U., Weller, A., and Moura, J.M. (2020). Evaluating and aggregating feature-based model explanations. arXiv.","DOI":"10.24963\/ijcai.2020\/417"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Sokol, K., and Flach, P. (2020, January 27\u201330). Explainability fact sheets: A framework for systematic assessment of explainable approaches. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, Barcelona, Spain.","DOI":"10.1145\/3351095.3372870"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Atanasova, P., Simonsen, J.G., Lioma, C., and Augenstein, I. (2020). A diagnostic study of explainability techniques for text classification. arXiv.","DOI":"10.18653\/v1\/2020.emnlp-main.263"},{"key":"ref_18","unstructured":"Liu, Y., Khandagale, S., White, C., and Neiswanger, W. (2021). Synthetic benchmarks for scientific research in explainable machine learning. arXiv."},{"key":"ref_19","first-page":"15784","article-title":"Openxai: Towards a transparent evaluation of model explanations","volume":"35","author":"Agarwal","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_20","unstructured":"Belaid, M.K., H\u00fcllermeier, E., Rabus, M., and Krestel, R. (2022). Do We Need Another Explainable AI Method? Toward Unifying Post-hoc XAI Evaluation Methods into an Interactive and Multi-dimensional Benchmark. arXiv."},{"key":"ref_21","unstructured":"Li, X., Du, M., Chen, J., Chai, Y., Lakkaraju, H., and Xiong, H. (2023, January 10\u201316). M4: A Unified XAI Benchmark for Faithfulness Evaluation of Feature Attribution Methods across Metrics, Modalities and Models. Proceedings of the NeurIPS, New Orleans, LA, USA."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"103821","DOI":"10.1016\/j.autcon.2021.103821","article-title":"An engineer\u2019s guide to eXplainable Artificial Intelligence and Interpretable Machine Learning: Navigating causality, forced goodness, and the false perception of inference","volume":"129","author":"Naser","year":"2021","journal-title":"Autom. Constr."},{"key":"ref_23","unstructured":"Scikit Learn (2024, October 29). Toy Datasets. Available online: https:\/\/scikit-learn.org\/stable\/datasets\/toy_dataset.html."},{"key":"ref_24","unstructured":"Wolberg, W., Mangasarian, O., Street, N., and Street, W. (2024, October 28). Wisconsin Diagnostic Breast Cancer Database; UCI Machine Learning Repository. Available online: https:\/\/archive.ics.uci.edu\/dataset\/17\/breast+cancer+wisconsin+diagnostic."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Garris, M.D., Blue, J.L., Candela, G.T., Grother, P.J., Janet, S., and Wilson, C.L. (1997). NIST Form-Based Handprint Recognition System, National Institute of Standards and Technology.","DOI":"10.6028\/NIST.IR.5959"},{"key":"ref_26","unstructured":"Cortez, P., Cerdeira, A., Almeida, F., Matos, J., and Reis, J. (2024, October 28). Wine Quality. UCI Machine Learning Repository. Available online: https:\/\/archive.ics.uci.edu\/dataset\/186\/wine+quality."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"179","DOI":"10.1111\/j.1469-1809.1936.tb02137.x","article-title":"The use of multiple measurements in taxonomic problems","volume":"7","author":"Fisher","year":"1936","journal-title":"Ann. Eugen."},{"key":"ref_28","first-page":"4768","article-title":"A unified approach to interpreting model predictions","volume":"30","author":"Lundberg","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Ribeiro, M.T., Singh, S., and Guestrin, C. (2016, January 13\u201317). \u201cWhy should i trust you?\u201d Explaining the predictions of any classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA.","DOI":"10.1145\/2939672.2939778"}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/18\/2\/85\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T16:27:22Z","timestamp":1760027242000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/18\/2\/85"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,5]]},"references-count":29,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,2]]}},"alternative-id":["a18020085"],"URL":"https:\/\/doi.org\/10.3390\/a18020085","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,5]]}}}