{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T11:05:05Z","timestamp":1785409505617,"version":"3.56.0"},"reference-count":43,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T00:00:00Z","timestamp":1782691200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T00:00:00Z","timestamp":1782691200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100006256","name":"Universit\u00e0 degli Studi dell\u2019Aquila","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100006256","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2026,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Machine Unlearning (MU), the process of removing specific data influences from trained machine learning models, is critical for regulatory compliance (e.g., GDPR\u2019s right to be forgotten) and for addressing copyright and privacy concerns in large-scale models. While a wide range of methods and metrics have been proposed, systematic evaluations remain fragmented, typically limited in scope by modality, metric coverage, or the number of methods considered. Moreover, the lack of standardized benchmarks leaves several gaps in evaluation protocols, including how to efficiently compare methods, identify optimal hyperparameters, and determine which experimental settings are appropriate for fair and meaningful benchmarking. To address these gaps, we present the most comprehensive MU benchmark to date, evaluating 12 unlearning methods across 8 classification datasets, 4 modalities, several hyperparameters and settings. Based on previous literature and our empirical results, we formalize evaluation protocol desiderata to guide future MU benchmarking. Following these guidelines, we report benchmark results highlighting the best methods within and across domains. To help with method comparison, we also introduce LUMA, a unified metric that aggregates core unlearning dimensions into a single score. Our code is reproducible and extensible to serve as a benchmark for MU research.<\/jats:p>","DOI":"10.1007\/s10994-026-07094-y","type":"journal-article","created":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T10:21:04Z","timestamp":1782728464000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["On the Evaluation of Machine Unlearning Methods: A Multi-domain Classification Benchmark"],"prefix":"10.1007","volume":"115","author":[{"given":"Andrea","family":"D\u2019Angelo","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Claudio","family":"Savelli","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Flavio","family":"Giobergia","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Elena","family":"Baralis","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Giovanni","family":"Stilo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,6,29]]},"reference":[{"key":"7094_CR1","doi-asserted-by":"publisher","unstructured":"Becker, B., & Kohavi, R. (1996). Adult. UCI Machine Learning Repository. https:\/\/doi.org\/10.24432\/C5XW20","DOI":"10.24432\/C5XW20"},{"issue":"1","key":"7094_CR2","first-page":"152","volume":"17","author":"A Benavoli","year":"2016","unstructured":"Benavoli, A., Corani, G., & Mangili, F. (2016). Should we really use post-hoc tests based on mean-ranks? Journal of Machine Learning Research, 17(1), 152\u2013161.","journal-title":"Journal of Machine Learning Research"},{"issue":"3","key":"7094_CR3","doi-asserted-by":"publisher","first-page":"90","DOI":"10.1007\/s10462-024-11078-6","volume":"58","author":"A Blanco-Justicia","year":"2025","unstructured":"Blanco-Justicia, A., Jebreel, N., Manzanares-Salor, B., S\u00e1nchez, D., Domingo-Ferrer, J., Collell, G., & Eeik Tan, K. (2025). Digital forgetting in large language models: A survey of unlearning methods. Artificial Intelligence Review, 58(3), 90.","journal-title":"Artificial Intelligence Review"},{"key":"7094_CR4","doi-asserted-by":"crossref","unstructured":"Cadet, X. F., Borovykh, A., Malekzadeh, M., Ahmadi-Abhari, S., & Haddadi, H. (2024). Deep unlearn: Benchmarking machine unlearning arXiv:2410.01276 arXiv preprint.","DOI":"10.1109\/EuroSP63326.2025.00058"},{"key":"7094_CR5","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i10.28996","author":"S Cha","year":"2024","unstructured":"Cha, S., Cho, S., Hwang, D., Lee, H., Moon, T., & Lee, M. (2024). Learning to unlearn: instance-wise unlearning for pre-trained classifiers. AAAI Press. https:\/\/doi.org\/10.1609\/aaai.v38i10.28996","journal-title":"AAAI Press"},{"key":"7094_CR6","unstructured":"Cheng, J., & Amiri, H. (2024). Mu-bench: A multitask multimodal benchmark for machine unlearning arXiv:2406.14796 arXiv preprint."},{"key":"7094_CR7","unstructured":"Choi, D., & Na, D. (2023). Towards machine unlearning benchmarks: Forgetting the personal identities in facial recognition systems arXiv:2311.02240 arXiv preprint."},{"key":"7094_CR8","doi-asserted-by":"crossref","unstructured":"Chundawat, V. S., Tarun, A. K., Mandal, M., & Kankanhalli, M. (2023). Can bad teaching induce forgetting? unlearning in deep networks using an incompetent teacher. Proceedings of the AAAI Conference on Artificial Intelligence","DOI":"10.1609\/aaai.v37i6.25879"},{"key":"7094_CR9","doi-asserted-by":"crossref","unstructured":"Crawford, K. (2022). Atlas of AI: Power, politics, and the planetary costs of artificial intelligence. Yale University Press.","DOI":"10.12987\/9780300252392"},{"key":"7094_CR10","doi-asserted-by":"crossref","unstructured":"D\u2019Angelo, A., Savelli, C., Tagliente, G., Giobergia, F., Baralis, E., & Stilo, G. (2025a). ERASURE: A modular and extensible framework for machine unlearning. ACM","DOI":"10.1145\/3746252.3761627"},{"key":"7094_CR11","doi-asserted-by":"publisher","unstructured":"D\u2019Angelo, A., Savelli, C., Tagliente, G., Giobergia, F., Baralis, E., & Stilo, G. (2025b). How to make reproducible research in machine unlearning with ERASURE. In: International joint conferences on artificial intelligence organization. Demo Track . https:\/\/doi.org\/10.24963\/ijcai.2025\/1255","DOI":"10.24963\/ijcai.2025\/1255"},{"issue":"1","key":"7094_CR12","first-page":"1","volume":"7","author":"J Dem\u0161ar","year":"2006","unstructured":"Dem\u0161ar, J. (2006). Statistical comparisons of classifiers over multiple data sets. Journal of Machine Learning Research, 7(1), 1\u201330.","journal-title":"Journal of Machine Learning Research"},{"key":"7094_CR13","unstructured":"Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding arXiv:1810.04805."},{"key":"7094_CR14","doi-asserted-by":"crossref","unstructured":"Fan, C., Liu, J., Hero, A., & Liu, S. (2024). Challenging forgets: Unveiling the worst-case forget sets in machine unlearning. In: European conference on computer vision, pp. 278\u2013297. Springer","DOI":"10.1007\/978-3-031-72664-4_16"},{"key":"7094_CR15","unstructured":"Fan, C., Liu, J., Zhang, Y., Wong, E., Wei, D., & Liu, S. (2023). Salun: Empowering machine unlearning via gradient-based weight saliency in both image classification and generation. arXiv preprint arXiv:2310.12508"},{"key":"7094_CR16","doi-asserted-by":"crossref","unstructured":"Foster, J., Schoepf, S., & Brintrup, A. (2024). Fast machine unlearning without retraining through selective synaptic dampening. Proceedings of the AAAI conference on artificial intelligence (Vol. 38, pp. 12043\u201312051)","DOI":"10.1609\/aaai.v38i11.29092"},{"key":"7094_CR17","unstructured":"Giobergia, F. (2023). IMDb-ID Dataset. https:\/\/huggingface.co\/datasets\/fgiobergia\/imdb-id. Accessed: 2026-01-10"},{"key":"7094_CR18","unstructured":"Goel, S., Prabhu, A., Sanyal, A., Lim, S.-N., Torr, P., & Kumaraguru, P. (2022). Towards adversarial evaluations for inexact machine unlearning arXiv:2201.06640 arXiv preprint."},{"key":"7094_CR19","doi-asserted-by":"crossref","unstructured":"Golatkar, A., Achille, A., & Soatto, S. (2020). Eternal sunshine of the spotless net: Selective forgetting in deep networks. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 9304\u20139312)","DOI":"10.1109\/CVPR42600.2020.00932"},{"key":"7094_CR20","unstructured":"Grimes, K., Abidi, C., Frank, C., & Gallagher, S. (2024). Gone but not forgotten: Improved benchmarks for machine unlearning arXiv:2405.19211 arXiv preprint."},{"key":"7094_CR21","unstructured":"Grynbaum, M. M., & Mac, R. (2023). The times sues openai and microsoft over ai use of copyrighted work. The New York Times, 27."},{"key":"7094_CR22","unstructured":"Gulli, A. (2005). AG\u2019s Corpus of news articles (pp. 2026\u20132027)"},{"key":"7094_CR23","doi-asserted-by":"crossref","unstructured":"Hayes, J., Shumailov, I., Triantafillou, E., Khalifa, A., & Papernot, N. (2024). Inexact unlearning needs more careful evaluations to avoid a false sense of privacy arXiv:2403.01218 arXiv preprint.","DOI":"10.1109\/SaTML64287.2025.00034"},{"key":"7094_CR24","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) (pp. 770\u2013778)","DOI":"10.1109\/CVPR.2016.90"},{"key":"7094_CR25","doi-asserted-by":"crossref","unstructured":"Jung, D. (2025). Entun: Mitigating the forget-retain dilemma in unlearning via entropy. ICT Express.","DOI":"10.1016\/j.icte.2025.06.007"},{"key":"7094_CR26","doi-asserted-by":"publisher","unstructured":"Koudounas, A., Savelli, C., Giobergia, F., & Baralis, E. (2025). Alexa, can you forget me. Machine unlearning benchmark in spoken language understanding (pp. 2025\u20132607). https:\/\/doi.org\/10.21437\/Interspeech Interspeech 2025.","DOI":"10.21437\/Interspeech"},{"key":"7094_CR27","unstructured":"Krizhevsky, A. (2009). Learning multiple layers of features from tiny images. University of Toronto. Technical report."},{"key":"7094_CR28","doi-asserted-by":"crossref","unstructured":"Kurmanji, M., Triantafillou, P., Hayes, J., & Triantafillou, E. (2024). Towards unbounded machine unlearning. Advances in Neural Information Processing Systems 36","DOI":"10.52202\/075280-0095"},{"key":"7094_CR29","unstructured":"Lanyon, J., Finke, A., Andreou, P., & Cosma, G. (2025). On the limitation of evaluating machine unlearning using only a single training seed arXiv:abs\/2510.26714."},{"issue":"3","key":"7094_CR30","first-page":"1452","volume":"12","author":"T Le Quy","year":"2022","unstructured":"Le Quy, T., Roy, A., Iosifidis, V., Zhang, W., & Ntoutsi, E. (2022). A survey on datasets for fairness-aware machine learning. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 12(3), 1452.","journal-title":"Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery"},{"key":"7094_CR31","doi-asserted-by":"crossref","unstructured":"Liu, Z., Luo, P., Wang, X., & Tang, X. (2015). Deep learning face attributes in the wild. Proceedings of international conference on computer vision (ICCV)","DOI":"10.1109\/ICCV.2015.425"},{"key":"7094_CR32","unstructured":"Maharshipandya: Spotify tracks dataset. https:\/\/huggingface.co\/datasets\/maharshipandya\/spotify-tracks-dataset (2023)"},{"key":"7094_CR33","unstructured":"Maini, P., Feng, Z., Schwarzschild, A., Lipton, Z. C., & Kolter, J. Z. (2024). Tofu: A task of fictitious unlearning for llms arXiv:2401.06121 arXiv preprint."},{"issue":"3","key":"7094_CR34","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1016\/j.clsr.2013.03.010","volume":"29","author":"A Mantelero","year":"2013","unstructured":"Mantelero, A. (2013). The eu proposal for a general data protection regulation and the roots of the \u2018right to be forgotten\u2019. Computer Law & Security Review, 29(3), 229\u2013235.","journal-title":"Computer Law & Security Review"},{"key":"7094_CR35","unstructured":"Marrie, J., Arbel, M., Mairal, J., & Larlus, D. (2024). On good practices for task-specific distillation of large pretrained visual models"},{"issue":"24","key":"7094_CR36","doi-asserted-by":"publisher","first-page":"7428","DOI":"10.3390\/molecules26247428","volume":"26","author":"H Sakiyama","year":"2021","unstructured":"Sakiyama, H., Fukuda, M., & Okuno, Y. (2021). Prediction of blood-brain barrier penetration (bbbp) based on molecular descriptors of the free-form and in-blood-form datasets. Molecules, 26(24), 7428. https:\/\/doi.org\/10.3390\/molecules26247428","journal-title":"Molecules"},{"key":"7094_CR37","unstructured":"Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2020). DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter arXiv:1910.01108."},{"key":"7094_CR38","doi-asserted-by":"crossref","unstructured":"Savelli, C., La Quatra, M., Koudounas, A., & Giobergia, F. (2026). FAME: Fictional actors for multilingual erasure. Proceedings of the fifteenth language resources and evaluation conference european language resources association.","DOI":"10.63317\/3npbbsyj2dd7"},{"key":"7094_CR39","doi-asserted-by":"crossref","unstructured":"Tarun, A. K., Chundawat, V. S., Mandal, M., & Kankanhalli, M. (2023). Fast yet effective machine unlearning. IEEE transactions on neural networks and learning systems","DOI":"10.1109\/TIFS.2023.3265506"},{"issue":"2","key":"7094_CR40","doi-asserted-by":"publisher","first-page":"513","DOI":"10.1039\/C7SC02664A","volume":"9","author":"Z Wu","year":"2018","unstructured":"Wu, Z., Ramsundar, B., Feinberg, E. N., Gomes, J., Geniesse, C., Pappu, A. S., Leswing, K., & Pande, V. (2018). Moleculenet: a benchmark for molecular machine learning. Chemical Science, 9(2), 513\u2013530. https:\/\/doi.org\/10.1039\/C7SC02664A","journal-title":"Chemical Science"},{"key":"7094_CR41","unstructured":"Xu, H., Zhu, T., Zhang, L., Zhou, W., & Yu, P. S. (2024). Machine unlearning: A survey"},{"key":"7094_CR42","unstructured":"Yu, L., Zhao, Z., Wang, Y., Wang, P., Cao, X., Wang, B., & Wang, Y. (2026). Falw: A forgetting-aware loss reweighting for long-tailed unlearning. arXiv preprint arXiv:2601.18650"},{"key":"7094_CR43","doi-asserted-by":"crossref","unstructured":"Zhao, K., Kurmanji, M., Barbulescu, G.-O., Triantafillou, E., & Triantafillou, P. (2024). What makes unlearning hard and what to do about it. In: Advances in neural information processing systems (NeurIPS 2024)","DOI":"10.52202\/079017-0394"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-026-07094-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-026-07094-y","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-026-07094-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T10:25:21Z","timestamp":1785407121000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-026-07094-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,29]]},"references-count":43,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7]]}},"alternative-id":["7094"],"URL":"https:\/\/doi.org\/10.1007\/s10994-026-07094-y","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,29]]},"assertion":[{"value":"16 January 2026","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 May 2026","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 June 2026","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 June 2026","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no conflict of interest.","order":1,"name":"Ethics","label":"Conflict of interest","group":{"name":"EthicsHeading","label":"Declarations"}}],"article-number":"161"}}