{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T13:24:47Z","timestamp":1780061087289,"version":"3.54.0"},"reference-count":87,"publisher":"Springer Science and Business Media LLC","issue":"9","license":[{"start":{"date-parts":[[2025,8,12]],"date-time":"2025-08-12T00:00:00Z","timestamp":1754956800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,8,12]],"date-time":"2025-08-12T00:00:00Z","timestamp":1754956800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100019180","name":"HORIZON EUROPE European Research Council","doi-asserted-by":"publisher","award":["101093003"],"award-info":[{"award-number":["101093003"]}],"id":[{"id":"10.13039\/100019180","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100007601","name":"Horizon 2020","doi-asserted-by":"publisher","award":["965221"],"award-info":[{"award-number":["965221"]}],"id":[{"id":"10.13039\/501100007601","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002347","name":"Bundesministerium f\u00fcr Bildung und Forschung","doi-asserted-by":"publisher","award":["01IS18025A, 01IS180371I"],"award-info":[{"award-number":["01IS18025A, 01IS180371I"]}],"id":[{"id":"10.13039\/501100002347","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["KI-FOR 5363 \u2212 project ID: 459422098"],"award-info":[{"award-number":["KI-FOR 5363 \u2212 project ID: 459422098"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Fraunhofer-Institut f\u00fcr Nachrichtentechnik, Heinrich-Hertz-Institut HHI"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2025,9]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Deep neural networks are increasingly employed in high-stakes medical applications, despite their tendency for shortcut learning in the presence of spurious correlations, which can have potentially fatal consequences in practice. Whereas a multitude of works address either the detection or mitigation of such shortcut behavior in isolation, the Reveal2Revise approach provides a comprehensive bias mitigation framework combining these steps. However, effectively addressing these biases often requires substantial labeling efforts from domain experts. In this work, we review the steps of the Reveal2Revise framework and enhance it with semi-automated interpretability-based bias annotation capabilities. This includes methods for the sample- and feature-level bias annotation, providing valuable information for bias mitigation methods to unlearn the undesired shortcut behavior. We show the applicability of the framework using four medical datasets across two modalities, featuring controlled and real-world spurious correlations caused by data artifacts. We successfully identify and mitigate these biases in VGG16, ResNet50, and contemporary Vision Transformer models, ultimately increasing their robustness and applicability for real-world medical tasks. Our code is available at <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/frederikpahde\/medical-ai-safety\" ext-link-type=\"uri\">https:\/\/github.com\/frederikpahde\/medical-ai-safety<\/jats:ext-link>.<\/jats:p>","DOI":"10.1007\/s10994-025-06834-w","type":"journal-article","created":{"date-parts":[[2025,8,12]],"date-time":"2025-08-12T20:27:56Z","timestamp":1755030476000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Ensuring medical AI safety: interpretability-driven detection and mitigation of spurious model behavior and associated data"],"prefix":"10.1007","volume":"114","author":[{"given":"Frederik","family":"Pahde","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thomas","family":"Wiegand","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sebastian","family":"Lapuschkin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wojciech","family":"Samek","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,8,12]]},"reference":[{"key":"6834_CR1","unstructured":"Alain, G., & Bengio, Y. (2017). Understanding intermediate layers using linear classifier probes. ICLR"},{"issue":"9","key":"6834_CR2","doi-asserted-by":"publisher","first-page":"1006","DOI":"10.1038\/s42256-023-00711-8","volume":"5","author":"R Achtibat","year":"2023","unstructured":"Achtibat, R., Dreyer, M., Eisenbraun, I., Bosse, S., Wiegand, T., Samek, W., & Lapuschkin, S. (2023). From attribution maps to human-understandable explanations through concept relevance propagation. Nature Machine Intelligence, 5(9), 1006\u20131019.","journal-title":"Nature Machine Intelligence"},{"issue":"11","key":"6834_CR3","doi-asserted-by":"publisher","first-page":"2274","DOI":"10.1109\/TPAMI.2012.120","volume":"34","author":"R Achanta","year":"2012","unstructured":"Achanta, R., Shaji, A., Smith, K., Lucchi, A., Fua, P., & S\u00fcsstrunk, S. (2012). Slic superpixels compared to state-of-the-art superpixel methods. IEEE TPAMI, 34(11), 2274\u20132282.","journal-title":"IEEE TPAMI"},{"key":"6834_CR4","doi-asserted-by":"publisher","first-page":"261","DOI":"10.1016\/j.inffus.2021.07.015","volume":"77","author":"CJ Anders","year":"2022","unstructured":"Anders, C. J., Weber, L., Neumann, D., Samek, W., M\u00fcller, K.-R., & Lapuschkin, S. (2022). Finding and removing clever Hans: Using explanation methods to debug and improve deep models. Information Fusion, 77, 261\u2013295.","journal-title":"Information Fusion"},{"issue":"7","key":"6834_CR5","doi-asserted-by":"publisher","first-page":"0130140","DOI":"10.1371\/journal.pone.0130140","volume":"10","author":"S Bach","year":"2015","unstructured":"Bach, S., Binder, A., Montavon, G., Klauschen, F., M\u00fcller, K.-R., & Samek, W. (2015). On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLoS ONE, 10(7), 0130140.","journal-title":"PLoS ONE"},{"key":"6834_CR6","unstructured":"Bykov, K., Deb, M., Grinwald, D., M\u00fcller, K.-R., & H\u00f6hne, M. M. (2023). Dora: Exploring outlier representations in deep neural networks. In ICLR Workshops."},{"key":"6834_CR7","doi-asserted-by":"crossref","unstructured":"Bareeva, D., Dreyer, M., Pahde, F., Samek, W., & Lapuschkin, S. (2024). Reactive model correction: Mitigating harm to task-relevant features via conditional bias suppression. In CVPRW, (pp. 3532\u20133541).","DOI":"10.1109\/CVPRW63382.2024.00357"},{"issue":"1","key":"6834_CR8","doi-asserted-by":"publisher","first-page":"207","DOI":"10.1162\/coli_a_00422","volume":"48","author":"Y Belinkov","year":"2022","unstructured":"Belinkov, Y. (2022). Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1), 207\u2013219.","journal-title":"Computational Linguistics"},{"key":"6834_CR9","doi-asserted-by":"publisher","first-page":"47","DOI":"10.1016\/j.ejca.2019.04.001","volume":"113","author":"TJ Brinker","year":"2019","unstructured":"Brinker, T. J., Hekler, A., Enk, A. H., Klode, J., Hauschild, A., Berking, C., Schilling, B., Haferkamp, S., Schadendorf, D., et al. (2019). Deep learning outperformed 136 of 157 dermatologists in a head-to-head dermoscopic melanoma image classification task. European Journal of Cancer, 113, 47\u201354.","journal-title":"European Journal of Cancer"},{"key":"6834_CR10","unstructured":"Bykov, K., Kopf, L., Nakajima, S., Kloft, M., & H\u00f6hne, M. (2024). Labeling neural representations with inverse recognition. NeurIPS36"},{"key":"6834_CR11","doi-asserted-by":"crossref","unstructured":"Breunig, M. M., Kriegel, H.-P., Ng, R. T., & Sander, J. (2000). Lof: identifying density-based local outliers. In Proceedings of the 2000 ACM SIGMOD, (pp. 93\u2013104).","DOI":"10.1145\/342009.335388"},{"key":"6834_CR12","unstructured":"Begley, T., Schwedes, T., Frye, C., & Feige, I. (2020). Explainability for fair machine learning. arXiv preprint arXiv:2010.07389"},{"key":"6834_CR13","unstructured":"Belrose, N., Schneider-Joseph, D., Ravfogel, S., Cotterell, R., Raff, E., & Biderman, S. (2024). Leace: Perfect linear concept erasure in closed form. NeurIPS36"},{"key":"6834_CR14","unstructured":"Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., & Olah, C. (2023).Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread2"},{"issue":"1","key":"6834_CR15","doi-asserted-by":"publisher","first-page":"4314","DOI":"10.1038\/s41467-023-39902-7","volume":"14","author":"A Brown","year":"2023","unstructured":"Brown, A., Tomasev, N., Freyberg, J., Liu, Y., Karthikesalingam, A., & Schrouff, J. (2023). Detecting shortcut learning for fair medical AI using shortcut testing. Nature Communications, 14(1), 4314.","journal-title":"Nature Communications"},{"issue":"1","key":"6834_CR16","doi-asserted-by":"publisher","first-page":"283","DOI":"10.1038\/s41597-020-00622-y","volume":"7","author":"H Borgli","year":"2020","unstructured":"Borgli, H., Thambawita, V., Smedsrud, P. H., Hicks, S., Jha, D., Eskeland, S. L., Randel, K. R., et al. (2020). Hyperkvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy. Scientific data, 7(1), 283.","journal-title":"Scientific data"},{"key":"6834_CR17","doi-asserted-by":"crossref","unstructured":"Bissoto, A., Valle, E., & Avila, S. (2020). Debiasing skin lesion datasets and models? Not so fast. In CVPRW, (pp. 740\u2013741).","DOI":"10.1109\/CVPRW50498.2020.00378"},{"issue":"48","key":"6834_CR18","doi-asserted-by":"publisher","first-page":"30071","DOI":"10.1073\/pnas.1907375117","volume":"117","author":"D Bau","year":"2020","unstructured":"Bau, D., Zhu, J.-Y., Strobelt, H., Lapedriza, A., Zhou, B., & Torralba, A. (2020). Understanding the role of individual units in a deep neural network. Proceedings of the National Academy of Sciences, 117(48), 30071\u201330078.","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"6834_CR19","unstructured":"Borowski, J., Zimmermann, R. S., Schepers, J., Geirhos, R., Wallis, T. S., Bethge, M., & Brendel, W. (2020). Natural images are more informative for interpreting cnn activations than state-of-the-art synthetic feature visualizations. In NeurIPS 2020 Workshop SVRHM."},{"key":"6834_CR20","unstructured":"Combalia, M., Codella, N. C., Rotemberg, V., Helba, B., Vilaplana, V., Reiter, O., Carrera, C., & Malvehy, J. (2019). BCN20000: Dermoscopic Lesions in the Wild"},{"key":"6834_CR21","doi-asserted-by":"crossref","unstructured":"Codella, N. C., Gutman, D., Celebi, M. E., Helba, B., Marchetti, M. A., Dusza, S. W., Kalloo, A., & Halpern, A. (2018). Skin lesion analysis toward melanoma detection: A challenge at the 2017 international symposium on biomedical imaging (ISBI), hosted by the international skin imaging collaboration (isic). In 15th International Symposium on Biomedical Imaging (ISBI 2018), (pp. 168\u2013172). IEEE.","DOI":"10.1109\/ISBI.2018.8363547"},{"key":"6834_CR22","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2021.102305","volume":"75","author":"B Cassidy","year":"2022","unstructured":"Cassidy, B., Kendrick, C., Brodzicki, A., Jaworek-Korjakowska, J., & Yap, M. H. (2022). Analysis of the ISIC image datasets: Usage, benchmarks and recommendations. Medical Image Analysis, 75, Article 102305.","journal-title":"Medical Image Analysis"},{"key":"6834_CR23","doi-asserted-by":"publisher","first-page":"273","DOI":"10.1023\/A:1022627411411","volume":"20","author":"C Cortes","year":"1995","unstructured":"Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20, 273\u2013297.","journal-title":"Machine Learning"},{"key":"6834_CR24","first-page":"2590","volume":"35","author":"J Crabb\u00e9","year":"2022","unstructured":"Crabb\u00e9, J., & Schaar, M. (2022). Concept activation regions: A generalized framework for concept-based explanations. NeurIPS, 35, 2590\u20132607.","journal-title":"NeurIPS"},{"key":"6834_CR25","doi-asserted-by":"crossref","unstructured":"Dreyer, M., Achtibat, R., Samek, W., & Lapuschkin, S. (2024). Understanding the (extra-) ordinary: Validating deep model decisions with prototypical concept-based explanations. In CVPRW, (pp. 3491\u20133501).","DOI":"10.1109\/CVPRW63382.2024.00353"},{"key":"6834_CR26","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., & Houlsby, N. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. ICLR"},{"key":"6834_CR27","doi-asserted-by":"crossref","unstructured":"Dreyer, M., Berend, J., Labarta, T., Vielhaben, J., Wiegand, T., Lapuschkin, S., & Samek, W. (2025). Mechanistic understanding and validation of large ai models with semanticlens. arXiv preprint arXiv:2501.05398arXiv:2501.05398","DOI":"10.1038\/s42256-025-01084-w"},{"key":"6834_CR28","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. In CVPR, (pp. 248\u2013255). IEEE.","DOI":"10.1109\/CVPR.2009.5206848"},{"issue":"7","key":"6834_CR29","doi-asserted-by":"publisher","first-page":"610","DOI":"10.1038\/s42256-021-00338-7","volume":"3","author":"AJ DeGrave","year":"2021","unstructured":"DeGrave, A. J., Janizek, J. D., & Lee, S.-I. (2021). Ai for radiographic covid-19 detection selects shortcuts over signal. Nature Machine Intelligence, 3(7), 610\u2013619.","journal-title":"Nature Machine Intelligence"},{"key":"6834_CR30","doi-asserted-by":"publisher","first-page":"21046","DOI":"10.1609\/aaai.v38i19.30096","volume":"38","author":"M Dreyer","year":"2024","unstructured":"Dreyer, M., Pahde, F., Anders, C. J., Samek, W., & Lapuschkin, S. (2024). From hope to safety: Unlearning biases of deep models via gradient penalization in latent space. AAAI, 38, 21046\u201321054.","journal-title":"AAAI"},{"key":"6834_CR31","unstructured":"Dreyer, M., Purelku, E., Vielhaben, J., Samek, W., & Lapuschkin, S. (2024). Pure: Turning polysemantic neurons into pure features by identifying relevant circuits. In CVPRW, (pp. 8212\u20138217)."},{"key":"6834_CR32","doi-asserted-by":"crossref","unstructured":"De\u00a0Santis, A., Campi, R., Bianchi, M., & Brambilla, M. (2024). Visual-TCAV: Concept-based attribution and saliency maps for post-hoc explainability in image classification. arXiv preprint arXiv:2411.05698","DOI":"10.1016\/B978-0-44-340553-2.00008-3"},{"key":"6834_CR33","unstructured":"Denil, M., Shakibi, B., Dinh, L., Ranzato, M., & De\u00a0Freitas, N. (2013). Predicting parameters in deep learning. NeurIPS26"},{"issue":"3","key":"6834_CR34","first-page":"1","volume":"1341","author":"D Erhan","year":"2009","unstructured":"Erhan, D., Bengio, Y., Courville, A., & Vincent, P. (2009). Visualizing higher-layer features of a deep network. University of Montreal, 1341(3), 1.","journal-title":"University of Montreal"},{"key":"6834_CR35","unstructured":"Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., & Grosse, R., (2022). Toy models of superposition. arXiv preprint arXiv:2209.10652"},{"key":"6834_CR36","unstructured":"Fel, T., Boissin, T., Boutin, V., Picard, A., Novello, P., Colin, J., Linsley, D., Rousseau, T., Cadene, R., Goetschalckx, L., & Serre, T. (2024). Unlocking feature visualization for deep network with magnitude constrained optimization. NeurIPS36"},{"key":"6834_CR37","unstructured":"Fel, T., Boutin, V., B\u00e9thune, L., Cad\u00e8ne, R., Moayeri, M., And\u00e9ol, L., Chalvidal, M., & Serre, T. (2024). A holistic approach to unifying automatic concept extraction and concept importance estimation. NeurIPS36"},{"issue":"2","key":"6834_CR38","doi-asserted-by":"publisher","first-page":"179","DOI":"10.1111\/j.1469-1809.1936.tb02137.x","volume":"7","author":"RA Fisher","year":"1936","unstructured":"Fisher, R. A. (1936). The use of multiple measurements in taxonomic problems. Annals of Eugenics, 7(2), 179\u2013188.","journal-title":"Annals of Eugenics"},{"key":"6834_CR39","doi-asserted-by":"crossref","unstructured":"Fel, T., Picard, A., Bethune, L., Boissin, T., Vigouroux, D., Colin, J., Cad\u00e8ne, R., & Serre, T. (2023). Craft: Concept recursive activation factorization for explainability. In CVPR, (pp. 2711\u20132721).","DOI":"10.1109\/CVPR52729.2023.00266"},{"key":"6834_CR40","doi-asserted-by":"crossref","unstructured":"Fong, R., & Vedaldi, A. (2018). Net2Vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks. In CVPR, (pp. 8730\u20138738).","DOI":"10.1109\/CVPR.2018.00910"},{"issue":"11","key":"6834_CR41","doi-asserted-by":"publisher","first-page":"665","DOI":"10.1038\/s42256-020-00257-z","volume":"2","author":"R Geirhos","year":"2020","unstructured":"Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., et al. (2020). Shortcut learning in deep neural networks. Nature Machine Intelligence, 2(11), 665\u2013673.","journal-title":"Nature Machine Intelligence"},{"key":"6834_CR42","unstructured":"Graziani, M., Nguyen, A.-p., O\u2019Mahony, L., M\u00fcller, H., & Andrearczyk, V. (2023). Concept discovery and dataset exploration with singular value decomposition. In ICLR Workshops."},{"key":"6834_CR43","unstructured":"Ghorbani, A., Wexler, J., Zou, J. Y., & Kim, B. (2019). Towards automatic concept-based explanations. NeurIPS32."},{"key":"6834_CR44","unstructured":"Huben, R., Cunningham, H., Smith, L. R., Ewart, A., & Sharkey, L. (2023). Sparse autoencoders find highly interpretable features in language models. In ICLR."},{"key":"6834_CR45","unstructured":"Hernandez, E., Schwettmann, S., Bau, D., Bagashvili, T., Torralba, A., & Andreas, J. (2021). Natural language descriptions of deep visual features. In ICLR."},{"key":"6834_CR46","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In CVPR, (pp. 770\u2013778).","DOI":"10.1109\/CVPR.2016.90"},{"key":"6834_CR47","doi-asserted-by":"crossref","unstructured":"He, T., Zhang, Z., Zhang, H., Zhang, Z., Xie, J., & Li, M. (2019). Bag of tricks for image classification with convolutional neural networks. In CVPR, (pp. 558\u2013567).","DOI":"10.1109\/CVPR.2019.00065"},{"key":"6834_CR48","doi-asserted-by":"crossref","unstructured":"Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K., & Seekins, J. (2019). Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison. In AAAI.","DOI":"10.1609\/aaai.v33i01.3301590"},{"key":"6834_CR49","unstructured":"Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F. (2018). Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV). In ICML, (pp. 2668\u20132677). PMLR."},{"key":"6834_CR50","unstructured":"Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. NeurIPS30."},{"issue":"1","key":"6834_CR51","doi-asserted-by":"publisher","first-page":"1096","DOI":"10.1038\/s41467-019-08987-4","volume":"10","author":"S Lapuschkin","year":"2019","unstructured":"Lapuschkin, S., W\u00e4ldchen, S., Binder, A., Montavon, G., Samek, W., & M\u00fcller, K.-R. (2019). Unmasking clever hans predictors and assessing what machines really learn. Nature Communications, 10(1), 1096.","journal-title":"Nature Communications"},{"key":"6834_CR52","doi-asserted-by":"crossref","unstructured":"McInnes, L., Healy, J., Saul, N., & Gro\u00dfberger, L. (2018). UMAP: Uniform manifold approximation and projection. Journal of Open Source Software3(29).","DOI":"10.21105\/joss.00861"},{"key":"6834_CR53","doi-asserted-by":"crossref","unstructured":"Morch, N. J., Kjems, U., Hansen, L. K., Svarer, C., Law, I., Lautrup, B., Strother, S., & Rehm, K. (1995). Visualization of neural networks using saliency maps. In ICNN, vol. 4, (pp. 2085\u20132090). IEEE.","DOI":"10.1109\/ICNN.1995.488997"},{"key":"6834_CR54","unstructured":"Murdoch, W. J., Liu, P. J., & Yu, B. (2018). Beyond word importance: Contextual decomposition to extract interactions from LSTMs. ICLR."},{"key":"6834_CR55","doi-asserted-by":"crossref","unstructured":"Mikriukov, G., Schwalbe, G., Hellert, C., & Bade, K. (2023). Evaluating the stability of semantic concept representations in CNNs for robust explainability. In World Conference on Explainable Artificial Intelligence, (pp. 499\u2013524). Springer.","DOI":"10.1007\/978-3-031-44067-0_26"},{"key":"6834_CR56","doi-asserted-by":"crossref","unstructured":"Neuhaus, Y., Augustin, M., Boreiko, V., & Hein, M. (2023). Spurious features everywhere-large-scale detection of harmful spurious features in ImageNet. In ICCV.","DOI":"10.1109\/ICCV51070.2023.01851"},{"issue":"285\u2013296","key":"6834_CR57","first-page":"23","volume":"11","author":"N Otsu","year":"1975","unstructured":"Otsu, N., et al. (1975). A threshold selection method from gray-level histograms. Automatica, 11(285\u2013296), 23\u201327.","journal-title":"Automatica"},{"issue":"3","key":"6834_CR58","doi-asserted-by":"publisher","first-page":"00024","DOI":"10.23915\/distill.00024.001","volume":"5","author":"C Olah","year":"2020","unstructured":"Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., & Carter, S. (2020). Zoom in: An introduction to circuits. Distill, 5(3), 00024\u2013001.","journal-title":"Distill"},{"issue":"11","key":"6834_CR59","doi-asserted-by":"publisher","first-page":"7","DOI":"10.23915\/distill.00007","volume":"2","author":"C Olah","year":"2017","unstructured":"Olah, C., Mordvintsev, A., & Schubert, L. (2017). Feature visualization. Distill, 2(11), 7.","journal-title":"Distill"},{"key":"6834_CR60","unstructured":"Oikarinen, T., & Weng, T.-W. (2023). Clip-dissect: Automatic description of neuron representations in deep vision networks. In ICLR."},{"key":"6834_CR61","doi-asserted-by":"crossref","unstructured":"Pahde, F., Dreyer, M., Samek, W., & Lapuschkin, S. (2023). Reveal to revise: An explainable AI life cycle for iterative bias correction of deep models. In MICCAI.","DOI":"10.1007\/978-3-031-43895-0_56"},{"key":"6834_CR62","unstructured":"Pahde, F., Dreyer, M., Weber, L., Weckbecker, M., Anders, C. J., Wiegand, T., Samek, W., & Lapuschkin, S. (2025). Navigating neural space: Revisiting concept activation vectors to overcome directional divergence. In International conference on learning representations."},{"key":"6834_CR63","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., & Desmaison, A. (2019). PyTorch: An imperative style, high-performance deep learning library. NeurIPS32."},{"key":"6834_CR64","doi-asserted-by":"crossref","unstructured":"Pradhan, R., Zhu, J., Glavic, B., & Salimi, B. (2022). Interpretable data-based explanations for fairness debugging. In Proceedings of the 2022 International Conference on Management of Data, (pp. 247\u2013261).","DOI":"10.1145\/3514221.3517886"},{"key":"6834_CR65","doi-asserted-by":"crossref","unstructured":"Ravfogel, S., Elazar, Y., Gonen, H., Twiton, M., & Goldberg, Y. (2020). Null it out: Guarding protected attributes by iterative nullspace projection. In: Jurafsky, D., Chai, J., Schluter, N., Tetreault, J. (Eds.), ACL, (pp. 7237\u20137256).","DOI":"10.18653\/v1\/2020.acl-main.647"},{"key":"6834_CR66","doi-asserted-by":"crossref","unstructured":"Ross, A. S., Hughes, M. C., & Doshi-Velez, F. (2017). Right for the right reasons: training differentiable models by constraining their explanations. In IJCAI.","DOI":"10.24963\/ijcai.2017\/371"},{"key":"6834_CR67","unstructured":"Radford, A., Jozefowicz, R., & Sutskever, I. (2017). Learning to generate reviews and discovering sentiment. arXiv preprint arXiv:1704.01444"},{"key":"6834_CR68","unstructured":"Rieger, L., Singh, C., Murdoch, W., & Yu, B. (2020). Interpretations are useful: penalizing explanations to align neural networks with prior knowledge. In ICML."},{"key":"6834_CR69","doi-asserted-by":"crossref","unstructured":"Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-cam: Visual explanations from deep networks via gradient-based localization. In ICCV, (pp. 618\u2013626).","DOI":"10.1109\/ICCV.2017.74"},{"key":"6834_CR70","unstructured":"Singla, S., & Feizi, S. (2022). Salient ImageNet: How to discover spurious features in deep learning. In ICLR."},{"issue":"2","key":"6834_CR71","first-page":"1","volume":"3","author":"D Slijepcevic","year":"2021","unstructured":"Slijepcevic, D., Horst, F., Lapuschkin, S., Horsak, B., Raberger, A.-M., Kranzl, A., Samek, W., Breiteneder, C., Sch\u00f6llhorn, W. I., & Zeppelzauer, M. (2021). Explaining machine learning models for clinical gait analysis. ACM Transactions on Computing for Healthcare (HEALTH), 3(2), 1\u201327.","journal-title":"ACM Transactions on Computing for Healthcare (HEALTH)"},{"issue":"8","key":"6834_CR72","doi-asserted-by":"publisher","first-page":"476","DOI":"10.1038\/s42256-020-0212-3","volume":"2","author":"P Schramowski","year":"2020","unstructured":"Schramowski, P., Stammer, W., Teso, S., Brugger, A., Herbert, F., Shao, X., Luigs, H.-G., Mahlein, A.-K., & Kersting, K. (2020). Making deep neural networks right for the right scientific reasons by interacting with their explanations. Nature Machine Intelligence, 2(8), 476\u2013486.","journal-title":"Nature Machine Intelligence"},{"key":"6834_CR73","first-page":"23359","volume":"34","author":"S Santurkar","year":"2021","unstructured":"Santurkar, S., Tsipras, D., Elango, M., Bau, D., Torralba, A., & Madry, A. (2021). Editing a classifier by rewriting its prediction rules. NeurIPS, 34, 23359\u201323373.","journal-title":"NeurIPS"},{"issue":"5","key":"6834_CR74","doi-asserted-by":"publisher","first-page":"1519","DOI":"10.1109\/JBHI.2020.3022989","volume":"25","author":"N Strodthoff","year":"2020","unstructured":"Strodthoff, N., Wagner, P., Schaeffter, T., & Samek, W. (2020). Deep learning for ECG analysis: Benchmarks and insights from PTB-xl. IEEE Journal of Biomedical and Health Informatics, 25(5), 1519\u20131528.","journal-title":"IEEE Journal of Biomedical and Health Informatics"},{"key":"6834_CR75","unstructured":"Simonyan, K., & Zisserman, A. (2015). Very deep convolutional networks for large-scale image recognition. In Bengio, Y., LeCun, Y. (Eds.) CLR. 2015."},{"key":"6834_CR76","unstructured":"Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., & Fergus, R. (2014). Intriguing properties of neural networks. In ICLR."},{"issue":"1","key":"6834_CR77","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/sdata.2018.161","volume":"5","author":"P Tschandl","year":"2018","unstructured":"Tschandl, P., Rosendahl, C., & Kittler, H. (2018). The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Scientific Data, 5(1), 1\u20139.","journal-title":"Scientific Data"},{"key":"6834_CR78","unstructured":"Vielhaben, J., Bluecher, S., & Strodthoff, N. (2023). Multi-dimensional concept discovery (MCD): A unifying framework with completeness guarantees. TMLR."},{"key":"6834_CR79","unstructured":"Maaten, L., & Hinton, G. (2008). Visualizing data using T-SNE. JMLR9(11)."},{"key":"6834_CR80","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.4414861","author":"R Wightman","year":"2019","unstructured":"Wightman, R. (2019). PyTorch Image Models. GitHub. https:\/\/doi.org\/10.5281\/zenodo.4414861","journal-title":"GitHub"},{"key":"6834_CR81","doi-asserted-by":"publisher","DOI":"10.1016\/j.compbiomed.2024.108525","volume":"176","author":"P Wagner","year":"2024","unstructured":"Wagner, P., Mehari, T., Haverkamp, W., & Strodthoff, N. (2024). Explaining deep learning for ECG analysis: Building blocks for auditing and knowledge discovery. Computers in Biology and Medicine, 176, Article 108525.","journal-title":"Computers in Biology and Medicine"},{"key":"6834_CR82","doi-asserted-by":"crossref","unstructured":"Weng, N., Pegios, P., Petersen, E., Feragen, A., & Bigdeli, S. (2025). Fast diffusion-based counterfactuals for shortcut removal and generation. In ECCV.","DOI":"10.1007\/978-3-031-73016-0_20"},{"issue":"1","key":"6834_CR83","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/s41597-020-0495-6","volume":"7","author":"P Wagner","year":"2020","unstructured":"Wagner, P., Strodthoff, N., Bousseljot, R.-D., Kreiseler, D., Lunze, F. I., Samek, W., & Schaeffter, T. (2020). PTB-XL, a large publicly available electrocardiography dataset. Scientific Data, 7(1), 1\u201315.","journal-title":"Scientific Data"},{"key":"6834_CR84","unstructured":"Wu, S., Yuksekgonul, M., Zhang, L., & Zou, J. (2023). Discover and cure: Concept-aware mitigation of spurious correlation. In ICML."},{"key":"6834_CR85","doi-asserted-by":"crossref","unstructured":"Zech, J. R., Badgeley, M. A., Liu, M., Costa, A. B., Titano, J. J., & Oermann, E. K. (2018). Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study. PLoS Medicine15(11).","DOI":"10.1371\/journal.pmed.1002683"},{"key":"6834_CR86","doi-asserted-by":"crossref","unstructured":"Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., & Torralba, A. (2016). Learning deep features for discriminative localization. In CVPR, (pp. 2921\u20132929).","DOI":"10.1109\/CVPR.2016.319"},{"key":"6834_CR87","doi-asserted-by":"publisher","first-page":"11682","DOI":"10.1609\/aaai.v35i13.17389","volume":"35","author":"R Zhang","year":"2021","unstructured":"Zhang, R., Madumal, P., Miller, T., Ehinger, K. A., & Rubinstein, B. I. (2021). Invertible concept-based explanations for CNN models with non-negative concept activation vectors. AAAI, 35, 11682\u201311690.","journal-title":"AAAI"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06834-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-025-06834-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06834-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,9]],"date-time":"2025-09-09T21:16:45Z","timestamp":1757452605000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-025-06834-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,12]]},"references-count":87,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2025,9]]}},"alternative-id":["6834"],"URL":"https:\/\/doi.org\/10.1007\/s10994-025-06834-w","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,12]]},"assertion":[{"value":"20 January 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 June 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 July 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 August 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"206"}}