{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T21:28:20Z","timestamp":1783978100779,"version":"3.55.0"},"reference-count":56,"publisher":"American Association for the Advancement of Science (AAAS)","content-domain":{"domain":["spj.science.org"],"crossmark-restriction":true},"short-container-title":["Intell Comput"],"published-print":{"date-parts":[[2024,1]]},"abstract":"<jats:p>Artificial intelligence (AI) capabilities have grown rapidly with the introduction of cutting-edge deep-model architectures and learning strategies. Explainable AI (XAI) methods aim to make the capabilities of AI models beyond accuracy interpretable by providing explanations. The explanations are mainly used to increase model transparency, debug the model, and justify the model predictions to the end user. Most current XAI methods focus on providing visual and textual explanations that are prone to being present in visual media. However, audio explanations are crucial because of their intuitiveness in audio-based tasks and higher expressiveness than other modalities in specific scenarios, such as when understanding visual explanations requires expertise. In this review, we provide an overview of XAI methods for audio in 2 categories: exploiting generic XAI methods to explain audio models, and XAI methods specialised for the interpretability of audio models. Additionally, we discuss certain open problems and highlight future directions for the development of XAI techniques for audio modeling.<\/jats:p>","DOI":"10.34133\/icomputing.0074","type":"journal-article","created":{"date-parts":[[2023,12,18]],"date-time":"2023-12-18T10:49:04Z","timestamp":1702896544000},"update-policy":"https:\/\/doi.org\/10.34133\/aaas_crossmark_01","source":"Crossref","is-referenced-by-count":40,"title":["Audio Explainable Artificial Intelligence: A Review"],"prefix":"10.34133","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8010-6897","authenticated-orcid":true,"given":"Alican","family":"Akman","sequence":"first","affiliation":[{"name":"Department of Computing, \rImperial College London, London, UK."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bj\u00f6rn W.","family":"Schuller","sequence":"additional","affiliation":[{"name":"Department of Computing, \rImperial College London, London, UK."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"221","published-online":{"date-parts":[[2024,1,23]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Devlin J Chang M Lee K Toutanova K. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv. 2018. https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_2_3_2","doi-asserted-by":"crossref","unstructured":"Peters ME Neumann M Iyyer M Gardner M Clark C Lee K Zettlemoyer L. Deep contextualized word representations. arXiv. 2018. https:\/\/arxiv.org\/abs\/1802.05365","DOI":"10.18653\/v1\/N18-1202"},{"key":"e_1_3_2_4_2","unstructured":"Lan Z Chen M Goodman S Gimpel K Sharma P Soricut R. ALBERT: A lite BERT for self-supervised learning of language representations. arXiv. 2019. https:\/\/arxiv.org\/abs\/1909.11942"},{"key":"e_1_3_2_5_2","unstructured":"Touvron H Cord M Douze M Massa F Sablayrolles A J\u00e9gou H. Training data-efficient image transformers & distillation through attention. arXiv. 2021. https:\/\/arxiv.org\/abs\/2012.12877"},{"key":"e_1_3_2_6_2","unstructured":"Radford A Kim JW Hallacy C Ramesh A Goh G Agarwal S Sastry G Askell A Mishkin P Clark J et\u00a0al. Learning transferable visual models from natural language supervision. arXiv. 2021. https:\/\/arxiv.org\/abs\/2103.00020."},{"key":"e_1_3_2_7_2","unstructured":"Khan S Naseer M Hayat M Zamir SW Khan FS Shah M Transformers in vision: A survey. 2021. https:\/\/arxiv.org\/abs\/2101.01169"},{"key":"e_1_3_2_8_2","unstructured":"Brown TB Mann B Ryder N Subbiah M Kaplan J Dhariwal P Neelakantan A Shyam P Sastry G Askell A et\u00a0al. Language models are few-shot learners. 2020. https:\/\/arxiv.org\/abs\/2005.14165"},{"key":"e_1_3_2_9_2","unstructured":"Dosovitskiy A Beyer L Kolesnikov A Weissenborn D Zhai X Unterthiner T Dehghani M Minderer M Heigold G Gelly S et\u00a0al. An image is worth 16x16 words: Transformers for image recognition at scale. 2020. https:\/\/arxiv.org\/abs\/2010.11929"},{"issue":"1","key":"e_1_3_2_10_2","doi-asserted-by":"crossref","first-page":"310","DOI":"10.1186\/s12911-020-01332-6","article-title":"Explainability for artificial intelligence in healthcare: A multidisciplinary perspective","volume":"20","author":"Amann J","year":"2020","unstructured":"Amann J, Blasimme A, Vayena E, Frey D, Madai V. Explainability for artificial intelligence in healthcare: A multidisciplinary perspective. BMC Med Inform Decis Mak. 2020;20(1):310.","journal-title":"BMC Med Inform Decis Mak"},{"issue":"4","key":"e_1_3_2_11_2","doi-asserted-by":"crossref","DOI":"10.1016\/S2589-7500(22)00029-2","article-title":"Explainability and artificial intelligence in medicine","volume":"4","author":"Reddy S","year":"2022","unstructured":"Reddy S. Explainability and artificial intelligence in medicine. Lancet Digital Health. 2022;4(4): Article e214.","journal-title":"Lancet Digital Health"},{"issue":"7","key":"e_1_3_2_12_2","first-page":"1829","article-title":"The judicial demand for explainable artificial intelligence","volume":"119","author":"Deeks A","year":"2019","unstructured":"Deeks A. The judicial demand for explainable artificial intelligence. Columbia Law Rev. 2019;119(7):1829\u20131850.","journal-title":"Columbia Law Rev"},{"key":"e_1_3_2_13_2","unstructured":"Atakishiyev S Salameh M Yao H Goebel R. Explainable artificial intelligence for autonomous driving: A comprehensive overview and field guide for future research directions. arXiv. 2021. https:\/\/arxiv.org\/abs\/2112.11561"},{"key":"e_1_3_2_14_2","article-title":"Applications of explainable artificial intelligence in finance\u2013a systematic review of finance, information systems, and computer science literature","author":"Weber P","year":"2023","unstructured":"Weber P, Carl KV, Hinz O. Applications of explainable artificial intelligence in finance\u2013a systematic review of finance, information systems, and computer science literature. Manag Rev Q. 2023.","journal-title":"Manag Rev Q"},{"key":"e_1_3_2_15_2","unstructured":"Vilone G Longo L. Explainable artificial intelligence: A systematic review. arXiv. 2020. https:\/\/arxiv.org\/abs\/2006.00093"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1016\/j.inffus.2019.12.012","article-title":"Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI","volume":"58","author":"Barredo Arrieta A","year":"2020","unstructured":"Barredo Arrieta A, D\u00edaz-Rodr\u00edguez N, del Ser J, Bennetot A, Tabik S, Barbado A, Garcia S, Gil-Lopez S, Molina D, Benjamins R, et al. Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Inf Fusion. 2020;58:82\u2013115.","journal-title":"Inf Fusion"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","unstructured":"Schuller BW Virtanen T Riveiro M Rizos G Han J Mesaros A Drossos K. Towards sonification in multimodal and user-friendly explainable artificial intelligence. Paper Presented at: Proceedings of the 2021 International Conference on Multimodal Interaction ICMI \u201921; 2021; Montreal QC Canada; p. 788-792.","DOI":"10.1145\/3462244.3479879"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","unstructured":"Eyben F Weninger F Gross F Schuller B. Recent developments in OpenSMILE the Munich open-source multimedia feature extractor. Paper presented at: Proceedings of the 21st ACM International Conference on Multimedia MM \u201913; 2013; Barcelona Spain. p. 835\u2013838.","DOI":"10.1145\/2502081.2502224"},{"key":"e_1_3_2_19_2","first-page":"1","article-title":"Openxbow\u2014Introducing the Passau open-source crossmodal bag-of-words toolkit","volume":"18","author":"Schmitt M","year":"2017","unstructured":"Schmitt M, Schuller B. Openxbow\u2014Introducing the Passau open-source crossmodal bag-of-words toolkit. J Mach Learn Res. 2017;18:1\u20135.","journal-title":"J Mach Learn Res"},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","unstructured":"Amiriparian S Gerczuk M Ottl S Cummins N Pugachevskly S Schuller B. Bag-of-deep-features: Noise-robust deep feature representations for audio analysis. Paper presented at: 2018 International Joint Conference on Neural Networks (IJCNN); 2018 Jul 8\u201313; Rio de Janeiro Brazil.","DOI":"10.1109\/IJCNN.2018.8489416"},{"issue":"1","key":"e_1_3_2_21_2","first-page":"6340","article-title":"Audeep: Unsupervised learning of representations from audio with deep recurrent neural networks","volume":"18","author":"Freitag M","year":"2018","unstructured":"Freitag M, Amiriparian S, Pugachevskiy S, Cummins N, Schuller B. Audeep: Unsupervised learning of representations from audio with deep recurrent neural networks. J Mach Learn Res. 2018;18(1):6340\u20136344.","journal-title":"J Mach Learn Res"},{"key":"e_1_3_2_22_2","unstructured":"Springenberg JT Dosovitskiy A Brox T Riedmiller M. Striving for simplicity: The all convolutional net. arXiv. 2014. https:\/\/arxiv.org\/abs\/1412.6806"},{"issue":"7","key":"e_1_3_2_23_2","doi-asserted-by":"crossref","first-page":"e0130140","DOI":"10.1371\/journal.pone.0130140","article-title":"On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation","volume":"10","author":"Bach S","year":"2015","unstructured":"Bach S, Binder A, Montavon G, Klauschen F, M\u00fcller K-R, Samek W. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLOS ONE. 2015;10(7):e0130140.","journal-title":"PLOS ONE"},{"key":"e_1_3_2_24_2","unstructured":"Sundararajan M Taly A Yan Q. Axiomatic attribution for deep networks. arXiv. 2017. https:\/\/arxiv.org\/abs\/1703.01365"},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","unstructured":"Selvaraju RR Cogswell M Das A Vedantam R Parikh D Batra D. Grad-cam: Why did you say that? Visual explanations from deep networks via gradient-based localization. arXiv. 2016. https:\/\/arxiv.org\/abs\/1610.02391","DOI":"10.1109\/ICCV.2017.74"},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","unstructured":"Ribeiro MT Singh S Guestrin C. \u201cWhy should I trust you?\u201d: Explaining the predictions of any classifier. arXiv. 2016. https:\/\/arxiv.org\/abs\/1602.04938","DOI":"10.1145\/2939672.2939778"},{"key":"e_1_3_2_27_2","unstructured":"Lundberg SM Lee S. A unified approach to interpreting model predictions. arXiv. 2017. https:\/\/arxiv.org\/abs\/1705.07874"},{"key":"e_1_3_2_28_2","doi-asserted-by":"crossref","unstructured":"Wiegreffe S Pinter Y. Attention is not not explanation. arXiv. 2019. https:\/\/arxiv.org\/abs\/1908.04626","DOI":"10.18653\/v1\/D19-1002"},{"key":"e_1_3_2_29_2","doi-asserted-by":"crossref","unstructured":"Caron M Touvron H Misra I J\u00e9gou H Mairal J Bojanowski P Joulin A Emerging properties in self-supervised vision transformers. arXiv. 2021. https:\/\/arxiv.org\/abs\/2104.14294","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"e_1_3_2_30_2","doi-asserted-by":"crossref","unstructured":"Chefer H Gur S Wolf L. Transformer interpretability beyond attention visualization. arXiv. 2020. https:\/\/arxiv.org\/abs\/2012.09838","DOI":"10.1109\/CVPR46437.2021.00084"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2016.11.008"},{"key":"e_1_3_2_32_2","unstructured":"Koh PW Liang P. Understanding black-box predictions via influence functions. arXiv. 2020. https:\/\/arxiv.org\/abs\/1703.04730"},{"key":"e_1_3_2_33_2","article-title":"Feature visualization","author":"Olah C","unstructured":"Olah C, Mordvintsev A, Schubert L. Feature visualization. Distill. 2017.","journal-title":"Distill"},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","unstructured":"Bau D Zhou B Khosla A Oliva A Torralba A. Network dissection: Quantifying interpretability of deep visual representations. arXiv. 2017. https:\/\/arxiv.org\/abs\/1704.05796","DOI":"10.1109\/CVPR.2017.354"},{"key":"e_1_3_2_35_2","unstructured":"Kim B Wattenberg M Gilmer J Cai C Wexler J Viegas F Sayres R. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). arXiv. 2018. https:\/\/arxiv.org\/abs\/1711.11279"},{"issue":"12","key":"e_1_3_2_36_2","doi-asserted-by":"crossref","DOI":"10.1016\/j.patter.2022.100616","article-title":"Audio self-supervised learning: A survey","volume":"3","author":"Liu S","year":"2022","unstructured":"Liu S, Mallol-Ragolta A, Parada-Cabaleiro E, Qian K, Jing X, Kathan A, Hu B, Schuller BW. Audio self-supervised learning: A survey. Patterns. 2022;3(12): Article 100616.","journal-title":"Patterns"},{"key":"e_1_3_2_37_2","unstructured":"Baevski A Zhou H Mohamed A Auli M. wav2vec 2.0: A framework for self-supervised learning of speech representations. arXiv. 2020. https:\/\/arxiv.org\/abs\/2006.11477"},{"key":"e_1_3_2_38_2","unstructured":"Baevski A Schneider S Auli M. vq-wav2vec: Self-supervised learning of discrete speech representations. arXiv. 2019. https:\/\/arxiv.org\/abs\/1910.05453"},{"key":"e_1_3_2_39_2","doi-asserted-by":"crossref","unstructured":"Hsu W-N Bolte B Tsai Y-HH Lakhotia K Salakhutdinov R Mohamed A. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. arXiv. 2021. https:\/\/arxiv.org\/abs\/2106.07447","DOI":"10.1109\/TASLP.2021.3122291"},{"key":"e_1_3_2_40_2","doi-asserted-by":"crossref","unstructured":"Diwan A Choi E Harwath D. When to use efficient self attention? profiling text speech and image transformer variants. arXiv. 2023. https:\/\/arxiv.org\/abs\/2306.08667","DOI":"10.18653\/v1\/2023.acl-short.141"},{"key":"e_1_3_2_41_2","unstructured":"Becker S Ackermann M Lapuschkin S M\u00fcller KR Samek W. Interpreting and explaining deep neural networks for classification of audio signals. arXiv. 2019. https:\/\/arxiv.org\/abs\/1807.03418"},{"key":"e_1_3_2_42_2","unstructured":"Frommholz A Seipel F Lapuschkin S Samek WJ Vielhaben. Xai-based comparison of input representations for audio event classification. arXiv. 2023. https:\/\/arxiv.org\/abs\/2304.14019"},{"key":"e_1_3_2_43_2","doi-asserted-by":"crossref","unstructured":"Vielhaben J Lapuschkin S Montavon G. Samek W. Explainable AI for time series via virtual inspection layers. arXiv. 2023. https:\/\/arxiv.org\/abs\/2303.06365","DOI":"10.2139\/ssrn.4399242"},{"key":"e_1_3_2_44_2","first-page":"1342","article-title":"CoughLIME: Sonified explanations for the predictions of COVID-19 cough classifiers","volume":"2022","author":"Wullenweber A","year":"2022","unstructured":"Wullenweber A, Akman A, Schuller BW. CoughLIME: Sonified explanations for the predictions of COVID-19 cough classifiers. Annu Int Conf IEEE Eng Med Biol Soc. 2022;2022:1342\u20131345.","journal-title":"Annu Int Conf IEEE Eng Med Biol Soc"},{"key":"e_1_3_2_45_2","unstructured":"Haunschmid V Manilow E Widmer G. Audiolime: Listenable explanations using source separation. arXiv. 2020. https:\/\/arxiv.org\/abs\/2008.00582"},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","first-page":"2154","DOI":"10.21105\/joss.02154","article-title":"Spleeter: A fast and efficient music source separation tool with pre-trained models","volume":"5","author":"Hennequin R","year":"2020","unstructured":"Hennequin R, Khlif A, Voituret F, Moussallam M. Spleeter: A fast and efficient music source separation tool with pre-trained models. J Open Source Softw. 2020;5:2154.","journal-title":"J Open Source Softw"},{"key":"e_1_3_2_47_2","doi-asserted-by":"crossref","unstructured":"Parekh J Parekh S Mozharovskyi P Richard G. Listen to interpret: Post-hoc interpretability for audio networks with nmf. arXiv. 2022. https:\/\/arxiv.org\/abs\/2202.11479","DOI":"10.31219\/osf.io\/4rtjs"},{"key":"e_1_3_2_48_2","doi-asserted-by":"crossref","unstructured":"Wu X Bell P. Rajan A. Explanations for automatic speech recognition. arXiv. 2023. https:\/\/arxiv.org\/abs\/2302.14062","DOI":"10.1109\/ICASSP49357.2023.10094635"},{"key":"e_1_3_2_49_2","unstructured":"Sun Y Chockler H Huang X Kroening D. Explaining deep neural networks using spectrum-based fault localization. arXiv. 2019. https:\/\/arxiv.org\/abs\/1908.02374"},{"key":"e_1_3_2_50_2","doi-asserted-by":"crossref","unstructured":"Chockler H Kroening D Sun Y. Compositional explanations for image classifiers. arXiv. 2021. https:\/\/arxiv.org\/abs\/2103.03622","DOI":"10.1109\/ICCV48922.2021.00127"},{"key":"e_1_3_2_51_2","unstructured":"Pruthi G Liu F Sundararajan M Kale S. Estimating training data influence by tracking gradient descent. arXiv. 2020. https:\/\/arxiv.org\/abs\/2002.08484"},{"key":"e_1_3_2_52_2","doi-asserted-by":"crossref","unstructured":"Salamon J Jacoby C Bello JP. A dataset and taxonomy for urban sound research. Paper presented at: Proceedings of the 22nd ACM International Conference on Multimedia MM \u201914; 2014; Orlando Florida USA. p. 1041\u20131044.","DOI":"10.1145\/2647868.2655045"},{"key":"e_1_3_2_53_2","doi-asserted-by":"crossref","unstructured":"Muguli A Pinto L Nirmala R Sharma N Krishnan P Ghosh PK Kumar R Bhat S Chetupalli SR Ganapathy S et\u00a0al. Dicova challenge: Dataset task and baseline system for covid-19 diagnosis using acoustics. arXiv. 2021. https:\/\/arxiv.org\/abs\/2103.09148","DOI":"10.21437\/Interspeech.2021-74"},{"key":"e_1_3_2_54_2","unstructured":"Bertin-Mahieux T Ellis DP Whitman B Lamere P. The million song dataset. Paper presented at: Proceedings of the 12th International Conference on Music Information Retrieval (ISMIR 2011); 2011 Oct 24\u201328; Miami Florida USA."},{"key":"e_1_3_2_55_2","unstructured":"Piczak KJ Dataset for environmental sound classification. Paper presented at: Proceedings of the 23rd Annual ACM Conference on Multimedia; 2015; Brisbane Australia. p. 1015\u20131018."},{"key":"e_1_3_2_56_2","unstructured":"Cartwright M Cramer J Mendez AEM Wang Y Wu H-H Lostanlen V Fuentes M Dove G Mydlarz C Salamon J. SONYC-UST-V2: An urban sound tagging dataset with spatiotemporal context. arXiv. 2020. https:\/\/arxiv.org\/abs\/2009.05188"},{"key":"e_1_3_2_57_2","unstructured":"Ardila R Branson M Davis K Henretty M Kohler M Meyer J Morais R Saunders L Tyers FM Weber G. Common voice: A massively-multilingual speech corpus. arXiv. 2019. https:\/\/arxiv.org\/abs\/1912.06670"}],"container-title":["Intelligent Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/spj.science.org\/doi\/pdf\/10.34133\/icomputing.0074","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,2,12]],"date-time":"2024-02-12T17:14:57Z","timestamp":1707758097000},"score":1,"resource":{"primary":{"URL":"https:\/\/spj.science.org\/doi\/10.34133\/icomputing.0074"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1]]},"references-count":56,"alternative-id":["10.34133\/icomputing.0074"],"URL":"https:\/\/doi.org\/10.34133\/icomputing.0074","relation":{},"ISSN":["2771-5892"],"issn-type":[{"value":"2771-5892","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1]]},"assertion":[{"value":"2023-07-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-28","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"0074"}}