{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T19:01:34Z","timestamp":1782846094089,"version":"3.54.5"},"reference-count":37,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2025,10,5]],"date-time":"2025-10-05T00:00:00Z","timestamp":1759622400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>This paper presents an automatic system for the classification of musical instruments from audio recordings. The project leverages deep learning (DL) techniques to achieve its objective, exploring three different classification approaches based on distinct input representations. The first method involves the extraction of Mel-Frequency Cepstral Coefficients (MFCCs) from the audio files, which are then fed into a two-dimensional convolutional neural network (Conv2D). The second approach makes use of mel-spectrogram images as input to a similar Conv2D architecture. The third approach employs conventional machine learning (ML) classifiers, including Logistic Regression, K-Nearest Neighbors, and Random Forest, trained on MFCC-derived feature vectors. To gain insight into the behavior of the DL model, explainability techniques were applied to the Conv2D model using mel-spectrograms, allowing for a better understanding of how the network interprets relevant features for classification. Additionally, t-distributed stochastic neighbor embedding (t-SNE) was employed on the MFCC vectors to visualize how instrument classes are organized in the feature space. One of the main challenges encountered was the class imbalance within the dataset, which was addressed by assigning class-specific weights during training. The results, in terms of classification accuracy, were very satisfactory across all approaches, with the convolutional models and Random Forest achieving around 97\u201398%, and Logistic Regression yielding slightly lower performance. In conclusion, the proposed methods proved effective for the selected dataset, and future work may focus on further improving class balance techniques.<\/jats:p>","DOI":"10.3390\/info16100864","type":"journal-article","created":{"date-parts":[[2025,10,6]],"date-time":"2025-10-06T15:05:06Z","timestamp":1759763106000},"page":"864","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Explainable Instrument Classification: From MFCC Mean-Vector Models to CNNs on MFCC and Mel-Spectrograms with t-SNE and Grad-CAM Insights"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-5521-7315","authenticated-orcid":false,"given":"Tommaso","family":"Senatori","sequence":"first","affiliation":[{"name":"Department of Civil, Computer Science and Aeronautical Technologies Engineering, Universit\u00e0 degli Studi Roma Tre, Via V. Volterra 62, 00146 Roma, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-1279-4966","authenticated-orcid":false,"given":"Daniela","family":"Nardone","sequence":"additional","affiliation":[{"name":"Department of Civil, Computer Science and Aeronautical Technologies Engineering, Universit\u00e0 degli Studi Roma Tre, Via V. Volterra 62, 00146 Roma, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2660-0405","authenticated-orcid":false,"given":"Michele","family":"Lo Giudice","sequence":"additional","affiliation":[{"name":"Department of Civil, Computer Science and Aeronautical Technologies Engineering, Universit\u00e0 degli Studi Roma Tre, Via V. Volterra 62, 00146 Roma, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5825-1019","authenticated-orcid":false,"given":"Alessandro","family":"Salvini","sequence":"additional","affiliation":[{"name":"Department of Civil, Computer Science and Aeronautical Technologies Engineering, Universit\u00e0 degli Studi Roma Tre, Via V. Volterra 62, 00146 Roma, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,10,5]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Saggese, A., Strisciuglio, N., Vento, M., and Petkov, N. (2016, January 23\u201326). Time-frequency analysis for audio event detection in real scenarios. Proceedings of the 2016 13th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), Colorado Springs, CO, USA.","DOI":"10.1109\/AVSS.2016.7738082"},{"key":"ref_2","unstructured":"Esmaili, S., Krishnan, S., and Raahemifar, K. (2004, January 17\u201321). Content based audio classification and retrieval using joint time-frequency analysis. Proceedings of the 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing, Montreal, QC, Canada."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"3580","DOI":"10.1121\/10.0026219","article-title":"Speech preprocessing and enhancement based on joint time domain and time-frequency domain analysis","volume":"155","author":"Zhang","year":"2024","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"78","DOI":"10.1016\/j.culher.2024.07.012","article-title":"Deep learning for the detection and classification of adhesion defects in antique plaster layers","volume":"69","author":"Mariani","year":"2024","journal-title":"J. Cult. Herit."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Lo Giudice, M., Mariani, F., Caliano, G., and Salvini, A. (2025). Enhancing Defect Detection on Surfaces Using Transfer Learning and Acoustic Non-Destructive Testing. Information, 16.","DOI":"10.3390\/info16070516"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"122136","DOI":"10.1109\/ACCESS.2022.3223444","article-title":"Mel Frequency Cepstral Coefficient and Its Applications: A Review","volume":"10","author":"Abdul","year":"2022","journal-title":"IEEE Access"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Allamy, S., and Koerich, A.L. (2021, January 5\u20137). 1D CNN architectures for music genre classification. Proceedings of the 2021 IEEE Symposium Series on Computational Intelligence (SSCI), Orlando, FL, USA.","DOI":"10.1109\/SSCI50451.2021.9659979"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"106620","DOI":"10.1109\/ACCESS.2023.3318015","article-title":"A survey of audio classification using deep learning","volume":"11","author":"Zaman","year":"2023","journal-title":"IEEE Access"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Bose, A., and Tripathy, B. (2020). Deep learning for audio signal classification. Deep Learning\u2014Research and Applications, De Gruyter.","DOI":"10.1515\/9783110670905-006"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"9","DOI":"10.21512\/ijcshai.v2i1.13019","article-title":"Systematic Literature Review of The Use of Music Information Retrieval in Music Genre Classification","volume":"2","author":"Budyputra","year":"2025","journal-title":"Int. J. Comput. Sci. Humanit. AI"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"2372","DOI":"10.1016\/j.procs.2025.04.500","article-title":"Enhanced Audio Signal Classification with Explainable AI: Deep Learning Approach in Time and Frequency Domain Analysis","volume":"258","author":"Jenifer","year":"2025","journal-title":"Procedia Comput. Sci."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"418","DOI":"10.1016\/j.jfranklin.2023.11.038","article-title":"AudioMNIST: Exploring Explainable Artificial Intelligence for audio analysis on a simple benchmark","volume":"361","author":"Becker","year":"2024","journal-title":"J. Frankl. Inst."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Lo Giudice, M., Mammone, N., Ieracitano, C., Aguglia, U., Mandic, D., and Morabito, F.C. (2022, January 1\u20133). Explainable deep learning classification of respiratory sound for telemedicine applications. Proceedings of the International Conference on Applied Intelligence and Informatics, Reggio Calabria, Italy.","DOI":"10.1007\/978-3-031-24801-6_28"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"0074","DOI":"10.34133\/icomputing.0074","article-title":"Audio Explainable Artificial Intelligence: A Review","volume":"3","author":"Akman","year":"2024","journal-title":"Intell. Comput."},{"key":"ref_15","unstructured":"Kaminskyj, I. (December, January 27). Automatic Source Identification of Monophonic Musical Instrument Sounds. Proceedings of the IEEE International Conference on Neural Networks, Perth, Australia."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1933","DOI":"10.1121\/1.426728","article-title":"Computer identification of musical instruments using pattern recognition with cepstral coefficients as features","volume":"105","author":"Brown","year":"1999","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_17","first-page":"143","article-title":"A Study of Musical Instrument Classification Using Gaussian Mixture Models and Support Vector Machines","volume":"4","author":"Marques","year":"1999","journal-title":"Camb. Res. Lab. Tech. Rep. Ser. CRL"},{"key":"ref_18","unstructured":"Eronen, A., and Klapuri, A. (2000, January 5\u20139). Musical Instrument Recognition Using Cepstral Coefficients and Temporal Features. Proceedings of the 2000 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Istanbul, Turkey."},{"key":"ref_19","first-page":"768","article-title":"Representing Musical Instrument Sounds for Their Automatic Classification","volume":"49","author":"Kostek","year":"2001","journal-title":"J. Audio Eng. Soc."},{"key":"ref_20","unstructured":"Livshin, A., and Rodet, X. (2003, January 22). The Importance of Cross-Database Evaluation in Sound Classification. Proceedings of the 4th International Conference on Music Information Retrieval (ISMIR), Washington, DC, USA."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"943279","DOI":"10.1155\/S1110865703210118","article-title":"Musical Instrument Timbres Classification with Spectral Features","volume":"2003","author":"Agostini","year":"2003","journal-title":"EURASIP J. Adv. Signal Process."},{"key":"ref_22","unstructured":"Krishna, A.G., and Sreenivas, T.V. (2004, January 17\u201321). Music Instrument Recognition: From Isolated Notes to Solo Phrases. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Montreal, QC, Canada."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1401","DOI":"10.1109\/TSA.2005.860842","article-title":"Musical instrument recognition by pairwise classification strategies","volume":"14","author":"Essid","year":"2006","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_24","unstructured":"Diment, A., Rajan, P., Heittola, T., and Virtanen, T. (2013, January 15\u201318). Modified Group Delay Feature for Musical Instrument Recognition. Proceedings of the 10th International Symposium on Computer Music Multidisciplinary Research (CMMR), Marseille, France."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Yu, L., Su, L., and Yang, Y. (2014, January 4\u20139). Sparse Cepstral Codes and Power Scale for Instrument Identification. Proceedings of the 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6855050"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"208","DOI":"10.1109\/TASLP.2016.2632307","article-title":"Deep convolutional neural networks for predominant instrument recognition in polyphonic music","volume":"25","author":"Han","year":"2017","journal-title":"Ieee\/Acm Trans. Audio Speech Lang. Process."},{"key":"ref_27","unstructured":"Lee, J., Kim, T., Park, J., and Nam, J. (2017). Raw Waveform-based Audio Classification Using Sample-level CNN Architectures. arXiv."},{"key":"ref_28","unstructured":"Haidar-Ahmad, L. (2025, October 01). Music and Instrument Classification using Deep Learning Techniques. Available online: https:\/\/cs230.stanford.edu\/projects_fall_2019\/reports\/26225883.pdf."},{"key":"ref_29","unstructured":"Gururani, S., Sharma, M., and Lerch, A. (2019). An Attention Mechanism for Musical Instrument Recognition. arXiv."},{"key":"ref_30","first-page":"1659","article-title":"Music instrument recognition using deep convolutional neural networks","volume":"14","author":"Solanki","year":"2022","journal-title":"Int. J. Inf. Technol."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Blaszke, M., and Kostek, B. (2022). Musical Instrument Identification Using Deep Learning Approach. Sensors, 22.","DOI":"10.3390\/s22083033"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1186\/s13636-022-00245-8","article-title":"Transformer-based ensemble method for multiple predominant instruments recognition in polyphonic music","volume":"2022","author":"Reghunath","year":"2022","journal-title":"EURASIP J. Audio Speech Music Process."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1499913","DOI":"10.3389\/frai.2024.1499913","article-title":"Interpreting CNN models for musical instrument recognition using multi-spectrogram heatmap analysis: A preliminary study","volume":"7","author":"Chen","year":"2024","journal-title":"Front. Artif. Intell."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"18","DOI":"10.25080\/Majora-7b98e3ed-003","article-title":"librosa: Audio and music signal analysis in python","volume":"2015","author":"McFee","year":"2015","journal-title":"SciPy"},{"key":"ref_35","unstructured":"Logan, B. (2000, January 23\u201325). Mel Frequency Cepstral Coefficients for Music Modeling. Proceedings of the 1st International Symposium Music Information Retrieval, Plymouth, MA, USA."},{"key":"ref_36","unstructured":"Roberts, L. (2025, October 01). Understanding the Mel Spectrogram. Available online: https:\/\/medium.com\/analytics-vidhya\/understanding-the-mel-spectrogram-fca2afa2ce53."},{"key":"ref_37","first-page":"351","article-title":"Deep Neural Network for Musical Instrument Recognition Using MFCCs","volume":"25","author":"Mahanta","year":"2021","journal-title":"Comput. Y Sist."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/10\/864\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T10:07:43Z","timestamp":1760004463000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/10\/864"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,5]]},"references-count":37,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2025,10]]}},"alternative-id":["info16100864"],"URL":"https:\/\/doi.org\/10.3390\/info16100864","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,5]]}}}