{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,3]],"date-time":"2026-08-03T17:10:43Z","timestamp":1785777043493,"version":"3.56.0"},"reference-count":32,"publisher":"MDPI AG","issue":"15","license":[{"start":{"date-parts":[[2023,8,5]],"date-time":"2023-08-05T00:00:00Z","timestamp":1691193600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Chaoyang University of Technology, Taiwan","award":["110F0021110"],"award-info":[{"award-number":["110F0021110"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Sound classification has been widely used in many fields. Unlike traditional signal-processing methods, using deep learning technology for sound classification is one of the most feasible and effective methods. However, limited by the quality of the training dataset, such as cost and resource constraints, data imbalance, and data annotation issues, the classification performance is affected. Therefore, we propose a sound classification mechanism based on convolutional neural networks and use the sound feature extraction method of Mel-Frequency Cepstral Coefficients (MFCCs) to convert sound signals into spectrograms. Spectrograms are suitable as input for CNN models. To provide the function of data augmentation, we can increase the number of spectrograms by setting the number of triangular bandpass filters. The experimental results show that there are 50 semantic categories in the ESC-50 dataset, the types are complex, and the amount of data is insufficient, resulting in a classification accuracy of only 63%. When using the proposed data augmentation method (K = 5), the accuracy is effectively increased to 97%. Furthermore, in the UrbanSound8K dataset, the amount of data is sufficient, so the classification accuracy can reach 90%, and the classification accuracy can be slightly increased to 92% via data augmentation. However, when only 50% of the training dataset is used, along with data augmentation, the establishment of the training model can be accelerated, and the classification accuracy can reach 91%.<\/jats:p>","DOI":"10.3390\/s23156972","type":"journal-article","created":{"date-parts":[[2023,8,5]],"date-time":"2023-08-05T10:25:36Z","timestamp":1691231136000},"page":"6972","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":41,"title":["A CNN Sound Classification Mechanism Using Data Augmentation"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1732-981X","authenticated-orcid":false,"given":"Hung-Chi","family":"Chu","sequence":"first","affiliation":[{"name":"Department of Information and Communication Engineering, Chaoyang University of Technology, Taichung 41349, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Young-Lin","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Information and Communication Engineering, Chaoyang University of Technology, Taichung 41349, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hao-Chu","family":"Chiang","sequence":"additional","affiliation":[{"name":"Department of Information and Communication Engineering, Chaoyang University of Technology, Taichung 41349, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,8,5]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Han, W., Zhang, Z., Zhang, Y., Yu, J., Chiu, C.-C., Qin, J., Gulati, A., Pang, R., and Wu, Y. (2020, January 25\u201329). ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context. Proceedings of the Interspeech 2020, 21st Annual Conference of the International Speech Communication Association, Shanghai, China.","DOI":"10.21437\/Interspeech.2020-2059"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"8337","DOI":"10.1038\/s41598-022-12260-y","article-title":"A study of transformer-based end-to-end speech recognition system for Kazakh language","volume":"12","author":"Orken","year":"2022","journal-title":"Sci. Rep."},{"key":"ref_3","first-page":"3387598","article-title":"Music Recommendation System and Recommendation Model Based on Convolutional Neural Network","volume":"2022","author":"Zhang","year":"2022","journal-title":"Mob. Inf. Syst."},{"key":"ref_4","first-page":"107","article-title":"Music recommender system based on graph convolutional neural networks with attention mechanism","volume":"135","author":"Huang","year":"2021","journal-title":"Neural Netw."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"107041","DOI":"10.1016\/j.apacoust.2019.107041","article-title":"Environmental sound monitoring using machine learning on mobile devices","volume":"159","author":"Marc","year":"2020","journal-title":"Appl. Acoust."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Nogueira, A.F.R., Oliveira, H.S., Machado, J.J.M., and Tavares, J.M.R.S. (2022). Sound Classification and Processing of Urban Environments: A Systematic Literature Review. Sensors, 22.","DOI":"10.3390\/s22228608"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Nishida, T., Dohi, K., Endo, T., Yamamoto, M., and Kawaguchi, Y. (September, January 29). Anomalous Sound Detection Based on Machine Activity Detection. Proceedings of the 2022 30th European Signal Processing Conference (EUSIPCO), Belgrade, Serbia.","DOI":"10.23919\/EUSIPCO55093.2022.9909901"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Wang, Y., Zheng, Y., Zhang, Y., Xie, Y., Xu, S., Hu, Y., and He, L. (2021). Unsupervised Anomalous Sound Detection for Machine Condition Monitoring Using Classification-Based Methods. Appl. Sci., 11.","DOI":"10.3390\/app112311128"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Shriram, K. (2021). Vasudevan, Sini Raj Pulari, Subashri Vasudevan, Deep Learning: A Comprehensive Guide, Chapman & Hall. [1st ed.].","DOI":"10.1201\/9781003185635"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Viarbitskaya, T., and Dobrucki, A. (2018, January 19\u201321). Audio processing with using Python language science libraries. Proceedings of the 2018 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA), Poznan, Poland.","DOI":"10.23919\/SPA.2018.8563430"},{"key":"ref_11","unstructured":"Eric, W. (2014). Hansen, Fourier Transforms: Principles and Applications, Wiley."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"153","DOI":"10.1016\/j.dsp.2007.12.004","article-title":"Time-frequency feature representation using energy concentration: An overview of recent advances","volume":"19","author":"Jiang","year":"2009","journal-title":"Digit. Signal Process."},{"key":"ref_13","unstructured":"Franzese, M., Iuliano, A., and Models, H.M. (2019). Encyclopedia of Bioinformatics and Computational Biology, Elsevier."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Wan, H., Wang, H., Scotney, B., and Liu, J. (2019, January 6\u20139). A Novel Gaussian Mixture Model for Classification. Proceedings of the 2019 IEEE International Conference on Systems, Man and Cybernetics (SMC), Bari, Italy.","DOI":"10.1109\/SMC.2019.8914215"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1685","DOI":"10.1109\/LSP.2015.2424991","article-title":"Location Estimation of Predominant Sound Source with Embedded Source Separation in Amplitude-Panned Stereo Signal","volume":"22","author":"Han","year":"2015","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"P\u00f3\u0142rolniczak, E., and Kramarczyk, M. (2017, January 20\u201322). Estimation of singing voice types based on voice parameters analysis. Proceedings of the 2017 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA), Poznan, Poland.","DOI":"10.23919\/SPA.2017.8166839"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Thwe, K.Z., and War, N. (2017, January 26\u201328). Environmental sound classification based on time-frequency representation. Proceedings of the 2017 18th IEEE\/ACIS International Conference on Software Engineering, Artificial Intelligence, Networking and Parallel\/Distributed Computing (SNPD), Kanazawa, Japan.","DOI":"10.1109\/SNPD.2017.8022729"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"122136","DOI":"10.1109\/ACCESS.2022.3223444","article-title":"Mel Frequency Cepstral Coefficient and its Applications: A Review","volume":"10","author":"Zrar","year":"2022","journal-title":"IEEE Access"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1016\/j.neunet.2014.09.003","article-title":"Deep learning in neural networks: An overview","volume":"61","author":"Schmidhuber","year":"2015","journal-title":"Neural Netw."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"175353","DOI":"10.1109\/ACCESS.2019.2957572","article-title":"Investigation of Different CNN-Based Models for Improved Bird Sound Classification","volume":"7","author":"Xie","year":"2019","journal-title":"IEEE Access"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Das, J.K., Ghosh, A., Pal, A.K., Dutta, S., and Chakrabarty, A. (2020, January 21\u201323). Urban Sound Classification Using Convolutional Neural Network and Long Short Term Memory Based on Multiple Features. Proceedings of the 2020 Fourth International Conference On Intelligent Computing in Data Sciences (ICDS), Fez, Morocco.","DOI":"10.1109\/ICDS50568.2020.9268723"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Chi, Z., Li, Y., and Chen, C. (2019, January 19\u201320). Deep Convolutional Neural Network Combined with Concatenated Spectrogram for Environmental Sound Classification. Proceedings of the 2019 IEEE 7th International Conference on Computer Science and Network Technology (ICCSNT), Dalian, China.","DOI":"10.1109\/ICCSNT47585.2019.8962462"},{"key":"ref_23","unstructured":"Piczak, K.J. (2023, July 08). ESC-50: Dataset for Environmental Sound Classification. Available online: https:\/\/github.com\/karolpiczak\/ESC-50."},{"key":"ref_24","unstructured":"Salamon, J., Jacoby, C., and Bello, J.P. (2023, July 08). Urban Sound Datasets. Available online: https:\/\/urbansounddataset.weebly.com\/urbansound8k.html."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"109025","DOI":"10.1016\/j.patcog.2022.109025","article-title":"Environmental Sound Classification on the Edge: A Pipeline for Deep Acoustic Networks on Extremely Resource-Constrained Devices","volume":"133","author":"Mohaimenuzzaman","year":"2023","journal-title":"Pattern Recognit."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Takahashi, N., Gygli, M., Pfister, B., and Gool, L.V. (2016, January 8\u201312). Deep Convolutional Neural Networks and Data Augmentation for Acoustic Event Recognition. Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, San Francisco, CA, USA.","DOI":"10.21437\/Interspeech.2016-805"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1109\/LSP.2017.2657381","article-title":"Deep convolutional neural networks and data augmentation for environmental sound classification","volume":"24","author":"Salamon","year":"2017","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Nam, H., Kim, S.-H., and Park, Y.-H. (2022, January 23\u201327). Filteraugment: An Acoustic Environmental Data Augmentation Method. Proceedings of the ICASSP 2022\u20132022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore.","DOI":"10.1109\/ICASSP43922.2022.9747680"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Park, D.S., Zhang, W.Y., Chiu, C.-C., Zoph, B., Cubuk, E.D., and Le, Q.V. (2019, January 15\u201319). SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. Proceedings of the 20th Annual Conference of the International Speech Communication Association INTERSPEECH, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-2680"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Nanni, L., Maguolo, G., Brahnam, S., and Paci, M. (2021). An Ensemble of Convolutional Neural Networks for Audio Classification. Appl. Sci., 11.","DOI":"10.3390\/app11135796"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Garcia-Balboa, J.L., Alba-Fernandez, M.V., Ariza-L\u00f3pez, F.J., and Rodriguez-Avi, J. (2018, January 22\u201327). Homogeneity Test for Confusion Matrices: A Method and an Example. Proceedings of the IGARSS 2018\u20132018 IEEE International Geoscience and Remote Sensing Symposium, Valencia, Spain.","DOI":"10.1109\/IGARSS.2018.8517924"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Chicco, D., and Jurman, G. (2020). The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genom., 21.","DOI":"10.1186\/s12864-019-6413-7"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/15\/6972\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:26:33Z","timestamp":1760127993000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/15\/6972"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,5]]},"references-count":32,"journal-issue":{"issue":"15","published-online":{"date-parts":[[2023,8]]}},"alternative-id":["s23156972"],"URL":"https:\/\/doi.org\/10.3390\/s23156972","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,5]]}}}