{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,8]],"date-time":"2026-08-08T18:06:36Z","timestamp":1786212396335,"version":"3.56.0"},"reference-count":87,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2025,3,15]],"date-time":"2025-03-15T00:00:00Z","timestamp":1741996800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,3,15]],"date-time":"2025-03-15T00:00:00Z","timestamp":1741996800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100009410","name":"Universitat de Lleida","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100009410","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Intell Rev"],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Sound recognition has a wide range of applications beyond speech and music, including environmental monitoring, sound source classification, mechanical fault diagnosis, audio fingerprinting, and event detection. These applications often require real-time data processing, making them well-suited for embedded systems. However, embedded devices face significant challenges due to limited computational power, memory, and low power consumption. Despite these constraints, achieving high performance in environmental sound recognition typically requires complex algorithms. Deep Learning models have demonstrated high accuracy on existing datasets, making them a popular choice for such tasks. However, these models are resource-intensive, posing challenges for real-time edge applications. This paper presents a comprehensive review of integrating Deep Learning models into embedded systems, examining their state-of-the-art applications, key components, and steps involved. It also explores strategies to optimise performance in resource-constrained environments through a comparison of various implementation approaches such as knowledge distillation, pruning, and quantization, with studies achieving a reduction in complexity of up to 97% compared to the unoptimized model. Overall, we conclude that in spite of the availability of lightweight deep learning models, input features, and compression techniques, their integration into low-resource devices, such as microcontrollers, remains limited. Furthermore, more complex tasks, such as general sound classification, especially with expanded frequency bands and real-time operation have yet to be effectively implemented on these devices. These findings highlight the need for a standardised research framework to evaluate these technologies applied to resource-constrained devices, and for further development to realise the wide range of potential applications.<\/jats:p>","DOI":"10.1007\/s10462-025-11106-z","type":"journal-article","created":{"date-parts":[[2025,3,15]],"date-time":"2025-03-15T05:08:57Z","timestamp":1742015337000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Environmental sound recognition on embedded devices using deep learning: a review"],"prefix":"10.1007","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-8265-4617","authenticated-orcid":false,"given":"Pau","family":"Gair\u00ed","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8495-8643","authenticated-orcid":false,"given":"Tom\u00e0s","family":"Pallej\u00e0","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7965-0086","authenticated-orcid":false,"given":"Marcel","family":"Tresanchez","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,3,15]]},"reference":[{"key":"11106_CR1","doi-asserted-by":"publisher","first-page":"1155","DOI":"10.3390\/electronics9071155","volume":"9","author":"RM Alsina-Pag\u00e8s","year":"2020","unstructured":"Alsina-Pag\u00e8s RM, Herv\u00e1s M, Duboc L, Carbassa J (2020) Design of a low-cost configurable acoustic sensor for the rapid development of sound recognition applications. Electronics 9:1155. https:\/\/doi.org\/10.3390\/electronics9071155","journal-title":"Electronics"},{"key":"11106_CR2","doi-asserted-by":"publisher","first-page":"661","DOI":"10.1080\/08839514.2018.1430469","volume":"31","author":"E Babaee","year":"2017","unstructured":"Babaee E, Anuar NB, Abdul Wahab AW, Shamshirband S, Chronopoulos AT (2017) An overview of audio event detection methods from feature extraction to classification. Appl Artif Intell 31:661\u2013714. https:\/\/doi.org\/10.1080\/08839514.2018.1430469","journal-title":"Appl Artif Intell"},{"key":"11106_CR3","doi-asserted-by":"crossref","unstructured":"Bahai A (2024) Making sense at the edge. In: 2024 IEEE symposium on VLSI technology and circuits (VLSI Technology and Circuits). pp 1\u20132. https:\/\/doi.org\/10.1109\/VLSITECHNOLOGYANDCIR46783.2024.10631378","DOI":"10.1109\/VLSITechnologyandCir46783.2024.10631378"},{"key":"11106_CR4","doi-asserted-by":"publisher","first-page":"200115","DOI":"10.1016\/j.iswa.2022.200115","volume":"16","author":"A Bansal","year":"2022","unstructured":"Bansal A, Garg NK (2022) Environmental sound classification: a descriptive review of the literature. Intell Syst Appl 16:200115. https:\/\/doi.org\/10.1016\/j.iswa.2022.200115","journal-title":"Intell Syst Appl"},{"key":"11106_CR5","doi-asserted-by":"publisher","first-page":"3876","DOI":"10.1109\/JIOT.2018.2845099","volume":"5","author":"I Bisio","year":"2018","unstructured":"Bisio I, Delfino A, Grattarola A, Lavagetto F, Sciarrone A (2018) Ultrasounds-based context sensing method and applications over the internet of things. IEEE Internet Things J 5:3876\u20133890. https:\/\/doi.org\/10.1109\/JIOT.2018.2845099","journal-title":"IEEE Internet Things J"},{"key":"11106_CR6","doi-asserted-by":"publisher","first-page":"475","DOI":"10.1038\/s41597-024-03301-4","volume":"11","author":"J Branding","year":"2024","unstructured":"Branding J, von H\u00f6rsten D, B\u00f6ckmann E, Wegener JK, Hartung E (2024) InsectSound1000 an insect sound dataset for deep learning based acoustic insect recognition. Sci Data 11:475. https:\/\/doi.org\/10.1038\/s41597-024-03301-4","journal-title":"Sci Data"},{"key":"11106_CR7","doi-asserted-by":"crossref","unstructured":"Brighente A, Conti M, Peruzzi G, Pozzebon A (2023) ADASS: anti-drone audio surveillance sentinel via embedded machine learning. In: 2023 IEEE sensors applications symposium (SAS). pp 1\u20136. https:\/\/doi.org\/10.1109\/SAS58821.2023.10254008","DOI":"10.1109\/SAS58821.2023.10254008"},{"key":"11106_CR8","doi-asserted-by":"publisher","unstructured":"Canziani A, Paszke A, Culurciello E (2017) An Analysis of Deep Neural Network Models for Practical Applications. ArXiv.  https:\/\/doi.org\/10.48550\/arXiv.1605.07678","DOI":"10.48550\/arXiv.1605.07678"},{"key":"11106_CR9","doi-asserted-by":"crossref","unstructured":"Cao S, Li D, Lee SI, Xiong J (2023) PowerPhone: unleashing the acoustic sensing capability of smartphones. In: Proceedings of the 29th annual international conference on mobile computing and networking. Association for Computing Machinery, New York, pp 1\u201316. https:\/\/doi.org\/10.1145\/3570361.3613270","DOI":"10.1145\/3570361.3613270"},{"key":"11106_CR10","doi-asserted-by":"publisher","first-page":"654","DOI":"10.1109\/JSTSP.2020.2969775","volume":"14","author":"G Cerutti","year":"2020","unstructured":"Cerutti G, Prasad R, Brutti A, Farella E (2020) Compact recurrent neural networks for acoustic event detection on low-energy low-complexity platforms. IEEE J Sel Top Signal Process 14:654\u2013664. https:\/\/doi.org\/10.1109\/JSTSP.2020.2969775","journal-title":"IEEE J Sel Top Signal Process"},{"key":"11106_CR11","doi-asserted-by":"publisher","first-page":"e14","DOI":"10.1017\/ATSIP.2014.12","volume":"3","author":"S Chachada","year":"2014","unstructured":"Chachada S, Kuo C-CJ (2014) Environmental sound recognition: a survey. APSIPA Trans Signal Inf Process 3:e14. https:\/\/doi.org\/10.1017\/ATSIP.2014.12","journal-title":"APSIPA Trans Signal Inf Process"},{"key":"11106_CR12","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3322240","volume":"52:63","author":"S Chandrakala","year":"2019","unstructured":"Chandrakala S, Jayalakshmi SL (2019) Environmental audio scene and sound event recognition for autonomous surveillance: a survey and comparative studies. ACM Comput Surv 52:63:1\u20136334. https:\/\/doi.org\/10.1145\/3322240","journal-title":"ACM Comput Surv"},{"key":"11106_CR13","doi-asserted-by":"publisher","unstructured":"Chen T, Moreau T, Jiang Z, Zheng L, Yan E, Cowan M, Shen H, Wang L, Hu Y, Ceze L, Guestrin C, Krishnamurthy A (2018) TVM: an automated end- to-end optimizing compiler for deep learning. ArXiv.  https:\/\/doi.org\/10.48550\/arXiv.1802.04799","DOI":"10.48550\/arXiv.1802.04799"},{"key":"11106_CR14","doi-asserted-by":"publisher","unstructured":"Choudhary S, Karthik CR, Lakshmi PS, Kumar S (2022) LEAN: light and efficient audio classification network. In: 2022 IEEE 19th India council international conference (INDICON). pp 1\u20136. https:\/\/doi.org\/10.1109\/INDICON56171.2022.10039921","DOI":"10.1109\/INDICON56171.2022.10039921"},{"key":"11106_CR15","doi-asserted-by":"publisher","first-page":"23","DOI":"10.3390\/informatics7030023","volume":"7","author":"G Ciaburro","year":"2020","unstructured":"Ciaburro G, Iannace G (2020) Improving Smart cities Safety using sound events detection based on deep neural network algorithms. Informatics 7:23. https:\/\/doi.org\/10.3390\/informatics7030023","journal-title":"Informatics"},{"key":"11106_CR16","doi-asserted-by":"publisher","first-page":"197","DOI":"10.1561\/2000000039","volume":"7","author":"L Deng","year":"2014","unstructured":"Deng L, Yu D (2014) Deep learning: methods and applications. Found Trends\u00ae Signal Process 7:197\u2013387. https:\/\/doi.org\/10.1561\/2000000039","journal-title":"Found Trends\u00ae Signal Process"},{"key":"11106_CR17","doi-asserted-by":"publisher","DOI":"10.15837\/ijccc.2024.4.6632","author":"M Doinea","year":"2024","unstructured":"Doinea M, Trandafir I, Toma C-V, Popa M, Zamfiroiu A (2024) IoT embedded smart monitoring system with edge machine learning for beehive management. Int J Comput Commun Control. https:\/\/doi.org\/10.15837\/ijccc.2024.4.6632","journal-title":"Int J Comput Commun Control"},{"key":"11106_CR18","doi-asserted-by":"publisher","first-page":"102637","DOI":"10.1016\/j.ecoinf.2024.102637","volume":"81","author":"L Duan","year":"2024","unstructured":"Duan L, Yang L, Guo Y (2024) SIAlex: species identification and monitoring based on bird sound features. Ecol Inf 81:102637. https:\/\/doi.org\/10.1016\/j.ecoinf.2024.102637","journal-title":"Ecol Inf"},{"key":"11106_CR19","doi-asserted-by":"publisher","first-page":"100848","DOI":"10.1016\/j.iot.2023.100848","volume":"23","author":"SS Hammad","year":"2023","unstructured":"Hammad SS, Iskandaryan D, Trilles S (2023) An unsupervised TinyML approach applied to the detection of urban noise anomalies under the smart cities environment. Internet Things 23:100848. https:\/\/doi.org\/10.1016\/j.iot.2023.100848","journal-title":"Internet Things"},{"key":"11106_CR20","doi-asserted-by":"publisher","first-page":"109833","DOI":"10.1016\/j.apacoust.2023.109833","volume":"217","author":"X Han","year":"2024","unstructured":"Han X, Peng J (2024) Bird sound detection based on sub-band features and the perceptron model. Appl Acoust 217:109833. https:\/\/doi.org\/10.1016\/j.apacoust.2023.109833","journal-title":"Appl Acoust"},{"key":"11106_CR21","doi-asserted-by":"publisher","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp 770\u2013778.  https:\/\/doi.org\/10.1109\/CVPR.2016.90","DOI":"10.1109\/CVPR.2016.90"},{"key":"11106_CR22","doi-asserted-by":"publisher","first-page":"25791","DOI":"10.1109\/JIOT.2022.3199085","volume":"9","author":"C He","year":"2022","unstructured":"He C, Tan J, Jian X, Zhong G, Wu H, Cheng L, Lin J (2022) A novel snore detection and suppression method for a flexible patch with MEMS microphone and accelerometer. IEEE Internet Things J 9:25791\u201325804. https:\/\/doi.org\/10.1109\/JIOT.2022.3199085","journal-title":"IEEE Internet Things J"},{"key":"11106_CR23","doi-asserted-by":"publisher","unstructured":"Hinton G, Vinyals O, Dean J (2015) Distilling the knowledge in a neural network. ArXiv.  https:\/\/doi.org\/10.48550\/arXiv.1503.02531","DOI":"10.48550\/arXiv.1503.02531"},{"key":"11106_CR24","doi-asserted-by":"publisher","first-page":"3630","DOI":"10.3390\/s23073630","volume":"23","author":"L Hou","year":"2023","unstructured":"Hou L, Duan W, Xuan G, Xiao S, Li Y, Li Y, Zhao J (2023) Intelligent microsystem for sound event recognition in edge computing using end-to-end mesh networking. Sensors 23:3630. https:\/\/doi.org\/10.3390\/s23073630","journal-title":"Sensors"},{"key":"11106_CR25","doi-asserted-by":"publisher","unstructured":"Huang Z, Tousnakhoff A, Kozyr P, Rehausen R, Bie\u00dfmann F, Lachlan R, Adjih C, Baccelli E (2024) TinyChirp: bird song recognition using TinyML models on low-power wireless acoustic sensors. In: 2024 IEEE 5th international symposium on the internet of sounds (IS2). pp 1\u201310. https:\/\/doi.org\/10.1109\/IS262782.2024.10704131","DOI":"10.1109\/IS262782.2024.10704131"},{"key":"11106_CR26","doi-asserted-by":"publisher","first-page":"550","DOI":"10.1109\/JTEHM.2024.3433448","volume":"12","author":"D Hyun Choi","year":"2024","unstructured":"Hyun Choi D, Ha Joo Y, Hong Kim K, Ho Park J, Joo H, Kong H-J, Lee H, Jun Song K, Kim S (2024) A development of a sound recognition-based cardiopulmonary resuscitation training system. IEEE J Transl Eng Health Med 12:550\u2013557. https:\/\/doi.org\/10.1109\/JTEHM.2024.3433448","journal-title":"IEEE J Transl Eng Health Med"},{"key":"11106_CR27","doi-asserted-by":"publisher","first-page":"107610","DOI":"10.1016\/j.nanoen.2022.107610","volume":"101","author":"YH Jung","year":"2022","unstructured":"Jung YH, Pham TX, Issa D, Wang HS, Lee JH, Chung M, Lee B-Y, Kim G, Yoo CD, Lee KJ (2022) Deep learning-based noise robust flexible piezoelectric acoustic sensors for speech processing. Nano Energy 101:107610. https:\/\/doi.org\/10.1016\/j.nanoen.2022.107610","journal-title":"Nano Energy"},{"key":"11106_CR28","doi-asserted-by":"publisher","first-page":"431","DOI":"10.1016\/j.jmsy.2020.12.020","volume":"58","author":"J Kim","year":"2021","unstructured":"Kim J, Lee H, Jeong S, Ahn S-H (2021) Sound-based remote real-time multi-device operational monitoring system using a convolutional neural network (CNN). J Manuf Syst 58:431\u2013441. https:\/\/doi.org\/10.1016\/j.jmsy.2020.12.020","journal-title":"J Manuf Syst"},{"key":"11106_CR29","doi-asserted-by":"publisher","first-page":"4650","DOI":"10.3390\/s22124650","volume":"22","author":"J Ko","year":"2022","unstructured":"Ko J, Kim H, Kim J (2022) Real-time sound source localization for low-power IoT devices based on Multi-stream CNN. Sensors 22:4650. https:\/\/doi.org\/10.3390\/s22124650","journal-title":"Sensors"},{"key":"11106_CR30","doi-asserted-by":"publisher","first-page":"194","DOI":"10.1016\/j.apacoust.2018.12.028","volume":"148","author":"O K\u00fcc\u0323\u00fcktopcu","year":"2019","unstructured":"K\u00fcc\u0323\u00fcktopcu O, Masazade E, \u00dcnsalan C, Varshney PK (2019) A real-time bird sound recognition system using a low-cost microcontroller. Appl Acoust 148:194\u2013201. https:\/\/doi.org\/10.1016\/j.apacoust.2018.12.028","journal-title":"Appl Acoust"},{"key":"11106_CR31","doi-asserted-by":"crossref","unstructured":"Kumari S, Roy D, Cartwright M, Bello JP, Arora A (2019) EdgeL^3: compressing L^3-Net for mote scale urban noise monitoring. In: 2019 IEEE international parallel and distributed processing symposium workshops (IPDPSW). pp 877\u2013884. https:\/\/doi.org\/10.1109\/IPDPSW.2019.00145","DOI":"10.1109\/IPDPSW.2019.00145"},{"key":"11106_CR32","doi-asserted-by":"crossref","unstructured":"Laksono BSP, Prasetio BH (2023) Speaker recognition on low power device using fully convolutional QuartzNet. In: Proceedings of the 8th international conference on sustainable information engineering and technology. Association for Computing Machinery, New York, pp 619\u2013624. https:\/\/doi.org\/10.1145\/3626641.3626946","DOI":"10.1145\/3626641.3626946"},{"key":"11106_CR33","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Y LeCun","year":"2015","unstructured":"LeCun Y, Bengio Y, Hinton G (2015) Deep learning. Nature 521:436\u2013444. https:\/\/doi.org\/10.1038\/nature14539","journal-title":"Nature"},{"key":"11106_CR34","doi-asserted-by":"crossref","unstructured":"Li J, Dai W, Metze F, Qu S, Das S (2017) A comparison of deep learning methods for environmental sound detection. In: 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP). pp 126\u2013130. https:\/\/doi.org\/10.1109\/ICASSP.2017.7952131","DOI":"10.1109\/ICASSP.2017.7952131"},{"key":"11106_CR35","doi-asserted-by":"publisher","first-page":"60","DOI":"10.3390\/computers12030060","volume":"12","author":"Z Li","year":"2023","unstructured":"Li Z, Li H, Meng L (2023) Model compression for deep neural networks: a survey. Computers 12:60. https:\/\/doi.org\/10.3390\/computers12030060","journal-title":"Computers"},{"key":"11106_CR36","doi-asserted-by":"publisher","first-page":"5389","DOI":"10.3390\/s24165389","volume":"24","author":"U Libal","year":"2024","unstructured":"Libal U, Biernacki P (2024) Non-intrusive system for honeybee recognition based on audio signals and maximum likelihood classification by autoencoder. Sensors 24:5389. https:\/\/doi.org\/10.3390\/s24165389","journal-title":"Sensors"},{"key":"11106_CR37","doi-asserted-by":"publisher","first-page":"056006","DOI":"10.1088\/1361-6501\/ad2665","volume":"35","author":"T-H Lin","year":"2024","unstructured":"Lin T-H, Chang C-T, Zhuang T-H, Putranto A (2024) Real-time hollow defect detection in tiles using on-device tiny machine learning. Meas Sci Technol 35:056006. https:\/\/doi.org\/10.1088\/1361-6501\/ad2665","journal-title":"Meas Sci Technol"},{"key":"11106_CR38","doi-asserted-by":"crossref","unstructured":"Liu J, Liu J, Du W, Li D (2019) Performance analysis and characterization of training deep learning models on mobile device. In: 2019 IEEE 25th international conference on parallel and distributed systems (ICPADS). pp 506\u2013515. https:\/\/doi.org\/10.1109\/ICPADS47876.2019.00077","DOI":"10.1109\/ICPADS47876.2019.00077"},{"key":"11106_CR39","doi-asserted-by":"publisher","first-page":"8","DOI":"10.1007\/s44163-023-00051-x","volume":"3","author":"M Maayah","year":"2023","unstructured":"Maayah M, Abunada A, Al-Janahi K, Ahmed ME, Qadir J (2023) LimitAccess: on-device TinyML based robust speech recognition and age classification. Discov Artif Intell 3:8. https:\/\/doi.org\/10.1007\/s44163-023-00051-x","journal-title":"Discov Artif Intell"},{"key":"11106_CR40","doi-asserted-by":"publisher","unstructured":"Marciniak F, Marciniak W, Marciniak T (2023) Analysis of fast prototyping of microcontroller-based ML software for acoustic signal classification. In: 2023 Signal processing: algorithms, architectures, arrangements, and applications (SPA). pp 36\u201341. https:\/\/doi.org\/10.23919\/SPA59660.2023.10274443","DOI":"10.23919\/SPA59660.2023.10274443"},{"key":"11106_CR41","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3618104","volume":"56","author":"D Meedeniya","year":"2023","unstructured":"Meedeniya D, Ariyarathne I, Bandara M, Jayasundara R, Perera C (2023) A survey on deep learning based forest environment sound classification at the edge. ACM Comput Surv 56:1\u20136636. https:\/\/doi.org\/10.1145\/3618104","journal-title":"ACM Comput Surv"},{"key":"11106_CR42","doi-asserted-by":"publisher","first-page":"6696","DOI":"10.1109\/ACCESS.2022.3140807","volume":"10","author":"M Mohaimenuzzaman","year":"2022","unstructured":"Mohaimenuzzaman M, Bergmeir C, Meyer B (2022) Pruning vs XNOR-net: a comprehensive study of deep learning for audio classification on edge-devices. IEEE Access 10:6696\u20136707. https:\/\/doi.org\/10.1109\/ACCESS.2022.3140807","journal-title":"IEEE Access"},{"key":"11106_CR43","doi-asserted-by":"publisher","first-page":"109025","DOI":"10.1016\/j.patcog.2022.109025","volume":"133","author":"M Mohaimenuzzaman","year":"2023","unstructured":"Mohaimenuzzaman M, Bergmeir C, West IT, Meyer B (2023) Environmental sound classification on the edge: a pipeline for deep acoustic networks on extremely resource-constrained devices. Pattern Recognit 133:109025. https:\/\/doi.org\/10.1016\/j.patcog.2022.109025","journal-title":"Pattern Recognit"},{"key":"11106_CR44","doi-asserted-by":"publisher","first-page":"103","DOI":"10.32622\/ijrat.76201926","volume":"7","author":"A Mohammad","year":"2019","unstructured":"Mohammad A, Tripathi DMM (2019) Audio analysis and classification: a review. Int J Res Advent Technol 7:103\u2013109. https:\/\/doi.org\/10.32622\/ijrat.76201926","journal-title":"Int J Res Advent Technol"},{"key":"11106_CR45","doi-asserted-by":"publisher","unstructured":"Montino P, Pau D (2019) Environmental intelligence for embedded real-time traffic sound classification. In: 2019 IEEE 5th international forum on research and technology for society and industry (RTSI). pp 45\u201350. https:\/\/doi.org\/10.1109\/RTSI.2019.8895517","DOI":"10.1109\/RTSI.2019.8895517"},{"key":"11106_CR46","doi-asserted-by":"publisher","first-page":"21","DOI":"10.3390\/sci6020021","volume":"6","author":"A Mou","year":"2024","unstructured":"Mou A, Milanova M (2024) Performance analysis of deep learning model-compression techniques for audio classification on edge devices. Science 6:21. https:\/\/doi.org\/10.3390\/sci6020021","journal-title":"Science"},{"key":"11106_CR47","doi-asserted-by":"publisher","first-page":"61950","DOI":"10.1109\/ACCESS.2023.3287093","volume":"11","author":"A Mukhamediya","year":"2023","unstructured":"Mukhamediya A, Fazli S, Zollanvari A (2023) On the effect of log-mel spectrogram parameter tuning for deep learning-based speech emotion recognition. IEEE Access 11:61950\u201361957. https:\/\/doi.org\/10.1109\/ACCESS.2023.3287093","journal-title":"IEEE Access"},{"key":"11106_CR48","doi-asserted-by":"crossref","unstructured":"Munirathinam R, Vitek S (2024) Sound source localization and classification for emergency vehicle siren detection using resource constrained systems. In: 2024 34th International conference radioelektronika (RADIOELEKTRONIKA). pp 1\u20135. https:\/\/doi.org\/10.1109\/RADIOELEKTRONIKA61599.2024.10524053","DOI":"10.1109\/RADIOELEKTRONIKA61599.2024.10524053"},{"key":"11106_CR49","volume-title":"Detection and classification of acoustic scenes and events","author":"F Naccari","year":"2020","unstructured":"Naccari F, Guarneri I, Curti S, Savi AA (2020) Embedded acoustic scene classification for low power microcontroller devices. In: Detection and classification of acoustic scenes and events, DCASE2020. Tokyo, Japan.\u00a0"},{"key":"11106_CR50","doi-asserted-by":"publisher","first-page":"8608","DOI":"10.3390\/s22228608","volume":"22","author":"AFR Nogueira","year":"2022","unstructured":"Nogueira AFR, Oliveira HS, Machado JJM, Tavares JMRS (2022a) Sound classification and processing of urban environments: a systematic literature review. Sensors 22:8608. https:\/\/doi.org\/10.3390\/s22228608","journal-title":"Sensors"},{"key":"11106_CR51","doi-asserted-by":"publisher","first-page":"8874","DOI":"10.3390\/s22228874","volume":"22","author":"AFR Nogueira","year":"2022","unstructured":"Nogueira AFR, Oliveira HS, Machado JJM, Tavares JMRS (2022b) Transformers for urban sound classification\u2014a comprehensive performance evaluation. Sensors 22:8874. https:\/\/doi.org\/10.3390\/s22228874","journal-title":"Sensors"},{"key":"11106_CR52","doi-asserted-by":"publisher","first-page":"2984","DOI":"10.3390\/s21092984","volume":"21","author":"P-E Novac","year":"2021","unstructured":"Novac P-E, Boukli Hacene G, Pegatoquet A, Miramond B, Gripon V (2021) Quantization and deployment of deep neural networks on microcontrollers. Sensors 21:2984. https:\/\/doi.org\/10.3390\/s21092984","journal-title":"Sensors"},{"key":"11106_CR53","doi-asserted-by":"publisher","first-page":"6978","DOI":"10.3390\/app11156978","volume":"11","author":"A Polo-Rodriguez","year":"2021","unstructured":"Polo-Rodriguez A, Vilchez Chiachio JM, Paggetti C, Medina-Quero J (2021) Ambient sound Recognition of Daily Events by means of Convolutional Neural Networks and fuzzy temporal restrictions. Appl Sci 11:6978. https:\/\/doi.org\/10.3390\/app11156978","journal-title":"Appl Sci"},{"key":"11106_CR54","doi-asserted-by":"publisher","first-page":"8129","DOI":"10.1007\/s11042-023-15891-z","volume":"83","author":"A Prashanth","year":"2024","unstructured":"Prashanth A, Jayalakshmi SL, Vedhapriyavadhana R (2024) A review of deep learning techniques in audio event recognition (AER) applications. Multimed Tools Appl 83:8129\u20138143. https:\/\/doi.org\/10.1007\/s11042-023-15891-z","journal-title":"Multimed Tools Appl"},{"key":"11106_CR55","doi-asserted-by":"publisher","first-page":"2046","DOI":"10.3390\/s24072046","volume":"24","author":"D Priebe","year":"2024","unstructured":"Priebe D, Ghani B, Stowell D (2024) Efficient speech detection in environmental audio using acoustic recognition and knowledge distillation. Sensors 24:2046. https:\/\/doi.org\/10.3390\/s24072046","journal-title":"Sensors"},{"key":"11106_CR56","doi-asserted-by":"publisher","first-page":"206","DOI":"10.1109\/JSTSP.2019.2908700","volume":"13","author":"H Purwins","year":"2019","unstructured":"Purwins H, Li B, Virtanen T, Schl\u00fcter J, Chang S, Sainath T (2019) Deep learning for audio signal processing. IEEE J Sel Top Signal Process 13:206\u2013219. https:\/\/doi.org\/10.1109\/JSTSP.2019.2908700","journal-title":"IEEE J Sel Top Signal Process"},{"key":"11106_CR57","doi-asserted-by":"publisher","first-page":"3888","DOI":"10.3390\/s22103888","volume":"22","author":"A Qurthobi","year":"2022","unstructured":"Qurthobi A, Maskeli\u016bnas R, Dama\u0161evi\u010dius R (2022) Detection of mechanical failures in Industrial machines using overlapping acoustic anomalies: a systematic literature review. Sensors 22:3888. https:\/\/doi.org\/10.3390\/s22103888","journal-title":"Sensors"},{"key":"11106_CR58","doi-asserted-by":"publisher","first-page":"21362","DOI":"10.1109\/JSEN.2022.3210773","volume":"22","author":"SS Saha","year":"2022","unstructured":"Saha SS, Sandha SS, Srivastava M (2022) Machine learning for microcontroller-class hardware: a review. IEEE Sens J 22:21362\u201321390. https:\/\/doi.org\/10.1109\/JSEN.2022.3210773","journal-title":"IEEE Sens J"},{"key":"11106_CR59","doi-asserted-by":"crossref","unstructured":"Sammarco M, Stellantis TZ, Gantert L, Campista MEM (2024) Sound event detection via pervasive devices for mobility surveillance in smart cities. In: 2024 IEEE international conference on pervasive computing and communications workshops and other affiliated events (PerCom Workshops). pp 581\u2013586. https:\/\/doi.org\/10.1109\/PerComWorkshops59983.2024.10503381","DOI":"10.1109\/PerComWorkshops59983.2024.10503381"},{"key":"11106_CR60","doi-asserted-by":"publisher","first-page":"107020","DOI":"10.1016\/j.apacoust.2019.107020","volume":"158","author":"G Sharma","year":"2020","unstructured":"Sharma G, Umapathy K, Krishnan S (2020) Trends in audio signal feature extraction methods. Appl Acoust 158:107020. https:\/\/doi.org\/10.1016\/j.apacoust.2019.107020","journal-title":"Appl Acoust"},{"key":"11106_CR61","doi-asserted-by":"publisher","first-page":"108382","DOI":"10.1016\/j.engappai.2024.108382","volume":"133","author":"R Shi","year":"2024","unstructured":"Shi R, Zhang F, Li Y (2024) Lightweight network based features fusion for steel rolling ambient sound classification. Eng Appl Artif Intell 133:108382. https:\/\/doi.org\/10.1016\/j.engappai.2024.108382","journal-title":"Eng Appl Artif Intell"},{"key":"11106_CR62","doi-asserted-by":"publisher","first-page":"2313","DOI":"10.1140\/epjst\/e2019-900046-x","volume":"228","author":"K Smagulova","year":"2019","unstructured":"Smagulova K, James AP (2019) A survey on LSTM memristive neural network architectures and applications. Eur Phys J Spec Top 228:2313\u20132324. https:\/\/doi.org\/10.1140\/epjst\/e2019-900046-x","journal-title":"Eur Phys J Spec Top"},{"key":"11106_CR63","doi-asserted-by":"crossref","unstructured":"Somwong B, Kumphet K, Massagram W (2023) Acoustic monitoring system with ai threat detection system for forest protection. In: 2023 20th International joint conference on computer science and software engineering (JCSSE). pp 253\u2013257. https:\/\/doi.org\/10.1109\/JCSSE58229.2023.10202043","DOI":"10.1109\/JCSSE58229.2023.10202043"},{"key":"11106_CR64","doi-asserted-by":"publisher","first-page":"9658","DOI":"10.3390\/s22249658","volume":"22","author":"K Strantzalis","year":"2022","unstructured":"Strantzalis K, Gioulekas F, Katsaros P, Symeonidis A (2022) Operational state recognition of a DC motor using edge artificial intelligence. Sensors 22:9658. https:\/\/doi.org\/10.3390\/s22249658","journal-title":"Sensors"},{"key":"11106_CR65","doi-asserted-by":"crossref","unstructured":"S\u00fcer S, K\u00f6seo\u011flu \u0130, \u00d6ner R, \u00cfnce G (2023) Detection of clips failures in manufacturing using audio signals. In: 2023 5th International congress on human-computer interaction, optimization and robotic applications (HORA). pp 01\u201305. https:\/\/doi.org\/10.1109\/HORA58378.2023.10156765","DOI":"10.1109\/HORA58378.2023.10156765"},{"key":"11106_CR66","doi-asserted-by":"publisher","DOI":"10.1007\/s00107-024-02139-2","author":"S Svrzi\u0107","year":"2024","unstructured":"Svrzi\u0107 S, Djurkovi\u0107 M, Vuki\u0107evi\u0107 A, Nikoli\u0107 Z, Mihailovi\u0107 V, Dedi\u0107 A (2024) Sound classification and power consumption to sound intensity relation as a tool for wood machining monitoring. Eur J Wood Wood Prod. https:\/\/doi.org\/10.1007\/s00107-024-02139-2","journal-title":"Eur J Wood Wood Prod"},{"key":"11106_CR67","doi-asserted-by":"publisher","first-page":"113294","DOI":"10.1016\/j.measurement.2023.113294","volume":"220","author":"L Tang","year":"2023","unstructured":"Tang L, Tian H, Huang H, Shi S, Ji Q (2023) A survey of mechanical fault diagnosis based on audio signal analysis. Measurement 220:113294. https:\/\/doi.org\/10.1016\/j.measurement.2023.113294","journal-title":"Measurement"},{"key":"11106_CR68","unstructured":"Tolstikhin I, Houlsby N, Kolesnikov A, Beyer L, Zhai X, Unterthiner T, Yung J, Steiner A, Keysers D, Uszkoreit J, Lucic M, Dosovitskiy A (2021) MLP-mixer: an all-MLP architecture for vision. In: arXiv.org. https:\/\/arxiv.org\/abs\/2105.01601v4. Accessed 23 Jul 2024"},{"key":"11106_CR69","doi-asserted-by":"publisher","first-page":"1100","DOI":"10.1109\/TASLP.2023.3244507","volume":"31","author":"AM Tripathi","year":"2023","unstructured":"Tripathi AM, Pandey OJ (2023) Divide and distill: new outlooks on knowledge distillation for environmental sound classification. IEEEACM Trans Audio Speech Lang Process 31:1100\u20131113. https:\/\/doi.org\/10.1109\/TASLP.2023.3244507","journal-title":"IEEEACM Trans Audio Speech Lang Process"},{"key":"11106_CR70","doi-asserted-by":"publisher","first-page":"10233","DOI":"10.1109\/JIOT.2020.2997047","volume":"7","author":"L Turchet","year":"2020","unstructured":"Turchet L, Fazekas G, Lagrange M, Ghadikolaei HS, Fischione C (2020) The internet of audio things: state of the art, vision, and challenges. IEEE Internet Things J 7:10233\u201310249. https:\/\/doi.org\/10.1109\/JIOT.2020.2997047","journal-title":"IEEE Internet Things J"},{"key":"11106_CR71","doi-asserted-by":"publisher","first-page":"2622","DOI":"10.3390\/electronics10212622","volume":"10","author":"J Vandendriessche","year":"2021","unstructured":"Vandendriessche J, Wouters N, da Silva B, Lamrini M, Chkouri MY, Touhafi A (2021) Environmental sound recognition on embedded systems: from FPGAs to TPUs. Electronics 10:2622. https:\/\/doi.org\/10.3390\/electronics10212622","journal-title":"Electronics"},{"key":"11106_CR72","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I (2023) Attention is all you need. In: NIPS\u201917: Proceedings of the 31st International Conference on Neural Information Processing Systems. pp 6000-6010."},{"key":"11106_CR73","doi-asserted-by":"publisher","first-page":"1627","DOI":"10.1007\/s12274-014-0652-3","volume":"8","author":"Y Wang","year":"2015","unstructured":"Wang Y, Yang T, Lao J, Zhang R, Zhang Y, Zhu M, Li X, Zang X, Wang K, Yu W, Jin H, Wang L, Zhu H (2015) Ultra-sensitive graphene strain sensor for sound signal acquisition and recognition. Nano Res 8:1627\u20131636. https:\/\/doi.org\/10.1007\/s12274-014-0652-3","journal-title":"Nano Res"},{"key":"11106_CR74","doi-asserted-by":"publisher","first-page":"5922","DOI":"10.3390\/s24185922","volume":"24","author":"J-J Wang","year":"2024","unstructured":"Wang J-J, Sharma AK, Liu S-H, Zhang H, Chen W, Lee T-L (2024) Prediction of vascular access stenosis by lightweight convolutional neural network using blood flow sound signals. Sensors 24:5922. https:\/\/doi.org\/10.3390\/s24185922","journal-title":"Sensors"},{"key":"11106_CR75","doi-asserted-by":"publisher","first-page":"110178","DOI":"10.1016\/j.apacoust.2024.110178","volume":"225","author":"P Wi\u00dfbrock","year":"2024","unstructured":"Wi\u00dfbrock P, Ren Z, Pelkmann D (2024) More than spectrograms: deep representation learning for machinery fault detection. Appl Acoust 225:110178. https:\/\/doi.org\/10.1016\/j.apacoust.2024.110178","journal-title":"Appl Acoust"},{"key":"11106_CR76","doi-asserted-by":"crossref","unstructured":"Wyatt S, Elliott D, Aravamudan A, Otero CE, Otero LD, Anagnostopoulos GC, Smith AO, Peter AM, Jones W, Leung S, Lam E (2021) Environmental sound classification with tiny transformers in noisy edge environments. In: 2021 IEEE 7th world forum on internet of things (WF-IoT). pp 309\u2013314. https:\/\/doi.org\/10.1109\/WF-IoT51360.2021.9596007","DOI":"10.1109\/WF-IoT51360.2021.9596007"},{"key":"11106_CR77","doi-asserted-by":"publisher","DOI":"10.3390\/buildings12111947","author":"W Xiong","year":"2022","unstructured":"Xiong W, Xu X, Chen L, Yang J (2022) Sound-based construction activity monitoring with deep learning. Buildings. https:\/\/doi.org\/10.3390\/buildings12111947","journal-title":"Buildings"},{"key":"11106_CR78","doi-asserted-by":"publisher","first-page":"182","DOI":"10.1016\/j.ymssp.2018.07.039","volume":"119","author":"Y Yang","year":"2019","unstructured":"Yang Y, Peng Z, Zhang W, Meng G (2019) Parameterised time-frequency analysis methods and their engineering applications: a review of recent advances. Mech Syst Signal Process 119:182\u2013221. https:\/\/doi.org\/10.1016\/j.ymssp.2018.07.039","journal-title":"Mech Syst Signal Process"},{"key":"11106_CR79","doi-asserted-by":"publisher","first-page":"85189","DOI":"10.1109\/ACCESS.2022.3198104","volume":"10","author":"F Yang","year":"2022","unstructured":"Yang F, Jiang Y, Xu Y (2022a) Design of bird sound recognition model based on lightweight. IEEE Access 10:85189\u201385198. https:\/\/doi.org\/10.1109\/ACCESS.2022.3198104","journal-title":"IEEE Access"},{"key":"11106_CR80","doi-asserted-by":"publisher","first-page":"23267","DOI":"10.1109\/JSEN.2022.3211098","volume":"22","author":"Z Yang","year":"2022","unstructured":"Yang Z, Wang Y, Pan Y, Huan R, Liang R (2022b) Inaudible sounds from appliances as anchors: a new signal of opportunity for indoor localization. IEEE Sens J 22:23267\u201323276. https:\/\/doi.org\/10.1109\/JSEN.2022.3211098","journal-title":"IEEE Sens J"},{"key":"11106_CR81","doi-asserted-by":"publisher","first-page":"412","DOI":"10.3390\/e25030412","volume":"25","author":"X Yu","year":"2023","unstructured":"Yu X, Li X (2023) Sound recognition method of coal mine gas and coal dust explosion based on GoogLeNet. Entropy 25:412. https:\/\/doi.org\/10.3390\/e25030412","journal-title":"Entropy"},{"key":"11106_CR82","doi-asserted-by":"publisher","first-page":"106620","DOI":"10.1109\/ACCESS.2023.3318015","volume":"11","author":"K Zaman","year":"2023","unstructured":"Zaman K, Sah M, Direkoglu C, Unoki M (2023) A survey of audio classification using deep learning. IEEE Access 11:106620\u2013106649. https:\/\/doi.org\/10.1109\/ACCESS.2023.3318015","journal-title":"IEEE Access"},{"key":"11106_CR83","doi-asserted-by":"publisher","first-page":"484","DOI":"10.3390\/mi11050484","volume":"11","author":"SA Zawawi","year":"2020","unstructured":"Zawawi SA, Hamzah AA, Majlis BY, Mohd-Yasin F (2020) A review of MEMS capacitive microphones. Micromachines 11:484. https:\/\/doi.org\/10.3390\/mi11050484","journal-title":"Micromachines"},{"key":"11106_CR84","doi-asserted-by":"publisher","first-page":"03007","DOI":"10.1051\/shsconf\/202213903007","volume":"139","author":"A Zelios","year":"2022","unstructured":"Zelios A, Grammenos A, Papatsimouli M, Asimopoulos N, Fragulis G (2022) Recursive neural networks: recent results and applications. SHS Web Conf 139:03007. https:\/\/doi.org\/10.1051\/shsconf\/202213903007","journal-title":"SHS Web Conf"},{"key":"11106_CR85","doi-asserted-by":"crossref","unstructured":"Zhang Z, Zhang R, Li Z, Bengio Y, Paull L (2020) Perceptual generative autoencoders. In: Proceedings of the 37th international conference on machine learning. PMLR, pp 11298\u201311306","DOI":"10.1609\/aaai.v37i9.26337"},{"key":"11106_CR86","doi-asserted-by":"crossref","unstructured":"Zhang Z, Shen Y, Valdes JJ, Huq S, Wallace B, Green J, Xi P, Goubran R (2023) Domestic sound classification with deep learning. In: 2023 IEEE sensors applications symposium (SAS). pp 01\u201306. https:\/\/doi.org\/10.1109\/SAS58821.2023.10254050","DOI":"10.1109\/SAS58821.2023.10254050"},{"key":"11106_CR87","doi-asserted-by":"publisher","first-page":"108835","DOI":"10.1016\/j.engappai.2024.108835","volume":"135","author":"Z Zhang","year":"2024","unstructured":"Zhang Z, Liu H, Shao Y, Yang J, Liu S, Yuan G (2024) CFENet: a contrastive frequency-sensitive learning method for gas-insulated switch-gear fault detection under varying operating conditions using acoustic signals. Eng Appl Artif Intell 135:108835. https:\/\/doi.org\/10.1016\/j.engappai.2024.108835","journal-title":"Eng Appl Artif Intell"}],"container-title":["Artificial Intelligence Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11106-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10462-025-11106-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11106-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,4,17]],"date-time":"2025-04-17T19:32:07Z","timestamp":1744918327000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10462-025-11106-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,15]]},"references-count":87,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2025,6]]}},"alternative-id":["11106"],"URL":"https:\/\/doi.org\/10.1007\/s10462-025-11106-z","relation":{},"ISSN":["1573-7462"],"issn-type":[{"value":"1573-7462","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,15]]},"assertion":[{"value":"7 January 2025","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 March 2025","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"163"}}