{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T05:07:17Z","timestamp":1780722437500,"version":"3.54.1"},"reference-count":35,"publisher":"MDPI AG","issue":"15","license":[{"start":{"date-parts":[[2023,7,31]],"date-time":"2023-07-31T00:00:00Z","timestamp":1690761600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Key Research and Development Projects in Zhejiang Province","award":["2022C01005"],"award-info":[{"award-number":["2022C01005"]}]},{"name":"Key Research and Development Projects in Zhejiang Province","award":["2023C01034"],"award-info":[{"award-number":["2023C01034"]}]},{"name":"Key Research and Development Projects in Zhejiang Province","award":["2023C1032"],"award-info":[{"award-number":["2023C1032"]}]},{"name":"Key Research and Development Projects in Zhejiang Province","award":["2021C03192"],"award-info":[{"award-number":["2021C03192"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>With the continuous promotion of \u201csmart cities\u201d worldwide, the approach to be used in combining smart cities with modern advanced technologies (Internet of Things, cloud computing, artificial intelligence) has become a hot topic. However, due to the non-stationary nature of environmental sound and the interference of urban noise, it is challenging to fully extract features from the model with a single input and achieve ideal classification results, even with deep learning methods. To improve the recognition accuracy of ESC (environmental sound classification), we propose a dual-branch residual network (dual-resnet) based on feature fusion. Furthermore, in terms of data pre-processing, a loop-padding method is proposed to patch shorter data, enabling it to obtain more useful information. At the same time, in order to prevent the occurrence of overfitting, we use the time-frequency data enhancement method to expand the dataset. After uniform pre-processing of all the original audio, the dual-branch residual network automatically extracts the frequency domain features of the log-Mel spectrogram and log-spectrogram. Then, the two different audio features are fused to make the representation of the audio features more comprehensive. The experimental results show that compared with other models, the classification accuracy of the UrbanSound8k dataset has been improved to different degrees.<\/jats:p>","DOI":"10.3390\/s23156823","type":"journal-article","created":{"date-parts":[[2023,7,31]],"date-time":"2023-07-31T10:08:14Z","timestamp":1690798094000},"page":"6823","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["An Automatic Classification System for Environmental Sound in Smart Cities"],"prefix":"10.3390","volume":"23","author":[{"given":"Dongping","family":"Zhang","sequence":"first","affiliation":[{"name":"Key Laboratory of Electromagnetic Wave Information Technology and Metrology of Zhejiang Province, China Jiliang University, Hangzhou 310018, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ziyin","family":"Zhong","sequence":"additional","affiliation":[{"name":"Key Laboratory of Electromagnetic Wave Information Technology and Metrology of Zhejiang Province, China Jiliang University, Hangzhou 310018, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuejian","family":"Xia","sequence":"additional","affiliation":[{"name":"Key Laboratory of Electromagnetic Wave Information Technology and Metrology of Zhejiang Province, China Jiliang University, Hangzhou 310018, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhutao","family":"Wang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Electromagnetic Wave Information Technology and Metrology of Zhejiang Province, China Jiliang University, Hangzhou 310018, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenbo","family":"Xiong","sequence":"additional","affiliation":[{"name":"Hangzhou Aihua Intelligent Technology Co., Ltd., 359 Shuxin Road, Hangzhou 311100, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,7,31]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Pan, X., Ge, C., Lu, R., Song, S., Chen, G., Huang, Z., and Huang, G. (2022, January 18\u201324). On the Integration of Self-Attention and Convolution. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00089"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Yu, R., Du, D., LaLonde, R., Davila, D., Funk, C., Hoogs, A., and Clipp, B. (2022, January 18\u201324). Cascade Transformers for End-to-End Person Search. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00712"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"531","DOI":"10.1016\/j.dcan.2022.03.023","article-title":"Chiller faults detection and diagnosis with sensor network and adaptive 1D CNN","volume":"8","author":"Yan","year":"2022","journal-title":"Digit. Commun. Netw."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Nagrani, A., Albanie, S., and Zisserman, A. (2018, January 18\u201323). Seeing voices and hearing faces: Cross-modal biometric matching. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00879"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"75702","DOI":"10.1109\/ACCESS.2020.2988986","article-title":"Acoustic-Based Emergency Vehicle Detection Using Convolutional Neural Networks","volume":"8","author":"Tran","year":"2020","journal-title":"IEEE Access"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1875","DOI":"10.1109\/TASLP.2020.2964959","article-title":"Sound Events Recognition and Retrieval Using Multi-Convolutional-Channel Sparse Coding Convolutional Neural Networks","volume":"28","author":"Wang","year":"2020","journal-title":"IEEE ACM Trans. Audio, Speech, Lang. Process."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Avramidis, K., Kratimenos, A., Garoufis, C., Zlatintsi, A., and Maragos, P. (2021, January 6\u201311). Deep Convolutional and Recurrent Networks for Polyphonic Instrument Classification from Monophonic Raw Audio Waveforms. Proceedings of the ICASSP 2021\u20142021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9413479"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Piczak, K.J. (2015, January 17\u201320). Environmental sound classification with convolutional neural networks. Proceedings of the Machine Learning for Signal Processing (MLSP), Boston, MA, USA.","DOI":"10.1109\/MLSP.2015.7324337"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhang, J., Liu, W., Lan, J., Hu, Y., and Zhang, F. (2021, January 4\u20136). Audio Fault Analysis for Industrial Equipment Based on Feature Metric Engineering with CNNs. Proceedings of the 2021 4th International Conference on Robotics, Control and Automation Engineering (RCAE), Wuhan, China.","DOI":"10.1109\/RCAE53607.2021.9638896"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Abdoli, S., Cardinal, P., and Koerich, A.L. (2019). End-to-End Environmental Sound Classification using a 1D Convolutional Neural Network. arXiv.","DOI":"10.1016\/j.eswa.2019.06.040"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"21552","DOI":"10.1038\/s41598-021-01045-4","article-title":"Environmental sound classification using temporal-frequency attention based convolutional neural network","volume":"11","author":"Mu","year":"2021","journal-title":"Sci. Rep."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, Y., Feng, C., and Anderson, D.V. (2021, January 6\u201311). A Multi-Channel Temporal Attention Convolutional Neural Network Model for Environmental Sound Classification. Proceedings of the ICASSP 2021\u20142021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9413498"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"16","DOI":"10.1109\/MSP.2014.2326181","article-title":"Acoustic scene classification: Classifying environments from the sounds they produce","volume":"32","author":"Barchiesi","year":"2015","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1109\/TASLP.2014.2367814","article-title":"Random regression forests for acoustic event detection and classification","volume":"23","author":"Phan","year":"2015","journal-title":"IEEE ACM Trans. Audio Speech Lang. Process."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1016\/j.ins.2013.04.014","article-title":"Very short time environmental sound classification based on spectrogram pattern matching","volume":"243","author":"Khunarsal","year":"2013","journal-title":"Inf. Sci."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1494","DOI":"10.1109\/TNSRE.2022.3178476","article-title":"AI Empowered Virtual Reality Integrated Systems for Sleep Stage Classification and Quality Enhancement","volume":"30","author":"Huang","year":"2022","journal-title":"IEEE Trans. Neural Syst. Rehabil. Eng."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Yan, K., Zhou, X., and Yang, B. (2022). AI and IoT Applications of Smart Buildings and Smart Environment Design, Construction and Maintenance. Build. Environ., 109968.","DOI":"10.1016\/j.buildenv.2022.109968"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zaw, T.H., and War, N. (2017, January 22\u201324). The combination of spectral entropy, zero crossing rate, short time energy and linear prediction error for voice activity detection. Proceedings of the 2017 20th International Conference of Computer and Information Technology (ICCIT), Dhaka, Bangladesh.","DOI":"10.1109\/ICCITECHN.2017.8281794"},{"key":"ref_20","unstructured":"Lartillot, O., and Toiviainen, P. (2007, January 10\u201315). A Matlab toolbox for musical feature extraction from audio. Proceedings of the 10th International Conference on Digital Audio Effects (DAFx-07), Bordeaux, France."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Cotton, C.V., and Ellis, D.P. (2011, January 16\u201319). Spectral vs. spectro-temporal features for acoustic event detection. Proceedings of the 2011 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), New Paltz, NY, USA.","DOI":"10.1109\/ASPAA.2011.6082331"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Giannoulis, D., Benetos, E., Stowell, D., Rossignol, M., Lagrange, M., and Plumbley, M.D. (2013, January 20\u201323). Detection and classification of acoustic scenes and events: An IEEE AASP challenge. Proceedings of the Applications of Signal Processing to Audio and Acoustics (WASPAA), New Paltz, NY, USA.","DOI":"10.1109\/WASPAA.2013.6701819"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1993","DOI":"10.1109\/TASLP.2014.2359159","article-title":"A feature study for classification-based speech separation at low signal-to-noise ratios","volume":"22","author":"Chen","year":"2014","journal-title":"IEEE ACM Trans. Audio Speech Lang. Process."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Li, R., Yin, B., Cui, Y., Du, Z., and Li, K. (2020, January 11\u201313). Research on Environmental Sound Classification Algorithm Based on Multi-feature Fusion. Proceedings of the 2020 IEEE 9th Joint International Information Technology and Artificial Intelligence Conference (ITAIC), Chongqing, China.","DOI":"10.1109\/ITAIC49862.2020.9338926"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Salamon, J., Jacoby, C., and Bello, J.P. (2014, January 3\u20137). A dataset and taxonomy for urban sound research. Proceedings of the 22nd ACM International Conference on Multimedia, Orlando, FL, USA.","DOI":"10.1145\/2647868.2655045"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"2823","DOI":"10.1109\/TASLP.2020.3030489","article-title":"Interpretable representation learning for speech and audio signals based on relevance weighting","volume":"28","author":"Agrawal","year":"2020","journal-title":"IEEE ACM Trans. Audio Speech Lang. Process."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201322). Squeeze-and-Excitation Networks. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Piczak, K.J. (2015, January 26\u201330). ESC: Dataset for environmental sound classification. Proceedings of the 23rd ACM Multimedia Conference, Brisbane, Australia.","DOI":"10.1145\/2733373.2806390"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Park, D.S., Chan, W., Zhang, Y., Chiu, C.C., Zoph, B., Cubuk, E.D., and Le, Q.V. (2019). Specaugment: A simple data augmentation method for automatic speech recognition. Proc. Interspeech, 2613\u20132617.","DOI":"10.21437\/Interspeech.2019-2680"},{"key":"ref_30","first-page":"1929","article-title":"Dropout: A Simple Way to Prevent Neural Networks from Overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"123","DOI":"10.1016\/j.apacoust.2018.12.019","article-title":"Environmental sound classification with dilated convolutions","volume":"148","author":"Chen","year":"2018","journal-title":"Appl. Acoust."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Tokozume, Y., and Harada, T. (2017, January 5\u20139). Learning environmental sounds with end-to-end convolutional neural network. Proceedings of the ICASSP 2017\u20142017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7952651"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Sang, J., Park, S., and Lee, J. (2018, January 3\u20137). Convolutional Recurrent Neural Networks for Urban Sound Classification Using Raw Waveforms. Proceedings of the 2018 26th European Signal Processing Conference (EUSIPCO), Rome, Italy.","DOI":"10.23919\/EUSIPCO.2018.8553247"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Hojjati, H., and Armanfard, N. (2022, January 7\u201313). Self-Supervised Acoustic Anomaly Detection Via Contrastive Learning. Proceedings of the ICASSP 2022\u20142022 IEEE International Conference on Acoustics, Speech and Signal Processing), Singapore.","DOI":"10.1109\/ICASSP43922.2022.9746207"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Chen, H., Song, Y., Dai, L.-R., McLoughlin, I., and Liu, L. (2022, January 7\u201313). Self-Supervised Representation Learning for Unsupervised Anomalous Sound Detection Under Domain Shift. Proceedings of the ICASSP 2022\u20142022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore.","DOI":"10.1109\/ICASSP43922.2022.9747863"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/15\/6823\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:23:13Z","timestamp":1760127793000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/15\/6823"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,31]]},"references-count":35,"journal-issue":{"issue":"15","published-online":{"date-parts":[[2023,8]]}},"alternative-id":["s23156823"],"URL":"https:\/\/doi.org\/10.3390\/s23156823","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,31]]}}}