{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T18:49:16Z","timestamp":1777142956394,"version":"3.51.4"},"reference-count":42,"publisher":"MDPI AG","issue":"24","license":[{"start":{"date-parts":[[2021,12,15]],"date-time":"2021-12-15T00:00:00Z","timestamp":1639526400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100003654","name":"Korea Environmental Industry and Technology Institute","doi-asserted-by":"publisher","award":["2021002280004"],"award-info":[{"award-number":["2021002280004"]}],"id":[{"id":"10.13039\/501100003654","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Weakly labeled sound event detection (WSED) is an important task as it can facilitate the data collection efforts before constructing a strongly labeled sound event dataset. Recent high performance in deep learning-based WSED\u2019s exploited using a segmentation mask for detecting the target feature map. However, achieving accurate detection performance was limited in real streaming audio due to the following reasons. First, the convolutional neural networks (CNN) employed in the segmentation mask extraction process do not appropriately highlight the importance of feature as the feature is extracted without pooling operations, and, concurrently, a small size kernel forces the receptive field small, making it difficult to learn various patterns. Second, as feature maps are obtained in an end-to-end fashion, the WSED model would be weak to unknown contents in the wild. These limitations would lead to generating undesired feature maps, such as noise in the unseen environment. This paper addresses these issues by constructing a more efficient model by employing a gated linear unit (GLU) and dilated convolution to improve the problems of de-emphasizing importance and lack of receptive field. In addition, this paper proposes pseudo-label-based learning for classifying target contents and unknown contents by adding \u2019noise label\u2019 and \u2019noise loss\u2019 so that unknown contents can be separated as much as possible through the noise label. The experiment is performed by mixing DCASE 2018 task1 acoustic scene data and task2 sound event data. The experimental results show that the proposed SED model achieves the best F1 performance with 59.7% at 0 SNR, 64.5% at 10 SNR, and 65.9% at 20 SNR. These results represent an improvement of 17.7%, 16.9%, and 16.5%, respectively, over the baseline.<\/jats:p>","DOI":"10.3390\/s21248375","type":"journal-article","created":{"date-parts":[[2021,12,15]],"date-time":"2021-12-15T21:47:36Z","timestamp":1639604856000},"page":"8375","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Sound Event Detection by Pseudo-Labeling in Weakly Labeled Dataset"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6712-6594","authenticated-orcid":false,"given":"Chungho","family":"Park","sequence":"first","affiliation":[{"name":"Department of Electronics and Electrical Engineering, Korea University Seoul, Seoul 136-713, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Donghyeon","family":"Kim","sequence":"additional","affiliation":[{"name":"Department of Electronics and Electrical Engineering, Korea University Seoul, Seoul 136-713, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8744-4514","authenticated-orcid":false,"given":"Hanseok","family":"Ko","sequence":"additional","affiliation":[{"name":"Department of Electronics and Electrical Engineering, Korea University Seoul, Seoul 136-713, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,12,15]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"2078","DOI":"10.1007\/s11771-020-4530-8","article-title":"Discrimination of mining microseismic events and blasts using convolutional neural networks and original waveform","volume":"27","author":"Dong","year":"2020","journal-title":"J. Cent. South Univ."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1109\/MSP.2021.3090678","article-title":"Sound event detection: A tutorial","volume":"38","author":"Mesaros","year":"2021","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"3433","DOI":"10.1007\/s00034-019-01094-1","article-title":"A survey: Neural network-based deep learning for acoustic event detection","volume":"38","author":"Xia","year":"2019","journal-title":"Circuits Syst. Signal Process."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2871183","article-title":"Audio surveillance: A systematic review","volume":"48","author":"Crocco","year":"2016","journal-title":"ACM Comput. Surv."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"640","DOI":"10.1109\/LSP.2020.2988199","article-title":"Amphibian Sounds Generating Network Based on Adversarial Learning","volume":"27","author":"Park","year":"2020","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"2615","DOI":"10.1587\/transinf.2019EDL8128","article-title":"Channel and Frequency Attention Module for Diverse Animal Sound Classification","volume":"102","author":"Ko","year":"2019","journal-title":"IEICE Trans. Inf. Syst."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Schr\u00f6der, J., Goetze, S., Gr\u00fctzmacher, V., and Anem\u00fcller, J. (2013, January 26\u201331). Automatic acoustic siren detection in traffic noise by part-based models. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6637696"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1291","DOI":"10.1109\/TASLP.2017.2690575","article-title":"Convolutional recurrent neural networks for polyphonic sound event detection","volume":"25","author":"Parascandolo","year":"2017","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Parascandolo, G., Huttunen, H., and Virtanen, T. (2016, January 20\u201325). Recurrent neural networks for polyphonic sound event detection in real life recordings. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472917"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1016\/S0004-3702(96)00034-3","article-title":"Solving the multiple instance problem with axis-parallel rectangles","volume":"89","author":"Dietterich","year":"1997","journal-title":"Artif. Intell."},{"key":"ref_11","first-page":"55","article-title":"FrameCNN: A weakly-supervised learning framework for frame-wise acoustic event detection and classification","volume":"14","author":"Chou","year":"2017","journal-title":"Recall"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Kumar, A., and Raj, B. (2017). Deep cnn framework for audio event recognition using weakly labeled web data. arXiv.","DOI":"10.1145\/2964284.2964310"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Xu, Y., Kong, Q., Wang, W., and Plumbley, M.D. (2018, January 15\u201320). Large-scale weakly supervised audio classification using gated convolutional neural network. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461975"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"777","DOI":"10.1109\/TASLP.2019.2895254","article-title":"Sound event detection and time-frequency segmentation from weakly labelled data","volume":"27","author":"Kong","year":"2019","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_15","unstructured":"Dauphin, Y.N., Fan, A., Auli, M., and Grangier, D. (2017, January 6\u201311). Language modeling with gated convolutional networks. Proceedings of the International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_16","unstructured":"Yu, F., and Koltun, V. (2015). Multi-scale context aggregation by dilated convolutions. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Graves, A., Fern\u00e1ndez, S., Gomez, F., and Schmidhuber, J. (2006, January 25\u201329). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. Proceedings of the 23rd International Conference on Machine Learning, Pittsburgh, PA, USA.","DOI":"10.1145\/1143844.1143891"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Wang, Y., and Metze, F. (2017, January 5\u20139). A first attempt at polyphonic sound event detection using connectionist temporal classification. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7952704"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Hou, Y., Kong, Q., Li, S., and Plumbley, M.D. (2019, January 12\u201317). Sound event detection with sequentially labelled data based on connectionist temporal classification and unsupervised clustering. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683627"},{"key":"ref_20","unstructured":"Lee, D.H. (2013, January 16\u201321). Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. Proceedings of the ICML 2013 Workshop: Challenges in Representation Learning (WREPL), Atlanta, GA, USA."},{"key":"ref_21","unstructured":"Mesaros, A., Heittola, T., and Virtanen, T. (2018). A multi-device dataset for urban acoustic scene classification. arXiv."},{"key":"ref_22","unstructured":"Fonseca, E., Plakal, M., Font, F., Ellis, D.P., Favory, X., Pons, J., and Serra, X. (2018). General-purpose tagging of freesound audio with audioset labels: Task description, dataset, and baseline. arXiv."},{"key":"ref_23","unstructured":"Maron, O., and Lozano-P\u00e9rez, T. (1998). A framework for multiple-instance learning. Adv. Neural Inf. Process. Syst., 570\u2013576. Available online: https:\/\/citeseerx.ist.psu.edu\/viewdoc\/download?doi=10.1.1.51.7638&rep=rep1&type=pdf."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"213","DOI":"10.1109\/RBME.2017.2651164","article-title":"Multiple-instance learning for medical image and video analysis","volume":"10","author":"Quellec","year":"2017","journal-title":"IEEE Rev. Biomed. Eng."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Xu, Y., Mo, T., Feng, Q., Zhong, P., Lai, M., Eric, I., and Chang, C. (2014, January 4\u20139). Deep learning of feature representation with multiple instance learning for medical image analysis. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6853873"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Papandreou, G., Chen, L.C., Murphy, K.P., and Yuille, A.L. (2015, January 7\u201313). Weakly-and semi-supervised learning of a deep convolutional network for semantic image segmentation. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.203"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Wu, J., Zhao, Y., Zhu, J.Y., Luo, S., and Tu, Z. (2014, January 23\u201328). Milcut: A sweeping line multiple instance learning paradigm for interactive image segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.40"},{"key":"ref_28","first-page":"1","article-title":"Acoustic scene classification with squeeze-excitation residual networks","volume":"48","author":"Zuccarello","year":"2016","journal-title":"ACM Comput. Surv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"McDonnell, M.D., and Gao, W. (2020, January 4\u20138). Acoustic scene classification using deep residual networks with late fusion of separated high and low frequency paths. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9053274"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"123","DOI":"10.1016\/j.apacoust.2018.12.019","article-title":"Environmental sound classification with dilated convolutions","volume":"148","author":"Chen","year":"2019","journal-title":"Appl. Acoust."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Wei, Y., Xiao, H., Shi, H., Jie, Z., Feng, J., and Huang, T.S. (2018, January 18\u201323). Revisiting dilated convolution: A simple approach for weakly-and semi-supervised semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00759"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"1057","DOI":"10.1007\/s11042-019-08208-6","article-title":"Multi-scale dilated convolution of convolutional neural network for crowd counting","volume":"79","author":"Wang","year":"2020","journal-title":"Multimed. Tools Appl."},{"key":"ref_33","unstructured":"Agarap, A.F. (2018). Deep learning using rectified linear units (relu). arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Li, Y., Liu, M., Drossos, K., and Virtanen, T. (2020, January 4\u20138). Sound event detection via dilated convolutional recurrent neural networks. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9054433"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Kolesnikov, A., and Lampert, C.H. (2016, January 11\u201314). Seed, expand and constrain: Three principles for weakly-supervised image segmentation. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46493-0_42"},{"key":"ref_36","unstructured":"Kong, Q., Iqbal, T., Xu, Y., Wang, W., and Plumbley, M.D. (2018). DCASE 2018 challenge surrey cross-task convolutional neural network baseline. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Miyazaki, K., Komatsu, T., Hayashi, T., Watanabe, S., Toda, T., and Takeda, K. (2020, January 4\u20138). Weakly-Supervised Sound Event Detection with Self-Attention. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9053609"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Mesaros, A., Heittola, T., and Virtanen, T. (2016). Metrics for polyphonic sound event detection. Appl. Sci., 6.","DOI":"10.3390\/app6060162"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1148\/radiology.143.1.7063747","article-title":"The meaning and use of the area under a receiver operating characteristic (ROC) curve","volume":"142","author":"Hanley","year":"1982","journal-title":"Radiology"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_41","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_42","unstructured":"Ioffe, S., and Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/24\/8375\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:48:56Z","timestamp":1760168936000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/24\/8375"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,15]]},"references-count":42,"journal-issue":{"issue":"24","published-online":{"date-parts":[[2021,12]]}},"alternative-id":["s21248375"],"URL":"https:\/\/doi.org\/10.3390\/s21248375","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12,15]]}}}