{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T16:36:29Z","timestamp":1781714189552,"version":"3.54.5"},"reference-count":41,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,6,16]],"date-time":"2025-06-16T00:00:00Z","timestamp":1750032000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,6,16]],"date-time":"2025-06-16T00:00:00Z","timestamp":1750032000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100012239","name":"Hubei Province Technological Innovation Major Project","doi-asserted-by":"crossref","award":["NO. 2021BAA034"],"award-info":[{"award-number":["NO. 2021BAA034"]}],"id":[{"id":"10.13039\/501100012239","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"crossref","award":["NO.62172306 and NO. 61872275"],"award-info":[{"award-number":["NO.62172306 and NO. 61872275"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Sound Event Detection (SED) is a pivotal task in audio signal processing with widespread applications, requiring the classification and temporal localization of sound events. However, there proves to be a challenge in balancing global features for event classification with local features for temporal localization. This paper introduces Group Feature Calibration (GFC), a plug-and-play solution designed to effectively representing diverse time-frequency characteristics of sound events. GFC consists of two sub-modules: Group Feature Learning (GFL) and Task-aware Activation (TA). GFL enhances the network\u2019s ability to capture and integrate both global and local sound event features, while TA adaptively refines feature activation, facilitating more precise event classification and time localization. This approach effectively recalibrates the features, significantly improving the performance of various SED models on both two sub-tasks. Experimental results demonstrate that implementing GFC module in three mainstream SED networks improves the Polyphonic Sound Event Detection Score (PSDS) on the DCASE 2022 Task 4 dataset by up to 9.94%. Through class-wise analysis of the results, the consistent improvement of the GFC module in detecting sound events with diverse time-frequency characteristics is proven. Further ablation study also validates the effectiveness of the two sub-modules within the GFC module.<\/jats:p>","DOI":"10.1186\/s13636-025-00405-6","type":"journal-article","created":{"date-parts":[[2025,6,16]],"date-time":"2025-06-16T17:58:29Z","timestamp":1750096709000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Group feature calibration for sound event detection"],"prefix":"10.1186","volume":"2025","author":[{"given":"Yanzhen","family":"Ren","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2790-3069","authenticated-orcid":false,"given":"Wuyang","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chenyu","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tingting","family":"Zhu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,6,16]]},"reference":[{"issue":"5","key":"405_CR1","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1109\/MSP.2021.3090678","volume":"38","author":"A Mesaros","year":"2021","unstructured":"A. Mesaros, T. Heittola, T. Virtanen, M.D. Plumbley, Sound event detection: A tutorial. IEEE Signal Process. Mag. 38(5), 67\u201383 (2021)","journal-title":"IEEE Signal Process. Mag."},{"issue":"4","key":"405_CR2","doi-asserted-by":"publisher","first-page":"854","DOI":"10.3390\/s17040854","volume":"17","author":"RM Alsina-Pag\u00e8s","year":"2017","unstructured":"R.M. Alsina-Pag\u00e8s, J. Navarro, F. Al\u00edas, M. Herv\u00e1s, homesound: Real-time audio event detection based on high performance computing for behaviour and surveillance remote monitoring. Sensors 17(4), 854 (2017)","journal-title":"Sensors"},{"key":"405_CR3","doi-asserted-by":"crossref","unstructured":"F. Angulo, S. Essid, G. Peeters, C. Mietlicki, in 2023 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2023, Rhodes Island, Greece, May 4-10, 2023. Cosmopolite sound monitoring (cosmo): A study of urban sound event detection systems generalizing to multiple cities (IEEE, Piscataway, NJ, USA, 2023), pp. 1\u20135.","DOI":"10.1109\/ICASSP49357.2023.10095833"},{"issue":"6","key":"405_CR4","doi-asserted-by":"publisher","first-page":"1291","DOI":"10.1109\/TASLP.2017.2690575","volume":"25","author":"E \u00c7ak\u0131r","year":"2017","unstructured":"E. \u00c7ak\u0131r, G. Parascandolo, T. Heittola, H. Huttunen, T. Virtanen, Convolutional recurrent neural networks for polyphonic sound event detection. IEEE\/ACM Trans. Audio Speech Lang. Process. 25(6), 1291\u20131303 (2017)","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"405_CR5","doi-asserted-by":"crossref","unstructured":"Y. Li, M. Liu, K. Drossos, T. Virtanen, in 2020 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2020, Barcelona, Spain, May 4-8, 2020. Sound event detection via dilated convolutional recurrent neural networks (IEEE, Piscataway, NJ, USA, 2020), pp. 286\u2013290","DOI":"10.1109\/ICASSP40776.2020.9054433"},{"key":"405_CR6","doi-asserted-by":"crossref","unstructured":"X. Zheng, Y. Song, I. McLoughlin, L. Liu, L. Dai, in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2021, Toronto, ON, Canada, June 6-11, 2021. An improved mean teacher based method for large scale weakly labeled semi-supervised sound event detection (IEEE, Piscataway, NJ, USA, 2021), pp. 356\u2013360","DOI":"10.1109\/ICASSP39728.2021.9414931"},{"key":"405_CR7","doi-asserted-by":"crossref","unstructured":"H. Nam, S. Kim, B. Ko, Y. Park, in Interspeech 2022, 23rd Annual Conference of the International Speech Communication Association, Incheon, Korea, 18-22 September 2022, eds. by H. Ko, J.H.L. Hansen. Frequency dynamic convolution: Frequency-adaptive pattern recognition for sound event detection (ISCA, Baixas, France, 2022), pp. 2763\u20132767","DOI":"10.21437\/Interspeech.2022-10127"},{"key":"405_CR8","doi-asserted-by":"crossref","unstructured":"W. Xia, K. Koishida, in Interspeech 2019, 20th Annual Conference of the International Speech Communication Association, Graz, Austria, 15-19 September 2019, eds. by G. Kubin, Z. Kacic. Sound event detection in multichannel audio using convolutional time-frequency-channel squeeze and excitation (ISCA, Baixas, France, 2019), pp. 3629\u20133633","DOI":"10.21437\/Interspeech.2019-1860"},{"key":"405_CR9","doi-asserted-by":"crossref","unstructured":"X. Zheng, Y. Song, L. Dai, I. McLoughlin, L. Liu, in Interspeech 2021, 22nd Annual Conference of the International Speech Communication Association, Brno, Czechia, 30 August - 3 September 2021, eds. by H. Hermansky, H. Cernock\u00fd, L. Burget, L. Lamel, O. Scharenborg, P. Motl\u00edcek. An effective mutual mean teaching based domain adaptation method for sound event detection (ISCA, Baixas, France, 2021), pp. 556\u2013560","DOI":"10.21437\/Interspeech.2021-281"},{"key":"405_CR10","doi-asserted-by":"crossref","unstructured":"J.W. Kim, G.W. Lee, C. Park, H.K. Kim, in IEEE International Conference on Consumer Electronics, ICCE 2023, Las Vegas, NV, USA, January 6-8, 2023. Sound event detection using efficientnet-b2 with an attentional pyramid network (IEEE, Piscataway, NJ, USA, 2023), pp. 1\u20132","DOI":"10.1109\/ICCE56470.2023.10043590"},{"key":"405_CR11","doi-asserted-by":"crossref","unstructured":"H. Wang, Y. Zou, D. Chong, W. Wang, in Interspeech 2020, 21st Annual Conference of the International Speech Communication Association, Virtual Event, Shanghai, China, 25-29 October 2020, eds. by H. Meng, B. Xu, T.F. Zheng. Environmental sound classification with parallel temporal-spectral attention (ISCA, Baixas, France, 2020), pp. 821\u2013825","DOI":"10.21437\/Interspeech.2020-1219"},{"key":"405_CR12","doi-asserted-by":"crossref","unstructured":"X. Zheng, Y. Song, J. Yan, L. Dai, I. McLoughlin, L. Liu, in Interspeech 2020, 21st Annual Conference of the International Speech Communication Association, Virtual Event, Shanghai, China, 25-29 October 2020, eds. by H. Meng, B. Xu, T.F. Zheng. An effective perturbation based semi-supervised learning method for sound event detection (ISCA, Baixas, France, 2020), pp. 841\u2013845","DOI":"10.21437\/Interspeech.2020-2329"},{"key":"405_CR13","unstructured":"X. Zheng, H. Chen, Y. Song, Zheng USTC team\u2019s submission for DCASE 2021 task4\u2013semi-supervised sound event detection. Tech. Rep. DCASE Challenge 2021 Task4 (2021),\u00a0 pp. 1\u20133"},{"key":"405_CR14","doi-asserted-by":"crossref","unstructured":"C. Koh, Y. Chen, Y. Liu, M.R. Bai, in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2021, Toronto, ON, Canada, June 6-11, 2021. Sound event detection by consistency training and pseudo-labeling with feature-pyramid convolutional recurrent neural networks (IEEE, Piscataway, NJ, USA, 2021), pp. 376\u2013380","DOI":"10.1109\/ICASSP39728.2021.9414350"},{"key":"405_CR15","doi-asserted-by":"crossref","unstructured":"L. Lin, X. Wang, H. Liu, Y. Qian, in 2020 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2020, Barcelona, Spain, May 4-8, 2020. Guided learning for weakly-labeled semi-supervised sound event detection (IEEE, Piscataway, NJ, USA, 2020), pp. 626\u2013630","DOI":"10.1109\/ICASSP40776.2020.9053584"},{"key":"405_CR16","doi-asserted-by":"crossref","unstructured":"N. Turpault, R. Serizel, A. Parag Shah, J. Salamon, in Workshop on Detection and Classification of Acoustic Scenes and Events. Sound event detection in domestic environments with weakly labeled data and soundscape synthesis (New York City, 2019). https:\/\/hal.inria.fr\/hal-02160855. Accessed: 2024 Aug 15","DOI":"10.33682\/006b-jx26"},{"key":"405_CR17","doi-asserted-by":"crossref","unstructured":"Y. Hao, H. Zhang, C. Ngo, X. He, in IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022. Group contextualization for video recognition (IEEE, Piscataway, NJ, USA, 2022), pp. 918\u2013928","DOI":"10.1109\/CVPR52688.2022.00100"},{"key":"405_CR18","doi-asserted-by":"publisher","unstructured":"J. Hu, L. Shen and G. Sun, in CVPR 2018, IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 2018. Squeeze-and-Excitation Networks (IEEE, Piscataway, NJ, USA, 2018), pp. 7132-7141, https:\/\/doi.org\/10.1109\/CVPR.2018.00745","DOI":"10.1109\/CVPR.2018.00745"},{"key":"405_CR19","doi-asserted-by":"crossref","unstructured":"H. Tak, J. Jung, J. Patino, M. Todisco, N.W.D. Evans, in Interspeech 2021, 22nd Annual Conference of the International Speech Communication Association, Brno, Czechia, 30 August - 3 September 2021, eds. by H. Hermansky, H. Cernock\u00fd, L. Burget, L. Lamel, O. Scharenborg, P. Motl\u00edcek. Graph attention networks for anti-spoofing (ISCA, Baixas, France, 2021), pp. 2356\u20132360","DOI":"10.21437\/Interspeech.2021-993"},{"key":"405_CR20","doi-asserted-by":"crossref","unstructured":"Y. Chen, X. Dai, M. Liu, D. Chen, L. Yuan, Z. Liu, in Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XIX, eds. by A. Vedaldi, H. Bischof, T. Brox, J. Frahm. Dynamic relu. Lecture Notes in Computer Science, vol. 12364 (Springer, Cham, Switzerland, 2020), pp. 351\u2013367","DOI":"10.1007\/978-3-030-58529-7_21"},{"issue":"3","key":"405_CR21","doi-asserted-by":"publisher","first-page":"2587","DOI":"10.1109\/TIE.2020.2972458","volume":"68","author":"M Zhao","year":"2021","unstructured":"M. Zhao, S. Zhong, X. Fu, B. Tang, S. Dong, M.G. Pecht, Deep residual networks with adaptively parametric rectifier linear units for fault diagnosis. IEEE Trans. Ind. Electron. 68(3), 2587\u20132597 (2021)","journal-title":"IEEE Trans. Ind. Electron."},{"key":"405_CR22","doi-asserted-by":"crossref","unstructured":"R. Serizel, N. Turpault, A. Shah, J. Salamon, in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Sound event detection in synthetic domestic environments (IEEE, Piscataway, NJ, USA, 2020), pp. 86\u201390","DOI":"10.1109\/ICASSP40776.2020.9054478"},{"key":"405_CR23","unstructured":"S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, F. Wei, in Proceedings of the 40th International Conference on Machine Learning, ICML'23. Beats: Audio pre-training with acoustic tokenizers. PMLR, vol. 202 (2023), pp. 5178\u20135193"},{"key":"405_CR24","doi-asserted-by":"crossref","unstructured":"C. Bilen, G. Ferroni, F. Tuveri, J. Azcarreta, S. Krstulovic, in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). A framework for the robust evaluation of sound event detection (IEEE, Piscataway, NJ, USA, 2020), pp. 61\u201365","DOI":"10.1109\/ICASSP40776.2020.9052995"},{"key":"405_CR25","unstructured":"J.W. Kim, S.W. Son, Y. Song, H.K. Kim, I.H. Song, J.E. Lim, Semi-supervised learning-based sound event detection using frequency dynamic convolution with large kernel attention for DCASE challenge 2023 task 4. Technical report, DCASE2023 Challenge (2023)"},{"key":"405_CR26","unstructured":"S. Xiao, J. Shen, A. Hu, X. Zhang, P. Zhang, Y. Yan, Sound event detection with weak prediction for dcase 2023 challenge task4a. Technical report, DCASE2023 Challenge (2023)"},{"key":"405_CR27","unstructured":"F.X. Duo, Wenxin, J. Li, Semi-supervised sound event detection system for DCASE 2023 task4a. Technical report, DCASE2023 Challenge (2023)"},{"key":"405_CR28","unstructured":"Y. Xiao, T. Khandelwal, R.K. Das, FMSG submission for DCASE 2023 challenge task 4 on sound event detection with weak labels and synthetic soundscapes. Technical report, DCASE2023 Challenge (2023)"},{"key":"405_CR29","unstructured":"Y. Guan, Q. Shang, Semi-supervised sound event detection system for DCASE 2023 task 4. Technical report, DCASE2023 Challenge (2023)"},{"key":"405_CR30","doi-asserted-by":"publisher","first-page":"902","DOI":"10.1109\/TASLP.2022.3233468","volume":"31","author":"I Mart\u00edn-Morat\u00f3","year":"2023","unstructured":"I. Mart\u00edn-Morat\u00f3, A. Mesaros, Strong labeling of sound events using crowdsourced weak labels and annotator competence estimation. IEEE\/ACM Trans. Audio Speech Lang. Process. 31, 902\u2013914 (2023)","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"405_CR31","unstructured":"H. Yin, J. Bai, S. Huang, J. Chen, How information on soft labels and hard labels mutually benefits sound event detection tasks. Technical report, DCASE2023 Challenge (2023)"},{"key":"405_CR32","unstructured":"Y. Jin, M. Chen, J. Shao, Y. Liu, B. Peng, J. Chen, DCASE 2023 challenge task4 technical report. Technical report, DCASE2023 Challenge (2023)"},{"key":"405_CR33","unstructured":"X. Xuenan, M. Ziyang, Y. Fei, Y. Guanrou, W. Mengyue, C. Xie, Sound event detection by aggregating pre-trained embeddings from different layers. Technical report, DCASE2023 Challenge (2023)"},{"key":"405_CR34","unstructured":"D. Min, H. Nam, P. Yong-Hwa, Application of spectro-temporal receptive field for DCASE 2023 challenge task4 b. Technical report, DCASE2023 Challenge (2023)"},{"key":"405_CR35","unstructured":"T.D. Nhan, B. Param, Y. Zhang, Sound event detection with soft labels using self-attention mechanisms for global scene feature extraction. Technical report, DCASE2023 Challenge (2023)"},{"key":"405_CR36","unstructured":"F. Schmid, P. Primus, T. Morocutti, J. Greif, G. Widmer, Improving audio spectrogram transformers for sound event detection through multi-stage training. Technical report, DCASE2024 Challenge (2024)"},{"key":"405_CR37","unstructured":"H. Nam, D. Min, S. Choi, I. Choi, Y.H. Park, Self training and ensembling frequency dependent networks with coarse prediction pooling and sound event bounding boxes. Technical report, DCASE2024 Challenge (2024)"},{"key":"405_CR38","unstructured":"H. Yue, Z. Wang, D. Mu, H. Sun, Y. Jiang, Z. Zhanf, J. Yin, Local and global features fusion for sound event detection with heterogeneous training dataset and potentially missing labels. Technical report, DCASE2024 Challenge (2024)"},{"key":"405_CR39","unstructured":"W.Y. Chen, C.L. Lu, H.F. Chuang, Y.H. Cheng, B.C.C. Chan, Sound event detection with heterogeneous training dataset and potentially missing labels for DCASE 2024 task 4. Technical report, DCASE2024 Challenge (2024)"},{"key":"405_CR40","unstructured":"S.W. Son, J. Park, H.K. Kim, S. Vesal, J.E. Lim, Sound event detection based on auxiliary decoder and maximum probability aggregation for DCASE challenge 2024 task 4. Technical report, DCASE2024 Challenge (2024)"},{"key":"405_CR41","unstructured":"A. Miech, I. Laptev, J. Sivic, Learnable pooling with context gating for video classification. arXiv preprint (2017), arXiv:1706.06905."}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-025-00405-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-025-00405-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-025-00405-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,16]],"date-time":"2025-06-16T17:58:38Z","timestamp":1750096718000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-025-00405-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,16]]},"references-count":41,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["405"],"URL":"https:\/\/doi.org\/10.1186\/s13636-025-00405-6","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,16]]},"assertion":[{"value":"6 May 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 April 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 June 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"23"}}