{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T12:43:20Z","timestamp":1781354600035,"version":"3.54.1"},"reference-count":30,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,6,30]],"date-time":"2023-06-30T00:00:00Z","timestamp":1688083200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,6,30]],"date-time":"2023-06-30T00:00:00Z","timestamp":1688083200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The goal of sound event detection and localization (SELD) is to identify each individual sound event class and its activity time from a piece of audio, while estimating its spatial location at the time of activity. Conformer combines the advantages of convolutional layers and Transformer, which is effective in tasks such as speech recognition. However, it achieves high performance relying on complex network structure and a large number of computations. In the SELD task of this paper, we propose to use an encoder with a simpler network structure, called the dual-branch attention module (DBAM). The module is improved based on the conformer using two parallel branches of attention and convolution, which can model both global and local contextual information. We also blend low-level and high-level features of the localization task. In addition, we add soft parameter sharing to the joint SELD network, which can efficiently exploit the potential relationship between the two subtasks, SED and DOA. The proposed method can effectively detect two sound events with overlapping occurrence in the same time period. We experimented with the open dataset DCASE 2020 task 3 proving that the proposed method achieves better SELD performance than the baseline model. Furthermore, we conducted ablation experiments for verifying the effectiveness of the dual-branch attention module and soft parameter sharing.<\/jats:p>","DOI":"10.1186\/s13636-023-00292-9","type":"journal-article","created":{"date-parts":[[2023,6,30]],"date-time":"2023-06-30T15:02:57Z","timestamp":1688137377000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Dual-branch attention module-based network with parameter sharing for joint sound event detection and localization"],"prefix":"10.1186","volume":"2023","author":[{"given":"Yuting","family":"Zhou","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4982-0689","authenticated-orcid":false,"given":"Hongjie","family":"Wan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,6,30]]},"reference":[{"issue":"1","key":"292_CR1","doi-asserted-by":"publisher","first-page":"279","DOI":"10.1109\/TITS.2015.2470216","volume":"17","author":"P Foggia","year":"2016","unstructured":"P. Foggia, N. Petkov, A. Saggese, N. Strisciuglio, M. Vento, Audio surveillance of roads: a system for detecting anomalous sounds. J IEEE Transactions on Intelligent Transportation Systems. 17(1), 279\u2013288 (2016)","journal-title":"J IEEE Transactions on Intelligent Transportation Systems."},{"key":"292_CR2","first-page":"6119","volume-title":"IECON 2017\u201343rd Annual Conference of the IEEE Industrial Electronics Society. Sound based localization and identification in industrial environments","author":"C Grobler","year":"2017","unstructured":"C. Grobler, C. Kruger, B. Silva, G. Hancke, IECON 2017\u201343rd Annual Conference of the IEEE Industrial Electronics Society. Sound based localization and identification in industrial environments (Beijing, IEEE, 2017), pp.6119\u20136124"},{"key":"292_CR3","first-page":"21","volume-title":"2007 Conference on Advanced Video and Signal Based Surveillance. Scream and gunshot detection and localization for audio-surveillance systems","author":"G Valenzise","year":"2007","unstructured":"G. Valenzise, L. Gerosa, M. Tagliasacchi, F. Antonacci, A. Sarti, 2007 Conference on Advanced Video and Signal Based Surveillance. Scream and gunshot detection and localization for audio-surveillance systems (IEEE, London, 2007), pp.21\u201326"},{"key":"292_CR4","doi-asserted-by":"crossref","unstructured":"C. Busso, S. Hernanz, C. W.Chu et al., IEEE International Conference on Acoustics, Speech, and Signal Processing. Smart room: participant and speaker localization and identification. (ICASSP, Philadelphia, PA, USA, 2005), pp. ii\/1117-ii\/1120, Vol. 2","DOI":"10.1109\/ICASSP.2005.1415605"},{"issue":"3","key":"292_CR5","doi-asserted-by":"publisher","first-page":"16","DOI":"10.1109\/MSP.2014.2326181","volume":"32","author":"D Barchiesi","year":"2015","unstructured":"D. Barchiesi, D. Giannoulis, D. Stowell, M.D. Plumbley, Acoustic scene classification: classifying environments from the sounds they produce. J IEEE Signal Processing Magazine. 32(3), 16\u201334 (2015)","journal-title":"J IEEE Signal Processing Magazine."},{"key":"292_CR6","first-page":"187","volume-title":"1997 IEEE International Conference on Acoustics, Speech, and Signal Processing. Voice source localization for automatic camera pointing system in videoconferencing","author":"H Wang","year":"1997","unstructured":"H. Wang, P. Chu, 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing. Voice source localization for automatic camera pointing system in videoconferencing (ICASSP, Munich, 1997), pp.187\u2013190"},{"key":"292_CR7","first-page":"1","volume-title":"2015IEEE 25th International Workshop on Machine Learning for Signal Processing.\u00a0Environmental sound classification with convolutional neural networks.","author":"KJ Piczak","year":"2015","unstructured":"K.J. Piczak, 2015IEEE 25th International Workshop on Machine Learning for Signal Processing.\u00a0Environmental sound classification with convolutional neural networks. (MLSP, Boston, 2015), pp.1\u20136"},{"key":"292_CR8","first-page":"6440","volume-title":"2016 IEEE International Conference on Acoustics Speech and Signal Processing. Recurrent neural networks for polyphonic sound event detection in real life recordings","author":"G Parascandolo","year":"2016","unstructured":"G. Parascandolo, H. Huttunen, T. Virtanen, 2016 IEEE International Conference on Acoustics Speech and Signal Processing. Recurrent neural networks for polyphonic sound event detection in real life recordings (ICASSP, Shanghai, 2016), pp.6440\u20136444"},{"key":"292_CR9","volume-title":"Audio Engineering Society 138th Convention, Classification of spatial audio location and content using convolutional neural networks","author":"T Hirvonen","year":"2015","unstructured":"T. Hirvonen, Audio Engineering Society 138th Convention, Classification of spatial audio location and content using convolutional neural networks (AES, Warsaw, 2015)"},{"key":"292_CR10","first-page":"771","volume-title":"2017 IEEE International Conference on Acoustics Speech and Signal Processing. Sound event detection using spatial features and convolutional recurrent neural network","author":"S Adavanne","year":"2017","unstructured":"S. Adavanne, P. Pertil\u00e4, T. Virtanen, 2017 IEEE International Conference on Acoustics Speech and Signal Processing. Sound event detection using spatial features and convolutional recurrent neural network (ICASSP, New Orleans, 2017), pp.771\u2013775"},{"key":"292_CR11","doi-asserted-by":"crossref","unstructured":"Y. Cao, Q. Kong, T. Iqbal, et al., polyphonic sound event detection and localization using a two-stage strategy. (2019). ArXiv Preprint arXiv:1905.00268","DOI":"10.33682\/4jhy-bj81"},{"key":"292_CR12","unstructured":"A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, et al., Attention is all you need. Advances in neural information processing systems, 2017, pp. 5998\u20136008. ArXiv Preprint arXiv: 1706.03762"},{"key":"292_CR13","doi-asserted-by":"publisher","unstructured":"Heittola, T., Mesaros, A., Eronen, A. et al. Context-dependent sound event detection. J EURASIP Journal on Audio, Speech, and Music Processing. (2013). https:\/\/doi.org\/10.1186\/1687-4722-2013-1","DOI":"10.1186\/1687-4722-2013-1"},{"key":"292_CR14","first-page":"5884","volume-title":"2018 IEEE International Conference on Acoustics, Speech and Signal Processing. Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition","author":"L Dong","year":"2018","unstructured":"L. Dong, S. Xu, B. Xu, 2018 IEEE International Conference on Acoustics, Speech and Signal Processing. Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition (ICASSP, Calgary, 2018), pp.5884\u20135888"},{"key":"292_CR15","doi-asserted-by":"crossref","unstructured":"Gulati, A., Qin, J., Chiu, C.-C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., and Pang, R. Conformer: Convolution-augmented Transformer for Speech Recognition. In Proceedings of INTERSPEECH, 2020. ArXiv Preprint arXiv:2005.08100","DOI":"10.21437\/Interspeech.2020-3015"},{"key":"292_CR16","unstructured":"Wu, Z., Liu, Z., Lin, J., Lin, Y., and Han, S. Lite Transformer with long-short range attention. In Proceedings of ICLR, 2020. ArXiv Preprint arXiv: 2004.11886"},{"key":"292_CR17","unstructured":"Yifan Peng, Siddharth Dalmia, Ian Lane, Shinji Watanabe, Branchformer: Parallel MLP-Attention architectures to capture local and global context for speech recognition and understanding. ICML 2022. ArXiv Preprint arXiv:2207.02971."},{"key":"292_CR18","unstructured":"Hendrycks, D. and Gimpel, K. Gaussian error linear units (GELUs). (2016) ArXiv preprint arXiv:1606.08415"},{"key":"292_CR19","doi-asserted-by":"crossref","unstructured":"Yang Zhang, Zhiqiang Lv, Haibin Wu, Shanshan Zhang, Pengfei Hu, MFA-Conformer: Multi-scale feature aggregation conformer for automatic speaker verification. INTERSPEECH 2022. ArXiv preprint arXiv: 2203.15249","DOI":"10.21437\/Interspeech.2022-563"},{"key":"292_CR20","first-page":"333","volume-title":"2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics. Joint measurement of localization and detection of sound events","author":"A Mesaros","year":"2019","unstructured":"A. Mesaros, S. Adavanne, A. Politis, T. Heittola, T. Virtanen, 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics. Joint measurement of localization and detection of sound events (WASPAA, New Paltz, 2019), pp.333\u2013337"},{"key":"292_CR21","first-page":"1131","volume-title":"2017 IEEE Conference on Computer Vision and Pattern Recognition. Fully-adaptive feature sharing in multi-task networks with applications in person attribute classification","author":"Y Lu","year":"2017","unstructured":"Y. Lu, A. Kumar, S. Zhai, Y. Cheng, T. Javidi, R. Feris, 2017 IEEE Conference on Computer Vision and Pattern Recognition. Fully-adaptive feature sharing in multi-task networks with applications in person attribute classification (CVPR, Honolulu, 2017), pp.1131\u20131140"},{"key":"292_CR22","first-page":"3994","volume-title":"2016 IEEE Conference on Computer Vision and Pattern Recognition. Cross-stitch networks for multi-task learning","author":"I Mistra","year":"2016","unstructured":"I. Mistra, A. Shrivastava, A. Gupta, M. Hebert, 2016 IEEE Conference on Computer Vision and Pattern Recognition. Cross-stitch networks for multi-task learning (CVPR, Las Vegas, 2016), pp.3994\u20134003"},{"key":"292_CR23","first-page":"1923","volume-title":"2017 Conference on Empirical Methods in Natural Language Processing. a joint many-task model: growing a neural network for multiple NLP tasks","author":"K Hashimoto","year":"2017","unstructured":"K. Hashimoto, C. Xiong, Y. Tsuruoka, R. Socher, 2017 Conference on Empirical Methods in Natural Language Processing. a joint many-task model: growing a neural network for multiple NLP tasks (EMNLP, Copenhagen, 2017), pp.1923\u20131933"},{"key":"292_CR24","doi-asserted-by":"publisher","first-page":"7482","DOI":"10.1109\/CVPR.2018.00781","volume-title":"2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics","author":"R Cipolla","year":"2018","unstructured":"R. Cipolla, Y. Gal, A. Kendall, 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics (CVPR, Salt Lake City, 2018), pp.7482\u20137491"},{"key":"292_CR25","unstructured":"Ruder, Sebastian, Joachim Bingel, Isabelle Augenstein and Anders S\u00f8gaard, Sluice networks: learning what to share between loosely related tasks. (2017). ArXiv preprint arXiv: 1705.08142"},{"key":"292_CR26","first-page":"885","volume-title":"2021 IEEE International Conference on Acoustics, Speech and Signal Processing. An improved event-independent network for polyphonic sound event localization and detection","author":"Y Cao","year":"2021","unstructured":"Y. Cao, T. Iqbal, Q. Kong, F. An, W. Wang, M. D. Plumbley, 2021 IEEE International Conference on Acoustics, Speech and Signal Processing. An improved event-independent network for polyphonic sound event localization and detection (ICASSP, Toronto, 2021), pp.885\u2013889"},{"key":"292_CR27","first-page":"241","volume-title":"2017 IEEE International Conference on Acoustics, Speech and Signal Processing. Permutation invariant training of deep models for speaker-independent multi-talker speech separation","author":"D Yu","year":"2017","unstructured":"D. Yu, M. Kolb\u00e6k, Z.- H. Tan, J. Jensen, 2017 IEEE International Conference on Acoustics, Speech and Signal Processing. Permutation invariant training of deep models for speaker-independent multi-talker speech separation (ICASSP, New Orleans, 2017), pp.241\u2013245"},{"key":"292_CR28","unstructured":"A. Politis, S. Adavanne and T. Virtanen, A dataset of reverberant spatial sound scenes with moving sources for sound event localization and detection. Proc. DCASE 2020 Workshop, 2020, pp. 165\u2013169. ArXiv preprint arXiv: 2006.01919"},{"key":"292_CR29","unstructured":"Y Cao T Iqbal Q Kong Y Zhong W Wang MD Plumbley Event-independent network for polyphonic sound event localization and detection. Proc. DCASE, 2020 Workshop, 2020 ArXiv preprint arXiv 2010 00140"},{"key":"292_CR30","doi-asserted-by":"crossref","unstructured":"T. N. T. Nguyen, Douglas L. Jones, W. Gan, DCASE 2020 TASK 3: Ensemble of sequence matching networks for dynamic sound event localization, detection and tracking. Proc. DCASE 2020 Workshop, 2020, pp. 120\u2013124","DOI":"10.1109\/ICASSP40776.2020.9053045"}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-023-00292-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-023-00292-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-023-00292-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,23]],"date-time":"2024-10-23T12:01:48Z","timestamp":1729684908000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-023-00292-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,30]]},"references-count":30,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["292"],"URL":"https:\/\/doi.org\/10.1186\/s13636-023-00292-9","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,30]]},"assertion":[{"value":"4 February 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 June 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 June 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"The authors declare that they have no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"27"}}