{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,20]],"date-time":"2025-11-20T12:58:58Z","timestamp":1763643538269,"version":"3.37.3"},"reference-count":49,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2022,6,18]],"date-time":"2022-06-18T00:00:00Z","timestamp":1655510400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,6,18]],"date-time":"2022-06-18T00:00:00Z","timestamp":1655510400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100003130","name":"Fonds Wetenschappelijk Onderzoek","doi-asserted-by":"publisher","award":["11G0721N"],"award-info":[{"award-number":["11G0721N"]}],"id":[{"id":"10.13039\/501100003130","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003130","name":"Fonds Wetenschappelijk Onderzoek","doi-asserted-by":"publisher","award":["G081420N"],"award-info":[{"award-number":["G081420N"]}],"id":[{"id":"10.13039\/501100003130","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"published-print":{"date-parts":[[2022,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>By means of spatial clustering and time-frequency masking, a mixture of multiple speakers and noise can be separated into the underlying signal components. The parameters of a model, such as a complex angular central Gaussian mixture model (cACGMM), can be determined based on the given signal mixture itself. Then, no misfit between training and testing conditions arises, as opposed to approaches that require labeled datasets to be trained. Whereas the separation can be performed in a completely unsupervised way, it may be beneficial to take advantage of a priori knowledge. The parameter estimation is sensitive to the initialization, and it is necessary to address the frequency permutation problem. In this paper, we therefore consider three techniques to overcome these limitations using direction of arrival (DOA) estimates. First, we propose an initialization with simple DOA-based masks. Secondly, we derive speaker specific time annotations from the same masks in order to constrain the cACGMM. Thirdly, we employ an approach where the mixture components are specific to each DOA instead of each speaker. We conduct experiments with sudden DOA changes, as well as a gradually moving speaker. The results demonstrate that particularly the DOA-based initialization is effective to overcome both of the described limitations. In this case, even methods based on normally unavailable oracle information are not observed to be more beneficial to the permutation resolution or the initialization. Lastly, we also show that the proposed DOA-guided source separation works quite robustly in the presence of adverse conditions and realistic DOA estimation errors.<\/jats:p>","DOI":"10.1186\/s13636-022-00246-7","type":"journal-article","created":{"date-parts":[[2022,6,18]],"date-time":"2022-06-18T08:02:53Z","timestamp":1655539373000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["DOA-guided source separation with direction-based initialization and time annotations using complex angular central Gaussian mixture models"],"prefix":"10.1186","volume":"2022","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1819-0482","authenticated-orcid":false,"given":"Alexander","family":"Bohlender","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lucas Van","family":"Severen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jonathan","family":"Sterckx","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nilesh","family":"Madhu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,6,18]]},"reference":[{"key":"246_CR1","doi-asserted-by":"publisher","unstructured":"S. Rickard, O. Yilmaz, in Proc. IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 1. On the approximate W-disjoint orthogonality of speech, (2002), pp. 529\u2013532. https:\/\/doi.org\/10.1109\/ICASSP.2002.5743771.","DOI":"10.1109\/ICASSP.2002.5743771"},{"key":"246_CR2","doi-asserted-by":"publisher","unstructured":"D. Yu, M. Kolb\u00e6k, Z. -H. Tan, J. Jensen, in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Permutation invariant training of deep models for speaker-independent multi-talker speech separation, (2017), pp. 241\u2013245. https:\/\/doi.org\/10.1109\/ICASSP.2017.7952154.","DOI":"10.1109\/ICASSP.2017.7952154"},{"key":"246_CR3","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s13636-016-0085-x","volume":"1","author":"Y. Yu","year":"2016","unstructured":"Y. Yu, W. Wang, P. Han, Localization based stereo speech source separation using probabilistic time-frequency masking and deep neural networks. EURASIP J. Audio Speech Music Process.1:, 1\u201318 (2016). https:\/\/doi.org\/10.1186\/s13636-016-0085-x.","journal-title":"EURASIP J. Audio Speech Music Process."},{"key":"246_CR4","doi-asserted-by":"publisher","unstructured":"S. E. Chazan, H. Hammer, G. Hazan, J. Goldberger, S. Gannot, in Proc. 27th European Signal Processing Conference (EUSIPCO). Multi-microphone speaker separation based on deep DOA estimation, (2019), pp. 1\u20135. https:\/\/doi.org\/10.23919\/EUSIPCO.2019.8903121.","DOI":"10.23919\/EUSIPCO.2019.8903121"},{"key":"246_CR5","doi-asserted-by":"publisher","unstructured":"Z. Chen, X. Xiao, T. Yoshioka, H. Erdogan, J. Li, Y. Gong, in Proc. IEEE Spoken Language Technology Workshop (SLT). Multi-channel overlapped speech recognition with location guided speech extraction network, (2018), pp. 558\u2013565. https:\/\/doi.org\/10.1109\/SLT.2018.8639593.","DOI":"10.1109\/SLT.2018.8639593"},{"key":"246_CR6","doi-asserted-by":"publisher","unstructured":"A. Bohlender, A. Spriet, W. Tirry, N. Madhu, in Proc. 29th European Signal Processing Conference (EUSIPCO). Neural networks using full-band and subband spatial features for mask based source separation, (2021), pp. 346\u2013350. https:\/\/doi.org\/10.23919\/EUSIPCO54536.2021.9616138.","DOI":"10.23919\/EUSIPCO54536.2021.9616138"},{"key":"246_CR7","doi-asserted-by":"publisher","unstructured":"A. Aroudi, S. Braun, in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). DBnet: Doa-driven beamforming network for end-to-end reverberant sound source separation, (2021), pp. 211\u2013215. https:\/\/doi.org\/10.1109\/ICASSP39728.2021.9414187.","DOI":"10.1109\/ICASSP39728.2021.9414187"},{"key":"246_CR8","doi-asserted-by":"publisher","unstructured":"J. R. Hershey, Z. Chen, J. Le Roux, S. Watanabe, in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Deep clustering: Discriminative embeddings for segmentation and separation, (2016), pp. 31\u201335. https:\/\/doi.org\/10.1109\/ICASSP.2016.7471631.","DOI":"10.1109\/ICASSP.2016.7471631"},{"key":"246_CR9","doi-asserted-by":"publisher","unstructured":"Z. -Q. Wang, J. Le Roux, J. R. Hershey, in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Multi-channel deep clustering: Discriminative spectral and spatial embeddings for speaker-independent speech separation, (2018), pp. 1\u20135. https:\/\/doi.org\/10.1109\/ICASSP.2018.8461639.","DOI":"10.1109\/ICASSP.2018.8461639"},{"key":"246_CR10","doi-asserted-by":"publisher","unstructured":"Z. Chen, Y. Luo, N. Mesgarani, in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Deep attractor network for single-microphone speaker separation, (2017), pp. 246\u2013250. https:\/\/doi.org\/10.1109\/ICASSP.2017.7952155.","DOI":"10.1109\/ICASSP.2017.7952155"},{"key":"246_CR11","doi-asserted-by":"publisher","unstructured":"N. Ito, S. Araki, T. Nakatani, in Proc. 24th European Signal Processing Conference (EUSIPCO). Complex angular central gaussian mixture model for directional statistics in mask-based microphone array signal processing, (2016), pp. 1153\u20131157. https:\/\/doi.org\/10.1109\/EUSIPCO.2016.7760429.","DOI":"10.1109\/EUSIPCO.2016.7760429"},{"issue":"3","key":"246_CR12","doi-asserted-by":"publisher","first-page":"516","DOI":"10.1109\/TASL.2010.2051355","volume":"19","author":"H. Sawada","year":"2011","unstructured":"H. Sawada, S. Araki, S. Makino, Underdetermined convolutive blind source separation via frequency bin-wise clustering and permutation alignment. IEEE Trans. Audio Speech Lang. Process.19(3), 516\u2013527 (2011).","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"246_CR13","doi-asserted-by":"publisher","unstructured":"N. Ito, S. Araki, T. Nakatani, in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing. Permutation-free convolutive blind source separation via full-band clustering based on frequency-independent source presence priors, (2013), pp. 3238\u20133242. https:\/\/doi.org\/10.1109\/ICASSP.2013.6638256.","DOI":"10.1109\/ICASSP.2013.6638256"},{"key":"246_CR14","doi-asserted-by":"publisher","unstructured":"J. Azcarreta, N. Ito, S. Araki, T. Nakatani, in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Permutation-free cGMM: Complex gaussian mixture model with inverse wishart mixture model based spatial prior for permutation-free source separation and source counting, (2018), pp. 51\u201355. https:\/\/doi.org\/10.1109\/ICASSP.2018.8461934.","DOI":"10.1109\/ICASSP.2018.8461934"},{"key":"246_CR15","doi-asserted-by":"publisher","unstructured":"C. Boeddeker, J. Heitkaemper, J. Schmalenstroeer, L. Drude, J. Heymann, R. Haeb-Umbach, in Proc. 5th International Workshop on Speech Processing in Everyday Environments (CHiME). Front-end processing for the chime-5 dinner party scenario, (2018). https:\/\/doi.org\/10.21437\/CHiME.2018-8.","DOI":"10.21437\/CHiME.2018-8"},{"key":"246_CR16","doi-asserted-by":"publisher","unstructured":"L. Drude, D. Hasenklever, R. Haeb-Umbach, in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Unsupervised training of a deep clustering model for multichannel blind source separation, (2019), pp. 695\u2013699. https:\/\/doi.org\/10.1109\/ICASSP.2019.8683520.","DOI":"10.1109\/ICASSP.2019.8683520"},{"key":"246_CR17","doi-asserted-by":"publisher","unstructured":"T. Nakatani, R. Takahashi, T. Ochiai, K. Kinoshita, R. Ikeshita, M. Delcroix, S. Araki, in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). DNN-supported mask-based convolutional beamforming for simultaneous denoising, dereverberation, and source separation, (2020), pp. 6399\u20136403. https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9053343.","DOI":"10.1109\/ICASSP40776.2020.9053343"},{"key":"246_CR18","doi-asserted-by":"crossref","unstructured":"T. Nakatani, N. Ito, T. Higuchi, S. Araki, K. Kinoshita, in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamforming, (2017), pp. 286\u2013290.","DOI":"10.1109\/ICASSP.2017.7952163"},{"issue":"4","key":"246_CR19","doi-asserted-by":"publisher","first-page":"815","DOI":"10.1109\/JSTSP.2019.2912565","volume":"13","author":"L. Drude","year":"2019","unstructured":"L. Drude, R. Haeb-Umbach, Integration of neural networks and probabilistic spatial models for acoustic blind source separation. IEEE J. Sel. Top. Signal Process.13(4), 815\u2013826 (2019). https:\/\/doi.org\/10.1109\/JSTSP.2019.2912565.","journal-title":"IEEE J. Sel. Top. Signal Process."},{"key":"246_CR20","doi-asserted-by":"publisher","unstructured":"D. H. T. Vu, R. Haeb-Umbach, in Proc. 12th Annual Conference of the International Speech Communication Association (INTERSPEECH). On initial seed selection for frequency domain blind speech separation, (2011), pp. 1757\u20131760. https:\/\/doi.org\/10.21437\/Interspeech.2011-494.","DOI":"10.21437\/Interspeech.2011-494"},{"key":"246_CR21","doi-asserted-by":"publisher","unstructured":"Y. Bando, Y. Sasaki, K. Yoshii, in 2019 IEEE 29th International Workshop on Machine Learning for Signal Processing (MLSP). Deep bayesian unsupervised source separation based on a complex gaussian mixture model, (2019), pp. 1\u20136. https:\/\/doi.org\/10.1109\/MLSP.2019.8918699.","DOI":"10.1109\/MLSP.2019.8918699"},{"key":"246_CR22","doi-asserted-by":"publisher","unstructured":"J. Barker, S. Watanabe, E. Vincent, J. Trmal, in Proc. 19th Annual Conference of the International Speech Communication Association (INTERSPEECH). The fifth \u2019CHiME\u2019 speech separation and recognition challenge: Dataset, task and baselines, (2018), pp. 1561\u20131565. https:\/\/doi.org\/10.21437\/Interspeech.2018-1768.","DOI":"10.21437\/Interspeech.2018-1768"},{"issue":"2","key":"246_CR23","doi-asserted-by":"publisher","first-page":"666","DOI":"10.1109\/TSA.2005.855832","volume":"14","author":"H. Saruwatari","year":"2006","unstructured":"H. Saruwatari, T. Kawamura, T. Nishikawa, A. Lee, K. Shikano, Blind source separation based on a fast-convergence algorithm combining ICA and beamforming. IEEE Trans. Audio Speech Lang. Processing. 14(2), 666\u2013678 (2006). https:\/\/doi.org\/10.1109\/TSA.2005.855832.","journal-title":"IEEE Trans. Audio Speech Lang. Processing"},{"key":"246_CR24","doi-asserted-by":"publisher","unstructured":"J. Heymann, L. Drude, R. Haeb-Umbach, in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Neural network based spectral mask estimation for acoustic beamforming, (2016), pp. 196\u2013200. https:\/\/doi.org\/10.1109\/ICASSP.2016.7471664.","DOI":"10.1109\/ICASSP.2016.7471664"},{"key":"246_CR25","doi-asserted-by":"publisher","unstructured":"T. Higuchi, N. Ito, T. Yoshioka, T. Nakatani, in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Robust MVDR beamforming using time-frequency masks for online\/offline ASR in noise, (2016), pp. 5210\u20135214. https:\/\/doi.org\/10.1109\/ICASSP.2016.7472671.","DOI":"10.1109\/ICASSP.2016.7472671"},{"key":"246_CR26","doi-asserted-by":"publisher","DOI":"10.1002\/0470031743","volume-title":"Digital Speech Transmission - Enhancement, Coding & Error Concealment","author":"P. Vary","year":"2006","unstructured":"P. Vary, R. Martin, Digital Speech Transmission - Enhancement, Coding & Error Concealment (Wiley, Chichester, 2006)."},{"issue":"2","key":"246_CR27","doi-asserted-by":"publisher","first-page":"260","DOI":"10.1109\/TASL.2009.2025790","volume":"18","author":"M. Souden","year":"2010","unstructured":"M. Souden, J. Benesty, S. Affes, On optimal frequency-domain multichannel linear filtering for noise reduction. IEEE Trans. Audio Speech Lang. Process.18(2), 260\u2013276 (2010). https:\/\/doi.org\/10.1109\/TASL.2009.2025790.","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"246_CR28","doi-asserted-by":"publisher","first-page":"135","DOI":"10.1002\/9780470727188.ch6","volume-title":"Advances in Digital Speech Transmission","author":"N. Madhu","year":"2008","unstructured":"N. Madhu, R. Martin, in Advances in Digital Speech Transmission, ed. by R. Martin, U. Heute, and C. Antweiler. Acoustic source localization with microphone arrays (WileyNew York, 2008), pp. 135\u2013170."},{"key":"246_CR29","unstructured":"J. H. DiBiase, A high-accuracy, low-latency technique for talker localization in reverberant environments using microphone arrays. PhD thesis, Brown University, Providence, RI, USA (2000)."},{"issue":"1","key":"246_CR30","doi-asserted-by":"publisher","first-page":"66","DOI":"10.1016\/j.dsp.2010.04.003","volume":"21","author":"M. Cobos","year":"2011","unstructured":"M. Cobos, J. J. Lopez, D. Martinez, Two-microphone multi-speaker localization based on a Laplacian mixture model. Digit. Signal Process.21(1), 66\u201376 (2011). https:\/\/doi.org\/10.1016\/j.dsp.2010.04.003.","journal-title":"Digit. Signal Process."},{"issue":"3","key":"246_CR31","doi-asserted-by":"publisher","first-page":"276","DOI":"10.1109\/TAP.1986.1143830","volume":"34","author":"R. Schmidt","year":"1986","unstructured":"R. Schmidt, Multiple emitter location and signal parameter estimation. IEEE Trans. Antennas Propag.34(3), 276\u2013280 (1986). https:\/\/doi.org\/10.1109\/TAP.1986.1143830.","journal-title":"IEEE Trans. Antennas Propag."},{"key":"246_CR32","unstructured":"P. -A. Grumiaux, S. Kiti\u0107, L. Girin, A. Gu\u00e9rin, A survey of sound source localization with deep learning methods (2021). arXiv:2109.03465. http:\/\/arxiv.org\/abs\/2109.03465. Accessed 28 May 2022."},{"key":"246_CR33","volume-title":"Introduction to Linear Algebra, 5th ed.","author":"G. Strang","year":"2016","unstructured":"G. Strang, Introduction to Linear Algebra, 5th ed. (Wellesley-Cambridge Press, Wellesley, 2016)."},{"issue":"6","key":"246_CR34","doi-asserted-by":"publisher","first-page":"709","DOI":"10.1109\/TSA.2003.818212","volume":"11","author":"I. A. McCowan","year":"2003","unstructured":"I. A. McCowan, H. Bourlard, Microphone array post-filter based on noise field coherence. IEEE Trans. Speech Audio Process.11(6), 709\u2013716 (2003).","journal-title":"IEEE Trans. Speech Audio Process."},{"key":"246_CR35","doi-asserted-by":"publisher","unstructured":"R. Zelinski, in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). A microphone array with adaptive post-filtering for noise reduction in reverberant rooms, (1988), pp. 2578\u20132581. https:\/\/doi.org\/10.1109\/ICASSP.1988.197172.","DOI":"10.1109\/ICASSP.1988.197172"},{"key":"246_CR36","unstructured":"P. Kabal, TSP speech database. Technical report, McGill University, Montreal, Quebec, Canada (2002)."},{"issue":"3","key":"246_CR37","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1109\/TAU.1969.1162058","volume":"17","author":"E. H. Rothauser","year":"1969","unstructured":"E. H. Rothauser, W. D. Chapman, et al, IEEE recommended practice for speech quality measurements. IEEE Trans. Audio Electroacoustics. 17(3), 225\u2013246 (1969). https:\/\/doi.org\/10.1109\/TAU.1969.1162058.","journal-title":"IEEE Trans. Audio Electroacoustics"},{"key":"246_CR38","unstructured":"miniDSP, UMA-16 USB microphone array. https:\/\/www.minidsp.com\/products\/usb-audio-interface\/uma-16-microphone-array. Accessed22 Feb 2022."},{"key":"246_CR39","unstructured":"European Telecommunications Standards Institute, Speech processing, transmission and quality aspects (STQ); speech quality performance in the presence of background noise; part 1: Background noise simulation technique and background noise database (Standard, ETSI ES 202 396-1, 2005). Current version (V1.8.1, published in 2022). https:\/\/www.etsi.org\/deliver\/etsi_es\/202300_202399\/20239601\/01.08.01_60\/es_20239601v010801p.pdf. Accessed 28 May 2022."},{"issue":"10","key":"246_CR40","doi-asserted-by":"publisher","first-page":"2707","DOI":"10.1109\/TASL.2012.2210879","volume":"20","author":"T. Yoshioka","year":"2012","unstructured":"T. Yoshioka, T. Nakatani, Generalization of multi-channel linear prediction methods for blind MIMO impulse response shortening. IEEE Trans. Audio Speech Lang. Process.20(10), 2707\u20132720 (2012). https:\/\/doi.org\/10.1109\/TASL.2012.2210879.","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"246_CR41","unstructured":"L. Drude, J. Heymann, C. Boeddeker, R. Haeb-Umbach, in Proc. 13th ITG Conference on Speech Communication. NARA-WPE: A python package for weighted prediction error dereverberation in numpy and tensorflow for online and offline processing, (2018), pp. 1\u20135. https:\/\/ieeexplore.ieee.org\/document\/8578026."},{"key":"246_CR42","unstructured":"R. Haeb-Umbach, et al., Blind Source Separation (BSS) algorithms. https:\/\/github.com\/fgnt\/pb_bss. Accessed 21 May 2021."},{"key":"246_CR43","doi-asserted-by":"publisher","first-page":"1594","DOI":"10.1109\/TASLP.2021.3067113","volume":"29","author":"A. Bohlender","year":"2021","unstructured":"A. Bohlender, A. Spriet, W. Tirry, N. Madhu, Exploiting temporal context in CNN based multisource DOA estimation. IEEE\/ACM Trans. Audio Speech Lang. Process.29:, 1594\u20131608 (2021).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"1","key":"246_CR44","doi-asserted-by":"publisher","first-page":"8","DOI":"10.1109\/JSTSP.2019.2901664","volume":"13","author":"S. Chakrabarty","year":"2019","unstructured":"S. Chakrabarty, E. A. P. Habets, Multi-speaker DOA estimation using deep convolutional networks trained with noise signals. IEEE J. Sel. Top. Signal Process.13(1), 8\u201321 (2019). https:\/\/doi.org\/10.1109\/JSTSP.2019.2901664.","journal-title":"IEEE J. Sel. Top. Signal Process."},{"issue":"7","key":"246_CR45","doi-asserted-by":"publisher","first-page":"2125","DOI":"10.1109\/TASL.2011.2114881","volume":"19","author":"C. H. Taal","year":"2011","unstructured":"C. H. Taal, R. C. Hendriks, R. Heusdens, J. Jensen, An algorithm for intelligibility prediction of time-frequency weighted noisy speech. IEEE Trans. Audio Speech Lang. Process.19(7), 2125\u20132136 (2011). https:\/\/doi.org\/10.1109\/TASL.2011.2114881.","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"246_CR46","volume-title":"Wideband extension to recommendation P.862 for the assessment of wideband telephone networks and speech codecs","author":"International Telecommunication Union","year":"2007","unstructured":"International Telecommunication Union, Wideband extension to recommendation P.862 for the assessment of wideband telephone networks and speech codecs (Standard, ITU-R P.862.2, Geneva, 2007). https:\/\/www.itu.int\/rec\/T-REC-P.862.2-200711-I\/en. Accessed 28 May 2022."},{"issue":"4","key":"246_CR47","doi-asserted-by":"publisher","first-page":"1462","DOI":"10.1109\/TSA.2005.858005","volume":"14","author":"E. Vincent","year":"2006","unstructured":"E. Vincent, R. Gribonval, C. Fevotte, Performance measurement in blind audio source separation. IEEE Trans. Audio Speech Lang. Process.14(4), 1462\u20131469 (2006).","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"246_CR48","unstructured":"E. A. P. Habets, Signal Generator. https:\/\/github.com\/ehabets\/Signal-Generator. Accessed 28 Oct 2021."},{"issue":"4","key":"246_CR49","doi-asserted-by":"publisher","first-page":"943","DOI":"10.1121\/1.382599","volume":"65","author":"J. B. Allen","year":"1979","unstructured":"J. B. Allen, D. A. Berkley, Image method for efficiently simulating small-room acoustics. J. Acoust. Soc. Am.65(4), 943\u2013950 (1979). https:\/\/doi.org\/10.1121\/1.382599.","journal-title":"J. Acoust. Soc. Am."}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-022-00246-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-022-00246-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-022-00246-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,6,18]],"date-time":"2022-06-18T08:08:29Z","timestamp":1655539709000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-022-00246-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,18]]},"references-count":49,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,12]]}},"alternative-id":["246"],"URL":"https:\/\/doi.org\/10.1186\/s13636-022-00246-7","relation":{},"ISSN":["1687-4722"],"issn-type":[{"type":"electronic","value":"1687-4722"}],"subject":[],"published":{"date-parts":[[2022,6,18]]},"assertion":[{"value":"3 January 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 May 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 June 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"16"}}