{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T13:17:44Z","timestamp":1740143864773,"version":"3.37.3"},"reference-count":45,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,3,3]],"date-time":"2021-03-03T00:00:00Z","timestamp":1614729600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,3,3]],"date-time":"2021-03-03T00:00:00Z","timestamp":1614729600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Localization of multiple speakers using microphone arrays remains a challenging problem, especially in the presence of noise and reverberation. State-of-the-art localization algorithms generally exploit the sparsity of speech in some representation for this purpose. Whereas the broadband approaches exploit time-domain sparsity for multi-speaker localization, narrowband approaches can additionally exploit sparsity and disjointness in the time-frequency representation. Broadband approaches are robust to spatial aliasing but do not optimally exploit the frequency domain sparsity, leading to poor localization performance for arrays with short inter-microphone distances. Narrowband approaches, on the other hand, are vulnerable to spatial aliasing, making them unsuitable for arrays with large inter-microphone spacing. Proposed here is an approach that decomposes a signal spectrum into a weighted sum of <jats:italic>broadband<\/jats:italic> spectral components (atoms) and then exploits signal sparsity in the <jats:italic>time-atom<\/jats:italic> representation for simultaneous multiple source localization. The decomposition into atoms is performed in situ using non-negative matrix factorization (NMF) of the short-term amplitude spectra and the localization estimate is obtained via a broadband steered-response power (SRP) approach for each active atom of a time frame. This SRP-NMF approach thereby combines the advantages of the narrowband and broadband approaches and performs well on the multi-speaker localization task for a broad range of inter-microphone spacings. On tests conducted on real-world data from public challenges such as SiSEC and LOCATA, and on data generated from recorded room impulse responses, the SRP-NMF approach outperforms the commonly used variants of narrowband and broadband localization approaches in terms of source detection capability and localization accuracy.<\/jats:p>","DOI":"10.1186\/s13636-021-00201-y","type":"journal-article","created":{"date-parts":[[2021,3,3]],"date-time":"2021-03-03T21:03:34Z","timestamp":1614805414000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["NMF-weighted SRP for multi-speaker direction of arrival estimation: robustness to spatial aliasing while exploiting sparsity in the atom-time domain"],"prefix":"10.1186","volume":"2021","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1742-5668","authenticated-orcid":false,"given":"Sushmita","family":"Thakallapalli","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Suryakanth V.","family":"Gangashetty","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nilesh","family":"Madhu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,3,3]]},"reference":[{"key":"201_CR1","doi-asserted-by":"publisher","unstructured":"S. Rickard, O. Yilmaz, in 2002 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 1. On the approximate W-disjoint orthogonality of speech, (2002), pp. 529\u2013532. https:\/\/doi.org\/10.1109\/ICASSP.2002.5743771.","DOI":"10.1109\/ICASSP.2002.5743771"},{"issue":"4","key":"201_CR2","doi-asserted-by":"publisher","first-page":"320","DOI":"10.1109\/TASSP.1976.1162830","volume":"24","author":"C. Knapp","year":"1976","unstructured":"C. Knapp, G. Carter, The generalized correlation method for estimation of time delay. IEEE Trans. Acoust. Speech Signal Proc. (TASSP). 24(4), 320\u2013327 (1976).","journal-title":"IEEE Trans. Acoust. Speech Signal Proc. (TASSP)"},{"issue":"2","key":"201_CR3","doi-asserted-by":"publisher","first-page":"525","DOI":"10.1109\/78.193195","volume":"41","author":"G. Jacovitti","year":"1993","unstructured":"G. Jacovitti, G. Scarano, Discrete time techniques for time delay estimation. IEEE Trans. Signal Proc. (TSP). 41(2), 525\u2013533 (1993).","journal-title":"IEEE Trans. Signal Proc. (TSP)"},{"issue":"1","key":"201_CR4","doi-asserted-by":"publisher","first-page":"384","DOI":"10.1121\/1.428310","volume":"107","author":"J. Benesty","year":"2000","unstructured":"J. Benesty, Adaptive eigenvalue decomposition algorithm for passive acoustic source localization. J. Acoust. Soc. Am.107(1), 384\u2013391 (2000).","journal-title":"J. Acoust. Soc. Am."},{"key":"201_CR5","doi-asserted-by":"publisher","first-page":"561","DOI":"10.1109\/LSP.2005.849546","volume":"12","author":"F. Talantzis","year":"2005","unstructured":"F. Talantzis, A. G. Constantinides, L. C. Polymenakos, Estimation of direction of arrival using information theory. IEEE Signal Proc. Lett.12:, 561\u2013564 (2005).","journal-title":"IEEE Signal Proc. Lett."},{"key":"201_CR6","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1007\/978-3-662-04619-7_8","volume-title":"Microphone arrays: signal processing techniques and applications","author":"J. DiBiase","year":"2001","unstructured":"J. DiBiase, H. F. Silverman, M. S. Brandstein, in Microphone arrays: signal processing techniques and applications, ed. by M. Brandstein, D. Ward. Robust localization in reverberant rooms (SpringerNew York, 2001), pp. 157\u2013180."},{"key":"201_CR7","unstructured":"C. Zhang, D. Florencio, Z. Zhang, in 2008 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Why does PHAT work well in lownoise, reverberative environments? (2008), pp. 2565\u20132568."},{"key":"201_CR8","unstructured":"J. Valin, F. Michaud, J. Rouat, D. Letourneau, in Proceedings 2003 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS 2003), 2. Robust sound source localization using a microphone array on a mobile robot, (2003), pp. 1228\u20131233."},{"key":"201_CR9","unstructured":"Y. Rui, D. Florencio, in 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2. Time delay estimation in the presence of correlated noise and reverberation, (2004), p. 133."},{"key":"201_CR10","unstructured":"H. Kang, M. Graczyk, J. Skoglund, in 2016 IEEE International Workshop on Acoustic Signal Enhancement (IWAENC). On pre-filtering strategies for the GCC-PHAT algorithm, (2016), pp. 1\u20135."},{"issue":"1","key":"201_CR11","doi-asserted-by":"publisher","first-page":"178","DOI":"10.1109\/TASLP.2018.2876169","volume":"27","author":"Z. Wang","year":"2019","unstructured":"Z. Wang, X. Zhang, D. Wang, Robust speaker localization guided by deep learning-based time-frequency masking. IEEE Trans. Audio Speech Lang. Process. (TASLP). 27(1), 178\u2013188 (2019).","journal-title":"IEEE Trans. Audio Speech Lang. Process. (TASLP)"},{"issue":"8","key":"201_CR12","doi-asserted-by":"publisher","first-page":"698","DOI":"10.1016\/j.apacoust.2012.02.002","volume":"73","author":"J. M. Perez-Lorenzo","year":"2012","unstructured":"J. M. Perez-Lorenzo, R. Viciana-Abad, P. Reche-Lopez, F. Rivas, J. Escolano, Evaluation of generalized cross-correlation methods for direction of arrival estimation using two microphones in real environments. Appl. Acoust.73(8), 698\u2013712 (2012).","journal-title":"Appl. Acoust."},{"key":"201_CR13","unstructured":"B. Loesch, B. Yang, in IEEE International Workshop on Acoustic Signal Enhancement (IWAENC). Source number estimation and clustering for underdetermined blind source separation, (2008), pp. 1\u20134."},{"key":"201_CR14","unstructured":"M. I. Mandel, D. P. W. Ellis, T. Jebara, in Proceedings of the Annual Conference on Neural Information Processing Systems. An em algorithm for localizing multiple sound: sources in reverberant environments, (2006), pp. 953\u2013960."},{"issue":"2","key":"201_CR15","doi-asserted-by":"publisher","first-page":"392","DOI":"10.1109\/TASLP.2013.2292361","volume":"22","author":"O. Schwartz","year":"2014","unstructured":"O. Schwartz, S. Gannot, Speaker tracking using recursive EM algorithms. IEEE Trans. Audio Speech Lang. Process. (TASLP). 22(2), 392\u2013402 (2014).","journal-title":"IEEE Trans. Audio Speech Lang. Process. (TASLP)"},{"key":"201_CR16","unstructured":"N. Madhu, R. Martin, in IEEE International Workshop on Acoustic Signal Enhancement (IWAENC). A scalable framework for multiple speaker localization and tracking, (2008), pp. 1\u20134."},{"issue":"1","key":"201_CR17","doi-asserted-by":"publisher","first-page":"66","DOI":"10.1016\/j.dsp.2010.04.003","volume":"21","author":"M. Cobos","year":"2011","unstructured":"M. Cobos, J. J. Lopez, D. Martinez, Two-microphone multi-speaker localization based on a Laplacian mixture model. Digit. Signal Process.21(1), 66\u201376 (2011).","journal-title":"Digit. Signal Process."},{"issue":"8","key":"201_CR18","doi-asserted-by":"publisher","first-page":"1781","DOI":"10.1016\/j.sigpro.2011.02.002","volume":"91","author":"M. Swartling","year":"2011","unstructured":"M. Swartling, B. S\u00e4llberg, N. Grbi\u0107, Source localization for multiple speech sources using low complexity non-parametric source separation and clustering. Signal Process.91(8), 1781\u20131788 (2011).","journal-title":"Signal Process."},{"issue":"8","key":"201_CR19","doi-asserted-by":"publisher","first-page":"1950","DOI":"10.1016\/j.sigpro.2011.09.032","volume":"92","author":"C. Blandin","year":"2012","unstructured":"C. Blandin, A. Ozerov, E. Vincent, Multi-source TDOA estimation in reverberant audio using angular spectra and clustering. Signal Process.92(8), 1950\u20131960 (2012).","journal-title":"Signal Process."},{"issue":"3","key":"201_CR20","doi-asserted-by":"publisher","first-page":"683","DOI":"10.1016\/j.csl.2012.08.003","volume":"27","author":"P. Pertil\u00e4","year":"2013","unstructured":"P. Pertil\u00e4, Online blind speech separation using multiple acoustic speaker tracking and time-frequency masking. Comput. Speech Lang.27(3), 683\u2013702 (2013).","journal-title":"Comput. Speech Lang."},{"key":"201_CR21","unstructured":"E. Hadad, S. Gannot, in 2018 IEEE International Conference on the Science of Electrical Engineering in Israel (ICSEE). Multi-speaker direction of arrival estimation using SRP-PHAT algorithm with a weighted histogram, (2018), pp. 1\u20135."},{"key":"201_CR22","doi-asserted-by":"publisher","first-page":"135","DOI":"10.1002\/9780470727188.ch6","volume-title":"Advances in digital speech transmission","author":"N. Madhu","year":"2008","unstructured":"N. Madhu, R. Martin, in Advances in digital speech transmission, ed. by R. Martin, U. Heute, and C. Antweiler. Acoustic source localization with microphone arrays (John Wiley & Sons, Ltd.New York, USA, 2008), pp. 135\u2013170."},{"key":"201_CR23","unstructured":"D. Bechler, K. Kroschel, in IEEE International Workshop on Acoustic Signal Enhancement (IWAENC). Considering the second peak in the GCC function for multi-source TDOA estimation with a microphone array, (2003), pp. 315\u2013318."},{"issue":"5","key":"201_CR24","doi-asserted-by":"publisher","first-page":"3075","DOI":"10.1121\/1.1791872","volume":"116","author":"C. Faller","year":"2004","unstructured":"C. Faller, J. Merimaa, Source localization in complex listening situations: selection of binaural cues based on interaural coherence. J. Acoust. Soc. Am.116(5), 3075\u20133089 (2004).","journal-title":"J. Acoust. Soc. Am."},{"key":"201_CR25","unstructured":"M. Togami, T. Sumiyoshi, A. Amano, in 2007 IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP), 1. Stepwise phase difference restoration method for sound source localization using multiple microphone pairs, (2007), pp. 117\u2013120."},{"key":"201_CR26","unstructured":"M. Togami, A. Amano, T. Sumiyoshi, Y. Obuchi, in 2009 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). DOA estimation method based on sparseness of speech sources for human symbiotic robots, (2009), pp. 3693\u20133696."},{"key":"201_CR27","unstructured":"J. Traa, P. Smaragdis, N. D. Stein, D. Wingate, in 2015 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics. Directional NMF for joint source localization and separation, (2015), pp. 1\u20135."},{"key":"201_CR28","unstructured":"H. Kayser, J. Anem\u00fcller, K. Adilo\u011flu, in 2014 IEEE 8th Sensor Array and Multichannel Signal Processing Workshop (SAM). Estimation of inter-channel phase differences using non-negative matrix factorization, (2014), pp. 77\u201380."},{"key":"201_CR29","unstructured":"A. Mu\u00f1oz-Montoro, V. Montiel-Zafra, J. Carabias-Orti, J. Torre-Cruz, F. Canadas-Quesada, P. Vera-Candeas, in Proceedings of the International Congress on Acoustics (ICA). Source localization using a spatial kernel based covariance model and supervised complex nonnegative matrix factorization, (2019), pp. 3321\u20133328."},{"issue":"4","key":"201_CR30","doi-asserted-by":"publisher","first-page":"745","DOI":"10.1109\/TASLP.2017.2656805","volume":"25","author":"S. U. N. Wood","year":"2017","unstructured":"S. U. N. Wood, J. Rouat, S. Dupont, G. Pironkov, Blind speech separation and enhancement with GCC-NMF. IEEE Trans. Audio Speech Lang. Process. (TASLP). 25(4), 745\u2013755 (2017).","journal-title":"IEEE Trans. Audio Speech Lang. Process. (TASLP)"},{"key":"201_CR31","volume-title":"A high-accuracy, low-latency technique for talker localization in reverberant environments. Ph.D. dissertation","author":"J. DiBiase","year":"2000","unstructured":"J. DiBiase, A high-accuracy, low-latency technique for talker localization in reverberant environments. Ph.D. dissertation (Brown University, Providence RI, USA, 2000)."},{"key":"201_CR32","doi-asserted-by":"publisher","first-page":"125","DOI":"10.1109\/MSP.2013.2288990","volume":"32","author":"T. Virtanen","year":"2015","unstructured":"T. Virtanen, J. F. Gemmeke, B. Raj, P. Smaragdis, Compositional models for audio processing: uncovering the structure of sound mixtures. IEEE Signal Process. Mag.32:, 125\u2013144 (2015).","journal-title":"IEEE Signal Process. Mag."},{"issue":"3","key":"201_CR33","doi-asserted-by":"publisher","first-page":"184","DOI":"10.1109\/TSA.2003.811542","volume":"11","author":"J. Tchorz","year":"2003","unstructured":"J. Tchorz, B. Kollmeier, SNR estimation based on amplitude modulation analysis with applications to noise suppression. IEEE Trans. Speech Audio Process. (TSAP). 11(3), 184\u2013192 (2003).","journal-title":"IEEE Trans. Speech Audio Process. (TSAP)"},{"issue":"8","key":"201_CR34","doi-asserted-by":"publisher","first-page":"1592","DOI":"10.1109\/TASLP.2017.2702385","volume":"25","author":"S. Elshamy","year":"2017","unstructured":"S. Elshamy, N. Madhu, W. Tirry, T. Fingscheidt, Instantaneous a priori SNR estimation by cepstral excitation manipulation. IEEE Trans. Audio Speech Lang. Process. (TASLP). 25(8), 1592\u20131605 (2017).","journal-title":"IEEE Trans. Audio Speech Lang. Process. (TASLP)"},{"key":"201_CR35","unstructured":"D. D. Lee, H. S. Seung, in Advances in Neural Information Processing Systems 13, ed. by T. K. Leen, T. G. Dietterich, and V. Tresp. Algorithms for non-negative matrix factorization, (2001), pp. 556\u2013562."},{"issue":"1","key":"201_CR36","doi-asserted-by":"publisher","first-page":"29","DOI":"10.1016\/j.laa.2005.06.025","volume":"416","author":"V. P. Pauca","year":"2006","unstructured":"V. P. Pauca, J. Piper, R. J. Plemmons, Nonnegative matrix factorization for spectral data analysis. Linear Algebra Appl.416(1), 29\u201347 (2006). Special Issue devoted to the Haifa 2005 conference on matrix theory.","journal-title":"Linear Algebra Appl."},{"key":"201_CR37","unstructured":"R. Lebarbenchon, E. Camberlein, Multi-Channel BSS Locate (2018). https:\/\/bass-db.gforge.inria.fr\/bss-locate\/bss-locate. Accessed 4 2020."},{"key":"201_CR38","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1007\/978-3-642-15995-4_6","volume-title":"9th International Conference on Latent variable analysis and signal separation (LVA\/ICA)","author":"B. Loesch","year":"2010","unstructured":"B. Loesch, B. Yang, in 9th International Conference on Latent variable analysis and signal separation (LVA\/ICA). Adaptive segmentation and separation of determined convolutive mixtures under dynamic conditions (SpringerBerlin, Heidelberg, 2010), pp. 41\u201348."},{"key":"201_CR39","unstructured":"N. Ono, Z. Koldovsk\u00fd, S. Miyabe, N. Ito, in 2013 IEEE International Workshop on Machine Learning for Signal Processing (MLSP). The 2013 signal separation evaluation campaign, (2013), pp. 1\u20136."},{"key":"201_CR40","unstructured":"H. W. L\u00f6llmann, C. Evers, A. Schmidt, H. Mellmann, H. Barfuss, P. A. Naylor, W. Kellermann, in 2018 IEEE 10th Sensor Array and Multichannel Signal Processing Workshop (SAM). The LOCATA challenge data corpus for acoustic source localization and tracking, (2018), pp. 410\u2013414."},{"key":"201_CR41","doi-asserted-by":"publisher","unstructured":"C. Veaux, J. Yamagishi, K. MacDonald, English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (University of Edinburgh. The Centre for Speech Technology Research (CSTR), 2019). https:\/\/doi.org\/10.7488\/ds\/2645.","DOI":"10.7488\/ds\/2645"},{"key":"201_CR42","unstructured":"Multi-channel impulse response database. https:\/\/www.iks.rwth-aachen.de\/en\/research\/tools-downloads\/databases\/multi-channel-impulse-response-database\/. Accessed 12 2020."},{"key":"201_CR43","volume-title":"TSP speech database. Technical report","author":"P. Kabal","year":"2002","unstructured":"P. Kabal, TSP speech database. Technical report (Telecommunications and Signal Processing Laboratory, McGill University, Canada, 2002)."},{"key":"201_CR44","first-page":"1","volume-title":"Audio Source Separation","author":"C. F\u00e9votte","year":"2018","unstructured":"C. F\u00e9votte, E. Vincent, A. Ozerov, in Audio Source Separation, ed. by S. Makino. Single-channel audio source separation with NMF: divergences, constraints and algorithms (SpringerCham, 2018), pp. 1\u201324."},{"key":"201_CR45","unstructured":"T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, J. Dean, in Advances in neural information processing systems 26, ed. by C. J. C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Q. Weinberger. Distributed representations of words and phrases and their compositionality, (2013), pp. 3111\u20133119."}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-021-00201-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/s13636-021-00201-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-021-00201-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,3,3]],"date-time":"2021-03-03T21:12:24Z","timestamp":1614805944000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-021-00201-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,3]]},"references-count":45,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["201"],"URL":"https:\/\/doi.org\/10.1186\/s13636-021-00201-y","relation":{},"ISSN":["1687-4722"],"issn-type":[{"type":"electronic","value":"1687-4722"}],"subject":[],"published":{"date-parts":[[2021,3,3]]},"assertion":[{"value":"26 August 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 February 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 March 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors state that they have no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"13"}}