{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,3,31]],"date-time":"2022-03-31T11:40:34Z","timestamp":1648726834510},"reference-count":42,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2014,4,15]],"date-time":"2014-04-15T00:00:00Z","timestamp":1397520000000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/2.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"published-print":{"date-parts":[[2014,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Previously, a dereverberation method based on generalized spectral subtraction (GSS) using multi-channel least mean-squares (MCLMS) has been proposed. The results of speech recognition experiments showed that this method achieved a significant improvement over conventional methods. In this paper, we apply this method to distant-talking (far-field) speaker recognition. However, for far-field speech, the GSS-based dereverberation method using clean speech models degrades the speaker recognition performance. This may be because GSS-based dereverberation causes some distortion between clean speech and dereverberant speech. In this paper, we address this problem by training speaker models using dereverberant speech obtained by suppressing reverberation from arbitrary artificial reverberant speech. Furthermore, we propose an efficient computational method for a combination of the likelihood of dereverberant speech using multiple compensation parameter sets. This addresses the problem of determining optimal compensation parameters for GSS. We report the results of a speaker recognition experiment performed on large-scale far-field speech with different reverberant environments to the training environments. The proposed GSS-based dereverberation method achieves a recognition rate of 92.2%, which compares well with conventional cepstral mean normalization with delay-and-sum beamforming using a clean speech model (49.0%) and a reverberant speech model (88.4%). We also compare the proposed method with another dereverberation technique, multi-step linear prediction-based spectral subtraction (MSLP-GSS). The proposed method achieves a better recognition rate than the 90.6% of MSLP-GSS. The use of multiple compensation parameters further improves the speech recognition performance, giving our approach a recognition rate of 93.6%. We implement this method in a real environment using the optimal compensation parameters estimated from an artificial environment. The results show a recognition rate of 87.8% compared with 72.5% for delay-and-sum beamforming using a reverberant speech model.<\/jats:p>","DOI":"10.1186\/1687-4722-2014-15","type":"journal-article","created":{"date-parts":[[2014,4,15]],"date-time":"2014-04-15T09:02:40Z","timestamp":1397552560000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Distant-talking speaker identification by generalized spectral subtraction-based dereverberation and its efficient computation"],"prefix":"10.1186","volume":"2014","author":[{"given":"Zhaofeng","family":"Zhang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Longbiao","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Atsuhiko","family":"Kai","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2014,4,15]]},"reference":[{"key":"111_CR1","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-540-37631-6","volume-title":"Acoustic MIMO Signal Processing","author":"Y Huang","year":"2006","unstructured":"Huang Y, Benesty J, Chen J: Acoustic MIMO Signal Processing. Berlin: Springer-Verlag; 2006."},{"key":"111_CR2","doi-asserted-by":"crossref","first-page":"570","DOI":"10.21437\/Interspeech.2010-225","volume-title":"Proceedings of INTERSPEECH-2010","author":"H Maganti","year":"2010","unstructured":"Maganti H, Matassoni M: An auditory based modulation spectral feature for reverberant speech recognition. In Proceedings of INTERSPEECH-2010. Makuhari, Chiba, 26-30 September, Curran Associates, Inc., Red Hook, NY; 2010:570-573."},{"key":"111_CR3","first-page":"1133","volume-title":"Proceedings of the 2006 ICASSP Toulouse, France, 14-19","author":"C Raut","year":"2006","unstructured":"Raut C, Nishimoto T, Sagayama S: Adaptation for long convolutional distortion by maximum likelihood based state filtering approach. In Proceedings of the 2006 ICASSP Toulouse, France, 14-19 May 2006 vol. 1. IEEE, Piscataway, 2006; 1133-1136."},{"issue":"6","key":"111_CR4","doi-asserted-by":"publisher","first-page":"114","DOI":"10.1109\/MSP.2012.2205029","volume":"29","author":"T Yoshioka","year":"2012","unstructured":"Yoshioka T, Sehr A, Delcroix M, Kinoshita K, Maas R, Nakatani T, Kellermann W: Making machines understand us in reverberant rooms: robustness against reverberation for automatic speech recognition. IEEE Signal Process. Mag 2012, 29(6):114-126.","journal-title":"IEEE Signal Process. Mag"},{"issue":"3","key":"111_CR5","doi-asserted-by":"publisher","first-page":"346","DOI":"10.1109\/89.759045","volume":"7","author":"TB Hughes","year":"1999","unstructured":"Hughes TB, Kim HS, DiBiase JH, Silverman HF: Performance of an an HMM speech recognizer using a real-time tracking microphone array as input. IEEE Trans. Speech Audio Process 1999, 7(3):346-349. 10.1109\/89.759045","journal-title":"IEEE Trans. Speech Audio Process"},{"issue":"2","key":"111_CR6","doi-asserted-by":"publisher","first-page":"254","DOI":"10.1109\/TASSP.1981.1163530","volume":"29","author":"S Furui","year":"1981","unstructured":"Furui S: Cepstral analysis technique for automatic speaker verification. IEEE Trans. Acoust. Speech Signal Process 1981, 29(2):254-272. 10.1109\/TASSP.1981.1163530","journal-title":"IEEE Trans. Acoust. Speech Signal Process"},{"key":"111_CR7","doi-asserted-by":"publisher","first-page":"69","DOI":"10.3115\/1075671.1075688","volume-title":"Proceedings of the workshop on Human Language Technology Princeton","author":"F Liu","year":"1993","unstructured":"Liu F, Stern R, Huang X, Acero A: Efficient cepstral normalization for robust speech recognition. Proceedings of the workshop on Human Language Technology Princeton, 69\u201374 (Association for Computational Linguistics, Stroudsburg, 1993)"},{"key":"111_CR8","first-page":"359","volume":"87","author":"K Lebart","year":"2001","unstructured":"Lebart K, Boucher J, Denbigh P: A new method based on spectral subtraction for speech dereverberation. Acta Acustica 2001, 87: 359-366.","journal-title":"Acta Acustica"},{"key":"111_CR9","first-page":"968","volume-title":"Double the trouble: handling noise and reverberation in far-field automatic speech recognition","author":"D Gelbart","year":"2002","unstructured":"Gelbart D, Morgan N: Double the trouble: handling noise and reverberation in far-field automatic speech recognition. In INTERSPEECH 2002. Denver, 16-20 September, 2002; 968-971."},{"key":"111_CR10","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU.2001.1034598","volume-title":"Evaluating long-term spectral subtraction for reverberant ASR","author":"D Gelbart","year":"2001","unstructured":"Gelbart D, Morgan N: Evaluating long-term spectral subtraction for reverberant ASR. In ASRU 2001. Madonna di Campiglio, Italy, 9-13 December 2001;"},{"issue":"3","key":"111_CR11","first-page":"774","volume":"14","author":"M Wu","year":"2006","unstructured":"Wu M, Wang D: A two-stage algorithm for one-microphone reverberant speech enhancement. IEEE Trans. ASLP 2006, 14(3):774-784.","journal-title":"IEEE Trans. ASLP"},{"key":"111_CR12","first-page":"173","volume-title":"Proceedings of IEEE ICASSP","author":"EA Habets","year":"2005","unstructured":"Habets EA: Multi-channel speech dereverberation based on a statistical model of late reverberation. In Proceedings of IEEE ICASSP. Philadelphia, 18-23 March vol. 4, IEEE, Piscataway; 2005:173-176."},{"key":"111_CR13","first-page":"5448","volume-title":"Hilbert envelope based features for robust speaker identification under reverberant mismatched conditions","author":"SO Sadjadi","year":"2011","unstructured":"Sadjadi SO, Hasnen JHL: Hilbert envelope based features for robust speaker identification under reverberant mismatched conditions. In Proceedings of IEEE ICASSP. Prague, Czech Republic, 22-27 May 2011; 5448-5451."},{"issue":"1","key":"111_CR14","doi-asserted-by":"publisher","first-page":"1074","DOI":"10.1155\/S1110865703305049","volume":"2003","author":"S Gannot","year":"2003","unstructured":"Gannot S, Moonen M: Subspace methods for multimicrophone speech dereverberation. EURASIP J. Appl. Signal Processv 2003, 2003(1):1074-1090.","journal-title":"EURASIP J. Appl. Signal Processv"},{"key":"111_CR15","first-page":"817","volume-title":"Spectral subtraction steered by multi-step forward linear prediction for single channel speech dereverberation","author":"K Kinoshita","year":"2006","unstructured":"Kinoshita K, Delcroix M, Nakatani T, Miyoshi M: Spectral subtraction steered by multi-step forward linear prediction for single channel speech dereverberation. In Proceedings of IEEE ICASSP 2006. Toulouse, France, 14-19 May 2006; 817-820."},{"issue":"2","key":"111_CR16","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1109\/TASSP.1979.1163209","volume":"27","author":"S Boll","year":"1979","unstructured":"Boll S: Suppression of acoustic noise in speech using spectral subtraction. IEEE Trans. Acoustics Speech Signal Process 1979, 27(2):113-120. 10.1109\/TASSP.1979.1163209","journal-title":"IEEE Trans. Acoustics Speech Signal Process"},{"issue":"2","key":"111_CR17","first-page":"430","volume":"15","author":"M Delcroix","year":"2007","unstructured":"Delcroix M, Hikichi T, Miyoshi M: Precise dereverberation using multi-channel linear prediction. IEEE Trans. ASLP 2007, 15(2):430-440.","journal-title":"IEEE Trans. ASLP"},{"issue":"7","key":"111_CR18","first-page":"2023","volume":"15","author":"Q Jin","year":"2007","unstructured":"Jin Q, Schultz T: A Waibel, Far-field speaker recognition. IEEE Trans. ASLP 2007, 15(7):2023-2032.","journal-title":"IEEE Trans. ASLP"},{"issue":"5","key":"111_CR19","doi-asserted-by":"publisher","first-page":"392","DOI":"10.1109\/89.536934","volume":"4","author":"S Subramaniam","year":"1996","unstructured":"Subramaniam S, Petropulu AP, Wendt C: Cepstrum-based deconvolution for speech dereverberation. IEEE Trans. Speech Audio Process 1996, 4(5):392-396. 10.1109\/89.536934","journal-title":"IEEE Trans. Speech Audio Process"},{"issue":"4","key":"111_CR20","doi-asserted-by":"publisher","first-page":"534","DOI":"10.1109\/TASL.2008.2009015","volume":"17","author":"K Kinoshita","year":"2009","unstructured":"Kinoshita K, Delcroix M, Nakatani T, Miyoshi M: Suppression of late reverberation effect on speech signal using long-term multiple-step linear prediction. IEEE Trans. Audio Speech Lang. Process 2009, 17(4):534-545.","journal-title":"IEEE Trans. Audio Speech Lang. Process"},{"key":"111_CR21","first-page":"937","volume-title":"Proceedings ICASSP 2006","author":"Q Jin","year":"2006","unstructured":"Jin Q, Pan Y, Schultz T: Far-field speaker recognition. In Proceedings ICASSP 2006. Toulouse, France, 14-19 May vol. 1 IEEE, Piscataway; 2006:937-940."},{"key":"111_CR22","doi-asserted-by":"crossref","unstructured":"Wang L, Odani K, Kai A: Dereverberation and denoising based on generalized spectral subtraction by nutil-channel LMS algorithm using a small-scale microphone array. Eurasip J. Adv. Signal Process 2012., 2012(12):","DOI":"10.1186\/1687-6180-2012-12"},{"issue":"4","key":"111_CR23","doi-asserted-by":"publisher","first-page":"328","DOI":"10.1109\/89.701361","volume":"6","author":"BL Sim","year":"1998","unstructured":"Sim BL, Tong YC, Chang JS, Tan CT: A parametric formulation of the generalized spectral subtraction method. IEEE Trans. Speech Audio Process 1998, 6(4):328-337. 10.1109\/89.701361","journal-title":"IEEE Trans. Speech Audio Process"},{"issue":"6","key":"111_CR24","doi-asserted-by":"publisher","first-page":"1770","DOI":"10.1109\/TASL.2010.2098871","volume":"19","author":"T Inoue","year":"2011","unstructured":"Inoue T, Saruwatari H, Takahashi Y, Shikano K, Kondo K: Theoretical analysis of musical noise in generalized spectral subtraction based on higher-order statistics. IEEE Trans. Audio Speech Lang. Process 2011, 19(6):1770-1779.","journal-title":"IEEE Trans. Audio Speech Lang. Process"},{"key":"111_CR25","doi-asserted-by":"crossref","first-page":"1032","DOI":"10.21437\/Interspeech.2008-299","volume-title":"Proceedings of InterSpeech 2008","author":"L Wang","year":"2008","unstructured":"Wang L, Nakagawa S, Kitaoka N: Blind dereverberation based on CMN and spectral subtraction by multi-channel LMS algorithm. In Proceedings of InterSpeech 2008. Brisbane, 22-26; September 2008:1032-1035."},{"key":"111_CR26","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1016\/0167-6393(95)00009-D","volume":"17","author":"DA Reynolds","year":"1995","unstructured":"Reynolds DA: Speaker identification and verification using Gaussian mixture speaker models. Speech Commun 1995, 17: 91-108. 10.1016\/0167-6393(95)00009-D","journal-title":"Speech Commun"},{"issue":"1-3","key":"111_CR27","doi-asserted-by":"publisher","first-page":"19","DOI":"10.1006\/dspr.1999.0361","volume":"10","author":"DA Reynolds","year":"2000","unstructured":"Reynolds DA, Quatieri TF, Dunn R: Speaker verification using adapted Gaussian mixture models. Dig. Signal Process 2000, 10(1-3):19-41. 10.1006\/dspr.1999.0361","journal-title":"Dig. Signal Process"},{"issue":"6","key":"111_CR28","doi-asserted-by":"publisher","first-page":"501","DOI":"10.1016\/j.specom.2007.04.004","volume":"49","author":"L Wang","year":"2007","unstructured":"Wang L, Kitaoka N, Nakagawa S: Robust distant speaker recognition based on position-dependent CMN by combining speaker-specific GMM with speaker-adapted HMM. Speech Commun 2007, 49(6):501-513. 10.1016\/j.specom.2007.04.004","journal-title":"Speech Commun"},{"issue":"1","key":"111_CR29","doi-asserted-by":"publisher","first-page":"194","DOI":"10.1109\/89.260362","volume":"2","author":"K Farrell","year":"1994","unstructured":"Farrell K, Mammone R, Assaleh K: Speaker recognition using neural networks and conventional classifiers. IEEE Trans. on Speech Audio Process 1994, 2(1):194-205. 10.1109\/89.260362","journal-title":"IEEE Trans. on Speech Audio Process"},{"issue":"2\u20133","key":"111_CR30","doi-asserted-by":"publisher","first-page":"210","DOI":"10.1016\/j.csl.2005.06.003","volume":"20","author":"W Campbell","year":"2006","unstructured":"Campbell W, Campbell J, Reynolds D, Singer E, Torres-Carrasquillo P: Support vector machines for speaker and language recognition. Comput. Speech Lang 2006, 20(2\u20133):210-229.","journal-title":"Comput. Speech Lang"},{"issue":"7","key":"111_CR31","doi-asserted-by":"publisher","first-page":"980","DOI":"10.1109\/TASL.2008.925147","volume":"15","author":"P Kenny","year":"2008","unstructured":"Kenny P, Ouellet P, Dehak N, Gupta V, Dumouchel P: A study of inter-speaker variability in speaker verification. IEEE Trans. Audio Speech Lang. Process 2008, 15(7):980-988.","journal-title":"IEEE Trans. Audio Speech Lang. Process"},{"issue":"4","key":"111_CR32","doi-asserted-by":"publisher","first-page":"788","DOI":"10.1109\/TASL.2010.2064307","volume":"19","author":"N Dehak","year":"2011","unstructured":"Dehak N, Kenny P, Dehak R, Dumouchel P, Ouellet P: Front-end factor analysis for speaker verification. IEEE Trans. Audio Speech Lang. Process 2011, 19(4):788-798.","journal-title":"IEEE Trans. Audio Speech Lang. Process"},{"key":"111_CR33","first-page":"1259","volume-title":"Proceedings of IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP)","author":"B Kingsbury","year":"1997","unstructured":"Kingsbury B, Morgan N: Recognizing reverberant speech with RASTA-PLP. In Proceedings of IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP). Munich, 21-24 April vol.2 IEEE, Piscataway; 1997:1259-1262."},{"issue":"5","key":"111_CR34","doi-asserted-by":"publisher","first-page":"3261","DOI":"10.1121\/1.411005","volume":"96","author":"AC Surendran","year":"1994","unstructured":"Surendran AC, Flanagan JL: Stable dereverberation using microphone arrays for speaker verification. J. Acoust. Soc. Am 1994, 96(5):3261-3262.","journal-title":"J. Acoust. Soc. Am"},{"key":"111_CR35","first-page":"1637","volume-title":"Adaptive blind channel identification: multi-channel least mean square and Newton algorithms","author":"Y Huang","year":"2002","unstructured":"Huang Y, Benesty J: Adaptive blind channel identification: multi-channel least mean square and Newton algorithms. In ICASSP Orlando, 13-17 May vol. 2. IEEE, Piscataway, 2002; 1637\u20131640"},{"key":"111_CR36","doi-asserted-by":"publisher","first-page":"1127","DOI":"10.1016\/S0165-1684(02)00247-5","volume":"82","author":"Y Huang","year":"2002","unstructured":"Huang Y, Benesty J: Adaptive multichannel least mean square and Newton algorithms for blind channel identification. Signal Process 2002, 82: 1127-1138. 10.1016\/S0165-1684(02)00247-5","journal-title":"Signal Process"},{"issue":"3","key":"111_CR37","doi-asserted-by":"publisher","first-page":"173","DOI":"10.1109\/LSP.2004.842286","volume":"12","author":"Y Huang","year":"2005","unstructured":"Huang Y, Benesty J, Chen J: Optimal step size of the adaptive multi-channel LMS algorithm for blind SIMO identification. IEEE Signal Process. Lett 2005, 12(3):173-175.","journal-title":"IEEE Signal Process. Lett"},{"issue":"3","key":"111_CR38","doi-asserted-by":"publisher","first-page":"659","DOI":"10.1587\/transinf.E94.D.659","volume":"E94-D","author":"L Wang","year":"2011","unstructured":"Wang L, Kitaoka N, Nakagawa S: Distant-talking speech recognition based on spectral subtraction by multi-channel LMS algorithm. IEICE Trans. Inf. Syst. 2011, E94-D(3):659-667. 10.1587\/transinf.E94.D.659","journal-title":"IEICE Trans. Inf. Syst"},{"key":"111_CR39","first-page":"965","volume-title":"Acoustical sound database in real environments for sound scene understanding and hands-free speech recognition","author":"S Nakamura","year":"2000","unstructured":"Nakamura S, Hiyane K, Asano F, Nishiura T, Yamada T: Acoustical sound database in real environments for sound scene understanding and hands-free speech recognition. In Proceedings of LREC 2000. Athens, Greece, 31 May - 2 June 2000; 965-968."},{"key":"111_CR40","first-page":"968","volume-title":"Evaluation framework for distant-talking speech recognition under reverberant environments","author":"T Nishiura","year":"2008","unstructured":"Nishiura T, Nakayama M, Denda Y, Kitaoka N, Yamamoto K, Yamada T, Tsuge S, Miyajima C, Fujimoto M, Takiguchi T, Tamura S, Kuroiwa S, Takeda K, Nakamura S: Evaluation framework for distant-talking speech recognition under reverberant environments. In Proceedings of INTERSPEECH 2008. Brisbane, Australia, 22-26 September 2008; 968-971."},{"issue":"3","key":"111_CR41","doi-asserted-by":"publisher","first-page":"199","DOI":"10.1250\/ast.20.199","volume":"20","author":"K Itou","year":"1999","unstructured":"Itou K, Takeda K, Kakezawa T, Matsuoka T, Kobayashi T, Shikano K, Itahashi S, M Yamamoto: Janpanese speech corpus for large vocabulary continuous speech recognition research. J. Acoust. Soc. Jpn. (E) 1999, 20(3):199-206. 10.1250\/ast.20.199","journal-title":"J. Acoust. Soc. Jpn. (E)"},{"key":"111_CR42","doi-asserted-by":"crossref","unstructured":"Patrick A Naylor: Signal-based performance evaluation of dereverberation algorithms. J. Electrical Comput. Eng 2010., 2010(5): Article ID 127513. doi:10.1155\/2010\/127513","DOI":"10.1155\/2010\/127513"}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/1687-4722-2014-15.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/1687-4722-2014-15\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1687-4722-2014-15.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,2]],"date-time":"2021-09-02T11:08:25Z","timestamp":1630580905000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/1687-4722-2014-15"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,4,15]]},"references-count":42,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2014,12]]}},"alternative-id":["111"],"URL":"https:\/\/doi.org\/10.1186\/1687-4722-2014-15","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2014,4,15]]},"assertion":[{"value":"4 July 2013","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 December 2013","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 April 2014","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"15"}}