{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,4,4]],"date-time":"2022-04-04T16:46:17Z","timestamp":1649090777691},"reference-count":15,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2011,11,15]],"date-time":"2011-11-15T00:00:00Z","timestamp":1321315200000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/2.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"published-print":{"date-parts":[[2011,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>In this article, a novel technique based on the empirical mode decomposition methodology for processing speech features is proposed and investigated. The empirical mode decomposition generalizes the Fourier analysis. It decomposes a signal as the sum of intrinsic mode functions. In this study, we implement an iterative algorithm to find the intrinsic mode functions for any given signal. We design a novel speech feature post-processing method based on the extracted intrinsic mode functions to achieve noise-robustness for automatic speech recognition. Evaluation results on the noisy-digit Aurora 2.0 database show that our method leads to significant performance improvement. The relative improvement over the baseline features increases from 24.0 to 41.1% when the proposed post-processing method is applied on mean-variance normalized speech features. The proposed method also improves over the performance achieved by a very noise-robust frontend when the test speech data are highly mismatched.<\/jats:p>","DOI":"10.1186\/1687-4722-2011-9","type":"journal-article","created":{"date-parts":[[2011,12,6]],"date-time":"2011-12-06T19:23:39Z","timestamp":1323199419000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Noise-robust speech feature processing with empirical mode decomposition"],"prefix":"10.1186","volume":"2011","author":[{"given":"Kuo-Hau","family":"Wu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chia-Ping","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bing-Feng","family":"Yeh","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2011,11,15]]},"reference":[{"issue":"2","key":"30_CR1","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1109\/TASSP.1979.1163209","volume":"27","author":"S Boll","year":"1979","unstructured":"Boll S: Suppression of acoustic noise in speech using spectral subtraction. IEEE Trans Acoust Speech Signal Process 1979,27(2):113-120. 10.1109\/TASSP.1979.1163209","journal-title":"IEEE Trans Acoust Speech Signal Process"},{"key":"30_CR2","first-page":"913","volume-title":"ICASSP","author":"A Berstein","year":"1991","unstructured":"Berstein A, Shallom I: A hypothesized Wiener filtering approach to noisy speech recognition, in. ICASSP 1991, 913-916."},{"key":"30_CR3","first-page":"617","volume-title":"Proceedings of the IEEE International Conference on Signal Processing","author":"W Zhu","year":"2004","unstructured":"Zhu W, O'Shaughnessy D: Incorporating frequency masking filtering in a standard MFCC feature extraction algorithm, in. Proceedings of the IEEE International Conference on Signal Processing 2004, 617-620."},{"issue":"5","key":"30_CR4","doi-asserted-by":"publisher","first-page":"451","DOI":"10.1109\/89.622569","volume":"5","author":"B Strope","year":"1997","unstructured":"Strope B, Alwan A: A model of dynamic auditory perception and its application to robust word recognition. IEEE Trans Speech Audio Process 1997,5(5):451-464. 10.1109\/89.622569","journal-title":"IEEE Trans Speech Audio Process"},{"issue":"2","key":"30_CR5","doi-asserted-by":"publisher","first-page":"254","DOI":"10.1109\/TASSP.1981.1163530","volume":"29","author":"S Furui","year":"1981","unstructured":"Furui S: Cepstral analysis technique for automatic speaker verification. IEEE Trans Acoust Speech Signal Process 1981,29(2):254-272. 10.1109\/TASSP.1981.1163530","journal-title":"IEEE Trans Acoust Speech Signal Process"},{"key":"30_CR6","first-page":"733","volume-title":"Proceedings of the ICASSP","author":"O Viikki","year":"1998","unstructured":"Viikki O, Bye D, Laurila K: A recursive feature vector normalization approach for robust speech recognition in noise, in. Proceedings of the ICASSP 1998, 733-736."},{"issue":"3","key":"30_CR7","doi-asserted-by":"publisher","first-page":"355","DOI":"10.1109\/TSA.2005.845805","volume":"13","author":"A de La Torre","year":"2005","unstructured":"de La Torre A, Peinado A, Segura J, Perez-Cordoba J, Benitez M, Rubio A: Histogram equalization of speech representation for robust speech recognition. IEEE Trans Speech Audio Process 2005,13(3):355-366.","journal-title":"IEEE Trans Speech Audio Process"},{"key":"30_CR8","doi-asserted-by":"publisher","first-page":"903","DOI":"10.1098\/rspa.1998.0193","volume":"454","author":"N Huang","year":"1998","unstructured":"Huang N, Shen Z, Long S, Wu M, Shih H, Zheng Q, Yen N, Tung C, Liu H: The empirical mode decomposition and the Hilbert spectrum for nonlinear and non-stationary time series analysis. Proc R Soc London Ser A Math Phys Eng Sci 1998, 454: 903-995. 10.1098\/rspa.1998.0193","journal-title":"Proc R Soc London Ser A Math Phys Eng Sci"},{"key":"30_CR9","volume-title":"The FPGA implementation of robust speech recognition system by combining genetic algorithm and empirical mode decomposition","author":"XY Li","year":"2009","unstructured":"Li XY: The FPGA implementation of robust speech recognition system by combining genetic algorithm and empirical mode decomposition. Master's thesis, National Kaohsiung University; 2009."},{"issue":"4","key":"30_CR10","doi-asserted-by":"publisher","first-page":"578","DOI":"10.1109\/89.326616","volume":"2","author":"H Hermansky","year":"1994","unstructured":"Hermansky H, Morgan N: RASTA processing of speech. IEEE Trans Speech Audio Process 1994,2(4):578-589. 10.1109\/89.326616","journal-title":"IEEE Trans Speech Audio Process"},{"key":"30_CR11","first-page":"1647","volume-title":"Proceedings of the ICASSP","author":"S Greenberg","year":"1997","unstructured":"Greenberg S, Kingsbury BED: The modulation spectrogram: in pursuit of an invariant representation of speech, in. Proceedings of the ICASSP 1997, 1647-1650."},{"key":"30_CR12","doi-asserted-by":"crossref","first-page":"36","DOI":"10.21437\/Interspeech.2009-7","volume-title":"Proceedings of the INTERSPEECH","author":"H You","year":"2009","unstructured":"You H, Alwan A: Temporal modulation processing of speech signals for noise robust ASR, in. Proceedings of the INTERSPEECH 2009, 36-39."},{"key":"30_CR13","volume-title":"Interpolating Cubic Splines","author":"GD Knoty","year":"1999","unstructured":"Knoty GD: Interpolating Cubic Splines. Birkh\u00e4user, Boston; 1999."},{"key":"30_CR14","volume-title":"ICSA ITRW ASR2000","author":"D Pearce","year":"2000","unstructured":"Pearce D, Hirsch H: The AURORA experimental framework for the performance evaluation of speech recognition systems under noisy conditions, in. ICSA ITRW ASR2000 2000."},{"key":"30_CR15","unstructured":"ETSI Standard ETSI ES 202 050: Speech processing, transmission and quality aspects (STQ); distributed speech recognition; advanced front-end feature extraction algorithm; compression algorithms 2007."}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/1687-4722-2011-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/1687-4722-2011-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1687-4722-2011-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T17:38:16Z","timestamp":1630517896000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/1687-4722-2011-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,11,15]]},"references-count":15,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2011,12]]}},"alternative-id":["30"],"URL":"https:\/\/doi.org\/10.1186\/1687-4722-2011-9","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2011,11,15]]},"assertion":[{"value":"4 May 2011","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 November 2011","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 November 2011","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"9"}}