{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,29]],"date-time":"2025-10-29T03:40:17Z","timestamp":1761709217408},"reference-count":59,"publisher":"Cambridge University Press (CUP)","issue":"3","license":[{"start":{"date-parts":[[2016,4,12]],"date-time":"2016-04-12T00:00:00Z","timestamp":1460419200000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2017,5]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper introduces deep neural network (DNN)\u2013hidden Markov model (HMM)-based methods to tackle speech recognition in heterogeneous groups of speakers including children. We target three speaker groups consisting of children, adult males and adult females. Two different kind of approaches are introduced here: approaches based on DNN adaptation and approaches relying on vocal-tract length normalisation (VTLN). First, the recent approach that consists in adapting a general DNN to domain\/language specific data is extended to target age\/gender groups in the context of DNN\u2013HMM. Then, VTLN is investigated by training a DNN\u2013HMM system by using either mel frequency cepstral coefficients normalised with standard VTLN or mel frequency cepstral coefficients derived acoustic features combined with the posterior probabilities of the VTLN warping factors. In this later, novel, approach the posterior probabilities of the warping factors are obtained with a separate DNN and the decoding can be operated in a single pass when the VTLN approach requires two decoding passes. Finally, the different approaches presented here are combined to take advantage of their complementarity. The combination of several approaches is shown to improve the baseline phone error rate performance by thirty per cent to thirty-five per cent relative and the baseline word error rate performance by about ten per cent relative.<\/jats:p>","DOI":"10.1017\/s135132491600005x","type":"journal-article","created":{"date-parts":[[2016,4,12]],"date-time":"2016-04-12T05:55:06Z","timestamp":1460440506000},"page":"325-350","source":"Crossref","is-referenced-by-count":34,"title":["Deep-neural network approaches for speech recognition with heterogeneous groups of speakers including children"],"prefix":"10.1017","volume":"23","author":[{"given":"ROMAIN","family":"SERIZEL","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"DIEGO","family":"GIULIANI","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2016,4,12]]},"reference":[{"key":"S135132491600005X_ref055","doi-asserted-by":"crossref","unstructured":"Welling L. , Kanthak S. , and Ney H. 1999. Improved methods for vocal tract normalization. In Proceedings of ICASSP, IEEE, New York (NY), United States (ICASSP, ASRU, SLT, IJCNN), vol. 2, 761\u20134.","DOI":"10.1109\/ICASSP.1999.759780"},{"key":"S135132491600005X_ref052","first-page":"6704","volume-title":"Proceedings of ICASSP","author":"Thomas","year":"2013"},{"key":"S135132491600005X_ref050","first-page":"321","volume-title":"Proceedings of ICASSP","author":"Stolcke","year":"2006"},{"key":"S135132491600005X_ref049","first-page":"600","volume-title":"Proceedings of Pattern Recognition, 25th DAGM Symposium","author":"Steidl","year":"2003"},{"key":"S135132491600005X_ref047","doi-asserted-by":"crossref","unstructured":"Serizel R. , and Giuliani D. 2014b. Vocal tract length normalisation approaches to DNN-based children's and adults\u2019 speech recognition. In Proceedings of SLT. IEEE, New York (NY), United States (ICASSP, ASRU, SLT, IJCNN).","DOI":"10.1109\/SLT.2014.7078563"},{"key":"S135132491600005X_ref042","first-page":"55","volume-title":"Proceedings of ASRU","author":"Saon","year":"2013"},{"key":"S135132491600005X_ref054","unstructured":"Wegmann S. , McAllaster D. , Orloff J. , and Peskin B. 1996. Speaker normalisation on conversational telephone speech. In Proceedings of ICASSP, IEEE, New York (NY), United States (ICASSP, ASRU, SLT, IJCNN), vol. 1, 339\u201341."},{"key":"S135132491600005X_ref046","doi-asserted-by":"crossref","unstructured":"Serizel R. , and Giuliani D. 2014a. Deep neural network adaptation for children's and adults\u2019 speech recognition. In Proceedings of CLIC-It. Pisa University Press, Pisa, Italy (CLIC-It).","DOI":"10.12871\/clicit2014166"},{"key":"S135132491600005X_ref045","doi-asserted-by":"crossref","unstructured":"Senior A. , and Lopez-Moreno I. 2014. Improving DNN speaker independence with I-vector inputs. In Proceedings of ICASSP. IEEE, New York (NY), United States (ICASSP, ASRU, SLT, IJCNN).","DOI":"10.1109\/ICASSP.2014.6853591"},{"key":"S135132491600005X_ref043","first-page":"24","volume-title":"Proceedings of ASRU","author":"Seide","year":"2011"},{"key":"S135132491600005X_ref038","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2011.2109382"},{"key":"S135132491600005X_ref037","doi-asserted-by":"crossref","first-page":"3009","DOI":"10.21437\/Interspeech.2005-141","volume-title":"Proceedings of INTERSPEECH","author":"Miguel","year":"2005"},{"key":"S135132491600005X_ref008","first-page":"4334","volume-title":"Proceedings of ICASSP","author":"Burget","year":"2010"},{"key":"S135132491600005X_ref033","doi-asserted-by":"crossref","unstructured":"Li Q. , and Russell M. 2001. Why is automatic recognition of children's speech difficult? In Proceedings of EUROSPEECH. ISCA, Grenoble, France (Interspeech, ICSLP, Workshop on Child, Computer and Interaction).","DOI":"10.21437\/Eurospeech.2001-625"},{"key":"S135132491600005X_ref032","doi-asserted-by":"crossref","first-page":"526","DOI":"10.21437\/Interspeech.2010-214","volume-title":"Proceedings of INTERSPEECH","author":"Li","year":"2010"},{"key":"S135132491600005X_ref030","first-page":"353","volume-title":"Proceedings of ICASSP","author":"Lee","year":"1996"},{"key":"S135132491600005X_ref029","doi-asserted-by":"crossref","unstructured":"Lee C.-H. , and Gauvain J.-L. 1993. Speaker adaptation based on map estimation of hmm parameters. In Proceedings of ICASSP, IEEE, New York (NY), United States (ICASSP, ASRU, SLT, IJCNN), vol. 2, 558\u201361.","DOI":"10.1109\/ICASSP.1993.319368"},{"key":"S135132491600005X_ref027","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-6393(98)00061-2"},{"key":"S135132491600005X_ref022","doi-asserted-by":"crossref","unstructured":"Hagen A. , Pellom B. , and Cole R. 2003. Children's speech recognition with application to interactive books and tutors. In Proceedings of ASRU. IEEE, New York (NY), United States (ICASSP, ASRU, SLT, IJCNN).","DOI":"10.1109\/ASRU.2003.1318426"},{"key":"S135132491600005X_ref056","first-page":"349","volume-title":"Proceedings of ICASSP","author":"Wilpon","year":"1996"},{"key":"S135132491600005X_ref009","doi-asserted-by":"publisher","DOI":"10.1109\/89.725321"},{"key":"S135132491600005X_ref051","first-page":"246","volume-title":"Proceedings of SLT","author":"Swietojanski","year":"2012"},{"key":"S135132491600005X_ref039","doi-asserted-by":"crossref","unstructured":"Nisimura R. , Lee A. , Saruwatari H. , and Shikano K. 2004. Public speech-oriented guidance system with adult and child discrimination capability. In Proceedings of ICASSP. IEEE, New York (NY), United States (ICASSP, ASRU, SLT, IJCNN).","DOI":"10.1109\/ICASSP.2004.1326015"},{"key":"S135132491600005X_ref020","doi-asserted-by":"crossref","unstructured":"Gillick L. , and Cox S. 1989. Some statistical issues in the comparison of speech recognition algorithms. In Proceedings of ICASSP, IEEE, New York (NY), United States (ICASSP, ASRU, SLT, IJCNN), 1: 532\u20135.","DOI":"10.1109\/ICASSP.1989.266481"},{"key":"S135132491600005X_ref005","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.50"},{"key":"S135132491600005X_ref014","unstructured":"Erhan D. , Bengio Y. , Courville A. , Manzagol P.-A. , Vincent P. , and Bengio S. 2010. Why does unsupervised pre-training help deep learning? The Journal of Machine Learning Research 11 (2): 625\u201360."},{"key":"S135132491600005X_ref024","doi-asserted-by":"publisher","DOI":"10.1162\/neco.2006.18.7.1527"},{"key":"S135132491600005X_ref040","first-page":"365","volume-title":"Proceedings of ASRU","author":"Pinto","year":"2009"},{"key":"S135132491600005X_ref011","doi-asserted-by":"crossref","unstructured":"Das S. , Nix D. , and Picheny M. 1998. Improvements in Children's speech recognition performance. In Proceedings of ICASSP, IEEE, New York (NY), United States (ICASSP, ASRU, SLT, IJCNN).","DOI":"10.1109\/ICASSP.1998.674460"},{"key":"S135132491600005X_ref041","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2003.818026"},{"key":"S135132491600005X_ref010","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2011.2134090"},{"key":"S135132491600005X_ref018","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2009.01.006"},{"key":"S135132491600005X_ref026","first-page":"332","volume-title":"Proceedings of ASRU","author":"Imseng","year":"2013"},{"key":"S135132491600005X_ref015","doi-asserted-by":"publisher","DOI":"10.1121\/1.427148"},{"key":"S135132491600005X_ref044","doi-asserted-by":"crossref","unstructured":"Seltzer M. L. , Yu D. , and Wang Y. 2013. An investigation of deep neural networks for noise robust speech recognition. In Proceedings of ICASSP. IEEE, New York (NY), United States (ICASSP, ASRU, SLT, IJCNN).","DOI":"10.1109\/ICASSP.2013.6639100"},{"key":"S135132491600005X_ref013","first-page":"346","volume-title":"Proceedings of ICASSP","author":"Eide","year":"1996"},{"key":"S135132491600005X_ref035","doi-asserted-by":"crossref","first-page":"1365","DOI":"10.21437\/Interspeech.2008-397","volume-title":"Proceedings of INTERSPEECH","author":"Maragakis","year":"2008"},{"key":"S135132491600005X_ref034","first-page":"7947","volume-title":"Proceedings of ICASSP","author":"Liao","year":"2013"},{"key":"S135132491600005X_ref028","first-page":"4866","volume-title":"Proceedings of ICASSP","author":"Le","year":"2010"},{"key":"S135132491600005X_ref004","first-page":"6975","volume-title":"Proceedings of ICASSP","author":"Bell","year":"2013"},{"key":"S135132491600005X_ref021","unstructured":"Giuliani D. , and Gerosa M. 2003. Investigating recognition of children speech. In Proceedings of ICASSP IEEE, New York (NY), United States (ICASSP, ASRU, SLT, IJCNN), 2: 137\u201340."},{"key":"S135132491600005X_ref023","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2012.2205597"},{"key":"S135132491600005X_ref001","first-page":"7942","volume-title":"Proceedings of ICASSP","author":"Abdel-Hamid","year":"2013"},{"key":"S135132491600005X_ref017","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2007.01.002"},{"key":"S135132491600005X_ref002","doi-asserted-by":"crossref","first-page":"1248","DOI":"10.21437\/Interspeech.2013-336","volume-title":"Proceedings of INTERSPEECH","author":"Abdel-Hamid","year":"2013"},{"key":"S135132491600005X_ref059","first-page":"332","volume-title":"Proceedings of Iternational Joint Conference on Neural Networks","author":"Yochai","year":"1992"},{"key":"S135132491600005X_ref003","first-page":"1391","volume-title":"Proceedings of ICSLP","author":"Angelini","year":"1994"},{"key":"S135132491600005X_ref057","doi-asserted-by":"crossref","unstructured":"W\u00f6llmer M. , Schuller B. , Batliner A. , Steidl S. , and Seppi D. 2011. Tandem decoding of children's speech for keyword detection in a child-robot interaction scenario. ACM Transasctions Speech Language Processing 7 (4): 12:1\u201312:22.","DOI":"10.1145\/1998384.1998386"},{"key":"S135132491600005X_ref007","doi-asserted-by":"crossref","unstructured":"Bourlard H. , and Morgan N. 1994. Connectionist Speech Recognition: A Hybrid Approach, Springer, Berlin, Germany, vol. 247, Springer.","DOI":"10.1007\/978-1-4615-3210-1"},{"key":"S135132491600005X_ref058","first-page":"125","volume-title":"Proceedings of ICASSP","author":"Woodland","year":"1994"},{"key":"S135132491600005X_ref048","doi-asserted-by":"crossref","unstructured":"Sivadas S. , and Hermansky H. 2004. On use of task independent training data in tandem feature extraction. In Proceedings of ICASSP vol. 1, IEEE, New York (NY), United States (ICASSP, ASRU, SLT, IJCNN), pp. 541\u20134.","DOI":"10.1109\/ICASSP.2004.1326042"},{"key":"S135132491600005X_ref012","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2010.2064307"},{"key":"S135132491600005X_ref016","doi-asserted-by":"publisher","DOI":"10.1006\/csla.1998.0043"},{"key":"S135132491600005X_ref019","first-page":"1","volume-title":"Proceedings of the 2nd Workshop on Child, Computer and Interaction","author":"Gerosa","year":"2009"},{"key":"S135132491600005X_ref025","doi-asserted-by":"publisher","DOI":"10.1121\/1.427150"},{"key":"S135132491600005X_ref036","first-page":"1468","volume-title":"Proceedings of INTERSPEECH","author":"Metallinou","year":"2014"},{"key":"S135132491600005X_ref053","doi-asserted-by":"crossref","unstructured":"Vesel\u1ef3 K. , Burget L. , and Gr\u00e9zl F. 2010. Parallel training of neural networks for speech recognition. In Text, Speech and Dialogue, Springer, Berlin, Germany, pp. 439\u201346. Springer.","DOI":"10.1007\/978-3-642-15760-8_56"},{"key":"S135132491600005X_ref006","unstructured":"Bengio Y. , Lamblin P. , Popovici D. , and Larochelle H. 2007. Greedy layer-wise training of deep networks. In Proceedings of NIPS, NIPS, La Jolla (CA), United States, vol. 19, 153\u201360."},{"key":"S135132491600005X_ref031","doi-asserted-by":"publisher","DOI":"10.1121\/1.426686"}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S135132491600005X","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,15]],"date-time":"2024-06-15T17:22:40Z","timestamp":1718472160000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S135132491600005X\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,4,12]]},"references-count":59,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2017,5]]}},"alternative-id":["S135132491600005X"],"URL":"https:\/\/doi.org\/10.1017\/s135132491600005x","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,4,12]]}}}