{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T10:31:01Z","timestamp":1777458661319,"version":"3.51.4"},"reference-count":42,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2017,9,23]],"date-time":"2017-09-23T00:00:00Z","timestamp":1506124800000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Sign Process Syst"],"published-print":{"date-parts":[[2018,7]]},"DOI":"10.1007\/s11265-017-1291-1","type":"journal-article","created":{"date-parts":[[2017,9,23]],"date-time":"2017-09-23T11:15:02Z","timestamp":1506165302000},"page":"985-997","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["CTC Regularized Model Adaptation for Improving LSTM RNN Based Multi-Accent Mandarin Speech Recognition"],"prefix":"10.1007","volume":"90","author":[{"given":"Jiangyan","family":"Yi","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhengqi","family":"Wen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianhua","family":"Tao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hao","family":"Ni","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bin","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2017,9,23]]},"reference":[{"issue":"2","key":"1291_CR1","doi-asserted-by":"crossref","first-page":"141","DOI":"10.1023\/B:IJST.0000017014.52972.1d","volume":"7","author":"C Huang","year":"2004","unstructured":"Huang, C., Chen, T., & Chang, E. (2004). Accent Issues in Large Vocabulary Continuous Speech Recognition. Int J Speech Technol, 7(2), 141\u2013153.","journal-title":"Int J Speech Technol"},{"key":"1291_CR2","unstructured":"Wang, Z., Schultz, T., & Waibel, A. (2013). Comparison of Acoustic Model Adaptation Techniques on Non-native Speech. In the Proceedings of the 2013 I.E. International Conference on Acoustics, Speech, and Signal Processing (ICASSP)."},{"issue":"1","key":"1291_CR3","doi-asserted-by":"crossref","first-page":"28","DOI":"10.1121\/1.419608","volume":"102","author":"LM Arslan","year":"1997","unstructured":"Arslan, L. M., & Hansen, J. L. (1997). A study of the temporal features and frequency characteristics in American english foreign accent. Journal of the Acoustical Society of America, 102(1), 28\u201340.","journal-title":"Journal of the Acoustical Society of America"},{"key":"1291_CR4","doi-asserted-by":"crossref","unstructured":"Liu, Y., & P. Fung (2006). Multi-accent Chinese Speech Recognition. In the Proceedings of Interspeech.","DOI":"10.21437\/Interspeech.2006-34"},{"issue":"4","key":"1291_CR5","doi-asserted-by":"crossref","first-page":"3279","DOI":"10.1121\/1.2035588","volume":"118","author":"P Fung","year":"2005","unstructured":"Fung, P., & Liu, Y. (Nov. 2005). Effects and Modeling of Phonetic and Acoustic Confusions in Accented Speech. J Acoust Soc Amer, 118(4), 3279\u20133293.","journal-title":"J Acoust Soc Amer"},{"key":"1291_CR6","unstructured":"Leading Group Office of Survey of Language Use in China (2006). In survey of language use in China. Beijing: Yu Wen Press (in Chinese)."},{"key":"1291_CR7","unstructured":"Davis S. B., & Mermelstein, P. (2013) Reliable Accent-Specific Unit Generation With Discriminative Dynamic Gaussian Mixture Selection for Multi-Accent Chinese Speech Recognition. IEEE Trans Acoustics Speech Signal Process, 21 (10), 2073\u20132084."},{"key":"1291_CR8","doi-asserted-by":"crossref","unstructured":"Zheng, Y. L., Sproat, R., Gu, L., Shafran, I., Zhou, H., Su, Y., Jurafsky, D., Starr, R., & Yoon, S. (2005). Accent Detection and Speech Recognition for Shanghai-accented Mandarin. In the Proceedings of Interspeech.","DOI":"10.21437\/Interspeech.2005-112"},{"key":"1291_CR9","doi-asserted-by":"crossref","unstructured":"Vergyri, D., Lamel, L., & Gauvain, L. (2010). Automatic Speech Recognition of Multiple Accented English Data. In the Proceedings of Interspeech.","DOI":"10.21437\/Interspeech.2010-477"},{"key":"1291_CR10","doi-asserted-by":"crossref","unstructured":"Ding, G. H. (2008). Phonetic Confusion Analysis and Robust Phone Set Generation for Shanghai-Accented Mandarin Speech Recognition. In the Proceedings of Interspeech.","DOI":"10.21437\/Interspeech.2008-344"},{"issue":"2","key":"1291_CR11","doi-asserted-by":"crossref","first-page":"153","DOI":"10.1016\/j.specom.2005.03.003","volume":"46","author":"E Fosler-Lussier","year":"2005","unstructured":"Fosler-Lussier, E., Amdal, I., & Kuo, H.-K. J. (2005). A Framework for Predicting Speech Recognition Errors. Speech Communication, 46(2), 153\u2013170.","journal-title":"Speech Communication"},{"key":"1291_CR12","unstructured":"Fosler-Lussier, E. (1999). Dynamic Pronunciation Models for Automatic Speech Recognition. Ph.D. dissertation, Int. Comput. Sci. Inst., Berkeley, CA, USA."},{"key":"1291_CR13","doi-asserted-by":"crossref","unstructured":"Hain, T., & Woodland, P. C. (1999). Dynamic HMM Selection for Continuous Speech Recognition. In Proc. Eurospeech, pp. 1327\u20131330.","DOI":"10.21437\/Eurospeech.1999-339x"},{"key":"1291_CR14","unstructured":"V. Fisher et al. (1998). Speaker-Independent Upfront Dialect Adaptation in A Large Vocabulary Continuous Speech Recognition. In Proc. Int. Conf. Spoken Lang. Process."},{"key":"1291_CR15","unstructured":"Wang, Z., Schultz, T., & Waibel, A. (2003). Comparison of Acoustic Model Adaptation Techniques on Non-Native Speech. In ICASSP 2003. IEEE, pp. 540\u2013543."},{"key":"1291_CR16","unstructured":"Mayfield Tomokiyo, L., & Waibel, A. (2001). Adaptation Methods for Non-Native Speech,\u201d in Proceedings of Multilinguality in Spoken Language Processing, Aalborg."},{"key":"1291_CR17","doi-asserted-by":"crossref","unstructured":"Huang, C., Chang, E., Zhou, J., & Lee, K.-F. (2000). Accent Modeling Based on Pronunciation Dictionary Adaptation for Large Vocabulary Mandarin Speech Recognition. In ICSLP 2000, Beijing, pp. 818\u2013821.","DOI":"10.21437\/ICSLP.2000-660"},{"issue":"1","key":"1291_CR18","first-page":"33","volume":"1","author":"GE Dahl","year":"2012","unstructured":"Dahl, G. E., Yu, D., Deng, L., & Acero, A. (2012). Context-Dependent Pre-trained Deep Neural Networks for Large Vocabulary Speech Recognition. IEEE Trans Audio Speech Lang Process, 1(1), 33\u201342.","journal-title":"IEEE Trans Audio Speech Lang Process"},{"key":"1291_CR19","unstructured":"Seide, F., Li, G., & Yu, D. (2012). Conversational Speech Transcription Using Context-Dependent Deep Neural Networks. In the Proceedings of Interspeech."},{"issue":"6","key":"1291_CR20","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1109\/MSP.2012.2205597","volume":"29","author":"G Hinton","year":"2012","unstructured":"Hinton, G., Deng, L., Yu, D., Dahl, G., Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., NSainath, T., et al. (2012). Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal Process Mag, 29(6), 82\u201397.","journal-title":"IEEE Signal Process Mag"},{"key":"1291_CR21","unstructured":"Yu, D., Seltzer, M., Li, J., Huang, J., & Seide, F. (2013). Feature learning in Deep Neural Networks - Studies on Speech Recognition Tasks. In the Proceedings of 2013 International Confernece on Learning Representation."},{"key":"1291_CR22","unstructured":"Goodfellow, I. J., Le, Q. V., Saxe, A. M., Lee, H., & Ng, A. Y. (2009). Measuring Invariances in Deep Networks. Advances in Neural Information Processing Systems (NIPS) 22."},{"key":"1291_CR23","doi-asserted-by":"crossref","unstructured":"Huang, Y., Yu, D., Liu, C. J., & Gong, Y. F. (2014). Multi-Accent Deep Neural Network Acoustic Model with Accent-Specific Top Layer Using the KLD-Regularized Model Adaptation. In the Proceedings of Interspeech.","DOI":"10.21437\/Interspeech.2014-497"},{"key":"1291_CR24","doi-asserted-by":"crossref","unstructured":"Huang, J., Li, J., Yu, D., Deng, L., & Gong, Y. F. (2013). Cross-Language Knowledge Transfer Using Multilingual Deep Neural Network With Shared Hidden Layers. In the Proceedings of the 2013 I.E. International Conference on Acoustics, Speech and Signal Processing (ICASSP).","DOI":"10.1109\/ICASSP.2013.6639081"},{"key":"1291_CR25","doi-asserted-by":"crossref","unstructured":"Chen, M. M., Yang, Z. Y., Liang, J. Z., Li, Y. P., Liu, W. J. (2015). Improving Deep Neural Networks Based Multi-Accent Mandarin Speech Recognition Using I-Vectors and Accent-Specific Top layer. In the Proceedings of Interspeech.","DOI":"10.21437\/Interspeech.2015-718"},{"key":"1291_CR26","unstructured":"Sak, H., Senior, A., & Beaufays, F. (2014). Long Short-Term Memory Based Recurrent Neural Network Architectures for Large Vocabulary Speech Recognition. In the Proceedings of Interspeech."},{"key":"1291_CR27","doi-asserted-by":"crossref","unstructured":"Liu, C., Wang, Y., Kumar, K., & Gong, Y. F. (2016). Investigations on Speaker Adaptation of LSTM RNN Models for Speech Recognition. In the Proceedings of the 2016 I.E. International Conference on Acoustics, Speech and Signal Processing (ICASSP).","DOI":"10.1109\/ICASSP.2016.7472633"},{"key":"1291_CR28","doi-asserted-by":"crossref","unstructured":"Huang, Z., Tang, J., Xue, S., & Dai, L. (2016). Speaker Adaptation of RNN-BLSTM for Speech Recognition Based on Speaker Code. In the Proceedings of the 2016 I.E. International Conference on Acoustics, Speech and Signal Processing (ICASSP).","DOI":"10.1109\/ICASSP.2016.7472690"},{"key":"1291_CR29","doi-asserted-by":"crossref","unstructured":"Tan, T., Qian, Y., Yu, D., Kundu, S., & Lu, L. (2016). Speaker-Aware Training of LSTM-RNNs for Acoustic Modelling. In the Proceedings of the 2016 I.E. International Conference on Acoustics, Speech and Signal Processing (ICASSP).","DOI":"10.1109\/ICASSP.2016.7472685"},{"key":"1291_CR30","doi-asserted-by":"crossref","unstructured":"Yi, J., Ni, H., Wen, Z. H., & Tao, J. (2016). Improving BLSTM RNN Based Mandarin Speech Recognition Using Accent Dependent Bottleneck Features. Asia-Pacific Signal and Information Processing Association Annual Summit and Conference.","DOI":"10.1109\/APSIPA.2016.7820723"},{"key":"1291_CR31","doi-asserted-by":"crossref","unstructured":"Graves, A., Fernandez, S., Gomez, F., & Schmidhuber, J. (2006). Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks. In ICML, Pittsburgh, USA.","DOI":"10.1145\/1143844.1143891"},{"key":"1291_CR32","doi-asserted-by":"crossref","unstructured":"Graves, A., Mohamed, A., & Hinton, G. (2013). Speech Recognition With Deep Recurrent Neural Networks. In the Proceedings of the 2013 I.E. International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, pp. 6645\u20136649.","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"1291_CR33","unstructured":"Graves, A., & Jaitly, N. (2014). Towards End-To-End Speech Recognition with Recurrent Neural Networks. In Proceedings of the 31st International Conference on Machine Learning (ICML-14), pp. 1764\u20131772."},{"key":"1291_CR34","unstructured":"Hannun, A., Case, C., Casper, J., Catanzaro, B., Diamos, G., Elsen, E., Prenger, R., Satheesh, S., Sengupta, S., Coates, A., et al. (2014). Deepspeech: Scaling up End-To-End Speech Recognition. arXiv preprint arXiv:1412.5567."},{"key":"1291_CR35","unstructured":"Hannun, A. Y., Maas, A. L., Jurafsky, D., & Ng, A. Y. (2014). First-Pass Large Vocabulary Continuous Speech Recognition Using Bi-Directional Recurrent DNNs. arXiv preprint arXiv:1408.2873."},{"key":"1291_CR36","doi-asserted-by":"crossref","unstructured":"Miao, Y. J., Gowayyed, M. & Metze, F. (2015). EESEN: End-to-End Speech Recognition using Deep RNN Models and WFST-based Decoding. In the Proceedings of ASRU.","DOI":"10.1109\/ASRU.2015.7404790"},{"key":"1291_CR37","doi-asserted-by":"crossref","unstructured":"Yu, D., Yao, K., Su, H., Li, G., & Seide, F. (2013). KL-Divergence Regularized Deep Neural Network Adaptation for Improved Large Vocabulary Speech Recognition. In the Proceedings of the 2013 I.E. International Conference on Acoustics, Speech and Signal Processing (ICASSP).","DOI":"10.1109\/ICASSP.2013.6639201"},{"key":"1291_CR38","unstructured":"(2003). RASC863: 863 annotated 4 regional accent speech corpus. Chinese Academy of Social Sciences. Available: http:\/\/www.chineseldc.org\/doc\/CLDC-SPC-2004-005\/intro.htm ."},{"key":"1291_CR39","unstructured":"(2003). CASIA: CASIA northern accent speech corpus. Chinese Academy of Sciences. Available: http:\/\/www.chineseldc.org\/doc\/CLDC-SPC-2004-015\/intro.htm ."},{"key":"1291_CR40","unstructured":"Povey, D., Ghoshal, A., Boulianne, G., Burget, L., Glembek, O., Goel, N., Hannemann, M., Motlicek, P., Qian, Y. M., Schwarz, P., Silovsky, J., Stemmer, G., & Vesely, K. (2011). The Kaldi SpeechRecognition Toolkit. In the Proceedings of ASRU."},{"key":"1291_CR41","unstructured":"Li, X., & Bilmes, J. (2006). Regularized adaptation of discriminative classifiers. In the Proceedings of the 2013 I.E. International Conference on Acoustics, Speech and Signal Processing (ICASSP)."},{"key":"1291_CR42","doi-asserted-by":"crossref","unstructured":"Miao, Y., Metze, F. (2015). On Speaker Adaptation of Long Short-Term Memory Recurrent Neural Networks. In the Proceedings of Interspeech.","DOI":"10.21437\/Interspeech.2015-290"}],"container-title":["Journal of Signal Processing Systems"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11265-017-1291-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11265-017-1291-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11265-017-1291-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,25]],"date-time":"2025-06-25T21:42:53Z","timestamp":1750887773000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11265-017-1291-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,9,23]]},"references-count":42,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2018,7]]}},"alternative-id":["1291"],"URL":"https:\/\/doi.org\/10.1007\/s11265-017-1291-1","relation":{},"ISSN":["1939-8018","1939-8115"],"issn-type":[{"value":"1939-8018","type":"print"},{"value":"1939-8115","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,9,23]]}}}