{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,30]],"date-time":"2026-01-30T03:16:24Z","timestamp":1769742984374,"version":"3.49.0"},"reference-count":33,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2018,11,21]],"date-time":"2018-11-21T00:00:00Z","timestamp":1542758400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61673395"],"award-info":[{"award-number":["61673395"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61403415"],"award-info":[{"award-number":["61403415"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100006407","name":"Natural Science Foundation of Henan Province","doi-asserted-by":"publisher","award":["162300410331"],"award-info":[{"award-number":["162300410331"]}],"id":[{"id":"10.13039\/501100006407","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"published-print":{"date-parts":[[2018,12]]},"DOI":"10.1186\/s13636-018-0141-9","type":"journal-article","created":{"date-parts":[[2018,11,21]],"date-time":"2018-11-21T12:03:44Z","timestamp":1542801824000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":30,"title":["Towards end-to-end speech recognition with transfer learning"],"prefix":"10.1186","volume":"2018","author":[{"given":"Chu-Xiong","family":"Qin","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dan","family":"Qu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lian-Hai","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2018,11,21]]},"reference":[{"key":"141_CR1","doi-asserted-by":"crossref","first-page":"369","DOI":"10.1145\/1143844.1143891","volume-title":"International Conference on Machine learning (ICML)","author":"A Graves","year":"2006","unstructured":"Graves, A., & Gomez, F. (2006). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. In International Conference on Machine learning (ICML) (pp. 369\u2013376)."},{"key":"141_CR2","first-page":"1602","volume":"v1412","author":"J Chorowski","year":"2014","unstructured":"Chorowski, J., Bahdanau, D., Cho, K., & Bengio, Y. (2014). End-to-end continuous speech recognition using attention-based recurrent NN: First results. arXiv preprint arXiv, v1412, 1602.","journal-title":"arXiv preprint arXiv"},{"key":"141_CR3","first-page":"1764","volume-title":"International Conference on Machine Learning (ICML)","author":"A Graves","year":"2014","unstructured":"Graves, A., & Jaitly, N. (2014). Towards end-to-end speech recognition with recurrent neural networks. In International Conference on Machine Learning (ICML) (pp. 1764\u20131772)."},{"key":"141_CR4","volume-title":"\u201cDeep speech 2: End-to-end speech recognition in English and mandarin,\u201d Computer Science","author":"D Amodei","year":"2015","unstructured":"D. Amodei, R. Anubhai, E. Battenberg, et al, \u201cDeep speech 2: End-to-end speech recognition in English and mandarin,\u201d Computer Science, 2015."},{"key":"141_CR5","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"V Panayotov","year":"2015","unstructured":"Panayotov, V., Chen, G., Povey, D., & Khudanpur, S. (2015). Librispeech: An ASR corpus based on public domain audio books. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)."},{"key":"141_CR6","first-page":"577","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"JK Chorowski","year":"2015","unstructured":"Chorowski, J. K., Bahdanau, D., Serdyuk, D., Cho, K., & Bengio, Y. (2015). Attention-based models for speech recognition. In Advances in Neural Information Processing Systems (NIPS) (pp. 577\u2013585)."},{"key":"141_CR7","volume-title":"International Conference on Statistical Language and Speech Processing (ICASSP)","author":"J Van\u011bk","year":"2017","unstructured":"Van\u011bk, J., Zelinka, J., Soutner, D., & Psutka, J. (2017). A regularization post layer: An additional way how to make deep neural networks robust. In International Conference on Statistical Language and Speech Processing (ICASSP)."},{"key":"141_CR8","first-page":"5567","volume":"1412","author":"A Hannun","year":"2014","unstructured":"Hannun, A., Case, C., Casper, J., et al. (2014). Deep speech: Scaling up end-to-end speech recognition. arXiv preprint arXiv, 1412, 5567.","journal-title":"arXiv preprint arXiv"},{"key":"141_CR9","first-page":"4960","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"W Chan","year":"2016","unstructured":"Chan, W., Jaitly, N., Le, Q., & Vinyals, O. (2016). Listen, attend and spell: A neural network for large vocabulary conversational speech recognition. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 4960\u20134964)."},{"key":"141_CR10","first-page":"01769","volume":"1712","author":"CC Chiu","year":"2018","unstructured":"Chiu, C. C., Sainath, T. N., Wu, Y., Prabhavalkar, R., Nguyen, P., Chen, Z., Kannan, A., Weiss, R. J., Rao, K., Gonina, K., et al. (2018). State-of-the-art speech recognition with sequence-to-sequence models. arXiv preprint arXiv, 1712, 01769.","journal-title":"arXiv preprint arXiv"},{"key":"141_CR11","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Y Zhang","year":"2017","unstructured":"Zhang, Y., Chan, W., & Jaitly, N. (2017). Very deep convolutional networks for end-to-end speech recognition. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)."},{"key":"141_CR12","first-page":"09288","volume":"1611","author":"T Sercu","year":"2016","unstructured":"Sercu, T., & Goel, V. (2016). Dense prediction on sequences with time-dilated convolutions for speech recognition. arXiv preprint arXiv, 1611, 09288.","journal-title":"arXiv preprint arXiv"},{"key":"141_CR13","first-page":"07793","volume":"1702","author":"Y Wang","year":"2017","unstructured":"Wang, Y., Deng, X., Pu, S., & Huang, Z. (2017). Residual convolutional CTC networks for automatic speech recognition. arXiv preprint arXiv, 1702, 07793.","journal-title":"arXiv preprint arXiv"},{"key":"141_CR14","first-page":"06378","volume":"1702","author":"L Lu","year":"2017","unstructured":"Lu, L., Kong, L., Dyer, C., & Smith, N. A. (2017). Multi-task learning with CTC and segmental CRF for speech recognition. arXiv preprint arXiv, 1702, 06378.","journal-title":"arXiv preprint arXiv"},{"key":"141_CR15","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"S Kim","year":"2017","unstructured":"Kim, S., Hori, T., & Watanabe, S. (2017). Joint CTC-attention based end-to-end speech recognition using multi-task learning. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)."},{"key":"141_CR16","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"JT Huang","year":"2013","unstructured":"Huang, J. T., Li, J., Yu, D., Deng, L., Gong, Y., et al. (2013). Cross-language knowledge transfer using multilingual deep neural network with shared hidden layers. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)."},{"key":"141_CR17","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"J Deng","year":"2014","unstructured":"Deng, J., Xia, R., Zhang, Z., Liu, Y., & Schuller, B. (2014). Introducing shared-hidden-layer autoencoders for transfer learning and their application in acoustic emotion recognition. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)."},{"key":"141_CR18","first-page":"02737","volume":"1706","author":"T Hori","year":"2017","unstructured":"Hori, T., Watanabe, S., Zhang, Y., & Chan, W. (2017). Advances in joint CTC-attention based end-to-end speech recognition with a deep CNN encoder and RNN-LM. arXiv preprint arXiv, 1706, 02737.","journal-title":"arXiv preprint arXiv"},{"key":"141_CR19","first-page":"03499","volume":"1609","author":"A Van Den Oord","year":"2016","unstructured":"Van Den Oord, A., Dieleman, S., Zen, H., et al. (2016). Wavenet: A generative model for raw audio. arXiv preprint arXiv, 1609, 03499.","journal-title":"arXiv preprint arXiv"},{"key":"141_CR20","volume-title":"Information Technology, Networking, Electronic and Automation Control Conference (ITNEC)","author":"C Qin","year":"2016","unstructured":"Qin, C., & Zhang, L. (2016). Deep neural network based feature extraction using convex-nonnegative matrix factorization for low-resource speech recognition. In Information Technology, Networking, Electronic and Automation Control Conference (ITNEC)."},{"key":"141_CR21","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1007\/0-306-47815-3_5","volume-title":"A practical approach to microarray data analysis","author":"ME Wall","year":"2003","unstructured":"Wall, M. E., Rechtsteiner, A., & Rocha, L. M. (2003). Singular value decomposition and principal component analysis. In A practical approach to microarray data analysis (pp. 91\u2013109). Boston, MA: Springer."},{"key":"141_CR22","first-page":"2365","volume-title":"Interspeech","author":"J Xue","year":"2013","unstructured":"Xue, J., Li, J., & Gong, Y. (2013). Restructuring of deep neural network acoustic models with singular value decomposition. In Interspeech (pp. 2365\u20132369)."},{"issue":"1","key":"141_CR23","doi-asserted-by":"publisher","first-page":"45","DOI":"10.1109\/TPAMI.2008.277","volume":"32","author":"CH Ding","year":"2010","unstructured":"Ding, C. H., Li, T., & Jordan, M. I. (2010). Convex and semi-nonnegative matrix factorizations. IEEE transactions on pattern analysis and machine intelligence, 32(1), 45\u201355.","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"141_CR24","volume-title":"\u201cJoint CTC\/attention decoding for end-to-end speech recognition,\u201d Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics. Vol. 1","author":"T Hori","year":"2017","unstructured":"T. Hori, S. Watanabe, and J. Hershey, \u201cJoint CTC\/attention decoding for end-to-end speech recognition,\u201d Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics. Vol. 1. 2017."},{"key":"141_CR25","first-page":"6980","volume":"1412","author":"DP Kingma","year":"2014","unstructured":"Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv, 1412, 6980.","journal-title":"arXiv preprint arXiv"},{"key":"141_CR26","first-page":"5063","volume":"1211","author":"R Pascanu","year":"2012","unstructured":"Pascanu, R., Mikolov, T., & Bengio, Y. (2012). On the difficulty of training recurrent neural networks. arXiv preprint arXiv, 1211, 5063.","journal-title":"arXiv preprint arXiv"},{"key":"141_CR27","first-page":"5701","volume":"1212","author":"MD Zeiler","year":"2012","unstructured":"Zeiler, M. D. (2012). Adadelta: an adaptive learning rate method. arXiv preprint arXiv, 1212, 5701.","journal-title":"arXiv preprint arXiv"},{"issue":"5","key":"141_CR28","doi-asserted-by":"publisher","first-page":"434","DOI":"10.1016\/j.specom.2008.01.002","volume":"50","author":"M Bisani","year":"2008","unstructured":"Bisani, M., & Ney, H. (2008). Joint-sequence models for grapheme-to-phoneme conversion. Speech Comm., 50(5), 434\u2013451.","journal-title":"Speech Comm."},{"key":"141_CR29","unstructured":"R. L. Weide, \u201cThe CMU pronouncing dictionary,\u201d URL: http:\/\/www.speech.cs.cmu.edu\/cgi-bin\/cmudict , 1998."},{"key":"141_CR30","doi-asserted-by":"crossref","unstructured":"T\u00f3th, L. (2015). Phone recognition with hierarchical convolutional deep maxout networks. EURASIP Journal on Audio, Speech, and Music Processing,\u00a0vol. 2015, no. 1, pp. 1\u201313.","DOI":"10.1186\/s13636-015-0068-3"},{"key":"141_CR31","first-page":"6645","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"A Graves","year":"2013","unstructured":"Graves, A., Mohamed, A. R., & Hinton, G. (2013). Speech recognition with deep recurrent neural networks. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6645\u20136649)."},{"key":"141_CR32","first-page":"02720","volume":"1701","author":"Y Zhang","year":"2017","unstructured":"Zhang, Y., Pezeshki, M., Brakel, P., Zhang, S., Bengio, C. L. Y., & Courville, A. (2017). Towards end-to-end speech recognition with deep convolutional neural networks. arXiv preprint arXiv, 1701, 02720.","journal-title":"arXiv preprint arXiv"},{"key":"141_CR33","first-page":"01161","volume":"1711","author":"N Zeghidour","year":"2017","unstructured":"Zeghidour, N., Usunier, N., Kokkinos, I., Schatz, T., Synnaeve, G., & Dupoux, E. (2017). Learning Filterbanks from Raw Speech for Phone Recognition. arXiv preprint arXiv, 1711, 01161.","journal-title":"arXiv preprint arXiv"}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-018-0141-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/s13636-018-0141-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-018-0141-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,9,6]],"date-time":"2022-09-06T17:06:47Z","timestamp":1662484007000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-018-0141-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,11,21]]},"references-count":33,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2018,12]]}},"alternative-id":["141"],"URL":"https:\/\/doi.org\/10.1186\/s13636-018-0141-9","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,11,21]]},"assertion":[{"value":"29 July 2018","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 September 2018","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 November 2018","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Chu-Xiong Qin was born in Shijiazhuang, China, in 1991. He received the B.S. and M.S. degrees in information and communication from the National Digital Switching System Engineering and Technological R&D Center, Zhengzhou, China, in 2013 and 2016, respectively. He is currently working towards the Ph.D. degree on speech recognition at the National Digital Switching System Engineering and Technological R&D Center. His research interests are in speech signal processing, continuous speech recognition, and machine learning.Dan Qu received the M.S. degree in communication and information system from Xi\u2019an Information Science and Technology Institute, Xi\u2019an, China, in 2000 and the Ph.D. degree in information and communication engineering from the National Digital Switching System Engineering and Technological R&D Center, Zhengzhou, China, in 2005. She is an Associate Professor at the National Digital Switching System Engineering and Technological R&D Center. Her research interests are in speech signal processing and pattern recognition.Lian-Hai Zhang received the M.S. degree in information and communication engineering from the National Digital Switching System Engineering and Technological R&D Center, Zhengzhou, China, in 2000. He is an Associate Professor at the National Digital Switching System Engineering and Technological R&D Center. His research interests are in speech signal processing.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Authors\u2019 information"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}},{"value":"Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Publisher\u2019s Note"}}],"article-number":"18"}}