{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,5,14]],"date-time":"2025-05-14T20:25:55Z","timestamp":1747254355713,"version":"3.37.3"},"reference-count":43,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2019,10,28]],"date-time":"2019-10-28T00:00:00Z","timestamp":1572220800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2019,10,28]],"date-time":"2019-10-28T00:00:00Z","timestamp":1572220800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No. 61673395","No. 61403415"],"award-info":[{"award-number":["No. 61673395","No. 61403415"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100006407","name":"Natural Science Foundation of Henan Province","doi-asserted-by":"publisher","award":["No. 162300410331"],"award-info":[{"award-number":["No. 162300410331"]}],"id":[{"id":"10.13039\/501100006407","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"publisher","award":["2016M602975"],"award-info":[{"award-number":["2016M602975"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"published-print":{"date-parts":[[2019,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n              <jats:p>A method called joint connectionist temporal classification (CTC)-attention-based speech recognition has recently received increasing focus and has achieved impressive performance. A hybrid end-to-end architecture that adds an extra CTC loss to the attention-based model could force extra restrictions on alignments. To explore better the end-to-end models, we propose improvements to the feature extraction and attention mechanism. First, we introduce a joint model trained with nonnegative matrix factorization (NMF)-based high-level features. Then, we put forward a hybrid attention mechanism by incorporating multi-head attentions and calculating attention scores over multi-level outputs. Experiments on TIMIT indicate that the new method achieves state-of-the-art performance with our best model. Experiments on WSJ show that our method exhibits a word error rate (WER) that is only 0.2% worse in absolute value than the best referenced method, which is trained on a much larger dataset, and it beats all present end-to-end methods. Further experiments on LibriSpeech show that our method is also comparable to the state-of-the-art end-to-end system in WER.<\/jats:p>","DOI":"10.1186\/s13636-019-0161-0","type":"journal-article","created":{"date-parts":[[2019,12,16]],"date-time":"2019-12-16T15:06:07Z","timestamp":1576508767000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["A new joint CTC-attention-based speech recognition model with multi-level multi-head attention"],"prefix":"10.1186","volume":"2019","author":[{"given":"Chu-Xiong","family":"Qin","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wen-Lin","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dan","family":"Qu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2019,10,28]]},"reference":[{"issue":"2","key":"161_CR1","doi-asserted-by":"publisher","first-page":"257","DOI":"10.1109\/5.18626","volume":"77","author":"LR Rabiner","year":"1989","unstructured":"L.R. Rabiner, A tutorial on hidden Markov models and selected applications in speech recognition. Proc. IEEE 77(2), 257\u2013286 (1989)","journal-title":"Proc. IEEE"},{"key":"161_CR2","volume-title":"Fundamentals of speech recognition","author":"LR Rabiner","year":"1993","unstructured":"L.R. Rabiner, B.H. Juang, Fundamentals of speech recognition, vol 14 (PTR Prentice Hall, Englewood Cliffs, 1993)"},{"key":"161_CR3","unstructured":"J.K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, Attention-based models for speech recognition n Advances in Neural Information Processing Systems (NIPS), 2015, pp. 577\u2013585"},{"key":"161_CR4","doi-asserted-by":"crossref","unstructured":"Y. Miao, M. Gowayyed, and F. Metze, EESEN: end-to-end speech recognition using deep RNN models and WFST-based decoding, Automatic Speech Recognition and Understanding (ASRU), 2015 IEEE Workshop on. IEEE, 2015","DOI":"10.1109\/ASRU.2015.7404790"},{"key":"161_CR5","doi-asserted-by":"crossref","unstructured":"C. C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, K. Gonina, et al., State-of-the-art speech recognition with sequence-to-sequence models, 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018","DOI":"10.1109\/ICASSP.2018.8462105"},{"key":"161_CR6","unstructured":"D. Bahdanau, K. Cho, and Y. Bengio, Neural machine translation by jointly learning to align and translate, arXiv preprint arXiv:1409.0473 (2014)"},{"key":"161_CR7","doi-asserted-by":"crossref","unstructured":"Luong, Minh-Thang, Hieu Pham, and Christopher D. Manning, Effective approaches to attention-based neural machine translation, arXiv preprint arXiv:1508.04025 (2015)","DOI":"10.18653\/v1\/D15-1166"},{"key":"161_CR8","doi-asserted-by":"crossref","unstructured":"W. Chan, N. Jaitly, Q. Le, and O. Vinyals, Listen, attend and spell: a neural network for large vocabulary conversational speech recognitione, in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016, pp. 4960\u20134964","DOI":"10.1109\/ICASSP.2016.7472621"},{"key":"161_CR9","unstructured":"P. Ren, Z. Chen, Z. Ren, et al., (2017), Leveraging contextual sentence relations for extractive summarization using a neural attention model, in Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 95\u2013104"},{"key":"161_CR10","doi-asserted-by":"crossref","unstructured":"Y. Cui, Z. Chen, S. Wei, et al., Attention-over-attention neural networks for reading comprehension arXiv preprint arXiv:1607.04423(2016)","DOI":"10.18653\/v1\/P17-1055"},{"key":"161_CR11","unstructured":"J. Gehring, M. Auli, D. Grangier, et al., Convolutional sequence to sequence learning, arXiv preprint arXiv:1705.03122, 2017"},{"issue":"8","key":"161_CR12","doi-asserted-by":"publisher","first-page":"1240","DOI":"10.1109\/JSTSP.2017.2763455","volume":"11","author":"S Watanabe","year":"2017","unstructured":"S. Watanabe, T. Hori, S. Kim, et al., Hybrid CTC\/attention architecture for end-to-end speech recognition. IEEE. J. Select. Topics. Signal. Process. 11(8), 1240\u20131253 (2017)","journal-title":"IEEE. J. Select. Topics. Signal. Process."},{"key":"161_CR13","doi-asserted-by":"crossref","unstructured":"T. Hori, S. Watanabe, and J. Hershey, Joint CTC\/attention decoding for end-to-end speech recognition, Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1. 2017","DOI":"10.18653\/v1\/P17-1048"},{"key":"161_CR14","doi-asserted-by":"crossref","unstructured":"S. Kim, T. Hori, S. Watanabe, Joint CTC-attention based end-to-end speech recognition using multi-task learning, Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on. IEEE, 2017: 4835\u20134839","DOI":"10.1109\/ICASSP.2017.7953075"},{"key":"161_CR15","volume-title":"Advances in joint CTC-attention based end-to-end speech recognition with a deep CNN encoder and RNN-LM, arXiv preprint arXiv:1706.02737","author":"T Hori","year":"2017","unstructured":"T. Hori, S. Watanabe, Y. Zhang, et al., Advances in joint CTC-attention based end-to-end speech recognition with a deep CNN encoder and RNN-LM, arXiv preprint arXiv:1706.02737 (2017)"},{"key":"161_CR16","doi-asserted-by":"crossref","unstructured":"C. Qin, and L. Zhang, Deep neural network based feature extraction using convex-nonnegative matrix factorization for low-resource speech recognition, in Information Technology, Networking, Electronic and Automation Control Conference (ITNEC), 2016","DOI":"10.1109\/ITNEC.2016.7560531"},{"key":"161_CR17","unstructured":"C.-X. Qin, Q. Dan, L.-H. Zhang, Towards end-to-end speech recognition with transfer learning. EURASIP J Audio. Speech. Music. Process 18, 1\u20139 (2018)"},{"key":"161_CR18","first-page":"5998","volume":"2017","author":"A Vaswani","year":"2017","unstructured":"A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, et al., Attention is all you need. Adv. Neural. Inf. Process. Syst.(NIPS) 2017, 5998\u20136008 (2017)","journal-title":"Adv. Neural. Inf. Process. Syst.(NIPS)"},{"issue":"1","key":"161_CR19","doi-asserted-by":"publisher","first-page":"45","DOI":"10.1109\/TPAMI.2008.277","volume":"32","author":"CH Ding","year":"2010","unstructured":"C.H. Ding, T. Li, M.I. Jordan, Convex and semi-nonnegative matrix factorizations. IEEE Trans. Pattern Anal. Mach. Intell. 32(1), 45\u201355 (2010)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"161_CR20","unstructured":"J. L. Ba, J. R. Kiros, and G. E. Hinton, Layer normalization, arXiv preprint arXiv:1607.06450 (2016)"},{"key":"161_CR21","volume-title":"The Kaldi speech recognition toolkit, IEEE 2011 workshop on automatic speech recognition and understanding. No. EPFL-CONF-192584. IEEE Signal Processing Society","author":"D Povey","year":"2011","unstructured":"D. Povey, A. Ghoshal, G. Boulianne, et al., The Kaldi speech recognition toolkit, IEEE 2011 workshop on automatic speech recognition and understanding. No. EPFL-CONF-192584. IEEE Signal Processing Society, 2011"},{"key":"161_CR22","volume-title":"Theano","author":"the Montreal Institute for Leaning Algorithms (MILA)","year":"2017","unstructured":"the Montreal Institute for Leaning Algorithms (MILA), Theano, \n                    http:\/\/www.deeplearning.net\/software\/theano\/\n                    \n                  , 2017"},{"key":"161_CR23","unstructured":"Christian Thurau, PyMF - Python Matrix Factorization Module, \n                    https:\/\/github.com\/cthurau\/pymf\n                    \n                  , 2013"},{"key":"161_CR24","volume-title":"Chainer: a next-generation open source framework for deep learning, Proceedings of Workshop on Machine Learning Systems (LearningSys) in The Twenty-ninth Annual Conference on Neural Information Processing Systems (NIPS)","author":"S Tokui","year":"2015","unstructured":"S. Tokui, K. Oono, S. Hido, J. Clayton, Chainer: a next-generation open source framework for deep learning, Proceedings of Workshop on Machine Learning Systems (LearningSys) in The Twenty-ninth Annual Conference on Neural Information Processing Systems (NIPS) (2015)"},{"key":"161_CR25","volume-title":"ESPnet: end-to-end speech processing toolkit, arXiv preprint arXiv:1804.00015","author":"S Watanabe","year":"2018","unstructured":"S. Watanabe, T. Hori, S. Karita, et al., ESPnet: end-to-end speech processing toolkit, arXiv preprint arXiv:1804.00015, 2018"},{"key":"161_CR26","doi-asserted-by":"crossref","unstructured":"L. T\u00f3th, Phone recognition with hierarchical convolutional deep maxout networks. EURASIP J Audio. Speech. Music. Process. (1), 25 (2015)","DOI":"10.1186\/s13636-015-0068-3"},{"key":"161_CR27","doi-asserted-by":"crossref","unstructured":"Y. Zhang, M. Pezeshki, P. Brakel, S. Zhang, C. L. Y. Bengio, and A. Courville, Towards end-to-end speech recognition with deep convolutional neural networks, arXiv preprint arXiv: 1701.02720, 2017","DOI":"10.21437\/Interspeech.2016-1446"},{"key":"161_CR28","doi-asserted-by":"crossref","unstructured":"A. Graves, A. R. Mohamed, and G. Hinton, Speech recognition with deep recurrent neural networks, in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6645\u20136649, 2013","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"161_CR29","doi-asserted-by":"crossref","unstructured":"L. Lu, L. Kong, C. Dyer, N. A. Smith, Multi-task learning with CTC and segmental CRF for speech recognition, arXiv preprint arXiv: 1702.06378, 2017","DOI":"10.21437\/Interspeech.2017-71"},{"key":"161_CR30","unstructured":"R. Pascanu, T. Mikolov, and Y. Bengio, On the difficulty of training recurrent neural networks, arXiv preprint arXiv: 1211.5063, 2012"},{"key":"161_CR31","doi-asserted-by":"crossref","unstructured":"T. Hori, J. Cho, and S. Watanabe, End-to-end speech recognition with word-based RNN language models, arXiv preprint arXiv:1808.02608, 2018","DOI":"10.1109\/SLT.2018.8639693"},{"key":"161_CR32","unstructured":"W. Chan, and I. Lane, Deep recurrent neural networks for acoustic modelling, arXiv preprint arXiv:1504.01482(2015)"},{"key":"161_CR33","doi-asserted-by":"crossref","unstructured":"D. Palaz, M. M. Doss, and R. Collobert, Convolutional neural networks-based continuous speech recognition using raw speech signal, Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on. IEEE, 2015","DOI":"10.1109\/ICASSP.2015.7178781"},{"key":"161_CR34","unstructured":"D. Amodei, R. Anubhai, E. Battenberg, et al, Deep Speech 2: end-to-end speech recognition in English and Mandarin, Computer Science, 2015"},{"key":"161_CR35","doi-asserted-by":"crossref","unstructured":"T. Hori, S. Watanabe, and J. R. Hershey, Multi-level language modeling and decoding for open vocabulary end-to-end speech recognition, Automatic Speech Recognition and Understanding Workshop (ASRU), 2017 IEEE. IEEE, 2017","DOI":"10.1109\/ASRU.2017.8268948"},{"key":"161_CR36","doi-asserted-by":"crossref","unstructured":"J. Chorowski, and N. Jaitly, Towards better decoding and language model integration in sequence to sequence models, arXiv preprint arXiv:1612.02695, 2016","DOI":"10.21437\/Interspeech.2017-343"},{"key":"161_CR37","doi-asserted-by":"crossref","unstructured":"A. Kannan, Y. Wu, P. Nguyen, T. N. Sainath, Z. Chen, and R. Prabhavalkar, An analysis of incorporating an external language model into a sequence-to-sequence model, In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018","DOI":"10.1109\/ICASSP.2018.8462682"},{"key":"161_CR38","doi-asserted-by":"crossref","unstructured":"D. Bahdanau, J. Chorowski, D. Serdyuk, et al., End-to-end attention-based large vocabulary speech recognition, Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on. IEEE, 2016: 4945\u20134949","DOI":"10.1109\/ICASSP.2016.7472618"},{"key":"161_CR39","doi-asserted-by":"crossref","unstructured":"R. Sennrich, B. Haddow, and A. Birch, Neural machine translation of rare words with subword units, in ACL, Berlin, 2016, pp. 1715\u20131725","DOI":"10.18653\/v1\/P16-1162"},{"key":"161_CR40","unstructured":"K. J. Han, A. Chandrashekaran, J. Kim, & I. Lane, The CAPIO 2017 conversational speech recognition system. arXiv preprint arXiv:1801.00059, 2017"},{"key":"161_CR41","doi-asserted-by":"crossref","unstructured":"D. Povey, G. Cheng, Y. Wang, K. Li, H. Xu, M. Yarmohamadi, and S. Khudanpur, Semi-orthogonal low-rank matrix factorization for deep neural networks. In Proceedings of the 19th Annual Conference of the International Speech Communication Association (INTERSPEECH 2018), Hyderabad, India","DOI":"10.21437\/Interspeech.2018-1417"},{"key":"161_CR42","unstructured":"N. Zeghidour, Q. Xu, V. Liptchinsky, N. Usunier, G. Synnaeve, and R. Collobert. Fully convolutional speech recognition. arXiv preprint arXiv:1812.06864, 2018"},{"key":"161_CR43","doi-asserted-by":"crossref","unstructured":"Y. Tang, G. Ding, J. Huang, X. He, & B. Zhou, Deep speaker embedding learning with multi-level pooling for text-independent speaker verification, arXiv preprint arXiv:1902.07821, 2019","DOI":"10.1109\/ICASSP.2019.8682712"}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-019-0161-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/s13636-019-0161-0\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-019-0161-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2020,10,27]],"date-time":"2020-10-27T00:08:12Z","timestamp":1603757292000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-019-0161-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,10,28]]},"references-count":43,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2019,12]]}},"alternative-id":["161"],"URL":"https:\/\/doi.org\/10.1186\/s13636-019-0161-0","relation":{},"ISSN":["1687-4722"],"issn-type":[{"type":"electronic","value":"1687-4722"}],"subject":[],"published":{"date-parts":[[2019,10,28]]},"assertion":[{"value":"2 April 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 September 2019","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 October 2019","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare that they have no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"18"}}