{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T14:31:03Z","timestamp":1784644263548,"version":"3.55.0"},"reference-count":65,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2019,8,7]],"date-time":"2019-08-07T00:00:00Z","timestamp":1565136000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Automatic speech recognition, especially large vocabulary continuous speech recognition, is an important issue in the field of machine learning. For a long time, the hidden Markov model (HMM)-Gaussian mixed model (GMM) has been the mainstream speech recognition framework. But recently, HMM-deep neural network (DNN) model and the end-to-end model using deep learning has achieved performance beyond HMM-GMM. Both using deep learning techniques, these two models have comparable performances. However, the HMM-DNN model itself is limited by various unfavorable factors such as data forced segmentation alignment, independent hypothesis, and multi-module individual training inherited from HMM, while the end-to-end model has a simplified model, joint training, direct output, no need to force data alignment and other advantages. Therefore, the end-to-end model is an important research direction of speech recognition. In this paper we review the development of end-to-end model. This paper first introduces the basic ideas, advantages and disadvantages of HMM-based model and end-to-end models, and points out that end-to-end model is the development direction of speech recognition. Then the article focuses on the principles, progress and research hotspots of three different end-to-end models, which are connectionist temporal classification (CTC)-based, recurrent neural network (RNN)-transducer and attention-based, and makes theoretically and experimentally detailed comparisons. Their respective advantages and disadvantages and the possible future development of the end-to-end model are finally pointed out. Automatic speech recognition is a pattern recognition task in the field of computer science, which is a subject area of Symmetry.<\/jats:p>","DOI":"10.3390\/sym11081018","type":"journal-article","created":{"date-parts":[[2019,8,7]],"date-time":"2019-08-07T10:56:38Z","timestamp":1565175398000},"page":"1018","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":205,"title":["An Overview of End-to-End Automatic Speech Recognition"],"prefix":"10.3390","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6992-7950","authenticated-orcid":false,"given":"Dong","family":"Wang","sequence":"first","affiliation":[{"name":"Science and Technology on Parallel and Distributed Processing Laboratory, National University of Defense Technology, Changsha 410073, China"},{"name":"College of Computer, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaodong","family":"Wang","sequence":"additional","affiliation":[{"name":"Science and Technology on Parallel and Distributed Processing Laboratory, National University of Defense Technology, Changsha 410073, China"},{"name":"College of Computer, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shaohe","family":"Lv","sequence":"additional","affiliation":[{"name":"Science and Technology on Parallel and Distributed Processing Laboratory, National University of Defense Technology, Changsha 410073, China"},{"name":"College of Computer, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2019,8,7]]},"reference":[{"key":"ref_1","first-page":"129","article-title":"Markovian models for sequential data","volume":"2","author":"Bengio","year":"1999","journal-title":"Neural Comput. Surv."},{"key":"ref_2","first-page":"15","article-title":"Automatic Speech Recognition Technology Review","volume":"49","author":"Chengyou","year":"1996","journal-title":"Acoust. Electr. Eng."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"637","DOI":"10.1121\/1.1906946","article-title":"Automatic recognition of spoken digits","volume":"24","author":"Davis","year":"1952","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1072","DOI":"10.1121\/1.1908561","article-title":"Phonetic typewriter","volume":"28","author":"Olson","year":"1956","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1480","DOI":"10.1121\/1.1907653","article-title":"Results obtained from a vowel recognition computer program","volume":"31","author":"Forgie","year":"1959","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_6","first-page":"67","article-title":"Automatic speech recognition\u2013a brief history of the technology development","volume":"1","author":"Juang","year":"2005","journal-title":"Georgia Inst. Technol."},{"key":"ref_7","first-page":"193","article-title":"Recognition of Japanese vowels\u2014Preliminary to the recognition of speech","volume":"37","author":"Suzuki","year":"1961","journal-title":"J. Radio Res. Lab."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1664","DOI":"10.1121\/1.1936652","article-title":"Phonetic Typewriter","volume":"33","author":"Sakai","year":"1961","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_9","unstructured":"Nagata, K., Kato, Y., and Chiba, S. (2018, August 04). Spoken digit recognizer for Japanese language. Audio Engineering Society Convention 16. Audio Engineering Society. Available online: http:\/\/www.aes.org\/e-lib\/browse.cfm?elib=603."},{"key":"ref_10","first-page":"36","article-title":"A statistical method for estimation of speech spectral density and formant frequencies","volume":"53","author":"Itakura","year":"1970","journal-title":"Electr. Commun. Jpn. A"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"52","DOI":"10.1007\/BF01074755","article-title":"Speech discrimination by dynamic programming","volume":"4","author":"Vintsyuk","year":"1968","journal-title":"Cybern. Syst. Anal."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"637","DOI":"10.1121\/1.1912679","article-title":"Speech analysis and synthesis by linear prediction of the speech wave","volume":"50","author":"Atal","year":"1971","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"375","DOI":"10.1016\/0167-6393(88)90053-2","article-title":"On large-vocabulary speaker-independent continuous speech recognition","volume":"7","author":"Lee","year":"1988","journal-title":"Speech Commun."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"73081","DOI":"10.1109\/ACCESS.2018.2881119","article-title":"Joint Implicit and Explicit Neural Networks for Question Recommendation in CQA Services","volume":"6","author":"Tu","year":"2018","journal-title":"IEEE Access"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"667","DOI":"10.1631\/FITEE.1500389","article-title":"Exploiting a depth context model in visual tracking with correlation filter","volume":"18","author":"Chen","year":"2017","journal-title":"Front. Inf. Technol. Electr. Eng."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"30","DOI":"10.1109\/TASL.2011.2134090","article-title":"Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition","volume":"20","author":"Dahl","year":"2011","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Rao, K., Sak, H., and Prabhavalkar, R. (2017, January 16\u201320). Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer. Proceedings of the 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Okinawa, Japan.","DOI":"10.1109\/ASRU.2017.8268935"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Lu, L., Zhang, X., Cho, K., and Renals, S. (2015, January 6\u201310). A study of the recurrent neural network encoder-decoder for large vocabulary speech recognition. Proceedings of the Sixteenth Annual Conference of the International Speech Communication Association, Dresden, Germany.","DOI":"10.21437\/Interspeech.2015-654"},{"key":"ref_19","unstructured":"Hannun, A., Case, C., Casper, J., Catanzaro, B., Diamos, G., Elsen, E., Prenger, R., Satheesh, S., Sengupta, S., and Coates, A. (2014). DeepSpeech: Scaling up end-to-end speech recognition. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Miao, Y., Gowayyed, M., and Metze, F. (2015, January 13\u201317). EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding. Proceedings of the 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), Scottsdale, AZ, USA.","DOI":"10.1109\/ASRU.2015.7404790"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Pezeshki, M., Brakel, P., Zhang, S., Laurent, C., Bengio, Y., and Courville, A. (2017). Towards End-to-End Speech Recognition with Deep Convolutional Neural Networks. arXiv.","DOI":"10.21437\/Interspeech.2016-1446"},{"key":"ref_22","unstructured":"Graves, A., and Jaitly, N. (June, January 21). Towards end-to-end speech recognition with recurrent neural networks. Proceedings of the International Conference on Machine Learning, Beijing, China."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Hori, T., Watanabe, S., Zhang, Y., and Chan, W. (2017). Advances in Joint CTC-Attention Based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM. arXiv, 949\u2013953.","DOI":"10.21437\/Interspeech.2017-1296"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Kim, S., Hori, T., and Watanabe, S. (2017, January 5\u20139). Joint CTC-attention based end-to-end speech recognition using multi-task learning. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7953075"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Graves, A., Fern\u00e1ndez, S., Gomez, F., and Schmidhuber, J. (2006, January 25\u201329). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. Proceedings of the 23rd international conference on Machine learning, Pittsburgh, PA, USA.","DOI":"10.1145\/1143844.1143891"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Eyben, F., W\u00f6llmer, M., Schuller, B., and Graves, A. (2009, January 13\u201317). From speech to letters-using a novel neural network architecture for grapheme based asr. Proceedings of the 2009 IEEE Workshop on Automatic Speech Recognition & Understanding, Merano\/Meran, Italy.","DOI":"10.1109\/ASRU.2009.5373257"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Li, J., Zhang, H., Cai, X., and Xu, B. (2015, January 6\u201310). Towards end-to-end speech recognition for Chinese Mandarin using long short-term memory recurrent neural networks. Proceedings of the Sixteenth Annual Conference of the International Speech Communication Association, Dresden, Germany.","DOI":"10.21437\/Interspeech.2015-717"},{"key":"ref_28","unstructured":"Hannun, A.Y., Maas, A.L., Jurafsky, D., and Ng, A.Y. (2014). First-pass large vocabulary continuous speech recognition using bi-directional recurrent DNNs. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Maas, A., Xie, Z., Jurafsky, D., and Ng, A. (June, January 31). Lexicon-free conversational speech recognition with neural networks. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Denver, CO, USA.","DOI":"10.3115\/v1\/N15-1038"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Sak, H., Senior, A.W., Rao, K., and Beaufays, F. (2015). Fast and accurate recurrent neural network acoustic models for speech recognition. arXiv, 1468\u20131472.","DOI":"10.21437\/Interspeech.2015-350"},{"key":"ref_31","unstructured":"Song, W., and Cai, J. (2015). End-to-end deep neural network for automatic speech recognition. Standford CS224D Rep., Available online: https:\/\/cs224d.stanford.edu\/reports\/SongWilliam.pdf."},{"key":"ref_32","unstructured":"Amodei, D., Ananthanarayanan, S., Anubhai, R., Bai, J., Battenberg, E., Case, C., Casper, J., Catanzaro, B., Cheng, Q., and Chen, G. (2016, January 19\u201324). Deep speech 2: End-to-end speech recognition in english and mandarin. Proceedings of the International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Soltau, H., Liao, H., and Sak, H. (2017). Neural Speech Recognizer: Acoustic-to-Word LSTM Model for Large Vocabulary Speech Recognition. arXiv, 3707\u20133711.","DOI":"10.21437\/Interspeech.2017-1566"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zweig, G., Yu, C., Droppo, J., and Stolcke, A. (2017, January 5\u20139). Advances in all-neural speech recognition. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7953069"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Audhkhasi, K., Ramabhadran, B., Saon, G., Picheny, M., and Nahamoo, D. (2017). Direct Acoustics-to-Word Models for English Conversational Speech Recognition. arXiv, 959\u2013963.","DOI":"10.21437\/Interspeech.2017-546"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Li, J., Ye, G., Zhao, R., Droppo, J., and Gong, Y. (2017, January 16\u201320). Acoustic-to-word model without OOV. Proceedings of the 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Okinawa, Japan.","DOI":"10.1109\/ASRU.2017.8268924"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Audhkhasi, K., Kingsbury, B., Ramabhadran, B., Saon, G., and Picheny, M. (2018, January 15\u201320). Building Competitive Direct Acoustics-to-Word Models for English Conversational Speech Recognition. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461935"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Prabhavalkar, R., Rao, K., Sainath, T.N., Li, B., Johnson, L., and Jaitly, N. (2017, January 20\u201324). A comparison of sequence-to-sequence models for speech recognition. Proceedings of the Interspeech, Stockholm, Sweden.","DOI":"10.21437\/Interspeech.2017-233"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Hwang, K., and Sung, W. (2016, January 20\u201325). Character-level incremental speech recognition with recurrent neural networks. Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472696"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Wang, Y., Zhang, L., Zhang, B., and Li, Z. (2018, January 8\u201310). End-to-End Mandarin Recognition based on Convolution Input. Proceedings of the MATEC Web of Conferences, Lille, France.","DOI":"10.1051\/matecconf\/201821401004"},{"key":"ref_41","first-page":"235","article-title":"Sequence Transduction with Recurrent Neural Networks","volume":"58","author":"Graves","year":"2012","journal-title":"Comput. Sci."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"351","DOI":"10.1016\/0167-6393(90)90010-7","article-title":"Speech database development at MIT: TIMIT and beyond","volume":"9","author":"Zue","year":"1990","journal-title":"Speech Commun."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Graves, A., Mohamed, A.r., and Hinton, G. (2013, January 26\u201331). Speech recognition with deep recurrent neural networks. Proceedings of the 2013 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Sak, H., Shannon, M., Rao, K., and Beaufays, F. (2017, January 20\u201324). Recurrent Neural Aligner: An Encoder-Decoder Neural Network Model for Sequence to Sequence Mapping. Proceedings of the INTERSPEECH, ISCA, Stockholm, Sweden.","DOI":"10.21437\/Interspeech.2017-1705"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Dong, L., Zhou, S., Chen, W., and Xu, B. (2018). Extending Recurrent Neural Aligner for Streaming End-to-End Speech Recognition in Mandarin. arXiv, 816\u2013820.","DOI":"10.21437\/Interspeech.2018-1086"},{"key":"ref_46","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv."},{"key":"ref_47","unstructured":"Chorowski, J., Bahdanau, D., Cho, K., and Bengio, Y. (2014, January 12). End-to-end continuous speech recognition using attention-based recurrent nn: First results. Proceedings of the NIPS 2014 Workshop on Deep Learning, Montreal, QC, Canada."},{"key":"ref_48","unstructured":"Cortes, C., Lawrence, N.D., Lee, D.D., Sugiyama, M., and Garnett, R. (2015). Attention-Based Models for Speech Recognition. Advances in Neural Information Processing Systems 28, Curran Associates, Inc."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Bahdanau, D., Chorowski, J., Serdyuk, D., Brakel, P., and Bengio, Y. (2016, January 20\u201325). End-to-end attention-based large vocabulary speech recognition. Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472618"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Cho, K., Van Merri\u00ebnboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Chan, W., Jaitly, N., Le, Q., and Vinyals, O. (2016, January 20\u201325). Listen, attend and spell: A neural network for large vocabulary conversational speech recognition. Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472621"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Chan, W., and Lane, I. (2016, January 8\u201312). On Online Attention-Based Speech Recognition and Joint Mandarin Character-Pinyin Training. Proceedings of the Interspeech, San Francisco, CA, USA.","DOI":"10.21437\/Interspeech.2016-334"},{"key":"ref_53","unstructured":"Chan, W., Zhang, Y., Le, Q.V., and Jaitly, N. (2017, January 24\u201326). Latent Sequence Decompositions. Proceedings of the 5th International Conference on Learning Representations, ICLR 2017, Toulon, France."},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"1240","DOI":"10.1109\/JSTSP.2017.2763455","article-title":"Hybrid CTC\/attention architecture for end-to-end speech recognition","volume":"11","author":"Watanabe","year":"2017","journal-title":"IEEE J. Sel. Top. Signal Process."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Hayashi, T., Watanabe, S., Toda, T., and Takeda, K. (2018). Multi-Head Decoder for End-to-End Speech Recognition. arXiv.","DOI":"10.21437\/Interspeech.2018-1655"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Chan, W., and Jaitly, N. (2017, January 5\u20139). Very deep convolutional networks for end-to-end speech recognition. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7953077"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Chorowski, J., and Jaitly, N. (2016). Towards better decoding and language model integration in sequence to sequence models. arXiv.","DOI":"10.21437\/Interspeech.2017-343"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Lu, L., Zhang, X., and Renais, S. (2016, January 20\u201325). On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition. Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472641"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Chiu, C., Sainath, T.N., Wu, Y., Prabhavalkar, R., Nguyen, P., Chen, Z., Kannan, A., Weiss, R.J., Rao, K., and Gonina, E. (2018, January 15\u201320). State-of-the-Art Speech Recognition with Sequence-to-Sequence Models. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8462105"},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Weng, C., Cui, J., Wang, G., Wang, J., Yu, C., Su, D., and Yu, D. (2018, January 2\u20136). Improving Attention Based Sequence-to-Sequence Models for End-to-End English Conversational Speech Recognition. Proceedings of the Interspeech ISCA, Hyderabad, Indian.","DOI":"10.21437\/Interspeech.2018-1030"},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Hou, J., Zhang, S., and Dai, L. (2017, January 20\u201324). Gaussian Prediction Based Attention for Online End-to-End Speech Recognition. Proceedings of the INTERSPEECH, ISCA, Stockholm, Sweden.","DOI":"10.21437\/Interspeech.2017-751"},{"key":"ref_62","doi-asserted-by":"crossref","unstructured":"Prabhavalkar, R., Sainath, T.N., Wu, Y., Nguyen, P., Chen, Z., Chiu, C., and Kannan, A. (2018, January 15\u201320). Minimum Word Error Rate Training for Attention-Based Sequence-to-Sequence Models. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461809"},{"key":"ref_63","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"Kingsbury, B. (2009, January 19\u201324). Lattice-based optimization of sequence classification criteria for neural-network acoustic modeling. Proceedings of the 2009 IEEE International Conference on Acoustics, Speech and Signal Processing, Taipei, Taiwan.","DOI":"10.1109\/ICASSP.2009.4960445"},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Battenberg, E., Chen, J., Child, R., Coates, A., Li, Y.G.Y., Liu, H., Satheesh, S., Sriram, A., and Zhu, Z. (2017, January 16\u201320). Exploring neural transducers for end-to-end speech recognition. Proceedings of the 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Okinawa, Japan.","DOI":"10.1109\/ASRU.2017.8268937"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/11\/8\/1018\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:09:21Z","timestamp":1760188161000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/11\/8\/1018"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,8,7]]},"references-count":65,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2019,8]]}},"alternative-id":["sym11081018"],"URL":"https:\/\/doi.org\/10.3390\/sym11081018","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,8,7]]}}}