{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T15:27:43Z","timestamp":1777649263947,"version":"3.51.4"},"reference-count":29,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2019,5,7]],"date-time":"2019-05-07T00:00:00Z","timestamp":1557187200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Science and Technology on Parallel and Distributed Processing Laboratory Fund","award":["6142110180405"],"award-info":[{"award-number":["6142110180405"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Since conventional Automatic Speech Recognition (ASR) systems often contain many modules and use varieties of expertise, it is hard to build and train such models. Recent research show that end-to-end ASRs can significantly simplify the speech recognition pipelines and achieve competitive performance with conventional systems. However, most end-to-end ASR systems are neither reproducible nor comparable because they use specific language models and in-house training databases which are not freely available. This is especially common for Mandarin speech recognition. In this paper, we propose a CNN+BLSTM+CTC end-to-end Mandarin ASR. This CNN+BLSTM+CTC ASR uses Convolutional Neural Net (CNN) to learn local speech features, uses Bidirectional Long-Short Time Memory (BLSTM) to learn history and future contextual information, and uses Connectionist Temporal Classification (CTC) for decoding. Our model is completely trained on the by-far-largest open-source Mandarin speech corpus AISHELL-1, using neither any in-house databases nor external language models. Experiments show that our CNN+BLSTM+CTC model achieves a WER of 19.2%, outperforming the exiting best work. Because all the data corpora we used are freely available, our model is reproducible and comparable, providing a new baseline for further Mandarin ASR research.<\/jats:p>","DOI":"10.3390\/sym11050644","type":"journal-article","created":{"date-parts":[[2019,5,9]],"date-time":"2019-05-09T11:22:35Z","timestamp":1557400955000},"page":"644","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":35,"title":["End-to-End Mandarin Speech Recognition Combining CNN and BLSTM"],"prefix":"10.3390","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6992-7950","authenticated-orcid":false,"given":"Dong","family":"Wang","sequence":"first","affiliation":[{"name":"College of Computer, National University of Defense Technology, Changsha 410073, China"},{"name":"Science and Technology on Parallel and Distributed Processing Laboratory, National University of Defense Technology, Changsha 410073, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaodong","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Computer, National University of Defense Technology, Changsha 410073, China"},{"name":"Science and Technology on Parallel and Distributed Processing Laboratory, National University of Defense Technology, Changsha 410073, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shaohe","family":"Lv","sequence":"additional","affiliation":[{"name":"College of Computer, National University of Defense Technology, Changsha 410073, China"},{"name":"Science and Technology on Parallel and Distributed Processing Laboratory, National University of Defense Technology, Changsha 410073, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,5,7]]},"reference":[{"key":"ref_1","unstructured":"Hannun, A., Case, C., Casper, J., Catanzaro, B., Diamos, G., Elsen, E., Prenger, R., Satheesh, S., Sengupta, S., and Coates, A. (2014). DeepSpeech: Scaling up end-to-end speech recognition. arXiv."},{"key":"ref_2","unstructured":"Graves, A., and Jaitly, N. (2014, January 21\u201326). Towards end-to-end speech recognition with recurrent neural networks. Proceedings of the International Conference on Machine Learning, Beijing, China."},{"key":"ref_3","unstructured":"Rousseau, A., Del\u00e9glise, P., and Est\u00e8ve, Y. (2012, January 23\u201325). TED-LIUM: An Automatic Speech Recognition dedicated corpus. Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC-2012), Istanbul, Turkey."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. (2015, January 19\u201324). Librispeech: An ASR corpus based on public domain audio books. Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, Australia.","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Bu, H., Du, J., Na, X., Wu, B., and Zheng, H. (2017, January 1\u20133). AIShell-1: An open-source Mandarin speech corpus and a speech recognition baseline. Proceedings of the 2017 20th Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I\/O Systems and Assessment (O-COCOSDA), Seoul, Korea.","DOI":"10.1109\/ICSDA.2017.8384449"},{"key":"ref_6","unstructured":"Wang, Y., Zhang, L., Zhang, B., and Li, Z. (2018, January 27\u201329). End-to-End Mandarin Recognition based on Convolution Input. Proceedings of the 2018 2nd International Conference on Information Processing and Control Engineering (ICIPCE 2018), Shanghai, China."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Li, M., and Liu, M. (2018). End-to-end speech recognition with adaptive computation steps. arXiv.","DOI":"10.1109\/ICASSP.2019.8682500"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Graves, A., Fern\u00e1ndez, S., Gomez, F., and Schmidhuber, J. (2006, January 25\u201329). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. Proceedings of the 23rd International Conference on Machine Learning, Pittsburgh, PA, USA.","DOI":"10.1145\/1143844.1143891"},{"key":"ref_9","unstructured":"Hannun, A.Y., Maas, A.L., Jurafsky, D., and Ng, A.Y. (2014). First-pass large vocabulary continuous speech recognition using bi-directional recurrent DNNs. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Maas, A., Xie, Z., Jurafsky, D., and Ng, A. (June, January 31). Lexicon-free conversational speech recognition with neural networks. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Denver, CO, USA.","DOI":"10.3115\/v1\/N15-1038"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Sak, H., Senior, A., Rao, K., and Beaufays, F. (2015). Fast and Accurate Recurrent Neural Network Acoustic Models for Speech Recognition. arXiv.","DOI":"10.21437\/Interspeech.2015-350"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Soltau, H., Liao, H., and Sak, H. (2017, January 20\u201324). Neural Speech Recognizer: Acoustic-to-Word LSTM Model for Large Vocabulary Speech Recognition. Proceedings of the Interspeech 2017, Stockholm, Sweden.","DOI":"10.21437\/Interspeech.2017-1566"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Audhkhasi, K., Ramabhadran, B., Saon, G., Picheny, M., and Nahamoo, D. (2017). Direct Acoustics-to-Word Models for English Conversational Speech Recognition. arXiv.","DOI":"10.21437\/Interspeech.2017-546"},{"key":"ref_14","unstructured":"Amodei, D., Ananthanarayanan, S., Anubhai, R., Bai, J., Battenberg, E., Case, C., Casper, J., Catanzaro, B., Cheng, Q., and Chen, G. (2016, January 19\u201324). Deep speech 2: End-to-end speech recognition in english and mandarin. Proceedings of the International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_15","unstructured":"Li, A., Yin, Z., Wang, T., Fang, Q., and Hu, F. (2004, January 17\u201319). RASC863-A Chinese speech corpus with four regional accents. Proceedings of the ICSLT-o-COCOSDA, New Delhi, India."},{"key":"ref_16","unstructured":"Wang, D., and Zhang, X. (2015). THCHS-30: A free Chinese speech corpus. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Wang, D., Tang, Z., Tang, D., and Chen, Q. (2016, January 26\u201328). OC16-CE80: A Chinese-English mixlingual database and a speech recognition baseline. Proceedings of the 2016 Conference of The Oriental Chapter of International Committee for Coordination and Standardization of Speech Databases and Assessment Techniques (O-COCOSDA), Bali, Indonesia.","DOI":"10.1109\/ICSDA.2016.7918989"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"501","DOI":"10.1109\/TASLP.2017.2782360","article-title":"Multitask Learning for Phone Recognition of Underresourced Languages Using Mismatched Transcription","volume":"26","author":"Chen","year":"2018","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhou, J., Jiang, T., Li, L., Hong, Q., Wang, Z., and Xia, B. (2018). Training Multi-Task Adversarial Network for Extracting Noise-Robust Speaker Embedding. arXiv.","DOI":"10.1109\/ICASSP.2019.8683828"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Tu, M., Grabek, A., Liss, J., and Berisha, V. (2018, January 2\u20136). Investigating the Role of L1 in Automatic Pronunciation Evaluation of L2 Speech. Proceedings of the Interspeech 2018, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-1350"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Zhang, P., and Yan, Y. (2018, January 2\u20136). Improving Language Modeling with an Adversarial Critic for Automatic Speech Recognition. Proceedings of the Interspeech 2018, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-1111"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Lugosch, L., and Tomar, V.S. (2018, January 2\u20136). Tone Recognition Using Lifters and CTC. Proceedings of the Interspeech 2018, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-2293"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Li, J., Wang, X., Zhao, Y., and Li, Y. (2018, January 2\u20136). Gated Recurrent Unit Based Acoustic Modeling with Future Context. Proceedings of the Interspeech 2018, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-1544"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Li, J., Shan, Y., Wang, X., and Li, Y. (2018). Improving Gated Recurrent Unit Based Acoustic Modeling with Batch Normalization and Enlarged Context. arXiv.","DOI":"10.21437\/Interspeech.2018-1544"},{"key":"ref_25","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 6\u201311). Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_26","unstructured":"Glorot, X., Bordes, A., and Bengio, Y. (2011, January 11\u201313). Deep sparse rectifier neural networks. Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, Fort Lauderdale, FL, USA."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Pezeshki, M., Brakel, P., Zhang, S., Laurent, C., Bengio, Y., and Courville, A. (2016, January 8\u201312). Towards End-to-End Speech Recognition with Deep Convolutional Neural Networks. Proceedings of the Interspeech 2016, San Francisco, CA, USA.","DOI":"10.21437\/Interspeech.2016-1446"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Battenberg, E., Chen, J., Child, R., Coates, A., Li, Y.G.Y., Liu, H., Satheesh, S., Sriram, A., and Zhu, Z. (2017, January 16\u201320). Exploring neural transducers for end-to-end speech recognition. Proceedings of the 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Okinawa, Japan.","DOI":"10.1109\/ASRU.2017.8268937"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Chang, E., Zhou, J., Di, S., Huang, C., and Lee, K.F. (2000, January 16\u201320). Large vocabulary Mandarin speech recognition with different approaches in modeling tones. Proceedings of the Sixth International Conference on Spoken Language Processing, Beijing, China.","DOI":"10.21437\/ICSLP.2000-436"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/11\/5\/644\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:49:51Z","timestamp":1760186991000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/11\/5\/644"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,5,7]]},"references-count":29,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2019,5]]}},"alternative-id":["sym11050644"],"URL":"https:\/\/doi.org\/10.3390\/sym11050644","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,5,7]]}}}