{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T04:34:13Z","timestamp":1779338053289,"version":"3.51.4"},"reference-count":39,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2022,1,29]],"date-time":"2022-01-29T00:00:00Z","timestamp":1643414400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Foundation of Fujian Key Laboratory of Automotive Electronics and Electric Drive (Fujian University of Technology)","award":["KF-X18002"],"award-info":[{"award-number":["KF-X18002"]}]},{"DOI":"10.13039\/501100001809","name":"National Science Foundation of China","doi-asserted-by":"publisher","award":["41971340, 41471333"],"award-info":[{"award-number":["41971340, 41471333"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100005270","name":"Fujian Provincial Department of Science and Technology","doi-asserted-by":"publisher","award":["2021Y4019,2020D002, 2020L3014, 2019I0019"],"award-info":[{"award-number":["2021Y4019,2020D002, 2020L3014, 2019I0019"]}],"id":[{"id":"10.13039\/501100005270","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Transformers have become popular in building end-to-end automatic speech recognition (ASR) systems. However, transformer ASR systems are usually trained to give output sequences in the left-to-right order, disregarding the right-to-left context. Currently, the existing transformer-based ASR systems that employ two decoders for bidirectional decoding are complex in terms of computation and optimization. The existing ASR transformer with a single decoder for bidirectional decoding requires extra methods (such as a self-mask) to resolve the problem of information leakage in the attention mechanism This paper explores different options for the development of a speech transformer that utilizes a single decoder equipped with bidirectional context embedding (BCE) for bidirectional decoding. The decoding direction, which is set up at the input level, enables the model to attend to different directional contexts without extra decoders and also alleviates any information leakage. The effectiveness of this method was verified with a bidirectional beam search method that generates bidirectional output sequences and determines the best hypothesis according to the output score. We achieved a word error rate (WER) of 7.65%\/18.97% on the clean\/other LibriSpeech test set, outperforming the left-to-right decoding style in our work by 3.17%\/3.47%. The results are also close to, or better than, other state-of-the-art end-to-end models.<\/jats:p>","DOI":"10.3390\/info13020069","type":"journal-article","created":{"date-parts":[[2022,1,29]],"date-time":"2022-01-29T23:02:06Z","timestamp":1643497326000},"page":"69","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["A Bidirectional Context Embedding Transformer for Automatic Speech Recognition"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5337-9083","authenticated-orcid":false,"given":"Lyuchao","family":"Liao","sequence":"first","affiliation":[{"name":"Fujian Key Laboratory of Automotive Electronics and Electric Drive, Fujian University of Technology, Fuzhou 350118, China"},{"name":"Fujian Provincial Universities Engineering Research Center for Intelligent Driving Technology, Fujian University of Technology, Fuzhou 350118, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2530-1067","authenticated-orcid":false,"given":"Francis","family":"Afedzie Kwofie","sequence":"additional","affiliation":[{"name":"Fujian Key Laboratory of Automotive Electronics and Electric Drive, Fujian University of Technology, Fuzhou 350118, China"},{"name":"Fujian Provincial Universities Engineering Research Center for Intelligent Driving Technology, Fujian University of Technology, Fuzhou 350118, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhifeng","family":"Chen","sequence":"additional","affiliation":[{"name":"Fujian Key Laboratory of Automotive Electronics and Electric Drive, Fujian University of Technology, Fuzhou 350118, China"},{"name":"Fujian Provincial Universities Engineering Research Center for Intelligent Driving Technology, Fujian University of Technology, Fuzhou 350118, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6921-7369","authenticated-orcid":false,"given":"Guangjie","family":"Han","sequence":"additional","affiliation":[{"name":"Fujian Provincial Universities Engineering Research Center for Intelligent Driving Technology, Fujian University of Technology, Fuzhou 350118, China"},{"name":"College of Internet of Things Engineering, Hohai University, Changzhou 213022, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongqiang","family":"Wang","sequence":"additional","affiliation":[{"name":"Fujian Key Laboratory of Automotive Electronics and Electric Drive, Fujian University of Technology, Fuzhou 350118, China"},{"name":"Fujian Provincial Universities Engineering Research Center for Intelligent Driving Technology, Fujian University of Technology, Fuzhou 350118, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuyuan","family":"Lin","sequence":"additional","affiliation":[{"name":"Fujian Key Laboratory of Automotive Electronics and Electric Drive, Fujian University of Technology, Fuzhou 350118, China"},{"name":"Fujian Provincial Universities Engineering Research Center for Intelligent Driving Technology, Fujian University of Technology, Fuzhou 350118, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dongmei","family":"Hu","sequence":"additional","affiliation":[{"name":"Fujian Key Laboratory of Automotive Electronics and Electric Drive, Fujian University of Technology, Fuzhou 350118, China"},{"name":"College of Environmental Science and Engineering, North China Electric Power University, Beijing 102206, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,1,29]]},"reference":[{"key":"ref_1","first-page":"749","article-title":"A real-time end-to-end multilingual speech recognition architecture","volume":"9","author":"Eustis","year":"2014","journal-title":"IEEE J. Sel. Top. Signal Processing"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Bosch, L.T., Boves, L., and Ernestus, M. (2013, January 25\u201329). Towards an end-to-end computational model of speech comprehension: Simulating a lexical decision task. Proceedings of the INTERSPEECH, Lyon, France.","DOI":"10.21437\/Interspeech.2013-645"},{"key":"ref_3","unstructured":"Chorowski, J., Bahdanau, D., Cho, K., and Bengio, Y. (2014). End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results. arXiv."},{"key":"ref_4","unstructured":"Chan, W., Jaitly, N., Le, Q.V., and Vinyals, O. (2015). Listen, Attend and Spell. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Emiru, E.D., Xiong, S., Li, Y., Fesseha, A., and Diallo, M. (2021). Improving Amharic Speech Recognition System Using Connectionist Temporal Classification with Attention Model and Phoneme-Based Byte-Pair-Encodings. Information, 12.","DOI":"10.3390\/info12020062"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Wang, X., and Zhao, C. (2021). A 2D Convolutional Gating Mechanism for Mandarin Streaming Speech Recognition. Information, 12.","DOI":"10.3390\/info12040165"},{"key":"ref_7","unstructured":"Vaswani, A., Shazeer, N.M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017). Attention is All you Need. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zhou, S., Dong, L., Xu, S., and Xu, B. (2018). Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin Chinese. arXiv.","DOI":"10.21437\/Interspeech.2018-1107"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zhang, Q., Lu, H., Sak, H., Tripathi, A., McDermott, E., Koo, S., and Kumar, S. (2020, January 4\u20138). Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9053896"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Karita, S., Yalta, N., Watanabe, S., Delcroix, M., Ogawa, A., and Nakatani, T. (2019, January 15\u201319). Improving Transformer-Based End-to-End Speech Recognition with Connectionist Temporal Classification and Language Model Integration. Proceedings of the INTERSPEECH, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-1938"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Miao, H., Cheng, G., Gao, C., Zhang, P., and Yan, Y. (2020, January 4\u20138). Transformer-based online CTC\/attention end-to-end speech recognition architecture. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9053165"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Chen, X., Zhang, S., Song, D., Ouyang, P., and Yin, S. (2020, January 25\u201329). Transformer with Bidirectional Decoder for Speech Recognition. Proceedings of the INTERSPEECH, Shanghai, China.","DOI":"10.21437\/Interspeech.2020-2677"},{"key":"ref_13","unstructured":"Wu, D., Zhang, B., Yang, C., Peng, Z., Xia, W., Chen, X., and Lei, X. (2021). U2++: Unified Two-pass Bidirectional End-to-end Model for Speech Recognition. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zhang, C.-F., Liu, Y., Zhang, T.-H., Chen, S.-L., Chen, F., and Yin, X.-C. (2021). Non-autoregressive Transformer with Unified Bidirectional Decoder for Automatic Speech Recognition. arXiv.","DOI":"10.1109\/ICASSP43922.2022.9746903"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. (2015, January 19\u201324). Librispeech: An ASR corpus based on public domain audio books. Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, QLD, Australia.","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Dong, L., Xu, S., and Xu, B. (2018, January 15\u201320). Speech-Transformer: A No-Recurrence Sequence-to-Sequence Model for Speech Recognition. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8462506"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Paul, D.B., and Baker, J.M. (1992, January 23\u201326). The Design for the Wall Street Journal-based CSR Corpus. Proceedings of the HLT, Harriman, NY, USA.","DOI":"10.3115\/1075527.1075614"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Le, H., Pino, J., Wang, C., Gu, J., Schwab, D., and Besacier, L. (2020). Dual-decoder Transformer for Joint Automatic Speech Recognition and Multilingual Speech Translation. arXiv preprint.","DOI":"10.18653\/v1\/2020.coling-main.314"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Shi, Y., Wang, Y., Wu, C., Fuegen, C., Zhang, F., Le, D., Yeh, C.-F., and Seltzer, M.L. (2020). Weak-Attention Suppression For Transformer Based Speech Recognition. arXiv preprint.","DOI":"10.21437\/Interspeech.2020-1363"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Xu, M., Li, S., and Zhang, X.-L. (2021, January 6\u201311). Transformer-based end-to-end speech recognition with local dense synthesizer attention. Proceedings of the ICASSP 2021\u20142021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9414353"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Luo, H., Zhang, S., Lei, M., and Xie, L. (2021, January 19\u201322). Simplified self-attention for transformer-based end-to-end speech recognition. Proceedings of the 2021 IEEE Spoken Language Technology Workshop (SLT), Shenzhen, China.","DOI":"10.1109\/SLT48900.2021.9383581"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Karita, S., Chen, N., Hayashi, T., Hori, T., Inaguma, H., Jiang, Z., Someki, M., Soplin, N.E.Y., Yamamoto, R., and Wang, X. (2019, January 14\u201318). A comparative study on transformer vs. rnn in speech applications. Proceedings of the 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Sentosa, Singapore.","DOI":"10.1109\/ASRU46091.2019.9003750"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wang, Y., Mohamed, A., Le, D., Liu, C., Xiao, A., Mahadeokar, J., Huang, H., Tjandra, A., Zhang, X., and Zhang, F. (2020, January 4\u20138). Transformer-based acoustic modeling for hybrid speech recognition. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9054345"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Tsunoo, E., Kashiwagi, Y., Kumakura, T., and Watanabe, S. (2019, January 14\u201318). Transformer ASR with contextual block processing. Proceedings of the 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Sentosa, Singapore.","DOI":"10.1109\/ASRU46091.2019.9003749"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Wu, C., Wang, Y., Shi, Y., Yeh, C.-F., and Zhang, F. (2020). Streaming transformer-based acoustic models using self-attention with augmented memory. arXiv preprint.","DOI":"10.21437\/Interspeech.2020-2079"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Li, M., Zorila, C., and Doddipatla, R. (2021, January 19\u201322). Transformer-Based Online Speech Recognition with Decoder-end Adaptive Computation Steps. Proceedings of the 2021 IEEE Spoken Language Technology Workshop (SLT), Shenzhen, China.","DOI":"10.1109\/SLT48900.2021.9383613"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Huang, W., Hu, W., Yeung, Y.T., and Chen, X. (2020). Conv-Transformer Transducer: Low Latency, Low Frame Rate, Streamable End-to-End Speech Recognition. arXiv.","DOI":"10.21437\/Interspeech.2020-2361"},{"key":"ref_28","unstructured":"Jiang, D., Lei, X., Li, W., Luo, N., Hu, Y., Zou, W., and Li, X. (2019). Improving Transformer-based Speech Recognition Using Unsupervised Pre-training. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Lu, L., Liu, C., Li, J., and Gong, Y. (2020). Exploring transformers for large-scale speech recognition. arXiv preprint.","DOI":"10.21437\/Interspeech.2020-2638"},{"key":"ref_30","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv preprint."},{"key":"ref_31","unstructured":"Bleeker, M., and de Rijke, M. (2020). Bidirectional Scene Text Recognition with a Single Decoder. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wang, C., Wu, Y., Du, Y., Li, J., Liu, S., Lu, L., Ren, S., Ye, G., Zhao, S., and Zhou, M. (2020). Semantic Mask for Transformer based End-to-End Speech Recognition. arXiv.","DOI":"10.21437\/Interspeech.2020-1778"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Park, D.S., Chan, W., Zhang, Y., Chiu, C.-C., Zoph, B., Cubuk, E.D., and Le, Q.V. (2019, January 15\u201319). SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. Proceedings of the INTERSPEECH, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-2680"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"795","DOI":"10.1162\/tacl_a_00346","article-title":"Best-First Beam Search","volume":"8","author":"Meister","year":"2020","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_35","unstructured":"Loshchilov, I., and Hutter, F. (2017). Fixing Weight Decay Regularization in Adam. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Hsu, W.-N., Lee, A., Synnaeve, G., and Hannun, A.Y. (2020). Semi-Supervised Speech Recognition via Local Prior Matching. arXiv.","DOI":"10.1109\/SLT48900.2021.9383552"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Kahn, J., Lee, A., and Hannun, A.Y. (2020, January 4\u20138). Self-Training for End-to-End Speech Recognition. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9054295"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"L\u00fcscher, C., Beck, E., Irie, K., Kitza, M., Michel, W., Zeyer, A., Schl\u00fcter, R., and Ney, H. (2019, January 15\u201319). RWTH ASR Systems for LibriSpeech: Hybrid vs. Attention\u2014w\/o Data Augmentation. Proceedings of the INTERSPEECH, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-1780"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Ling, S., Liu, Y., Salazar, J., and Kirchhoff, K. (2020, January 4\u20138). Deep Contextualized Acoustic Representations for Semi-Supervised Speech Recognition. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9053176"}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/13\/2\/69\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:10:42Z","timestamp":1760134242000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/13\/2\/69"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,29]]},"references-count":39,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2022,2]]}},"alternative-id":["info13020069"],"URL":"https:\/\/doi.org\/10.3390\/info13020069","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,29]]}}}