{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T00:52:32Z","timestamp":1760230352206,"version":"build-2065373602"},"reference-count":55,"publisher":"MDPI AG","issue":"15","license":[{"start":{"date-parts":[[2022,7,23]],"date-time":"2022-07-23T00:00:00Z","timestamp":1658534400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001659","name":"German Research Foundation DFG","doi-asserted-by":"publisher","award":["KO3434\/4-2"],"award-info":[{"award-number":["KO3434\/4-2"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Audio-visual speech recognition (AVSR) can significantly improve performance over audio-only recognition for small or medium vocabularies. However, current AVSR, whether hybrid or end-to-end (E2E), still does not appear to make optimal use of this secondary information stream as the performance is still clearly diminished in noisy conditions for large-vocabulary systems. We, therefore, propose a new fusion architecture\u2014the decision fusion net (DFN). A broad range of time-variant reliability measures are used as an auxiliary input to improve performance. The DFN is used in both hybrid and E2E models. Our experiments on two large-vocabulary datasets, the Lip Reading Sentences 2 and 3 (LRS2 and LRS3) corpora, show highly significant improvements in performance over previous AVSR systems for large-vocabulary datasets. The hybrid model with the proposed DFN integration component even outperforms oracle dynamic stream-weighting, which is considered to be the theoretical upper bound for conventional dynamic stream-weighting approaches. Compared to the hybrid audio-only model, the proposed DFN achieves a relative word-error-rate reduction of 51% on average, while the E2E-DFN model, with its more competitive audio-only baseline system, achieves a relative word error rate reduction of 43%, both showing the efficacy of our proposed fusion architecture.<\/jats:p>","DOI":"10.3390\/s22155501","type":"journal-article","created":{"date-parts":[[2022,7,25]],"date-time":"2022-07-25T04:52:47Z","timestamp":1658724767000},"page":"5501","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Reliability-Based Large-Vocabulary Audio-Visual Speech Recognition"],"prefix":"10.3390","volume":"22","author":[{"given":"Wentao","family":"Yu","sequence":"first","affiliation":[{"name":"Institute of Communication Acoustics, Ruhr University Bochum, 44801 Bochum, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Steffen","family":"Zeiler","sequence":"additional","affiliation":[{"name":"Institute of Communication Acoustics, Ruhr University Bochum, 44801 Bochum, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0678-3053","authenticated-orcid":false,"given":"Dorothea","family":"Kolossa","sequence":"additional","affiliation":[{"name":"Institute of Communication Acoustics, Ruhr University Bochum, 44801 Bochum, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,7,23]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"9888","DOI":"10.1523\/JNEUROSCI.1396-16.2016","article-title":"Eye can hear clearly now: Inverse effectiveness in natural audiovisual speech processing relies on long-term crossmodal temporal integration","volume":"36","author":"Crosse","year":"2016","journal-title":"J. Neurosci."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"746","DOI":"10.1038\/264746a0","article-title":"Hearing lips and seeing voices","volume":"264","author":"McGurk","year":"1976","journal-title":"Nature"},{"key":"ref_3","unstructured":"Potamianos, G., Neti, C., Luettin, J., and Matthews, I. (2004). Audio-Visual Automatic Speech Recognition: An Overview. Issues in Visual and Audio-Visual Speech Processing, MIT Press."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wand, M., and Schmidhuber, J. (2017). Improving speaker-independent lipreading with domain-adversarial training. arXiv.","DOI":"10.21437\/Interspeech.2017-421"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Meutzner, H., Ma, N., Nickel, R., Schymura, C., and Kolossa, D. (2017, January 5\u20139). Improving audio-visual speech recognition using deep neural networks with dynamic stream reliability estimates. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7953172"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Gurban, M., Thiran, J.P., Drugman, T., and Dutoit, T. (2008, January 20\u201322). Dynamic modality weighting for multi-stream hmms inaudio-visual speech recognition. Proceedings of the Tenth International Conference on Multimodal Interfaces, Chania, Crete, Greece.","DOI":"10.1145\/1452392.1452442"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Kolossa, D., Chong, J., Zeiler, S., and Keutzer, K. (2010, January 26\u201330). Efficient manycore chmm speech recognition for audiovisual and multistream data. Proceedings of the Eleventh Annual Conference of the International Speech Communication Association, Makuhari, Chiba, Japan.","DOI":"10.21437\/Interspeech.2010-715"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Thangthai, K., and Harvey, R.W. (2018, January 2\u20136). Building large-vocabulary speaker-independent lipreading systems. Proceedings of the 19th Annual Conference of the International Speech Communication Association, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-2112"},{"key":"ref_9","unstructured":"Afouras, T., Chung, J.S., Senior, A., Vinyals, O., and Zisserman, A. (2018). Deep audio-visual speech recognition. IEEE Trans. Pattern Anal. Mach. Intell., 1."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"175","DOI":"10.1109\/TCYB.2013.2250954","article-title":"Robust audio-visual speech recognition under noisy audio-video conditions","volume":"44","author":"Stewart","year":"2013","journal-title":"IEEE Trans. Cybern."},{"key":"ref_11","first-page":"863","article-title":"Learning dynamic stream weights for coupled-hmm-based audio-visual speech recognition","volume":"23","author":"Abdelaziz","year":"2015","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1306","DOI":"10.1109\/JPROC.2003.817150","article-title":"Recent advances in the automatic recognition of audiovisual speech","volume":"91","author":"Potamianos","year":"2003","journal-title":"Proc. IEEE"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Luettin, J., Potamianos, G., and Neti, C. (2001, January 7\u201311). Asynchronous stream modeling for large vocabulary audio-visual speech recognition. Proceedings of the 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing, Salt Lake City, UT, USA.","DOI":"10.1109\/ICASSP.2001.940794"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1155\/S1110865702206083","article-title":"Dynamic bayesian networks for audio-visual speech recognition","volume":"2002","author":"Nefian","year":"2002","journal-title":"EURASIP J. Adv. Signal Process."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Wand, M., and Schmidhuber, J. (2020, January 25\u201329). Fusion architectures for word-based audiovisual speech recognition. Proceedings of the 21st Annual Conference of the International Speech Communication Association, Shanghai, China.","DOI":"10.21437\/Interspeech.2020-2117"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zhou, P., Yang, W., Chen, W., Wang, Y., and Jia, J. (2019, January 12\u201317). Modality attention for end-to-end audio-visual speech recognition. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683733"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Yu, J., Zhang, S.X., Wu, J., Ghorbani, S., Wu, B., Kang, S., Liu, S., Liu, X., Meng, H., and Yu, D. (2020, January 4\u20138). Audio-visual recognition of overlapped speech for the LRS2 dataset. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9054127"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"10209","DOI":"10.1007\/s00521-019-04559-1","article-title":"Gated multimodal networks","volume":"32","author":"Arevalo","year":"2020","journal-title":"Neural Comput. Appl."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhang, S., Lei, M., Ma, B., and Xie, L. (2019, January 12\u201317). Robust audio-visual speech recognition using bimodal DFSMN with multi-condition training and dropout regularization. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8682566"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wand, M., Schmidhuber, J., and Vu, N.T. (2018, January 15\u201320). Investigations on end-to-end audiovisual fusion. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461900"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Riva, M., Wand, M., and Schmidhuber, J. (2020, January 4\u20138). Motion dynamics improve speaker-independent lipreading. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9053535"},{"key":"ref_22","unstructured":"Yu, W., Zeiler, S., and Kolossa, D. (September, January 30). Fusing information streams in end-to-end audio-visual speech recognition. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brno, Czech Republic."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Yu, W., Zeiler, S., and Kolossa, D. (2021, January 6\u20138). Large-vocabulary audio-visual speech recognition in noisy environments. Proceedings of the IEEE 23rd International Workshop on Multimedia Signal Processing (MMSP), Tampere, Finland.","DOI":"10.1109\/MMSP53017.2021.9733452"},{"key":"ref_24","unstructured":"Afouras, T., Chung, J.S., and Zisserman, A. (2018). LRS2-TED: A large-scale dataset for visual speech recognition. arXiv."},{"key":"ref_25","unstructured":"Bourlard, H.A., and Morgan, N. (2012). Connectionist Speech Recognition: A Hybrid Approach, Springer."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"L\u00fcscher, C., Beck, E., Irie, K., Kitza, M., Michel, W., Zeyer, A., Schl\u00fcter, R., and Ney, H. (2019). RWTH ASR systems for LibriSpeech: Hybrid vs. attention\u2013w\/o data augmentation. arXiv.","DOI":"10.21437\/Interspeech.2019-1780"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1155\/S1110865702206150","article-title":"Noise adaptive stream weighting in audio-visual speech recognition","volume":"2002","author":"Heckmann","year":"2002","journal-title":"EURASIP J. Adv. Signal Process."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"131","DOI":"10.1002\/ima.20046","article-title":"A multimodal fusion system for people detection and tracking","volume":"15","author":"Yang","year":"2005","journal-title":"Int. J. Imaging Syst. Technol."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"937","DOI":"10.1109\/TMM.2006.879876","article-title":"Experiential sampling in multimedia systems","volume":"8","author":"Kankanhalli","year":"2006","journal-title":"IEEE Trans. Multimed."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Yu, W., Zeiler, S., and Kolossa, D. (2021, January 18\u201321). Multimodal integration for large-vocabulary audio-visual speech recognition. Proceedings of the 28th European Signal Processing Conference (EUSIPCO), Amsterdam, The Netherlands.","DOI":"10.23919\/Eusipco47968.2020.9287841"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"1076","DOI":"10.1109\/JPROC.2012.2236871","article-title":"Multistream recognition of speech: Dealing with unknown unknowns","volume":"101","author":"Hermansky","year":"2013","journal-title":"Proc. IEEE"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Vorwerk, A., Zeiler, S., Kolossa, D., Astudillo, R.F., and Lerch, D. (2011). Use of missing and unreliable data for audiovisual speech recognition. Robust Speech Recognition of Uncertain or Missing Data, Springer.","DOI":"10.1007\/978-3-642-21317-5_13"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Seymour, R., Ming, J., and Stewart, D. (2005, January 4\u20138). A new posterior based audio-visual integration method for robust speech recognition. Proceedings of the Ninth European Conference on Speech Communication and Technology, Lisbon, Portugal.","DOI":"10.21437\/Interspeech.2005-375"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"846","DOI":"10.1109\/TASLP.2016.2520364","article-title":"Turbo automatic speech recognition","volume":"24","author":"Receveur","year":"2016","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Chan, W., Jaitly, N., Le, Q., and Vinyals, O. (2016, January 20\u201325). Listen, attend and spell: A neural network for large vocabulary conversational speech recognition. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472621"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Son Chung, J., Senior, A., Vinyals, O., and Zisserman, A. (2017, January 21\u201326). Lip reading sentences in the wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.367"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Higuchi, Y., Watanabe, S., Chen, N., Ogawa, T., and Kobayashi, T. (2020). Mask CTC: Non-autoregressive end-to-end ASR with CTC and mask predict. arXiv.","DOI":"10.21437\/Interspeech.2020-2404"},{"key":"ref_38","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, Neural Information Processing Systems Foundation."},{"key":"ref_39","unstructured":"Kawakami, K. (2008). Supervised Sequence Labelling with Recurrent Neural Networks. [Ph.D. Thesis, Technical University of Munich]."},{"key":"ref_40","first-page":"12449","article-title":"Wav2vec 2.0: A framework for self-supervised learning of speech representations","volume":"33","author":"Baevski","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_41","unstructured":"Nakatani, T. (2019, January 15\u201319). Improving transformer-based end-to-end speech recognition with connectionist temporal classification and language model integration. Proceedings of the Proc. Interspeech, Graz, Austria."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Mohri, M., Pereira, F., and Riley, M. (2008). Speech recognition with weighted finite-state transducers. Springer Handbook of Speech Processing, Springer.","DOI":"10.1007\/978-3-540-49127-9_28"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Povey, D., Hannemann, M., Boulianne, G., Burget, L., Ghoshal, A., Janda, M., Karafi\u00e1t, M., Kombrink, S., Motl\u00ed\u010dek, P., and Qian, Y. (2012, January 25\u201330). Generating exact lattices in the WFST framework. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Kyoto, Japan.","DOI":"10.1109\/ICASSP.2012.6288848"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Stafylakis, T., and Tzimiropoulos, G. (2017). Combining residual networks with LSTMs for lipreading. arXiv.","DOI":"10.21437\/Interspeech.2017-85"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"259","DOI":"10.1145\/357311.357312","article-title":"Using program transformations to derive line-drawing algorithms","volume":"1","author":"Sproull","year":"1982","journal-title":"ACM Trans. Graph."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1016\/j.specom.2019.06.002","article-title":"Deep learning for minimum mean-square error approaches to speech enhancement","volume":"111","author":"Nicolson","year":"2019","journal-title":"Speech Commun."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"224","DOI":"10.1109\/TASL.2006.876776","article-title":"Robust feature extraction for continuous speech recognition using the MVDR spectrum estimation method","volume":"15","author":"Dharanipragada","year":"2006","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Ghai, S., and Sinha, R. (2011, January 27\u201331). A study on the effect of pitch on LPCC and PLPC features for children\u2019s ASR in comparison to MFCC. Proceedings of the Twelfth Annual Conference of the International Speech Communication Association, Florence, Italy.","DOI":"10.21437\/Interspeech.2011-662"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Baltru\u0161aitis, T., Robinson, P., and Morency, L.P. (2016, January 7\u201310). Openface: An open source facial behavior analysis toolkit. Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Placid, NY, USA.","DOI":"10.1109\/WACV.2016.7477553"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"1052","DOI":"10.1109\/TASLP.2020.2980436","article-title":"How to teach DNNs to pay attention to the visual modality in speech recognition","volume":"28","author":"Sterpu","year":"2020","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_51","unstructured":"Snyder, D., Chen, G., and Povey, D. (2015). Musan: A music, speech, and noise corpus. arXiv."},{"key":"ref_52","unstructured":"Povey, D., Ghoshal, A., Boulianne, G., Burget, L., Glembek, O., Goel, N., Hannemann, M., Motlicek, P., Qian, Y., and Schwarz, P. (2011, January 11\u201315). The kaldi speech recognition toolkit. Proceedings of the IEEE 2011 Workshop on Automatic Speech Recognition and Understanding, Waikoloa, HI, USA."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Zhang, X., Trmal, J., Povey, D., and Khudanpur, S. (2014, January 4\u20139). Improving deep neural network acoustic models using generalized maxout networks. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6853589"},{"key":"ref_54","first-page":"8026","article-title":"Pytorch: An imperative style, high-performance deep learning library","volume":"32","author":"Paszke","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. (2015, January 19\u201324). Librispeech: An asr corpus based on public domain audio books. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, QLD, Australia.","DOI":"10.1109\/ICASSP.2015.7178964"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/15\/5501\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:55:26Z","timestamp":1760140526000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/15\/5501"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,23]]},"references-count":55,"journal-issue":{"issue":"15","published-online":{"date-parts":[[2022,8]]}},"alternative-id":["s22155501"],"URL":"https:\/\/doi.org\/10.3390\/s22155501","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2022,7,23]]}}}