{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T23:09:39Z","timestamp":1784243379503,"version":"3.55.0"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2023,11,29]],"date-time":"2023-11-29T00:00:00Z","timestamp":1701216000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,11,29]],"date-time":"2023-11-29T00:00:00Z","timestamp":1701216000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2024,2]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>With the rapid development of information technology in modern society, the application of multimedia integration platform is more and more extensive. Speech recognition has become an important subject in the process of multimedia visual interaction. The accuracy of speech recognition is dependent on a number of elements, two of which are the acoustic characteristics of speech and the speech recognition model. Speech data is complex and changeable. Most methods only extract a single type of feature of the signal to represent the speech signal. This single feature cannot express the hidden information. And, the excellent speech recognition model can also better learn the characteristic speech information to improve performance. This work proposes a new method for speech recognition in multimedia visual interaction. First of all, this work considers the problem that a single feature cannot fully represent complex speech information. This paper proposes three kinds of feature fusion structures to extract speech information from different angles. This extracts three different fusion features based on the low-level features and higher-level sparse representation. Secondly, this work relies on the strong learning ability of neural network and the weight distribution mechanism of attention model. In this paper, the fusion feature is combined with the bidirectional long and short memory network with attention. The extracted fusion features contain more speech information with strong discrimination. When the weight increases, it can further improve the influence of features on the predicted value and improve the performance. Finally, this paper has carried out systematic experiments on the proposed method, and the results verify the feasibility.<\/jats:p>","DOI":"10.1007\/s00521-023-08959-2","type":"journal-article","created":{"date-parts":[[2023,11,29]],"date-time":"2023-11-29T04:30:34Z","timestamp":1701232234000},"page":"2371-2383","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Intelligent speech recognition algorithm in multimedia visual interaction via BiLSTM and attention mechanism"],"prefix":"10.1007","volume":"36","author":[{"given":"Yican","family":"Feng","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,11,29]]},"reference":[{"issue":"5","key":"8959_CR1","doi-asserted-by":"publisher","first-page":"711","DOI":"10.3390\/math10050711","volume":"10","author":"A Zgank","year":"2022","unstructured":"Zgank A (2022) Influence of highly inflected word forms and acoustic background on the robustness of automatic speech recognition for human-computer interaction. Mathematics 10(5):711","journal-title":"Mathematics"},{"issue":"2","key":"8959_CR2","doi-asserted-by":"publisher","first-page":"391","DOI":"10.1007\/s10772-021-09955-4","volume":"25","author":"M Liu","year":"2022","unstructured":"Liu M (2022) English speech emotion recognition method based on speech recognition. Int J Speech Technol 25(2):391\u2013398","journal-title":"Int J Speech Technol"},{"issue":"1","key":"8959_CR3","doi-asserted-by":"publisher","first-page":"20","DOI":"10.3390\/s22010020","volume":"22","author":"B \u0160umak","year":"2022","unstructured":"\u0160umak B, Brdnik S, Pu\u0161nik M (2022) Sensors and artificial intelligence methods and algorithms for human\u2013computer intelligent interaction: a systematic mapping study. Sensors 22(1):20","journal-title":"Sensors"},{"key":"8959_CR4","first-page":"1","volume":"1","author":"Y Liu","year":"2022","unstructured":"Liu Y, Sivaparthipan CB, Shankar A (2022) Human\u2013computer interaction based visual feedback system for augmentative and alternative communication. Int J Speech Technol 1:1\u201310","journal-title":"Int J Speech Technol"},{"key":"8959_CR5","doi-asserted-by":"crossref","unstructured":"Sang Y, Chen X (2022) Human-computer interactive physical education teaching method based on speech recognition engine technology. Front Public Health 10:941083\u2013941097","DOI":"10.3389\/fpubh.2022.941083"},{"key":"8959_CR6","unstructured":"Markl N, Lai C(2021) Context-sensitive evaluation of automatic speech recognition: considering user experience & language variation[C]. In: Proceedings of the First Workshop on Bridging Human\u2013Computer Interaction and Natural Language Processing, pp 34\u201340"},{"issue":"2","key":"8959_CR7","doi-asserted-by":"publisher","first-page":"861","DOI":"10.1007\/s11423-020-09910-1","volume":"69","author":"EY Oh","year":"2021","unstructured":"Oh EY, Song D (2021) Developmental research on an interactive application for language speaking practice using speech recognition technology. Educ Tech Res Dev 69(2):861\u2013884","journal-title":"Educ Tech Res Dev"},{"issue":"2","key":"8959_CR8","doi-asserted-by":"publisher","first-page":"3513","DOI":"10.3233\/JIFS-189388","volume":"40","author":"D Ran","year":"2021","unstructured":"Ran D, Yingli W, Haoxin Q (2021) Artificial intelligence speech recognition model for correcting spoken English teaching. J Intell Fuzzy Syst 40(2):3513\u20133524","journal-title":"J Intell Fuzzy Syst"},{"issue":"20","key":"8959_CR9","doi-asserted-by":"publisher","first-page":"23471","DOI":"10.1109\/JSEN.2021.3107949","volume":"21","author":"Q Fu","year":"2021","unstructured":"Fu Q, Fu J, Zhang S et al (2021) Design of intelligent human-computer interaction system for hard of hearing and non-disabled people. IEEE Sens J 21(20):23471\u201323479","journal-title":"IEEE Sens J"},{"key":"8959_CR10","unstructured":"Pei J, Yu Z, Li J, et al (2022) TKAGFL: a federated communication framework under data heterogeneity. IEEE Trans Netw Sci Eng 1:1\u201311"},{"key":"8959_CR11","doi-asserted-by":"crossref","unstructured":"Weng Z, Qin Z, Tao X, et al 2023 () Deep learning enabled semantic communications with speech recognition and synthesis. IEEE Trans Wirel Commun 1:6227\u20136240","DOI":"10.1109\/TWC.2023.3240969"},{"key":"8959_CR12","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2022.101360","volume":"75","author":"AS Subramanian","year":"2022","unstructured":"Subramanian AS, Weng C, Watanabe S et al (2022) Deep learning based multi-source localization with source splitting and its effectiveness in multi-talker speech recognition. Comput Speech Lang 75:101360","journal-title":"Comput Speech Lang"},{"key":"8959_CR13","doi-asserted-by":"publisher","first-page":"30069","DOI":"10.1109\/ACCESS.2022.3159339","volume":"10","author":"J Oruh","year":"2022","unstructured":"Oruh J, Viriri S, Adegun A (2022) Long short-term Memory Recurrent neural network for Automatic speech recognition. IEEE Access 10:30069\u201330079","journal-title":"IEEE Access"},{"issue":"1","key":"8959_CR14","doi-asserted-by":"publisher","first-page":"2095039","DOI":"10.1080\/08839514.2022.2095039","volume":"36","author":"JLKE Fendji","year":"2022","unstructured":"Fendji JLKE, Tala DCM, Yenke BO et al (2022) Automatic speech recognition using limited vocabulary: a survey. Appl Artif Intell 36(1):2095039","journal-title":"Appl Artif Intell"},{"issue":"2","key":"8959_CR15","doi-asserted-by":"publisher","first-page":"1913","DOI":"10.1007\/s11277-022-09640-y","volume":"125","author":"KB Bhangale","year":"2022","unstructured":"Bhangale KB, Kothandaraman M (2022) Survey of deep learning paradigms for speech processing. Wirel Pers Commun 125(2):1913\u20131949","journal-title":"Wirel Pers Commun"},{"issue":"12","key":"8959_CR16","doi-asserted-by":"publisher","first-page":"6223","DOI":"10.3390\/app12126223","volume":"12","author":"S Dua","year":"2022","unstructured":"Dua S, Kumar SS, Albagory Y et al (2022) Developing a speech recognition system for recognizing tonal speech signals using a convolutional neural network. Appl Sci 12(12):6223","journal-title":"Appl Sci"},{"key":"8959_CR17","doi-asserted-by":"crossref","first-page":"1","DOI":"10.57255\/intellect.v1i1.9","volume":"1","author":"AK Gupta","year":"2022","unstructured":"Gupta AK, Gupta P, Rahtu E (2022) FATALRead-fooling visual speech recognition models: put words on lips. Appl Intell 1:1\u201316","journal-title":"Appl Intell"},{"key":"8959_CR18","doi-asserted-by":"crossref","unstructured":"Lu Y J, Chang X, Li C, et al (2022) ESPnet-SE++: Speech enhancement for robust speech recognition, translation, and understanding. arXiv preprint arXiv:2207.09514","DOI":"10.21437\/Interspeech.2022-10727"},{"issue":"4","key":"8959_CR19","doi-asserted-by":"publisher","first-page":"672","DOI":"10.4218\/etrij.2021-0118","volume":"44","author":"P Agarwal","year":"2022","unstructured":"Agarwal P, Kumar S (2022) Electroencephalography-based imagined speech recognition using deep long short-term memory network. ETRI J 44(4):672\u2013685","journal-title":"ETRI J"},{"issue":"7","key":"8959_CR20","first-page":"3425","volume":"14","author":"R Shashidhar","year":"2022","unstructured":"Shashidhar R, Patilkulkarni S, Puneeth SB (2022) Combining audio and visual speech recognition using LSTM and deep convolutional neural network. Int J Inf Technol 14(7):3425\u20133436","journal-title":"Int J Inf Technol"},{"issue":"1","key":"8959_CR21","first-page":"1","volume":"3","author":"LE Baum","year":"1972","unstructured":"Baum LE (1972) An inequality and associated maximization technique in statistical estimation for probabilistic functions of Markov processes. Inequalities 3(1):1\u20138","journal-title":"Inequalities"},{"key":"8959_CR22","first-page":"61","volume":"1","author":"A Graves","year":"2012","unstructured":"Graves A, Graves A (2012) Connectionist temporal classification. Supervised Seq Labell Recurr Neural Netw 1:61\u201393","journal-title":"Supervised Seq Labell Recurr Neural Netw"},{"key":"8959_CR23","doi-asserted-by":"crossref","unstructured":"Chan W, Jaitly N, Le Q, et al (2016) Listen, attend and spell: a neural network for large vocabulary conversational speech recognition. In: IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, pp 4960\u20134964","DOI":"10.1109\/ICASSP.2016.7472621"},{"key":"8959_CR24","doi-asserted-by":"crossref","unstructured":"Graves A (2012) Sequence transduction with recurrent neural networks. arXiv preprint arXiv:1211.3711","DOI":"10.1007\/978-3-642-24797-2"},{"issue":"3","key":"8959_CR25","doi-asserted-by":"publisher","first-page":"328","DOI":"10.1109\/29.21701","volume":"37","author":"A Waibel","year":"1989","unstructured":"Waibel A, Hanazawa T, Hinton G et al (1989) Phoneme recognition using time-delay neural networks. IEEE Trans Acoust Speech Signal Process 37(3):328\u2013339","journal-title":"IEEE Trans Acoust Speech Signal Process"},{"key":"8959_CR26","doi-asserted-by":"publisher","first-page":"4840","DOI":"10.1007\/s00034-019-01092-3","volume":"38","author":"H Liu","year":"2019","unstructured":"Liu H, Zhao L (2019) A speaker verification method based on TDNN\u2013LSTMP. Circ Syst Signal Process 38:4840\u20134854","journal-title":"Circ Syst Signal Process"},{"key":"8959_CR27","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1007\/978-1-4613-1367-0_3","volume":"1","author":"Y Normandin","year":"1996","unstructured":"Normandin Y (1996) Maximum mutual information estimation of hidden Markov models. Autom Speech Speak Recog: Adv Top 1:57\u201381","journal-title":"Autom Speech Speak Recog: Adv Top"},{"key":"8959_CR28","doi-asserted-by":"crossref","unstructured":"Cho K, Van Merri\u00ebnboer B, Gulcehre C, et al (2014) Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078","DOI":"10.3115\/v1\/D14-1179"},{"key":"8959_CR29","doi-asserted-by":"crossref","unstructured":"Bahdanau D, Chorowski J, Serdyuk D, et al (2016) End-to-end attention-based large vocabulary speech recognition. In: IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, pp 4945\u20134949","DOI":"10.1109\/ICASSP.2016.7472618"},{"key":"8959_CR30","unstructured":"Vaswani A, Shazeer N, Parmar N, et al (2017) Attention is all you need. Adv in Neural Information Processing Systems 30:1\u201311"},{"key":"8959_CR31","doi-asserted-by":"crossref","unstructured":"Zhou S, Dong L, Xu S, et al (2018) Syllable-based sequence-to-sequence speech recognition with the transformer in mandarin Chinese. arXiv preprint arXiv:1804.10752","DOI":"10.21437\/Interspeech.2018-1107"},{"key":"8959_CR32","doi-asserted-by":"crossref","unstructured":"Zhang Y, Lu X (2018) A speech recognition acoustic model based on LSTM-CTC[C]. In: IEEE 18th International Conference on Communication Technology (ICCT). IEEE, pp 1052\u20131055","DOI":"10.1109\/ICCT.2018.8599961"},{"key":"8959_CR33","doi-asserted-by":"crossref","unstructured":"Zhang S, Lei M, Yan Z, et al (2018) Deep-FSMN for large vocabulary continuous speech recognition. In: IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, pp 5869\u20135873","DOI":"10.1109\/ICASSP.2018.8461404"},{"key":"8959_CR34","doi-asserted-by":"publisher","first-page":"28","DOI":"10.1017\/ATSIP.2020.26","volume":"9","author":"X Cheng","year":"2020","unstructured":"Cheng X, Xu M, Zheng TF (2020) A multi-branch ResNet with discriminative features for detection of replay speech signals. APSIPA Trans Signal Inform Process 9:28","journal-title":"APSIPA Trans Signal Inform Process"},{"key":"8959_CR35","doi-asserted-by":"crossref","unstructured":"Sivaram G, Nemala S K, Elhilali M, et al (2010) Sparse Coding for Speech Recognition. In: IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, pp 4346\u20134349","DOI":"10.1109\/ICASSP.2010.5495649"},{"issue":"1","key":"8959_CR36","doi-asserted-by":"publisher","first-page":"129","DOI":"10.1137\/S003614450037906X","volume":"43","author":"S Chen","year":"2001","unstructured":"Chen S, Saunders D (2001) Atomic decomposition by basis pursuit. SIAM Rev 43(1):129\u2013159","journal-title":"SIAM Rev"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-023-08959-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-023-08959-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-023-08959-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,1,19]],"date-time":"2024-01-19T15:10:56Z","timestamp":1705677056000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-023-08959-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,29]]},"references-count":36,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,2]]}},"alternative-id":["8959"],"URL":"https:\/\/doi.org\/10.1007\/s00521-023-08959-2","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,29]]},"assertion":[{"value":"17 February 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 August 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 November 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no conflict of interest exists.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}