{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T23:02:39Z","timestamp":1784674959599,"version":"3.55.0"},"reference-count":53,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2023,6,9]],"date-time":"2023-06-09T00:00:00Z","timestamp":1686268800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"GRRC program of Gyeonggi province","award":["GRRC-Gachon2021(B03)"],"award-info":[{"award-number":["GRRC-Gachon2021(B03)"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Methods for detecting emotions that employ many modalities at the same time have been found to be more accurate and resilient than those that rely on a single sense. This is due to the fact that sentiments may be conveyed in a wide range of modalities, each of which offers a different and complementary window into the thoughts and emotions of the speaker. In this way, a more complete picture of a person\u2019s emotional state may emerge through the fusion and analysis of data from several modalities. The research suggests a new attention-based approach to multimodal emotion recognition. This technique integrates facial and speech features that have been extracted by independent encoders in order to pick the aspects that are the most informative. It increases the system\u2019s accuracy by processing speech and facial features of various sizes and focuses on the most useful bits of input. A more comprehensive representation of facial expressions is extracted by the use of both low- and high-level facial features. These modalities are combined using a fusion network to create a multimodal feature vector which is then fed to a classification layer for emotion recognition. The developed system is evaluated on two datasets, IEMOCAP and CMU-MOSEI, and shows superior performance compared to existing models, achieving a weighted accuracy WA of 74.6% and an F1 score of 66.1% on the IEMOCAP dataset and a WA of 80.7% and F1 score of 73.7% on the CMU-MOSEI dataset.<\/jats:p>","DOI":"10.3390\/s23125475","type":"journal-article","created":{"date-parts":[[2023,6,9]],"date-time":"2023-06-09T10:52:15Z","timestamp":1686307935000},"page":"5475","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":113,"title":["Multimodal Emotion Detection via Attention-Based Fusion of Extracted Facial and Speech Features"],"prefix":"10.3390","volume":"23","author":[{"given":"Dilnoza","family":"Mamieva","sequence":"first","affiliation":[{"name":"Department of Computer Engineering, Gachon University, Seongnam-si 13120, Republic of Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5923-8695","authenticated-orcid":false,"given":"Akmalbek Bobomirzaevich","family":"Abdusalomov","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Gachon University, Seongnam-si 13120, Republic of Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alpamis","family":"Kutlimuratov","sequence":"additional","affiliation":[{"name":"Department of AI. Software, Gachon University, Seongnam-si 13120, Republic of Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bahodir","family":"Muminov","sequence":"additional","affiliation":[{"name":"Department of Artificial Intelligence, Tashkent State University of Economics, Tashkent 100066, Uzbekistan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Taeg Keun","family":"Whangbo","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Gachon University, Seongnam-si 13120, Republic of Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,6,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Biele, C., Kacprzyk, J., Kope\u0107, W., Owsi\u0144ski, J.W., Romanowski, A., and Sikorski, M. (2022). Digital Interaction and Machine Intelligence, 9th Machine Intelligence and Digital Interaction Conference, Warsaw, Poland, 9\u201310 December 2021, Springer. Lecture Notes in Networks and Systems.","DOI":"10.1007\/978-3-031-11432-8"},{"key":"ref_2","first-page":"200171","article-title":"A systematic survey on multimodal emotion recognition using learning algorithms","volume":"17","author":"Ahmed","year":"2023","journal-title":"Intell. Syst. Appl."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Gu, X., Shen, Y., and Xu, J. (2021, January 18\u201321). Multimodal Emotion Recognition in Deep Learning:a Survey. Proceedings of the 2021 International Conference on Culture-Oriented Science & Technology (ICCST), Beijing, China.","DOI":"10.1109\/ICCST53801.2021.00027"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"16359","DOI":"10.1007\/s11042-022-14185-0","article-title":"Multimodal emotion recognition from facial expression and speech based on feature fusion","volume":"82","author":"Tang","year":"2022","journal-title":"Multimedia Tools Appl."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Luna-Jim\u00e9nez, C., Griol, D., Callejas, Z., Kleinlein, R., Montero, J.M., and Fern\u00e1ndez-Mart\u00ednez, F. (2021). Multimodal Emotion Recognition on RAVDESS Dataset Using Transfer Learning. Sensors, 21.","DOI":"10.3390\/s21227665"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"817","DOI":"10.1016\/j.aej.2023.01.017","article-title":"A comprehensive survey on deep facial expression recognition: Challenges, applications, and future guidelines","volume":"68","author":"Sajjad","year":"2023","journal-title":"Alex. Eng. J."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"759485","DOI":"10.3389\/fpsyg.2021.759485","article-title":"Facial Expression Emotion Recognition Model Integrating Philosophy and Machine Learning Theory","volume":"12","author":"Song","year":"2021","journal-title":"Front. Psychol."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Abdusalomov, A.B., Safarov, F., Rakhimov, M., Turaev, B., and Whangbo, T.K. (2022). Improved Feature Parameter Extraction from Speech Signals Using Machine Learning Algorithm. Sensors, 22.","DOI":"10.3390\/s22218122"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1675","DOI":"10.1109\/TASLP.2021.3076364","article-title":"Speech emotion recognition considering nonverbal vocalization in affective con-versations","volume":"29","author":"Hsu","year":"2021","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_10","first-page":"5511","article-title":"Automatic speaker recognition using mel-frequency cepstral coefficients through machine learning","volume":"71","author":"Ayvaz","year":"2022","journal-title":"Comput. Mater. Contin."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"2050052","DOI":"10.1142\/S0219691320500526","article-title":"Improvement of the end-to-end scene text recognition method for \u201ctext-to-speech\u201d conversion","volume":"18","author":"Makhmudov","year":"2020","journal-title":"Int. J. Wavelets Multiresolution Inf. Process."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"28349","DOI":"10.1007\/s11042-021-10997-8","article-title":"Selective shallow models strength integration for emotion detection using GloVe and LSTM","volume":"80","author":"Vijayvergia","year":"2021","journal-title":"Multimed. Tools Appl."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Farkhod, A., Abdusalomov, A., Makhmudov, F., and Cho, Y.I. (2021). LDA-Based Topic Modeling Sentiment Analysis Using Topic\/Document\/Sentence (TDS) Model. Appl. Sci., 11.","DOI":"10.3390\/app112311091"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Pan, J., Fang, W., Zhang, Z., Chen, B., Zhang, Z., and Wang, S. (2023). Multimodal Emotion Recognition based on Facial Expressions, Speech, and EEG. IEEE Open J. Eng. Med. Biol., 1\u20138.","DOI":"10.1109\/OJEMB.2023.3240280"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1155\/2023\/7091407","article-title":"Multimodal Emotion Recognition Based on Cascaded Multichannel and Hierarchical Fusion","volume":"2023","author":"Liu","year":"2023","journal-title":"Comput. Intell. Neurosci."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Farkhod, A., Abdusalomov, A.B., Mukhiddinov, M., and Cho, Y.-I. (2022). Development of Real-Time Landmark-Based Emotion Recognition CNN for Masked Faces. Sensors, 22.","DOI":"10.3390\/s22228704"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Chaudhari, A., Bhatt, C., Krishna, A., and Travieso-Gonz\u00e1lez, C.M. (2023). Facial Emotion Recognition with Inter-Modality-Attention-Transformer-Based Self-Supervised Learning. Electronics, 12.","DOI":"10.3390\/electronics12020288"},{"key":"ref_18","first-page":"4243","article-title":"Multimodal Emotion Recognition Using Cross-Modal Attention and 1D Convolutional Neural Networks","volume":"2020","author":"Krishna","year":"2020","journal-title":"Interspeech"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"951","DOI":"10.1007\/s40747-022-00841-3","article-title":"A novel dual-modal emotion recognition algorithm with fusing hybrid features of audio signal and speech context","volume":"9","author":"Xu","year":"2023","journal-title":"Complex Intell. Syst."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Dai, W., Cahyawijaya, S., Liu, Z., and Fung, P. (2021, January 6\u201311). Multimodal end-to-end sparse model for emotion recognition. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics, Online.","DOI":"10.18653\/v1\/2021.naacl-main.417"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1086380","DOI":"10.3389\/fnins.2022.1086380","article-title":"Multimodal interaction enhanced representation learning for video emotion recognition","volume":"16","author":"Xia","year":"2022","journal-title":"Front. Neurosci."},{"key":"ref_22","first-page":"64516","article-title":"Can We Exploit All Datasets?","volume":"10","author":"Yoon","year":"2022","journal-title":"Multimodal Emotion Recognition Using Cross-Modal Translation. IEEE Access"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1082","DOI":"10.1109\/TAFFC.2021.3100868","article-title":"Behavioral and Physiological Signals-Based Deep Multimodal Approach for Mobile Emotion Recognition","volume":"14","author":"Yang","year":"2021","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Tashu, T.M., Hajiyeva, S., and Horvath, T. (2021). Multimodal Emotion Recognition from Art Using Sequential Co-Attention. J. Imaging, 7.","DOI":"10.3390\/jimaging7080157"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Kutlimuratov, A., Abdusalomov, A., and Whangbo, T.K. (2020). Evolving Hierarchical and Tag Information via the Deeply Enhanced Weighted Non-Negative Matrix Factorization of Rating Predictions. Symmetry, 12.","DOI":"10.3390\/sym12111930"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Dang, X., Chen, Z., Hao, Z., Ga, M., Han, X., Zhang, X., and Yang, J. (2023). Wireless Sensing Technology Combined with Facial Expression to Realize Multimodal Emotion Recognition. Sensors, 23.","DOI":"10.3390\/s23010338"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"10535","DOI":"10.1007\/s00521-023-08248-y","article-title":"Meta-transfer learning for emotion recognition","volume":"35","author":"Nguyen","year":"2023","journal-title":"Neural Comput. Appl."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Dresvyanskiy, D., Ryumina, E., Kaya, H., Markitantov, M., Karpov, A., and Minker, W. (2022). End-to-End Modeling and Transfer Learning for Audiovisual Emotion Recognition in-the-Wild. Multimodal Technol. Interact., 6.","DOI":"10.3390\/mti6020011"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1007\/s12193-019-00308-9","article-title":"Multi-modal facial expression feature based on deep-neural networks","volume":"14","author":"Wei","year":"2019","journal-title":"J. Multimodal User Interfaces"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"11365","DOI":"10.1007\/s11042-022-13558-9","article-title":"Facial emotion recognition based real-time learner engagement detection system in online learning context using deep learning models","volume":"82","author":"Gupta","year":"2023","journal-title":"Multimedia Tools Appl."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Chowdary, M.K., Nguyen, T.N., and Hemanth, D.J. (2021). Deep learning-based facial emotion recognition for human\u2013computer inter-action applications. Neural Comput. Appl., 1\u201318.","DOI":"10.1007\/s00521-021-06012-8"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Li, J., Zhang, X., Huang, L., Li, F., Duan, S., and Sun, Y. (2022). Speech Emotion Recognition Using a Dual-Channel Complementary Spectrogram and the CNN-SSAE Neutral Network. Appl. Sci., 12.","DOI":"10.3390\/app12199518"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Kutlimuratov, A., Abdusalomov, A.B., Oteniyazov, R., Mirzakhalilov, S., and Whangbo, T.K. (2022). Modeling and Applying Implicit Dormant Features for Recommendation via Clustering and Deep Factorization. Sensors, 22.","DOI":"10.3390\/s22218224"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zou, H., Si, Y., Chen, C., Rajan, D., and Chng, E.S. (2022, January 23\u201327). Speech Emotion Recognition with Co-Attention Based Multi-Level Acoustic Information. Proceedings of the ICASSP 2022\u20132022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore.","DOI":"10.1109\/ICASSP43922.2022.9747095"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"e1091","DOI":"10.7717\/peerj-cs.1091","article-title":"Feature selection enhancement and feature space visualization for speech-based emotion recognition","volume":"8","author":"Kanwal","year":"2022","journal-title":"PeerJ Comput. Sci."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Du, X., Yang, J., and Xie, X. (2023, January 24\u201326). Multimodal emotion recognition based on feature fusion and residual connection. Proceedings of the 2023 IEEE 2nd International Conference on Electrical Engineering, Big Data and Algorithms (EEBDA), Changchun, China.","DOI":"10.1109\/EEBDA56825.2023.10090537"},{"key":"ref_37","first-page":"112","article-title":"Attention-based Multi-modal Sentiment Analysis and Emotion Detection in Conversation using RNN","volume":"6","author":"Huddar","year":"2021","journal-title":"Int. J. Interact. Multimed. Artif. Intell."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"e12201","DOI":"10.1049\/sil2.12201","article-title":"Attention-based sensor fusion for emotion recognition from human motion by combining convolutional neural network and weighted kernel support vector machine and using inertial measurement unit signals","volume":"17","author":"Zhao","year":"2023","journal-title":"IET Signal Process."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"53","DOI":"10.5772\/54002","article-title":"Towards Efficient Multi-Modal Emotion Recognition","volume":"10","year":"2013","journal-title":"Int. J. Adv. Robot. Syst."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Mamieva, D., Abdusalomov, A.B., Mukhiddinov, M., and Whangbo, T.K. (2023). Improved Face Detection Method via Learning Small Faces on Hard Images Based on a Deep Learning Approach. Sensors, 23.","DOI":"10.3390\/s23010502"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Makhmudov, F., Kutlimuratov, A., Akhmedov, F., Abdallah, M.S., and Cho, Y.-I. (2022). Modeling Speech Emotion Recognition via At-tention-Oriented Parallel CNN Encoders. Electronics, 11.","DOI":"10.3390\/electronics11234047"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"e21","DOI":"10.23915\/distill.00021","article-title":"Computing Receptive Fields of Convolutional Neural Networks","volume":"4","author":"Araujo","year":"2019","journal-title":"Distill"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Wang, C., Sun, H., Zhao, R., and Cao, X. (2020). Research on Bearing Fault Diagnosis Method Based on an Adaptive Anti-Noise Network under Long Time Series. Sensors, 20.","DOI":"10.3390\/s20247031"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Hsu, S.-M., Chen, S.-H., and Huang, T.-R. (2021). Personal Resilience Can Be Well Estimated from Heart Rate Variability and Paralinguistic Features during Human\u2013Robot Conversations. Sensors, 21.","DOI":"10.3390\/s21175844"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Mirsamadi, S., Barsoum, E., and Zhang, C. (2017, January 5\u20139). Automatic speech emotion recognition using recurrent neural networks with local attention. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7952552"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Ayetiran, E.F. (2022). Attention-based aspect sentiment classification using enhanced learning through cnn-Bilstm networks. Knowl. Based Syst., 252.","DOI":"10.1016\/j.knosys.2022.109409"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Poria, S., Cambria, E., Hazarika, D., Mazumder, N., Zadeh, A., and Morency, L.-P. (2017, January 18\u201321). Multi-level Multiple Attentions for Contextual Multimodal Sentiment Analysis. Proceedings of the 2017 IEEE International Conference on Data Mining (ICDM), Orleans, LA, USA.","DOI":"10.1109\/ICDM.2017.134"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1007\/s10579-008-9076-6","article-title":"IEMOCAP: Interactive emotional dyadic motion capture database","volume":"42","author":"Busso","year":"2008","journal-title":"Lang. Resour. Evaluation"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1109\/MIS.2018.2882362","article-title":"Multimodal Sentiment Analysis: Addressing Key Issues and Setting Up the Baselines","volume":"33","author":"Poria","year":"2018","journal-title":"IEEE Intell. Syst."},{"key":"ref_51","unstructured":"Zadeh, A., and Pu, P. (2018, January 15\u201320). Multimodal language analysis in the wild: CMU-mosei dataset and interpretable dynamic fusion graph. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Long Papers), Melbourne, VIC, Australia."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Ilyosov, A., Kutlimuratov, A., and Whangbo, T.-K. (2021). Deep-Sequence\u2013Aware Candidate Generation for e-Learning System. Processes, 9.","DOI":"10.3390\/pr9081454"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Safarov, F., Kutlimuratov, A., Abdusalomov, A.B., Nasimov, R., and Cho, Y.-I. (2023). Deep Learning Recommendations of E-Education Based on Clustering and Sequence. Electronics, 12.","DOI":"10.3390\/electronics12040809"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/12\/5475\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:52:00Z","timestamp":1760125920000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/12\/5475"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,9]]},"references-count":53,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2023,6]]}},"alternative-id":["s23125475"],"URL":"https:\/\/doi.org\/10.3390\/s23125475","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,9]]}}}