{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T04:17:26Z","timestamp":1782965846826,"version":"3.54.5"},"reference-count":43,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2023,10,13]],"date-time":"2023-10-13T00:00:00Z","timestamp":1697155200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["BDCC"],"abstract":"<jats:p>Emotion recognition is crucial in artificial intelligence, particularly in the domain of human\u2013computer interaction. The ability to accurately discern and interpret emotions plays a critical role in helping machines to effectively decipher users\u2019 underlying intentions, allowing for a more streamlined interaction process that invariably translates into an elevated user experience. The recent increase in social media usage, as well as the availability of an immense amount of unstructured data, has resulted in a significant demand for the deployment of automated emotion recognition systems. Artificial intelligence (AI) techniques have emerged as a powerful solution to this pressing concern in this context. In particular, the incorporation of multimodal AI-driven approaches for emotion recognition has proven beneficial in capturing the intricate interplay of diverse human expression cues that manifest across multiple modalities. The current study aims to develop an effective multimodal emotion recognition system known as MM-EMOR in order to improve the efficacy of emotion recognition efforts focused on audio and text modalities. The use of Mel spectrogram features, Chromagram features, and the Mobilenet Convolutional Neural Network (CNN) for processing audio data are central to the operation of this system, while an attention-based Roberta model caters to the text modality. The methodology of this study is based on an exhaustive evaluation of this approach across three different datasets. Notably, the empirical findings show that MM-EMOR outperforms competing models across the same datasets. This performance boost is noticeable, with accuracy gains of an impressive 7% on one dataset and a substantial 8% on another. Most significantly, the observed increase in accuracy for the final dataset was an astounding 18%.<\/jats:p>","DOI":"10.3390\/bdcc7040164","type":"journal-article","created":{"date-parts":[[2023,10,13]],"date-time":"2023-10-13T10:04:39Z","timestamp":1697191479000},"page":"164","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["MM-EMOR: Multi-Modal Emotion Recognition of Social Media Using Concatenated Deep Learning Networks"],"prefix":"10.3390","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-8742-3801","authenticated-orcid":false,"given":"Omar","family":"Adel","sequence":"first","affiliation":[{"name":"Department of Computer Engineering, Faculty of Engineering and Technology, Arab Academy for Science, Technology and Maritime Transport (AAST), Alexandria 1029, Egypt"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Karma M.","family":"Fathalla","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Faculty of Engineering and Technology, Arab Academy for Science, Technology and Maritime Transport (AAST), Alexandria 1029, Egypt"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ahmed","family":"Abo ElFarag","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Faculty of Engineering and Technology, Arab Academy for Science, Technology and Maritime Transport (AAST), Alexandria 1029, Egypt"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,10,13]]},"reference":[{"key":"ref_1","unstructured":"Li, J., Mishra, S., El-Kishky, A., Mehta, S., and Kulkarni, V. (2022). NTULM: Enriching social media text representations with non-textual units. arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1093\/iwc\/iwt057","article-title":"Dynamic facial emotion recognition oriented to HCI applications","volume":"27","author":"Pablos","year":"2015","journal-title":"Interact. Comput."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Makiuchi, M.R., Uto, K., and Shinoda, K. (2021, January 13\u201317). Multimodal emotion recognition with high-level speech and text features. Proceedings of the 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Cartagena, Colombia.","DOI":"10.1109\/ASRU51503.2021.9688036"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Kandali, A.B., Routray, A., and Basu, T.K. (2008, January 19\u201321). Emotion recognition from Assamese speeches using MFCC features and GMM classifier. Proceedings of the TENCON 2008\u20142008 IEEE Region 10 Conference, Hyderabad, India.","DOI":"10.1109\/TENCON.2008.4766487"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"603","DOI":"10.1016\/S0167-6393(03)00099-2","article-title":"Speech emotion recognition using hidden Markov models","volume":"41","author":"Nwe","year":"2003","journal-title":"Speech Commun."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Graves, A., Mohamed, A.R., and Hinton, G. (2013, January 26\u201331). Speech recognition with deep recurrent neural networks. Proceedings of the 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2203","DOI":"10.1109\/TMM.2014.2360798","article-title":"Learning salient features for speech emotion recognition using convolutional neural networks","volume":"16","author":"Mao","year":"2014","journal-title":"IEEE Trans. Multimed."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1109\/ICDAR.2017.12","article-title":"High performance text recognition using a hybrid convolutional-lstm implementation","volume":"Volume 1","author":"Breuel","year":"2017","journal-title":"Proceedings of the 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR)"},{"key":"ref_9","unstructured":"Jaderberg, M., Simonyan, K., Vedaldi, A., and Zisserman, A. (2014). Deep structured output learning for unconstrained text recognition. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"94","DOI":"10.1109\/MMUL.2022.3161411","article-title":"Emotion Recognition with Multimodal Transformer Fusion Framework Based on Acoustic and Lexical Information","volume":"29","author":"Guo","year":"2022","journal-title":"IEEE MultiMedia"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"63373","DOI":"10.1109\/ACCESS.2019.2916887","article-title":"Deep multimodal representation learning: A survey","volume":"7","author":"Guo","year":"2019","journal-title":"IEEE Access"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Xu, H., Zhang, H., Han, K., Wang, Y., Peng, Y., and Li, X. (2019, January 15\u201319). Learning alignment for multimodal emotion recognition from speech. Proceedings of the 20th Annual Conference of the International Speech Communication Association (INTERSPEECH), Graz, Austria.","DOI":"10.21437\/Interspeech.2019-3247"},{"key":"ref_13","unstructured":"Tripathi, S., Tripathi, S., and Beigi, H. (2018). Multimodal Emotion Recognition on IEMOCAP Dataset using Deep Learning. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"13059","DOI":"10.1007\/s11042-020-10285-x","article-title":"Attention-based multimodal contextual fusion for sentiment and emotion classification using bidirectional LSTM","volume":"80","author":"Huddar","year":"2021","journal-title":"Multimed. Tools Appl."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Eyben, F., Wollmer, M., and Schuller, B. (2010, January 25\u201329). Opensmile: The munich versatile and fast open-source audio feature extractor. Proceedings of the 18th ACM International Conference on Multimedia, Firenze, Italy.","DOI":"10.1145\/1873951.1874246"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1007\/s10579-008-9076-6","article-title":"IEMOCAP: Interactive emotional dyadic motion capture database","volume":"42","author":"Busso","year":"2008","journal-title":"J. Lang. Resour. Eval."},{"key":"ref_17","unstructured":"Kumar, P., Kaushik, V., and Raman, B. (September, January 30). Towards the Explainability of Multimodal Speech Emotion Recognition. Proceedings of the Interspeech, Brno, Czech Republic."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"107316","DOI":"10.1016\/j.knosys.2021.107316","article-title":"A multimodal hierarchical approach to speech emotion recognition from audio and text","volume":"229","author":"Singh","year":"2021","journal-title":"Knowl. Based Syst."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1181598","DOI":"10.3389\/fnbot.2023.1181598","article-title":"Multimodal transformer augmented fusion for speech emotion recognition","volume":"17","author":"Wang","year":"2023","journal-title":"Front. Neurorobotics"},{"key":"ref_20","unstructured":"Zaidi, S.A.M., Latif, S., and Qadi, J. (2023). Cross-Language Speech Emotion Recognition Using Multimodal Dual Attention Transformers. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"593","DOI":"10.1016\/j.ins.2021.10.005","article-title":"A survey on facial emotion recognition techniques: A state-of-the-art literature review","volume":"582","author":"Canal","year":"2022","journal-title":"Inf. Sci."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"345","DOI":"10.1007\/s00530-010-0182-0","article-title":"Multimodal fusion for multimedia analysis: A survey","volume":"16","author":"Atrey","year":"2010","journal-title":"Multimed. Syst."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Huang, J., Li, Y., Tao, J., Lian, Z., Wen, Z., Yang, M., and Yi, J. (2017, January 23\u201327). Continuous multimodal emotion prediction based on long short term memory recurrent neural network. Proceedings of the 7th Annual Workshop on Audio\/Visual Emotion Challenge, Mountain View, CA, USA.","DOI":"10.1145\/3133944.3133946"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Stappen, L., Baird, A., Christ, L., Schumann, L., Sertolli, B., Messner, E.M., Cambria, E., Zhao, G., and Schuller, B.W. (2021, January 24). The MuSe 2021 multimodal sentiment analysis challenge: Sentiment, emotion, physiological-emotion, and stress. Proceedings of the 2nd on Multimodal Sentiment Analysis Challenge, Virtual Event.","DOI":"10.1145\/3475957.3484450"},{"key":"ref_25","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv."},{"key":"ref_26","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_27","unstructured":"Venkataramanan, K., and Rajamohan, H.R. (2019). Emotion recognition from speech. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1016\/j.neucom.2020.01.048","article-title":"Visual-audio emotion recognition based on multi-task and ensemble learning with multiple features","volume":"391","author":"Hao","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"McFee, B., Raffel, C., Liang, D., Ellis, D.P., McVicar, M., Battenberg, E., and Nieto, O. (2015, January 6\u201312). librosa: Audio and music signal analysis in python. Proceedings of the 14th Python in Science Conference, Austin, TX, USA.","DOI":"10.25080\/Majora-7b98e3ed-003"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Solovyev, R.A., Vakhrushev, M., Radionov, A., Romanova, I.I., Amerikanov, A.A., Aliev, V., and Shvets, A.A. (2020, January 22\u201324). Deep learning approaches for understanding simple speech commands. Proceedings of the 2020 IEEE 40th International Conference on Electronics and Nanotechnology (ELNANO), Kyiv, Ukraine.","DOI":"10.1109\/ELNANO50318.2020.9088863"},{"key":"ref_31","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv."},{"key":"ref_32","unstructured":"Sifre, L. (2014). Rigid-Motion Scattering for Image Classification. [Ph.D. Thesis, CMAP Ecole Polytechnique]."},{"key":"ref_33","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Barbieri, F., Camacho-Collados, J., Neves, L., and Espinosa-Anke, L. (2020). Tweeteval: Unified benchmark and comparative evaluation for tweet classification. arXiv.","DOI":"10.18653\/v1\/2020.findings-emnlp.148"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Poria, S., Hazarika, D., Majumder, N., Naik, G., Cambria, E., and Mihalcea, R. (2018). Meld: A multimodal multi-party dataset for emotion recognition in conversations. arXiv.","DOI":"10.18653\/v1\/P19-1050"},{"key":"ref_36","first-page":"344","article-title":"The Nature of Emotions","volume":"89","author":"Plutchik","year":"2001","journal-title":"J. Storage (JSTOR) Digit. Libr. Am. Sci. J."},{"key":"ref_37","unstructured":"Wang, Y., Shen, G., Xu, Y., Li, J., and Zhao, Z. (September, January 30). Learning Mutual Correlation in Multimodal Transformer for Speech Emotion Recognition. Proceedings of the Interspeech, Brno, Czech Republic."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Sahoo, S., Kumar, P., Raman, B., and Roy, P.P. (2019, January 26\u201329). A Segment Level Approach to Speech Emotion Recognition using Transfer Learning. Proceedings of the 5th Asian Conference on Pattern Recognition (ACPR), Auckland, New Zealand.","DOI":"10.1007\/978-3-030-41299-9_34"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Feng, H., Ueno, S., and Kawahara, T. (2020, January 25\u201329). End-to-end Speech Emotion Recognition Combined with Acoustic-to-Word ASR Model. Proceedings of the Interspeech, Shanghai, China.","DOI":"10.21437\/Interspeech.2020-1180"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"3008","DOI":"10.11591\/eei.v12i5.5031","article-title":"Data augmentation and enhancement for multimodal speech emotion recognition","volume":"12","author":"Setyono","year":"2023","journal-title":"Bull. Electr. Eng. Inform."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"1305","DOI":"10.1109\/TAI.2022.3201809","article-title":"M2R2: Missing-Modality Robust emotion Recognition framework with iterative data augmentation","volume":"4","author":"Wang","year":"2022","journal-title":"IEEE Trans. Artif. Intell."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"985","DOI":"10.1109\/TASLP.2021.3049898","article-title":"CTNet: Conversational transformer network for emotion recognition","volume":"29","author":"Lian","year":"2021","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"61672","DOI":"10.1109\/ACCESS.2020.2984368","article-title":"Multimodal approach of speech emotion recognition using multi-level multihead fusion attention-based recurrent neural network","volume":"8","author":"Ho","year":"2020","journal-title":"IEEE Access"}],"container-title":["Big Data and Cognitive Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2504-2289\/7\/4\/164\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:06:32Z","timestamp":1760130392000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2504-2289\/7\/4\/164"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,13]]},"references-count":43,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["bdcc7040164"],"URL":"https:\/\/doi.org\/10.3390\/bdcc7040164","relation":{},"ISSN":["2504-2289"],"issn-type":[{"value":"2504-2289","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,13]]}}}