{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T09:57:07Z","timestamp":1778234227885,"version":"3.51.4"},"reference-count":45,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2025,1,10]],"date-time":"2025-01-10T00:00:00Z","timestamp":1736467200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>The development of emotionally intelligent computers depends on emotion recognition based on richer multimodal inputs, such as text, speech, and visual cues, as multiple modalities complement one another. The effectiveness of complex relationships between modalities for emotion recognition has been demonstrated, but these relationships are still largely unexplored. Various fusion mechanisms using simply concatenated information have been the mainstay of previous research in learning multimodal representations for emotion classification, rather than fully utilizing the benefits of deep learning. In this paper, a unique deep multimodal emotion model is proposed, which uses the meaningful neural network to learn meaningful multimodal representations while classifying data. Specifically, the proposed model concatenates multimodality inputs using a graph convolutional network to extract acoustic modality, a capsule network to generate the textual modality, and vision transformer to acquire the visual modality. Despite the effectiveness of MNN, we have used it as a methodological innovation that will be fed with the previously generated vector parameters to produce better predictive results. Our suggested approach for more accurate multimodal emotion recognition has been shown through extensive examinations, producing state-of-the-art results with accuracies of 69% and 56% on two public datasets, MELD and MOSEI, respectively.<\/jats:p>","DOI":"10.3390\/info16010040","type":"journal-article","created":{"date-parts":[[2025,1,10]],"date-time":"2025-01-10T06:25:54Z","timestamp":1736490354000},"page":"40","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Meaningful Multimodal Emotion Recognition Based on Capsule Graph Transformer Architecture"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7623-2113","authenticated-orcid":false,"given":"Hajar","family":"Filali","sequence":"first","affiliation":[{"name":"LISAC, Department of Computer Science, Faculty of Science Dhar El Mahraz, Sidi Mohamed Ben Abdellah University, Fez 30000, Morocco"},{"name":"Laboratory of \u0130nnovation in Management and Engineering (LIMIE), ISGA, Fez 30000, Morocco"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2097-9098","authenticated-orcid":false,"given":"Chafik","family":"Boulealam","sequence":"additional","affiliation":[{"name":"LISAC, Department of Computer Science, Faculty of Science Dhar El Mahraz, Sidi Mohamed Ben Abdellah University, Fez 30000, Morocco"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Khalid","family":"El Fazazy","sequence":"additional","affiliation":[{"name":"LISAC, Department of Computer Science, Faculty of Science Dhar El Mahraz, Sidi Mohamed Ben Abdellah University, Fez 30000, Morocco"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Adnane Mohamed","family":"Mahraz","sequence":"additional","affiliation":[{"name":"LISAC, Department of Computer Science, Faculty of Science Dhar El Mahraz, Sidi Mohamed Ben Abdellah University, Fez 30000, Morocco"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hamid","family":"Tairi","sequence":"additional","affiliation":[{"name":"LISAC, Department of Computer Science, Faculty of Science Dhar El Mahraz, Sidi Mohamed Ben Abdellah University, Fez 30000, Morocco"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jamal","family":"Riffi","sequence":"additional","affiliation":[{"name":"LISAC, Department of Computer Science, Faculty of Science Dhar El Mahraz, Sidi Mohamed Ben Abdellah University, Fez 30000, Morocco"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,1,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"248","DOI":"10.1037\/h0024648","article-title":"Inference of attitudes from nonverbal communication in two channels","volume":"31","author":"Mehrabian","year":"1967","journal-title":"J. Consult. Psychol."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Liu, J., Liu, Z., Wang, L., Guo, L., and Dang, J. (2020, January 4\u20138). Speech emotion recognition with local-global aware deep representation learning. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9053192"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"58","DOI":"10.1016\/j.future.2020.06.050","article-title":"Transformer based deep intelligent contextual embedding for twitter sentiment analysis","volume":"113","author":"Naseem","year":"2020","journal-title":"Future Gener. Comput. Syst."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Alreshidi, A., and Ullah, M. (2020). Facial emotion recognition using hybrid features. Informatics, 7.","DOI":"10.3390\/informatics7010006"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"374","DOI":"10.1109\/TAFFC.2017.2714671","article-title":"Emotions recognition using EEG signals: A survey","volume":"10","author":"Alarcao","year":"2017","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_6","unstructured":"Poria, S., Hazarika, D., Majumder, N., Naik, G., Cambria, E., and Mihalcea, R. (August, January 28). MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy."},{"key":"ref_7","unstructured":"Chen, S.-Y., Hsu, C.-C., Kuo, C.-C., and Ku, L.-W. (2018). Emotionlines: An emotion corpus of multi-party conversations. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1007\/s10579-008-9076-6","article-title":"IEMOCAP: Interactive emotional dyadic motion capture database","volume":"42","author":"Busso","year":"2008","journal-title":"Lang. Resour. Eval."},{"key":"ref_9","unstructured":"Zadeh, A., Zellers, R., Pincus, E., and Morency, L.-P. (2016). Mosi: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Ouyang, X., Nagisetty, S., Goh, E.G.H., Shen, S., Ding, W., Ming, H., and Huang, D.-Y. (2018, January 20\u201322). Audio-visual emotion recognition with capsule-like feature representation and model-based reinforcement learning. Proceedings of the 2018 First Asian Conference on Affective Computing and Intelligent Interaction (ACII Asia), Beijing, China.","DOI":"10.1109\/ACIIAsia.2018.8470316"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Liu, J., Chen, S., Wang, L., Liu, Z., Fu, Y., Guo, L., and Dang, J. (2021, January 6\u201311). Multimodal emotion recognition with capsule graph convolutional based representation fusion. Proceedings of the ICASSP 2021\u20142021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9413608"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Tang, S., Luo, Z., Nan, G., Baba, J., Yoshikawa, Y., and Ishiguro, H. (2022, January 7\u201310). Fusion with Hierarchical Graphs for Multimodal Emotion Recognition. Proceedings of the 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Chiang Mai, Thailand.","DOI":"10.23919\/APSIPAASC55919.2022.9979932"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Jia, Z., Lin, Y., Wang, J., Feng, Z., Xie, X., and Chen, C. (2021, January 20\u201324). HetEmotionNet: Two-stream heterogeneous graph recurrent neural network for multi-modal emotion recognition. Proceedings of the 29th ACM International Conference on Multimedia, Virtual Event.","DOI":"10.1145\/3474085.3475583"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"29821","DOI":"10.1109\/ACCESS.2022.3159346","article-title":"Simple and Effective Multimodal Learning Based on Pre-Trained Transformer Models","volume":"10","author":"Miyazawa","year":"2022","journal-title":"IEEE Access"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"985","DOI":"10.1109\/TASLP.2021.3049898","article-title":"CTNet: Conversational transformer network for emotion recognition","volume":"29","author":"Lian","year":"2021","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Filali, H., Riffi, J., Boulealam, C., Mahraz, M.A., and Tairi, H. (2022). Multimodal Emotional Classification Based on Meaningful Learning. Big Data Cogn. Comput., 6.","DOI":"10.3390\/bdcc6030095"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"387","DOI":"10.1007\/s11063-021-10636-1","article-title":"Meaningful Learning for Deep Facial Emotional Features","volume":"54","author":"Filali","year":"2021","journal-title":"Neural Process. Lett."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1109\/TCSS.2022.3221128","article-title":"An Emotion Recognition Method Based on Eye Movement and Audiovisual Features in MOOC Learning Environment","volume":"11","author":"Bao","year":"2022","journal-title":"IEEE Trans. Comput. Soc. Syst."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Dhuheir, M., Albaseer, A., Baccour, E., Erbad, A., Abdallah, M., and Hamdi, M. (July, January 28). Emotion recognition for healthcare surveillance systems using neural networks: A survey. Proceedings of the 2021 International Wireless Communications and Mobile Computing (IWCMC), Harbin, China.","DOI":"10.1109\/IWCMC51323.2021.9498861"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"104483","DOI":"10.1016\/j.imavis.2022.104483","article-title":"MEmoR: A multimodal emotion recognition using affective biomarkers for smart prediction of emotional health for people analytics in smart industries","volume":"123","author":"Kumar","year":"2022","journal-title":"Image Vis. Comput."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Wu, Y., Peng, P., Zhang, Z., Zhao, Y., and Qin, B. (2022). An Efficient End-to-End Transformer with Progressive Tri-modal Attention for Multi-modal Emotion Recognition. arXiv.","DOI":"10.1007\/978-981-99-8540-1_32"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Huang, J., Tao, J., Liu, B., Lian, Z., and Niu, M. (2020, January 4\u20138). Multimodal transformer fusion for continuous emotion recognition. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9053762"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Xie, B., Sidulova, M., and Park, C.H. (2021). Robust multimodal emotion recognition from conversation with transformer-based crossmodality fusion. Sensors, 21.","DOI":"10.3390\/s21144913"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"187","DOI":"10.1109\/TCYB.2022.3185119","article-title":"Modeling Hierarchical Uncertainty for Multimodal Emotion Recognition in Conversation","volume":"54","author":"Chen","year":"2022","journal-title":"IEEE Trans. Cybern."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Li, Z., Tang, F., Zhao, M., and Zhu, Y. (2022). EmoCaps: Emotion capsule based model for conversational emotion recognition. arXiv.","DOI":"10.18653\/v1\/2022.findings-acl.126"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Joshi, A., Bhat, A., Jain, A., Singh, A.V., and Modi, A. (2022). COGMEN: COntextualized GNN based multimodal emotion recognitioN. arXiv.","DOI":"10.18653\/v1\/2022.naacl-main.306"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1109\/MMUL.2022.3173430","article-title":"Context-and Knowledge-Aware Graph Convolutional Network for Multimodal Emotion Recognition","volume":"29","author":"Fu","year":"2022","journal-title":"IEEE Multimed."},{"key":"ref_28","unstructured":"Vaswani, A., Shazeer, N., ParmaR, N., Uszkoreit, J., Jone, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017). Attention is all you need. Adv. Neural Inf. Process Syst., 30, Available online: https:\/\/papers.nips.cc\/paper_files\/paper\/2017\/hash\/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html."},{"key":"ref_29","unstructured":"Dosovitskiy, A. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_31","unstructured":"Sabour, S., Frosst, N., and Hinton, G.E. (2017). Dynamic routing between capsules. Adv. Neural Inf. Process Syst., 30, Available online: https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2017\/hash\/2cad8fa47bbef282badbb8de5374b894-Abstract.html."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wang, Y., Huang, L., Jiang, S., Wang, Y., Zou, J., Fu, H., and Yang, S. (2020). Capsule networks showed excellent performance in the classification of hERG blockers\/nonblockers. Front. Pharmacol., 10.","DOI":"10.3389\/fphar.2019.01631"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1109\/TNN.2008.2005605","article-title":"The graph neural network model","volume":"20","author":"Scarselli","year":"2008","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_34","unstructured":"Chen, C., Li, K., Teo, S.G., Zou, X., Wang, K., Wang, J., and Zeng, Z. (February, January 27). Gated residual recurrent graph neural networks for traffic prediction. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_35","unstructured":"Li, Y., Tarlow, D., Brockschmidt, M., and Zemel, R. (2015). Gated graph sequence neural networks. arXiv."},{"key":"ref_36","unstructured":"Kipf, T.N., and Welling, M. (2016). Semi-supervised classification with graph convolutional networks. arXiv."},{"key":"ref_37","unstructured":"Kipf, T.N., and Welling, M. (2016). Variational graph auto-encoders. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Salmam, F.Z., Madani, A., and Kissi, M. (April, January 29). Facial expression recognition using decision trees. Proceedings of the 2016 13th International Conference on Computer Graphics, Imaging and Visualization (CGiV), Beni Mellal, Morocco.","DOI":"10.1109\/CGiV.2016.33"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"1301","DOI":"10.1109\/JSTSP.2017.2764438","article-title":"End-to-end multimodal emotion recognition using deep neural networks","volume":"11","author":"Tzirakis","year":"2017","journal-title":"IEEE J. Sel. Top Signal Process"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"1565","DOI":"10.1038\/nbt1206-1565","article-title":"What is a support vector machine?","volume":"24","author":"Noble","year":"2006","journal-title":"Nat Biotechnol."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"6472","DOI":"10.1109\/TAI.2024.3445325","article-title":"Deep imbalanced learning for multimodal emotion recognition in conversations","volume":"5","author":"Meng","year":"2024","journal-title":"IEEE Trans. Artif. Intell."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"1919","DOI":"10.1109\/TAFFC.2024.3389453","article-title":"CFN-ESA: A Cross-Modal Fusion Network with Emotion-Shift Awareness for Dialogue Emotion Recognition","volume":"15","author":"Li","year":"2024","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Hu, G., Lin, T.-E., Zhao, Y., Lu, G., Wu, Y., and Li, Y. (2022). UniMSE: Towards unified multimodal sentiment analysis and emotion recognition. arXiv.","DOI":"10.18653\/v1\/2022.emnlp-main.534"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"1686","DOI":"10.1162\/tacl_a_00628","article-title":"MissModal: Increasing Robustness to Missing Modality in Multimodal Sentiment Analysis","volume":"11","author":"Lin","year":"2023","journal-title":"Trans. Assoc. Comput. Linguistics"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1613\/jair.953","article-title":"SMOTE: Synthetic minority over-sampling technique","volume":"16","author":"Chawla","year":"2002","journal-title":"J. Artif. Intell. Res."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/1\/40\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,8]],"date-time":"2025-10-08T10:26:29Z","timestamp":1759919189000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/1\/40"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,10]]},"references-count":45,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,1]]}},"alternative-id":["info16010040"],"URL":"https:\/\/doi.org\/10.3390\/info16010040","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,10]]}}}