{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T01:05:16Z","timestamp":1760058316700,"version":"build-2065373602"},"reference-count":55,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2025,3,31]],"date-time":"2025-03-31T00:00:00Z","timestamp":1743379200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["MTI"],"abstract":"<jats:p>Artificial agents are expected to increasingly interact with humans and to demonstrate multimodal adaptive emotional responses. Such social integration requires both perception and production mechanisms, thus enabling a more realistic approach to emotional alignment than existing systems. Indeed, existing emotion recognition methods rely on behavioral signals, predominantly facial expressions, as well as non-invasive brain recordings, such as Electroencephalograms (EEGs) and functional Magnetic Resonance Imaging (fMRI), to identify humans\u2019 emotions, but accurate labeling remains a challenge. This paper introduces a novel approach examining how behavioral and physiological signals can be used to predict activity in emotion-related regions of the brain. To this end, we propose a multimodal deep learning network that processes two categories of signals recorded alongside brain activity during conversations: two behavioral signals (video and audio) and one physiological signal (blood pulse). Our network enables (1) the prediction of brain activity from these multimodal inputs, and (2) the assessment of our model\u2019s performance depending on the nature of interlocutor (human or robot) and the brain region of interest. Results demonstrate that the proposed architecture outperforms existing models in anterior insula and hypothalamus regions, for interactions with a human or a robot. An ablation study evaluating subsets of input modalities indicates that local brain activity prediction was reduced when one or two modalities are omitted. However, they also revealed that the physiological data (blood pulse) achieve similar levels of predictions alone compared to the full model, further underscoring the importance of somatic markers in the central nervous system\u2019s processing of social emotions.<\/jats:p>","DOI":"10.3390\/mti9040031","type":"journal-article","created":{"date-parts":[[2025,3,31]],"date-time":"2025-03-31T05:21:04Z","timestamp":1743398464000},"page":"31","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Predicting Activity in Brain Areas Associated with Emotion Processing Using Multimodal Behavioral Signals"],"prefix":"10.3390","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-6047-8798","authenticated-orcid":false,"given":"Lahoucine","family":"Kdouri","sequence":"first","affiliation":[{"name":"International Artificial Intelligence Center of Morocco, University Mohammed VI Polytechnique, Rabat 11103, Morocco"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-1831-7295","authenticated-orcid":false,"given":"Youssef","family":"Hmamouche","sequence":"additional","affiliation":[{"name":"International Artificial Intelligence Center of Morocco, University Mohammed VI Polytechnique, Rabat 11103, Morocco"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Amal","family":"El Fallah Seghrouchni","sequence":"additional","affiliation":[{"name":"International Artificial Intelligence Center of Morocco, University Mohammed VI Polytechnique, Rabat 11103, Morocco"},{"name":"Laboratoire de Recherche en Informatique, UMR 7606, CNRS-Sorbonne University, 75252 Paris, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4952-1467","authenticated-orcid":false,"given":"Thierry","family":"Chaminade","sequence":"additional","affiliation":[{"name":"Institut de Neurosciences de la Timone, UMR 7289, CNRS-Aix-Marseille University, 13385 Marseille, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,3,31]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1016\/j.aiopen.2022.02.001","article-title":"Learning towards conversational AI: A survey","volume":"3","author":"Fu","year":"2022","journal-title":"AI Open"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Spezialetti, M., Placidi, G., and Rossi, S. (2020). Emotion Recognition for Human-Robot Interaction: Recent Advances and Future Perspectives. Front. Robot. AI, 7.","DOI":"10.3389\/frobt.2020.532279"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"597","DOI":"10.1016\/j.procs.2020.07.086","article-title":"Emotions recognition as innovative tool for improving students\u2019 performance and learning approaches","volume":"175","author":"Bouhlal","year":"2020","journal-title":"Procedia Comput. Sci."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"184","DOI":"10.1007\/978-3-319-67401-8_20","article-title":"A Psychotherapy Training Environment with Virtual Patients Implemented Using the Furhat Robot Platform","volume":"Volume 10498","author":"Johansson","year":"2017","journal-title":"Intelligent Virtual Agents"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Barros, P., Weber, C., and Wermter, S. (2015, January 3\u20135). Emotional expression recognition with a cross-channel convolutional neural network for human-robot interaction. Proceedings of the 2015 IEEE-RAS 15th International Conference on Humanoid Robots (Humanoids), Seoul, Republic of Korea.","DOI":"10.1109\/HUMANOIDS.2015.7363421"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"715","DOI":"10.1109\/TCDS.2021.3071170","article-title":"Comparing Recognition Performance and Robustness of Multimodal Deep Learning Models for Multimodal Emotion Recognition","volume":"14","author":"Liu","year":"2022","journal-title":"IEEE Trans. Cogn. Dev. Syst."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"121692","DOI":"10.1016\/j.eswa.2023.121692","article-title":"Deep learning-based multimodal emotion recognition from audio, visual, and text modalities: A systematic review of recent advancements and future prospects","volume":"237","author":"Zhang","year":"2024","journal-title":"Expert Syst. Appl."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"122579","DOI":"10.1016\/j.eswa.2023.122579","article-title":"Deep CNN with late fusion for real time multimodal emotion recognition","volume":"240","author":"Dixit","year":"2024","journal-title":"Expert Syst. Appl."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"e12607","DOI":"10.1111\/coin.12607","article-title":"A joint hierarchical cross attention graph convolutional network for multimodal facial expression recognition","volume":"40","author":"Xu","year":"2024","journal-title":"Comput. Intell."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"107708","DOI":"10.1016\/j.engappai.2023.107708","article-title":"Multimodal Emotion Recognition via Convolutional Neural Networks: Comparison of different strategies on two multimodal datasets","volume":"130","author":"Bilotti","year":"2024","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"357","DOI":"10.1016\/j.tics.2017.03.002","article-title":"A Network Model of the Emotional Brain","volume":"21","author":"Pessoa","year":"2017","journal-title":"Trends Cogn. Sci."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"20180033","DOI":"10.1098\/rstb.2018.0033","article-title":"Brain activity during reciprocal social interaction investigated using conversational robots as control condition","volume":"374","author":"Rauchbauer","year":"2019","journal-title":"Phil. Trans. R. Soc. B"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"83","DOI":"10.1016\/S0165-0173(97)00064-7","article-title":"Emotion in the perspective of an integrated nervous system","volume":"26","author":"Damasio","year":"1998","journal-title":"Brain Res. Rev."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Liang, P.P., Zadeh, A., and Morency, L.P. (2018, January 16\u201320). Multimodal Local-Global Ranking Fusion for Emotion Recognition. Proceedings of the 20th ACM International Conference on Multimodal Interaction, Boulder, CO, USA.","DOI":"10.1145\/3242969.3243019"},{"key":"ref_15","unstructured":"Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A.Y. (2011, January 28). Multimodal deep learning. Proceedings of the 28th International Conference on International Conference on Machine Learning, ICML\u201911, Madison, WI, USA."},{"key":"ref_16","unstructured":"Pereira, F., Burges, C., Bottou, L., and Weinberger, K. (2012). Multimodal Learning with Deep Boltzmann Machines. Advances in Neural Information Processing Systems, Curran Associates, Inc."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"829","DOI":"10.1162\/neco_a_01273","article-title":"A Survey on Deep Learning for Multimodal Data Fusion","volume":"32","author":"Gao","year":"2020","journal-title":"Neural Comput."},{"key":"ref_18","unstructured":"Akkus, C., Chu, L., Djakovic, V., Jauch-Walser, S., Koch, P., Loss, G., Marquardt, C., Moldovan, M., Sauter, N., and Schneider, M. (2023). Multimodal Deep Learning. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"126866","DOI":"10.1016\/j.neucom.2023.126866","article-title":"A review of multimodal emotion recognition from datasets, preprocessing, features, and fusion methods","volume":"561","author":"Pan","year":"2023","journal-title":"Neurocomputing"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Sosea, T., and Caragea, C. (2020, January 16\u201320). CancerEmo: A Dataset for Fine-Grained Emotion Detection. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, Online.","DOI":"10.18653\/v1\/2020.emnlp-main.715"},{"key":"ref_21","first-page":"2782","article-title":"CAS(ME)3: A Third Generation Facial Spontaneous Micro-Expression Database with Depth Information and High Ecological Validity","volume":"45","author":"Li","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"108580","DOI":"10.1016\/j.knosys.2022.108580","article-title":"Deep learning based multimodal emotion recognition using model-level fusion of audio\u2013visual modalities","volume":"244","author":"Middya","year":"2022","journal-title":"Knowl.-Based Syst."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Deschamps-Berger, T., Lamel, L., and Devillers, L. (2022, January 7\u201311). Investigating Transformer Encoders and Fusion Strategies for Speech Emotion Recognition in Emergency Call Center Conversations. Proceedings of the International Conference on Multimodal Interaction, New York, NY, USA.","DOI":"10.1145\/3536220.3558038"},{"key":"#cr-split#-ref_24.1","unstructured":"Ranchordas, A.N. (2008). VISAPP 2008: Proceedings of the Third International Conference on Vision Theory and Applications"},{"key":"#cr-split#-ref_24.2","unstructured":"Funchal, Madeira, Portugal, January 22-25, 2008. VISIGRAPP 2008, INSTICC Press."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Etienne, C., Fidanza, G., Petrovskii, A., Devillers, L., and Schmauch, B. (2018, January 1). CNN+LSTM Architecture for Speech Emotion Recognition with Data Augmentation. Proceedings of the Workshop on Speech, Music and Mind (SMM 2018), Hyderabad, India.","DOI":"10.21437\/SMM.2018-5"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"111077","DOI":"10.1016\/j.knosys.2023.111077","article-title":"Spatio-temporal representation learning enhanced speech emotion recognition with multi-head attention mechanisms","volume":"281","author":"Chen","year":"2023","journal-title":"Knowl.-Based Syst."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Wang, S., Wang, W., Zhao, J., Chen, S., Jin, Q., Zhang, S., and Qin, Y. (2017, January 13\u201317). Emotion recognition with multimodal features and temporal models. Proceedings of the 19th ACM International Conference on Multimodal Interaction, Glasgow, UK.","DOI":"10.1145\/3136755.3143016"},{"key":"ref_28","unstructured":"Krishna, D.N., and Patil, A. (2020, January 25\u201329). Multimodal Emotion Recognition Using Cross-Modal Attention and 1D Convolutional Neural Networks. Proceedings of the Interspeech 2020, ISCA, Shanghai, China."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Schroff, F., Kalenichenko, D., and Philbin, J. (2015, January 7\u201312). FaceNet: A Unified Embedding for Face Recognition and Clustering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"2593036","DOI":"10.1155\/2019\/2593036","article-title":"Audio-Textual Emotion Recognition Based on Improved Neural Networks","volume":"2019","author":"Cai","year":"2019","journal-title":"Math. Probl. Eng."},{"key":"ref_31","unstructured":"Ghauri, J.A., Hakimov, S., and Ewerth, R. (2020). Classification of Important Segments in Educational Videos using Multimodal Features. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Poria, S., Cambria, E., Hazarika, D., Mazumder, N., Zadeh, A., and Morency, L.P. (2017, January 18\u201321). Multi-level Multiple Attentions for Contextual Multimodal Sentiment Analysis. Proceedings of the 2017 IEEE International Conference on Data Mining (ICDM), New Orleans, LA, USA.","DOI":"10.1109\/ICDM.2017.134"},{"key":"ref_33","unstructured":"Tsai, Y.H.H., Bai, S., Liang, P.P., Kolter, J.Z., Morency, L.P., and Salakhutdinov, R. (August, January 28). Multimodal Transformer for Unaligned Multimodal Language Sequences. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Florence, Italy."},{"key":"ref_34","first-page":"24206","article-title":"VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text","volume":"Volume 34","author":"Ranzato","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1017\/S0140525X11000446","article-title":"The brain basis of emotion: A meta-analytic review","volume":"35","author":"Lindquist","year":"2012","journal-title":"Behav. Brain. Sci."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Aybek, S., Nicholson, T.R., O\u2019Daly, O., Zelaya, F., Kanaan, R.A., and David, A.S. (2015). Emotion-Motion Interactions in Conversion Disorder: An fMRI Study. PLoS ONE, 10.","DOI":"10.1371\/journal.pone.0123273"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"858","DOI":"10.1038\/s41593-023-01304-9","article-title":"Semantic reconstruction of continuous language from non-invasive brain recordings","volume":"26","author":"Tang","year":"2023","journal-title":"Nat. Neurosci."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"393","DOI":"10.1038\/s41583-023-00713-w","article-title":"Non-invasive continuous language decoding","volume":"24","author":"Rogers","year":"2023","journal-title":"Nat. Rev. Neurosci."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Takagi, Y., and Nishimoto, S. (2023, January 18\u201322). High-Resolution Image Reconstruction with Latent Diffusion Models From Human Brain Activity. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01389"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Zhang, J., Li, C., Liu, G., Min, M., Wang, C., Li, J., Wang, Y., Yan, H., Zuo, Z., and Huang, W. (2022). A CNN-transformer hybrid approach for decoding visual neural activity into text. Comput. Methods Programs Biomed., 214.","DOI":"10.1016\/j.cmpb.2021.106586"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1007\/s13311-022-01190-2","article-title":"Brain-Computer Interface: Applications to Speech Decoding and Synthesis to Augment Communication","volume":"19","author":"Luo","year":"2022","journal-title":"Neurotherapeutics"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"1017","DOI":"10.1097\/WNR.0000000000000461","article-title":"Focal atrophy of the hypothalamus associated with third ventricle enlargement in autism spectrum disorder","volume":"26","author":"Wolfe","year":"2015","journal-title":"NeuroReport"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"140","DOI":"10.1006\/nimg.2001.0795","article-title":"Bayesian modeling of the hemodynamic response function in BOLD fMRI","volume":"14","author":"Fahrmeir","year":"2001","journal-title":"NeuroImage"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"3508","DOI":"10.1093\/cercor\/bhw157","article-title":"The Human Brainnetome Atlas: A New Brain Atlas Based on Connectional Architecture","volume":"26","author":"Fan","year":"2016","journal-title":"Cereb. Cortex"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Sun, M., Li, J., Feng, H., Gou, W., Shen, H., Tang, J., Yang, Y., and Ye, J. (2020, January 25\u201329). Multi-modal Fusion Using Spatio-temporal and Static Features for Group Emotion Recognition. Proceedings of the 2020 International Conference on Multimodal Interaction, Online.","DOI":"10.1145\/3382507.3417971"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Zhu, B., Lan, X., Guo, X., Barner, K.E., and Boncelet, C. (2020, January 25\u201329). Multi-rate Attention Based GRU Model for Engagement Prediction. Proceedings of the 2020 International Conference on Multimodal Interaction, Online.","DOI":"10.1145\/3382507.3417965"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Oliveira, L.M.R., Shuen, L.C., Da Cruz, A.K.B.S., and Soares, C.D.S. (2023, January 25\u201329). Summarization of Educational Videos with Transformers Networks. Proceedings of the 29th Brazilian Symposium on Multimedia and the Web, Online.","DOI":"10.1145\/3617023.3617042"},{"key":"ref_48","unstructured":"Dror, R., Shlomov, S., and Reichart, R. (August, January 28). Deep Dominance - How to Properly Compare Deep Neural Models. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1007\/978-3-319-73848-2_3","article-title":"An Optimal Transportation Approach for Assessing Almost Stochastic Order","volume":"Volume 142","author":"Gil","year":"2018","journal-title":"The Mathematics of the Uncertain"},{"key":"ref_50","unstructured":"Ulmer, D., Hardmeier, C., and Frellsen, J. (2022). deep-significance - Easy and Meaningful Statistical Significance Testing in the Age of Neural Networks. arXiv."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"71","DOI":"10.1007\/s11760-023-02707-8","article-title":"Multimodal modelling of human emotion using sound, image and text fusion","volume":"18","author":"Hosseini","year":"2024","journal-title":"SIViP Signal Image Video Process."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Chaminade, T., and Spatola, N. (2022). Perceived facial happiness during conversation correlates with insular and hypothalamus activity for humans, not robots. Front. Psychol., 13.","DOI":"10.3389\/fpsyg.2022.871676"},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"577","DOI":"10.1016\/S1053-8100(03)00076-X","article-title":"When the self represents the other: A new cognitive neuroscience view on psychological identification","volume":"12","author":"Decety","year":"2003","journal-title":"Conscious. Cogn."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Alsabhan, W. (2023). Human\u2013Computer Interaction with a Real-Time Speech Emotion Recognition with Ensembling Techniques 1D Convolution Neural Network and Attention. Sensors, 23.","DOI":"10.3390\/s23031386"}],"container-title":["Multimodal Technologies and Interaction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2414-4088\/9\/4\/31\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T17:06:27Z","timestamp":1760029587000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2414-4088\/9\/4\/31"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,31]]},"references-count":55,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2025,4]]}},"alternative-id":["mti9040031"],"URL":"https:\/\/doi.org\/10.3390\/mti9040031","relation":{},"ISSN":["2414-4088"],"issn-type":[{"type":"electronic","value":"2414-4088"}],"subject":[],"published":{"date-parts":[[2025,3,31]]}}}