{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,4]],"date-time":"2026-08-04T17:40:01Z","timestamp":1785865201236,"version":"3.56.0"},"reference-count":51,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2024,6,7]],"date-time":"2024-06-07T00:00:00Z","timestamp":1717718400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Shenzhen Science and Technology Research and Development Fund for Sustainable Development Project","award":["KCXFZ20201221173613036"],"award-info":[{"award-number":["KCXFZ20201221173613036"]}]},{"name":"Shenzhen Science and Technology Research and Development Fund for Sustainable Development Project","award":["B2023078"],"award-info":[{"award-number":["B2023078"]}]},{"name":"Shenzhen Science and Technology Research and Development Fund for Sustainable Development Project","award":["RKX20220705152815035"],"award-info":[{"award-number":["RKX20220705152815035"]}]},{"DOI":"10.13039\/501100003785","name":"Medical Scientific Research Foundation of Guangdong Province of China","doi-asserted-by":"publisher","award":["KCXFZ20201221173613036"],"award-info":[{"award-number":["KCXFZ20201221173613036"]}],"id":[{"id":"10.13039\/501100003785","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003785","name":"Medical Scientific Research Foundation of Guangdong Province of China","doi-asserted-by":"publisher","award":["B2023078"],"award-info":[{"award-number":["B2023078"]}],"id":[{"id":"10.13039\/501100003785","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003785","name":"Medical Scientific Research Foundation of Guangdong Province of China","doi-asserted-by":"publisher","award":["RKX20220705152815035"],"award-info":[{"award-number":["RKX20220705152815035"]}],"id":[{"id":"10.13039\/501100003785","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Shenzhen Soft Science Research Program Project","award":["KCXFZ20201221173613036"],"award-info":[{"award-number":["KCXFZ20201221173613036"]}]},{"name":"Shenzhen Soft Science Research Program Project","award":["B2023078"],"award-info":[{"award-number":["B2023078"]}]},{"name":"Shenzhen Soft Science Research Program Project","award":["RKX20220705152815035"],"award-info":[{"award-number":["RKX20220705152815035"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Depression is a major psychological disorder with a growing impact worldwide. Traditional methods for detecting the risk of depression, predominantly reliant on psychiatric evaluations and self-assessment questionnaires, are often criticized for their inefficiency and lack of objectivity. Advancements in deep learning have paved the way for innovations in depression risk detection methods that fuse multimodal data. This paper introduces a novel framework, the Audio, Video, and Text Fusion-Three Branch Network (AVTF-TBN), designed to amalgamate auditory, visual, and textual cues for a comprehensive analysis of depression risk. Our approach encompasses three dedicated branches\u2014Audio Branch, Video Branch, and Text Branch\u2014each responsible for extracting salient features from the corresponding modality. These features are subsequently fused through a multimodal fusion (MMF) module, yielding a robust feature vector that feeds into a predictive modeling layer. To further our research, we devised an emotion elicitation paradigm based on two distinct tasks\u2014reading and interviewing\u2014implemented to gather a rich, sensor-based depression risk detection dataset. The sensory equipment, such as cameras, captures subtle facial expressions and vocal characteristics essential for our analysis. The research thoroughly investigates the data generated by varying emotional stimuli and evaluates the contribution of different tasks to emotion evocation. During the experiment, the AVTF-TBN model has the best performance when the data from the two tasks are simultaneously used for detection, where the F1 Score is 0.78, Precision is 0.76, and Recall is 0.81. Our experimental results confirm the validity of the paradigm and demonstrate the efficacy of the AVTF-TBN model in detecting depression risk, showcasing the crucial role of sensor-based data in mental health detection.<\/jats:p>","DOI":"10.3390\/s24123714","type":"journal-article","created":{"date-parts":[[2024,6,7]],"date-time":"2024-06-07T08:05:17Z","timestamp":1717747517000},"page":"3714","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":59,"title":["Multimodal Sensing for Depression Risk Detection: Integrating Audio, Video, and Text Data"],"prefix":"10.3390","volume":"24","author":[{"given":"Zhenwei","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Biomedical Engineering, Health Science Center, Shenzhen University, Shenzhen 518060, China"},{"name":"Guangdong Provincial Key Laboratory of Biomedical Measurements and Ultrasound Imaging, Shenzhen 518060, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5749-2964","authenticated-orcid":false,"given":"Shengming","family":"Zhang","sequence":"additional","affiliation":[{"name":"Affiliated Mental Health Center, Southern University of Science and Technology, Shenzhen 518055, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dong","family":"Ni","sequence":"additional","affiliation":[{"name":"School of Biomedical Engineering, Health Science Center, Shenzhen University, Shenzhen 518060, China"},{"name":"Guangdong Provincial Key Laboratory of Biomedical Measurements and Ultrasound Imaging, Shenzhen 518060, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhaoguo","family":"Wei","sequence":"additional","affiliation":[{"name":"Shenzhen Kangning Hospital, Shenzhen 518020, China"},{"name":"Shenzhen Mental Health Center, Shenzhen 518020, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kongjun","family":"Yang","sequence":"additional","affiliation":[{"name":"Shenzhen Kangning Hospital, Shenzhen 518020, China"},{"name":"Shenzhen Mental Health Center, Shenzhen 518020, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shan","family":"Jin","sequence":"additional","affiliation":[{"name":"Shenzhen Kangning Hospital, Shenzhen 518020, China"},{"name":"Shenzhen Mental Health Center, Shenzhen 518020, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9895-6163","authenticated-orcid":false,"given":"Gan","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Biomedical Engineering, Health Science Center, Shenzhen University, Shenzhen 518060, China"},{"name":"Guangdong Provincial Key Laboratory of Biomedical Measurements and Ultrasound Imaging, Shenzhen 518060, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhen","family":"Liang","sequence":"additional","affiliation":[{"name":"School of Biomedical Engineering, Health Science Center, Shenzhen University, Shenzhen 518060, China"},{"name":"Guangdong Provincial Key Laboratory of Biomedical Measurements and Ultrasound Imaging, Shenzhen 518060, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Li","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Biomedical Engineering, Health Science Center, Shenzhen University, Shenzhen 518060, China"},{"name":"Guangdong Provincial Key Laboratory of Biomedical Measurements and Ultrasound Imaging, Shenzhen 518060, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Linling","family":"Li","sequence":"additional","affiliation":[{"name":"School of Biomedical Engineering, Health Science Center, Shenzhen University, Shenzhen 518060, China"},{"name":"Guangdong Provincial Key Laboratory of Biomedical Measurements and Ultrasound Imaging, Shenzhen 518060, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3388-4928","authenticated-orcid":false,"given":"Huijun","family":"Ding","sequence":"additional","affiliation":[{"name":"School of Biomedical Engineering, Health Science Center, Shenzhen University, Shenzhen 518060, China"},{"name":"Guangdong Provincial Key Laboratory of Biomedical Measurements and Ultrasound Imaging, Shenzhen 518060, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhiguo","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen 518055, China"},{"name":"Peng Cheng Laboratory, Shenzhen 518055, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianhong","family":"Wang","sequence":"additional","affiliation":[{"name":"Shenzhen Kangning Hospital, Shenzhen 518020, China"},{"name":"Shenzhen Mental Health Center, Shenzhen 518020, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,6,7]]},"reference":[{"key":"ref_1","unstructured":"World Health Organization (2023, December 30). Depressive Disorder (Depression). Available online: https:\/\/www.who.int\/zh\/news-room\/fact-sheets\/detail\/depression."},{"key":"ref_2","unstructured":"Institute of Health Metrics and Evaluation (2023, December 30). Global Health Data Exchange (GHDx). Available online: https:\/\/vizhub.healthdata.org\/gbd-results."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Perez, J.E., and Riggio, R.E. (2003). Nonverbal social skills and psychopathology. Nonverbal Behavior in Clinical Settings, Oxford University Press.","DOI":"10.1093\/med:psych\/9780195141092.003.0002"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"319","DOI":"10.1037\/h0036706","article-title":"Nonverbal cues for depression","volume":"83","author":"Waxer","year":"1974","journal-title":"J. Abnorm. Psychol."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1016\/j.specom.2015.03.004","article-title":"A review of depression and suicide risk assessment using speech analysis","volume":"71","author":"Cummins","year":"2015","journal-title":"Speech Commun."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"379","DOI":"10.1016\/S0272-7358(98)00104-4","article-title":"Social skills deficits associated with depression","volume":"20","author":"Segrin","year":"2000","journal-title":"Clin. Psychol. Rev."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"181","DOI":"10.1016\/j.psychres.2010.04.011","article-title":"Analysis of syntax and word use to predict successful participation in guided self-help for anxiety and depression","volume":"179","author":"Zinken","year":"2010","journal-title":"Psychiatry Res."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"464","DOI":"10.1176\/ajp.145.4.464","article-title":"Diagnostic classification through content analysis of patients\u2019 speech","volume":"145","author":"Oxman","year":"1988","journal-title":"Am. J. Psychiatry"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Yang, L., Jiang, D., Xia, X., Pei, E., Oveneke, M.C., and Sahli, H. (2017, January 23). Multimodal measurement of depression using deep learning models. Proceedings of the 7th Annual Workshop on Audio\/Visual Emotion Challenge, Mountain View, CA, USA.","DOI":"10.1145\/3133944.3133948"},{"key":"ref_10","unstructured":"Gratch, J., Artstein, R., Lucas, G.M., Stratou, G., Scherer, S., Nazarian, A., Wood, R., Boberg, J., DeVault, D., and Marsella, S. (2014, January 26\u201331). The distress analysis interview corpus of human and computer interviews. Proceedings of the 2014 International Conference on Language Resources and Evaluation, Reykjavik, Iceland."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Valstar, M., Schuller, B., Smith, K., Eyben, F., Jiang, B., Bilakhia, S., Schnieder, S., Cowie, R., and Pantic, M. (2013, January 21). Avec 2013: The continuous audio\/visual emotion and depression recognition challenge. Proceedings of the 3rd ACM International Workshop on Audio\/Visual Emotion Challenge, Barcelona, Spain.","DOI":"10.1145\/2512530.2512533"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"228","DOI":"10.1016\/j.jvcir.2018.11.003","article-title":"Facial expression video analysis for depression detection in Chinese patients","volume":"57","author":"Wang","year":"2018","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_13","unstructured":"Mehrabian, A., and Russell, J.A. (1974). An Approach to Environmental Psychology, MIT Press."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"641","DOI":"10.1016\/j.imavis.2013.12.007","article-title":"Nonverbal social withdrawal in depression: Evidence from manual and automatic analyses","volume":"32","author":"Girard","year":"2014","journal-title":"Image Vis. Comput."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Alghowinem, S., Goecke, R., Wagner, M., Parker, G., and Breakspear, M. (2013, January 15\u201318). Eye movement analysis for depression detection. Proceedings of the 2013 IEEE International Conference on Image Processing, Melbourne, VIC, Australia.","DOI":"10.1109\/ICIP.2013.6738869"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Jan, A., Meng, H., Gaus, Y.F.A., Zhang, F., and Turabzadeh, S. (2014, January 7). Automatic depression scale prediction using facial expression dynamics and regression. Proceedings of the 4th International Workshop on Audio\/Visual Emotion Challenge, Orlando, FL, USA.","DOI":"10.1145\/2661806.2661812"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"971","DOI":"10.1109\/TPAMI.2002.1017623","article-title":"Multiresolution gray-scale and rotation invariant texture classification with local binary patterns","volume":"24","author":"Ojala","year":"2002","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1339","DOI":"10.1007\/s11517-021-02358-2","article-title":"AutoDep: Automatic depression detection using facial expressions based on linear binary pattern descriptor","volume":"59","author":"Tadalagi","year":"2021","journal-title":"Med. Biol. Eng. Comput."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"915","DOI":"10.1109\/TPAMI.2007.1110","article-title":"Dynamic texture recognition using local binary patterns with an application to facial expressions","volume":"29","author":"Zhao","year":"2007","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1476","DOI":"10.1109\/TMM.2018.2877129","article-title":"Automatic depression analysis using dynamic facial appearance descriptor and dirichlet process fisher encoding","volume":"21","author":"He","year":"2018","journal-title":"IEEE Trans. Multimed."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"239","DOI":"10.1109\/TAFFC.2018.2870398","article-title":"Integrating deep and shallow models for multi-modal depression analysis\u2014Hybrid architectures","volume":"12","author":"Yang","year":"2018","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_22","first-page":"525","article-title":"Dynamic multimodal measurement of depression severity using deep autoencoding","volume":"22","author":"Hammal","year":"2017","journal-title":"IEEE J. Biomed. Health Inform."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"578","DOI":"10.1109\/TAFFC.2017.2650899","article-title":"Automated depression diagnosis based on deep networks to encode facial appearance and dynamics","volume":"9","author":"Zhu","year":"2017","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1016\/j.neucom.2020.10.015","article-title":"Automatic depression recognition using CNN with attention mechanism from videos","volume":"422","author":"He","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Song, S., Shen, L., and Valstar, M. (2018, January 15\u201319). Human behaviour-based automatic depression analysis using hand-crafted statistics and deep learned spectral features. Proceedings of the 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018), Xi\u2019an, China.","DOI":"10.1109\/FG.2018.00032"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"117512","DOI":"10.1016\/j.eswa.2022.117512","article-title":"Depressioner: Facial dynamic representation for automatic depression level prediction","volume":"204","author":"Niu","year":"2022","journal-title":"Expert Syst. Appl."},{"key":"ref_27","unstructured":"Xu, J., Song, S., Kusumam, K., Gunes, H., and Valstar, M. (2021). Two-stage temporal modelling framework for video-based depression recognition using graph representation. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"30","DOI":"10.1016\/j.bandc.2004.05.003","article-title":"Voice acoustical measurement of the severity of major depression","volume":"56","author":"Cannizzaro","year":"2004","journal-title":"Brain Cogn."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1109\/TBME.2007.900562","article-title":"Critical analysis of the impact of glottal features in the classification of clinical depression in speech","volume":"55","author":"Moore","year":"2007","journal-title":"IEEE Trans. Biomed. Eng."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Chen, W., Xing, X., Xu, X., Pang, J., and Du, L. (2022). SpeechFormer: A hierarchical efficient framework incorporating the characteristics of speech. arXiv.","DOI":"10.21437\/Interspeech.2022-74"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"684037","DOI":"10.3389\/fnbot.2021.684037","article-title":"Multi-head attention-based long short-term memory for depression detection from speech","volume":"15","author":"Zhao","year":"2021","journal-title":"Front. Neurorobotics"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zhao, Z., Li, Q., Cummins, N., Liu, B., Wang, H., Tao, J., and Schuller, B. (2020, January 25\u201329). Hybrid network feature extraction for depression assessment from speech. Proceedings of the INTERSPEECH 2020, Shanghai, China.","DOI":"10.21437\/Interspeech.2020-2396"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"116076","DOI":"10.1016\/j.eswa.2021.116076","article-title":"Audio based depression detection using Convolutional Autoencoder","volume":"189","author":"Sardari","year":"2022","journal-title":"Expert Syst. Appl."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Hosseini-Saravani, S.H., Besharati, S., Calvo, H., and Gelbukh, A. (2020). Depression detection in social media using a psychoanalytical technique for feature extraction and a cognitive based classifier. Mexican International Conference on Artificial Intelligence, Springer.","DOI":"10.1007\/978-3-030-60887-3_25"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"1121","DOI":"10.1080\/02699930441000030","article-title":"Language use of depressed and depression-vulnerable college students","volume":"18","author":"Rude","year":"2004","journal-title":"Cogn. Emot."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Chiong, R., Budhi, G.S., Dhakal, S., and Chiong, F. (2021). A textual-based featuring approach for depression detection using machine learning classifiers and social media texts. Comput. Biol. Med., 135.","DOI":"10.1016\/j.compbiomed.2021.104499"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Jang, B., Kim, M., Harerimana, G., Kang, S.-U., and Kim, J.W. (2020). Bi-LSTM model to increase accuracy in text classification: Combining Word2vec CNN and attention mechanism. Appl. Sci., 10.","DOI":"10.3390\/app10175841"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1109\/TCSS.2022.3154442","article-title":"Ensemble hybrid learning methods for automated depression detection","volume":"10","author":"Ansari","year":"2022","journal-title":"IEEE Trans. Comput. Soc. Syst."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"668","DOI":"10.1109\/TCDS.2017.2721552","article-title":"Artificial intelligent system for automatic depression level analysis through visual and vocal expressions","volume":"10","author":"Jan","year":"2017","journal-title":"IEEE Trans. Cogn. Dev. Syst."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"294","DOI":"10.1109\/TAFFC.2020.3031345","article-title":"Multimodal spatiotemporal representation for automatic depression level detection","volume":"14","author":"Niu","year":"2020","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Dai, Z., Li, Q., Shang, Y., and Wang, X.A. (2023, January 24\u201326). Depression Detection Based on Facial Expression, Audio and Gait. Proceedings of the 2023 IEEE 6th Information Technology, Networking, Electronic and Automation Control Conference (ITNEC), Chongqing, China.","DOI":"10.1109\/ITNEC56291.2023.10082163"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Solieman, H., and Pustozerov, E.A. (2021, January 26\u201329). The detection of depression using multimodal models based on text and voice quality features. Proceedings of the 2021 IEEE Conference of Russian Young Researchers in Electrical and Electronic Engineering (ElConRus), St. Petersburg, Russia.","DOI":"10.1109\/ElConRus51938.2021.9396540"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Shen, Y., Yang, H., and Lin, L. (2022, January 23\u201327). Automatic depression detection: An emotional audio-textual corpus and a GRU\/BiLSTM-based model. Proceedings of the ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore.","DOI":"10.1109\/ICASSP43922.2022.9746569"},{"key":"ref_44","unstructured":"Arandjelovic, R., Gronat, P., Torii, A., Pajdla, T., and Sivic, J. (July, January 26). NetVLAD: CNN architecture for weakly supervised place recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Paradise, NV, USA."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Fang, M., Peng, S., Liang, Y., Hung, C.-C., and Liu, S. (2023). A multimodal fusion model with multi-level attention mechanism for depression detection. Biomed. Signal Process. Control, 82.","DOI":"10.1016\/j.bspc.2022.104561"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Zheng, W., Yan, L., Gou, C., and Wang, F.-Y. (2020, January 6\u201310). Graph attention model embedded with multi-modal knowledge for depression detection. Proceedings of the 2020 IEEE International Conference on Multimedia and Expo (ICME), London, UK.","DOI":"10.1109\/ICME46284.2020.9102872"},{"key":"ref_47","unstructured":"Sudhan, H.M., and Kumar, S.S. (2021, January 13\u201314). Multimodal Depression Severity Detection Using Deep Neural Networks and Depression Assessment Scale. Proceedings of the International Conference on Computational Intelligence and Data Engineering: ICCIDE 2021, Vijayawada, India."},{"key":"ref_48","first-page":"6000","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Zhang, S., Zhao, Z., and Guan, C. (2023, January 17\u201324). Multimodal continuous emotion recognition: A technical report for abaw5. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPRW59228.2023.00611"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Sun, H., Wang, H., Liu, J., Chen, Y.-W., and Lin, L. (2022, January 10\u201314). CubeMLP: An MLP-based model for multimodal sentiment analysis and depression estimation. Proceedings of the 30th ACM International Conference on Multimedia, Lisboa, Portugal.","DOI":"10.1145\/3503161.3548025"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Rajan, V., Brutti, A., and Cavallaro, A. (2022, January 23\u201327). Is cross-attention preferable to self-attention for multi-modal emotion recognition?. Proceedings of the ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore.","DOI":"10.1109\/ICASSP43922.2022.9746924"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/12\/3714\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:55:22Z","timestamp":1760108122000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/12\/3714"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,7]]},"references-count":51,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2024,6]]}},"alternative-id":["s24123714"],"URL":"https:\/\/doi.org\/10.3390\/s24123714","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,7]]}}}