{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T14:38:42Z","timestamp":1780583922786,"version":"3.54.1"},"reference-count":50,"publisher":"Wiley","license":[{"start":{"date-parts":[[2024,1,29]],"date-time":"2024-01-29T00:00:00Z","timestamp":1706486400000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["82072008"],"award-info":[{"award-number":["82072008"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["N2224001-10"],"award-info":[{"award-number":["N2224001-10"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["82072008"],"award-info":[{"award-number":["82072008"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["N2224001-10"],"award-info":[{"award-number":["N2224001-10"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Shenzhen Jingmei Health Technology Company Ltd","award":["82072008"],"award-info":[{"award-number":["82072008"]}]},{"name":"Shenzhen Jingmei Health Technology Company Ltd","award":["N2224001-10"],"award-info":[{"award-number":["N2224001-10"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["International Journal of Intelligent Systems"],"published-print":{"date-parts":[[2024,1,29]]},"abstract":"<jats:p>Background and Objective. Currently, depression is a widespread global issue that imposes a significant burden and disability on individuals, families, and society. Deep learning (DL) has emerged as a valuable approach for automatically detecting depression by extracting cues from audiovisual data and making a diagnosis. PHQ-8 is considered a validated diagnostic tool for depressive disorders in clinical studies, and the objective of this experiment is to improve the accuracy of PHQ-8 prediction. Furthermore, this paper aims to demonstrate the effectiveness of expert knowledge in depression diagnosis and discuss a novel multimodal network architecture. Methods. This research paper focuses on multimodal depression analysis, proposing a flexible parallel transformer (FPT) model capable of extracting data from three distinct modalities (i.e., one video and two audio descriptors). The FPT-Former model incorporates three paths, each using expert-knowledge-based descriptors from one modality as inputs. These descriptors are represented into 32 features by the encoder part of a transformer module, and these features are fused to realize the final regression of PHQ-8 score. The extended distress analysis interview corpus (E-DAIC) is an expansion of WOZ-DAIC which comprises semiclinical interviews intended to assist in the diagnosis of psychological distress conditions. It encompasses a sample size of 275 participants, and in this study, it was utilized to test the model in a way of 10-fold cross-validation. Results. The FPT presented herein achieved comparable performance to the state-of-the-art works, with a root mean square error (RMSE) of 4.80 and a mean absolute error (MAE) of 4.58. The ablation experiments demonstrate that the three-modality-fused model outperforms other two-modality-fused and single-modality models. While using a PHQ-8 score threshold of 10, the accuracy of the depression classification is 0.79. Conclusions. Leveraging the strength of expert-knowledge-based multimodal measures and parallel transformer structure, the FPT model exhibits promising performance in depression detection. This model improved the accuracy of depression diagnosis through audio and video, and it also proved the effectiveness of using expert-knowledge in the diagnosis of depression. The traits of flexible structure, high predictive efficiency, and secure privacy protection make our model a promotable intelligent system in mental healthcare.<\/jats:p>","DOI":"10.1155\/2024\/1564574","type":"journal-article","created":{"date-parts":[[2024,1,29]],"date-time":"2024-01-29T23:20:09Z","timestamp":1706570409000},"page":"1-13","source":"Crossref","is-referenced-by-count":23,"title":["FPT-Former: A Flexible Parallel Transformer of Recognizing Depression by Using Audiovisual Expert-Knowledge-Based Multimodal Measures"],"prefix":"10.1155","volume":"2024","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-5991-9926","authenticated-orcid":true,"given":"Yifu","family":"Li","sequence":"first","affiliation":[{"name":"College of Medicine and Biological Information Engineering, Northeastern University, Shenyang, China"},{"name":"Key Laboratory of Intelligent Computing in Medical Image, Ministry of Education, Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0681-5812","authenticated-orcid":true,"given":"Xueping","family":"Yang","sequence":"additional","affiliation":[{"name":"Department of Psychology, The People\u2019s Hospital of Liaoning Province, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-8449-8935","authenticated-orcid":true,"given":"Meng","family":"Zhao","sequence":"additional","affiliation":[{"name":"College of Medicine and Biological Information Engineering, Northeastern University, Shenyang, China"},{"name":"Key Laboratory of Intelligent Computing in Medical Image, Ministry of Education, Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-3617-7814","authenticated-orcid":true,"given":"Zihao","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Medicine and Biological Information Engineering, Northeastern University, Shenyang, China"},{"name":"Key Laboratory of Intelligent Computing in Medical Image, Ministry of Education, Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3868-0593","authenticated-orcid":true,"given":"Yudong","family":"Yao","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, Stevens Institute of Technology, Hoboken, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9563-721X","authenticated-orcid":true,"given":"Wei","family":"Qian","sequence":"additional","affiliation":[{"name":"College of Medicine and Biological Information Engineering, Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0977-1939","authenticated-orcid":true,"given":"Shouliang","family":"Qi","sequence":"additional","affiliation":[{"name":"College of Medicine and Biological Information Engineering, Northeastern University, Shenyang, China"},{"name":"Key Laboratory of Intelligent Computing in Medical Image, Ministry of Education, Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","reference":[{"key":"1","doi-asserted-by":"publisher","DOI":"10.1016\/s0140-6736(18)31612-x"},{"key":"2","doi-asserted-by":"crossref","DOI":"10.9783\/9780812290882","volume-title":"Depression: Causes and Treatment","author":"A. T. Beck","year":"2009"},{"key":"3","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pmed.0030442"},{"key":"4","doi-asserted-by":"publisher","DOI":"10.1001\/jama.289.23.3095"},{"key":"5","volume-title":"Artificial Intelligence in Behavioral and Mental Health Care","author":"D. D. Luxton","year":"2015"},{"key":"6","doi-asserted-by":"publisher","DOI":"10.1017\/s0033291719000151"},{"key":"7","doi-asserted-by":"publisher","DOI":"10.1109\/taffc.2017.2724035"},{"key":"8","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-022-10290-6"},{"key":"9","doi-asserted-by":"publisher","DOI":"10.1109\/taffc.2020.3022732"},{"key":"10","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.10.015"},{"key":"11","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2022.09.025"},{"key":"12","first-page":"53","article-title":"Multimodal measurement of depression using deep learning models","author":"L. Yang"},{"key":"13","article-title":"Facial action coding system","author":"P. Ekman","year":"1978","journal-title":"Environmental Psychology and Nonverbal Behavior"},{"key":"14","doi-asserted-by":"publisher","DOI":"10.1109\/10.846676"},{"key":"15","doi-asserted-by":"publisher","DOI":"10.1109\/taffc.2015.2457417"},{"key":"16","first-page":"1","article-title":"A novel approach for MFCC feature extraction","author":"M. A. Hossan"},{"key":"17","first-page":"6247","article-title":"Automatic depression detection: an emotional audio-textual corpus and a GRU\/BiLSTM-based model","author":"Y. Shen"},{"key":"18","article-title":"Modma dataset: a multi-modal open dataset for mental-disorder analysis","author":"H. Cai","year":"2020"},{"key":"19","first-page":"3","article-title":"AVEC 2019 workshop and challenge: state-of-mind, detecting depression with AI, and cross-cultural affect recognition","author":"F. Ringeval"},{"key":"20","article-title":"Attention is all you need","volume":"30","author":"A. Vaswani","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"21","doi-asserted-by":"publisher","DOI":"10.1016\/j.jad.2008.06.026"},{"key":"22","doi-asserted-by":"publisher","DOI":"10.1016\/j.jad.2022.11.060"},{"key":"23","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2023.05.041"},{"key":"24","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2022.117512"},{"key":"25","doi-asserted-by":"publisher","DOI":"10.1109\/access.2022.3225684"},{"key":"26","doi-asserted-by":"publisher","DOI":"10.1109\/access.2022.3223705"},{"key":"27","doi-asserted-by":"publisher","DOI":"10.1007\/s13755-022-00197-5"},{"key":"28","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2023.3296318"},{"key":"29","doi-asserted-by":"publisher","DOI":"10.1109\/taffc.2022.3181210"},{"key":"30","doi-asserted-by":"publisher","DOI":"10.1109\/jiot.2023.3283616"},{"key":"31","volume-title":"The Distress Analysis Interview Corpus of Human and Computer Interviews","author":"J. Gratch","year":"2014"},{"key":"32","doi-asserted-by":"publisher","DOI":"10.1016\/j.bspc.2023.105248"},{"key":"33","doi-asserted-by":"publisher","DOI":"10.1002\/int.22704"},{"key":"34","doi-asserted-by":"publisher","DOI":"10.3390\/drones7020081"},{"key":"35","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2022.05.024"},{"key":"36","doi-asserted-by":"publisher","DOI":"10.1016\/j.aei.2023.102007"},{"key":"37","doi-asserted-by":"publisher","DOI":"10.1016\/j.compbiomed.2023.106734"},{"key":"38","first-page":"59","article-title":"Openface 2.0: facial behavior analysis toolkit","author":"T. Baltrusaitis"},{"key":"39","doi-asserted-by":"publisher","DOI":"10.1109\/34.908962"},{"key":"40","first-page":"1","article-title":"Using information theoretic vector quantization for inverted MFCC based speaker verification","author":"S. Memon"},{"key":"41","article-title":"Depitch and the role of fundamental frequency in speaker recognition","author":"R. Zilea"},{"key":"42","first-page":"1459","article-title":"Opensmile: the munich versatile and fast open-source audio feature extractor","author":"F. Eyben"},{"key":"43","doi-asserted-by":"crossref","DOI":"10.21437\/Interspeech.2018-2522","volume-title":"Detecting Depression with Audio\/Text Sequence Modeling of Interviews","author":"T. Al Hanai","year":"2018"},{"key":"44","first-page":"508","article-title":"Autoencoder based on cepstrum separation to detect depression from speech","author":"Y. Zhang"},{"key":"45","doi-asserted-by":"publisher","DOI":"10.1109\/access.2020.2970496"},{"key":"46","doi-asserted-by":"publisher","DOI":"10.1109\/TCDS.2023.3273614"},{"key":"47","doi-asserted-by":"publisher","DOI":"10.1016\/j.bspc.2022.104561"},{"key":"48","first-page":"3","article-title":"Avec 2014: 3d dimensional affect and depression recognition challenge","author":"M. Valstar"},{"key":"49","doi-asserted-by":"publisher","DOI":"10.1002\/acr.20556"},{"key":"50","first-page":"407","article-title":"Sensemood: depression detection on social media","author":"C. Lin"}],"container-title":["International Journal of Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/downloads.hindawi.com\/journals\/ijis\/2024\/1564574.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/ijis\/2024\/1564574.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/ijis\/2024\/1564574.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,1,29]],"date-time":"2024-01-29T23:20:24Z","timestamp":1706570424000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.hindawi.com\/journals\/ijis\/2024\/1564574\/"}},"subtitle":[],"editor":[{"given":"Said","family":"El Kafhali","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2024,1,29]]},"references-count":50,"alternative-id":["1564574","1564574"],"URL":"https:\/\/doi.org\/10.1155\/2024\/1564574","relation":{},"ISSN":["1098-111X","0884-8173"],"issn-type":[{"value":"1098-111X","type":"electronic"},{"value":"0884-8173","type":"print"}],"subject":[],"published":{"date-parts":[[2024,1,29]]}}}