{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T01:33:47Z","timestamp":1785288827528,"version":"3.55.0"},"reference-count":55,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2026,3,16]],"date-time":"2026-03-16T00:00:00Z","timestamp":1773619200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100006769","name":"Russian Science Foundation","doi-asserted-by":"publisher","award":["25-11-00319"],"award-info":[{"award-number":["25-11-00319"]}],"id":[{"id":"10.13039\/501100006769","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["BDCC"],"abstract":"<jats:p>Automated video-based detection of cognitive disorders can enable a scalable non-invasive health monitoring. However, existing methods focus on a single disease and provide limited interpretability, whereas real-world videos often contain co-occurring conditions. We propose a novel unified multi-task method to detect depression and Parkinson\u2019s disease (PD) from in-the-wild video data called DEPART (DEpression and PArkinson\u2019s Recognition Technique). It performs body region extraction, Contrastive Language-Image Pre-training (CLIP)-based visual encoding, Transformer-based temporal modeling, and prototype-aware classification with a gated fusion technique. Gradient-based attention maps are used to visualize task-specific regions that drive predictions. Experiments on the In-the-Wild Speech Medical (WSM) corpus demonstrate competitive performance: the multi-task model achieves Recall of 82.39% for depression and 78.20% for PD, compared with 87.76% and 78.20%, for the best single-task models. The multi-task learning initially increases false positives for healthy persons in the PD subset, mainly due to annotation\u2013modality mismatches, static visual content misinterpreted as motor impairments, and occasional body detection failures. After cleaning the test data, Recall for healthy individuals becomes comparable across models; the multi-task model improves Recall for both depression (from 82.39% to 87.50%) and PD (from 78.20% to 86.14%), suggesting better robustness for real-life clinical applications.<\/jats:p>","DOI":"10.3390\/bdcc10030089","type":"journal-article","created":{"date-parts":[[2026,3,16]],"date-time":"2026-03-16T13:32:07Z","timestamp":1773667927000},"page":"89","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["DEPART: Multi-Task Interpretable Depression and Parkinson\u2019s Disease Detection from In-the-Wild Video Data"],"prefix":"10.3390","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4135-6949","authenticated-orcid":false,"given":"Elena","family":"Ryumina","sequence":"first","affiliation":[{"name":"St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS), 199178 St. Petersburg, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7479-2851","authenticated-orcid":false,"given":"Alexandr","family":"Axyonov","sequence":"additional","affiliation":[{"name":"St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS), 199178 St. Petersburg, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4344-2330","authenticated-orcid":false,"given":"Mikhail","family":"Dolgushin","sequence":"additional","affiliation":[{"name":"St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS), 199178 St. Petersburg, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7935-0569","authenticated-orcid":false,"given":"Dmitry","family":"Ryumin","sequence":"additional","affiliation":[{"name":"St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS), 199178 St. Petersburg, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3424-652X","authenticated-orcid":false,"given":"Alexey","family":"Karpov","sequence":"additional","affiliation":[{"name":"St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS), 199178 St. Petersburg, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,3,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"161","DOI":"10.1186\/s13195-023-01308-4","article-title":"Stress, depression, and risk of dementia\u2014A cohort study in the total population between 18 and 65 years old in Region Stockholm","volume":"15","author":"Wallensten","year":"2023","journal-title":"Alzheimer\u2019s Res. Ther."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"132","DOI":"10.14412\/2074-2711-2021-6-132-138","article-title":"Treatment of noncognitive neuropsychiatric disorders in Alzheimer\u2019s disease","volume":"13","author":"Lokshina","year":"2021","journal-title":"Neurol. Neuropsychiatry Psychosom."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"323","DOI":"10.1038\/nrneurol.2011.60","article-title":"Depression and risk of developing dementia","volume":"7","author":"Byers","year":"2011","journal-title":"Nat. Rev. Neurol."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Sharma, D., Singh, J., Sehra, S.S., and Sehra, S.K. (2024). Demystifying Mental Health by Decoding Facial Action Unit Sequences. Big Data Cogn. Comput., 8.","DOI":"10.3390\/bdcc8070078"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Markitantov, M., Ryumina, E., Kaya, H., and Karpov, A. (2025). Multi-Modal Multi-Task Affective States Recognition Based on Label Encoder Fusion. Proceedings of the Interspeech, ISCA Archive.","DOI":"10.21437\/Interspeech.2025-2060"},{"key":"ref_6","unstructured":"Parikh, A., Sadeghi, M., and Eskofier, B. (2024). Exploring facial biomarkers for depression through temporal analysis of action units. arXiv."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"554","DOI":"10.1109\/TNSRE.2022.3204757","article-title":"Dual-stream multiple instance learning for depression detection with facial expression videos","volume":"31","author":"Shangguan","year":"2022","journal-title":"IEEE Trans. Neural Syst. Rehabil. Eng."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Yang, N., Liu, J., Sun, D., Ding, J., Sun, L., Qi, X., and Yan, W. (2025). Motor Symptoms of Parkinson\u2019s Disease: Critical Markers for Early AI-assisted Diagnosis. Front. Aging Neurosci., 17.","DOI":"10.3389\/fnagi.2025.1602426"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Rangel-Cascajosa, C., Luna-Perej\u00f3n, F., Vicente-Diaz, S., and Dom\u00ednguez-Morales, M. (2025). Gait-Based Parkinson\u2019s Disease Detection Using Recurrent Neural Networks for Wearable Systems. Big Data Cogn. Comput., 9.","DOI":"10.3390\/bdcc9070183"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"105316","DOI":"10.1016\/j.parkreldis.2023.105316","article-title":"Classification and staging of Parkinson\u2019s disease using video-based eye tracking","volume":"110","author":"Brien","year":"2023","journal-title":"Park. Relat. Disord."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Maddage, N.C., Senaratne, R., Low, L.S.A., Lech, M., and Allen, N. (2009). Video-based detection of the clinical depression in adolescents. Proceedings of the Annual International Conference of the IEEE Engineering in Medicine and Biology Society, IEEE.","DOI":"10.1109\/IEMBS.2009.5334815"},{"key":"ref_12","unstructured":"Mu, X., Seyedi, S., Zheng, I., Jiang, Z., Chen, L., Omofojoye, B., Hershenberg, R., Levey, A.I., Clifford, G.D., and Dodge, H.H. (2024). Detecting Cognitive Impairment and Psychological Well-being among Older Adults Using Facial, Acoustic, Linguistic, and Cardiovascular Patterns Derived from Remote Conversations. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"894","DOI":"10.1109\/TAFFC.2020.2973984","article-title":"Modeling, recognizing, and explaining apparent personality from videos","volume":"13","author":"Escalante","year":"2020","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Williamson, J.R., Godoy, E., Cha, M., Schwarzentruber, A., Khorrami, P., Gwon, Y., Kung, H.T., Dagli, C., and Quatieri, T.F. (2016). Detecting depression using vocal, facial and semantic communication cues. Proceedings of the International Workshop on Audio\/Visual Emotion Challenge, ACM.","DOI":"10.1145\/2988257.2988263"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Song, S., Shen, L., and Valstar, M. (2018). Human behaviour-based automatic depression analysis using hand-crafted statistics and deep learned spectral features. Proceedings of the IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018), IEEE.","DOI":"10.1109\/FG.2018.00032"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Wei, P.C., Peng, K., Roitberg, A., Yang, K., Zhang, J., and Stiefelhagen, R. (2022). Multi-modal depression estimation based on sub-attentional fusion. Proceedings of the European Conference on Computer Vision (ECCV), Springer.","DOI":"10.1007\/978-3-031-25075-0_42"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Gimeno-G\u00f3mez, D., Bucur, A.M., Cosma, A., Mart\u00ednez-Hinarejos, C.D., and Rosso, P. (2024). Reading between the frames: Multi-modal depression detection in videos from non-verbal cues. Proceedings of the European Conference on Information Retrieval, Springer.","DOI":"10.1007\/978-3-031-56027-9_12"},{"key":"ref_18","unstructured":"Jaegle, A., Gimeno, F., Brock, A., Vinyals, O., Zisserman, A., and Carreira, J. (2021). Perceiver: General perception with iterative attention. Proceedings of the International Conference on Machine Learning, PMLR."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Zhang, S., Ni, D., Wei, Z., Yang, K., Jin, S., Huang, G., Liang, Z., Zhang, L., and Li, L. (2024). Multimodal sensing for depression risk detection: Integrating audio, video, and text data. Sensors, 24.","DOI":"10.3390\/s24123714"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Kyprakis, I., Skaramagkas, V., Boura, I., Karamanis, G., Fotiadis, D.I., Kefalopoulou, Z., Spanaki, C., and Tsiknakis, M. (2025). A Deep Learning approach for Depressive Symptoms assessment in Parkinson\u2019s disease patients using facial videos. arXiv.","DOI":"10.1109\/EMBC58623.2025.11253137"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., and Hu, H. (2022). Video swin transformer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE.","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"12226","DOI":"10.1609\/aaai.v36i11.21483","article-title":"D-vlog: Multimodal vlog dataset for depression detection","volume":"Volume 36","author":"Yoon","year":"2022","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Dolgushin, M., Guseva, D., and Karpov, A. (2025). Investigation of Explainable Multimodal Methods for Detecting Mental Disorders. Proceedings of the International Conference on Speech and Computer (SPECOM), Springer.","DOI":"10.1007\/978-3-032-07956-5_12"},{"key":"ref_24","first-page":"55140","article-title":"Youtubepd: A multimodal benchmark for Parkinson\u2019s disease analysis","volume":"36","author":"Zhou","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Calvo-Ariza, N.R., G\u00f3mez-G\u00f3mez, L.F., and Orozco-Arroyave, J.R. (2022). Classical FE Analysis to Classify Parkinson\u2019s Disease Patients. Electronics, 11.","DOI":"10.3390\/electronics11213533"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"24","DOI":"10.1007\/s44163-025-00241-9","article-title":"A review of machine learning and deep learning for Parkinson\u2019s disease detection","volume":"5","author":"Rabie","year":"2025","journal-title":"Discov. Artif. Intell."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Valstar, M., Schuller, B., Smith, K., Eyben, F., Jiang, B., Bilakhia, S., Schnieder, S., Cowie, R., and Pantic, M. (2013). AVEC 2013: The continuous audio\/visual emotion and depression recognition challenge. Proceedings of the ACM International Workshop on Audio\/Visual Emotion Challenge, ACM.","DOI":"10.1145\/2512530.2512533"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Valstar, M., Schuller, B., Smith, K., Almaev, T., Eyben, F., Krajewski, J., Cowie, R., and Pantic, M. (2014). AVEC 2014: 3D Dimensional Affect and Depression Recognition Challenge. Proceedings of the International Workshop on Audio\/Visual Emotion Challenge, Association for Computing Machinery.","DOI":"10.1145\/2661806.2661807"},{"key":"ref_29","unstructured":"Gratch, J., Artstein, R., Lucas, G.M., Stratou, G., Scherer, S., Nazarian, A., Wood, R., Boberg, J., DeVault, D., and Marsella, S. (2014, January 26\u201331). The distress analysis interview corpus of human and computer interviews. Proceedings of the International Conference on Language Resources and Evaluation (LREC), Reykjavik, Iceland."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"DeVault, D., Artstein, R., Benn, G., Dey, T., Fast, E., Gainer, A., Georgila, K., Gratch, J., Hartholt, A., and Lhommet, M. (2014). SimSensei Kiosk: A virtual human interviewer for healthcare decision support. Proceedings of the International Conference on Autonomous Agents and Multi-Agent Systems, ACM.","DOI":"10.65109\/MXIV3169"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Baltru\u0161aitis, T., Robinson, P., and Morency, L.P. (2016). OpenFace: An open source facial behavior analysis toolkit. Proceedings of the 2016 IEEE Winter Conference on Applications of Computer Vision (WACV), IEEE.","DOI":"10.1109\/WACV.2016.7477553"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Correia, J., Teixeira, F., Botelho, C., Trancoso, I., and Raj, B. (2021). The in-the-wild speech medical corpus. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP39728.2021.9414230"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"gnad147","DOI":"10.1093\/geront\/gnad147","article-title":"Internet-based conversational engagement randomized controlled clinical trial (I-CONECT) among socially isolated adults 75+ years old with normal cognition or mild cognitive impairment: Topline results","volume":"64","author":"Dodge","year":"2024","journal-title":"Gerontol."},{"key":"ref_34","unstructured":"Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., and El-Nouby, A. (2023). Dinov2: Learning robust visual features without supervision. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"e929","DOI":"10.7717\/peerj-cs.929","article-title":"pyVHR: A Python framework for remote photoplethysmography","volume":"8","author":"Boccignone","year":"2022","journal-title":"PeerJ Comput. Sci."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Teng, S., Chai, S., Liu, J., Tateyama, T., Lin, L., and Chen, Y.W. (2024). Multi-Modal and Multi-Task Depression Detection with Sentiment Assistance. Proceedings of the 2024 IEEE International Conference on Consumer Electronics (ICCE), IEEE.","DOI":"10.1109\/ICCE59016.2024.10444213"},{"key":"ref_37","first-page":"100011","article-title":"Time Perspective-Enhanced Suicidal Ideation Detection Using Multi-Task Learning","volume":"3","author":"Yang","year":"2024","journal-title":"Int. J. Netw. Dyn. Intell."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"e66907","DOI":"10.2196\/66907","article-title":"Multimodal Multitask Learning for Predicting Depression Severity and Suicide Risk Using Pretrained Audio and Text Embeddings: Methodology Development and Application","volume":"13","author":"Hu","year":"2025","journal-title":"JMIR Med Inf."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Bia\u0142ek, K., Potulska-Chromik, A., Jakubowski, J., Nojszewska, M., and Kostera-Pruszczyk, A. (2024). Analysis of handwriting for recognition of Parkinson\u2019s disease: Current state and new study. Electronics, 13.","DOI":"10.3390\/electronics13193962"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1007\/s11063-025-11735-z","article-title":"Parkinsons detection from gait time series classification using modified metaheuristic optimized long short term memory","volume":"57","author":"Markovic","year":"2025","journal-title":"Neural Process. Lett."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1038\/s41531-022-00414-8","article-title":"An integrated biometric voice and facial features for early detection of Parkinson\u2019s disease","volume":"8","author":"Lim","year":"2022","journal-title":"NPJ Park. Dis."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"106480","DOI":"10.1016\/j.bspc.2024.106480","article-title":"Leveraging multimodal deep learning framework and a comprehensive audio-visual dataset to advance Parkinson\u2019s detection","volume":"95","author":"Lv","year":"2024","journal-title":"Biomed. Signal Process. Control"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"147818","DOI":"10.1109\/ACCESS.2025.3593254","article-title":"Multitask Deep Learning for Predicting Parkinson\u2019s Progression and Depression From Multimodal Time Series Data","volume":"13","author":"Junaid","year":"2025","journal-title":"IEEE Access"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014). Microsoft COCO: Common objects in context. Proceedings of the European Conference on Computer Vision (ECCV), Springer.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_45","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning (ICML), PmLR."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_47","first-page":"85523","article-title":"A closer look at the cls token for cross-domain few-shot learning","volume":"37","author":"Zou","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"103551","DOI":"10.1016\/j.media.2025.103551","article-title":"CLIP in medical imaging: A survey","volume":"102","author":"Zhao","year":"2025","journal-title":"Med Image Anal."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"8755","DOI":"10.1038\/s41598-024-58382-3","article-title":"Analyzing to discover origins of CNNs and ViT architectures in medical images","volume":"14","author":"Oh","year":"2024","journal-title":"Sci. Rep."},{"key":"ref_50","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA."},{"key":"ref_51","unstructured":"Gu, A., and Dao, T. (2024, January 7\u201311). Mamba: Linear-time sequence modeling with selective state spaces. Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria."},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"387","DOI":"10.1007\/s10462-025-11388-3","article-title":"CNNs, RNNs and Transformers in human action recognition: A survey and a hybrid model","volume":"58","author":"Alomar","year":"2025","journal-title":"Artif. Intell. Rev."},{"key":"ref_53","unstructured":"Snell, J., Swersky, K., and Zemel, R. (2017). Prototypical networks for few-shot learning. Proceedings of the International Conference on Neural Information Processing Systems, Curran Associates Inc."},{"key":"ref_54","first-page":"1","article-title":"Statistical comparisons of classifiers over multiple data sets","volume":"7","year":"2006","journal-title":"J. Mach. Learn. Res."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. (2017). Grad-cam: Visual explanations from deep networks via gradient-based localization. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), IEEE.","DOI":"10.1109\/ICCV.2017.74"}],"container-title":["Big Data and Cognitive Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2504-2289\/10\/3\/89\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,16]],"date-time":"2026-03-16T14:12:50Z","timestamp":1773670370000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2504-2289\/10\/3\/89"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,16]]},"references-count":55,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2026,3]]}},"alternative-id":["bdcc10030089"],"URL":"https:\/\/doi.org\/10.3390\/bdcc10030089","relation":{},"ISSN":["2504-2289"],"issn-type":[{"value":"2504-2289","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,16]]}}}