{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,28]],"date-time":"2026-06-28T01:10:53Z","timestamp":1782609053415,"version":"3.54.5"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>Clinical depression diagnosis relies heavily on both verbal and non-verbal cues in patient interviews, yet existing automated methods often operate as black-box models and fail to provide trustworthy explanations, limiting their clinical applicability. Moreover, depression datasets commonly impose strict privacy constraints that prohibit access to raw audio\u2013visual data, and the deployment of large LLMs in real-world medical environments is often constrained by computational cost limitations. To address these challenges, we propose MLlm-DR, a multimodal large language model for explainable depression recognition. We first employ knowledge distillation to transfer high-quality diagnostic rationales from a powerful LLMs to a smaller, deployable LLMs, thereby equipping it with clinically aligned reasoning abilities. We further introduce a lightweight query module (LQ-former) that extracts salient depression-related cues from pre-extracted audio and visual features and maps them into an LLM-compatible representation. MLlm-DR is trained in two stages: multimodal alignment via LQ-former pretraining, followed by multi-task optimization that jointly learns rationale generation and score regression, encouraging consistency and improving interpretability. Experimental results show that MLlm-DR achieves state-of-the-art performance on two interview-based benchmark datasets, CMDC and E-DAIC-WOZ, while simultaneously generating clinically meaningful and readable explanations.<\/jats:p>","DOI":"10.1145\/3796722","type":"journal-article","created":{"date-parts":[[2026,2,13]],"date-time":"2026-02-13T16:07:58Z","timestamp":1770998878000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["MLlm-DR: Towards Explainable Depression Recognition with MultiModal Large Language Models"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-8079-3349","authenticated-orcid":false,"given":"Wei","family":"Zhang","sequence":"first","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-8703-7276","authenticated-orcid":false,"given":"Juan","family":"Chen","sequence":"additional","affiliation":[{"name":"University of the Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2305-7555","authenticated-orcid":false,"given":"En","family":"Zhu","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4469-9442","authenticated-orcid":false,"given":"Wenhong","family":"Cheng","sequence":"additional","affiliation":[{"name":"Shanghai Mental Health Center, Shanghai Jiao Tong University School of Medicine, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-4338-4960","authenticated-orcid":false,"given":"Yunpeng","family":"Li","sequence":"additional","affiliation":[{"name":"Nanjing Industria Tenebris Information Technology Co., Ltd., Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-1301-3416","authenticated-orcid":false,"given":"Yuhan","family":"Li","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-0658-930X","authenticated-orcid":false,"given":"Yanbo J","family":"Wang","sequence":"additional","affiliation":[{"name":"National University of Uzbekistan named after Mirzo Ulugbek, Tashkent, Uzbekistan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,3,23]]},"reference":[{"issue":"1","key":"e_1_3_1_2_2","doi-asserted-by":"crossref","first-page":"163","DOI":"10.1016\/j.psychres.2011.09.008","article-title":"The Beck depression inventory and general health questionnaire as measures of depression in the general population: A validation study using the composite international diagnostic interview as the gold standard","volume":"197","author":"Aalto Anna-Mari","year":"2012","unstructured":"Anna-Mari Aalto, Marko Elovainio, Mika Kivim\u00e4ki, Antti Uutela, and Sami Pirkola. 2012. The Beck depression inventory and general health questionnaire as measures of depression in the general population: A validation study using the composite international diagnostic interview as the gold standard. Psychiatry Research 197, 1\u20132 (2012), 163\u2013171.","journal-title":"Psychiatry Research"},{"key":"e_1_3_1_3_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et al. 2023. GPT-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","first-page":"23716","DOI":"10.52202\/068431-1723","article-title":"Flamingo: A visual language model for few-shot learning","volume":"35","author":"Alayrac Jean-Baptiste","year":"2022","unstructured":"Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: A visual language model for few-shot learning. In Advances in Neural Information Processing Systems, Vol. 35, 23716\u201323736.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","DOI":"10.1176\/appi.books.9780890425596","volume-title":"Diagnostic and Statistical Manual of Mental Disorders: DSM-5","author":"DSMTF American Psychiatric Association","year":"2013","unstructured":"DSMTF American Psychiatric Association and DS American Psychiatric Association. 2013. Diagnostic and Statistical Manual of Mental Disorders: DSM-5, Vol. 5. American Psychiatric Association, Washington, DC."},{"key":"e_1_3_1_6_2","article-title":"The Claude 3 model family: Opus, sonnet, haiku","author":"AI Anthropic","year":"2024","unstructured":"AI Anthropic. 2024. The Claude 3 model family: Opus, sonnet, haiku. Claude-3 Model Card. Retrieved from https:\/\/www-cdn.anthropic.com\/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627\/Model_Card_Claude_3.pdf","journal-title":"Claude-3 Model Card"},{"key":"e_1_3_1_7_2","unstructured":"Shuai Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge Sibo Song Kai Dang Peng Wang Shijie Wang Jun Tang et al. 2025. Qwen2.5-VL technical report. arXiv:2502.13923. Retrieved from https:\/\/arxiv.org\/abs\/2502.13923"},{"key":"e_1_3_1_8_2","first-page":"59","article-title":"OpenFace 2.0: Facial behavior analysis toolkit","author":"Baltrusaitis Tadas","year":"2018","unstructured":"Tadas Baltrusaitis, Amir Zadeh, Yao Chong Lim, and Louis Philippe Morency. 2018. OpenFace 2.0: Facial behavior analysis toolkit. In Proceedings of the IEEE International Conference on Automatic Face & Gesture Recognition, 59\u201366.","journal-title":"Proceedings of the IEEE International Conference on Automatic Face & Gesture Recognition"},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","first-page":"102017","DOI":"10.1016\/j.inffus.2023.102017","article-title":"IIFDD: Intra and inter-modal fusion for depression detection with multi-modal information from internet of medical things","volume":"102","author":"Chen Jian","year":"2024","unstructured":"Jian Chen, Yuzhu Hu, Qifeng Lai, Wei Wang, Junxin Chen, Han Liu, Gautam Srivastava, Ali Kashif Bashir, and Xiping Hu. 2024. IIFDD: Intra and inter-modal fusion for depression detection with multi-modal information from internet of medical things. Information Fusion 102 (2024), 102017.","journal-title":"Information Fusion"},{"issue":"6","key":"e_1_3_1_10_2","doi-asserted-by":"crossref","first-page":"1505","DOI":"10.1109\/JSTSP.2022.3188113","article-title":"WavLM: Large-scale self-supervised pre-training for full stack speech processing","volume":"16","author":"Chen Sanyuan","year":"2022","unstructured":"Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al. 2022. WavLM: Large-scale self-supervised pre-training for full stack speech processing. IEEE Journal of Selected Topics in Signal Processing 16, 6 (2022), 1505\u20131518.","journal-title":"IEEE Journal of Selected Topics in Signal Processing"},{"key":"e_1_3_1_11_2","unstructured":"Jiaming Cheng Ruiyu Liang Chao Xu Ye Ni Wei Zhou Bj\u00f6rn W. Schuller and Xiaoshuai Hao. 2025. I2S-TFCKD: Intra-inter set knowledge distillation with time-frequency calibration for speech enhancement. arXiv:2506.13127. Retrieved from https:\/\/arxiv.org\/abs\/2506.13127"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","first-page":"110805","DOI":"10.52202\/079017-3518","article-title":"Emotion-LLaMA: Multimodal emotion recognition and reasoning with instruction tuning","volume":"37","author":"Cheng Zebang","year":"2024","unstructured":"Zebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang, Yuxiang Lin, Zheng Lian, Xiaojiang Peng, and Alexander Hauptmann. 2024. Emotion-LLaMA: Multimodal emotion recognition and reasoning with instruction tuning. In Advances in Neural Information Processing Systems, Vol. 37, 110805\u2013110853.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"3","key":"e_1_3_1_13_2","doi-asserted-by":"crossref","first-page":"1581","DOI":"10.1109\/TAFFC.2020.3021755","article-title":"A deep multiscale spatiotemporal network for assessing depression from facial dynamics","volume":"13","author":"De Melo Wheidima Carneiro","year":"2020","unstructured":"Wheidima Carneiro De Melo, Eric Granger, and Abdenour Hadid. 2020. A deep multiscale spatiotemporal network for assessing depression from facial dynamics. IEEE Transactions on Affective Computing 13, 3 (2020), 1581\u20131592.","journal-title":"IEEE Transactions on Affective Computing"},{"key":"e_1_3_1_14_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","volume":"1","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol. 1, 4171\u20134186."},{"key":"e_1_3_1_15_2","first-page":"3123","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation","author":"Gratch Jonathan","year":"2014","unstructured":"Jonathan Gratch, Ron Artstein, Gale M. Lucas, Giota Stratou, Stefan Scherer, Angela Nazarian, Rachel Wood, Jill Boberg, David DeVault, Stacy Marsella, et al. 2014. The distress analysis interview corpus of human and computer interviews. In Proceedings of the International Conference on Language Resources and Evaluation, 3123\u20133128."},{"key":"e_1_3_1_16_2","first-page":"770","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770\u2013778."},{"key":"e_1_3_1_17_2","doi-asserted-by":"crossref","unstructured":"Cheng-Yu Hsieh Chun-Liang Li Chih-Kuan Yeh Hootan Nakhost Yasuhisa Fujii Alexander Ratner Ranjay Krishna Chen-Yu Lee and Tomas Pfister. 2023. Distilling step-by-step! Outperforming larger language models with less training data and smaller model sizes. arXiv:2305.02301. Retrieved from https:\/\/arxiv.org\/abs\/2305.02301","DOI":"10.18653\/v1\/2023.findings-acl.507"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","first-page":"3451","DOI":"10.1109\/TASLP.2021.3122291","article-title":"Hubert: Self-supervised speech representation learning by masked prediction of hidden units","volume":"29","author":"Hsu Wei-Ning","year":"2021","unstructured":"Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE\/ACM Transactions on Audio, Speech, and Language Processing 29 (2021), 3451\u20133460.","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"e_1_3_1_19_2","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations Vol. 1 3."},{"issue":"3","key":"e_1_3_1_20_2","first-page":"1","article-title":"A multi-hop graph reasoning network for knowledge-based VQA","volume":"16","author":"Hu Zihan","year":"2025","unstructured":"Zihan Hu, Jiuxiang You, Zhenguo Yang, Xiaoping Li, Haoran Xie, Qing Li, and Wenyin Liu. 2025. A multi-hop graph reasoning network for knowledge-based VQA. ACM Transactions on Intelligent Systems and Technology 16, 3 (2025), 1\u201323.","journal-title":"ACM Transactions on Intelligent Systems and Technology"},{"key":"e_1_3_1_21_2","first-page":"6288","volume-title":"Proceedings of the 33rd International Joint Conference on Artificial Intelligence","author":"Huang Zhaopei","year":"2024","unstructured":"Zhaopei Huang, Jinming Zhao, and Qin Jin. 2024. ECR-Chain: Advancing generative language models to better emotion-cause reasoners through reasoning chains. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence, 6288\u20136296."},{"key":"e_1_3_1_22_2","first-page":"57","article-title":"Bootstrapping vision-language learning with decoupled language pre-training","volume":"36","author":"Jian Yiren","year":"2024","unstructured":"Yiren Jian, Chongyang Gao, and Soroush Vosoughi. 2024. Bootstrapping vision-language learning with decoupled language pre-training. In Advances in Neural Information Processing Systems, Vol. 36, 57\u201372.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","first-page":"1049","DOI":"10.1145\/3627673.3679797","volume-title":"Proceedings of the 33rd ACM International Conference on Information and Knowledge Management","author":"Jung Juho","year":"2024","unstructured":"Juho Jung, Chaewon Kang, Jeewoo Yoon, Seungbae Kim, and Jinyoung Han. 2024. HiQuE: Hierarchical question embedding network for multimodal depression detection. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, 1049\u20131059."},{"key":"e_1_3_1_24_2","first-page":"22199","article-title":"Large language models are zero-shot reasoners","volume":"35","author":"Kojima Takeshi","year":"2022","unstructured":"Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems, Vol. 35, 22199\u201322213.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_25_2","unstructured":"Shanglin Lei Guanting Dong Xiaoping Wang Keheng Wang and Sirui Wang. 2023. InstructERC: Reforming emotion recognition in conversation with a retrieval multi-task LLMs framework. arXiv:2309.11911. Retrieved from https:\/\/arxiv.org\/abs\/2309.11911"},{"key":"e_1_3_1_26_2","first-page":"6359","volume-title":"Proceedings of the 33rd International Joint Conference on Artificial Intelligence","author":"Li Guozheng","year":"2024","unstructured":"Guozheng Li, Zijie Xu, Ziyu Shang, Jiajun Liu, Ji Ke, and Yikai Guo. 2024. Empirical analysis of dialogue relation extraction with large language models. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence, 6359\u20136367."},{"key":"e_1_3_1_27_2","first-page":"19730","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the International Conference on Machine Learning, 19730\u201319742. PMLR."},{"key":"e_1_3_1_28_2","first-page":"3365","volume-title":"Proceedings of the 31st ACM International Conference on Multimedia","author":"Meng Liu Dayu","year":"2023","unstructured":"Dayu Meng Liu, Hao Yu Ke Liang, Yue Hu, Lingyuan Meng Liu, Wenxuan Tu, Sihang Zhou, and Xinwang Liu. 2023. TMac: Temporal multi-modal graph learning for acoustic event classification. In Proceedings of the 31st ACM International Conference on Multimedia, 3365\u20133374."},{"key":"e_1_3_1_29_2","first-page":"316","volume-title":"Proceedings of the International Conference on Intelligent Computing","author":"Liu Qiang","year":"2025","unstructured":"Qiang Liu, Mengxi Ying, Peng Xiao, Gan Li, and Xinpan Yuan. 2025. Prompting large models for knowledge and reasoning augmentation in KB-VQA. In Proceedings of the International Conference on Intelligent Computing, 316\u2013327. Springer."},{"key":"e_1_3_1_30_2","unstructured":"Junyu Lu Ruyi Gan Dixiang Zhang Xiaojun Wu Ziwei Wu Renliang Sun Jiaxing Zhang Pingjian Zhang and Yan Song. 2023. Lyrics: Boosting fine-grained language-vision alignment and comprehension via semantic-aware visual objects. arXiv:2312.05278. Retrieved from https:\/\/arxiv.org\/abs\/2312.05278"},{"issue":"18","key":"e_1_3_1_31_2","first-page":"7","article-title":"Librosa: Audio and music signal analysis in python","volume":"2015","author":"McFee Brian","year":"2015","unstructured":"Brian McFee, Colin Raffel, Dawen Liang, Daniel P. W. Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto. Librosa: Audio and music signal analysis in python. SciPy 2015, 18\u201324 (2015), 7.","journal-title":"SciPy"},{"key":"e_1_3_1_32_2","doi-asserted-by":"crossref","first-page":"106433","DOI":"10.1016\/j.cmpb.2021.106433","article-title":"End-to-end multimodal clinical depression recognition using deep neural networks: A comparative analysis","volume":"211","author":"Muzammel Muhammad","year":"2021","unstructured":"Muhammad Muzammel, Hanan Salam, and Alice Othmani. 2021. End-to-end multimodal clinical depression recognition using deep neural networks: A comparative analysis. Computer Methods and Programs in Biomedicine 211 (2021), 106433.","journal-title":"Computer Methods and Programs in Biomedicine"},{"issue":"1","key":"e_1_3_1_33_2","first-page":"294","article-title":"Multimodal spatiotemporal representation for automatic depression level detection","volume":"14","author":"Niu Mingyue","year":"2020","unstructured":"Mingyue Niu, Jianhua Tao, Bin Liu, Jian Huang, and Zheng Lian. 2020. Multimodal spatiotemporal representation for automatic depression level detection. IEEE Transactions on Affective Computing 14, 1 (2020), 294\u2013307.","journal-title":"IEEE Transactions on Affective Computing"},{"key":"e_1_3_1_34_2","volume-title":"Chinese Classification of Mental Disorders, 3rd Edition (CCMD-3)","author":"Chinese Society of Psychiatry","year":"2001","unstructured":"Chinese Society of Psychiatry. 2001. Chinese Classification of Mental Disorders, 3rd Edition (CCMD-3). Shandong Science and Technology Press, Jinan, China."},{"key":"e_1_3_1_35_2","volume-title":"International Classification of Diseases for Mortality and Morbidity Statistics (11th revision)","author":"World Health Organization.","year":"2018","unstructured":"World Health Organization. 2018. International Classification of Diseases for Mortality and Morbidity Statistics (11th revision). World Health Organization."},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","first-page":"110340","DOI":"10.1016\/j.knosys.2023.110340","article-title":"FADO: Feedback-aware double controlling network for emotional support conversation","volume":"264","author":"Peng Wei","year":"2023","unstructured":"Wei Peng, Ziyuan Qin, Yue Hu, Yuqiang Xie, and Yunpeng Li. 2023. FADO: Feedback-aware double controlling network for emotional support conversation. Knowledge-Based Systems 264 (2023), 110340.","journal-title":"Knowledge-Based Systems"},{"issue":"140","key":"e_1_3_1_37_2","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21, 140 (2020), 1\u201367.","journal-title":"Journal of Machine Learning Research"},{"issue":"1","key":"e_1_3_1_38_2","doi-asserted-by":"crossref","first-page":"66","DOI":"10.1038\/s44184-024-00112-8","article-title":"Harnessing multimodal approaches for depression detection using large language models and facial expressions","volume":"3","author":"Sadeghi Misha","year":"2024","unstructured":"Misha Sadeghi, Robert Richer, Bernhard Egger, Lena Schindler-Gmelch, Lydia Helene Rupp, Farnaz Rahimi, Matthias Berking, and Bjoern M. Eskofier. 2024. Harnessing multimodal approaches for depression detection using large language models and facial expressions. NPJ Mental Health Research 3, 1 (2024), 66.","journal-title":"NPJ Mental Health Research"},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","first-page":"3722","DOI":"10.1145\/3503161.3548025","volume-title":"Proceedings of the 30th ACM International Conference on Multimedia","author":"Sun Hao","year":"2022","unstructured":"Hao Sun, Hongyi Wang, Jiaqing Liu, Yen-Wei Chen, and Lanfen Lin. 2022. CubeMLP: An MLP-based model for multimodal sentiment analysis and depression estimation. In Proceedings of the 30th ACM International Conference on Multimedia, 3722\u20133729."},{"key":"e_1_3_1_40_2","first-page":"2912","article-title":"GraphIQA: Learning distortion graph representations for blind image quality assessment","volume":"25","author":"Sun Simeng","year":"2022","unstructured":"Simeng Sun, Tao Yu, Jiahua Xu, Wei Zhou, and Zhibo Chen. 2022. GraphIQA: Learning distortion graph representations for blind image quality assessment. IEEE Transactions on Multimedia 25 (2022), 2912\u20132925.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_41_2","unstructured":"Yu Sun Shuohuan Wang Shikun Feng Siyu Ding Chao Pang Junyuan Shang Jiaxiang Liu Xuyi Chen Yanbin Zhao Yuxiang Lu et al. 2021. Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation. arXiv:2107.02137. Retrieved from https:\/\/arxiv.org\/abs\/2107.02137"},{"key":"e_1_3_1_42_2","unstructured":"Gemini Team Rohan Anil Sebastian Borgeaud Jean-Baptiste Alayrac Jiahui Yu Radu Soricut Johan Schalkwyk Andrew M. Dai Anja Hauth Katie Millican et al. 2023. Gemini: A family of highly capable multimodal models. arXiv:2312.11805. Retrieved from https:\/\/arxiv.org\/abs\/2312.11805"},{"key":"e_1_3_1_43_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et al. 2023. LLaMA: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_1_44_2","first-page":"5998","article-title":"Attention is all you need","volume":"30","author":"Vaswani A.","year":"2017","unstructured":"A. Vaswani. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, Vol. 30, 5998\u20136008.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_45_2","first-page":"274","volume-title":"Proceedings of theEuropean Conference on Computer Vision","author":"Wang Haibo","year":"2024","unstructured":"Haibo Wang and Weifeng Ge. 2024. Q&A prompts: Discovering rich visual clues through mining question-answer prompts for VQA requiring diverse world knowledge. In Proceedings of the European Conference on Computer Vision, 274\u2013292. Springer."},{"key":"e_1_3_1_46_2","unstructured":"Shanmin Wang Chengguang Liu and Qingshan Liu. 2025. Multi-modality collaborative learning for sentiment analysis. arXiv:2501.12424. Retrieved from https:\/\/arxiv.org\/abs\/2501.12424"},{"key":"e_1_3_1_47_2","first-page":"121475","article-title":"CogVLM: Visual expert for pretrained language models","volume":"37","author":"Wang Weihan","year":"2024","unstructured":"Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Song XiXuan, et al. 2024. CogVLM: Visual expert for pretrained language models. In Advances in Neural Information Processing Systems, Vol. 37, 121475\u2013121499.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_48_2","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, Vol. 35, 24824\u201324837.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_49_2","first-page":"623","volume-title":"Proceedings of theEuropean Conference on Computer Vision","author":"Wei Ping-Cheng","year":"2022","unstructured":"Ping-Cheng Wei, Kunyu Peng, Alina Roitberg, Kailun Yang, Jiaming Zhang, and Rainer Stiefelhagen. 2022. Multi-modal depression estimation based on sub-attentional fusion. In Proceedings of the European Conference on Computer Vision, 623\u2013639. Springer."},{"key":"e_1_3_1_50_2","doi-asserted-by":"crossref","first-page":"128888","DOI":"10.1016\/j.neucom.2024.128888","article-title":"MAGO: Multi-knowledge aware and global strategy sequence optimizing network for emotional support conversation","volume":"618","author":"Xie Qijun","year":"2025","unstructured":"Qijun Xie and Wei Peng. 2025. MAGO: Multi-knowledge aware and global strategy sequence optimizing network for emotional support conversation. Neurocomputing 618 (2025), 128888.","journal-title":"Neurocomputing"},{"key":"e_1_3_1_51_2","unstructured":"An Yang Baosong Yang Beichen Zhang Binyuan Hui Bo Zheng Bowen Yu Chengyuan Li Dayiheng Liu Fei Huang Haoran Wei et al. 2024. Qwen2.5 technical report. arXiv:2412.15115. Retrieved from https:\/\/arxiv.org\/abs\/2412.15115"},{"key":"e_1_3_1_52_2","unstructured":"Li Yu Xuanzhe Sun Wei Zhou and Moncef Gabbouj. 2025. Text-audio-visual-conditioned diffusion model for video saliency prediction. arXiv:2504.14267. Retrieved from https:\/\/arxiv.org\/abs\/2504.14267"},{"key":"e_1_3_1_53_2","doi-asserted-by":"crossref","unstructured":"Li Yu Situo Wang Wei Zhou and Moncef Gabbouj. 2025. DVLTA-VQA: Decoupled vision-language modeling with text-guided adaptation for blind video quality assessment. arXiv:2504.11733. Retrieved from https:\/\/arxiv.org\/abs\/2504.11733","DOI":"10.1109\/TCSVT.2026.3657415"},{"key":"e_1_3_1_54_2","first-page":"56","volume-title":"Proceedings of theInternational Conference on Artificial Neural Networks","author":"Yuan Chengbo","year":"2024","unstructured":"Chengbo Yuan, Xuxu Liu, Qianhui Xu, Yongqian Li, Yong Luo, and Xin Zhou. 2024. Depression diagnosis and analysis via multimodal multi-order factor fusion. In Proceedings of the International Conference on Artificial Neural Networks, 56\u201370. Springer."},{"key":"e_1_3_1_55_2","first-page":"6558","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Zadeh Amir","year":"2019","unstructured":"Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2019. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 6558\u20136569."},{"key":"e_1_3_1_56_2","unstructured":"Aohan Zeng Xiao Liu Zhengxiao Du Zihan Wang Hanyu Lai Ming Ding Zhuoyi Yang Yifan Xu Wendi Zheng Xiao Xia et al. 2022. GLM-130B: An open bilingual pre-trained model. arXiv:2210.02414. Retrieved from https:\/\/arxiv.org\/abs\/2210.02414"},{"key":"e_1_3_1_57_2","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1145\/3664647.3681491","volume-title":"Proceedings of the 32nd ACM International Conference on Multimedia","author":"Zhang Wei","year":"2024","unstructured":"Wei Zhang, En Zhu, Juan Chen, and YunPeng Li. 2024. MDDR: Multi-modal dual-attention aggregation for depression recognition. In Proceedings of the 32nd ACM International Conference on Multimedia, 321\u2013329."},{"key":"e_1_3_1_58_2","unstructured":"Weixiang Zhao Yanyan Zhao Xin Lu Shilong Wang Yanpeng Tong and Bing Qin. 2023. Is ChatGPT equipped with emotional dialogue capabilities? arXiv:2304.09582. Retrieved from https:\/\/arxiv.org\/abs\/2304.09582"},{"issue":"4","key":"e_1_3_1_59_2","first-page":"2823","article-title":"Semi-structural interview-based Chinese multimodal depression corpus towards automatic preliminary screening of depressive disorders","volume":"14","author":"Zou Bochao","year":"2022","unstructured":"Bochao Zou, Jiali Han, Yingxue Wang, Rui Liu, Shenghui Zhao, Lei Feng, Xiangwen Lyu, and Huimin Ma. 2022. Semi-structural interview-based Chinese multimodal depression corpus towards automatic preliminary screening of depressive disorders. IEEE Transactions on Affective Computing 14, 4 (2022), 2823\u20132838.","journal-title":"IEEE Transactions on Affective Computing"},{"issue":"7","key":"e_1_3_1_60_2","doi-asserted-by":"crossref","first-page":"5575","DOI":"10.1109\/TCSVT.2024.3358547","article-title":"Weakly-supervised action learning in procedural task videos via process knowledge decomposition","volume":"34","author":"Zou Minghao","year":"2024","unstructured":"Minghao Zou, Qingtian Zeng, and Xue Zhang. 2024. Weakly-supervised action learning in procedural task videos via process knowledge decomposition. IEEE Transactions on Circuits and Systems for Video Technology 34, 7 (2024), 5575\u20135588.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3796722","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,24]],"date-time":"2026-03-24T05:16:28Z","timestamp":1774329388000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3796722"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,23]]},"references-count":59,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3796722"],"URL":"https:\/\/doi.org\/10.1145\/3796722","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,23]]},"assertion":[{"value":"2025-07-07","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-03","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-23","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}