{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,20]],"date-time":"2026-08-20T04:49:25Z","timestamp":1787201365795,"version":"build-2736575974"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T00:00:00Z","timestamp":1782259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100018542","name":"Natural Science Foundation of Sichuan","doi-asserted-by":"crossref","award":["2025YFHZ0124"],"award-info":[{"award-number":["2025YFHZ0124"]}],"id":[{"id":"10.13039\/501100018542","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Sichuan Science and Technology Program","award":["2024NSFTD0036"],"award-info":[{"award-number":["2024NSFTD0036"]}]},{"name":"Frontier Cross Innovation Team Project of Southwest Jiaotong University","award":["YH1500112432297"],"award-info":[{"award-number":["YH1500112432297"]}]},{"DOI":"10.13039\/100009110","name":"Natural Science Foundation of Xinjiang Uygur Autonomous Region","doi-asserted-by":"crossref","award":["2024D01A20"],"award-info":[{"award-number":["2024D01A20"]}],"id":[{"id":"10.13039\/100009110","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>\n                    Multimodal sentiment analysis aims to comprehensively understand human sentiment by integrating diverse modalities, such as text, audio, and vision. To improve modality complementarity, the recent Multimodal Multi-task Learning (MML) framework employs joint training of unimodal and multimodal sentiment analysis tasks using sub-annotations of modality. In this work, we further draw attention to the observation that integrating unimodal tasks may introduce conflicting task information, negatively affecting the multimodal task performance. Motivated by this issue, we propose the Multimodal Task Correlation-aware Learning (MTCL) framework to leverage beneficial task correlations and suppress harmful ones. Specifically, MTCL introduces a Correlation-Adaptive Training (CAT) strategy to learn a task-relation aware unimodal encoder for each modality. First, in order to distinguish whether a sample contains conflicting information, CAT strategy incorporates a Dual-Branch Contrast (DBC) module which divides the training set into a beneficial subset and a harmful subset. Based on this division, CAT strategy proposes an adaptive training loss to guide the model in understanding nuanced multitask correlations. The adaptive training loss has two components: (1) For the beneficial subset, a contrastive loss is utilized to improve the model\u2019s ability to extract complementary representations. (2) For the harmful subset, we apply a task-correction loss to mitigate the negative interference caused by harmful task associations. With the CAT strategy, our framework can effectively distinguish beneficial and harmful task correlations to extract distinctive and robust unimodal representations. The superiority of MTCL is verified via extensive experiments on several multimodal video sentiment analysis benchmarks. Our work is publicly available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/tiggers23\/MTCL\">https:\/\/github.com\/tiggers23\/MTCL<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3796718","type":"journal-article","created":{"date-parts":[[2026,3,17]],"date-time":"2026-03-17T20:19:13Z","timestamp":1773778753000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Rethinking the Effect of Unimodal Labels in Multimodal Sentiment Analysis"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3961-4881","authenticated-orcid":false,"given":"Lingli","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Computing and Artificial Intelligence, Southwest Jiaotong University, Chengdu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7780-104X","authenticated-orcid":false,"given":"Tianrui","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computing and Artificial Intelligence, Southwest Jiaotong University, Chengdu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-1137-5709","authenticated-orcid":false,"given":"Baiyu","family":"Lu","sequence":"additional","affiliation":[{"name":"School of Computing and Artificial Intelligence, Southwest Jiaotong University, Chengdu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-8162-3201","authenticated-orcid":false,"given":"Junlin","family":"Fang","sequence":"additional","affiliation":[{"name":"School of Computing and Artificial Intelligence, Southwest Jiaotong University, Chengdu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0994-4660","authenticated-orcid":false,"given":"Desheng","family":"Zheng","sequence":"additional","affiliation":[{"name":"School of Computer and Software, Southwest Petroleum University, Chengdu, China and Intelligent Manufacturing Institute, Kash Institute of Electronics and Information Industry, Kash, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3641-1429","authenticated-orcid":false,"given":"Wei","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Informatics, Cardiff University, Cardiff, United Kingdom of Great Britain and Northern Ireland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9855-4479","authenticated-orcid":false,"given":"Weide","family":"Liu","sequence":"additional","affiliation":[{"name":"Harvard Medical School, Harvard University, Cambridge, Massachusetts, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1640-0992","authenticated-orcid":false,"given":"Fengmao","family":"Lv","sequence":"additional","affiliation":[{"name":"School of Computing and Artificial Intelligence, Southwest Jiaotong University, Chengdu, China, State Key Laboratory of Bridge Intelligent and Green Construction, Southwest Jiaotong University, Chengdu, China, and Manufacturing Industry Chain Collaboration Industrial Software Key Laboratory of Sichuan Province, Chengdu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"59","volume-title":"Proceedings of the 13th IEEE International Conference on Automatic Face and Gesture Recognition (FG \u201918)","author":"Baltrusaitis Tadas","year":"2018","unstructured":"Tadas Baltrusaitis, Amir Zadeh, Yao Chong Lim, and Louis-Philippe Morency. 2018. OpenFace 2.0: Facial behavior analysis toolkit. In Proceedings of the 13th IEEE International Conference on Automatic Face and Gesture Recognition (FG \u201918), 59\u201366."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3136755.3136801"},{"key":"e_1_3_1_4_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT \u201919)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT \u201919), 4171\u20134186."},{"key":"e_1_3_1_5_2","first-page":"1459","volume-title":"Proceedings of the 18th International Conference on Multimedia (MM \u201910)","author":"Eyben Florian","year":"2010","unstructured":"Florian Eyben, Martin W\u00f6llmer, and Bj\u00f6rn W. Schuller. 2010. Opensmile: The munich versatile and fast open-source audio feature extractor. In Proceedings of the 18th International Conference on Multimedia (MM \u201910), 1459\u20131462."},{"key":"e_1_3_1_6_2","first-page":"14755","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201924)","author":"Feng Xinyu","year":"2024","unstructured":"Xinyu Feng, Yuming Lin, Lihua He, You Li, Liang Chang, and Ya Zhou. 2024. Knowledge-Guided dynamic modality attention fusion framework for multimodal sentiment analysis. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201924), 14755\u201314766."},{"key":"e_1_3_1_7_2","first-page":"1373","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo (ICME \u201923)","author":"Geng Wenxiu","year":"2023","unstructured":"Wenxiu Geng, Yulong Bian, and Xiangxian Li. 2023. A multi-view co-learning method for multimodal sentiment analysis. In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME \u201923), 1373\u20131378."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3591106.3592260"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413678"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3276075"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_12_2","unstructured":"Wenyi Hong Wenmeng Yu Xiaotao Gu Guo Wang Guobing Gan Haomiao Tang Jiale Cheng Ji Qi Junhui Ji Lihang Pan et al. 2025. GLM-4.5V and GLM-4.1V-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning. arXiv:2507.01006. Retrieved from https:\/\/arxiv.org\/abs\/2507.01006"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612295"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2023.111346"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612053"},{"key":"e_1_3_1_16_2","first-page":"2584","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917)","author":"Li Shan","year":"2017","unstructured":"Shan Li, Weihong Deng, and Junping Du. 2017. Reliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917), 2584\u20132593."},{"key":"e_1_3_1_17_2","first-page":"6631","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201923)","author":"Li Yong","year":"2023","unstructured":"Yong Li, Yuanzhi Wang, and Zhen Cui. 2023. Decoupled multimodal distilling for emotion recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201923), 6631\u20136640."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/18.61115"},{"key":"e_1_3_1_19_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201923)","author":"Liu Peipei","year":"2023","unstructured":"Peipei Liu, Xin Zheng, Hong Li, Jie Liu, Yimo Ren, Hongsong Zhu, and Limin Sun. 2023. Improving the modality representation with multi-view contrastive learning for multimodal sentiment analysis. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201923), 1\u20135."},{"key":"e_1_3_1_20_2","first-page":"247","volume-title":"Proceedings of the International Conference on Multimodal Interaction (ICMI \u201922)","author":"Liu Yihe","year":"2022","unstructured":"Yihe Liu, Ziqi Yuan, Huisheng Mao, Zhiyun Liang, Wanqiuyue Yang, Yuanzhe Qiu, Tie Cheng, Xiaoteng Li, Hua Xu, and Kai Gao. 2022. Make acoustic and visual cues matter: CH-SIMS v2.0 dataset and AV-mixup consistent module. In Proceedings of the International Conference on Multimodal Interaction (ICMI \u201922), 247\u2013258."},{"key":"e_1_3_1_21_2","first-page":"2247","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL \u201918)","author":"Liu Zhun","year":"2018","unstructured":"Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2018. Efficient low-rank multimodal fusion with modality-specific factors. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL \u201918), 2247\u20132256."},{"key":"e_1_3_1_22_2","first-page":"2554","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201921)","author":"Fengmao Lv","year":"2021","unstructured":"Lv Fengmao, Xiang Chen, Yanyong Huang, Lixin Duan, and Guosheng Lin. 2021. Progressive modality reinforcement for human multimodal emotion recognition from unaligned multimodal sequences. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201921), 2554\u20132562."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3068598"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3171679"},{"key":"e_1_3_1_25_2","first-page":"1359","volume-title":"Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI \u201920)","author":"Mittal Trisha","year":"2020","unstructured":"Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera, and Dinesh Manocha. 2020. M3ER: Multiplicative multimodal emotion recognition using facial, textual, and speech cues. In Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI \u201920), 1359\u20131367."},{"key":"e_1_3_1_26_2","first-page":"169","volume-title":"Proceedings of the 13th International Conference on Multimodal Interfaces (ICMI \u201911)","author":"Morency Louis-Philippe","year":"2011","unstructured":"Louis-Philippe Morency, Rada Mihalcea, and Payal Doshi. 2011. Towards multimodal sentiment analysis: Harvesting opinions from the web. In Proceedings of the 13th International Conference on Multimodal Interfaces (ICMI \u201911), 169\u2013176."},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","unstructured":"Hai Pham Thomas Manzini Paul Pu Liang and Barnab\u00e1s P\u00f3czos. 2018. Seq2Seq2Sentiment: Multimodal sequence to sequence models for sentiment analysis. arXiv:1807.03915. Retrieved from https:\/\/arxiv.org\/abs\/1807.03915","DOI":"10.18653\/v1\/W18-3308"},{"key":"e_1_3_1_28_2","first-page":"7925","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201924)","author":"Ruan Yu-Ping","year":"2024","unstructured":"Yu-Ping Ruan, Shoukang Han, Taihao Li, and Yanfeng Wu. 2024. Fusing modality-specific representations and decisions for multimodal emotion recognition. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201924), 7925\u20137929."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548025"},{"key":"e_1_3_1_30_2","first-page":"658","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL \u201923)","author":"Sun Jun","year":"2023","unstructured":"Jun Sun, Shoukang Han, Yu-Ping Ruan, Xiaoning Zhang, Shu-Kai Zheng, Yulong Liu, Yuxin Huang, and Taihao Li. 2023. Layer-wise fusion with modality independence modeling for multi-modal emotion recognition. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL \u201923), 658\u2013670."},{"key":"e_1_3_1_31_2","first-page":"8992","volume-title":"Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI \u201920)","author":"Sun Zhongkai","year":"2020","unstructured":"Zhongkai Sun, Prathusha Kameswara Sarma, William A. Sethares, and Yingyu Liang. 2020. Learning relationships between text, audio, and video via deep canonical correlation for multimodal language analysis. In Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI \u201920), 8992\u20138999."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3218018"},{"key":"e_1_3_1_33_2","first-page":"6558","volume-title":"Proceedings of the 57th Conference of the Association for Computational Linguistics (ACL \u201919)","author":"Tsai Yao-Hung Hubert","year":"2019","unstructured":"Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the 57th Conference of the Association for Computational Linguistics (ACL \u201919), 6558\u20136569."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.109259"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3183830"},{"key":"e_1_3_1_36_2","first-page":"21180","volume-title":"Proceedings of the 39th AAAI Conference on Artificial Intelligence (AAAI \u201925)","author":"Wang Pan","year":"2025","unstructured":"Pan Wang, Qiang Zhou, Yawen Wu, Tianlong Chen, and Jingtong Hu. 2025. DLF: Disentangled-language-focused multimodal sentiment analysis. In Proceedings of the 39th AAAI Conference on Artificial Intelligence (AAAI \u201925), 21180\u201321188."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-3302"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3517139"},{"key":"e_1_3_1_39_2","unstructured":"An Yang Anfeng Li Baosong Yang Beichen Zhang Binyuan Hui Bo Zheng Bowen Yu Chang Gao Chengen Huang Chenxu Lv et al. 2025. Qwen3 technical report. arXiv:2505.09388. Retrieved from https:\/\/arxiv.org\/abs\/2505.09388"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547754"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.421"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413690"},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","first-page":"2099","DOI":"10.18653\/v1\/2024.findings-naacl.135","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (NAACL \u201924)","author":"Yang Yang","year":"2024","unstructured":"Yang Yang, Xunde Dong, and Yupeng Qiang. 2024. CLGSI: A multimodal sentiment analysis framework based on contrastive learning guided by sentiment intensity. In Proceedings of the Findings of the Association for Computational Linguistics (NAACL \u201924), 2099\u20132110."},{"key":"e_1_3_1_44_2","first-page":"2822","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201919)","author":"Yoon Seunghyun","year":"2019","unstructured":"Seunghyun Yoon, Seokhyun Byun, Subhadeep Dey, and Kyomin Jung. 2019. Speech emotion recognition using multi-hop attention mechanism. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201919), 2822\u20132826."},{"key":"e_1_3_1_45_2","first-page":"3718","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL \u201920)","author":"Yu Wenmeng","year":"2020","unstructured":"Wenmeng Yu, Hua Xu, Fanyang Meng, Yilin Zhu, Yixiao Ma, Jiele Wu, Jiyun Zou, and Kaicheng Yang. 2020. CH-SIMS: A Chinese multimodal sentiment analysis dataset with fine-grained annotation of modality. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL \u201920), 3718\u20133727."},{"key":"e_1_3_1_46_2","first-page":"10790","volume-title":"Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI \u201921)","author":"Yu Wenmeng","year":"2021","unstructured":"Wenmeng Yu, Hua Xu, Ziqi Yuan, and Jiele Wu. 2021. Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis. In Proceedings of the 35th AAAI Conference on Artificial Intelligence (AAAI \u201921), 10790\u201310797."},{"key":"e_1_3_1_47_2","first-page":"1103","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201917)","author":"Zadeh Amir","year":"2017","unstructured":"Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2017. Tensor fusion network for multimodal sentiment analysis. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201917), 1103\u20131114."},{"key":"e_1_3_1_48_2","first-page":"5634","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence (AAAI \u201918)","author":"Zadeh Amir","year":"2018","unstructured":"Amir Zadeh, Paul Pu Liang, Navonil Mazumder, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018. Memory fusion network for multi-view sequential learning. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence (AAAI \u201918), 5634\u20135641."},{"key":"e_1_3_1_49_2","first-page":"2236","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL \u201918)","author":"Zadeh Amir","year":"2018","unstructured":"Amir Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL \u201918), 2236\u20132246."},{"key":"e_1_3_1_50_2","first-page":"6182","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201922)","author":"Zhang Binbin","year":"2022","unstructured":"Binbin Zhang, Lv Hang, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, et al. 2022. WENETSPEECH: A 10000+ hours multi-domain mandarin corpus for speech recognition. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201922), 6182\u20136186."},{"key":"e_1_3_1_51_2","doi-asserted-by":"crossref","first-page":"756","DOI":"10.18653\/v1\/2023.emnlp-main.49","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201923)","author":"Zhang Haoyu","year":"2023","unstructured":"Haoyu Zhang, Yu Wang, Guanghao Yin, Kejun Liu, Yuanyuan Liu, and Tianshu Yu. 2023. Learning language-guided adaptive hyper-modality representation for multimodal sentiment analysis. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP \u201923), 756\u2013767."},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2016.2603342"},{"key":"e_1_3_1_53_2","first-page":"9100","volume-title":"Proceedings of the 36th AAAI Conference on Artificial Intelligence (AAAI \u201922)","author":"Zhang Yi","year":"2022","unstructured":"Yi Zhang, Mingyuan Chen, Jundong Shen, and Chongjun Wang. 2022. Tailor versatile multi-modal learning for multi-label emotion recognition. In Proceedings of the 36th AAAI Conference on Artificial Intelligence (AAAI \u201922), 9100\u20139108."},{"key":"e_1_3_1_54_2","first-page":"1","volume-title":"Proceedings of the International Joint Conference on Neural Networks (IJCNN \u201924)","author":"Zhu Qingmeng","year":"2024","unstructured":"Qingmeng Zhu, Tianxing Lan, Jian Liang, Deliang Xiang, and Hao He. 2024. Feature alignment and reconstruction constraints for multimodal sentiment analysis. In Proceedings of the International Joint Conference on Neural Networks (IJCNN \u201924), 1\u20136."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3796718","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:48:55Z","timestamp":1782312535000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3796718"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":53,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3796718"],"URL":"https:\/\/doi.org\/10.1145\/3796718","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-06-20","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-03","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}