{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T09:56:03Z","timestamp":1780394163156,"version":"3.54.1"},"reference-count":71,"publisher":"Association for Computing Machinery (ACM)","issue":"3s","license":[{"start":{"date-parts":[[2023,2,24]],"date-time":"2023-02-24T00:00:00Z","timestamp":1677196800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61672068"],"award-info":[{"award-number":["61672068"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,10,31]]},"abstract":"<jats:p>Efficient recognition of emotions has attracted extensive research interest, which makes new applications in many fields possible, such as human-computer interaction, disease diagnosis, service robots, and so forth. Although existing work on sentiment analysis relying on sensors or unimodal methods performs well for simple contexts like business recommendation and facial expression recognition, it does far below expectations for complex scenes, such as sarcasm, disdain, and metaphors. In this article, we propose a novel two-stage multimodal learning framework, called AMSA, to adaptively learn correlation and complementarity between modalities for dynamic fusion, achieving more stable and precise sentiment analysis results. Specifically, a multiscale attention model with a slice positioning scheme is proposed to get stable quintuplets of sentiment in images, texts, and speeches in the first stage. Then a Transformer-based self-adaptive network is proposed to assign weights flexibly for multimodal fusion in the second stage and update the parameters of the loss function through compensation iteration. To quickly locate key areas for efficient affective computing, a patch-based selection scheme is proposed to iteratively remove redundant information through a novel loss function before fusion. Extensive experiments have been conducted on both machine weakly labeled and manually annotated datasets of self-made Video-SA, CMU-MOSEI, and CMU-MOSI. The results demonstrate the superiority of our approach through comparison with baselines.<\/jats:p>","DOI":"10.1145\/3572915","type":"journal-article","created":{"date-parts":[[2022,12,1]],"date-time":"2022-12-01T12:41:10Z","timestamp":1669898470000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":17,"title":["AMSA: Adaptive Multimodal Learning for Sentiment Analysis"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1782-8704","authenticated-orcid":false,"given":"Jingyao","family":"Wang","sequence":"first","affiliation":[{"name":"Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1551-4448","authenticated-orcid":false,"given":"Luntian","family":"Mou","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Multimedia and Intelligent Software Technology, Beijing Institute of Artificial Intelligence, Faculty of Information, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6024-3854","authenticated-orcid":false,"given":"Lei","family":"Ma","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4234-6099","authenticated-orcid":false,"given":"Tiejun","family":"Huang","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8894-1806","authenticated-orcid":false,"given":"Wen","family":"Gao","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,2,24]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3199668"},{"key":"e_1_3_1_3_2","volume-title":"25th AAAI Conference on Artificial Intelligence","author":"Chung Jessica Elan","year":"2011","unstructured":"Jessica Elan Chung and Eni Mustafaraj. 2011. Can collective sentiment expressed on twitter predict political elections? In 25th AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_1_4_2","first-page":"30","volume-title":"AAAI","volume":"6","author":"Cui Hang","year":"2006","unstructured":"Hang Cui, Vibhu Mittal, and Mayur Datar. 2006. Comparative experiments on sentiment classification for online product reviews. In AAAI, Vol. 6. 30."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2967210"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3363560"},{"key":"e_1_3_1_7_2","first-page":"5194","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Jiang Wentao","year":"2020","unstructured":"Wentao Jiang, Si Liu, Chen Gao, Jie Cao, Ran He, Jiashi Feng, and Shuicheng Yan. 2020. PSGAN: Pose and expression robust spatial-aware GAN for customizable makeup transfer. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 5194\u20135202."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2021.114693"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2017.08.003"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-019-09794-5"},{"key":"e_1_3_1_11_2","article-title":"Isotropic self-supervised learning for driver drowsiness detection with attention-based multimodal fusion","author":"Mou Luntian","year":"2021","unstructured":"Luntian Mou, Chao Zhou, Pengtao Xie, Pengfei Zhao, Ramesh C. Jain, Wen Gao, and Baocai Yin. 2021. Isotropic self-supervised learning for driver drowsiness detection with attention-based multimodal fusion. IEEE Transactions on Multimedia (2021).","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654930"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICOSC.2015.7050801"},{"issue":"1","key":"e_1_3_1_14_2","first-page":"43","article-title":"The impact of NLP on Turkish sentiment analysis","volume":"7","author":"Y\u0131ld\u0131r\u0131m Ezgi","year":"2015","unstructured":"Ezgi Y\u0131ld\u0131r\u0131m, Fatih Samet \u00c7etin, G\u00fcl\u015fen Eryi\u011fit, and Tanel Temel. 2015. The impact of NLP on Turkish sentiment analysis. T\u00fcrkiye Bili\u015fim Vakf\u0131 Bilgisayar Bilimleri ve M\u00fchendisli\u011fi Dergisi 7, 1 (2015), 43\u201351.","journal-title":"T\u00fcrkiye Bili\u015fim Vakf\u0131 Bilgisayar Bilimleri ve M\u00fchendisli\u011fi Dergisi"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2015.2432791"},{"issue":"3","key":"e_1_3_1_16_2","doi-asserted-by":"crossref","first-page":"632","DOI":"10.1109\/TMM.2016.2617741","article-title":"Continuous probability distribution prediction of image emotions via multitask shared sparse regression","volume":"19","author":"Zhao Sicheng","year":"2016","unstructured":"Sicheng Zhao, Hongxun Yao, Yue Gao, Rongrong Ji, and Guiguang Ding. 2016. Continuous probability distribution prediction of image emotions via multitask shared sparse regression. IEEE Transactions on Multimedia 19, 3 (2016), 632\u2013645.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jnca.2019.102447"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2018.11.001"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-010-0182-0"},{"key":"e_1_3_1_20_2","first-page":"973","volume-title":"Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"P\u00e9rez-Rosas Ver\u00f3nica","year":"2013","unstructured":"Ver\u00f3nica P\u00e9rez-Rosas, Rada Mihalcea, and Louis-Philippe Morency. 2013. Utterance-level multimodal sentiment analysis. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 973\u2013982."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.5555\/2969442.2969523"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2012.10.009"},{"key":"e_1_3_1_23_2","first-page":"3718","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Yu Wenmeng","year":"2020","unstructured":"Wenmeng Yu, Hua Xu, Fanyang Meng, Yilin Zhu, Yixiao Ma, Jiele Wu, Jiyun Zou, and Kaicheng Yang. 2020. Ch-sims: A Chinese multimodal sentiment analysis dataset with fine-grained annotation of modality. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 3718\u20133727."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3136755.3136801"},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","first-page":"1033","DOI":"10.1109\/ICDM.2017.134","volume-title":"2017 IEEE International Conference on Data Mining (ICDM\u201917)","author":"Poria Soujanya","year":"2017","unstructured":"Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Mazumder, Amir Zadeh, and Louis-Philippe Morency. 2017. Multi-level multiple attentions for contextual multimodal sentiment analysis. In 2017 IEEE International Conference on Data Mining (ICDM\u201917). IEEE, 1033\u20131038."},{"key":"e_1_3_1_26_2","first-page":"13","volume-title":"Proceedings of the 9th ACM International Conference on Web Search and Data Mining","author":"You Quanzeng","year":"2016","unstructured":"Quanzeng You, Jiebo Luo, Hailin Jin, and Jianchao Yang. 2016. Cross-modality consistent regression for joint visual-textual sentiment analysis of social multimedia. In Proceedings of the 9th ACM International Conference on Web Search and Data Mining. 13\u201322."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12021"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.3390\/s21144927"},{"key":"e_1_3_1_29_2","volume-title":"31st AAAI Conference on Artificial Intelligence","author":"Li Linghui","year":"2017","unstructured":"Linghui Li, Sheng Tang, Lixi Deng, Yongdong Zhang, and Qi Tian. 2017. Image caption with global-local attention. In 31st AAAI Conference on Artificial Intelligence."},{"issue":"4","key":"e_1_3_1_30_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3271485","article-title":"Image captioning via semantic guidance attention and consensus selection strategy","volume":"14","author":"Wu Jie","year":"2018","unstructured":"Jie Wu, Haifeng Hu, and Yi Wu. 2018. Image captioning via semantic guidance attention and consensus selection strategy. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) 14, 4 (2018), 1\u201319.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3231737"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2017.02.003"},{"issue":"3","key":"e_1_3_1_33_2","doi-asserted-by":"crossref","first-page":"327","DOI":"10.1080\/0305764X.2020.1831440","article-title":"Refining the teacher emotion model: Evidence from a review of literature published between 1985 and 2019","volume":"51","author":"Chen Junjun","year":"2021","unstructured":"Junjun Chen. 2021. Refining the teacher emotion model: Evidence from a review of literature published between 1985 and 2019. Cambridge Journal of Education 51, 3 (2021), 327\u2013357.","journal-title":"Cambridge Journal of Education"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324916000334"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2017.07.007"},{"issue":"1","key":"e_1_3_1_36_2","first-page":"80","article-title":"Advances in emotion recognition based on physiological big data","volume":"53","author":"Zhao Guozhen","year":"2016","unstructured":"Guozhen Zhao, Jinjing Song, Yan Ge, Yongjin Liu, Lin Yao, and Tao Wen. 2016. Advances in emotion recognition based on physiological big data. Journal of Computer Research and Development 53, 1 (2016), 80.","journal-title":"Journal of Computer Research and Development"},{"key":"e_1_3_1_37_2","unstructured":"ReadFace. 2020. ReadFace webpage on 36Kr. http:\/\/36kr.com\/p\/5038637.html. (2020)."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-017-1134-1"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2014.04.022"},{"key":"e_1_3_1_40_2","first-page":"5139","volume-title":"IJCAI","author":"Mao Qianren","year":"2019","unstructured":"Qianren Mao, Jianxin Li, Senzhang Wang, Yuanning Zhang, Hao Peng, Min He, and Lihong Wang. 2019. Aspect-based sentiment classification with attentive neural turing machines. In IJCAI. 5139\u20135145."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11280-013-0221-9"},{"key":"e_1_3_1_42_2","first-page":"4595","volume-title":"IJCAI","author":"Zhang Yuxiang","year":"2018","unstructured":"Yuxiang Zhang, Jiamei Fu, Dongyu She, Ying Zhang, Senzhang Wang, and Jufeng Yang. 2018. Text emotion distribution learning via multi-task convolutional neural network. In IJCAI. 4595\u20134601."},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","first-page":"4351","DOI":"10.18653\/v1\/2020.acl-main.401","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Chauhan Dushyant Singh","year":"2020","unstructured":"Dushyant Singh Chauhan, S. R. Dhanush, Asif Ekbal, and Pushpak Bhattacharyya. 2020. Sentiment and emotion help sarcasm? A multi-task learning framework for multi-modal sarcasm, sentiment and emotion analysis. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 4351\u20134360."},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.3390\/s21155135"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2017.2738401"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1046"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.3390\/s20082308"},{"key":"e_1_3_1_48_2","article-title":"Learning alignment for multimodal emotion recognition from speech","author":"Xu Haiyang","year":"2019","unstructured":"Haiyang Xu, Hui Zhang, Kun Han, Yun Wang, Yiping Peng, and Xiangang Li. 2019. Learning alignment for multimodal emotion recognition from speech. arXiv preprint arXiv:1909.05645 (2019).","journal-title":"arXiv preprint arXiv:1909.05645"},{"key":"e_1_3_1_49_2","article-title":"Uncertainty and surprisal jointly deliver the punchline: Exploiting incongruity-based features for humor recognition","author":"Xie Yubo","year":"2020","unstructured":"Yubo Xie, Junze Li, and Pearl Pu. 2020. Uncertainty and surprisal jointly deliver the punchline: Exploiting incongruity-based features for humor recognition. arXiv preprint arXiv:2012.12007 (2020).","journal-title":"arXiv preprint arXiv:2012.12007"},{"key":"e_1_3_1_50_2","first-page":"2364","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Wang Xiangyu","year":"2021","unstructured":"Xiangyu Wang and Chengqing Zong. 2021. Distributed representations of emotion categories in emotion space. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2364\u20132375."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2020.08.005"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240533"},{"key":"e_1_3_1_53_2","first-page":"371","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Xu Nan","year":"2019","unstructured":"Nan Xu, Wenji Mao, and Guandan Chen. 2019. Multi-interactive memory network for aspect based multimodal sentiment analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 371\u2013378."},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.108102"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.108498"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3057733"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3035277"},{"key":"e_1_3_1_58_2","first-page":"305","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Truong Quoc-Tuan","year":"2019","unstructured":"Quoc-Tuan Truong and Hady W. Lauw. 2019. Vistanet: Visual aspect attention network for multimodal sentiment analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 305\u2013312."},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3132847.3133142"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2019.06.010"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2020.06.050"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.3301996"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.345"},{"key":"e_1_3_1_64_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"31","author":"Mun Jonghwan","year":"2017","unstructured":"Jonghwan Mun, Minsu Cho, and Bohyung Han. 2017. Text-guided attention model for image captioning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 31."},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12272"},{"key":"e_1_3_1_66_2","article-title":"Evaluating the ability of LSTMs to learn context-free grammars","author":"Sennhauser Luzi","year":"2018","unstructured":"Luzi Sennhauser and Robert C. Berwick. 2018. Evaluating the ability of LSTMs to learn context-free grammars. arXiv preprint arXiv:1811.02611 (2018).","journal-title":"arXiv preprint arXiv:1811.02611"},{"key":"e_1_3_1_67_2","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Long Papers)","author":"Zadeh Amir","year":"2018","unstructured":"Amir Zadeh and Paul Pu. 2018. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Long Papers)."},{"key":"e_1_3_1_68_2","article-title":"MOSI: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos","author":"Zadeh Amir","year":"2016","unstructured":"Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency. 2016. MOSI: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos. arXiv preprint arXiv:1606.06259 (2016).","journal-title":"arXiv preprint arXiv:1606.06259"},{"key":"e_1_3_1_69_2","article-title":"DialogueCRN: Contextual reasoning networks for emotion recognition in conversations","author":"Hu Dou","year":"2021","unstructured":"Dou Hu, Lingwei Wei, and Xiaoyong Huai. 2021. DialogueCRN: Contextual reasoning networks for emotion recognition in conversations. arXiv preprint arXiv:2106.01978 (2021).","journal-title":"arXiv preprint arXiv:2106.01978"},{"key":"e_1_3_1_70_2","first-page":"6319","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Li Ruifan","year":"2021","unstructured":"Ruifan Li, Hao Chen, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang, and Eduard Hovy. 2021. Dual graph convolutional networks for aspect-based sentiment analysis. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 6319\u20136329."},{"issue":"6","key":"e_1_3_1_71_2","doi-asserted-by":"crossref","first-page":"1343","DOI":"10.3390\/s19061343","article-title":"Visnet: Deep convolutional neural networks for forecasting atmospheric visibility","volume":"19","author":"Palvanov Akmaljon","year":"2019","unstructured":"Akmaljon Palvanov and Young Im Cho. 2019. Visnet: Deep convolutional neural networks for forecasting atmospheric visibility. Sensors 19, 6 (2019), 1343.","journal-title":"Sensors"},{"key":"e_1_3_1_72_2","article-title":"Tensor fusion network for multimodal sentiment analysis","author":"Zadeh Amir","year":"2017","unstructured":"Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2017. Tensor fusion network for multimodal sentiment analysis. arXiv preprint arXiv:1707.07250 (2017).","journal-title":"arXiv preprint arXiv:1707.07250"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3572915","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3572915","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:51:38Z","timestamp":1750182698000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3572915"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,24]]},"references-count":71,"journal-issue":{"issue":"3s","published-print":{"date-parts":[[2023,10,31]]}},"alternative-id":["10.1145\/3572915"],"URL":"https:\/\/doi.org\/10.1145\/3572915","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,24]]},"assertion":[{"value":"2022-07-17","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-11-20","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-24","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}