{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T02:26:30Z","timestamp":1781144790080,"version":"3.54.1"},"reference-count":70,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,2,6]],"date-time":"2023-02-06T00:00:00Z","timestamp":1675641600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62076262"],"award-info":[{"award-number":["62076262"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,5,31]]},"abstract":"<jats:p>Multimodal sequence analysis aims to draw inferences from visual, language, and acoustic sequences. A majority of existing works focus on the aligned fusion of three modalities to explore inter-modal interactions, which is impractical in real-world scenarios. To overcome this issue, we seek to focus on analyzing unaligned sequences, which is still relatively underexplored and also more challenging. We propose Multimodal Graph, whose novelty mainly lies in transforming the sequential learning problem into graph learning problem. The graph-based structure enables parallel computation in time dimension (as opposed to recurrent neural network) and can effectively learn longer intra- and inter-modal temporal dependency in unaligned sequences. First, we propose multiple ways to construct the adjacency matrix for sequence to perform sequence to graph transformation. To learn intra-modal dynamics, a graph convolution network is employed for each modality based on the defined adjacency matrix. To learn inter-modal dynamics, given that the unimodal sequences are unaligned, the commonly considered word-level fusion does not pertain. To this end, we innovatively devise graph pooling algorithms to automatically explore the associations between various time slices from different modalities and learn high-level graph representation hierarchically. Multimodal Graph outperforms state-of-the-art models on three datasets under the same experimental setting.<\/jats:p>","DOI":"10.1145\/3542927","type":"journal-article","created":{"date-parts":[[2022,6,9]],"date-time":"2022-06-09T13:13:51Z","timestamp":1654780431000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":24,"title":["Multimodal Graph for Unaligned Multimodal Sequence Analysis via Graph Convolution and Graph Pooling"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9763-375X","authenticated-orcid":false,"given":"Sijie","family":"Mai","sequence":"first","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-sen University, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2734-1695","authenticated-orcid":false,"given":"Songlong","family":"Xing","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-sen University, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1419-3670","authenticated-orcid":false,"given":"Jiaxuan","family":"He","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-sen University, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8842-2045","authenticated-orcid":false,"given":"Ying","family":"Zeng","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-sen University, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4884-323X","authenticated-orcid":false,"given":"Haifeng","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Electronics and Information Technology, Sun Yat-sen University, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,2,6]]},"reference":[{"key":"e_1_3_3_2_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Bai Shaojie","year":"2019","unstructured":"Shaojie Bai, J. Kolter, and Vladlen Koltun. 2019. Trellis networks for sequence modeling. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_3_3_2","article-title":"An empirical evaluation of generic convolutional and recurrent networks for sequence modeling","author":"Bai Shaojie","year":"2018","unstructured":"Shaojie Bai, J. Zico Kolter, and Vladlen Koltun. 2018. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv: 1803.01271. Retrieved from https:\/\/arxiv.org\/abs\/1803.01271.","journal-title":"arXiv: 1803.01271"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2798607"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/72.279181"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-008-9076-6"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3136755.3136801"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1179"},{"key":"e_1_3_3_9_2","first-page":"960","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Degottex Gilles","year":"2014","unstructured":"Gilles Degottex, John Kane, Thomas Drugman, Tuomo Raitio, and Stefan Scherer. 2014. COVAREP: A collaborative voice analysis repository for speech technologies. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. 960\u2013964."},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_3_3_11_2","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Franceschi Luca","year":"2019","unstructured":"Luca Franceschi, Mathias Niepert, Massimiliano Pontil, and Xiao He. 2019. Learning discrete structures for graph neural networks. In Proceedings of the International Conference on Machine Learning."},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3412846"},{"key":"e_1_3_3_13_2","first-page":"12743","article-title":"Multi-modal graph neural network for joint reasoning on vision and scene text","author":"Gao Difei","year":"2020","unstructured":"Difei Gao, Ke Li, R. Wang, S. Shan, and X. Chen. 2020. Multi-modal graph neural network for joint reasoning on vision and scene text. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920), 12743\u201312753.","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2020.09.005"},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/72.286928"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143891"},{"key":"e_1_3_3_17_2","first-page":"1024","volume-title":"Advances in Neural Information Processing Systems","author":"Hamilton Will","year":"2017","unstructured":"Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive representation learning on large graphs. In Advances in Neural Information Processing Systems. 1024\u20131034."},{"key":"e_1_3_3_18_2","doi-asserted-by":"crossref","DOI":"10.1145\/3394171.3413678","article-title":"MISA: Modality-invariant and -specific representations for multimodal sentiment analysis","author":"Hazarika Devamanyu","year":"2020","unstructured":"Devamanyu Hazarika, R. Zimmermann, and Soujanya Poria. 2020. MISA: Modality-invariant and -specific representations for multimodal sentiment analysis. In Proceedings of the 28th ACM International Conference on Multimedia.","journal-title":"Proceedings of the 28th ACM International Conference on Multimedia"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_3_21_2","first-page":"12113","volume-title":"Advances in Neural Information Processing Systems","author":"Hou Ming","year":"2019","unstructured":"Ming Hou, Jiajia Tang, Jianhai Zhang, Wanzeng Kong, and Qibin Zhao. 2019. Deep multimodal multilinear fusion with high-order polynomial pooling. In Advances in Neural Information Processing Systems. 12113\u201312122."},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3388861"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2867718"},{"key":"e_1_3_3_24_2","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Kampman Onno","year":"2018","unstructured":"Onno Kampman, Elham J. Barezi, Dario Bertero, and Pascale Fung. 2018. Investigating audio, visual, and text fusion methods for end-to-end automatic personality prediction. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 606\u2013611."},{"key":"e_1_3_3_25_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201915)","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations (ICLR\u201915)."},{"key":"e_1_3_3_26_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Kipf Thomas N.","year":"2016","unstructured":"Thomas N. Kipf and Max Welling. 2016. Semi-supervised classification with graph convolutional networks. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2020.08.006"},{"key":"e_1_3_3_28_2","doi-asserted-by":"crossref","first-page":"1569","DOI":"10.18653\/v1\/P19-1152","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Liang Paul Pu","year":"2019","unstructured":"Paul Pu Liang, Zhun Liu, Yao-Hung Hubert Tsai, Qibin Zhao, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2019. Learning representations from imperfect time series data via tensor rank regularization. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 1569\u20131576."},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1014"},{"key":"e_1_3_3_30_2","first-page":"2247","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics","author":"Liu Zhun","year":"2018","unstructured":"Zhun Liu, Ying Shen, Paul Pu Liang, Amir Zadeh, and Louis Philippe Morency. 2018. Efficient low-rank multimodal fusion with modality-specific factors. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. 2247\u20132256."},{"key":"e_1_3_3_31_2","first-page":"481","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Mai Sijie","year":"2019","unstructured":"Sijie Mai, Haifeng Hu, and Songlong Xing. 2019. Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 481\u2013492."},{"key":"e_1_3_3_32_2","first-page":"164","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Mai Sijie","year":"2020","unstructured":"Sijie Mai, Haifeng Hu, and Songlong Xing. 2020. Modality to modality translation: An adversarial representation learning and graph fusion network for multimodal fusion. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 164\u2013172."},{"key":"e_1_3_3_33_2","doi-asserted-by":"crossref","unstructured":"S. Mai H. Hu J. Xu and S. Xing. 2022. Multi-fusion residual memory network for multimodal human sentiment comprehension. IEEE Trans. Affective Comput. 13 1 (2022) 320\u2013334.","DOI":"10.1109\/TAFFC.2020.3000510"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2925966"},{"key":"e_1_3_3_35_2","doi-asserted-by":"crossref","unstructured":"Sijie Mai Songlong Xing and Haifeng Hu. 2021. Analyzing multimodal language via acoustic-and visual-LSTM with channel-aware temporal convolution network. IEEE\/ACM Trans. Aud. Speech Lang. Process 29 (2021) 1424\u20131437.","DOI":"10.1109\/TASLP.2021.3068598"},{"key":"e_1_3_3_36_2","article-title":"Hybrid contrastive learning of tri-modal representation for multimodal sentiment analysis","author":"Mai Sijie","year":"2022","unstructured":"Sijie Mai, Ying Zeng, Shuangjia Zheng, and Haifeng Hu. 2022. Hybrid contrastive learning of tri-modal representation for multimodal sentiment analysis. IEEE Trans. Affect. Comput. (2022).","journal-title":"IEEE Trans. Affect. Comput."},{"key":"e_1_3_3_37_2","doi-asserted-by":"crossref","unstructured":"Sijie Mai Shuangjia Zheng Yuedong Yang and Haifeng Hu. 2021. Communicative message passing for inductive relation reasoning. In Proceedings of the 35th AAAI Conference on Artificial Intelligence . 4294\u20134302.","DOI":"10.1609\/aaai.v35i5.16554"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2008.2010350"},{"key":"e_1_3_3_39_2","article-title":"Multi-modal retrieval using graph neural networks","volume":"2010","author":"Misraa Aashish Kumar","year":"2020","unstructured":"Aashish Kumar Misraa, Ajinkya Kale, Pranav Aggarwal, and A. Aminian. 2020. Multi-modal retrieval using graph neural networks. arXiv: abs\/2010.01666. Retrieved from https:\/\/arxiv.org\/abs\/2010.01666.","journal-title":"arXiv:"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/2993148.2993176"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.17763\/haer.47.3.8840364413869005"},{"key":"e_1_3_3_42_2","first-page":"6875","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Pandey Ashutosh","year":"2019","unstructured":"Ashutosh Pandey and DeLiang Wang. 2019. TCNN: Temporal convolutional neural network for real-time speech enhancement in the time domain. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 6875\u20136879."},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_3_44_2","first-page":"6892","volume-title":"Proceedings of the 33rd AAAI Conference on Artificial Intelligence","author":"Pham Hai","year":"2019","unstructured":"Hai Pham, Paul Pu Liang, Thomas Manzini, Louis Philippe Morency, and Pocz\u01d2s Barnab\u01ces. 2019. Found in translation: Learning robust joint representations by cyclic translations between modalities. In Proceedings of the 33rd AAAI Conference on Artificial Intelligence. 6892\u20136899."},{"key":"e_1_3_3_45_2","first-page":"873","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics","author":"Poria Soujanya","year":"2017","unstructured":"Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis Philippe Morency. 2017. Context-dependent sentiment analysis in user-generated videos. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics. 873\u2013883."},{"key":"e_1_3_3_46_2","first-page":"439","volume-title":"Proceedings of the IEEE International Conference on Data Mining (ICDM\u201916)","author":"Poria Soujanya","year":"2016","unstructured":"Soujanya Poria, Iti Chaturvedi, Erik Cambria, and Amir Hussain. 2016. Convolutional MKL based multimodal emotion recognition and sentiment analysis. In Proceedings of the IEEE International Conference on Data Mining (ICDM\u201916). 439\u2013448."},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2008.2005605"},{"issue":"2","key":"e_1_3_3_48_2","first-page":"663","article-title":"Host\u2013parasite: Graph LSTM-in-LSTM for group activity recognition","volume":"32","author":"Shu Xiangbo","year":"2020","unstructured":"Xiangbo Shu, Liyan Zhang, Yunlian Sun, and Jinhui Tang. 2020. Host\u2013parasite: Graph LSTM-in-LSTM for group activity recognition. IEEE Trans. Neural Netw. Learn. Syst. 32, 2 (2020), 663\u2013674.","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2906603"},{"key":"e_1_3_3_50_2","doi-asserted-by":"crossref","first-page":"6558","DOI":"10.18653\/v1\/P19-1656","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Tsai Yao-Hung Hubert","year":"2019","unstructured":"Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 6558\u20136569."},{"key":"e_1_3_3_51_2","first-page":"1823","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Tsai Yao-Hung Hubert","year":"2020","unstructured":"Yao-Hung Hubert Tsai, Martin Ma, Muqiao Yang, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2020. Multimodal routing: Improving local and global interpretability of multimodal language analysis. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. 1823\u20131833."},{"key":"e_1_3_3_52_2","first-page":"5998","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_3_53_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Veli\u010dkovi\u0107 Petar","year":"2018","unstructured":"Petar Veli\u010dkovi\u0107, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2018. Graph attention networks. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2663324"},{"key":"e_1_3_3_55_2","first-page":"7216","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Wang Yansen","year":"2019","unstructured":"Yansen Wang, Ying Shen, Zhun Liu, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2019. Words can shift: Dynamically adjusting word representations using nonverbal behaviors. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 7216\u20137223."},{"key":"e_1_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/MIS.2013.34"},{"key":"e_1_3_3_57_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Xu Keyulu","year":"2019","unstructured":"Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. 2019. How powerful are graph neural networks? In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_3_58_2","article-title":"MTGAT: Multimodal temporal graph attention networks for unaligned human multimodal language sequences","author":"Yang Jianing","year":"2020","unstructured":"Jianing Yang, Yongxin Wang, Ruitao Yi, Yuying Zhu, Azaan Rehman, Amir Zadeh, Soujanya Poria, and Louis-Philippe Morency. 2020. MTGAT: Multimodal temporal graph attention networks for unaligned human multimodal language sequences. arXiv:2010.11985. Retrieved from https:\/\/arxiv.org\/abs\/2010.11985.","journal-title":"arXiv:2010.11985"},{"key":"e_1_3_3_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413690"},{"key":"e_1_3_3_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3035277"},{"key":"e_1_3_3_61_2","first-page":"4800","volume-title":"Advances in Neural Information Processing Systems","author":"Ying Zhitao","year":"2018","unstructured":"Zhitao Ying, Jiaxuan You, Christopher Morris, Xiang Ren, Will Hamilton, and Jure Leskovec. 2018. Hierarchical graph representation learning with differentiable pooling. In Advances in Neural Information Processing Systems. 4800\u20134810."},{"key":"e_1_3_3_62_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Yu Fisher","year":"2016","unstructured":"Fisher Yu and Vladlen Koltun. 2016. Multi-scale context aggregation by dilated convolutions. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_3_63_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Yuan Hao","year":"2020","unstructured":"Hao Yuan and Shuiwang Ji. 2020. StructPool: Structured graph pooling via conditional random fields. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_3_64_2","doi-asserted-by":"publisher","DOI":"10.1121\/1.2935783"},{"key":"e_1_3_3_65_2","first-page":"1114","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Zadeh Amir","year":"2017","unstructured":"Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis Philippe Morency. 2017. Tensor fusion network for multimodal sentiment analysis. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. 1114\u20131125."},{"key":"e_1_3_3_66_2","first-page":"5634","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Zadeh Amir","year":"2018","unstructured":"Amir Zadeh, Paul Pu Liang, Navonil Mazumder, Soujanya Poria, Erik Cambria, and Louis Philippe Morency. 2018. Memory fusion network for multi-view sequential learning. In Proceedings of the AAAI Conference on Artificial Intelligence. 5634\u20135641."},{"key":"e_1_3_3_67_2","first-page":"2236","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics","author":"Zadeh Amir","year":"2018","unstructured":"Amir Zadeh, Paul Pu Liang, Jonathan Vanbriesen, Soujanya Poria, Edmund Tong, Erik Cambria, Minghai Chen, and Louis Philippe Morency. 2018. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. 2236\u20132246."},{"key":"e_1_3_3_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/MIS.2016.94"},{"key":"e_1_3_3_69_2","first-page":"5165","volume-title":"Advances in Neural Information Processing Systems","author":"Zhang Muhan","year":"2018","unstructured":"Muhan Zhang and Yixin Chen. 2018. Link prediction based on graph neural networks. In Advances in Neural Information Processing Systems. 5165\u20135175."},{"key":"e_1_3_3_70_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11782"},{"key":"e_1_3_3_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/3363560"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3542927","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3542927","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:02:22Z","timestamp":1750186942000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3542927"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,6]]},"references-count":70,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,5,31]]}},"alternative-id":["10.1145\/3542927"],"URL":"https:\/\/doi.org\/10.1145\/3542927","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,6]]},"assertion":[{"value":"2022-01-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-05-31","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}