{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T05:17:33Z","timestamp":1784179053614,"version":"3.55.0"},"reference-count":108,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2025,3,15]],"date-time":"2025-03-15T00:00:00Z","timestamp":1741996800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62402491, U2336202, 62302345, and U23A20305"],"award-info":[{"award-number":["62402491, U2336202, 62302345, and U23A20305"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"CCF-ALIMAMA TECH Kangaroo","award":["CCF-ALIMAMA OF 2024009"],"award-info":[{"award-number":["CCF-ALIMAMA OF 2024009"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2025,5,31]]},"abstract":"<jats:p>Multimodal recommendation, which utilizes rich multimodal information to learn user preferences, has attracted significant attention. Most works focus on designing powerful encoders for extracting multimodal features, and simply aggregate the learned features together to make prediction. Consequently, they have a limited capacity to learn the inter-modality knowledge including the modality-shared and modality-unique knowledge. In fact, learning the modality-shared knowledge enables us to align cross-modality data for fusing heterogeneous modality features. Learning the modality-unique knowledge is equally important when recommendation tasks only involve a small amount of shared features and the necessary information is contained within specific modality. In this article, we propose Contrastive Modality-Disentangled Learning (CMDL) to overcome this critical limitation. CMDL exactly captures the inter-modality knowledge by achieving modality disentanglement. Specifically, CMDL first disentangles the initial representation into the modality-invariant and modality-specific representations. Afterwards, CMDL introduces a novel manner of contrastive learning to approximate the MI upper bounds for achieving disentanglement regularization. Building upon the proposed regularization, CMDL encourages the modality-invariant and modality-specific representations to capture the modality-shared and modality-unique knowledge respectively and to be statistically independent to each other. Empirically, extensive experiments are conducted on benchmark datasets, demonstrating the superior performance of CMDL compared with strong multimodal recommenders.<\/jats:p>","DOI":"10.1145\/3715876","type":"journal-article","created":{"date-parts":[[2025,1,28]],"date-time":"2025-01-28T15:44:14Z","timestamp":1738079054000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":41,"title":["Contrastive Modality-Disentangled Learning for Multimodal Recommendation"],"prefix":"10.1145","volume":"43","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-6645-0597","authenticated-orcid":false,"given":"Xixun","family":"Lin","sequence":"first","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3254-0170","authenticated-orcid":false,"given":"Rui","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Cyber Security, University of Chinese Academy of Sciences, Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3534-1094","authenticated-orcid":false,"given":"Yanan","family":"Cao","sequence":"additional","affiliation":[{"name":"School of Cyber Security, University of Chinese Academy of Sciences, Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6755-871X","authenticated-orcid":false,"given":"Lixin","family":"Zou","sequence":"additional","affiliation":[{"name":"School of Cyber Science and Engineering, Wuhan University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8308-9551","authenticated-orcid":false,"given":"Qian","family":"Li","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering, Computing and Mathematical Sciences, Curtin University, Perth, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-1622-5434","authenticated-orcid":false,"given":"Yongxuan","family":"Wu","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3791-4343","authenticated-orcid":false,"given":"Yang","family":"Liu","sequence":"additional","affiliation":[{"name":"Academy of Mathematics and Systems Science, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0684-6205","authenticated-orcid":false,"given":"Dawei","family":"Yin","sequence":"additional","affiliation":[{"name":"Baidu Inc, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4493-6663","authenticated-orcid":false,"given":"Guandong","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Computer Science, University of Technology Sydney, Sydney, Australia and The Education University of Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,3,15]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Alexander A. Alemi Ian Fischer Joshua V. Dillon and Kevin Murphy. 2016. Deep variational information bottleneck. arXiv:1612.00410. Retrieved from https:\/\/arxiv.org\/abs\/1612.00410"},{"key":"e_1_3_2_3_2","first-page":"4764","volume-title":"Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"An Zhicheng","year":"2024","unstructured":"Zhicheng An, Zhexu Gu, Li Yu, Ke Tu, Zhengwei Wu, Binbin Hu, Zhiqiang zhang, Lihong Gu, and Jinjie Gu. 2024. DCDR: A disentangle-based distillation framework for cross-domain drecommendation. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 4764\u20134773."},{"key":"e_1_3_2_4_2","first-page":"214","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Arjovsky Martin","year":"2017","unstructured":"Martin Arjovsky, Soumith Chintala, and L\u00e9on Bottou. 2017. Wasserstein generative adversarial networks. In Proceedings of the International Conference on Machine Learning. PMLR, 214\u2013223."},{"key":"e_1_3_2_5_2","volume-title":"Information Theory","author":"Ash Robert B.","year":"2012","unstructured":"Robert B. Ash. 2012. Information Theory. Courier Corporation."},{"issue":"2","key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"423","DOI":"10.1109\/TPAMI.2018.2798607","article-title":"Multimodal machine learning: A survey and taxonomy","volume":"41","author":"Baltru\u0161aitis Tadas","year":"2018","unstructured":"Tadas Baltru\u0161aitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2018. Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence 41, 2 (2018), 423\u2013443.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"320","key":"e_1_3_2_7_2","first-page":"201","article-title":"The im algorithm: A variational approach to information maximization","volume":"16","author":"Barber David","year":"2004","unstructured":"David Barber and Felix Agakov. 2004. The im algorithm: A variational approach to information maximization. Advances in Neural Information Processing Systems 16, 320 (2004), 201.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"518","key":"e_1_3_2_8_2","doi-asserted-by":"crossref","first-page":"859","DOI":"10.1080\/01621459.2017.1285773","article-title":"Variational inference: A review for statisticians","volume":"112","author":"Blei David M.","year":"2017","unstructured":"David M. Blei, Alp Kucukelbir, and Jon D. McAuliffe. 2017. Variational inference: A review for statisticians. Journal of the American Statistical Association 112, 518 (2017), 859\u2013877.","journal-title":"Journal of the American Statistical Association"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1016\/j.knosys.2013.03.012","article-title":"Recommender systems survey","volume":"46","author":"Bobadilla Jes\u00fas","year":"2013","unstructured":"Jes\u00fas Bobadilla, Fernando Ortega, Antonio Hernando, and Abraham Guti\u00e9rrez. 2013. Recommender systems survey. Knowledge-Based Systems 46 (2013), 109\u2013132.","journal-title":"Knowledge-Based Systems"},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","first-page":"418","DOI":"10.1142\/9789814447331_0040","volume-title":"Biocomputing 2000","author":"Butte Atul J.","year":"1999","unstructured":"Atul J. Butte and Isaac S. Kohane. 1999. Mutual information relevance networks: functional genomic clustering using pairwise entropy measurements. In Biocomputing 2000. World Scientific, 418\u2013429."},{"issue":"1","key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1080\/03610927408827101","article-title":"A dendrite method for cluster analysis","volume":"3","author":"Calinski T.","year":"1974","unstructured":"T. Calinski and J. Harabasz. 1974. A dendrite method for cluster analysis. Communications in Statistics \u2013 Theory and Methods 3, 1 (1974), 1\u201327.","journal-title":"Communications in Statistics \u2013 Theory and Methods"},{"key":"e_1_3_2_12_2","first-page":"267","volume-title":"Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Cao Jiangxia","year":"2022","unstructured":"Jiangxia Cao, Xixun Lin, Xin Cong, Jing Ya, Tingwen Liu, and Bin Wang. 2022. Disencdr: Learning disentangled representations for cross-domain recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, 267\u2013277."},{"key":"e_1_3_2_13_2","first-page":"9912","article-title":"Unsupervised learning of visual features by contrasting cluster assignments","volume":"33","author":"Caron Mathilde","year":"2020","unstructured":"Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. 2020. Unsupervised learning of visual features by contrasting cluster assignments. Advances in Neural Information Processing Systems 33 (2020), 9912\u20139924.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_14_2","first-page":"165","article-title":"Information bottleneck for gaussian variables","volume":"16","author":"Chechik Gal","year":"2003","unstructured":"Gal Chechik, Amir Globerson, Naftali Tishby, and Yair Weiss. 2003. Information bottleneck for gaussian variables. Advances in Neural Information Processing Systems 16 (2003), 165\u2013188.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_15_2","first-page":"9","article-title":"Fast approximate kNN graph construction for high dimensional data via recursive lanczos bisection","volume":"10","author":"Chen Jie","year":"2009","unstructured":"Jie Chen, Haw-Ren Fang, and Yousef Saad. 2009. Fast approximate kNN graph construction for high dimensional data via recursive lanczos bisection. Journal of Machine Learning Research 10 (2009), 9.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_16_2","first-page":"335","volume-title":"Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Chen Jingyuan","year":"2017","unstructured":"Jingyuan Chen, Hanwang Zhang, Xiangnan He, Liqiang Nie, Wei Liu, and Tat-Seng Chua. 2017. Attentive collaborative filtering: Multimedia recommendation with item-and component-level attention. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, 335\u2013344."},{"key":"e_1_3_2_17_2","first-page":"2615","article-title":"Isolating sources of disentanglement in variational autoencoders","volume":"31","author":"Chen Ricky T. Q.","year":"2018","unstructured":"Ricky T. Q. Chen, Xuechen Li, Roger B. Grosse, and David K. Duvenaud. 2018. Isolating sources of disentanglement in variational autoencoders. Advances in Neural Information Processing Systems 31 (2018), 2615\u20132625.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_18_2","first-page":"1597","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In Proceedings of the International Conference on Machine Learning. PMLR, 1597\u20131607."},{"key":"e_1_3_2_19_2","doi-asserted-by":"crossref","first-page":"765","DOI":"10.1145\/3331184.3331254","volume-title":"Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Chen Xu","year":"2019","unstructured":"Xu Chen, Hanxiong Chen, Hongteng Xu, Yongfeng Zhang, Yixin Cao, Zheng Qin, and Hongyuan Zha. 2019. Personalized fashion recommendation with visual explanations based on multimodal attention network: Towards visually explainable recommendation. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, 765\u2013774."},{"key":"e_1_3_2_20_2","first-page":"2180","article-title":"Infogan: Interpretable representation learning by information maximizing generative adversarial nets","volume":"29","author":"Chen Xi","year":"2016","unstructured":"Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel. 2016. Infogan: Interpretable representation learning by information maximizing generative adversarial nets. Advances in Neural Information Processing Systems 29 (2016), 2180\u20132188.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_21_2","first-page":"1779","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Cheng Pengyu","year":"2020","unstructured":"Pengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu, Zhe Gan, and Lawrence Carin. 2020. CLUB: A contrastive log-ratio upper bound of mutual information. In Proceedings of the International Conference on Machine Learning, 1779\u20131788."},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1016\/j.neunet.2020.11.011","article-title":"Bridging multimedia heterogeneity gap via graph representation learning for cross-modal retrieval","volume":"134","author":"Cheng Qingrong","year":"2021","unstructured":"Qingrong Cheng and Xiaodong Gu. 2021. Bridging multimedia heterogeneity gap via graph representation learning for cross-modal retrieval. Neural Networks: The Official Journal of the International Neural Network Society 134 (2021), 143\u2013162.","journal-title":"Neural Networks: The Official Journal of the International Neural Network Society"},{"key":"e_1_3_2_23_2","first-page":"1436","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Creager Elliot","year":"2019","unstructured":"Elliot Creager, David Madras, J\u00f6rn-Henrik Jacobsen, Marissa Weis, Kevin Swersky, Toniann Pitassi, and Richard Zemel. 2019. Flexibly fair representation learning by disentanglement. In Proceedings of the International Conference on Machine Learning. PMLR, 1436\u20131445."},{"issue":"2","key":"e_1_3_2_24_2","first-page":"317","article-title":"MV-RNN: A multi-view recurrent neural network for sequential recommendation","volume":"32","author":"Cui Qiang","year":"2018","unstructured":"Qiang Cui, Shu Wu, Qiang Liu, Wen Zhong, and Liang Wang. 2018. MV-RNN: A multi-view recurrent neural network for sequential recommendation. IEEE Transactions on Knowledge and Data Engineering 32, 2 (2018), 317\u2013331.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_3_2_25_2","unstructured":"Alexey Dosovitskiy. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_2_26_2","first-page":"2091","volume-title":"Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Du Jing","year":"2024","unstructured":"Jing Du, Zesheng Ye, Bin Guo, Zhiwen Yu, and Lina Yao. 2024. Identifiability of cross-domain recommendation via causal subspace disentanglement. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2091\u20132101."},{"key":"e_1_3_2_27_2","first-page":"1623","volume-title":"Proceedings of the 15th ACM International Conference on Web Search and Data Mining","author":"Gao Chen","year":"2022","unstructured":"Chen Gao, Xiang Wang, Xiangnan He, and Yong Li. 2022. Graph neural networks for recommender system. In Proceedings of the 15th ACM International Conference on Web Search and Data Mining, 1623\u20131625."},{"key":"e_1_3_2_28_2","unstructured":"Tianyu Gao Xingcheng Yao and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. arXiv:2104.08821. Retrieved from https:\/\/arxiv.org\/abs\/2104.08821"},{"key":"e_1_3_2_29_2","volume-title":"Proceedings of the International Conference on Artificial Intelligence and Statistics","author":"Glorot Xavier","year":"2010","unstructured":"Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the International Conference on Artificial Intelligence and Statistics. Retrieved from https:\/\/Api.semanticscholar.org\/CorpusID:5575601"},{"issue":"11","key":"e_1_3_2_30_2","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1145\/3422622","article-title":"Generative adversarial networks","volume":"63","author":"Goodfellow Ian","year":"2020","unstructured":"Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020. Generative adversarial networks. Communications of the ACM 63, 11 (2020), 139\u2013144.","journal-title":"Communications of the ACM"},{"key":"e_1_3_2_31_2","first-page":"2058","volume-title":"Proceedings of the ACM Web Conference 2022","author":"Han Tengyue","year":"2022","unstructured":"Tengyue Han, Pengfei Wang, Shaozhang Niu, and Chenliang Li. 2022. Modality matches modality: Pretraining modality-disentangled item representations for recommendation. In Proceedings of the ACM Web Conference 2022, 2058\u20132066."},{"key":"e_1_3_2_32_2","first-page":"9729","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"He Kaiming","year":"2020","unstructured":"Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 9729\u20139738."},{"key":"e_1_3_2_33_2","first-page":"507","volume-title":"Proceedings of the 25th International Conference on World Wide Web","author":"He Ruining","year":"2016","unstructured":"Ruining He and Julian McAuley. 2016. Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering. In Proceedings of the 25th International Conference on World Wide Web, 507\u2013517."},{"issue":"1","key":"e_1_3_2_34_2","first-page":"144","article-title":"VBPR: Visual Bayesian personalized ranking from implicit feedback","volume":"30","author":"He Ruining","year":"2016","unstructured":"Ruining He and Julian McAuley. 2016. VBPR: Visual Bayesian personalized ranking from implicit feedback. In Proceedings of the AAAI Conference on Artificial Intelligence 30, 1 (2016), 144\u2013150.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_2_35_2","first-page":"639","volume-title":"Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"He Xiangnan","year":"2020","unstructured":"Xiangnan He, Kuan Deng, Xiang Wang, Yan Li, Yongdong Zhang, and Meng Wang. 2020. Lightgcn: Simplifying and powering graph convolution network for recommendation. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, 639\u2013648."},{"key":"e_1_3_2_36_2","first-page":"173","volume-title":"Proceedings of the 26th International Conference on World Wide Web","author":"He Xiangnan","year":"2017","unstructured":"Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. 2017. Neural collaborative filtering. In Proceedings of the 26th International Conference on World Wide Web, 173\u2013182."},{"key":"e_1_3_2_37_2","article-title":"Beta-vae: Learning basic visual concepts with a constrained variational framework","volume":"3","author":"Higgins Irina","year":"2017","unstructured":"Irina Higgins, Loic Matthey, Arka Pal, Christopher P. Burgess, Xavier Glorot, Matthew M. Botvinick, Shakir Mohamed, and Alexander Lerchner. 2017. Beta-vae: Learning basic visual concepts with a constrained variational framework. In International Conference on Learning Representations (Poster), 3.","journal-title":"International Conference on Learning Representations (Poster)"},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","first-page":"128645","DOI":"10.1016\/j.neucom.2024.128645","article-title":"A comprehensive survey on contrastive learning","volume":"610","author":"Hu Haigen","year":"2024","unstructured":"Haigen Hu, Xiaoyuan Wang, Yan Zhang, Qi Chen, and Qiu Guan. 2024. A comprehensive survey on contrastive learning. Neurocomputing 610 (2024), 128645.","journal-title":"Neurocomputing"},{"issue":"6","key":"e_1_3_2_39_2","doi-asserted-by":"crossref","first-page":"10675","DOI":"10.1109\/JIOT.2019.2940709","article-title":"Multimodal representation learning for recommendation in internet of things","volume":"6","author":"Huang Zhenhua","year":"2019","unstructured":"Zhenhua Huang, Xin Xu, Juan Ni, Honghao Zhu, and Cheng Wang. 2019. Multimodal representation learning for recommendation in internet of things. IEEE Internet of Things Journal 6, 6 (2019), 10675\u201310685.","journal-title":"IEEE Internet of Things Journal"},{"issue":"1","key":"e_1_3_2_40_2","doi-asserted-by":"crossref","first-page":"2","DOI":"10.3390\/technologies9010002","article-title":"A survey on contrastive self-supervised learning","volume":"9","author":"Jaiswal Ashish","year":"2020","unstructured":"Ashish Jaiswal, Ashwin Ramesh Babu, Mohammad Zaki Zadeh, Debapriya Banerjee, and Fillia Makedon. 2020. A survey on contrastive self-supervised learning. Technologies 9, 1 (2020), 2.","journal-title":"Technologies"},{"key":"e_1_3_2_41_2","doi-asserted-by":"crossref","first-page":"4626","DOI":"10.1145\/3581783.3612006","volume-title":"Proceedings of the 31st ACM International Conference on Multimedia","author":"Jiang Chen","year":"2023","unstructured":"Chen Jiang, Hong Liu, Xuzheng Yu, Qing Wang, Yuan Cheng, Jia Xu, Zhongyi Liu, Qingpei Guo, Wei Chu, Ming Yang, et al. 2023. Dual-modal attention-enhanced text-video retrieval with triplet partial margin contrastive learning. In Proceedings of the 31st ACM International Conference on Multimedia, 4626\u20134636."},{"key":"e_1_3_2_42_2","first-page":"2649","volume-title":"International Conference on Machine Learning","author":"Kim Hyunjik","year":"2018","unstructured":"Hyunjik Kim and Andriy Mnih. 2018. Disentangling by factorising. In International Conference on Machine Learning. PMLR, 2649\u20132658."},{"key":"e_1_3_2_43_2","unstructured":"Diederik P. Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv:1312.6114. Retrieved from https:\/\/arxiv.org\/abs\/1312.6114"},{"key":"e_1_3_2_44_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Kipf Thomas N.","year":"2016","unstructured":"Thomas N. Kipf, and Max Welling. 2016. Semi-supervised classification with graph convolutional networks. In Proceedings of the International Conference on Learning Representations, 1\u201314."},{"key":"e_1_3_2_45_2","first-page":"84","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. Imagenet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems 25 (2012), 84\u201390.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_46_2","first-page":"3161","volume-title":"Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining","author":"Lei Chenyi","year":"2021","unstructured":"Chenyi Lei, Yong Liu, Lingzi Zhang, Guoxin Wang, Haihong Tang, Houqiang Li, and Chunyan Miao. 2021. Semi: A sequential multi-modal information transfer network for e-commerce micro-video recommendations. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, 3161\u20133171."},{"key":"e_1_3_2_47_2","unstructured":"Haochen Li Xin Zhou Luu Anh Tuan and Chunyan Miao. 2023. Rethinking negative pairs in code search. arXiv:2310.08069. Retrieved from https:\/\/arxiv.org\/abs\/2310.08069"},{"key":"e_1_3_2_48_2","first-page":"29461","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Li Jianxiong","year":"2024","unstructured":"Jianxiong Li, Jinliang Zheng, Yinan Zheng, Liyuan Mao, Xiao Hu, Sijie Cheng, Haoyi Niu, Jihao Liu, Yu Liu, Jingjing Liu, et al. 2024. DecisionNCE: Embodied multimodal representations via implicit preference learning. In Proceedings of the 41st International Conference on Machine Learning, 29461\u201329488."},{"key":"e_1_3_2_49_2","first-page":"32971","article-title":"Factorized contrastive learning: Going beyond multi-view redundancy","volume":"36","author":"Liang Paul Pu","year":"2024","unstructured":"Paul Pu Liang, Zihao Deng, Martin Q. Ma, James Y. Zou, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2024. Factorized contrastive learning: Going beyond multi-view redundancy. Advances in Neural Information Processing Systems 36 (2024), 32971\u201332998.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"1","key":"e_1_3_2_50_2","first-page":"1","article-title":"Multibench: Multiscale benchmarks for multimodal representation learning","volume":"2021","author":"Liang Paul Pu","year":"2021","unstructured":"Paul Pu Liang, Yiwei Lyu, Xiang Fan, Zetian Wu, Yun Cheng, Jason Wu, Leslie Chen, Peter Wu, Michelle A. Lee, Yuke Zhu, et al. 2021. Multibench: Multiscale benchmarks for multimodal representation learning. Advances in Neural Information Processing Systems 2021, DB1 (2021), 1\u201320.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_51_2","first-page":"360","volume-title":"Proceedings of the 2021 IEEE International Conference on Data Mining (ICDM)","author":"Lin Xixun","year":"2021","unstructured":"Xixun Lin, Jiangxia Cao, Peng Zhang, Chuan Zhou, Zhao Li, Jia Wu, and Bin Wang. 2021. Disentangled deep multivariate hawkes process for learning event sequences. In Proceedings of the 2021 IEEE International Conference on Data Mining (ICDM). IEEE, 360\u2013369."},{"issue":"4","key":"e_1_3_2_52_2","first-page":"1815","article-title":"Towards flexible and adaptive neural process for cold-start recommendation","volume":"36","author":"Lin Xixun","year":"2023","unstructured":"Xixun Lin, Chuan Zhou, Jia Wu, Lixin Zou, Shirui Pan, Yanan Cao, Bin Wang, Shuaiqiang Wang, and Dawei Yin. 2023. Towards flexible and adaptive neural process for cold-start recommendation. IEEE Transactions on Knowledge and Data Engineering 36, 4 (2023), 1815\u20131828.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_3_2_53_2","unstructured":"Zinan Lin Kiran Koshy Thekumparampil Giulia Fanti and Sewoong Oh. 2019. Infogan-cr: Disentangling generative adversarial networks with contrastive regularizers. arXiv:1906.06034. Retrieved from https:\/\/arxiv.org\/abs\/1906.06034"},{"key":"e_1_3_2_54_2","first-page":"7149","article-title":"Disentangled multimodal representation learning for recommendation","volume":"25","author":"Liu Fan","year":"2022","unstructured":"Fan Liu, Huilin Chen, Zhiyong Cheng, Anan Liu, Liqiang Nie, and Mohan Kankanhalli. 2022. Disentangled multimodal representation learning for recommendation. IEEE Transactions on Multimedia 25 (2022), 7149\u20137159.","journal-title":"IEEE Transactions on Multimedia"},{"issue":"2","key":"e_1_3_2_55_2","first-page":"1","article-title":"Dynamic multimodal fusion via meta-learning towards micro-video recommendation","volume":"42","author":"Liu Han","year":"2023","unstructured":"Han Liu, Yinwei Wei, Fan Liu, Wenjie Wang, Liqiang Nie, and Tat-Seng Chua. 2023. Dynamic multimodal fusion via meta-learning towards micro-video recommendation. ACM Transactions on Information Systems 42, 2 (2023), 1\u201326.","journal-title":"ACM Transactions on Information Systems"},{"issue":"6","key":"e_1_3_2_56_2","first-page":"5977","article-title":"HS-GCN: Hamming spatial graph convolutional networks for recommendation","volume":"35","author":"Liu Han","year":"2022","unstructured":"Han Liu, Yinwei Wei, Jianhua Yin, and Liqiang Nie.2022. HS-GCN: Hamming spatial graph convolutional networks for recommendation. IEEE Transactions on Knowledge and Data Engineering 35, 6 (2022), 5977\u20135990.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_3_2_57_2","first-page":"841","volume-title":"Proceedings of the 40th International ACM Sigir Conference on Research and Development in Information Retrieval","author":"Liu Qiang","year":"2017","unstructured":"Qiang Liu, Shu Wu, and Liang Wang. 2017. Deepstyle: Learning user preferences for visual recommendation. In Proceedings of the 40th International ACM Sigir Conference on Research and Development in Information Retrieval, 841\u2013844."},{"key":"e_1_3_2_58_2","unstructured":"Yinhan Liu. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_2_59_2","first-page":"4114","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Locatello Francesco","year":"2019","unstructured":"Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Sch\u00f6lkopf, and Olivier Bachem. 2019. Challenging common assumptions in the unsupervised learning of disentangled representations. In Proceedings of the International Conference on Machine Learning. PMLR, 4114\u20134124."},{"key":"e_1_3_2_60_2","unstructured":"Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv:1711.05101. Retrieved from https:\/\/arxiv.org\/abs\/1711.05101"},{"key":"e_1_3_2_61_2","first-page":"5711","article-title":"Learning disentangled representations for recommendation","volume":"32","author":"Ma Jianxin","year":"2019","unstructured":"Jianxin Ma, Chang Zhou, Peng Cui, Hongxia Yang, and Wenwu Zhu. 2019. Learning disentangled representations for recommendation. Advances in Neural Information Processing Systems 32 (2019), 5711\u20135722.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_62_2","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1145\/2766462.2767755","volume-title":"Proceedings of the 38th international ACM SIGIR Conference on Research and Development in Information Retrieval","author":"McAuley Julian","year":"2015","unstructured":"Julian McAuley, Christopher Targett, Qinfeng Shi, and Anton van den Hengel. 2015. Image-based recommendations on styles and substitutes. In Proceedings of the 38th international ACM SIGIR Conference on Research and Development in Information Retrieval, 43\u201352."},{"issue":"4","key":"e_1_3_2_63_2","doi-asserted-by":"crossref","first-page":"93","DOI":"10.1109\/TIT.1954.1057469","article-title":"Multivariate information transmission","volume":"4","author":"McGill William","year":"1954","unstructured":"William McGill. 1954. Multivariate information transmission. Transactions of the IRE Professional Group on Information Theory 4, 4 (1954), 93\u2013111.","journal-title":"Transactions of the IRE Professional Group on Information Theory"},{"key":"e_1_3_2_64_2","unstructured":"Leland McInnes John Healy and James Melville. 2018. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv:1802.03426. Retrieved from https:\/\/arxiv.org\/abs\/1802.03426"},{"key":"e_1_3_2_65_2","first-page":"689","volume-title":"Proceedings of the 28th International Conference on Machine Learning (ICML \u201911)","author":"Ngiam Jiquan","year":"2011","unstructured":"Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y. Ng. 2011. Multimodal deep learning. In Proceedings of the 28th International Conference on Machine Learning (ICML \u201911), 689\u2013696."},{"key":"e_1_3_2_66_2","first-page":"188","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Ni Jianmo","year":"2019","unstructured":"Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019. Justifying recommendations using distantly-labeled reviews and fine-grained aspects. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.), Association for Computational Linguistics, Hong Kong, China, 188\u2013197. DOI: 10.18653\/v1\/D19-1018"},{"key":"e_1_3_2_67_2","unstructured":"Aaron van den Oord Yazhe Li and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv:1807.03748. Retrieved from https:\/\/arxiv.org\/abs\/1807.03748"},{"key":"e_1_3_2_68_2","first-page":"5171","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Poole Ben","year":"2019","unstructured":"Ben Poole, Sherjil Ozair, Aaron Van Den Oord, Alex Alemi, and George Tucker. 2019. On variational bounds of mutual information. In Proceedings of the International Conference on Machine Learning. PMLR, 5171\u20135180."},{"key":"e_1_3_2_69_2","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1016\/j.eswa.2017.01.005","article-title":"Preference dynamics with multimodal user-item interactions in social media recommendation","volume":"74","author":"Rafailidis Dimitrios","year":"2017","unstructured":"Dimitrios Rafailidis, Pavlos Kefalas, and Yannis Manolopoulos. 2017. Preference dynamics with multimodal user-item interactions in social media recommendation. Expert Systems with Applications 74 (2017), 11\u201318.","journal-title":"Expert Systems with Applications"},{"key":"e_1_3_2_70_2","first-page":"3982","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Reimers Nils","year":"2019","unstructured":"Nils Reimers, and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 3982\u20133992."},{"key":"e_1_3_2_71_2","unstructured":"Steffen Rendle Christoph Freudenthaler Zeno Gantner and Lars Schmidt-Thieme. 2012. BPR: Bayesian personalized ranking from implicit feedback. arXiv:1205.2618. Retrieved from https:\/\/arxiv.org\/abs\/1205.2618"},{"key":"e_1_3_2_72_2","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1016\/0377-0427(87)90125-7","article-title":"Silhouettes: A graphical aid to the interpretation and validation of cluster analysis","volume":"20","author":"Rousseeuw Peter J.","year":"1987","unstructured":"Peter J. Rousseeuw. 1987. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics 20 (1987), 53\u201365.","journal-title":"Journal of Computational and Applied Mathematics"},{"key":"e_1_3_2_73_2","doi-asserted-by":"crossref","first-page":"285","DOI":"10.1145\/371920.372071","volume-title":"Proceedings of the 10th International Conference on World Wide Web","author":"Sarwar Badrul","year":"2001","unstructured":"Badrul Sarwar, George Karypis, Joseph Konstan, and John Riedl. 2001. Item-based collaborative filtering recommendation algorithms. In Proceedings of the 10th International Conference on World Wide Web, 285\u2013295."},{"issue":"3","key":"e_1_3_2_74_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3446374","article-title":"Generative adversarial networks (GANs) challenges, solutions, and future directions","volume":"54","author":"Saxena Divya","year":"2021","unstructured":"Divya Saxena and Jiannong Cao. 2021. Generative adversarial networks (GANs) challenges, solutions, and future directions. ACM Computing Surveys 54, 3 (2021), 1\u201342.","journal-title":"ACM Computing Surveys"},{"key":"e_1_3_2_75_2","first-page":"1","volume-title":"Proceedings of the 3rd International Conference on Learning Representations (ICLR \u201915)","author":"Simonyan K.","year":"2015","unstructured":"K. Simonyan and A. Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In Proceedings of the 3rd International Conference on Learning Representations (ICLR \u201915), 1\u201314."},{"key":"e_1_3_2_76_2","first-page":"2141","article-title":"Improved multimodal deep learning with variation of information","volume":"27","author":"Sohn Kihyuk","year":"2014","unstructured":"Kihyuk Sohn, Wenling Shang, and Honglak Lee. 2014. Improved multimodal deep learning with variation of information. Advances in Neural Information Processing Systems, 27 (2014), 2141\u20132149.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_77_2","first-page":"2222","article-title":"Multimodal learning with deep boltzmann machines","volume":"25","author":"Srivastava Nitish","year":"2012","unstructured":"Nitish Srivastava and Russ R. Salakhutdinov. 2012. Multimodal learning with deep boltzmann machines. Advances in Neural Information Processing Systems 25 (2012), 2222\u20132230.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3187556"},{"key":"e_1_3_2_79_2","first-page":"1","volume-title":"Proceedings of the 2015 IEEE Information Theory Workshop (ITW)","author":"Tishby Naftali","year":"2015","unstructured":"Naftali Tishby and Noga Zaslavsky. 2015. Deep learning and the information bottleneck principle. In Proceedings of the 2015 IEEE Information Theory Workshop (ITW). IEEE, 1\u20135."},{"key":"e_1_3_2_80_2","first-page":"4642","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Veit Andreas","year":"2015","unstructured":"Andreas Veit, Balazs Kovacs, Sean Bell, Julian McAuley, Kavita Bala, and Serge Belongie. 2015. Learning visual clothing style with heterogeneous dyadic co-Occurrences. In Proceedings of the IEEE International Conference on Computer Vision, 4642\u20134650."},{"key":"e_1_3_2_81_2","unstructured":"Petar Veli\u010dkovi\u0107 Guillem Cucurull Arantxa Casanova Adriana Romero Pietro Lio and Yoshua Bengio. 2017. Graph attention networks. arXiv:1710.10903. Retrieved from https:\/\/arxiv.org\/abs\/1710.10903"},{"key":"e_1_3_2_82_2","unstructured":"Kaiye Wang Qiyue Yin Wei Wang Shu Wu and Liang Wang. 2016. A comprehensive survey on cross-modal retrieval. arXiv:1607.06215. Retrieved from https:\/\/arxiv.org\/abs\/1607.06215"},{"key":"e_1_3_2_83_2","doi-asserted-by":"crossref","first-page":"1074","DOI":"10.1109\/TMM.2021.3138298","article-title":"Dualgnn: Dual graph neural network for multimedia recommendation","volume":"25","author":"Wang Qifan","year":"2021","unstructured":"Qifan Wang, Yinwei Wei, Jianhua Yin, Jianlong Wu, Xuemeng Song, and Liqiang Nie. 2021. Dualgnn: Dual graph neural network for multimedia recommendation. IEEE Transactions on Multimedia 25 (2021), 1074\u20131084.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_84_2","unstructured":"Xin Wang Hong Chen Si\u2019ao Tang Zihao Wu and Wenwu Zhu. 2022. Disentangled representation learning. arXiv:2211.11695. Retrieved from https:\/\/arxiv.org\/abs\/2211.11695"},{"key":"e_1_3_2_85_2","first-page":"1","volume-title":"Proceedings of the 2021 IEEE International Conference on Multimedia and Expo (ICME)","author":"Wang Xin","year":"2021","unstructured":"Xin Wang, Hong Chen, and Wenwu Zhu. 2021. Multimodal disentangled representation for recommendation. In Proceedings of the 2021 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1\u20136."},{"key":"e_1_3_2_86_2","doi-asserted-by":"crossref","first-page":"1154","DOI":"10.1145\/3477495.3532012","volume-title":"Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Wang Zhaobo","year":"2022","unstructured":"Zhaobo Wang, Yanmin Zhu, Haobing Liu, and Chunyang Wang. 2022. Learning graph-based disentangled representations for next POI recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, 1154\u20131163."},{"issue":"1","key":"e_1_3_2_87_2","doi-asserted-by":"crossref","first-page":"66","DOI":"10.1147\/rd.41.0066","article-title":"Information theoretical analysis of multivariate correlation","volume":"4","author":"Watanabe Satosi","year":"1960","unstructured":"Satosi Watanabe. 1960. Information theoretical analysis of multivariate correlation. IBM Journal of Research and Development 4, 1 (1960), 66\u201382.","journal-title":"IBM Journal of Research and Development"},{"key":"e_1_3_2_88_2","first-page":"5382","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"Wei Yinwei","year":"2021","unstructured":"Yinwei Wei, Xiang Wang, Qi Li, Liqiang Nie, Yan Li, Xuanping Li, and Tat-Seng Chua. 2021. Contrastive learning for cold-start recommendation. In Proceedings of the 29th ACM International Conference on Multimedia, 5382\u20135390."},{"key":"e_1_3_2_89_2","first-page":"3541","volume-title":"Proceedings of the 28th ACM International Conference on Multimedia","author":"Wei Yinwei","year":"2020","unstructured":"Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, and Tat-Seng Chua. 2020. Graph-refined convolutional network for multimedia recommendation with implicit feedback. In Proceedings of the 28th ACM International Conference on Multimedia, 3541\u20133549."},{"key":"e_1_3_2_90_2","first-page":"1437","volume-title":"Proceedings of the 27th ACM International Conference on Multimedia","author":"Wei Yinwei","year":"2019","unstructured":"Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, Richang Hong, and Tat-Seng Chua. 2019. MMGCN: Multi-modal graph convolution network for personalized recommendation of micro-Video. In Proceedings of the 27th ACM International Conference on Multimedia, 1437\u20131445."},{"key":"e_1_3_2_91_2","first-page":"4570","volume-title":"Proceedings of the 31st ACM International Conference on Information & Knowledge Management","author":"Wu Jiahao","year":"2022","unstructured":"Jiahao Wu, Wenqi Fan, Jingfan Chen, Shengcai Liu, Qing Li, and Ke Tang. 2022. Disentangled contrastive learning for social recommendation. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, 4570\u20134574."},{"issue":"5","key":"e_1_3_2_92_2","first-page":"1","article-title":"Graph neural networks in recommender systems: A survey","volume":"55","author":"Wu Shiwen","year":"2022","unstructured":"Shiwen Wu, Fei Sun, Wentao Zhang, Xu Xie, and Bin Cui. 2022. Graph neural networks in recommender systems: A survey. Computing Surveys 55, 5 (2022), 1\u201337.","journal-title":"Computing Surveys"},{"issue":"1","key":"e_1_3_2_93_2","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1109\/TNNLS.2020.2978386","article-title":"A comprehensive survey on graph neural networks","volume":"32","author":"Wu Zonghan","year":"2021","unstructured":"Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and S. Yu Philip. 2021. A comprehensive survey on graph neural networks. IEEE Transactions on Neural Networks and Learning Systems 32, 1 (2021), 4\u201324.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_2_94_2","first-page":"1259","volume-title":"Proceedings of the 2022 IEEE 38th International Conference on Data Engineering (ICDE)","author":"Xie Xu","year":"2022","unstructured":"Xu Xie, Fei Sun, Zhaoyang Liu, Shiwen Wu, Jinyang Gao, Jiandong Zhang, Bolin Ding, and Bin Cui. 2022. Contrastive learning for sequential recommendation. In Proceedings of the 2022 IEEE 38th International Conference on Data Engineering (ICDE). IEEE, 1259\u20131273."},{"key":"e_1_3_2_95_2","first-page":"9593","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yang Mengyue","year":"2021","unstructured":"Mengyue Yang, Furui Liu, Zhitang Chen, Xinwei Shen, Jianye Hao, and Wang Jun. 2021. Causalvae: Disentangled representation learning via neural structural causal models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 9593\u20139602."},{"key":"e_1_3_2_96_2","first-page":"1294","volume-title":"Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Yu Junliang","year":"2022","unstructured":"Junliang Yu, Hongzhi Yin, Xin Xia, Tong Chen, Lizhen Cui, and Quoc Viet Hung Nguyen. 2022. Are graph augmentations necessary? simple graph contrastive learning for recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, 1294\u20131303."},{"key":"e_1_3_2_97_2","first-page":"6576","volume-title":"Proceedings of the 31st ACM International Conference on Multimedia","author":"Yu Penghang","year":"2023","unstructured":"Penghang Yu, Zhiyi Tan, Guanming Lu, and Bing-Kun Bao. 2023. Multi-view graph convolutional network for multimedia recommendation. In Proceedings of the 31st ACM International Conference on Multimedia, 6576\u20136585."},{"issue":"8","key":"e_1_3_2_98_2","doi-asserted-by":"crossref","first-page":"2008","DOI":"10.1109\/TPAMI.2018.2889774","article-title":"Advances in variational inference","volume":"41","author":"Zhang Cheng","year":"2018","unstructured":"Cheng Zhang, Judith B\u00fctepage, Hedvig Kjellstr\u00f6m, and Stephan Mandt. 2018. Advances in variational inference. IEEE Transactions on Pattern Analysis and Machine Intelligence 41, 8 (2018), 2008\u20132026.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_99_2","doi-asserted-by":"crossref","first-page":"325","DOI":"10.1145\/2911451.2911502","volume-title":"Proceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Zhang Hanwang","year":"2016","unstructured":"Hanwang Zhang, Fumin Shen, Wei Liu, Xiangnan He, Huanbo Luan, and Tat-Seng Chua. 2016. Discrete collaborative filtering. In Proceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval, 325\u2013334."},{"key":"e_1_3_2_100_2","doi-asserted-by":"crossref","first-page":"3872","DOI":"10.1145\/3474085.3475259","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"Zhang Jinghao","year":"2021","unstructured":"Jinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu, Shuhui Wang, and Liang Wang. 2021. Mining latent structures for multimedia recommendation. In Proceedings of the 29th ACM International Conference on Multimedia, 3872\u20133880."},{"key":"e_1_3_2_101_2","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia (MM \u201921)","author":"Zhang Jinghao","year":"2021","unstructured":"Jinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu, Shuhui Wang, and Liang Wang. 2021. Mining latent structures for multimedia recommendation. In Proceedings of the 29th ACM International Conference on Multimedia (MM \u201921). ACM. DOI: 10.1145\/3474085.3475259"},{"issue":"1","key":"e_1_3_2_102_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3285029","article-title":"Deep learning based recommender system: A survey and new perspectives","volume":"52","author":"Zhang Shuai","year":"2019","unstructured":"Shuai Zhang, Lina Yao, Aixin Sun, and Yi Tay. 2019. Deep learning based recommender system: A survey and new perspectives. ACM Computing Surveys (CSUR) 52, 1 (2019), 1\u201338.","journal-title":"ACM Computing Surveys (CSUR)"},{"key":"e_1_3_2_103_2","first-page":"490","volume-title":"Proceedings of the 16th ACM International Conference on Web Search and Data Mining","author":"Zhang Xiaoying","year":"2023","unstructured":"Xiaoying Zhang, Hongning Wang, and Hang Li. 2023. Disentangled representation for diversified recommendations. In Proceedings of the 16th ACM International Conference on Web Search and Data Mining, 490\u2013498."},{"issue":"8","key":"e_1_3_2_104_2","doi-asserted-by":"crossref","first-page":"8454","DOI":"10.1609\/aaai.v38i8.28688","article-title":"LGMRec: Local and global graph learning for multimodal recommendation","volume":"38","author":"Li Guohui","year":"2024","unstructured":"Guohui Li, Chaoyang Wang, Si Shi, Bin Ruan, Zhiqiang Guo, and Jianjun Li. 2024. LGMRec: Local and global graph learning for multimodal recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence 38, 8 (2024), 8454\u20138462.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_2_105_2","unstructured":"Jie Zhou Ganqu Cui Zhengyan Zhang Cheng Yang Zhiyuan Liu and Maosong Sun. 2018. Graph neural networks: A review of methods and applications. arXiv:1812.08434. Retrieved from https:\/\/arxiv.org\/abs\/1812.08434"},{"key":"e_1_3_2_106_2","doi-asserted-by":"crossref","first-page":"935","DOI":"10.1145\/3581783.3611943","volume-title":"Proceedings of the 31st ACM International Conference on Multimedia","author":"Zhou Xin","year":"2023","unstructured":"Xin Zhou, and Zhiqi Shen. 2023. A tale of two graphs: Freezing and denoising graph structures for multimodal recommendation. In Proceedings of the 31st ACM International Conference on Multimedia, 935\u2013943."},{"key":"e_1_3_2_107_2","volume-title":"Proceedings of the ACM Web Conference 2023 (WWW \u201923)","author":"Zhou Xin","year":"2023","unstructured":"Xin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng, Chunyan Miao, Pengwei Wang, Yuan You, and Feijun Jiang.2023. Bootstrap latent representations for multi-modal recommendation. In Proceedings of the ACM Web Conference 2023 (WWW \u201923). ACM. DOI: 10.1145\/3543507.3583251"},{"key":"e_1_3_2_108_2","volume-title":"Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Zhou Yan","year":"2023","unstructured":"Yan Zhou, Jie Guo, Hao Sun, Bin Song, and Fei Richard Yu. 2023. Attention-guided multi-step fusion: A hierarchical fusion network for multimodal recommendation. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. Retrieved from https:\/\/Api.semanticscholar.org\/CorpusID:258298765."},{"key":"e_1_3_2_109_2","unstructured":"Tiangang Zhu Yue Wang Haoran Li Youzheng Wu Xiaodong He and Bowen Zhou. 2020. Multimodal joint attribute prediction and value extraction for e-commerce product. arXiv:2009.07162. Retrieved from https:\/\/arxiv.org\/abs\/2009.07162"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715876","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3715876","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:48Z","timestamp":1750295928000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715876"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,15]]},"references-count":108,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,5,31]]}},"alternative-id":["10.1145\/3715876"],"URL":"https:\/\/doi.org\/10.1145\/3715876","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,15]]},"assertion":[{"value":"2024-08-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-23","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}