{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T11:14:13Z","timestamp":1784805253256,"version":"3.55.0"},"reference-count":72,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,10,25]],"date-time":"2024-10-25T00:00:00Z","timestamp":1729814400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Alibaba Group through Alibaba Innovative Research (AIR) Program"},{"name":"Alibaba-NTU Singapore Joint Research Institute"},{"DOI":"10.13039\/501100001475","name":"Nanyang Technological University, Singapore","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001475","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Recomm. Syst."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>Sequential recommendation systems often suffer from data sparsity, leading to suboptimal performance. While multimodal content, such as images and text, has been utilized to mitigate this issue, its integration within sequential recommendation frameworks remains challenging. Current multimodal sequential recommendation models are often unable to effectively explore and capture correlations among behavior sequences of users and items across different modalities, either neglecting correlations among sequence representations or inadequately capturing associations between multimodal data and sequence data in their representations. To address this problem, we explore multimodal pre-training in the context of sequential recommendation, with the aim of enhancing fusion and utilization of multimodal information.<\/jats:p>\n          <jats:p>\n            We propose a novel Multimodal Pre-training for Sequential Recommendation (MP4SR) framework, which utilizes contrastive losses to capture the correlation among different modality sequences of users, as well as the correlation among different modality sequences of users and items. MP4SR consists of three key components: (1) multimodal feature extraction; (2) a backbone network, Multimodal Mixup Sequence Encoder (M\n            <jats:sup>2<\/jats:sup>\n            SE); and (3) pre-training tasks. After utilizing pre-trained encoders to generate initial multimodal features of items, M\n            <jats:sup>2<\/jats:sup>\n            SE adopts a complementary sequence mixup strategy to fuse different modality sequences, and leverages contrastive learning to capture modality interactions at the sequence-to-sequence and sequence-to-item levels. Extensive experiments on four real-world datasets demonstrate that MP4SR outperforms state-of-the-art approaches in both normal and cold-start settings. We further highlight the efficacy of incorporating multimodal pre-training in sequential recommendation representation learning, serving as an effective regularizer and optimizing the parameter space for the recommendation task.\n          <\/jats:p>","DOI":"10.1145\/3682075","type":"journal-article","created":{"date-parts":[[2024,7,29]],"date-time":"2024-07-29T11:10:02Z","timestamp":1722251402000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":27,"title":["Multimodal Pre-training for Sequential Recommendation via Contrastive Learning"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5319-875X","authenticated-orcid":false,"given":"Lingzi","family":"Zhang","sequence":"first","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0948-8033","authenticated-orcid":false,"given":"Xin","family":"Zhou","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7787-5644","authenticated-orcid":false,"given":"Zhiwei","family":"Zeng","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7626-7295","authenticated-orcid":false,"given":"Zhiqi","family":"Shen","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,10,25]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Hangbo Bao Wenhui Wang Li Dong Qiang Liu Owais Khan Mohammed Kriti Aggarwal Subhojit Som Songhao Piao and Furu Wei. 2022. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. Advances in Neural Information Processing Systems (2022) 32897\u201332912."},{"key":"e_1_3_1_3_2","doi-asserted-by":"crossref","first-page":"378","DOI":"10.1145\/3404835.3462968","volume-title":"Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Chang Jianxin","year":"2021","unstructured":"Jianxin Chang, Chen Gao, Yu Zheng, Yiqun Hui, Yanan Niu, Yang Song, Depeng Jin, and Yong Li. 2021. Sequential recommendation with graph neural networks. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 378\u2013387."},{"key":"e_1_3_1_4_2","first-page":"51","volume-title":"Proceedings of the IEEE International Conference on Data Mining (ICDM\u201921)","author":"Cheng Mingyue","year":"2021","unstructured":"Mingyue Cheng, Fajie Yuan, Qi Liu, Xin Xin, and Enhong Chen. 2021. Learning transferable user representations with sequential behaviors via contrastive pre-training. In Proceedings of the IEEE International Conference on Data Mining (ICDM\u201921). IEEE, 51\u201360."},{"issue":"2","key":"e_1_3_1_5_2","first-page":"317","article-title":"MV-RNN: A multi-view recurrent neural network for sequential recommendation","volume":"32","author":"Cui Qiang","year":"2018","unstructured":"Qiang Cui, Shu Wu, Qiang Liu, Wen Zhong, and Liang Wang. 2018. MV-RNN: A multi-view recurrent neural network for sequential recommendation. IEEE Transactions on Knowledge and Data Engineering 32, 2 (2018), 317\u2013331.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_3_1_6_2","first-page":"152","volume-title":"Proceedings of the 11th ACM Conference on Recommender Systems","author":"Donkers Tim","year":"2017","unstructured":"Tim Donkers, Benedikt Loepp, and J\u00fcrgen Ziegler. 2017. Sequential user-based recurrent neural network recommendations. In Proceedings of the 11th ACM Conference on Recommender Systems. 152\u2013160."},{"key":"e_1_3_1_7_2","first-page":"201","volume-title":"Proceedings of the 13th International Conference on Artificial Intelligence and Statistics","author":"Erhan Dumitru","year":"2010","unstructured":"Dumitru Erhan, Aaron Courville, Yoshua Bengio, and Pascal Vincent. 2010. Why does unsupervised pre-training help deep learning? In Proceedings of the 13th International Conference on Artificial Intelligence and Statistics. 201\u2013208."},{"key":"e_1_3_1_8_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"He Ruining","year":"2016","unstructured":"Ruining He and Julian McAuley. 2016. VBPR: Visual Bayesian personalized ranking from implicit feedback. In Proceedings of the AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_1_9_2","first-page":"639","volume-title":"Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"He Xiangnan","year":"2020","unstructured":"Xiangnan He, Kuan Deng, Xiang Wang, Yan Li, Yongdong Zhang, and Meng Wang. 2020. LightGCN: Simplifying and powering graph convolution network for recommendation. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval. 639\u2013648."},{"key":"e_1_3_1_10_2","first-page":"585","volume-title":"Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Hou Yupeng","year":"2022","unstructured":"Yupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li, Bolin Ding, and Ji-Rong Wen. 2022. Towards universal sequence representation learning for recommender systems. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 585\u2013593."},{"key":"e_1_3_1_11_2","first-page":"843","volume-title":"Proceedings of the 32nd ACM International Conference on Information and Knowledge Management","author":"Hu Hengchang","year":"2023","unstructured":"Hengchang Hu, Wei Guo, Yong Liu, and Min-Yen Kan. 2023. Adaptive multi-modalities fusion in sequential recommendation systems. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management. 843\u2013853."},{"key":"e_1_3_1_12_2","first-page":"791","volume-title":"Proceedings of the 31st ACM International Conference on Information & Knowledge Management","author":"Hu Yidan","year":"2022","unstructured":"Yidan Hu, Yong Liu, Chunyan Miao, and Yuan Miao. 2022. Memory bank augmented long-tail sequential recommendation. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management. 791\u2013801."},{"key":"e_1_3_1_13_2","first-page":"4904","volume-title":"International Conference on Machine Learning","author":"Jia Chao","year":"2021","unstructured":"Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning. PMLR, 4904\u20134916."},{"key":"e_1_3_1_14_2","first-page":"207","volume-title":"Proceedings of the IEEE International Conference on Data Mining (ICDM\u201917)","author":"Kang Wang-Cheng","year":"2017","unstructured":"Wang-Cheng Kang, Chen Fang, Zhaowen Wang, and Julian McAuley. 2017. Visually-aware fashion recommendation and design with generative image models. In Proceedings of the IEEE International Conference on Data Mining (ICDM\u201917). IEEE, 207\u2013216."},{"key":"e_1_3_1_15_2","first-page":"197","volume-title":"Proceedings of the IEEE International Conference on Data Mining (ICDM\u201918)","author":"Kang Wang-Cheng","year":"2018","unstructured":"Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recommendation. In Proceedings of the IEEE International Conference on Data Mining (ICDM\u201918). IEEE, 197\u2013206."},{"key":"e_1_3_1_16_2","first-page":"5583","volume-title":"International Conference on Machine Learning","author":"Kim Wonjae","year":"2021","unstructured":"Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning. PMLR, 5583\u20135594."},{"key":"e_1_3_1_17_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_18_2","first-page":"3161","volume-title":"Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining","author":"Lei Chenyi","year":"2021","unstructured":"Chenyi Lei, Yong Liu, Lingzi Zhang, Guoxin Wang, Haihong Tang, Houqiang Li, and Chunyan Miao. 2021. Semi: A sequential multi-modal information transfer network for e-commerce micro-video recommendations. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. 3161\u20133171."},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","first-page":"1109","DOI":"10.1145\/3543507.3583378","volume-title":"Proceedings of the ACM Web Conference 2023","author":"Liang Jiahao","year":"2023","unstructured":"Jiahao Liang, Xiangyu Zhao, Muyang Li, Zijian Zhang, Wanyu Wang, Haochen Liu, and Zitao Liu. 2023. MMMLP: Multi-modal multilayer perceptron for sequential recommendations. In Proceedings of the ACM Web Conference 2023. 1109\u20131117."},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","unstructured":"Xudong Lin Simran Tiwari Shiyuan Huang Manling Li Mike Zheng Shou Heng Ji and Shih-Fu Chang. 2022. Towards fast adaptation of pretrained contrastive models for multi-channel video-language retrieval. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 14846\u201314855.","DOI":"10.1109\/CVPR52729.2023.01426"},{"key":"e_1_3_1_21_2","first-page":"1526","volume-title":"Proceedings of the 27th ACM International Conference on Multimedia","author":"Liu Fan","year":"2019","unstructured":"Fan Liu, Zhiyong Cheng, Changchang Sun, Yinglong Wang, Liqiang Nie, and Mohan Kankanhalli. 2019. User diverse preference modeling by multimodal attentive metric learning. In Proceedings of the 27th ACM International Conference on Multimedia. 1526\u20131534."},{"key":"e_1_3_1_22_2","first-page":"3020","volume-title":"Proceedings of the Web Conference 2019","author":"Liu Shang","year":"2019","unstructured":"Shang Liu, Zhenzhong Chen, Hongyi Liu, and Xinghai Hu. 2019. User-video co-attention network for personalized micro-video recommendation. In Proceedings of the Web Conference 2019. 3020\u20133026."},{"key":"e_1_3_1_23_2","first-page":"2853","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"Liu Yong","year":"2021","unstructured":"Yong Liu, Susen Yang, Chenyi Lei, Guoxin Wang, Haihong Tang, Juyong Zhang, Aixin Sun, and Chunyan Miao. 2021. Pre-training graph transformer with multimodal side information for recommendation. In Proceedings of the 29th ACM International Conference on Multimedia. 2853\u20132861."},{"key":"e_1_3_1_24_2","first-page":"1608","volume-title":"Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Liu Zhiwei","year":"2021","unstructured":"Zhiwei Liu, Ziwei Fan, Yu Wang, and Philip S Yu. 2021. Augmenting sequential recommendation with pseudo-prior items via reversely pre-training transformer. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1608\u20131612."},{"key":"e_1_3_1_25_2","article-title":"Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks","author":"Lu Jiasen","year":"2019","unstructured":"Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_26_2","first-page":"188","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP\u201919)","author":"Ni Jianmo","year":"2019","unstructured":"Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019. Justifying recommendations using distantly-labeled reviews and fine-grained aspects. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP\u201919). 188\u2013197."},{"key":"e_1_3_1_27_2","first-page":"3421","volume-title":"Proceedings of the 31st ACM International Conference on Information & Knowledge Management","author":"Pan Xingyu","year":"2022","unstructured":"Xingyu Pan, Yushuo Chen, Changxin Tian, Zihan Lin, Jinpeng Wang, He Hu, and Wayne Xin Zhao. 2022. Multimodal meta-learning for cold-start sequential recommendation. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management. 3421\u20133430."},{"key":"e_1_3_1_28_2","first-page":"8026","article-title":"PyTorch: An imperative style, high-performance deep learning library","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et\u00a0al. 2019. PyTorch: An imperative style, high-performance deep learning library. In Proceedings of the Neural Information Processing Systems, 8026\u20138037.","journal-title":"Proceedings of the Neural Information Processing Systems"},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","unstructured":"Bo Peng Zhiyun Ren Srinivasan Parthasarathy and Xia Ning. 2021. HAM: Hybrid associations models for sequential recommendation. IEEE Transactions on Knowledge and Data Engineering 34 10 (2021) 4838\u20134853.","DOI":"10.1109\/TKDE.2021.3049692"},{"key":"e_1_3_1_30_2","first-page":"8748","volume-title":"International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et\u00a0al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning. 8748\u20138763."},{"key":"e_1_3_1_31_2","first-page":"3982","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP\u201919)","author":"Reimers Nils","year":"2019","unstructured":"Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using siamese BERT-networks. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP\u201919). 3982\u20133992."},{"key":"e_1_3_1_32_2","first-page":"811","volume-title":"Proceedings of the Web Conference 2010","author":"Rendle Steffen","year":"2010","unstructured":"Steffen Rendle, Christoph Freudenthaler, and Lars Schmidt-Thieme. 2010. Factorizing personalized markov chains for next-basket recommendation. In Proceedings of the Web Conference 2010. 811\u2013820."},{"key":"e_1_3_1_33_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Shazeer Noam","year":"2017","unstructured":"Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_34_2","doi-asserted-by":"crossref","first-page":"1283","DOI":"10.1145\/3477495.3531927","volume-title":"Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Shuai Jie","year":"2022","unstructured":"Jie Shuai, Kun Zhang, Le Wu, Peijie Sun, Richang Hong, Meng Wang, and Yong Li. 2022. A review-aware graph contrastive learning framework for recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1283\u20131293."},{"key":"e_1_3_1_35_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Su Weijie","year":"2020","unstructured":"Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020. Vl-bert: Pre-training of generic visual-linguistic representations. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_36_2","first-page":"1441","volume-title":"Proceedings of the 28th ACM International Conference on Information & Knowledge Management","author":"Sun Fei","year":"2019","unstructured":"Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang. 2019. BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer. In Proceedings of the 28th ACM International Conference on Information & Knowledge Management. 1441\u20131450."},{"key":"e_1_3_1_37_2","first-page":"598","volume-title":"Proceedings of the 14th ACM International Conference on Web Search and Data Mining","author":"Tan Qiaoyu","year":"2021","unstructured":"Qiaoyu Tan, Jianwei Zhang, Jiangchao Yao, Ninghao Liu, Jingren Zhou, Hongxia Yang, and Xia Hu. 2021. Sparse-interest network for sequential recommendation. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining. 598\u2013606."},{"key":"e_1_3_1_38_2","first-page":"565","volume-title":"Proceedings of the 11th ACM International Conference on Web Search and Data Mining","author":"Tang Jiaxi","year":"2018","unstructured":"Jiaxi Tang and Ke Wang. 2018. Personalized Top-N sequential recommendation via convolutional sequence embedding. In Proceedings of the 11th ACM International Conference on Web Search and Data Mining. 565\u2013573."},{"key":"e_1_3_1_39_2","unstructured":"Ilya O. Tolstikhin Neil Houlsby Alexander Kolesnikov Lucas Beyer Xiaohua Zhai Thomas Unterthiner Jessica Yung Andreas Steiner Daniel Keysers Jakob Uszkoreit et\u00a0al. 2021. Mlp-mixer: An all-mlp architecture for vision. Advances in Neural Information Processing Systems 34 (2021) 24261\u201324272."},{"key":"e_1_3_1_40_2","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems (2017) 5998\u20136008."},{"key":"e_1_3_1_41_2","first-page":"86","volume-title":"Proceedings of the 12th ACM Conference on Recommender Systems (RecSys\u201918)","author":"Wan Mengting","year":"2018","unstructured":"Mengting Wan and Julian J. McAuley. 2018. Item recommendation on monotonic behavior chains. In Proceedings of the 12th ACM Conference on Recommender Systems (RecSys\u201918), Sole Pera, Michael D. Ekstrand, Xavier Amatriain, and John O\u2019Donovan (Eds.). ACM, 86\u201394."},{"key":"e_1_3_1_42_2","first-page":"2605","volume-title":"Proceedings of the 57th Conference of the Association for Computational Linguistics, (ACL\u201919), Volume 1: Long Papers","author":"Wan Mengting","year":"2019","unstructured":"Mengting Wan, Rishabh Misra, Ndapa Nakashole, and Julian J. McAuley. 2019. Fine-grained spoiler detection from large-scale review corpora. In Proceedings of the 57th Conference of the Association for Computational Linguistics, (ACL\u201919), Volume 1: Long Papers, Anna Korhonen, David R. Traum, and Llu\u00eds M\u00e0rquez (Eds.). Association for Computational Linguistics, 2605\u20132610."},{"key":"e_1_3_1_43_2","unstructured":"Jie Wang Fajie Yuan Mingyue Cheng Joemon M Jose Chenyun Yu Beibei Kong Zhijin Wang Bo Hu and Zang Li. 2022. TransRec: Learning transferable recommendation from mixture-of-modality feedback. arXiv preprint arXiv:2206.06190 (2022)."},{"key":"e_1_3_1_44_2","doi-asserted-by":"crossref","unstructured":"Qifan Wang Yinwei Wei Jianhua Yin Jianlong Wu Xuemeng Song and Liqiang Nie. 2021. DualGNN: Dual graph neural network for multimedia recommendation. IEEE Transactions on Multimedia 25 (2021) 1074\u20131084.","DOI":"10.1109\/TMM.2021.3138298"},{"key":"e_1_3_1_45_2","first-page":"6332","volume-title":"Proceedings of the 28th International Joint Conference on Artificial Intelligence","author":"Wang S.","year":"2019","unstructured":"S. Wang, L. Hu, Y. Wang, L. Cao, Q. Z. Sheng, and M. Orgun. 2019. Sequential recommender systems: Challenges, progress and prospects. In Proceedings of the 28th International Joint Conference on Artificial Intelligence. 6332\u20136338."},{"key":"e_1_3_1_46_2","doi-asserted-by":"crossref","unstructured":"Wenhui Wang Hangbo Bao Li Dong Johan Bjorck Zhiliang Peng Qiang Liu Kriti Aggarwal Owais Khan Mohammed Saksham Singhal Subhojit Som et\u00a0al. 2022. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 19175\u201319186.","DOI":"10.1109\/CVPR52729.2023.01838"},{"key":"e_1_3_1_47_2","doi-asserted-by":"crossref","first-page":"2209","DOI":"10.1145\/3442381.3450038","volume-title":"Proceedings of the Web Conference 2021","author":"Wang Xi","year":"2021","unstructured":"Xi Wang, Iadh Ounis, and Craig Macdonald. 2021. Leveraging review properties for effective recommendation. In Proceedings of the Web Conference 2021. 2209\u20132219."},{"key":"e_1_3_1_48_2","unstructured":"Zirui Wang Jiahui Yu Adams Wei Yu Zihang Dai Yulia Tsvetkov and Yuan Cao. 2022. Simvlm: Simple visual language model pretraining with weak supervision. In International Conference on Learning Representations."},{"key":"e_1_3_1_49_2","first-page":"5382","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"Wei Yinwei","year":"2021","unstructured":"Yinwei Wei, Xiang Wang, Qi Li, Liqiang Nie, Yan Li, Xuanping Li, and Tat-Seng Chua. 2021. Contrastive learning for cold-start recommendation. In Proceedings of the 29th ACM International Conference on Multimedia. 5382\u20135390."},{"key":"e_1_3_1_50_2","first-page":"3541","volume-title":"Proceedings of the 28th ACM International Conference on Multimedia","author":"Wei Yinwei","year":"2020","unstructured":"Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, and Tat-Seng Chua. 2020. Graph-refined convolutional network for multimedia recommendation with implicit feedback. In Proceedings of the 28th ACM International Conference on Multimedia. 3541\u20133549."},{"key":"e_1_3_1_51_2","first-page":"1437","volume-title":"Proceedings of the 27th ACM International Conference on Multimedia","author":"Wei Yinwei","year":"2019","unstructured":"Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, Richang Hong, and Tat-Seng Chua. 2019. MMGCN: Multi-modal graph convolution network for personalized recommendation of micro-video. In Proceedings of the 27th ACM International Conference on Multimedia. 1437\u20131445."},{"key":"e_1_3_1_52_2","unstructured":"Yiqing Wu Ruobing Xie Yongchun Zhu Fuzhen Zhuang Xu Zhang Leyu Lin and Qing He. 2022. Personalized prompts for sequential recommendation. IEEE Transactions on Knowledge and Data Engineering (2024)."},{"key":"e_1_3_1_53_2","first-page":"1259","volume-title":"Proceedings of the IEEE 38th International Conference on Data Engineering (ICDE\u201922)","author":"Xie Xu","year":"2022","unstructured":"Xu Xie, Fei Sun, Zhaoyang Liu, Shiwen Wu, Jinyang Gao, Jiandong Zhang, Bolin Ding, and Bin Cui. 2022. Contrastive learning for sequential recommendation. In Proceedings of the IEEE 38th International Conference on Data Engineering (ICDE\u201922). 1259\u20131273."},{"key":"e_1_3_1_54_2","first-page":"1611","volume-title":"Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Xie Yueqi","year":"2022","unstructured":"Yueqi Xie, Peilin Zhou, and Sunghun Kim. 2022. Decoupled side information fusion for sequential recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1611\u20131621."},{"key":"e_1_3_1_55_2","first-page":"981","volume-title":"Proceedings of the 41st International ACM SIGIR Conference on Research & Development in Information Retrieval","author":"Xu Qidi","year":"2018","unstructured":"Qidi Xu, Fumin Shen, Li Liu, and Heng Tao Shen. 2018. Graphcar: Content-aware multimedia recommendation with graph autoencoder. In Proceedings of the 41st International ACM SIGIR Conference on Research & Development in Information Retrieval. 981\u2013984."},{"key":"e_1_3_1_56_2","first-page":"1067","article-title":"Multi-modal variational graph auto-encoder for recommendation systems","volume":"24","author":"Yi Jing","year":"2021","unstructured":"Jing Yi and Zhenzhong Chen. 2021. Multi-modal variational graph auto-encoder for recommendation systems. IEEE Trans. Multimedia 24 (2021), 1067\u20131079.","journal-title":"IEEE Trans. Multimedia"},{"key":"e_1_3_1_57_2","first-page":"582","volume-title":"Proceedings of the 12th ACM International Conference on Web Search and Data Mining","author":"Yuan Fajie","year":"2019","unstructured":"Fajie Yuan, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. Jose, and Xiangnan He. 2019. A simple convolutional generative network for next item recommendation. In Proceedings of the 12th ACM International Conference on Web Search and Data Mining. 582\u2013590."},{"key":"e_1_3_1_58_2","doi-asserted-by":"crossref","first-page":"2639","DOI":"10.1145\/3539618.3591932","volume-title":"Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Yuan Zheng","year":"2023","unstructured":"Zheng Yuan, Fajie Yuan, Yu Song, Youhua Li, Junchen Fu, Fei Yang, Yunzhu Pan, and Yongxin Ni. 2023. Where to go next for recommender systems? id-vs. modality-based recommender models revisited. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2639\u20132649."},{"key":"e_1_3_1_59_2","doi-asserted-by":"crossref","first-page":"3872","DOI":"10.1145\/3474085.3475259","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"Zhang Jinghao","year":"2021","unstructured":"Jinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu, Shuhui Wang, and Liang Wang. 2021. Mining latent structures for multimedia recommendation. In Proceedings of the 29th ACM International Conference on Multimedia. 3872\u20133880."},{"key":"e_1_3_1_60_2","doi-asserted-by":"crossref","first-page":"232","DOI":"10.1007\/978-3-031-00126-0_15","volume-title":"International Conference on Database Systems for Advanced Applications","author":"Zhang Lingzi","year":"2022","unstructured":"Lingzi Zhang, Yong Liu, Xin Zhou, Chunyan Miao, Guoxin Wang, and Haihong Tang. 2022. Diffusion-based graph contrastive learning for recommendation with implicit feedback. In International Conference on Database Systems for Advanced Applications. Springer, 232\u2013247."},{"key":"e_1_3_1_61_2","doi-asserted-by":"crossref","first-page":"625","DOI":"10.1145\/3589335.3651516","volume-title":"Companion Proceedings of the ACM on Web Conference 2024","author":"Zhang Lingzi","year":"2024","unstructured":"Lingzi Zhang, Yinan Zhang, Xin Zhou, and Zhiqi Shen. 2024. GreenRec: A large-scale dataset for green food recommendation. In Companion Proceedings of the ACM on Web Conference 2024. 625\u2013628."},{"key":"e_1_3_1_62_2","volume-title":"Proceedings of the IEEE 40th International Conference on Data Engineering (ICDE\u201924)","author":"Zhang Lingzi","year":"2024","unstructured":"Lingzi Zhang, Xin Zhou, Zhiwei Zeng, and Zhiqi Shen. 2024. Are ID embeddings necessary? Whitening pre-trained text embeddings for effective sequential recommendation. In Proceedings of the IEEE 40th International Conference on Data Engineering (ICDE\u201924)."},{"key":"e_1_3_1_63_2","first-page":"9332","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"38","author":"Zhang Lingzi","year":"2024","unstructured":"Lingzi Zhang, Xin Zhou, Zhiwei Zeng, and Zhiqi Shen. 2024. Dual-view whitening on pre-trained text embeddings for sequential recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 9332\u20139340."},{"key":"e_1_3_1_64_2","first-page":"4320","volume-title":"Proceedings of the 28th International Joint Conference on Artificial Intelligence","author":"Zhang Tingting","year":"2019","unstructured":"Tingting Zhang, Pengpeng Zhao, Yanchi Liu, Victor S. Sheng, Jiajie Xu, Deqing Wang, Guanfeng Liu, and Xiaofang Zhou. 2019. Feature-level deeper self-attention network for sequential recommendation. In Proceedings of the 28th International Joint Conference on Artificial Intelligence. 4320\u20134326."},{"key":"e_1_3_1_65_2","first-page":"2398","volume-title":"Proceedings of the 31st International Joint Conference on Artificial Intelligence","author":"Zhang Yixin","year":"2022","unstructured":"Yixin Zhang, Yong Liu, Yonghui Xu, Hao Xiong, Chenyi Lei, Wei He, Lizhen Cui, and Chunyan Miao. 2022. Enhancing sequential recommendation with graph contrastive learning. In Proceedings of the 31st International Joint Conference on Artificial Intelligence. 2398\u20132405."},{"key":"e_1_3_1_66_2","doi-asserted-by":"crossref","first-page":"4653","DOI":"10.1145\/3459637.3482016","volume-title":"Proceedings of the 30th ACM International Conference on Information & Knowledge Management","author":"Zhao Wayne Xin","year":"2021","unstructured":"Wayne Xin Zhao, Shanlei Mu, Yupeng Hou, Zihan Lin, Yushuo Chen, Xingyu Pan, Kaiyuan Li, Yujie Lu, Hui Wang, Changxin Tian, et\u00a0al. 2021. RecBole: Towards a unified, comprehensive and efficient framework for recommendation algorithms. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management. 4653\u20134664."},{"key":"e_1_3_1_67_2","unstructured":"Hongyu Zhou Xin Zhou Zhiwei Zeng Lingzi Zhang and Zhiqi Shen. 2023. A comprehensive survey on multimodal recommender systems: Taxonomy evaluation and future directions. arXiv preprint arXiv:2302.04473 (2023)."},{"key":"e_1_3_1_68_2","first-page":"3123","volume-title":"Proceedings of the European Conference on Artificial Intelligence (ECAI\u201923)","author":"Zhou Hongyu","year":"2023","unstructured":"Hongyu Zhou, Xin Zhou, Lingzi Zhang, and Zhiqi Shen. 2023. Enhancing dyadic relations with homogeneous graphs for multimodal recommendation. In Proceedings of the European Conference on Artificial Intelligence (ECAI\u201923). IOS Press, 3123\u20133130."},{"key":"e_1_3_1_69_2","doi-asserted-by":"crossref","first-page":"1893","DOI":"10.1145\/3340531.3411954","volume-title":"Proceedings of the 29th ACM International Conference on Information & Knowledge Management","author":"Zhou Kun","year":"2020","unstructured":"Kun Zhou, Hui Wang, Wayne Xin Zhao, Yutao Zhu, Sirui Wang, Fuzheng Zhang, Zhongyuan Wang, and Ji-Rong Wen. 2020. S3-Rec: Self-supervised learning for sequential recommendation with mutual information maximization. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management. 1893\u20131902."},{"key":"e_1_3_1_70_2","first-page":"13041","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Zhou Luowei","year":"2020","unstructured":"Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao. 2020. Unified vision-language pre-training for image captioning and vqa. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 13041\u201313049."},{"key":"e_1_3_1_71_2","doi-asserted-by":"crossref","first-page":"935","DOI":"10.1145\/3581783.3611943","volume-title":"Proceedings of the 31st ACM International Conference on Multimedia","author":"Zhou Xin","year":"2023","unstructured":"Xin Zhou and Zhiqi Shen. 2023. A tale of two graphs: Freezing and denoising graph structures for multimodal recommendation. In Proceedings of the 31st ACM International Conference on Multimedia. 935\u2013943."},{"key":"e_1_3_1_72_2","doi-asserted-by":"crossref","first-page":"845","DOI":"10.1145\/3543507.3583251","volume-title":"Proceedings of the ACM Web Conference 2023","author":"Zhou Xin","year":"2023","unstructured":"Xin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng, Chunyan Miao, Pengwei Wang, Yuan You, and Feijun Jiang. 2023. Bootstrap latent representations for multi-modal recommendation. In Proceedings of the ACM Web Conference 2023. 845\u2013854."},{"key":"e_1_3_1_73_2","first-page":"2069","volume-title":"Proceedings of the Web Conference 2021","author":"Zhu Yanqiao","year":"2021","unstructured":"Yanqiao Zhu, Yichen Xu, Feng Yu, Qiang Liu, Shu Wu, and Liang Wang. 2021. Graph contrastive learning with adaptive augmentation. In Proceedings of the Web Conference 2021. 2069\u20132080."}],"container-title":["ACM Transactions on Recommender Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3682075","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3682075","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:10:03Z","timestamp":1750295403000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3682075"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,25]]},"references-count":72,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3682075"],"URL":"https:\/\/doi.org\/10.1145\/3682075","relation":{},"ISSN":["2770-6699"],"issn-type":[{"value":"2770-6699","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,25]]},"assertion":[{"value":"2023-07-09","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-07-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}