{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T14:48:21Z","timestamp":1787496501407,"version":"build-2736575974"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2025,4,10]],"date-time":"2025-04-10T00:00:00Z","timestamp":1744243200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Recomm. Syst."],"published-print":{"date-parts":[[2025,12,31]]},"abstract":"<jats:p>In recent years, the rapid growth of online multimedia services, such as e-commerce platforms, has necessitated the development of personalised recommendation approaches that can encode diverse content about each item. Indeed, modern multi-modal recommender systems exploit diverse features obtained from raw images and item descriptions to enhance the recommendation performance. However, the existing multi-modal recommender systems primarily depend on the features extracted individually from different media through pre-trained modality-specific encoders, and exhibit only shallow alignments between different modalities, thereby limiting these systems\u2019 ability to capture the underlying relationships between the modalities. In this article, we enhance the deep alignment of large multi-modal encoders to address the shallow alignment of modalities in multi-modal recommender systems. These encoders have previously demonstrated state-of-the-art effectiveness in ranking items across various domains. Specifically, we investigate the use of three state-of-the-art large multi-modal encoders \u2013 CLIP (dual-stream), VLMo and BEiT-3 (unified) \u2013 for recommendation tasks. We explore their benefits for recommendation through using a range of strategies, including the use of pre-trained and fine-tuned encoders, as well as the evaluation of the end-to-end training of these encoders. We show that pre-trained large multi-modal encoders generate more aligned and effective user\/item representations compared with existing modality-specific encoders across four existing multi-modal recommendation datasets. Furthermore, we show that fine-tuning these encoders further improves the recommendation performance, with end-to-end training emerging as the most effective paradigm, significantly outperforming both pre-trained and fine-tuned encoders with an improved recommendation performance. We also demonstrate the effectiveness of large multi-modal encoders in facilitating modality alignment by evaluating the contribution of each modality separately. Finally, we show that the dual-stream approach, specifically CLIP, is the most effective architecture for these large multi-modal encoders, outperforming the unified approaches (i.e., VLMo and BEiT3) in terms of effectiveness and efficiency.<\/jats:p>","DOI":"10.1145\/3718099","type":"journal-article","created":{"date-parts":[[2025,2,19]],"date-time":"2025-02-19T06:18:06Z","timestamp":1739945886000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Enhancing Recommender Systems: Deep Modality Alignment with Large Multi-Modal Encoders"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7752-8650","authenticated-orcid":false,"given":"Zixuan","family":"Yi","sequence":"first","affiliation":[{"name":"University of Glasgow, Glasgow, United Kingdom of Great Britain and Northern Ireland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2008-960X","authenticated-orcid":false,"given":"Zijun","family":"Long","sequence":"additional","affiliation":[{"name":"University of Glasgow, Glasgow, United Kingdom of Great Britain and Northern Ireland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4701-3223","authenticated-orcid":false,"given":"Iadh","family":"Ounis","sequence":"additional","affiliation":[{"name":"school of computing science, University of Glasgow, Glasgow, United Kingdom of Great Britain and Northern Ireland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3143-279X","authenticated-orcid":false,"given":"Craig","family":"Macdonald","sequence":"additional","affiliation":[{"name":"school of computing science, University of Glasgow, Glasgow, United Kingdom of Great Britain and Northern Ireland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2751-2087","authenticated-orcid":false,"given":"Richard","family":"McCreadie","sequence":"additional","affiliation":[{"name":"University of Glasgow, Glasgow, United Kingdom of Great Britain and Northern Ireland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,4,10]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"Proc. of NeurIPS","author":"Bao Hangbo","year":"2022","unstructured":"Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Songhao Piao, and Furu Wei. 2022. VLMo: Unified vision-language pre-training with mixture-of-modality-experts. In Proc. of NeurIPS. 1\u201322."},{"key":"e_1_3_2_3_2","volume-title":"Proc. of SIGIR","author":"Chen Xu","year":"2019","unstructured":"Xu Chen, Hanxiong Chen, Hongteng Xu, Yongfeng Zhang, Yixin Cao, Zheng Qin, and Hongyuan Zha. 2019. Personalised fashion recommendation with visual explanations based on multimodal attention network: Towards visually explainable recommendation. In Proc. of SIGIR. 765\u2013774."},{"key":"e_1_3_2_4_2","volume-title":"Proc. of ICML","author":"Donahue Jeff","year":"2014","unstructured":"Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell. 2014. DeCAF: A deep convolutional activation feature for generic visual recognition. In Proc. of ICML. 647\u2013656."},{"key":"e_1_3_2_5_2","article-title":"VSE++: Improving visual-semantic embeddings with hard negatives","author":"Faghri Fartash","year":"2017","unstructured":"Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2017. VSE++: Improving visual-semantic embeddings with hard negatives. arXiv preprint arXiv:1707.05612 (2017).","journal-title":"arXiv preprint arXiv:1707.05612"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952261"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v30i1.9973"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/582415.582418"},{"key":"e_1_3_2_11_2","volume-title":"Proc. of NAACL-HLT","author":"Kenton Jacob Devlin, Ming-Wei Chang,","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, KentonLee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proc. of NAACL-HLT. 4171\u20134186."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3511808.3557387"},{"key":"e_1_3_2_13_2","volume-title":"Proc. of ICLR","author":"Kingma Diederik P.","year":"2014","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. In Proc. of ICLR."},{"key":"e_1_3_2_14_2","unstructured":"Alexander Kolesnikov Alexey Dosovitskiy Dirk Weissenborn Georg Heigold Jakob Uszkoreit Lucas Beyer Matthias Minderer Mostafa Dehghani Neil Houlsby and Sylvain Gelly et\u00a0al. 2022. An image is worth \\(16\\times 16\\) words: Transformers for image recognition at scale. In Proc. of ICLR. 1\u201314."},{"key":"e_1_3_2_15_2","volume-title":"Proc. of ICML","author":"Le Quoc","year":"2014","unstructured":"Quoc Le and Tomas Mikolov. 2014. Distributed representations of sentences and documents. In Proc. of ICML. 1188\u20131196."},{"key":"e_1_3_2_16_2","volume-title":"Proc. of NeurIPS","author":"Liu Haotian","year":"2024","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. In Proc. of NeurIPS. 1\u201314."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3695461"},{"key":"e_1_3_2_18_2","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692 (2019).","journal-title":"arXiv preprint arXiv:1907.11692"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i4.28177"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3512527.3531378"},{"key":"e_1_3_2_22_2","volume-title":"Proc. of SIGIR","author":"Long Zijun","year":"2024","unstructured":"Zijun Long, Xuri Ge, Richard McCreadie, and Joemon M. Jose. 2024. CFIF: Fast and effective long-text to image retrieval for large corpora. In Proc. of SIGIR."},{"key":"e_1_3_2_23_2","volume-title":"Proc. of NeurIPS","author":"Lu Jiasen","year":"2019","unstructured":"Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. VilBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Proc. of NeurIPS. 1\u201310."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-56063-7_1"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3640457.3688186"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1018"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3511808.3557101"},{"key":"e_1_3_2_28_2","volume-title":"Proc. of ICML","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et\u00a0al. 2021. Learning transferable visual models from natural language supervision. In Proc. of ICML."},{"key":"e_1_3_2_29_2","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et\u00a0al. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1 (2019).","journal-title":"OpenAI Blog"},{"key":"e_1_3_2_30_2","volume-title":"Proc. of SIGIR","author":"Rao Jun","year":"2022","unstructured":"Jun Rao, Fei Wang, Liang Ding, Shuhan Qi, Yibing Zhan, Weifeng Liu, and Dacheng Tao. 2022. Where does the performance improvement come from? -A reproducibility concern about image-text retrieval. In Proc. of SIGIR. 1301\u20131310."},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_3_2_32_2","volume-title":"Proc. of UAI","author":"Rendle Steffen","year":"2009","unstructured":"Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme. 2009. BPR: Bayesian personalized ranking from implicit feedback. In Proc. of UAI. 452\u2013459."},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-85820-3_1"},{"key":"e_1_3_2_34_2","volume-title":"Proc. of ICLR","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In Proc. of ICLR."},{"issue":"2","key":"e_1_3_2_35_2","article-title":"Self-supervised learning for multimedia recommendation","volume":"24","author":"Tao Zhulin","year":"2022","unstructured":"Zhulin Tao, Xiaohao Liu, Yewei Xia, Xiang Wang, Lifang Yang, Xianglin Huang, and Tat-Seng Chua. 2022. Self-supervised learning for multimedia recommendation. Transactions on Multimedia 24, 2 (2022), 513\u2013525.","journal-title":"Transactions on Multimedia"},{"key":"e_1_3_2_36_2","volume-title":"Proc. of NeurIPS","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proc. of NeurIPS. 5998\u20136008."},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01566"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01838"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11633-022-1410-8"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591716"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3351034"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462862"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3640457.3688096"},{"key":"e_1_3_2_44_2","volume-title":"Proc. of ECIR","author":"Yi Zixuan","year":"2025","unstructured":"Zixuan Yi and Iadh Ounis. 2025. A multi-modal recipe for improved multi-domain recommendation. In Proc. of ECIR."},{"key":"e_1_3_2_45_2","article-title":"Contrastive graph prompt-tuning for cross-domain recommendation","author":"Yi Zixuan","year":"2023","unstructured":"Zixuan Yi, Iadh Ounis, and Craig Macdonald. 2023. Contrastive graph prompt-tuning for cross-domain recommendation. Transactions on Information Systems (2023), 1\u201320.","journal-title":"Transactions on Information Systems"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-28238-6_19"},{"key":"e_1_3_2_47_2","article-title":"A directional diffusion graph transformer for recommendation","author":"Yi Zixuan","year":"2024","unstructured":"Zixuan Yi, Xi Wang, and Iadh Ounis. 2024. A directional diffusion graph transformer for recommendation. arXiv preprint arXiv:2404.03326 (2024).","journal-title":"arXiv preprint arXiv:2404.03326"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3532027"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591932"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01100"},{"key":"e_1_3_2_51_2","volume-title":"Proc. of MM","year":"2021","unstructured":"Jinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu, Shuhui Wang, and Liang Wang 2021. Mining latent structures for multimedia recommendation. In Proc. of MM. 789\u2013798."},{"key":"e_1_3_2_52_2","article-title":"A comprehensive survey on multi-modal recommender systems: Taxonomy, evaluation, and future directions","author":"Zhou Hongyu","year":"2023","unstructured":"Hongyu Zhou, Xin Zhou, Zhiwei Zeng, Lingzi Zhang, and Zhiqi Shen. 2023. A comprehensive survey on multi-modal recommender systems: Taxonomy, evaluation, and future directions. arXiv preprint arXiv:2302.04473 (2023).","journal-title":"arXiv preprint arXiv:2302.04473"},{"key":"e_1_3_2_53_2","volume-title":"Proc. of ICLR","author":"Zhu Bin","year":"2023","unstructured":"Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, et\u00a0al. 2023. LanguageBind: Extending video-language pretraining to N-modality by language-based semantic alignment. In Proc. of ICLR."}],"container-title":["ACM Transactions on Recommender Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3718099","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3718099","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T21:18:37Z","timestamp":1750281517000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3718099"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,10]]},"references-count":52,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,12,31]]}},"alternative-id":["10.1145\/3718099"],"URL":"https:\/\/doi.org\/10.1145\/3718099","relation":{},"ISSN":["2770-6699"],"issn-type":[{"value":"2770-6699","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,10]]},"assertion":[{"value":"2024-05-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}