{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T15:46:56Z","timestamp":1782316016290,"version":"3.54.5"},"reference-count":63,"publisher":"Association for Computing Machinery (ACM)","issue":"7","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["No. U24A20328, 62476071, and 62376137"],"award-info":[{"award-number":["No. U24A20328, 62476071, and 62376137"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100007129","name":"Shandong Provincial Natural Science Foundation","doi-asserted-by":"crossref","award":["No. ZR2022YQ59"],"award-info":[{"award-number":["No. ZR2022YQ59"]}],"id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>Fashion vision-language pre-training (VLP) models have demonstrated remarkable capabilities in excelling at a wide range of fashion cross-modal tasks. However, current models still face three notable limitations (1) an inability to discern varying levels of consistency between textual descriptions and multi-view images, (2) a deficiency in explicit fine-grained alignment between images and text, and (3) a lack of specific supervision mechanisms for facilitating global joint embedding learning. To address these limitations, we propose a novel dual alignment-enhanced fashion VLP model. This model delves deeply into the rich resources of multi-view images and semantic attributes associated with each fashion item. Notably, we introduce two novel pre-training tasks: Multi-grained Adaptive Image-Text Alignment (MAITA) and Joint Embedding-oriented Alignment (JEA). MAITA focuses on optimizing the text\/image encoder by orchestrating adaptive alignment between multi-view images and input text. This encompasses both coarse-grained and fine-grained alignment strategies to enrich semantic understanding, while JEA is devised to supervise the fine-grained semantic learning process of the global multimodal joint embedding. Experimental results spanning four diverse downstream tasks, including cross-modal retrieval, text-guided image retrieval, category recognition, and subcategory recognition, substantiate the significant performance superiority of our model over prior state-of-the-art fashion VLP models.<\/jats:p>","DOI":"10.1145\/3774884","type":"journal-article","created":{"date-parts":[[2025,11,7]],"date-time":"2025-11-07T13:56:04Z","timestamp":1762523764000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Dual Alignment-Enhanced Fashion Vision-Language Pre-Training"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5658-5509","authenticated-orcid":false,"given":"Weili","family":"Guan","sequence":"first","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0595-6856","authenticated-orcid":false,"given":"Kejie","family":"Wang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5274-4197","authenticated-orcid":false,"given":"Xuemeng","family":"Song","sequence":"additional","affiliation":[{"name":"Southern University of Science and Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4317-660X","authenticated-orcid":false,"given":"Kaihao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7778-8807","authenticated-orcid":false,"given":"Xiaojun","family":"Chang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5200-3420","authenticated-orcid":false,"given":"Shengping","family":"Zhang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Weihai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"6077","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"Anderson Peter","year":"2018","unstructured":"Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 6077\u20136086."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_1_4_2","first-page":"4955","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"Baldrati Alberto","year":"2022","unstructured":"Alberto Baldrati, Marco Bertini, Tiberio Uricchio, and Alberto Del Bimbo. 2022. Conditioned and composed image retrieval combining and partially fine-tuning CLIP-based features. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 4955\u20134964."},{"key":"e_1_3_1_5_2","first-page":"1","article-title":"VLMo: Unified vision-language pre-training with mixture-of-modality-experts","volume":"35","author":"Bao Hangbo","year":"2022","unstructured":"Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Songhao Piao, and Furu Wei. 2022. VLMo: Unified vision-language pre-training with mixture-of-modality-experts. In Advances in Neural Information Processing System, Vol. 35, 1\u201316.","journal-title":"Advances in Neural Information Processing System"},{"key":"e_1_3_1_6_2","first-page":"38","article-title":"VLP: A survey on vision-language pre-training","volume":"20","author":"Chen Feilong","year":"2023","unstructured":"Feilong Chen, Duzhen Zhang, Minglun Han, Xiuyi Chen, Jing Shi, Shuang Xu, and Bo Xu. 2023. VLP: A survey on vision-language pre-training. Machine Learning Intelligence 20 (2023), 38\u201356.","journal-title":"Machine Learning Intelligence"},{"key":"e_1_3_1_7_2","first-page":"1597","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. 2020. A simple framework for contrastive learning of visual representations. In Proceedings of the International Conference on Machine Learning, 1597\u20131607."},{"key":"e_1_3_1_8_2","first-page":"2998","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"Chen Yanbei","year":"2020","unstructured":"Yanbei Chen, Shaogang Gong, and Loris Bazzani. 2020. Image search with text feedback by visiolinguistic attention learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 2998\u20133008."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58577-8_7"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3640345"},{"key":"e_1_3_1_11_2","first-page":"3008","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition Workshops","author":"Cubuk Ekin D.","year":"2020","unstructured":"Ekin D. Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V. Le. 2020. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition Workshops, 3008\u20133017."},{"key":"e_1_3_1_12_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4171\u20134186."},{"key":"e_1_3_1_13_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2022\/762"},{"key":"e_1_3_1_15_2","first-page":"1","article-title":"Large-scale adversarial training for vision-and-language representation learning","volume":"33","author":"Gan Zhe","year":"2020","unstructured":"Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu. 2020. Large-scale adversarial training for vision-and-language representation learning. In Advances in Neural Information Processing System, Vol. 33, 1\u201313.","journal-title":"Advances in Neural Information Processing System"},{"key":"e_1_3_1_16_2","first-page":"2251","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Gao Dehong","year":"2020","unstructured":"Dehong Gao, Linbo Jin, Ben Chen, Minghui Qiu, Peng Li, Yi Wei, Yi Hu, and Hao Wang. 2020. FashionBERT: Text and image matching with adaptive loss for cross-modal retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval, 2251\u20132260."},{"key":"e_1_3_1_17_2","first-page":"1","article-title":"LLM-enhanced composed image retrieval: An intent uncertainty-aware linguistic-visual dual channel matching model","volume":"43","author":"Ge Hongfei","year":"2024","unstructured":"Hongfei Ge, Yuanchun Jiang, Jianshan Sun, Kun Yuan, and Yezheng Liu. 2024. LLM-enhanced composed image retrieval: An intent uncertainty-aware linguistic-visual dual channel matching model. ACM Transactions on Information System 43 (2024), 1\u201330.","journal-title":"ACM Transactions on Information System"},{"key":"e_1_3_1_18_2","first-page":"14085","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"Goenka Sonam","year":"2022","unstructured":"Sonam Goenka, Zhaoheng Zheng, Ayush Jaiswal, Rakesh Chada, Yue Wu, Varsha Hedau, and Pradeep Natarajan. 2022. FashionVLP: Vision language transformer for fashion retrieval with feedback. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 14085\u201314095."},{"key":"e_1_3_1_19_2","first-page":"1472","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Han Xintong","year":"2017","unstructured":"Xintong Han, Zuxuan Wu, Phoenix X. Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, and Larry S. Davis. 2017. Automatic spatially-aware fashion concept discovery. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 1472\u20131480."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19833-5_37"},{"key":"e_1_3_1_21_2","first-page":"15028","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"Han Yunpeng","year":"2023","unstructured":"Yunpeng Han, Lisai Zhang, Qingcai Chen, Zhijian Chen, Zhonghua Li, Jianxin Yang, and Zhao Cao. 2023. FashionSAP: Symbols and attributes prompt for fine-grained fashion vision-language pre-training. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 15028\u201315038."},{"key":"e_1_3_1_22_2","first-page":"9726","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"He Kaiming","year":"2020","unstructured":"Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross B. Girshick. 2020. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 9726\u20139735."},{"key":"e_1_3_1_23_2","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Jia Qinjin","year":"2023","unstructured":"Qinjin Jia, Yang Liu, Daoping Wu, Shaoyuan Xu, Huidong Liu, Jinmiao Fu, Roland Vollgraf, and Bryan Wang. 2023. KG-FLIP: Knowledge-guided fashion-domain language-image pre-training for e-commerce. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics."},{"key":"e_1_3_1_24_2","first-page":"735","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Khan Zaid","year":"2022","unstructured":"Zaid Khan, B. G. Vijay Kumar, Xiang Yu, Samuel Schulter, Manmohan Chandraker, and Yun Fu. 2022. Single-stream multi-level alignment for vision-language pretraining. In Proceedings of the European Conference on Computer Vision, 735\u2013751."},{"key":"e_1_3_1_25_2","first-page":"5583","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Kim Wonjae","year":"2021","unstructured":"Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. ViLT: Vision-and-language transformer without convolution or region supervision. In Proceedings of the International Conference on Machine Learning, 5583\u20135594."},{"key":"e_1_3_1_26_2","first-page":"802","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"Lee Seungmin","year":"2021","unstructured":"Seungmin Lee, Dongwan Kim, and Bohyung Han. 2021. CoSMo: Content-Style modulation for image retrieval with text feedback. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 802\u2013812."},{"key":"e_1_3_1_27_2","first-page":"9694","article-title":"Align before fuse: Vision and language representation learning with momentum distillation","volume":"34","author":"Li Junnan","year":"2021","unstructured":"Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, Steven Chu, and Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. In Advances in Neural Information Processing System, Vol. 34, 9694\u20139705.","journal-title":"Advances in Neural Information Processing System"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58577-8_8"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3473140"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3286259"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3449986"},{"issue":"2","key":"e_1_3_1_32_2","first-page":"1","article-title":"Multimodal recommender systems: A survey","volume":"57","author":"Liu Qidong","year":"2024","unstructured":"Qidong Liu, Jiaxi Hu, Yutian Xiao, Xiangyu Zhao, Jingtong Gao, Wanyu Wang, Qing Li, and Jiliang Tang. 2024. Multimodal recommender systems: A survey. ACM Computing Survey 57, 2 (2024), 1\u201317.","journal-title":"ACM Computing Survey"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00213"},{"key":"e_1_3_1_34_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Loshchilov Ilya","year":"2019","unstructured":"Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_35_2","first-page":"18030","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"Ma Haoyu","year":"2022","unstructured":"Haoyu Ma, Handong Zhao, Zhe Lin, Ajinkya Kale, Zhangyang Wang, Tong Yu, Jiuxiang Gu, Sunav Choudhary, and Xiaohui Xie. 2022. EI-CLIP: Entity-aware interventional contrastive learning for e-commerce cross-modal retrieval. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 18030\u201318040."},{"key":"e_1_3_1_36_2","first-page":"38","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Ma Yihui","year":"2017","unstructured":"Yihui Ma, Jia Jia, Suping Zhou, Jingtian Fu, Yejun Liu, and Zijian Tong. 2017. Towards better understanding the clothing fashion styles: A multimodal deep learning approach. In Proceedings of the AAAI Conference on Artificial Intelligence, 38\u201344."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/219717.219748"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.716"},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3708349","article-title":"Vision-language models learn super images for efficient partially relevant video retrieval","volume":"21","author":"Nishimura Taichi","year":"2023","unstructured":"Taichi Nishimura, Shota Nakada, and Masayoshi Kondo. 2023. Vision-language models learn super images for efficient partially relevant video retrieval. ACM Transactions on Multimedia Computing, Communications and Applications 21 (2023), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_1_40_2","first-page":"2798","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"Park Jaeyoo","year":"2023","unstructured":"Jaeyoo Park and Bohyung Han. 2023. Multi-Modal representation learning with text-driven soft masks. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 2798\u20132807."},{"key":"e_1_3_1_41_2","first-page":"1","article-title":"PyTorch: An imperative style, high-performance deep learning library","volume":"32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing System, Vol. 32, 1\u201312.","journal-title":"Advances in Neural Information Processing System"},{"key":"e_1_3_1_42_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, 8748\u20138763."},{"key":"e_1_3_1_43_2","unstructured":"Negar Rostamzadeh Seyedarian Hosseini Thomas Boquet Wojciech Stokowiec Ying Zhang Christian Jauvin and Chris Pal. 2018. Fashion-Gen: The generative fashion dataset and challenge. arXiv:1806.08317. Retrieved from https:\/\/arxiv.org\/abs\/1806.08317"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3708991"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.74"},{"key":"e_1_3_1_46_2","unstructured":"Bin Shan Weichong Yin Yu Sun Hao Tian Hua Wu and Haifeng Wang. 2022. ERNIE-ViL 2.0: Multi-view contrastive learning for image-text pre-training. arXiv:2209.15270. Retrieved from https:\/\/arxiv.org\/abs\/2209.15270"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58621-8_45"},{"key":"e_1_3_1_48_2","unstructured":"A\u00e4ron van den Oord Yazhe Li and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv:1807.03748. Retrieved from https:\/\/arxiv.org\/abs\/1807.03748"},{"key":"e_1_3_1_49_2","first-page":"3156","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"Vinyals Oriol","year":"2015","unstructured":"Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015. Show and tell: A neural image caption generator. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 3156\u20133164."},{"key":"e_1_3_1_50_2","first-page":"8384","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"Wan Yongquan","year":"2024","unstructured":"Yongquan Wan, Wenhai Wang, Guobing Zou, and Bofeng Zhang. 2024. Cross-modal feature alignment and fusion for composed image retrieval. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 8384\u20138388."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3690640"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00306"},{"key":"e_1_3_1_53_2","first-page":"121475","article-title":"CogVLM: Visual expert for pretrained language models","volume":"37","author":"Wang Weihan","year":"2025","unstructured":"Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Song XiXuan, et al. 2025. CogVLM: Visual expert for pretrained language models. In Advances in Neural Information Processing System, Vol. 37, 121475\u2013121499.","journal-title":"Advances in Neural Information Processing System"},{"key":"e_1_3_1_54_2","first-page":"149","volume-title":"Proceedings of the 61st Annual Meeting of the Association For Computational Linguistics","author":"Wang Xiaodan","year":"2023","unstructured":"Xiaodan Wang, Chengyu Wang, Lei Li, Zhixu Li, Ben Chen, Linbo Jin, Jun Huang, Yanghua Xiao, and Ming Gao. 2023. FashionKLIP: Enhancing e-commerce image-text retrieval with fashion multi-modal conceptual knowledge graph. In Proceedings of the 61st Annual Meeting of the Association For Computational Linguistics, 149\u2013158."},{"key":"e_1_3_1_55_2","first-page":"11307","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"Wu Hui","year":"2021","unstructured":"Hui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah, Steven Rennie, Kristen Grauman, and Rog\u00e9rio Feris. 2021. Fashion IQ: A new dataset towards retrieving images by natural language feedback. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 11307\u201311317."},{"key":"e_1_3_1_56_2","unstructured":"Ning Xie Farley Lai Derek Doran and Asim Kadav. 2019. Visual entailment: A novel task for fine-grained image understanding. arxiv:1901.06706. Retrieved from https:\/\/arxiv.org\/abs\/1901.06706"},{"key":"e_1_3_1_57_2","unstructured":"Chang Xu Dacheng Tao and Chao Xu. 2013. A survey on multi-view learning. arXiv:1304.5634. Retrieved from https:\/\/arxiv.org\/abs\/1304.5634"},{"key":"e_1_3_1_58_2","first-page":"15650","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"Yang Jinyu","year":"2022","unstructured":"Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi, and Junzhou Huang. 2022. Vision-language pre-training with triple contrastive learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 15650\u201315659."},{"key":"e_1_3_1_59_2","first-page":"1","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Yang Xuewen","year":"2020","unstructured":"Xuewen Yang, Heming Zhang, Di Jin, Yingru Liu, Chi-Hao Wu, Jianchao Tan, Dongliang Xie, Jue Wang, and Xin Wang. 2020. Fashion captioning: Towards generating accurate descriptions with semantic rewards. In Proceedings of the European Conference on Computer Vision, 1\u201317."},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01100"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3369699"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3532047"},{"key":"e_1_3_1_63_2","first-page":"73","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Zheng Kecheng","year":"2024","unstructured":"Kecheng Zheng, Yifei Zhang, Wei Wu, Fan Lu, Shuailei Ma, Xin Jin, Wei Chen, and Yujun Shen. 2024. DreamLIP: Language-image pre-training with long captions. In Proceedings of the European Conference on Computer Vision. Springer, 73\u201390."},{"key":"e_1_3_1_64_2","first-page":"12647","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition","author":"Zhuge Mingchen","year":"2021","unstructured":"Mingchen Zhuge, Dehong Gao, Deng-Ping Fan, Linbo Jin, Ben Chen, Haoming Zhou, Minghui Qiu, and Ling Shao. 2021. Kaleido-BERT: Vision-language pre-training on fashion domain. In Proceedings of the IEEE\/CVF Conference on Computer Vision Pattern Recognition, 12647\u201312657."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3774884","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:49:15Z","timestamp":1782312555000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3774884"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":63,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3774884"],"URL":"https:\/\/doi.org\/10.1145\/3774884","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-03-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-31","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}