{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T05:00:41Z","timestamp":1784178041468,"version":"3.55.0"},"reference-count":247,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62376137, 62206157, and 624B2047"],"award-info":[{"award-number":["62376137, 62206157, and 624B2047"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100007129","name":"Natural Science Foundation of Shandong Province","doi-asserted-by":"crossref","award":["ZR2022YQ59 and ZR2022QF047"],"award-info":[{"award-number":["ZR2022YQ59 and ZR2022QF047"]}],"id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2026,1,31]]},"abstract":"<jats:p>Composed Image Retrieval (CIR) is an emerging yet challenging task that allows users to search for target images using a multimodal query, comprising a reference image and a modification text specifying the user\u2019s desired changes to the reference image. Given its significant academic and practical value, CIR has become a rapidly growing area of interest in the computer vision and machine learning communities, particularly with the advances in deep learning. To the best of our knowledge, there is currently no comprehensive review of CIR to provide a timely overview of this field. Therefore, we synthesize insights from over 150 publications in top conferences and journals, including ACM TOIS, SIGIR, and CVPR. In particular, we systematically categorize existing supervised CIR and zero-shot CIR models using a fine-grained taxonomy. For a comprehensive review, we also briefly discuss approaches for tasks closely related to CIR, such as attribute-based CIR and dialog-based CIR. Additionally, we summarize benchmark datasets for evaluation and analyze existing supervised and zero-shot CIR methods by comparing experimental results across multiple datasets. Furthermore, we present promising future directions in this field, offering practical insights for researchers interested in further exploration.<\/jats:p>","DOI":"10.1145\/3767328","type":"journal-article","created":{"date-parts":[[2025,9,16]],"date-time":"2025-09-16T13:18:09Z","timestamp":1758028689000},"page":"1-54","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":16,"title":["A Comprehensive Survey on Composed Image Retrieval"],"prefix":"10.1145","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5274-4197","authenticated-orcid":false,"given":"Xuemeng","family":"Song","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Shandong University, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-5768-5467","authenticated-orcid":false,"given":"Haoqiang","family":"Lin","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Shandong University, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0633-3722","authenticated-orcid":false,"given":"Haokun","family":"Wen","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology (Shenzhen), Shenzhen, China and City University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-0751-166X","authenticated-orcid":false,"given":"Bohan","family":"Hou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Shandong University, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1492-0970","authenticated-orcid":false,"given":"Mingzhu","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1476-0273","authenticated-orcid":false,"given":"Liqiang","family":"Nie","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,11,14]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Lorenzo Agnolucci Alberto Baldrati Marco Bertini and Alberto Del Bimbo. 2024. iSEARLE: Improving textual inversion for zero-shot composed image retrieval. arXiv:2405.02951. Retrieved from https:\/\/arxiv.org\/abs\/2405.02951"},{"key":"e_1_3_2_3_2","first-page":"7708","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Ak Kenan E.","year":"2018","unstructured":"Kenan E. Ak, Ashraf A. Kassim, Joo Hwee Lim, and Jo Yew Tham. 2018. Learning attribute representations with localization for flexible fashion search. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 7708\u20137717."},{"key":"e_1_3_2_4_2","first-page":"1671","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Ak Kenan E.","year":"2018","unstructured":"Kenan E. Ak, Joo Hwee Lim, Jo Yew Tham, and Ashraf A. Kassim. 2018. Efficient multi-attribute similarity learning towards attribute-based fashion search. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. IEEE, 1671\u20131679."},{"key":"e_1_3_2_5_2","unstructured":"Rohan Anil Andrew M. Dai Orhan Firat Melvin Johnson Dmitry Lepikhin Alexandre Passos Siamak Shakeri Emanuel Taropa Paige Bailey Zhifeng Chen et al. 2023. PaLM 2 technical report. arXiv:2305.10403. Retrieved from https:\/\/arxiv.org\/abs\/2305.10403"},{"key":"e_1_3_2_6_2","first-page":"1140","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Anwaar Muhammad Umer","year":"2021","unstructured":"Muhammad Umer Anwaar, Egor Labintcev, and Martin Kleinsteuber. 2021. Compositional learning of image-text query for image retrieval. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. IEEE, 1140\u20131149."},{"key":"e_1_3_2_7_2","first-page":"15338","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Baldrati Alberto","year":"2023","unstructured":"Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini, and Alberto Del Bimbo. 2023. Zero-shot composed image retrieval with textual inversion. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, 15338\u201315347."},{"key":"e_1_3_2_8_2","first-page":"4959","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Baldrati Alberto","year":"2022","unstructured":"Alberto Baldrati, Marco Bertini, Tiberio Uricchio, and Alberto Del Bimbo. 2022. Conditioned and composed image retrieval combining and partially fine-tuning clip-based features. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 4959\u20134968."},{"key":"e_1_3_2_9_2","first-page":"21466","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Baldrati Alberto","year":"2022","unstructured":"Alberto Baldrati, Marco Bertini, Tiberio Uricchio, and Alberto Del Bimbo. 2022. Effective conditioned and composed image retrieval combining CLIP-based features. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 21466\u201321474."},{"issue":"3","key":"e_1_3_2_10_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3617597","article-title":"Composed image retrieval using contrastive learning and task-oriented clip-based features","volume":"20","author":"Baldrati Alberto","year":"2023","unstructured":"Alberto Baldrati, Marco Bertini, Tiberio Uricchio, and Alberto Del Bimbo. 2023. Composed image retrieval using contrastive learning and task-oriented clip-based features. ACM Transactions on Multimedia Computing, Communications and Applications 20, 3 (2023), 1\u201324.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_11_2","first-page":"1839","volume-title":"Proceedings of the International Conference on Computational Linguistics","author":"Bao Tong","year":"2025","unstructured":"Tong Bao, Che Liu, Derong Xu, Zhi Zheng, and Tong Xu. 2025. MLLM-I2W: Harnessing multimodal large language model for zero-shot composed image retrieval. In Proceedings of the International Conference on Computational Linguistics, 1839\u20131849."},{"key":"e_1_3_2_12_2","first-page":"1201","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE","author":"Barbany Oriol","year":"2024","unstructured":"Oriol Barbany, Michael Huang, Xinliang Zhu, and Arnab Dhua. 2024. Leveraging large language models for multimodal search. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1201\u20131210."},{"key":"e_1_3_2_13_2","first-page":"663","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Berg Tamara L.","year":"2010","unstructured":"Tamara L. Berg, Alexander C. Berg, and Jonathan Shih. 2010. Automatic attribute discovery and characterization from noisy web data. In Proceedings of the European Conference on Computer Vision. Springer, 663\u2013676."},{"key":"e_1_3_2_14_2","first-page":"18392","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Brooks Tim","year":"2023","unstructured":"Tim Brooks, Aleksander Holynski, and Alexei A. Efros. 2023. InstructPix2Pix: Learning to follow image editing instructions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 18392\u201318402."},{"key":"e_1_3_2_15_2","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 33, 1877\u20131901.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_16_2","unstructured":"Jaeseok Byun Seokhyeon Jeong Wonjae Kim Sanghyuk Chun and Taesup Moon. 2024. Reducing task discrepancy of text encoders for zero-shot composed image retrieval. arXiv:2406.09188. Retrieved from https:\/\/arxiv.org\/abs\/2406.09188"},{"key":"e_1_3_2_17_2","first-page":"1209","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Caesar Holger","year":"2018","unstructured":"Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. 2018. COCO-Stuff: Thing and stuff classes in context. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1209\u20131218."},{"key":"e_1_3_2_18_2","first-page":"3978","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Chawla Pranit","year":"2021","unstructured":"Pranit Chawla, Surgan Jandial, Pinkesh Badjatiya, Ayush Chopra, Mausoom Sarkar, and Balaji Krishnamurthy. 2021. Leveraging style and content features for text conditioned image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 3978\u20133982."},{"key":"e_1_3_2_19_2","first-page":"397","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Chefer Hila","year":"2021","unstructured":"Hila Chefer, Shir Gur, and Lior Wolf. 2021. Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, 397\u2013406."},{"key":"e_1_3_2_20_2","unstructured":"Junyang Chen and Hanjiang Lai. 2023. Pretrain like you inference: Masked tuning improves zero-shot composed image retrieval. arXiv:2311.07622. Retrieved from https:\/\/arxiv.org\/abs\/2311.07622"},{"key":"e_1_3_2_21_2","unstructured":"Junyang Chen and Hanjiang Lai. 2023. Ranking-aware uncertainty for text-guided image retrieval. arXiv:2308.08131. Retrieved from https:\/\/arxiv.org\/abs\/2308.08131"},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","first-page":"69","DOI":"10.1145\/3607827.3616844","volume-title":"Proceedings of the Workshop on Large Generative Models Meet Multimodal Applications","author":"Chen Qianqian","year":"2023","unstructured":"Qianqian Chen, Tianyi Zhang, Maowen Nie, Zheng Wang, Shihao Xu, Wei Shi, and Zhao Cao. 2023. Fashion-GPT: Integrating LLMs with fashion retrieval system. In Proceedings of the Workshop on Large Generative Models Meet Multimodal Applications. ACM, 69\u201378."},{"key":"e_1_3_2_23_2","unstructured":"Xi Chen Josip Djolonga Piotr Padlewski Basil Mustafa Soravit Changpinyo Jialin Wu Carlos Riquelme Ruiz Sebastian Goodman Xiao Wang Yi Tay et al. 2023. PaLI-X: On scaling up a multilingual vision and language model. arXiv:2305.18565. Retrieved from https:\/\/arxiv.org\/abs\/2305.18565"},{"key":"e_1_3_2_24_2","first-page":"136","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Chen Yanbei","year":"2020","unstructured":"Yanbei Chen and Loris Bazzani. 2020. Learning joint visual semantic matching embeddings for language-guided retrieval. In Proceedings of the European Conference on Computer Vision. Springer, 136\u2013152."},{"key":"e_1_3_2_25_2","first-page":"3001","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Chen Yanbei","year":"2020","unstructured":"Yanbei Chen, Shaogang Gong, and Loris Bazzani. 2020. Image search with text feedback by visiolinguistic attention learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 3001\u20133011."},{"key":"e_1_3_2_26_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Chen Yanzhe","year":"2025","unstructured":"Yanzhe Chen, Zhiwen Yang, Jinglin Xu, and Yuxin Peng. 2025. MAI: A multi-turn aggregation-iteration model for composed image retrieval. In Proceedings of the International Conference on Learning Representations. OpenReview.net, 1\u201320."},{"key":"e_1_3_2_27_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Chen Yiyang","year":"2024","unstructured":"Yiyang Chen, Zhedong Zheng, Wei Ji, Leigang Qu, and Tat-Seng Chua. 2024. Composed image retrieval with text feedback via multi-grained uncertainty regularization. In Proceedings of the International Conference on Learning Representations. OpenReview.net, 1\u201313."},{"key":"e_1_3_2_28_2","first-page":"1228","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Chen Yanzhe","year":"2024","unstructured":"Yanzhe Chen, Huasong Zhong, Xiangteng He, Yuxin Peng, Jiahuan Zhou, and Lele Cheng. 2024. FashionERN: Enhance-and-refine network for composed fashion image retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 1228\u20131236."},{"issue":"6","key":"e_1_3_2_29_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3640345","article-title":"SPIRIT: Style-guided patch interaction for fashion image retrieval with text feedback","volume":"20","author":"Chen Yanzhe","year":"2024","unstructured":"Yanzhe Chen, Jiahuan Zhou, and Yuxin Peng. 2024. SPIRIT: Style-guided patch interaction for fashion image retrieval with text feedback. ACM Transactions on Multimedia Computing, Communications and Applications 20, 6 (2024), 1\u201317.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_30_2","unstructured":"Zining Chen Zhicheng Zhao Fei Su Xiaoqin Zhang and Shijian Lu. 2025. Data-efficient generalization for zero-shot composed image retrieval. arXiv:2503.05204. Retrieved from https:\/\/arxiv.org\/abs\/2503.05204"},{"issue":"10","key":"e_1_3_2_31_2","doi-asserted-by":"crossref","first-page":"1865","DOI":"10.1109\/JPROC.2017.2675998","article-title":"Remote sensing image scene classification: Benchmark and state of the art","volume":"105","author":"Cheng Gong","year":"2017","unstructured":"Gong Cheng, Junwei Han, and Xiaoqiang Lu. 2017. Remote sensing image scene classification: Benchmark and state of the art. Proceedings of the IEEE 105, 10 (2017), 1865\u20131883.","journal-title":"Proceedings of the IEEE"},{"key":"e_1_3_2_32_2","first-page":"1724","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Cho Kyunghyun","year":"2014","unstructured":"Kyunghyun Cho. 2014. Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. ACL, 1724\u20131734."},{"key":"e_1_3_2_33_2","first-page":"10972","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Chowdhury Pinaki Nath","year":"2023","unstructured":"Pinaki Nath Chowdhury, Ayan Kumar Bhunia, Aneeshan Sain, Subhadeep Koley, Tao Xiang, and Yi-Zhe Song. 2023. SceneTrilogy: On human scene-sketch and its complementarity with photo and text. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 10972\u201310983."},{"key":"e_1_3_2_34_2","first-page":"253","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Chowdhury Pinaki Nath","year":"2022","unstructured":"Pinaki Nath Chowdhury, Aneeshan Sain, Ayan Kumar Bhunia, Tao Xiang, Yulia Gryaditskaya, and Yi-Zhe Song. 2022. FS-COCO: Towards understanding of freehand sketches of common objects in context. In Proceedings of the European Conference on Computer Vision. Springer, 253\u2013270."},{"key":"e_1_3_2_35_2","first-page":"211","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Chua T.-S.","year":"1994","unstructured":"T.-S. Chua, S.-K. Lim, and H.-K. Pung. 1994. Content-based retrieval of segmented images. In Proceedings of the ACM International Conference on Multimedia. ACM, 211\u2013218."},{"issue":"70","key":"e_1_3_2_36_2","first-page":"1","article-title":"Scaling instruction-finetuned language models","volume":"25","author":"Chung Hyung Won","year":"2024","unstructured":"Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research 25, 70 (2024), 1\u201353.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_37_2","first-page":"558","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Cohen Niv","year":"2022","unstructured":"Niv Cohen, Rinon Gal, Eli A. Meirom, Gal Chechik, and Yuval Atzmon. 2022. \u201cThis is my unicorn, fluffy\u201d: personalizing frozen vision-language representations. In Proceedings of the European Conference on Computer Vision. Springer, 558\u2013577."},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","unstructured":"Ritendra Datta Dhiraj Joshi Jia Li and James Z. Wang. 2008. Image retrieval: Ideas influences and trends of the new age. ACM Computing Surveys 40 2 (2008) 1\u201360.","DOI":"10.1145\/1348246.1348248"},{"key":"e_1_3_2_39_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations. OpenReview.net","author":"Delmas Ginger","year":"2022","unstructured":"Ginger Delmas, Rafael S. Rezende, Gabriela Csurka, and Diane Larlus. 2022. ARTEMIS: Attention-based retrieval with text-explicit matching and implicit similarity. In Proceedings of the International Conference on Learning Representations. OpenReview.net, 1\u201312."},{"key":"e_1_3_2_40_2","first-page":"248","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Deng Jia","year":"2009","unstructured":"Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. ImageNet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 248\u2013255."},{"key":"e_1_3_2_41_2","first-page":"4171","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics. ACL, 4171\u20134186."},{"key":"e_1_3_2_42_2","unstructured":"Eric Dodds Jack Culpepper Simao Herdade Yang Zhang and Kofi Boakye. 2020. Modality-agnostic attention fusion for visual search with text feedback. arXiv:2007.00145. Retrieved from https:\/\/arxiv.org\/abs\/2007.00145"},{"key":"e_1_3_2_43_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy. 2021. An image is worth 16\u2009\u00d7\u200916 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations. OpenReview.net, 1\u201322."},{"key":"e_1_3_2_44_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations. OpenReview.net","author":"Du Yongchao","year":"2024","unstructured":"Yongchao Du, Min Wang, Wengang Zhou, Shuping Hui, and Houqiang Li. 2024. Image2Sentence based asymmetrical zero-shot composed image retrieval. In Proceedings of the International Conference on Learning Representations. OpenReview.net, 1\u201321."},{"key":"e_1_3_2_45_2","first-page":"1003","volume-title":"Proceedings of the ACM International Conference on Web Search and Data Mining","author":"Du Yali","year":"2023","unstructured":"Yali Du, Yinwei Wei, Wei Ji, Fan Liu, Xin Luo, and Liqiang Nie. 2023. Multi-queue momentum contrast for microvideo-product retrieval. In Proceedings of the ACM International Conference on Web Search and Data Mining. ACM, 1003\u20131011."},{"key":"e_1_3_2_46_2","unstructured":"Yiqun Duan Sameera Ramasinghe Stephen Gould and Ajanthan Thalaiyasingam. 2025. Scaling prompt instructed zero shot composed image retrieval with image-only data. arXiv:2504.00812. Retrieved from https:\/\/arxiv.org\/abs\/2504.00812"},{"key":"e_1_3_2_47_2","first-page":"1723","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Efthymiadis Nikos","year":"2025","unstructured":"Nikos Efthymiadis, Bill Psomas, Zakaria Laskar, Konstantinos Karantzalos, Yannis Avrithis, Ond\u0159ej Chum, and Giorgos Tolias. 2025. Composed image retrieval for training-FREE DOMain conversion. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. IEEE, 1723\u20131733."},{"key":"e_1_3_2_48_2","unstructured":"Chun-Mei Feng Yang Bai Tao Luo Zhen Li Salman Khan Wangmeng Zuo Xinxing Xu Rick Siow Mong Goh and Yong Liu. 2023. VQA4CIR: Boosting composed image retrieval with visual question answering. arXiv:2312.12273. Retrieved from https:\/\/arxiv.org\/abs\/2312.12273"},{"key":"e_1_3_2_49_2","first-page":"1","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Feng Zhangchi","year":"2024","unstructured":"Zhangchi Feng, Richong Zhang, and Zhijie Nie. 2024. Improving composed image retrieval via contrastive learning with scaling positives and negatives. In Proceedings of the ACM International Conference on Multimedia. ACM, 1\u201310."},{"key":"e_1_3_2_50_2","first-page":"708","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Forbes Maxwell","year":"2019","unstructured":"Maxwell Forbes, Christine Kaeser-Chen, Piyush Sharma, and Serge J. Belongie. 2019. Neural naturalist: Generating fine-grained image comparisons. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. ACL, 708\u2013717."},{"issue":"6","key":"e_1_3_2_51_2","first-page":"2938","article-title":"DVG-Face: Dual variational generation for heterogeneous face recognition","volume":"44","author":"Fu Chaoyou","year":"2021","unstructured":"Chaoyou Fu, Xiang Wu, Yibo Hu, Huaibo Huang, and Ran He. 2021. DVG-Face: Dual variational generation for heterogeneous face recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 6 (2021), 2938\u20132952.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_52_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Fu Zhiheng","year":"2025","unstructured":"Zhiheng Fu, Zixu Li, Zhiwei Chen, Chunxiao Wang, Xuemeng Song, Yupeng Hu, and Liqiang Nie. 2025. PAIR: Complementarity-guided disentanglement for composed image retrieval. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 1\u20135."},{"key":"e_1_3_2_53_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations. OpenReview.net","author":"Gal Rinon","year":"2023","unstructured":"Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit Haim Bermano, Gal Chechik, and Daniel Cohen-Or. 2023. An image is worth one word: Personalizing text-to-image generation using textual inversion. In Proceedings of the International Conference on Learning Representations. OpenReview.net, 1\u201318."},{"key":"e_1_3_2_54_2","first-page":"5173","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Gao Chengying","year":"2020","unstructured":"Chengying Gao, Qi Liu, Qi Xu, Limin Wang, Jianzhuang Liu, and Changqing Zou. 2020. SketchyCOCO: Image generation from freehand scene sketches. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 5173\u20135182."},{"key":"e_1_3_2_55_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Gao Peng","year":"2025","unstructured":"Peng Gao, Yujian Lee, Zailong Chen, Xubo Liu, Hui Zhang, Yiyang Hu, and Guquan Jing. 2025. NCL-CIR: Noise-aware contrastive learning for composed image retrieval. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 1\u20135."},{"key":"e_1_3_2_56_2","unstructured":"Yunfan Gao Yun Xiong Xinyu Gao Kangxiang Jia Jinliu Pan Yuxi Bi Yixin Dai Jiawei Sun Haofen Wang and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv:2312.10997. Retrieved from https:\/\/arxiv.org\/abs\/2312.10997"},{"key":"e_1_3_2_57_2","first-page":"1869","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Gatti Prajwal","year":"2024","unstructured":"Prajwal Gatti, Kshitij Parikh, Dhriti Prasanna Paul, Manish Gupta, and Anand Mishra. 2024. Composite sketch+ text queries for retrieving objects with elusive names and complex interactions. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 1869\u20131877."},{"issue":"2","key":"e_1_3_2_58_2","first-page":"1","article-title":"LLM-enhanced composed image retrieval: An intent uncertainty-aware linguistic-visual dual channel matching model","volume":"43","author":"Ge Hongfei","year":"2025","unstructured":"Hongfei Ge, Yuanchun Jiang, Jianshan Sun, Kun Yuan, and Yezheng Liu. 2025. LLM-enhanced composed image retrieval: An intent uncertainty-aware linguistic-visual dual channel matching model. ACM Transactions on Information Systems 43, 2 (2025), 1\u201330.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_2_59_2","first-page":"14105","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Goenka Sonam","year":"2022","unstructured":"Sonam Goenka, Zhaoheng Zheng, Ayush Jaiswal, Rakesh Chada, Yue Wu, Varsha Hedau, and Pradeep Natarajan. 2022. FashionVLP: Vision language transformer for fashion retrieval with feedback. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 14105\u201314115."},{"key":"e_1_3_2_60_2","first-page":"6325","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Goyal Yash","year":"2017","unstructured":"Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 6325\u20136334."},{"key":"e_1_3_2_61_2","unstructured":"Alex Graves. 2014. Neural turing machines. arXiv:1410.5401. Retrieved from https:\/\/arxiv.org\/abs\/1410.5401"},{"key":"e_1_3_2_62_2","first-page":"4600","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Gu Chunbin","year":"2021","unstructured":"Chunbin Gu, Jiajun Bu, Zhen Zhang, Zhi Yu, Dongfang Ma, and Wei Wang. 2021. Image search with text feedback by deep hierarchical attention mutual information maximization. In Proceedings of the ACM International Conference on Multimedia. ACM, 4600\u20134609."},{"key":"e_1_3_2_63_2","first-page":"1","article-title":"CompoDiff: Versatile composed image retrieval with latent diffusion","volume":"2024","author":"Gu Geonmo","year":"2024","unstructured":"Geonmo Gu, Sanghyuk Chun, Wonjae Kim, HeeJae Jun, Yoohoon Kang, and Sangdoo Yun. 2024. CompoDiff: Versatile composed image retrieval with latent diffusion. Transactions on Machine Learning Research 2024 (2024), 1\u201330.","journal-title":"Transactions on Machine Learning Research"},{"key":"e_1_3_2_64_2","first-page":"13225","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Gu Geonmo","year":"2024","unstructured":"Geonmo Gu, Sanghyuk Chun, Wonjae Kim, Yoohoon Kang, and Sangdoo Yun. 2024. Language-only training of zero-shot composed image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 13225\u201313234."},{"key":"e_1_3_2_65_2","article-title":"Dialog-based interactive image retrieval","volume":"31","author":"Guo Xiaoxiao","year":"2018","unstructured":"Xiaoxiao Guo, Hui Wu, Yu Cheng, Steven Rennie, Gerald Tesauro, and Rogerio Feris. 2018. Dialog-based interactive image retrieval. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 31.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_66_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Ha David","year":"2018","unstructured":"David Ha and Douglas Eck. 2018. A neural representation of sketch drawings. In Proceedings of the International Conference on Learning Representations. OpenReview.net, 1\u201315."},{"key":"e_1_3_2_67_2","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Han Xintong","year":"2017","unstructured":"Xintong Han, Zuxuan Wu, Phoenix X. Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, and Larry S. Davis. 2017. Automatic spatially-aware fashion concept discovery. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE."},{"key":"e_1_3_2_68_2","first-page":"634","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Han Xiao","year":"2022","unstructured":"Xiao Han, Licheng Yu, Xiatian Zhu, Li Zhang, Yi-Zhe Song, and Tao Xiang. 2022. FashionViL: Fashion-focused vision-and-language representation learning. In Proceedings of the European Conference on Computer Vision. Springer, 634\u2013651."},{"key":"e_1_3_2_69_2","first-page":"2669","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Han Xiao","year":"2023","unstructured":"Xiao Han, Xiatian Zhu, Licheng Yu, Li Zhang, Yi-Zhe Song, and Tao Xiang. 2023. FAME-ViL: Multi-tasking vision-language model for heterogeneous fashion tasks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 2669\u20132680."},{"issue":"8","key":"e_1_3_2_70_2","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter S.","year":"1997","unstructured":"S. Hochreiter and J. Schmidhuber. 1997. Long short-term memory. Neural Computation 9, 8 (1997), 1735\u20131780.","journal-title":"Neural Computation"},{"key":"e_1_3_2_71_2","first-page":"3596","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Hosseinzadeh Mehrdad","year":"2020","unstructured":"Mehrdad Hosseinzadeh and Yang Wang. 2020. Composed query image retrieval using locally bounded features. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 3596\u20133605."},{"key":"e_1_3_2_72_2","first-page":"1","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Hou Bohan","year":"2025","unstructured":"Bohan Hou, Haoqiang Lin, Xuemeng Song, Haokun Wen, Meng Liu, Yupeng Hu, and Xiangyu Zhao. 2025. FiRE: Enhancing MLLMs with fine-grained context learning for complex image retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 1\u201311."},{"key":"e_1_3_2_73_2","first-page":"1","volume-title":"Proceedings of the 2024 International Joint Conference on Neural Networks","author":"Hou Bohan","year":"2025","unstructured":"Bohan Hou, Haoqiang Lin, Haokun Wen, Meng Liu, and Xuemeng Song. 2025. Pseudo-triplet guided few-shot composed image retrieval. In Proceedings of the 2024 International Joint Conference on Neural Networks. IEEE, 1\u20138."},{"key":"e_1_3_2_74_2","first-page":"12147","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Hou Yuxin","year":"2021","unstructured":"Yuxin Hou, Eleonora Vig, Michael Donoser, and Loris Bazzani. 2021. Learning attribute-driven disentangled representations for interactive fashion retrieval. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, 12147\u201312157."},{"key":"e_1_3_2_75_2","first-page":"1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Howard Andrew G.","year":"2017","unstructured":"Andrew G. Howard. 2017. MobileNets: Efficient convolutional neural networks for mobile vision applications. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1\u20139."},{"key":"e_1_3_2_76_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations. OpenReview.net","author":"Hu Edward J.","year":"2021","unstructured":"Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. LoRA: Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations. OpenReview.net, 1\u201326."},{"key":"e_1_3_2_77_2","first-page":"2772","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Hu Zhizhang","year":"2023","unstructured":"Zhizhang Hu, Xinliang Zhu, Son Tran, Ren\u00e9 Vidal, and Arnab Dhua. 2023. ProVLA: Compositional image search with progressive vision-language alignment and multimodal fusion. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, 2772\u20132777."},{"key":"e_1_3_2_78_2","first-page":"6104","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Huang Fuxiang","year":"2023","unstructured":"Fuxiang Huang and Lei Zhang. 2023. Language guided local infiltration for interactive image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 6104\u20136113."},{"key":"e_1_3_2_79_2","first-page":"2303","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Huang Fuxiang","year":"2024","unstructured":"Fuxiang Huang, Lei Zhang, Xiaowei Fu, and Suqi Song. 2024. Dynamic weighted combiner for mixed-modal image retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 2303\u20132311."},{"key":"e_1_3_2_80_2","doi-asserted-by":"crossref","first-page":"7415","DOI":"10.1109\/TMM.2022.3222624","article-title":"Adversarial and isotropic gradient augmentation for image retrieval with text feedback","volume":"25","author":"Huang Fuxiang","year":"2022","unstructured":"Fuxiang Huang, Lei Zhang, Yuhang Zhou, and Xinbo Gao. 2022. Adversarial and isotropic gradient augmentation for image retrieval with text feedback. IEEE Transactions on Multimedia 25 (2022), 7415\u20137427.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_81_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Huang Qinlei","year":"2025","unstructured":"Qinlei Huang, Zhiwei Chen, Zixu Li, Chunxiao Wang, Xuemeng Song, Yupeng Hu, and Liqiang Nie. 2025. MEDIAN: Adaptive intermediate-grained aggregation network for composed image retrieval. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 1\u20135."},{"key":"e_1_3_2_82_2","unstructured":"Yi Huang Jiancheng Huang Yifan Liu Mingfu Yan Jiaxi Lv Jianzhuang Liu Wei Xiong He Zhang Shifeng Chen and Liangliang Cao. 2024. Diffusion model-based image editing: A survey. arXiv:2402.17525. Retrieved from https:\/\/arxiv.org\/abs\/2402.17525"},{"key":"e_1_3_2_83_2","first-page":"1","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Hummel Thomas","year":"2024","unstructured":"Thomas Hummel, Shyamgopal Karthik, Mariana-Iuliana Georgescu, and Zeynep Akata. 2024. EgoCVR: An egocentric benchmark for fine-grained composed video retrieval. In Proceedings of the European Conference on Computer Vision. Springer, 1\u201317."},{"key":"e_1_3_2_84_2","first-page":"3994","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Huynh Chuong","year":"2025","unstructured":"Chuong Huynh, Jinyu Yang, Ashish Tawari, Mubarak Shah, Son Tran, Raffay Hamid, Trishul Chilimbi, and Abhinav Shrivastava. 2025. CoLLM: A large language model for composed image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 3994\u20134004."},{"key":"e_1_3_2_85_2","first-page":"1383","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Isola Phillip","year":"2015","unstructured":"Phillip Isola, Joseph J. Lim, and Edward H. Adelson. 2015. Discovering states and transformations in image collections. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1383\u20131391."},{"issue":"1","key":"e_1_3_2_86_2","doi-asserted-by":"crossref","first-page":"79","DOI":"10.1162\/neco.1991.3.1.79","article-title":"Adaptive mixtures of local experts","volume":"3","author":"Jacobs Robert A.","year":"1991","unstructured":"Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton. 1991. Adaptive mixtures of local experts. Neural Computation 3, 1 (1991), 79\u201387.","journal-title":"Neural Computation"},{"key":"e_1_3_2_87_2","first-page":"4021","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Jandial Surgan","year":"2022","unstructured":"Surgan Jandial, Pinkesh Badjatiya, Pranit Chawla, Ayush Chopra, Mausoom Sarkar, and Balaji Krishnamurthy. 2022. SAC: Semantic attention composition for text-conditioned image retrieval. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. IEEE, 4021\u20134030."},{"key":"e_1_3_2_88_2","unstructured":"Young Kyun Jang Dat Huynh Ashish Shah Wen-Kai Chen and Ser-Nam Lim. 2024. Spherical linear interpolation and text-anchoring for zero-shot composed image retrieval. arXiv:2405.00571. Retrieved from https:\/\/arxiv.org\/abs\/2405.00571"},{"key":"e_1_3_2_89_2","first-page":"1","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Jawade Bhavin","year":"2025","unstructured":"Bhavin Jawade, Joao V. B. Soares, Kapil Thadani, Deen Dayal Mohan, Amir Erfan Eshratifar, Benjamin Culpepper, Paloma de Juan, Srirangaraj Setlur, and Venu Govindaraju. 2025. SCOT: Self-supervised contrastive pretraining for zero-shot compositional retrieval. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. IEEE, 1\u201311."},{"key":"e_1_3_2_90_2","first-page":"4024","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Jhamtani H.","year":"2018","unstructured":"H. Jhamtani and T. Berg-Kirkpatrick. 2018. Learning to describe differences between pairs of similar images. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. ACL, 4024\u20134034."},{"key":"e_1_3_2_91_2","unstructured":"Ting Jiang Minghui Song Zihan Zhang Haizhen Huang Weiwei Deng Feng Sun Qi Zhang Deqing Wang and Fuzhen Zhuang. 2024. E5-V: Universal embeddings with multimodal large language models. arXiv:2407.12580. Retrieved from https:\/\/arxiv.org\/abs\/2407.12580"},{"key":"e_1_3_2_92_2","first-page":"2177","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Jiang Xintong","year":"2024","unstructured":"Xintong Jiang, Yaxiong Wang, Mengjian Li, Yujiao Wu, Bingwen Hu, and Xueming Qian. 2024. CaLa: Complementary association learning for augmenting composed image retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 2177\u20132187."},{"key":"e_1_3_2_93_2","unstructured":"Yingying Jiang Hanchao Jia Xiaobing Wang and Peng Hao. 2024. HyCIR: Boosting zero-shot composed image retrieval with synthetic labels. arXiv:2407.05795. Retrieved from https:\/\/arxiv.org\/abs\/2407.05795"},{"key":"e_1_3_2_94_2","first-page":"1988","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Johnson Justin","year":"2017","unstructured":"Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross B. Girshick. 2017. CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1988\u20131997."},{"key":"e_1_3_2_95_2","first-page":"1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Karthik Shyamgopal","year":"2023","unstructured":"Shyamgopal Karthik, Karsten Roth, Massimiliano Mancini, and Zeynep Akata. 2023. Vision-by-language for training-free compositional image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1\u201315."},{"key":"e_1_3_2_96_2","first-page":"1771","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Kim Jongseok","year":"2021","unstructured":"Jongseok Kim, Youngjae Yu, Hoeseong Kim, and Gunhee Kim. 2021. Dual compositional learning in interactive image retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 1771\u20131779."},{"key":"e_1_3_2_97_2","first-page":"16509","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Koley Subhadeep","year":"2024","unstructured":"Subhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Tao Xiang, and Yi-Zhe Song. 2024. You\u2019ll never walk alone: A sketch and text duet for fine-grained image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 16509\u201316519."},{"key":"e_1_3_2_98_2","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1007\/s11263-016-0981-7","article-title":"Visual genome: Connecting language and vision using crowdsourced dense image annotations","volume":"123","author":"Krishna Ranjay","year":"2017","unstructured":"Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, et al. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International Journal of Computer Vision 123 (2017), 32\u201373.","journal-title":"International Journal of Computer Vision"},{"issue":"6","key":"e_1_3_2_99_2","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1145\/3065386","article-title":"ImageNet classification with deep convolutional neural networks","volume":"60","author":"Krizhevsky Alex","year":"2017","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2017. ImageNet classification with deep convolutional neural networks. Communications of the ACM 60, 6 (2017), 84\u201390.","journal-title":"Communications of the ACM"},{"key":"e_1_3_2_100_2","first-page":"802","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Lee Seungmin","year":"2021","unstructured":"Seungmin Lee, Dongwan Kim, and Bohyung Han. 2021. CoSMo: Content-style modulation for image retrieval with text feedback. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 802\u2013812."},{"key":"e_1_3_2_101_2","first-page":"2991","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Levy Matan","year":"2024","unstructured":"Matan Levy, Rami Ben-Ari, Nir Darshan, and Dani Lischinski. 2024. Data roaming and quality assessment for composed image retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 2991\u20132999."},{"key":"e_1_3_2_102_2","first-page":"108","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo","author":"Li Dafeng","year":"2023","unstructured":"Dafeng Li and Yingying Zhu. 2023. Visual-linguistic alignment and composition for image retrieval with text feedback. In Proceedings of the IEEE International Conference on Multimedia and Expo. IEEE, 108\u2013113."},{"key":"e_1_3_2_103_2","first-page":"19730","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the International Conference on Machine Learning. PMLR, 19730\u201319742."},{"key":"e_1_3_2_104_2","first-page":"12888","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine Learning. PMLR, 12888\u201312900."},{"issue":"2","key":"e_1_3_2_105_2","first-page":"622","article-title":"Pose-guided representation learning for person re-identification","volume":"44","author":"Li Jianing","year":"2019","unstructured":"Jianing Li, Shiliang Zhang, Qi Tian, Meng Wang, and Wen Gao. 2019. Pose-guided representation learning for person re-identification. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 2 (2019), 622\u2013635.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_106_2","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Li Kunpeng","year":"2019","unstructured":"Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu. 2019. Visual semantic reasoning for image-text matching. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE."},{"key":"e_1_3_2_107_2","first-page":"1104","volume-title":"Proceedings of the ACM International Conference on Multimedia Retrieval. ACM","author":"Li Mingyong","year":"2024","unstructured":"Mingyong Li, Zongwei Zhao, Xiaolong Jiang, and Zheng Jiang. 2024. CLIP-ProbCR: CLIP-based probability embedding combination retrieval. In Proceedings of the ACM International Conference on Multimedia Retrieval. ACM, 1104\u20131109."},{"key":"e_1_3_2_108_2","first-page":"636","volume-title":"Proceedings of the ACM International Conference on Multimedia Retrieval","author":"Li Shenshen","year":"2023","unstructured":"Shenshen Li. 2023. Dual-path semantic construction network for composed query-based image retrieval. In Proceedings of the ACM International Conference on Multimedia Retrieval. ACM, 636\u2013639."},{"key":"e_1_3_2_109_2","first-page":"19628","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Li Shuxian","year":"2025","unstructured":"Shuxian Li, Changhao He, Xiting Liu, Joey Tianyi Zhou, Xi Peng, and Peng Hu. 2025. Learning with noisy triplet correspondence for composed image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 19628\u201319637."},{"issue":"4","key":"e_1_3_2_110_2","first-page":"2959","article-title":"Multi-grained attention network with mutual exclusion for composed query-based image retrieval","volume":"34","author":"Li Shenshen","year":"2023","unstructured":"Shenshen Li, Xing Xu, Xun Jiang, Fumin Shen, Xin Liu, and Heng Tao Shen. 2023. Multi-grained attention network with mutual exclusion for composed query-based image retrieval. IEEE Transactions on Circuits and Systems for Video Technology 34, 4 (2023), 2959\u20132972.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"issue":"6","key":"e_1_3_2_111_2","first-page":"1","article-title":"Cross-modal attention preservation with self-contrastive learning for composed query-based image retrieval","volume":"20","author":"Li Shenshen","year":"2024","unstructured":"Shenshen Li, Xing Xu, Xun Jiang, Fumin Shen, Zhe Sun, and Andrzej Cichocki. 2024. Cross-modal attention preservation with self-contrastive learning for composed query-based image retrieval. ACM Transactions on Multimedia Computing, Communications and Applications 20, 6 (2024), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_112_2","first-page":"1","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Wei","year":"2024","unstructured":"Wei Li, Hehe Fan, Yongkang Wong, Yi Yang, and Mohan Kankanhalli. 2024. Improving context understanding in multimodal large language models via multimodal composition learning. In Proceedings of the International Conference on Machine Learning. PMLR, 1\u201321."},{"key":"e_1_3_2_113_2","first-page":"3984","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Li You","year":"2025","unstructured":"You Li, Fan Ma, and Yi Yang. 2025. Imagine and seek: Improving composed image retrieval with an imagined proxy. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 3984\u20133993."},{"key":"e_1_3_2_114_2","first-page":"5101","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Li Zixu","year":"2025","unstructured":"Zixu Li, Zhiwei Chen, Haokun Wen, Zhiheng Fu, Yupeng Hu, and Weili Guan. 2025. Encoder: Entity mining and modification relation binding for composed image retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 5101\u20135109."},{"key":"e_1_3_2_115_2","unstructured":"Zixu Li Zhiheng Fu Yupeng Hu Zhiwei Chen Haokun Wen and Liqiang Nie. 2025. FineCIR: Explicit parsing of fine-grained modification semantics for composed image retrieval. arXiv:2503.21309. Retrieved from https:\/\/arxiv.org\/abs\/2503.21309"},{"key":"e_1_3_2_116_2","doi-asserted-by":"crossref","first-page":"1571","DOI":"10.1145\/3240508.3240646","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Liao Lizi","year":"2018","unstructured":"Lizi Liao, Xiangnan He, Bo Zhao, Chong-Wah Ngo, and Tat-Seng Chua. 2018. Interpretable multimodal retrieval for fashion products. In Proceedings of the ACM International Conference on Multimedia. ACM, 1571\u20131579."},{"key":"e_1_3_2_117_2","first-page":"190","volume-title":"Proceedings of the Australasian Joint Conference on Artificial Intelligence","author":"Lin Haoqiang","year":"2023","unstructured":"Haoqiang Lin, Haokun Wen, Xiaolin Chen, and Xuemeng Song. 2023. CLIP-based composed image retrieval with comprehensive fusion and data augmentation. In Proceedings of the Australasian Joint Conference on Artificial Intelligence. Springer, 190\u2013202."},{"key":"e_1_3_2_118_2","first-page":"240","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Lin Haoqiang","year":"2024","unstructured":"Haoqiang Lin, Haokun Wen, Xuemeng Song, Meng Liu, Yupeng Hu, and Liqiang Nie. 2024. Fine-grained textual inversion network for zero-shot composed image retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 240\u2013250."},{"key":"e_1_3_2_119_2","first-page":"740","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In Proceedings of the European Conference on Computer Vision. Springer, 740\u2013755."},{"key":"e_1_3_2_120_2","first-page":"1","article-title":"RemoteCLIP: A vision language foundation model for remote sensing","volume":"62","author":"Liu Fan","year":"2024","unstructured":"Fan Liu, Delong Chen, Zhangqingyun Guan, Xiaocong Zhou, Jiale Zhu, Qiaolin Ye, Liyong Fu, and Jun Zhou. 2024. RemoteCLIP: A vision language foundation model for remote sensing. IEEE Transactions on Geoscience and Remote Sensing 62 (2024), 1\u201316.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_2_121_2","first-page":"676","article-title":"Visual instruction tuning","volume":"36","author":"Liu Haotian","year":"2024","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 36, 676\u2013686.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_122_2","first-page":"15","volume-title":"Proceedings of the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Liu Meng","year":"2018","unstructured":"Meng Liu, Xiang Wang, Liqiang Nie, Xiangnan He, Baoquan Chen, and Tat-Seng Chua. 2018. Attentive moment retrieval in videos. In Proceedings of the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval, 15\u201324."},{"key":"e_1_3_2_123_2","first-page":"843","volume-title":"Proceedings of the 26th ACM International Conference on Multimedia (ACMMM)","author":"Liu Meng","year":"2018","unstructured":"Meng Liu, Xiang Wang, Liqiang Nie, Qi Tian, Baoquan Chen, and Tat-Seng Chua. 2018. Cross-modal moment localization in videos. In Proceedings of the 26th ACM International Conference on Multimedia (ACMMM), 843\u2013851."},{"key":"e_1_3_2_124_2","unstructured":"Yinhan Liu. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_2_125_2","first-page":"315","volume-title":"MultiMedia Modeling","author":"Liu Yating","year":"2021","unstructured":"Yating Liu and Yan Lu. 2021. Multi-grained fusion for conditional image retrieval. In MultiMedia Modeling. Jakub Loko\u010d, Tom\u00e1\u0161 Skopal, Klaus Schoeffmann, Vasileios Mezaris, Xirong Li, Stefanos Vrochidis, and Ioannis Patras (Eds.), Springer, 315\u2013327."},{"key":"e_1_3_2_126_2","first-page":"381","volume-title":"Proceedings of the British Machine Vision Conference","author":"Liu Yikun","year":"2023","unstructured":"Yikun Liu, Jiangchao Yao, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023. Zero-shot composed text-image retrieval. In Proceedings of the British Machine Vision Conference. BMVA Press, 381\u2013397."},{"key":"e_1_3_2_127_2","first-page":"10012","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Liu Ze","year":"2021","unstructured":"Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, 10012\u201310022."},{"key":"e_1_3_2_128_2","first-page":"2125","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Liu Zheyuan","year":"2021","unstructured":"Zheyuan Liu, Cristian Rodriguez-Opazo, Damien Teney, and Stephen Gould. 2021. Image retrieval on real-life images with pre-trained vision-and-language models. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, 2125\u20132134."},{"key":"e_1_3_2_129_2","first-page":"5753","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Liu Zheyuan","year":"2024","unstructured":"Zheyuan Liu, Weixuan Sun, Yicong Hong, Damien Teney, and Stephen Gould. 2024. Bi-directional training for composed image retrieval via text prompt learning. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. IEEE, 5753\u20135762."},{"key":"e_1_3_2_130_2","first-page":"1","article-title":"Candidate set re-ranking for composed image retrieval with dual multi-modal encoder","volume":"2024","author":"Liu Zheyuan","year":"2024","unstructured":"Zheyuan Liu, Weixuan Sun, Damien Teney, and Stephen Gould. 2024. Candidate set re-ranking for composed image retrieval with dual multi-modal encoder. Transactions on Machine Learning Research 2024 (2024), 1\u201319.","journal-title":"Transactions on Machine Learning Research"},{"key":"e_1_3_2_131_2","volume-title":"Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Long Zijun","year":"2025","unstructured":"Zijun Long, Kangheng Liang, Gerardo Aragon Camarasa, Richard Mccreadie, and Paul Henderson. 2025. Diffusion augmented retrieval: A training-free approach to interactive text-to-image retrieval. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval."},{"key":"e_1_3_2_132_2","first-page":"1666","volume-title":"Proceedings of the ACM on Web Conference","author":"Luo Pengfei","year":"2025","unstructured":"Pengfei Luo, Jingbo Zhou, Tong Xu, Yuan Xia, Linli Xu, and Enhong Chen. 2025. ImageScope: Unifying language-guided image retrieval via large multimodal model collective reasoning. In Proceedings of the ACM on Web Conference, 1666\u20131682."},{"key":"e_1_3_2_133_2","first-page":"10484","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Mirchandani Suvir","year":"2022","unstructured":"Suvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha, Wenwen Jiang, Tao Xiang, and Ning Zhang. 2022. FaD-VLP: Fashion vision-and-language pre-training towards unified retrieval and captioning. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. ACL, 10484\u201310497."},{"key":"e_1_3_2_134_2","first-page":"11323","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Pal Anwesan","year":"2023","unstructured":"Anwesan Pal, Sahil Wadhwa, Ayush Jaiswal, Xu Zhang, Yue Wu, Rakesh Chada, Pradeep Natarajan, and Henrik I. Christensen. 2023. FashionNTM: Multi-turn fashion image retrieval via cascaded memory. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, 11323\u201311334."},{"key":"e_1_3_2_135_2","first-page":"13018","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Pham Khoi","year":"2021","unstructured":"Khoi Pham, Kushal Kafle, Zhe Lin, Zhihong Ding, Scott Cohen, Quan Tran, and Abhinav Shrivastava. 2021. Learning to predict visual attributes in the wild. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 13018\u201313028."},{"key":"e_1_3_2_136_2","first-page":"8526","volume-title":"Proceedings of the IEEE International Geoscience and Remote Sensing Symposium (IGARSS)","author":"Psomas Bill","year":"2024","unstructured":"Bill Psomas, Ioannis Kakogeorgiou, Nikos Efthymiadis, Giorgos Tolias, Ond\u0159ej Chum, Yannis Avrithis, and Konstantinos Karantzalos. 2024. Composed image retrieval for remote sensing. In Proceedings of the IEEE International Geoscience and Remote Sensing Symposium (IGARSS). IEEE, 8526\u20138534."},{"key":"e_1_3_2_137_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Qu Leigang","year":"2025","unstructured":"Leigang Qu, Haochuan Li, Tan Wang, Wenjie Wang, Yongqi Li, Liqiang Nie, and Tat-Seng Chua. 2025. TIGeR: Unifying text-to-image generation and retrieval with large multimodal models. In Proceedings of the International Conference on Learning Representations. OpenReview.net, 1\u201330."},{"key":"e_1_3_2_138_2","first-page":"1047","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Qu Leigang","year":"2020","unstructured":"Leigang Qu, Meng Liu, Da Cao, Liqiang Nie, and Qi Tian. 2020. Context-aware multi-view summarization network for image-text matching. In Proceedings of the ACM International Conference on Multimedia. ACM, 1047\u20131055."},{"key":"e_1_3_2_139_2","first-page":"1104","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Qu Leigang","year":"2021","unstructured":"Leigang Qu, Meng Liu, Jianlong Wu, Zan Gao, and Liqiang Nie. 2021. Dynamic modality interaction modeling for image-text retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 1104\u20131113."},{"key":"e_1_3_2_140_2","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever. 2018. Improving Language Understanding by Generative Pre-Training. San Francisco CA USA. Retrieved from https:\/\/openai.com\/research\/language-unsupervised"},{"key":"e_1_3_2_141_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"issue":"140","key":"e_1_3_2_142_2","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21, 140 (2020), 1\u201367.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_143_2","first-page":"2727","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Rao Jun","year":"2022","unstructured":"Jun Rao, Fei Wang, Liang Ding, Shuhan Qi, Yibing Zhan, Weifeng Liu, and Dacheng Tao. 2022. Where does the performance improvement come from?\u2014A reproducibility concern about image-text retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 2727\u20132737."},{"issue":"6","key":"e_1_3_2_144_2","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren Shaoqing","year":"2017","unstructured":"Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2017. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 39, 6 (2017), 1137\u20131149.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_145_2","first-page":"82","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Ridnik Tal","year":"2021","unstructured":"Tal Ridnik, Emanuel Ben-Baruch, Nadav Zamir, Asaf Noy, Itamar Friedman, Matan Protter, and Lihi Zelnik-Manor. 2021. Asymmetric loss for multi-label classification. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, 82\u201391."},{"key":"e_1_3_2_146_2","first-page":"10684","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Rombach Robin","year":"2022","unstructured":"Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj\u00f6rn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 10684\u201310695."},{"key":"e_1_3_2_147_2","first-page":"19305","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Saito Kuniaki","year":"2023","unstructured":"Kuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li, Chen-Yu Lee, Kate Saenko, and Tomas Pfister. 2023. Pic2Word: Mapping pictures to words for zero-shot composed image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 19305\u201319314."},{"issue":"5","key":"e_1_3_2_148_2","doi-asserted-by":"crossref","first-page":"513","DOI":"10.1016\/0306-4573(88)90021-0","article-title":"Term-weighting approaches in automatic text retrieval","volume":"24","author":"Salton Gerard","year":"1988","unstructured":"Gerard Salton and Christopher Buckley. 1988. Term-weighting approaches in automatic text retrieval. Information Processing & Management 24, 5 (1988), 513\u2013523.","journal-title":"Information Processing & Management"},{"key":"e_1_3_2_149_2","first-page":"251","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Sangkloy Patsorn","year":"2022","unstructured":"Patsorn Sangkloy, Wittawat Jitkrittum, Diyi Yang, and James Hays. 2022. A sketch is worth a thousand words: Image retrieval with text and sketch. In Proceedings of the European Conference on Computer Vision. Springer, 251\u2013267."},{"key":"e_1_3_2_150_2","unstructured":"V. Sanh. 2019. DistilBERT a distilled version of BERT: Smaller faster cheaper and lighter. arXiv:1910.01108. Retrieved from https:\/\/arxiv.org\/abs\/1910.01108"},{"key":"e_1_3_2_151_2","doi-asserted-by":"crossref","first-page":"203","DOI":"10.1016\/j.patrec.2023.06.018","article-title":"Attribute disentanglement with gradient reversal for interactive fashion retrieval","volume":"172","author":"Scaramuzzino Giovanna","year":"2023","unstructured":"Giovanna Scaramuzzino, Federico Becattini, and Alberto Del Bimbo. 2023. Attribute disentanglement with gradient reversal for interactive fashion retrieval. Pattern Recognition Letters 172 (2023), 203\u2013212.","journal-title":"Pattern Recognition Letters"},{"issue":"6","key":"e_1_3_2_152_2","doi-asserted-by":"crossref","first-page":"964","DOI":"10.3390\/rs10060964","article-title":"Performance evaluation of single-label and multi-label remote sensing image retrieval using a dense labeling dataset","volume":"10","author":"Shao Zhenfeng","year":"2018","unstructured":"Zhenfeng Shao, Ke Yang, and Weixun Zhou. 2018. Performance evaluation of single-label and multi-label remote sensing image retrieval using a dense labeling dataset. Remote Sensing 10, 6 (2018), 964.","journal-title":"Remote Sensing"},{"key":"e_1_3_2_153_2","unstructured":"Minchul Shin Yoonjae Cho Byungsoo Ko and Geonmo Gu. 2021. RTIC: Residual learning for text and image composition using graph convolutional network. arXiv:2104.03015. Retrieved from https:\/\/arxiv.org\/abs\/2104.03015"},{"key":"e_1_3_2_154_2","first-page":"245","volume-title":"Proceedings of the Conference on Computer Graphics and Interactive Techniques","author":"Shoemake Ken","year":"1985","unstructured":"Ken Shoemake. 1985. Animating rotation with quaternion curves. In Proceedings of the Conference on Computer Graphics and Interactive Techniques. ACM, 245\u2013254."},{"key":"e_1_3_2_155_2","first-page":"13948","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Song Chull Hwan","year":"2024","unstructured":"Chull Hwan Song, Taebaek Hwang, Jooyoung Yoon, Shunghyun Choi, and Yeong Hyeon Gu. 2024. SyncMask: Synchronized attentional masking for fashion-centric vision-language pretraining. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 13948\u201313957."},{"key":"e_1_3_2_156_2","first-page":"1","volume-title":"Proceedings of the British Machine Vision Conference","author":"Song Jifei","year":"2017","unstructured":"Jifei Song, Yi-Zhe Song, Tao Xiang, and Timothy Hospedales. 2017. Fine-grained image retrieval: The text\/sketch input dilemma. In Proceedings of the British Machine Vision Conference. BMVA Press, 1\u201312."},{"key":"e_1_3_2_157_2","first-page":"5","volume-title":"Proceedings of the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR)","author":"Song Xuemeng","year":"2018","unstructured":"Xuemeng Song, Fuli Feng, Xianjing Han, Xin Yang, Wei Liu, and Liqiang Nie. 2018. Neural compatibility modeling with attentive knowledge distillation. In Proceedings of the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), 5\u201314."},{"key":"e_1_3_2_158_2","doi-asserted-by":"crossref","first-page":"753","DOI":"10.1145\/3123266.3123314","volume-title":"Proceedings of the 25th ACM International Conference on Multimedia (ACMMM)","author":"Song Xuemeng","year":"2017","unstructured":"Xuemeng Song, Fuli Feng, Jinhuan Liu, Zekun Li, Liqiang Nie, and Jun Ma. 2017. NeuroStylist: Neural compatibility modeling for clothing matching. In Proceedings of the 25th ACM International Conference on Multimedia (ACMMM), 753\u2013761."},{"key":"e_1_3_2_159_2","first-page":"6418","volume-title":"Proceedings of the Conference of the Association for Computational Linguistics","author":"Suhr Alane","year":"2019","unstructured":"Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019. A corpus for reasoning about natural language grounded in photographs. In Proceedings of the Conference of the Association for Computational Linguistics. ACL, 6418\u20136428."},{"key":"e_1_3_2_160_2","unstructured":"Shitong Sun Fanghua Ye and Shaogang Gong. 2023. Training-free zero-shot composed image retrieval with local concept reranking. arXiv:2312.08924. Retrieved from https:\/\/arxiv.org\/abs\/2312.08924"},{"key":"e_1_3_2_161_2","unstructured":"Zelong Sun Dong Jing and Zhiwu Lu. 2025. CoTMR: Chain-of-thought multi-scale reasoning for training-free zero-shot composed image retrieval. arXiv:2502.20826. Retrieved from https:\/\/arxiv.org\/abs\/2502.20826"},{"key":"e_1_3_2_162_2","first-page":"7149","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Sun Zelong","year":"2025","unstructured":"Zelong Sun, Dong Jing, Guoxing Yang, Nanyi Fei, and Zhiwu Lu. 2025. Leveraging large vision-language model as user intent-aware encoder for composed image retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 7149\u20137157."},{"key":"e_1_3_2_163_2","first-page":"26951","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Suo Yucheng","year":"2024","unstructured":"Yucheng Suo, Fan Ma, Linchao Zhu, and Yi Yang. 2024. Knowledge-enhanced dual-stream zero-shot composed image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 26951\u201326962."},{"key":"e_1_3_2_164_2","first-page":"1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Szegedy Christian","year":"2015","unstructured":"Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1\u20139."},{"key":"e_1_3_2_165_2","first-page":"24785","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Tang Yuanmin","year":"2025","unstructured":"Yuanmin Tang, Jing Yu, Keke Gai, Jiamin Zhuang, Gang Xiong, Gaopeng Gou, and Qi Wu. 2025. Missing target-relevant information prediction with world model for accurate zero-shot composed image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 24785\u201324795."},{"key":"e_1_3_2_166_2","first-page":"5180","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Tang Yuanmin","year":"2024","unstructured":"Yuanmin Tang, Jing Yu, Keke Gai, Jiamin Zhuang, Gang Xiong, Yue Hu, and Qi Wu. 2024. Context-I2W: Mapping images to context-dependent words for accurate zero-shot composed image retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 5180\u20135188."},{"key":"e_1_3_2_167_2","first-page":"14400","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Tang Yuanmin","year":"2025","unstructured":"Yuanmin Tang, Jue Zhang, Xiaoting Qin, Jing Yu, Gaopeng Gou, Gang Xiong, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, and Qi Wu. 2025. Reason-before-retrieve: One-stage reflective chain-of-thoughts for training-free zero-shot composed image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 14400\u201314410."},{"key":"e_1_3_2_168_2","unstructured":"Ivona Tautkute and Tomasz Trzcinski. 2021. I want this product but different: Multimodal retrieval with synthetic query expansion. arXiv:2102.08871. Retrieved from https:\/\/arxiv.org\/abs\/2102.08871"},{"key":"e_1_3_2_169_2","unstructured":"Gemini Team Rohan Anil Sebastian Borgeaud Jean-Baptiste Alayrac Jiahui Yu Radu Soricut Johan Schalkwyk Andrew M. Dai Anja Hauth Katie Millican et al. 2023. Gemini: A family of highly capable multimodal models. arXiv:2312.11805. Retrieved from https:\/\/arxiv.org\/abs\/2312.11805"},{"key":"e_1_3_2_170_2","first-page":"26896","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Thawakar Omkar","year":"2024","unstructured":"Omkar Thawakar, Muzammal Naseer, Rao Muhammad Anwer, Salman Khan, Michael Felsberg, Mubarak Shah, and Fahad Shahbaz Khan. 2024. Composed video retrieval via enriched context and discriminative embeddings. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 26896\u201326906."},{"key":"e_1_3_2_171_2","first-page":"3974","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Tian Likai","year":"2025","unstructured":"Likai Tian, Jian Zhao, Zechao Hu, Zhengwei Yang, Hao Li, Lei Jin, Zheng Wang, and Xuelong Li. 2025. CCIN: Compositional conflict identification and neutralization for composed image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 3974\u20133983."},{"key":"e_1_3_2_172_2","first-page":"1011","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Tian Yuxin","year":"2023","unstructured":"Yuxin Tian, Shawn Newsam, and Kofi Boakye. 2023. Fashion image retrieval with text feedback by additive attention compositional learning. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. IEEE, 1011\u20131021."},{"key":"e_1_3_2_173_2","unstructured":"Rong-Cheng Tu Zhao Jin Jingyi Liao Xiao Luo Yingjie Wang Li Shen and Dacheng Tao. 2025. MLLM-guided VLM fine-tuning with joint inference for zero-shot composed image retrieval. arXiv:2505.19707. Retrieved from https:\/\/arxiv.org\/abs\/2505.19707"},{"key":"e_1_3_2_174_2","unstructured":"Rong-Cheng Tu Wenhao Sun Hanzhe You Yingjie Wang Jiaxing Huang Li Shen and Dacheng Tao. 2025. Multimodal reasoning agent for zero-shot composed image retrieval. arXiv:2505.19952. Retrieved from https:\/\/arxiv.org\/abs\/2505.19952"},{"key":"e_1_3_2_175_2","unstructured":"Prateksha Udhayanan Srikrishna Karanam and Balaji Vasan Srinivasan. 2023. Learning with multi-modal gradient attention for explainable composed image retrieval. arXiv:2308.16649. Retrieved from https:\/\/arxiv.org\/abs\/2308.16649"},{"key":"e_1_3_2_176_2","first-page":"5998","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 30, 5998\u20136008.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_177_2","first-page":"6862","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Vaze Sagar","year":"2023","unstructured":"Sagar Vaze, Nicolas Carion, and Ishan Misra. 2023. GeneCIS: A benchmark for general conditional image similarity. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 6862\u20136872."},{"key":"e_1_3_2_178_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations. OpenReview","author":"Veli\u010dkovi\u0107 Petar","year":"2018","unstructured":"Petar Veli\u010dkovi\u0107, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Li\u00f2, and Yoshua Bengio. 2018. Graph attention networks. In Proceedings of the International Conference on Learning Representations. OpenReview.net, 1\u201312."},{"issue":"12","key":"e_1_3_2_179_2","first-page":"1","article-title":"CoVR-2: Automatic data construction for composed video retrieval","volume":"46","author":"Ventura Lucas","year":"2024","unstructured":"Lucas Ventura, Antoine Yang, Cordelia Schmid, and G\u00fcl Varol. 2024. CoVR-2: Automatic data construction for composed video retrieval. IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 12 (2024), 1\u201315.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_180_2","first-page":"5270","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Ventura Lucas","year":"2024","unstructured":"Lucas Ventura, Antoine Yang, Cordelia Schmid, and G\u00fcl Varol. 2024. CoVR: Learning composed video retrieval from web video captions. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 5270\u20135279."},{"key":"e_1_3_2_181_2","first-page":"6439","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Vo Nam","year":"2019","unstructured":"Nam Vo, Lu Jiang, Chen Sun, Kevin Murphy, Li-Jia Li, Li Fei-Fei, and James Hays. 2019. Composing text and image for image retrieval\u2014An empirical odyssey. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 6439\u20136448."},{"key":"e_1_3_2_182_2","first-page":"8384","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wan Yongquan","year":"2024","unstructured":"Yongquan Wan, Wenhai Wang, Guobing Zou, and Bofeng Zhang. 2024. Cross-modal feature alignment and fusion for composed image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 8384\u20138388."},{"key":"e_1_3_2_183_2","unstructured":"Chaoyang Wang Zeyu Zhang Long Teng Zijun Li and Shichao Kan. 2025. TMCIR: Token merge benefits composed image retrieval. arXiv:2504.10995. Retrieved from https:\/\/arxiv.org\/abs\/2504.10995"},{"key":"e_1_3_2_184_2","first-page":"1","article-title":"Scene graph-aware hierarchical fusion network for remote sensing image retrieval with text feedback","volume":"62","author":"Wang Fei","year":"2024","unstructured":"Fei Wang, Xianzhang Zhu, Xiaojian Liu, Yongjun Zhang, and Yansheng Li. 2024. Scene graph-aware hierarchical fusion network for remote sensing image retrieval with text feedback. IEEE Transactions on Geoscience and Remote Sensing 62 (2024), 1\u201316.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"issue":"4","key":"e_1_3_2_185_2","doi-asserted-by":"crossref","first-page":"769","DOI":"10.1109\/TPAMI.2017.2699960","article-title":"A survey on learning to hash","volume":"40","author":"Wang Jingdong","year":"2017","unstructured":"Jingdong Wang, Ting Zhang, Jingkuan Song, Nicu Sebe, and Heng Tao Shen. 2017. A survey on learning to hash. IEEE Transactions on Pattern Analysis and Machine Intelligence 40, 4 (2017), 769\u2013790.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_186_2","first-page":"29690","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang Lan","year":"2025","unstructured":"Lan Wang, Wei Ao, Vishnu Naresh Boddeti, and Ser-Nam Lim. 2025. Generative zero-shot composed image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 29690\u201329700."},{"key":"e_1_3_2_187_2","first-page":"1","volume-title":"Proceedings of the Asian Conference on Machine Learning","author":"Wang Peng","year":"2024","unstructured":"Peng Wang, Zining Chen, Zhicheng Zhao, and Fei Su. 2024. Prompting vision-language fusion for zero-shot composed image retrieval. In Proceedings of the Asian Conference on Machine Learning, 1\u201316."},{"key":"e_1_3_2_188_2","first-page":"1","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Wang Yifan","year":"2024","unstructured":"Yifan Wang, Wuliang Huang, Lei Li, and Chun Yuan. 2024. Semantic distillation from neighborhood for composed image retrieval. In Proceedings of the ACM International Conference on Multimedia. ACM, 1\u20139."},{"key":"e_1_3_2_189_2","doi-asserted-by":"crossref","first-page":"7608","DOI":"10.1109\/TMM.2024.3369898","article-title":"Negative-sensitive framework with semantic enhancement for composed image retrieval","volume":"26","author":"Wang Yifan","year":"2024","unstructured":"Yifan Wang, Liyuan Liu, Chun Yuan, Minbo Li, and Jing Liu. 2024. Negative-sensitive framework with semantic enhancement for composed image retrieval. IEEE Transactions on Multimedia 26 (2024), 7608\u20137621.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_190_2","unstructured":"Yabing Wang Zhuotao Tian Qingpei Guo Zheng Qin Sanping Zhou Ming Yang and Le Wang. 2025. From mapping to composing: A two-stage framework for zero-shot composed image retrieval. arXiv:2504.17990. Retrieved from https:\/\/arxiv.org\/abs\/2504.17990"},{"key":"e_1_3_2_191_2","first-page":"6390","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Wei Hao","year":"2023","unstructured":"Hao Wei, Shuhui Wang, Zhe Xue, Shengbo Chen, and Qingming Huang. 2023. Conversational composed retrieval with iterative sequence refinement. In Proceedings of the ACM International Conference on Multimedia. ACM, 6390\u20136399."},{"key":"e_1_3_2_192_2","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 35, 24824\u201324837.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_193_2","first-page":"229","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Wen Haokun","year":"2024","unstructured":"Haokun Wen, Xuemeng Song, Xiaolin Chen, Yinwei Wei, Liqiang Nie, and Tat-Seng Chua. 2024. Simple but effective raw-data level multimodal fusion for composed image retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 229\u2013239."},{"key":"e_1_3_2_194_2","first-page":"1369","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Wen Haokun","year":"2021","unstructured":"Haokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan, and Liqiang Nie. 2021. Comprehensive linguistic-visual composition network for image retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 1369\u20131378."},{"issue":"5","key":"e_1_3_2_195_2","doi-asserted-by":"crossref","first-page":"3665","DOI":"10.1109\/TPAMI.2023.3346434","article-title":"Self-training boosted multi-factor matching network for composed image retrieval","volume":"46","author":"Wen Haokun","year":"2024","unstructured":"Haokun Wen, Xuemeng Song, Jianhua Yin, Jianlong Wu, Weili Guan, and Liqiang Nie. 2024. Self-training boosted multi-factor matching network for composed image retrieval. IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 5 (2024), 3665\u20133678.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_196_2","first-page":"915","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Wen Haokun","year":"2023","unstructured":"Haokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei, and Liqiang Nie. 2023. Target-guided composed image retrieval. In Proceedings of the ACM International Conference on Multimedia. ACM, 915\u2013923."},{"key":"e_1_3_2_197_2","first-page":"11307","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wu Hui","year":"2021","unstructured":"Hui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah, Steven Rennie, Kristen Grauman, and Rogerio Feris. 2021. Fashion IQ: A new dataset towards retrieving images by natural language feedback. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 11307\u201311317."},{"key":"e_1_3_2_198_2","first-page":"4729","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Wu Junda","year":"2023","unstructured":"Junda Wu, Rui Wang, Handong Zhao, Ruiyi Zhang, Chaochao Lu, Shuai Li, and Ricardo Henao. 2023. Few-shot composition learning for image retrieval with prompt tuning. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 4729\u20134737."},{"key":"e_1_3_2_199_2","unstructured":"Ren-Di Wu Yu-Yen Lin and Huei-Fang Yang. 2024. Training-free zero-shot composed image retrieval via weighted modality fusion and similarity. arXiv:2409.04918. Retrieved from https:\/\/arxiv.org\/abs\/2409.04918"},{"issue":"11","key":"e_1_3_2_200_2","doi-asserted-by":"crossref","first-page":"3600","DOI":"10.1109\/TCAD.2024.3445809","article-title":"FIRM-Tree: A multidimensional index structure for reprogrammable flash memory","volume":"43","author":"Wu Shin-Ting","year":"2024","unstructured":"Shin-Ting Wu, Pin-Jung Chen, Po-Chun Huang, Wei-Kuan Shih, and Yuan-Hao Chang. 2024. FIRM-Tree: A multidimensional index structure for reprogrammable flash memory. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 43, 11 (2024), 3600\u20133613.","journal-title":"IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems"},{"key":"e_1_3_2_201_2","first-page":"3260","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Wu Yiming","year":"2024","unstructured":"Yiming Wu, Hangfei Li, Fangfang Wang, Yilong Zhang, and Ronghua Liang. 2024. Self-distilled dynamic fusion network for language-based fashion retrieval. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 3260\u20133264."},{"key":"e_1_3_2_202_2","first-page":"19638","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Xing Eric","year":"2025","unstructured":"Eric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou, and Nathan Jacobs. 2025. ConText-CIR: Learning from concepts in text for composed image retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 19638\u201319648."},{"key":"e_1_3_2_203_2","first-page":"2048","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu. 2015. Show, attend and tell: Neural image caption generation with visual attention. In Proceedings of the International Conference on Machine Learning. PMLR, 2048\u20132057."},{"key":"e_1_3_2_204_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Xu Xinxing","year":"2024","unstructured":"Xinxing Xu, Yong Liu, Salman Khan, Fahad Khan, Wangmeng Zuo, Rick Siow Mong Goh, Chun-Mei Feng, et al. 2024. Sentence-level prompts benefit composed image retrieval. In Proceedings of the International Conference on Learning Representations. OpenReview.net, 1\u201313."},{"key":"e_1_3_2_205_2","doi-asserted-by":"crossref","first-page":"8346","DOI":"10.1109\/TMM.2023.3235495","article-title":"Multi-modal transformer with global-local alignment for composed query image retrieval","volume":"25","author":"Xu Yahui","year":"2023","unstructured":"Yahui Xu, Yi Bin, Jiwei Wei, Yang Yang, Guoqing Wang, and Heng Tao Shen. 2023. Multi-modal transformer with global-local alignment for composed query image retrieval. IEEE Transactions on Multimedia 25 (2023), 8346\u20138357.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_206_2","first-page":"1","article-title":"Align and retrieve: Composition and decomposition learning in image retrieval with text feedback","volume":"26","author":"Xu Yahui","year":"2024","unstructured":"Yahui Xu, Yi Bin, Jiwei Wei, Yang Yang, Guoqing Wang, and Heng Tao Shen. 2024. Align and retrieve: Composition and decomposition learning in image retrieval with text feedback. IEEE Transactions on Multimedia 26 (2024), 1\u201313.","journal-title":"IEEE Transactions on Multimedia"},{"issue":"10","key":"e_1_3_2_207_2","doi-asserted-by":"crossref","first-page":"10494","DOI":"10.1109\/TCSVT.2024.3401006","article-title":"Set of diverse queries with uncertainty regularization for composed image retrieval","volume":"34","author":"Xu Yahui","year":"2024","unstructured":"Yahui Xu, Jiwei Wei, Yi Bin, Yang, Yang Zeyu Ma, and Heng Tao Shen. 2024. Set of diverse queries with uncertainty regularization for composed image retrieval. IEEE Transactions on Circuits and Systems for Video Technology 34, 10 (2024), 10494\u201310506.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_2_208_2","first-page":"239","volume-title":"MultiMedia Modeling","author":"Yan Cairong","year":"2024","unstructured":"Cairong Yan, Meng Ma, Yanting Zhang, and Yongquan Wan. 2024. Dual-path multimodal optimal transport for composed image retrieval. In MultiMedia Modeling. Minsu Cho, Ivan Laptev, Du Tran, Angela Yao and Hongbin Zha (Eds.), Springer Nature, Singapore, 239\u2013254."},{"key":"e_1_3_2_209_2","first-page":"447","volume-title":"Proceedings of the International Conference on Intelligent Computing","author":"Yan Cairong","year":"2024","unstructured":"Cairong Yan, Erhe Yang, Ran Tao, Yongquan Wan, and Derun Ai. 2024. SHAF: Semantic-guided hierarchical alignment and fusion for composed image retrieval. In Proceedings of the International Conference on Intelligent Computing. Springer, 447\u2013459."},{"key":"e_1_3_2_210_2","first-page":"1","article-title":"Language-empowered conversion for remote sensing image retrieval with text feedback","author":"Yang Jian","year":"2025","unstructured":"Jian Yang, Shengyang Li, Yuhan Sun, Han Wang, and Zhuang Zhou. 2025. Language-empowered conversion for remote sensing image retrieval with text feedback. IEEE Transactions on Geoscience and Remote Sensing 63 (2025), 1\u201315.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_2_211_2","doi-asserted-by":"crossref","first-page":"4543","DOI":"10.1109\/TIP.2023.3299791","article-title":"Composed image retrieval via cross relation network with hierarchical aggregation transformer","volume":"32","author":"Yang Qu","year":"2023","unstructured":"Qu Yang, Mang Ye, Zhaohui Cai, Kehua Su, and Bo Du. 2023. Composed image retrieval via cross relation network with hierarchical aggregation transformer. IEEE Transactions on Image Processing 32 (2023), 4543\u20134554.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_2_212_2","first-page":"6576","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Yang Xingyu","year":"2024","unstructured":"Xingyu Yang, Daqing Liu, Heng Zhang, Yong Luo, Chaoyue Wang, and Jing Zhang. 2024. Decomposing semantic shifts for composed image retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 6576\u20136584."},{"key":"e_1_3_2_213_2","first-page":"941","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Yang Xin","year":"2020","unstructured":"Xin Yang, Xuemeng Song, Xianjing Han, Haokun Wen, Jie Nie, and Liqiang Nie. 2020. Generative attribute manipulation scheme for flexible fashion search. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 941\u2013950."},{"key":"e_1_3_2_214_2","first-page":"270","volume-title":"Proceedings of the SIGSPATIAL International Conference on Advances in Geographic Information Systems","author":"Yang Yi","year":"2010","unstructured":"Yi Yang and Shawn Newsam. 2010. Bag-of-visual-words and spatial extensions for land-use classification. In Proceedings of the SIGSPATIAL International Conference on Advances in Geographic Information Systems. ACM, 270\u2013279."},{"key":"e_1_3_2_215_2","first-page":"3303","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Yang Yuchen","year":"2021","unstructured":"Yuchen Yang, Min Wang, Wengang Zhou, and Houqiang Li. 2021. Cross-modal joint prediction and alignment for composed query image retrieval. In Proceedings of the ACM International Conference on Multimedia. ACM, 3303\u20133311."},{"key":"e_1_3_2_216_2","first-page":"5250","volume-title":"Proceedings of the Conference of the Association for Computational Linguistics","author":"Yang Yuchen","year":"2024","unstructured":"Yuchen Yang, Yu Wang, and Yanfeng Wang. 2024. SDA: Semantic discrepancy alignment for text-conditioned image retrieval. In Proceedings of the Conference of the Association for Computational Linguistics. ACL, 5250\u20135261."},{"key":"e_1_3_2_217_2","unstructured":"Yuxin Yang Yinan Zhou Yuxin Chen Ziqi Zhang Zongyang Ma Chunfeng Yuan Bing Li Lin Song Jun Gao Peng Li et al. 2025. DetailFusion: A dual-branch framework with detail enhancement for composed image retrieval. arXiv:2505.17796. Retrieved from https:\/\/arxiv.org\/abs\/2505.17796"},{"key":"e_1_3_2_218_2","first-page":"1245","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Yang Zhenyu","year":"2024","unstructured":"Zhenyu Yang, Shengsheng Qian, Dizhan Xue, Jiahong Wu, Fan Yang, Weiming Dong, and Changsheng Xu. 2024. Semantic editing increment benefits zero-shot composed image retrieval. In Proceedings of the ACM International Conference on Multimedia. ACM, 1245\u20131254."},{"key":"e_1_3_2_219_2","first-page":"80","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Yang Zhenyu","year":"2024","unstructured":"Zhenyu Yang, Dizhan Xue, Shengsheng Qian, Weiming Dong, and Changsheng Xu. 2024. LDRE: LLM-based divergent reasoning and ensemble for zero-shot composed image retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 80\u201390."},{"key":"e_1_3_2_220_2","unstructured":"Yang Ye Xianyi He Zongjian Li Bin Lin Shenghai Yuan Zhiyuan Yan Bohan Hou and Li Yuan. 2025. ImgEdit: A unified image editing dataset and benchmark. arXiv:2505.20275. Retrieved from https:\/\/arxiv.org\/abs\/2505.20275"},{"key":"e_1_3_2_221_2","first-page":"839","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Yuan Yifei","year":"2021","unstructured":"Yifei Yuan and Wai Lam. 2021. Conversational fashion image retrieval via multiturn natural language feedback. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 839\u2013848."},{"key":"e_1_3_2_222_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations. OpenReview.net","author":"Yue W. U.","year":"2025","unstructured":"W. U. Yue, Zhaobo Qi, Yiling Wu, Junshu Sun, Yaowei Wang, and Shuhui Wang. 2025. Learning fine-grained representations through textual token disentanglement in composed video retrieval. In Proceedings of the International Conference on Learning Representations. OpenReview.net, 1\u201323."},{"key":"e_1_3_2_223_2","first-page":"12116","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zaeemzadeh Alireza","year":"2021","unstructured":"Alireza Zaeemzadeh, Shabnam Ghadar, Baldo Faieta, Zhe Lin, Nazanin Rahnavard, Mubarak Shah, and Ratheesh Kalarot. 2021. Face image retrieval with attribute manipulation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, 12116\u201312125."},{"issue":"12","key":"e_1_3_2_224_2","doi-asserted-by":"crossref","first-page":"15098","DOI":"10.1109\/TPAMI.2023.3305243","article-title":"Multimodal image synthesis and editing: The generative AI era","volume":"45","author":"Zhan Fangneng","year":"2023","unstructured":"Fangneng Zhan, Yingchen Yu, Rongliang Wu, Jiahui Zhang, Shijian Lu, Lingjie Liu, Adam Kortylewski, Christian Theobalt, and Eric Xing. 2023. Multimodal image synthesis and editing: The generative AI era. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 12 (2023), 15098\u201315119.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_225_2","first-page":"3367","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Zhang Feifei","year":"2020","unstructured":"Feifei Zhang, Mingliang Xu, Qirong Mao, and Changsheng Xu. 2020. Joint attribute manipulation and modality alignment learning for composing text and image to image retrieval. In Proceedings of the ACM International Conference on Multimedia. ACM, 3367\u20133376."},{"key":"e_1_3_2_226_2","doi-asserted-by":"crossref","first-page":"1000","DOI":"10.1109\/TIP.2021.3138302","article-title":"Geometry sensitive cross-modal reasoning for composed query based image retrieval","volume":"31","author":"Zhang Feifei","year":"2021","unstructured":"Feifei Zhang, Mingliang Xu, and Changsheng Xu. 2021. Geometry sensitive cross-modal reasoning for composed query based image retrieval. IEEE Transactions on Image Processing 31 (2021), 1000\u20131011.","journal-title":"IEEE Transactions on Image Processing"},{"issue":"2","key":"e_1_3_2_227_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3478642","article-title":"Tell, imagine, and search: End-to-end learning for composing text and image to image retrieval","volume":"18","author":"Zhang Feifei","year":"2022","unstructured":"Feifei Zhang, Mingliang Xu, and Changsheng Xu. 2022. Tell, imagine, and search: End-to-end learning for composing text and image to image retrieval. ACM Transactions on Multimedia Computing, Communications and Applications 18, 2 (2022), 1\u201323.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_228_2","doi-asserted-by":"crossref","first-page":"1149","DOI":"10.1109\/TIP.2024.3359062","article-title":"Multimodal composition example mining for composed query image retrieval","volume":"33","author":"Zhang Gangjian","year":"2024","unstructured":"Gangjian Zhang, Shikun Li, Shikui Wei, Shiming Ge, Na Cai, and Yao Zhao. 2024. Multimodal composition example mining for composed query image retrieval. IEEE Transactions on Image Processing 33 (2024), 1149\u20131161.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_2_229_2","doi-asserted-by":"crossref","first-page":"5976","DOI":"10.1109\/TIP.2022.3204213","article-title":"Composed image retrieval via explicit erasure and replenishment with semantic alignment","volume":"31","author":"Zhang Gangjian","year":"2022","unstructured":"Gangjian Zhang, Shikui Wei, Huaxin Pang, Shuang Qiu, and Yao Zhao. 2022. Composed image retrieval via explicit erasure and replenishment with semantic alignment. IEEE Transactions on Image Processing 31 (2022), 5976\u20135988.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_2_230_2","doi-asserted-by":"crossref","first-page":"916","DOI":"10.1109\/TMM.2023.3273466","article-title":"Enhance composed image retrieval via multi-level collaborative localization and semantic activeness perception","volume":"26","author":"Zhang Gangjian","year":"2023","unstructured":"Gangjian Zhang, Shikui Wei, Huaxin Pang, Shuang Qiu, and Yao Zhao. 2023. Enhance composed image retrieval via multi-level collaborative localization and semantic activeness perception. IEEE Transactions on Multimedia 26 (2023), 916\u2013928.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_231_2","first-page":"5353","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Zhang Gangjian","year":"2021","unstructured":"Gangjian Zhang, Shikui Wei, Huaxin Pang, and Yao Zhao. 2021. Heterogeneous feature fusion and cross-modal alignment for composed image retrieval. In Proceedings of the ACM International Conference on Multimedia. ACM, 5353\u20135362."},{"key":"e_1_3_2_232_2","first-page":"2431","volume-title":"Proceedings of the IEEE International Conference on Image Processing","author":"Zhang Huaying","year":"2024","unstructured":"Huaying Zhang, Rintaro Yanagi, Ren Togo, Takahiro Ogawa, and Miki Haseyama. 2024. Zero-shot composed image retrieval considering query-target relationship leveraging masked image-text pairs. In Proceedings of the IEEE International Conference on Image Processing. IEEE, 2431\u20132437."},{"key":"e_1_3_2_233_2","first-page":"411","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo","author":"Zhang Jiayan","year":"2023","unstructured":"Jiayan Zhang, Jie Zhang, Honghao Wu, Zongwei Zhao, Jinyu Hu, and Mingyong Li. 2023. PCaSM: Text-guided composed image retrieval with parallel content and style modules. In Proceedings of the IEEE International Conference on Multimedia and Expo. IEEE, 411\u2013416."},{"key":"e_1_3_2_234_2","first-page":"1","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Zhang Kai","year":"2024","unstructured":"Kai Zhang, Yi Luan, Hexiang Hu, Kenton Lee, Siyuan Qiao, Wenhu Chen, Yu Su, and Ming-Wei Chang. 2024. MagicLens: Self-supervised image retrieval with open-ended instructions. In Proceedings of the International Conference on Machine Learning. PMLR, 1\u201318."},{"key":"e_1_3_2_235_2","unstructured":"Susan Zhang Stephen Roller Naman Goyal Mikel Artetxe Moya Chen Shuohui Chen Christopher Dewan Mona Diab Xian Li Xi Victoria Lin et al. 2022. OPT: Open pre-trained transformer language models. arXiv:2205.01068. Retrieved from https:\/\/arxiv.org\/abs\/2205.01068"},{"key":"e_1_3_2_236_2","first-page":"9274","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhang Xin","year":"2025","unstructured":"Xin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li, Ziqi Dai, Dingkun Long, Pengjun Xie, Meishan Zhang, Wenjie Li, and Min Zhang. 2025. Bridging modalities: Improving universal multimodal retrieval by multimodal large language models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9274\u20139285."},{"key":"e_1_3_2_237_2","doi-asserted-by":"crossref","first-page":"112","DOI":"10.1016\/j.knosys.2024.112135","article-title":"Collaborative group: Composed image retrieval via consensus learning from noisy annotations","volume":"300","author":"Zhang Xu","year":"2024","unstructured":"Xu Zhang, Zhedong Zheng, Linchao Zhu, and Yi Yang. 2024. Collaborative group: Composed image retrieval via consensus learning from noisy annotations. Knowledge-Based Systems 300 (2024), 112\u2013135.","journal-title":"Knowledge-Based Systems"},{"key":"e_1_3_2_238_2","first-page":"1520","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE","author":"Zhao Bo","year":"2017","unstructured":"Bo Zhao, Jiashi Feng, Xiao Wu, and Shuicheng Yan. 2017. Memory-augmented attribute manipulation networks for interactive fashion search. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1520\u20131528."},{"key":"e_1_3_2_239_2","first-page":"47","volume-title":"Proceedings of the 1st Workshop on Unifying Representations in Neural Models (UniReps)","author":"Zhao Shu","year":"2024","unstructured":"Shu Zhao and Huijuan Xu. 2024. NEUCORE: Neural concept reasoning for composed image retrieval. In Proceedings of the 1st Workshop on Unifying Representations in Neural Models (UniReps). PMLR, 47\u201359."},{"key":"e_1_3_2_240_2","first-page":"1490","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Zhao Xiangyu","year":"2024","unstructured":"Xiangyu Zhao, Yuehan Zhang, Wenlong Zhang, and Xiao-Ming Wu. 2024. UniFashion: A unified vision-language model for multimodal fashion retrieval and generation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. ACL, 1490\u20131507."},{"key":"e_1_3_2_241_2","doi-asserted-by":"crossref","unstructured":"Xiangyu Zhao Yuehan Zhang Wenlong Zhang and Xiao-Ming Wu. 2024. UniFashion: A unified vision-language model for multimodal fashion retrieval and generation. arXiv:2408.11305. Retrieved from https:\/\/arxiv.org\/abs\/2408.11305","DOI":"10.18653\/v1\/2024.emnlp-main.89"},{"key":"e_1_3_2_242_2","first-page":"1012","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Zhao Yida","year":"2022","unstructured":"Yida Zhao, Yuqing Song, and Qin Jin. 2022. Progressive learning for image retrieval with hybrid-modality queries. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 1012\u20131021."},{"key":"e_1_3_2_243_2","unstructured":"Wenliang Zhong Weizhi An Feng Jiang Hehuan Ma Yuzhi Guo and Junzhou Huang. 2024. Compositional image retrieval via instruction-aware contrastive learning. arXiv:2412.05756. Retrieved from https:\/\/arxiv.org\/abs\/2412.05756"},{"key":"e_1_3_2_244_2","first-page":"6818","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhou Dewei","year":"2024","unstructured":"Dewei Zhou, You Li, Fan Ma, Xiaoting Zhang, and Yi Yang. 2024. MIGC: Multi-instance generation controller for text-to-image synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 6818\u20136828."},{"key":"e_1_3_2_245_2","first-page":"3185","volume-title":"Proceedings of the Conference of the Association for Computational Linguistics","author":"Zhou Junjie","year":"2024","unstructured":"Junjie Zhou, Zheng Liu, Shitao Xiao, Bo Zhao, and Yongping Xiong. 2024. VISTA: Visualized text embedding for universal multi-modal retrieval. In Proceedings of the Conference of the Association for Computational Linguistics. ACL, 3185\u20133200."},{"key":"e_1_3_2_246_2","doi-asserted-by":"crossref","first-page":"197","DOI":"10.1016\/j.isprsjprs.2018.01.004","article-title":"PatternNet: A benchmark dataset for performance evaluation of remote sensing image retrieval","volume":"145","author":"Zhou Weixun","year":"2018","unstructured":"Weixun Zhou, Shawn Newsam, Congmin Li, and Zhenfeng Shao. 2018. PatternNet: A benchmark dataset for performance evaluation of remote sensing image retrieval. ISPRS Journal of Photogrammetry and Remote Sensing 145 (2018), 197\u2013209.","journal-title":"ISPRS Journal of Photogrammetry and Remote Sensing"},{"key":"e_1_3_2_247_2","doi-asserted-by":"crossref","unstructured":"Yinan Zhou Yaxiong Wang Haokun Lin Chen Ma Li Zhu and Zhedong Zheng. 2025. Scale up composed image retrieval learning via modification text generation. arXiv preprint arXiv:2504.05316. Retrieved from https:\/\/arxiv.org\/abs\/2504.05316","DOI":"10.1109\/TMM.2025.3599088"},{"issue":"6","key":"e_1_3_2_248_2","first-page":"1","article-title":"AMC: Adaptive multi-expert collaborative network for text-guided image retrieval","volume":"19","author":"Zhu Hongguang","year":"2023","unstructured":"Hongguang Zhu, Yunchao Wei, Yao Zhao, Chunjie Zhang, and Shujuan Huang. 2023. AMC: Adaptive multi-expert collaborative network for text-guided image retrieval. ACM Transactions on Multimedia Computing, Communications and Applications 19, 6 (2023), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3767328","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,14]],"date-time":"2025-11-14T14:16:07Z","timestamp":1763129767000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3767328"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,14]]},"references-count":247,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1,31]]}},"alternative-id":["10.1145\/3767328"],"URL":"https:\/\/doi.org\/10.1145\/3767328","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,14]]},"assertion":[{"value":"2024-12-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-05","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}