{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T05:27:40Z","timestamp":1781587660033,"version":"3.54.5"},"reference-count":106,"publisher":"Association for Computing Machinery (ACM)","issue":"10","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,10,31]]},"abstract":"<jats:p>Image search is a pivotal task in multi-media and computer vision, finding applications across diverse domains, ranging from internet search to medical diagnostics. Conventional image search systems operate by accepting textual or visual queries and retrieving the top-relevant candidate results from the database. However, prevalent methods often rely on single-turn procedures, introducing potential inaccuracies and limited recall. These methods also face challenges, such as vocabulary mismatch and the semantic gap, constraining their overall effectiveness. To address these issues, we propose an interactive image retrieval system capable of refining queries based on user relevance feedback in a multi-turn setting. This system incorporates an image captioner based on a vision-language model (VLM) to enhance the quality of text-based queries, resulting in more informative queries with each iteration. Moreover, we introduce a denoiser based on a large language model (LLM) to refine text-based query expansions, mitigating inaccuracies in image descriptions generated by captioning models. To evaluate our system, we curate a new dataset by adapting the MSR-VTT and MSVD video retrieval datasets to the image retrieval task, offering multiple relevant ground-truth images for each query. Through comprehensive experiments, we validate the effectiveness of our proposed system against baseline methods, achieving state-of-the-art performance with a notable 10% improvement in terms of recall. Our contributions encompass the development of an innovative interactive image retrieval system, the integration of an LLM-based denoiser, the curation of a meticulously designed evaluation dataset, and thorough experimental validation.<\/jats:p>","DOI":"10.1145\/3744910","type":"journal-article","created":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T21:20:17Z","timestamp":1750281617000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Interactive Image Retrieval Meets Query Rewriting with Large Language and Vision Language Models"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-0298-0905","authenticated-orcid":false,"given":"Hongyi","family":"Zhu","sequence":"first","affiliation":[{"name":"University of Amsterdam, Amsterdam, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7943-2591","authenticated-orcid":false,"given":"Jia-Hong","family":"Huang","sequence":"additional","affiliation":[{"name":"University of Amsterdam, Amsterdam, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8447-872X","authenticated-orcid":false,"given":"Yixian","family":"Shen","sequence":"additional","affiliation":[{"name":"University of Amsterdam, Amsterdam, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1904-8736","authenticated-orcid":false,"given":"Stevan","family":"Rudinac","sequence":"additional","affiliation":[{"name":"University of Amsterdam, Amsterdam, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8312-0694","authenticated-orcid":false,"given":"Evangelos","family":"Kanoulas","sequence":"additional","affiliation":[{"name":"University of Amsterdam, Amsterdam, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,10,14]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"7708","volume-title":"2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ak Kenan E.","year":"2018","unstructured":"Kenan E. Ak, Ashraf Ali Kassim, Joo-Hwee Lim, and Jo Yew Tham. 2018. Learning attribute representations with localization for flexible fashion search. In 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 7708\u20137717."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/582415.582416"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/322017.322021"},{"key":"e_1_3_1_5_2","volume-title":"IEEE\/CVF International Conference on Computer Vision","author":"Baldrati Alberto","year":"2023","unstructured":"Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini, and Alberto Del Bimbo. 2023. Zero-shot composed image retrieval with textual inversion. In IEEE\/CVF International Conference on Computer Vision."},{"key":"e_1_3_1_6_2","article-title":"The IIR evaluation model: A framework for evaluation of interactive information retrieval systems","volume":"8","author":"Borlund Pia","year":"2003","unstructured":"Pia Borlund. 2003. The IIR evaluation model: A framework for evaluation of interactive information retrieval systems. Journal of Information Research 8, 3 (2003).","journal-title":"Journal of Information Research"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.5555\/2002472.2002497"},{"key":"e_1_3_1_8_2","unstructured":"Hyung Won Chung Le Hou S. Longpre Barret Zoph Yi Tay William Fedus Yunxuan Li Xuezhi Wang Mostafa Dehghani Siddhartha Brahma et al. 2022. Scaling instruction-finetuned language models. arXiv: 2210.11416. Retrieved from https:\/\/arxiv.org\/abs\/2210.11416"},{"key":"e_1_3_1_9_2","unstructured":"Wenliang Dai Junnan Li Dongxu Li Anthony Meng Huat Tiong Junqi Zhao Weisheng Wang Boyang Albert Li Pascale Fung and Steven C. H. Hoi. 2023. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. arXiv: 2305.06500. Retrieved from https:\/\/arxiv.org\/abs\/2305.06500"},{"key":"e_1_3_1_10_2","first-page":"11157","volume-title":"2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Desai Karan","year":"2020","unstructured":"Karan Desai and Justin Johnson. 2020. VirTex: Learning visual representations from textual annotations. In 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11157\u201311168."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.3166\/dn.17.1.61-84"},{"key":"e_1_3_1_12_2","volume-title":"North American Chapter of the Association for Computational Linguistics","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In North American Chapter of the Association for Computational Linguistics."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10593-2_13"},{"key":"e_1_3_1_14_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv: 2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_1_15_2","unstructured":"Wen Gang Zhou Houqiang Li and Qi Tian. 2017. Recent advance in content-based image retrieval: A literature survey. arXiv: 1706.06064. Retrieved from https:\/\/arxiv.org\/abs\/1706.06064"},{"key":"e_1_3_1_16_2","unstructured":"Xiaoxiao Guo Hui Wu Yu Cheng Steven J. Rennie and Rog\u00e9rio Schmidt Feris. 2018. Dialog-based interactive image retrieval. In 32nd International Conference on Neural Information Processing Systems 676\u2013686."},{"key":"e_1_3_1_17_2","unstructured":"Xiaoxiao Guo Hui Wu Yupeng Gao Steven J. Rennie and Rog\u00e9rio Schmidt Feris. 2019. The fashion IQ dataset: Retrieving images by combining side information and relative natural language feedback. arXiv:1905.12794. Retrieved from https:\/\/arxiv.org\/abs\/1905.12794"},{"key":"e_1_3_1_18_2","first-page":"1472","volume-title":"2017 IEEE International Conference on Computer Vision (ICCV)","author":"Han Xintong","year":"2017","unstructured":"Xintong Han, Zuxuan Wu, Phoenix X. Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, and Larry S. Davis. 2017. Automatic spatially-aware fashion concept discovery. In 2017 IEEE International Conference on Computer Vision (ICCV), 1472\u20131480."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2008.4587784"},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","first-page":"487","DOI":"10.1007\/978-3-030-98355-0_43","volume-title":"MultiMedia Modeling","author":"Hezel Nico","year":"2022","unstructured":"Nico Hezel, Konstantin Schall, Klaus Jung, and Kai Uwe Barthel. 2022. Efficient search and browsing of large-scale video collections with Vibro. In MultiMedia Modeling. Bj\u00f6rn \u00de\u00f3r J\u00f3nsson, Cathal Gurrin, Minh-Triet Tran, Duc-Tien Dang-Nguyen, Anita Min-Chun Hu, Binh Huynh Thi Thanh, and Benoit Huet (Eds.), Springer, 487\u2013492."},{"key":"e_1_3_1_21_2","first-page":"440","volume-title":"IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Hu Brian","year":"2022","unstructured":"Brian Hu, Bhavan Vasu, and Anthony Hoogs. 2022. X-mir: Explainable medical image retrieval. In IEEE\/CVF Winter Conference on Applications of Computer Vision, 440\u2013450."},{"key":"e_1_3_1_22_2","unstructured":"Hongyu Hu Jiyuan Zhang Minyi Zhao and Zhenbang Sun. 2023. CIEM: Contrastive instruction evaluation method for better instruction tuning. arXiv:2309.02301. Retrieved from https:\/\/arxiv.org\/abs\/2309.02301"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2890899"},{"key":"e_1_3_1_24_2","unstructured":"Jia-Hong Huang. 2017. Robustness Analysis of Visual Question Answering Models by Basic Questions. Master\u2019s Thesis. King Abdullah University of Science and Technology Thuwal Saudi Arabia."},{"key":"e_1_3_1_25_2","volume-title":"VQA Challenge Workshop (CVPR)","author":"Huang Jia-Hong","year":"2017","unstructured":"Jia-Hong Huang, Modar Alfadly, and Bernard Ghanem. 2017. VQABQ: Visual question answering by basic questions. VQA Challenge Workshop (CVPR)."},{"key":"e_1_3_1_26_2","volume-title":"VQA Challenge and Visual Dialog Workshop (CVPR)","author":"Huang Jia-Hong","year":"2018","unstructured":"Jia-Hong Huang, Modar Alfadly, and Bernard Ghanem. 2018. Robustness analysis of visual QA models by basic questions. VQA Challenge and Visual Dialog Workshop (CVPR)."},{"key":"e_1_3_1_27_2","unstructured":"Jia-Hong Huang Modar Alfadly Bernard Ghanem and Marcel Worring. 2019. Assessing the robustness of visual question answering. arXiv:1912.01452. Retrieved from https:\/\/arxiv.org\/abs\/1912.01452"},{"key":"e_1_3_1_28_2","unstructured":"Jia-Hong Huang Modar Alfadly Bernard Ghanem and Marcel Worring. 2023. Improving visual question answering models through robustness analysis and in-context learning with a chain of basic questions. arXiv:2304.03147. Retrieved from https:\/\/arxiv.org\/abs\/2304.03147"},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","unstructured":"Jia-Hong Huang Hongyi Zhu Yixian Shen Stevan Rudinac and Evangelos Kanoulas. 2024. Image2Text2Image: A novel framework for label-free evaluation of image-to-text generation with text-to-image diffusion models. arXiv:2411.05706. Retrieved from https:\/\/arxiv.org\/abs\/2411.05706","DOI":"10.1007\/978-981-96-2071-5_30"},{"key":"e_1_3_1_30_2","unstructured":"Jia-Hong Huang Hongyi Zhu Yixian Shen Stevan Rudinac Alessio M. Pacces and Evangelos Kanoulas. 2024. A novel evaluation framework for image2text generation. arXiv:2408.01723. Retrieved from https:\/\/arxiv.org\/abs\/2408.01723"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403305"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2008.916364"},{"key":"e_1_3_1_33_2","volume-title":"Annual International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Iwayama Makoto","year":"2000","unstructured":"Makoto Iwayama. 2000. Relevance feedback with a small number of relevance judgements: Incremental relevance feedback vs. document clustering. In the Annual International ACM SIGIR Conference on Research and Development in Information Retrieval."},{"key":"e_1_3_1_34_2","unstructured":"Rolf Jagerman Honglei Zhuang Zhen Qin Xuanhui Wang and Michael Bendersky.2023. Query expansion by prompting large language models. arXiv:2305.03653. Retrieved from https:\/\/arxiv.org\/abs\/2305.03653"},{"key":"e_1_3_1_35_2","first-page":"4904","volume-title":"International Conference on Machine Learning","author":"Jia Chao","year":"2021","unstructured":"Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning. PMLR, 4904\u20134916."},{"key":"e_1_3_1_36_2","volume-title":"5th ACM International Conference on Multimedia Retrieval","author":"Jiang Lu","year":"2015","unstructured":"Lu Jiang, Shoou-I Yu, Deyu Meng, Teruko Mitamura, and Alexander Hauptmann. 2015. Bridging the ultimate semantic gap: A semantic search engine for internet videos. In 5th ACM International Conference on Multimedia Retrieval."},{"key":"e_1_3_1_37_2","volume-title":"7th Annual ACM Workshop on the Lifelog Search Challenge","author":"Shahbaz Khan Omar","year":"2024","unstructured":"Omar Shahbaz Khan, Ujjwal Sharma, Hongyi Zhu, Stevan Rudinac, and Bj\u00f6rn \u00de\u00f3r J\u00f3nsson. 2024. Exquisitor at the lifelog search challenge 2024: Blending conversational search with user relevance feedback. In 7th Annual ACM Workshop on the Lifelog Search Challenge."},{"key":"e_1_3_1_38_2","volume-title":"International Conference on Multimedia Modeling","author":"Shahbaz Khan Omar","unstructured":"Omar Shahbaz Khan, Hongyi Zhu, Ujjwal Sharma, Evangelos Kanoulas, Stevan Rudinac, and Bj\u00f6rn \u00de\u00f3r J\u00f3nsson. 2024. Exquisitor at the video browser showdown 2024: Relevance feedback meets conversational search. In International Conference on Multimedia Modeling."},{"key":"e_1_3_1_39_2","volume-title":"Advances in Information Retrieval","author":"Shahbaz Khan Omar","year":"2020","unstructured":"Omar Shahbaz Khan, Bj\u00f6rn \u00de\u00f3r J\u00f3nsson, Stevan Rudinac, Jan Zah\u00e1lka, Hanna Ragnarsd\u00f3ttir, \u00de\u00f3rhildur \u00deorleiksd\u00f3ttir, Gylfi \u00de\u00f3r Gu\u00f0mundsson, Laurent Amsaleg, and Marcel Worring. 2020. Interactive learning for multimedia at large. In Advances in Information Retrieval."},{"key":"e_1_3_1_40_2","unstructured":"Takeshi Kojima Shixiang Shane Gu Machel Reid Yutaka Matsuo and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. arXiv:2205.11916. Retrieved from https:\/\/arxiv.org\/abs\/2205.11916"},{"key":"e_1_3_1_41_2","doi-asserted-by":"crossref","first-page":"2973","DOI":"10.1109\/CVPR.2012.6248026","volume-title":"2012 IEEE Conference on Computer Vision and Pattern Recognition","author":"Kovashka Adriana","year":"2012","unstructured":"Adriana Kovashka, Devi Parikh, and Kristen Grauman. 2012. WhittleSearch: Image search with relative attribute feedback. In 2012 IEEE Conference on Computer Vision and Pattern Recognition, 2973\u20132980."},{"key":"e_1_3_1_42_2","volume-title":"Proceedings of the 1st ACM International Conference on Multimedia Retrieval (ICMR \u201911)","author":"Larson Martha","year":"2011","unstructured":"Martha Larson, Mohammad Soleymani, Pavel Serdyukov, Stevan Rudinac, Christian Wartena, Vanessa Murdock, Gerald Friedland, Roeland Ordelman, and Gareth J. F. Jones. 2011. Automatic tagging and geotagging in video collections and communities. In Proceedings of the 1st ACM International Conference on Multimedia Retrieval (ICMR \u201911)."},{"key":"e_1_3_1_43_2","first-page":"61437","volume-title":"37th International Conference on Neural Information Processing Systems","author":"Levy Matan","year":"2024","unstructured":"Matan Levy, Rami Ben-Ari, Nir Darshan, and Dani Lischinski. 2024. Chatting makes perfect: Chat-based image retrieval. In 37th International Conference on Neural Information Processing Systems, 61437\u201361449."},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1108\/00220410810912451"},{"key":"e_1_3_1_45_2","volume-title":"International Conference on Machine Learning","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023. BLIP-2: Bootstrapping Language-Image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467101"},{"key":"e_1_3_1_47_2","unstructured":"Yifan Li Yifan Du Kun Zhou Jinpeng Wang Wayne Xin Zhao and Ji Rong Wen. 2023. Evaluating object hallucination in large vision-language models. arXiv:2305.10355. Retrieved from https:\/\/arxiv.org\/abs\/2305.10355"},{"key":"e_1_3_1_48_2","unstructured":"Yangguang Li Feng Liang Lichen Zhao Yufeng Cui Wanli Ouyang Jing Shao Fengwei Yu and Junjie Yan. 2021. Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm. arXiv:2110.05208. Retrieved from https:\/\/arxiv.org\/abs\/2110.05208"},{"key":"e_1_3_1_49_2","first-page":"5007","article-title":"Learning deep representations for ground-to-aerial geolocalization","author":"Lin Tsung-Yi","year":"2015","unstructured":"Tsung-Yi Lin, Yin Cui, Serge J. Belongie, and James Hays. 2015. Learning deep representations for ground-to-aerial geolocalization. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 5007\u20135015.","journal-title":"2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"e_1_3_1_50_2","article-title":"Microsoft COCO: Common objects in context","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In European Conference on Computer Vision.","journal-title":"European Conference on Computer Vision"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2942142"},{"key":"e_1_3_1_52_2","first-page":"26296","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu Haotian","year":"2024","unstructured":"Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved baselines with visual instruction tuning. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 26296\u201326306."},{"key":"e_1_3_1_53_2","first-page":"26296","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Loko\u010d Jakub","year":"2023","unstructured":"Jakub Loko\u010d, Zuzana Vop\u00e1lkov\u00e1, Patrik Dokoupil, and Ladislav Pe\u0161ka. 2023. Video search with CLIP and interactive text query reformulation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 26296\u201326306."},{"key":"e_1_3_1_54_2","first-page":"15692","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lu Haoyu","year":"2022","unstructured":"Haoyu Lu, Nanyi Fei, Yuqi Huo, Yizhao Gao, Zhiwu Lu, and Ji-Rong Wen. 2022. COTS: Collaborative two-stream vision-language pre-training model for cross-modal retrieval. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 15692\u201315701."},{"key":"e_1_3_1_55_2","volume-title":"Neural Information Processing Systems","author":"Lu Jiasen","year":"2019","unstructured":"Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Neural Information Processing Systems."},{"key":"e_1_3_1_56_2","unstructured":"Xiaopeng Lu Tiancheng Zhao and Kyusong Lee. 2021. VisualSparta: An embarrassingly simple approach to large-scale text-to-image search with weighted bag-of-words. arXiv:2101.00265. Retrieved from https:\/\/arxiv.org\/abs\/2101.00265"},{"key":"e_1_3_1_57_2","unstructured":"Iain Mackie Shubham Chatterjee and Jeffrey Dalton. 2023. Generative relevance feedback with large language models. arXiv:2304.13157. Retrieved from https:\/\/arxiv.org\/abs\/2304.13157"},{"key":"e_1_3_1_58_2","unstructured":"Kelong Mao Zhicheng Dou Haonan Chen Fengran Mo and Hongjin Qian. 2023. Large language models know your contextual search intent: A prompting framework for conversational search. arXiv:2303.06573. Retrieved from https:\/\/arxiv.org\/abs\/2303.06573"},{"key":"e_1_3_1_59_2","volume-title":"Findings of the Association for Computational Linguistics","author":"Mao Kelong","year":"2023","unstructured":"Kelong Mao, Zhicheng Dou, Bang Liu, Hongjin Qian, Fengran Mo, Xiangli Wu, Xiaohua Cheng, and Zhao Cao. 2023. Search-oriented conversational query editing. In Findings of the Association for Computational Linguistics. ACL."},{"key":"e_1_3_1_60_2","unstructured":"Tomas Mikolov Ilya Sutskever Kai Chen Greg Corrado and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. arXiv:1310.4546. Retrieved from https:\/\/arxiv.org\/abs\/1310.4546"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3372278.3390668"},{"key":"e_1_3_1_62_2","first-page":"3456","volume-title":"IEEE International Conference on Computer Vision","author":"Noh Hyeonwoo","year":"2017","unstructured":"Hyeonwoo Noh, Andre Araujo, Jack Sim, Tobias Weyand, and Bohyung Han. 2017. Large-scale image retrieval with attentive deep local features. In IEEE International Conference on Computer Vision, 3456\u20133465."},{"key":"e_1_3_1_63_2","article-title":"Extending faceted navigation for RDF data","author":"Oren Eyal","year":"2006","unstructured":"Eyal Oren, Renaud Delbru, and Stefan Decker. 2006. Extending faceted navigation for RDF data. In International Workshop on the Semantic Web.","journal-title":"International Workshop on the Semantic Web"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126281"},{"key":"e_1_3_1_65_2","article-title":"Deep face recognition","author":"Parkhi Omkar M.","year":"2015","unstructured":"Omkar M. Parkhi, Andrea Vedaldi, and Andrew Zisserman. 2015. Deep face recognition. In British Machine Vision Conference.","journal-title":"British Machine Vision Conference"},{"key":"e_1_3_1_66_2","unstructured":"Zhiliang Peng Wenhui Wang Li Dong Yaru Hao Shaohan Huang Shuming Ma and Furu Wei. 2023. Kosmos-2: Grounding multimodal large language models to the world. arXiv:2306.14824. Retrieved from https:\/\/arxiv.org\/abs\/2306.14824"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-022-14046-w"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0965-7"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2846566"},{"key":"e_1_3_1_70_2","volume-title":"International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning."},{"issue":"8","key":"e_1_3_1_71_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_1_72_2","first-page":"1","volume-title":"Recommender Systems Handbook","author":"Ricci Francesco","year":"2010","unstructured":"Francesco Ricci, Lior Rokach, and Bracha Shapira. 2010. Introduction to recommender systems handbook. In Recommender Systems Handbook. Springer, 1\u201335."},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1108\/eb026866"},{"key":"e_1_3_1_74_2","unstructured":"J. J. Rocchio. 1971. Relevance feedback in information retrieval. In Salton 313\u2013323."},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.2980944"},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.1007\/s13735-012-0018-0"},{"key":"e_1_3_1_77_2","volume-title":"Multimedia \u201999","author":"Rui Yong","year":"1999","unstructured":"Yong Rui and Thomas S. Huang. 1999. A novel relevance feedback technique in image retrieval. In Multimedia \u201999."},{"key":"e_1_3_1_78_2","first-page":"815","volume-title":"International Conference on Image Processing","volume":"2","author":"Rui Yong","year":"1997","unstructured":"Yong Rui, Thomas S. Huang, and Sharad Mehrotra. 1997. Content-based image retrieval with relevance feedback in MARS. In International Conference on Image Processing, Vol. 2, 815\u2013818."},{"key":"e_1_3_1_79_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539813.3545138"},{"key":"e_1_3_1_80_2","volume-title":"European Conference on Computer Vision","author":"Sariyildiz Mert Bulent","year":"2020","unstructured":"Mert Bulent Sariyildiz, Julien Perez, and Diane Larlus. 2020. Learning visual representations with caption annotations. In European Conference on Computer Vision."},{"key":"e_1_3_1_81_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"e_1_3_1_82_2","unstructured":"Christoph Schuhmann Richard Vencu Romain Beaumont Robert Kaczmarczyk Clayton Mullis Aarush Katta Theo Coombes Jenia Jitsev and Aran Komatsuzaki. 2021. LAION-400M: Open dataset of CLIP-filtered 400 million image-text pairs. arXiv:2111.02114. Retrieved from https:\/\/arxiv.org\/abs\/2111.02114"},{"key":"e_1_3_1_83_2","volume-title":"Annual International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Shen Xuehua","year":"2005","unstructured":"Xuehua Shen, Bin Tan, and Chengxiang Zhai. 2005. Context-sensitive information retrieval using implicit feedback. In Annual International ACM SIGIR Conference on Research and Development in Information Retrieval."},{"key":"e_1_3_1_84_2","doi-asserted-by":"publisher","DOI":"10.1109\/34.895972"},{"key":"e_1_3_1_85_2","first-page":"982","volume-title":"2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Sun Siqi","year":"2021","unstructured":"Siqi Sun, Yen-Chun Chen, Linjie Li, Shuohang Wang, Yuwei Fang, and Jingjing Liu. 2021. Lightningdot: Pre-training visual-semantic embeddings for real-time image-text retrieval. In 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 982\u2013997."},{"key":"e_1_3_1_86_2","doi-asserted-by":"publisher","DOI":"10.1145\/3591106.3592234"},{"key":"e_1_3_1_87_2","unstructured":"Hugo Touvron Louis Martin Kevin R. Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_1_88_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2024.3405638"},{"key":"e_1_3_1_89_2","first-page":"6432","article-title":"Composing text and image for image retrieval: An empirical odyssey","author":"Vo Nam S.","year":"2018","unstructured":"Nam S. Vo, Lu Jiang, Chen Sun, Kevin P. Murphy, Li-Jia Li, Li Fei-Fei, and James Hays. 2018. Composing text and image for image retrieval: An empirical odyssey. In 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6432\u20136441.","journal-title":"2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"e_1_3_1_90_2","volume-title":"International Conference on Multimedia Modeling","author":"Wang Shuai","unstructured":"Shuai Wang, Jiayi Shen, Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, and Marcel Worring. 2024. Prototype-enhanced hypergraph learning for heterogeneous information networks. In International Conference on Multimedia Modeling."},{"key":"e_1_3_1_91_2","first-page":"2794","volume-title":"IEEE International Conference on Computer Vision","author":"Wang Xiaolong","year":"2015","unstructured":"Xiaolong Wang and Abhinav Gupta. 2015. Unsupervised learning of visual representations using videos. In IEEE International Conference on Computer Vision, 2794\u20132802."},{"key":"e_1_3_1_92_2","volume-title":"Annual Meeting of the Association for Computational Linguistics","author":"Wang Yiming","year":"2023","unstructured":"Yiming Wang, Zhuosheng Zhang, and Rui Wang. 2023. Element-aware summarization with large language models: Expert-aligned evaluation and chain-of-thought method. In Annual Meeting of the Association for Computational Linguistics."},{"key":"e_1_3_1_93_2","volume-title":"Web Conference","author":"Wang Zhenduo","year":"2021","unstructured":"Zhenduo Wang and Qingyao Ai. 2021. Controlling the risk of conversational search via reinforcement learning. In Web Conference."},{"key":"e_1_3_1_94_2","unstructured":"Jason Wei Xuezhi Wang Dale Schuurmans Maarten Bosma Huai Hsin Chi F. Xia Quoc Le and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv:2201.11903. Retrieved from https:\/\/arxiv.org\/abs\/2201.11903"},{"key":"e_1_3_1_95_2","volume-title":"Annual International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Xu Jinxi","year":"1996","unstructured":"Jinxi Xu and W. Bruce Croft. 1996. Query expansion using local and global document analysis. In Annual International ACM SIGIR Conference on Research and Development in Information Retrieval."},{"key":"e_1_3_1_96_2","first-page":"5288","volume-title":"2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Xu Jun","year":"2016","unstructured":"Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. MSR-VTT: A large video description dataset for bridging video and language. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 5288\u20135296."},{"key":"e_1_3_1_97_2","unstructured":"Fanghua Ye Meng Fang Shenghui Li and Emine Yilmaz. 2023. Enhancing conversational search: Large language model-aided informative query rewriting. arXiv:2310.09716. Retrieved from https:\/\/arxiv.org\/abs\/2310.09716"},{"key":"e_1_3_1_98_2","volume-title":"23rd ACM International Conference on Multimedia","author":"Zah\u00e1lka Jan","year":"2015","unstructured":"Jan Zah\u00e1lka, Stevan Rudinac, and Marcel Worring. 2015. Analytic quality: Evaluation of performance and insight in multimedia collection analysis. In 23rd ACM International Conference on Multimedia."},{"issue":"3","key":"e_1_3_1_99_2","doi-asserted-by":"crossref","first-page":"687","DOI":"10.1109\/TMM.2017.2755986","article-title":"Blackthorn: Large-scale interactive multimodal learning","volume":"20","author":"Zah\u00e1lka Jan","year":"2018","unstructured":"Jan Zah\u00e1lka, Stevan Rudinac, Bj\u00f6rn \u00de\u00f3r J\u00f3nsson, Dennis C. Koelma, and Marcel Worring. 2018. Blackthorn: Large-scale interactive multimodal learning. IEEE Transactions on Multimedia 20, 3 (2018), 687\u2013698.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_100_2","unstructured":"Bohan Zhai Shijia Yang Xiangchen Zhao Chenfeng Xu Sheng Shen Dongdi Zhao Kurt Keutzer Manling Li Tan Yan and Xiangjun Fan. 2023. HallE-Switch: Rethinking and controlling object existence hallucinations in large vision language models for detailed caption. arXiv:2310.01779. Retrieved from https:\/\/arxiv.org\/abs\/10.01779"},{"key":"e_1_3_1_101_2","volume-title":"International Conference on Information and Knowledge Management","author":"Zhai ChengXiang","year":"2001","unstructured":"ChengXiang Zhai and John D. Lafferty. 2001. Model-based feedback in the language modeling approach to information retrieval. In International Conference on Information and Knowledge Management."},{"key":"e_1_3_1_102_2","first-page":"1059","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhang Ke","year":"2016","unstructured":"Ke Zhang, Wei-Lun Chao, Fei Sha, and Kristen Grauman. 2016. Summary transfer: Exemplar-based subset selection for video summarization. In IEEE Conference on Computer Vision and Pattern Recognition, 1059\u20131067."},{"key":"e_1_3_1_103_2","first-page":"2","volume-title":"7th Machine Learning for Healthcare Conference (Proceedings of Machine Learning Research, Vol. 182)","author":"Zhang Yuhao","year":"2022","unstructured":"Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D. Manning, and Curtis P. Langlotz. 2022. Contrastive learning of medical visual representations from paired images and text. In 7th Machine Learning for Healthcare Conference (Proceedings of Machine Learning Research, Vol. 182), Zachary Lipton, Rajesh Ranganath, Mark Sendak, Michael Sjoding, and Serena Yeung (Eds.). PMLR, 2\u201325."},{"key":"e_1_3_1_104_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123328"},{"key":"e_1_3_1_105_2","unstructured":"Lianmin Zheng Wei Lin Chiang Ying Sheng Siyuan Zhuang Zhanghao Wu Yonghao Zhuang Zi Lin Zhuohan Li Dacheng Li Eric P. Xing et al. 2023. Judging LLM-as-a-judge with MT-bench and chatbot arena. arXiv:2306.05685. Retrieved from https:\/\/arxiv.org\/abs\/2306.05685"},{"key":"e_1_3_1_106_2","article-title":"Enhancing interactive image retrieval with query rewriting using large language models and vision language models","author":"Zhu Hongyi","year":"2024","unstructured":"Hongyi Zhu, Jia-Hong Huang, Stevan Rudinac, and Evangelos Kanoulas. 2024. Enhancing interactive image retrieval with query rewriting using large language models and vision language models. In 2024 International Conference on Multimedia Retrieval (ICMR \u201924).","journal-title":"2024 International Conference on Multimedia Retrieval (ICMR \u201924)"},{"key":"e_1_3_1_107_2","volume-title":"2023 11th International Conference on Affective Computing and Intelligent Interaction (ACII)","author":"Zhu Hongyi","year":"2023","unstructured":"Hongyi Zhu, Yasemin Salg\u0131rl\u0131, P\u0131nar Can, Durmu\u015f At\u0131lgan, and Albert Ali Salah. 2023. Video-based estimation of pain indicators in dogs. In 2023 11th International Conference on Affective Computing and Intelligent Interaction (ACII)."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3744910","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,14]],"date-time":"2025-10-14T21:25:00Z","timestamp":1760477100000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3744910"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,14]]},"references-count":106,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2025,10,31]]}},"alternative-id":["10.1145\/3744910"],"URL":"https:\/\/doi.org\/10.1145\/3744910","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,14]]},"assertion":[{"value":"2024-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-05","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}