{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T14:22:23Z","timestamp":1780928543632,"version":"3.54.1"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T00:00:00Z","timestamp":1780876800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Chat-based image retrieval uses Large Language Models (LLM) to guide user input to enable more specific and precise search results, where LLM can enhance this process by asking user retrieval-oriented questions eliciting additional details about the target image. Despite the potential of this approach, no specialized Questioner model has been developed for this task due to the following significant challenges: (a) the difficulty of determining the optimal questions to ask; (b) the lack of a suitable protocol for fair model comparison; and (c) the notable scarcity of dialog-to-image retrieval data. To address these challenges, two fundamental principles are developed in this article to ensure the simplicity and effectiveness of the generated questions while enabling a fair comparison and accurate estimation of data quality and model performance. A bootstrap training methodology is introduced to collect retrieval-oriented dialog data and concurrently train the Questioner and the image Retriever. Under a fair comparison protocol, our extensive experiments have demonstrated that our proposed method can not only address the critical data gap but also achieve state-of-the-art results, which substantially surpass GPT-4o and GPT-4-Turbo through the fine-tuning of an 8B model.<\/jats:p>","DOI":"10.1145\/3807947","type":"journal-article","created":{"date-parts":[[2026,4,15]],"date-time":"2026-04-15T14:52:32Z","timestamp":1776264752000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["A Bootstrap Pipeline for Chat-Based Image Retrieval with Effective Question Generation"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7821-7056","authenticated-orcid":false,"given":"Shikai","family":"Chen","sequence":"first","affiliation":[{"name":"Lenovo Group Ltd, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-0492-3909","authenticated-orcid":false,"given":"Yicheng","family":"Jiang","sequence":"additional","affiliation":[{"name":"Southeast University, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9954-0693","authenticated-orcid":false,"given":"Jin","family":"Yuan","sequence":"additional","affiliation":[{"name":"Lenovo Group Ltd., Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7260-5098","authenticated-orcid":false,"given":"Yang","family":"Zhang","sequence":"additional","affiliation":[{"name":"Lenovo Group Ltd., Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5216-3827","authenticated-orcid":false,"given":"Zhongchao","family":"Shi","sequence":"additional","affiliation":[{"name":"Lenovo Group Ltd., Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2290-1785","authenticated-orcid":false,"given":"Jianping","family":"Fan","sequence":"additional","affiliation":[{"name":"Lenovo Group Ltd., Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7729-0622","authenticated-orcid":false,"given":"Xin","family":"Geng","sequence":"additional","affiliation":[{"name":"Southeast University, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9142-5914","authenticated-orcid":false,"given":"Yong","family":"Rui","sequence":"additional","affiliation":[{"name":"Southeast University, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,8]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et al. 2023. GPT-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_1_3_2","doi-asserted-by":"crossref","first-page":"10427","DOI":"10.1609\/aaai.v36i10.21285","article-title":"Cross-modal coherence for text-to-image retrieval","volume":"36","author":"Alikhani Malihe","year":"2022","unstructured":"Malihe Alikhani, Fangda Han, Hareesh Ravi, Mubbasir Kapadia, Vladimir Pavlovic, and Matthew Stone. 2022. Cross-modal coherence for text-to-image retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36-10, 10427\u201310435.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_1_4_2","first-page":"3952","volume-title":"Proceedings of the Computer Vision and Pattern Recognition Conference","author":"Bai Yang","year":"2025","unstructured":"Yang Bai, Yucheng Ji, Min Cao, Jinqiao Wang, and Mang Ye. 2025. Chat-based person retrieval via dialogue-refined cross-modal alignment. In Proceedings of the Computer Vision and Pattern Recognition Conference, 3952\u20133962."},{"key":"e_1_3_1_5_2","unstructured":"Y. Bai A. Jones K. Ndousse A. Askell A. Chen N. DasSarma D. Drain S. Fort D. Ganguli T. Henighan et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv:2204.05862. Retrieved from https:\/\/arxiv.org\/abs\/2204.05862"},{"key":"e_1_3_1_6_2","article-title":"Deep reinforcement learning from human preferences","volume":"30","author":"Christiano Paul F.","year":"2017","unstructured":"Paul F. Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.121"},{"key":"e_1_3_1_8_2","unstructured":"Kawin Ethayarajh Winnie Xu Niklas Muennighoff Dan Jurafsky and Douwe Kiela. 2024. KTO: Model alignment as prospect theoretic optimization. arXiv:2402.01306. Retrieved from https:\/\/arxiv.org\/abs\/2402.01306"},{"key":"e_1_3_1_9_2","unstructured":"Dongyoung Go Tomasz Korbak Germ\u00e1n Kruszewski Jos Rozen Nahyeon Ryu and Marc Dymetman. 2023. Aligning language models with preferences through f-divergence minimization. arXiv:2302.08215. Retrieved from https:\/\/arxiv.org\/abs\/2302.08215"},{"key":"e_1_3_1_10_2","unstructured":"Mandy Guo Yinfei Yang Daniel Cer Qinlan Shen and Noah Constant. 2020. MultiReQa: A cross-domain evaluation for retrieval question answering models. arXiv:2005.02507. Retrieved from https:\/\/arxiv.org\/abs\/2005.02507"},{"key":"e_1_3_1_11_2","article-title":"Dialog-based interactive image retrieval","volume":"31","author":"Guo Xiaoxiao","year":"2018","unstructured":"Xiaoxiao Guo, Hui Wu, Yu Cheng, Steven Rennie, Gerald Tesauro, and Rogerio Feris. 2018. Dialog-based interactive image retrieval. In Advances in Neural Information Processing Systems, Vol. 31.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_12_2","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv:2106.09685. Retrieved from https:\/\/arxiv.org\/abs\/2106.09685"},{"key":"e_1_3_1_13_2","first-page":"4904","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Jia Chao","year":"2021","unstructured":"Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 4904\u20134916."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.lindif.2023.102274"},{"key":"e_1_3_1_15_2","unstructured":"Dong-Jin Kim Jae Won Cho Jinsoo Choi Yunjae Jung and In So Kweon. 2021. Single-modal entropy based active learning for visual question answering. arXiv:2110.10906. Retrieved from https:\/\/arxiv.org\/abs\/2110.10906"},{"key":"e_1_3_1_16_2","article-title":"Chatting makes perfect: Chat-based image retrieval","volume":"36","author":"Levy Matan","year":"2024","unstructured":"Matan Levy, Rami Ben-Ari, Nir Darshan, and Dani Lischinski. 2024. Chatting makes perfect: Chat-based image retrieval. In Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_17_2","first-page":"19730","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the International Conference on Machine Learning. PMLR, 19730\u201319742."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_1_19_2","unstructured":"Xiao Lin and Devi Parikh. 2017. Active learning for visual question answering: An empirical study. arXiv:1711.01732. Retrieved from https:\/\/arxiv.org\/abs\/1711.01732"},{"key":"e_1_3_1_20_2","volume-title":"Advances in Neural Information Processing Systems 33 (NeurIPS 2020)","author":"Stiennon Nisan","year":"2020","unstructured":"Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. 2020. Learning to summarize from human feedback. In Advances in Neural Information Processing Systems 33 (NeurIPS 2020)."},{"key":"e_1_3_1_21_2","volume-title":"Advances in Neural Information Processing Systems","volume":"36","author":"Liu Haotian","year":"2024","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. In Advances in Neural Information Processing Systems, Vol. 36."},{"key":"e_1_3_1_22_2","unstructured":"Yuan Liu Haodong Duan Yuanhan Zhang Bo Li Songyang Zhang Wangbo Zhao Yike Yuan Jiaqi Wang Conghui He Ziwei Liu et al. 2023. MMBench: Is your multi-modal model an all-around player? arXiv:2307.06281. Retrieved from https:\/\/arxiv.org\/abs\/2307.06281"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01524"},{"key":"e_1_3_1_24_2","unstructured":"Xiaopeng Lu Tiancheng Zhao and Kyusong Lee. 2021. VisualSparta: An embarrassingly simple approach to large-scale text-to-image search with weighted bag-of-words. arXiv:2101.00265. Retrieved from https:\/\/arxiv.org\/abs\/2101.00265"},{"key":"e_1_3_1_25_2","first-page":"2421","volume-title":"Proceedings of 47th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Ma Xueguang","year":"2024","unstructured":"Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2024. Fine-tuning llama for multi-stage text retrieval. In Proceedings of 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2421\u20132425."},{"key":"e_1_3_1_26_2","first-page":"1898","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Matsumori Shoya","year":"2021","unstructured":"Shoya Matsumori, Kosuke Shingyouchi, Yuki Abe, Yosuke Fukuchi, Komei Sugiura, and Michita Imai. 2021. Unified questioner transformer for descriptive question generation in goal-oriented visual dialogue. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 1898\u20131907."},{"key":"e_1_3_1_27_2","first-page":"27730","article-title":"Training language models to follow instructions with human feedback","volume":"35","author":"Ouyang Long","year":"2022","unstructured":"Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, Vol. 35, 27730\u201327744.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_28_2","unstructured":"Aldo Pacchiano Aadirupa Saha and Jonathan Lee. 2021. Dueling RL: Reinforcement learning with trajectory preferences. arXiv:2111.04850. Retrieved from https:\/\/arxiv.org\/abs\/2111.04850"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.303"},{"key":"e_1_3_1_30_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_1_31_2","article-title":"Direct preference optimization: Your language model is secretly a reward model","volume":"36","author":"Rafailov Rafael","year":"2024","unstructured":"Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_32_2","unstructured":"Rajkumar Ramamurthy Prithviraj Ammanabrolu Kiant\u00e9 Brantley Jack Hessel Rafet Sifa Christian Bauckhage Hannaneh Hajishirzi and Yejin Choi. 2022. Is reinforcement learning (not) for natural language processing: Benchmarks baselines and building blocks for natural language policy optimization. arXiv:2210.01241. Retrieved from https:\/\/arxiv.org\/abs\/2210.01241"},{"key":"e_1_3_1_33_2","first-page":"12977","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ranasinghe Kanchana","year":"2024","unstructured":"Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo, and Tsung-Yu Lin. 2024. Learning to localize objects improves spatial reasoning in visual-LLMs. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 12977\u201312987."},{"key":"e_1_3_1_34_2","article-title":"Cola: A benchmark for compositional text-to-image retrieval","volume":"36","author":"Ray Arijit","year":"2024","unstructured":"Arijit Ray, Filip Radenovic, Abhimanyu Dubey, Bryan Plummer, Ranjay Krishna, and Kate Saenko. 2024. Cola: A benchmark for compositional text-to-image retrieval. In Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"1","key":"e_1_3_1_35_2","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1006\/jvci.1999.0413","article-title":"Image retrieval: Current techniques, promising directions, and open issues","volume":"10","author":"Rui Yong","year":"1999","unstructured":"Yong Rui, Thomas S. Huang, and Shih-Fu Chang. 1999. Image retrieval: Current techniques, promising directions, and open issues. Journal of Visual Communication and Image Representation 10, 1 (1999), 39\u201362.","journal-title":"Journal of Visual Communication and Image Representation"},{"issue":"5","key":"e_1_3_1_36_2","doi-asserted-by":"crossref","first-page":"644","DOI":"10.1109\/76.718510","article-title":"Relevance feedback: A power tool for interactive content-based image retrieval","volume":"8","author":"Rui Yong","year":"1998","unstructured":"Yong Rui, Thomas S. Huang, Michael Ortega, and Sharad Mehrotra. 1998. Relevance feedback: A power tool for interactive content-based image retrieval. IEEE Transactions on Circuits and Systems for Video Technology 8, 5 (1998), 644\u2013655.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_37_2","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv:1707.06347. Retrieved from https:\/\/arxiv.org\/abs\/1707.06347"},{"issue":"1","key":"e_1_3_1_38_2","doi-asserted-by":"crossref","first-page":"10785","DOI":"10.1038\/s41598-024-60405-y","article-title":"Evaluating the strengths and weaknesses of large language models in answering neurophysiology questions","volume":"14","author":"Shojaee-Mend Hassan","year":"2024","unstructured":"Hassan Shojaee-Mend, Reza Mohebbati, Mostafa Amiri, and Alireza Atarodi. 2024. Evaluating the strengths and weaknesses of large language models in answering neurophysiology questions. Scientific Reports 14, 1 (2024), 10785.","journal-title":"Scientific Reports"},{"key":"e_1_3_1_39_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et al. 2023. Llama: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_1_40_2","first-page":"37-2, 2644","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Wang Shijie","year":"2023","unstructured":"Shijie Wang, Jianlong Chang, Zhihui Wang, Haojie Li, Wanli Ouyang, and Qi Tian. 2023. Fine-grained retrieval prompt tuning. In Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 37-2, 2644\u20132652."},{"key":"e_1_3_1_41_2","unstructured":"Wenhao Yu Dan Iter Shuohang Wang Yichong Xu Mingxuan Ju Soumya Sanyal Chenguang Zhu Michael Zeng and Meng Jiang. 2022. Generate rather than retrieve: Large language models are strong context generators. arXiv:2209.10063. Retrieved from https:\/\/arxiv.org\/abs\/2209.10063"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3478642"},{"key":"e_1_3_1_43_2","unstructured":"Yuetong Zhao Hongyu Cao Xianyu Zhao and Zhijian Ou. 2024. An empirical study of retrieval augmented generation with chain-of-thought. arXiv:2407.15569. Retrieved from https:\/\/arxiv.org\/abs\/2407.15569"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3652583.3658032"},{"key":"e_1_3_1_45_2","unstructured":"Yutao Zhu Huaying Yuan Shuting Wang Jiongnan Liu Wenhan Liu Chenlong Deng Zhicheng Dou and Ji-Rong Wen. 2023. Large language models for information retrieval: A survey. arXiv:2308.07107. Retrieved from https:\/\/arxiv.org\/abs\/2308.07107"},{"key":"e_1_3_1_46_2","unstructured":"Daniel M. Ziegler Nisan Stiennon Jeffrey Wu Tom B. Brown Alec Radford Dario Amodei Paul Christiano and Geoffrey Irving. 2019. Fine-tuning language models from human preferences. arXiv:1909.08593. Retrieved from https:\/\/arxiv.org\/abs\/1909.08593"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3807947","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T13:35:29Z","timestamp":1780925729000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3807947"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,8]]},"references-count":45,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3807947"],"URL":"https:\/\/doi.org\/10.1145\/3807947","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,8]]},"assertion":[{"value":"2024-12-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-19","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}