{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T01:47:48Z","timestamp":1787017668876,"version":"build-2736575974"},"reference-count":150,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2025,7,31]]},"abstract":"<jats:p>\n                    Query performance prediction (QPP) aims to estimate the retrieval quality of a search system for a query without human relevance judgments. Previous QPP methods typically return a single scalar value and do not require the predicted values to approximate a specific information retrieval (IR) evaluation measure, leading to certain drawbacks: (i) a single scalar is insufficient to accurately represent different IR evaluation measures, especially when metrics do not highly correlate, and (ii) a single scalar limits the interpretability of QPP methods because solely using a scalar is insufficient to explain QPP results. To address these issues, we propose a QPP framework using automatically\n                    <jats:italic toggle=\"yes\">gen<\/jats:italic>\n                    erated\n                    <jats:italic toggle=\"yes\">re<\/jats:italic>\n                    levance judgments (QPP-GenRE), which decomposes QPP into independent subtasks of predicting the relevance of each item in a ranked list to a given query. This allows us to predict any IR evaluation measure using the generated relevance judgments as pseudo-labels. This also allows us to interpret predicted IR evaluation measures, and identify, track, and rectify errors in generated relevance judgments to improve QPP quality. We predict an item\u2019s relevance by using\n                    <jats:italic toggle=\"yes\">open source<\/jats:italic>\n                    large language models (LLMs) to ensure scientific reproducibility.\n                  <\/jats:p>\n                  <jats:p>We face two main challenges: (i) excessive computational costs of judging an entire corpus for predicting a metric considering recall, and (ii) limited performance in prompting open source LLMs in a zero-\/few-shot manner. To solve the challenges, we devise an approximation strategy to predict an IR measure considering recall and propose to fine-tune open source LLMs using human-labeled relevance judgments. Experiments on the TREC 2019\u20132022 deep learning tracks and CAsT-19\u201320 datasets show that QPP-GenRE achieves state-of-the-art QPP quality for both lexical and neural rankers.<\/jats:p>","DOI":"10.1145\/3736402","type":"journal-article","created":{"date-parts":[[2025,5,19]],"date-time":"2025-05-19T06:48:36Z","timestamp":1747637316000},"page":"1-35","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":19,"title":["Query Performance Prediction Using Relevance Judgments Generated by Large Language Models"],"prefix":"10.1145","volume":"43","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1434-7596","authenticated-orcid":false,"given":"Chuan","family":"Meng","sequence":"first","affiliation":[{"name":"University of Amsterdam, Amsterdam, Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4411-7089","authenticated-orcid":false,"given":"Negar","family":"Arabzadeh","sequence":"additional","affiliation":[{"name":"University of Waterloo, Waterloo, Ontario, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4712-832X","authenticated-orcid":false,"given":"Arian","family":"Askari","sequence":"additional","affiliation":[{"name":"Leiden University, Leiden, Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9447-4172","authenticated-orcid":false,"given":"Mohammad","family":"Aliannejadi","sequence":"additional","affiliation":[{"name":"University of Amsterdam, Amsterdam, Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1086-0202","authenticated-orcid":false,"given":"Maarten de","family":"Rijke","sequence":"additional","affiliation":[{"name":"University of Amsterdam, Amsterdam, Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,7,12]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"Zahra Abbasiantaeb Chuan Meng Leif Azzopardi and Mohammad Aliannejadi. 2024. Can we use large language models to fill relevance judgment holes? arXiv:2405.05600. Retrieved from https:\/\/arxiv.org\/abs\/2405.05600"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-88708-6_13"},{"key":"e_1_3_3_4_2","volume-title":"TREC","author":"Abbasiantaeb Zahra","year":"2023","unstructured":"Zahra Abbasiantaeb, Chuan Meng, David Rau, Antonis Krasakis, Hossein A. Rahmani, and Mohammad Aliannejadi. 2023. LLM-based retrieval and generation pipelines for TREC interactive knowledge assistance track (iKAT). In TREC."},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-24752-4_10"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-72240-1_15"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3583780.3615270"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482063"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-56069-9_51"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3673791.3698438"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3701551.3703480"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.623"},{"key":"e_1_3_3_13_2","unstructured":"Arian Askari Chuan Meng Mohammad Aliannejadi Zhaochun Ren Evangelos Kanoulas and Suzan Verberne. 2024. Generative retrieval with few-shot indexing. arXiv:2408.02152. Retrieved from https:\/\/arxiv.org\/abs\/2408.02152"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.findings-naacl.357"},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-71496-5_20"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.1111\/nyas.15007"},{"key":"e_1_3_3_17_2","first-page":"1877","volume-title":"NeurIPS","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et\u00a0al. 2020. Language models are few-shot learners. In NeurIPS, 1877\u20131901."},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.2200\/S00235ED1V01Y201004ICR015"},{"key":"e_1_3_3_19_2","unstructured":"Jerry Chee Yaohui Cai Volodymyr Kuleshov and Christopher De Sa. 2023. QuIP: 2-Bit quantization of large language models with guarantees. arXiv:2307.13304. Retrieved from https:\/\/arxiv.org\/abs\/2307.13304"},{"key":"e_1_3_3_20_2","doi-asserted-by":"crossref","unstructured":"Nuo Chen Jiqun Liu Xiaoyu Dong Qijiong Liu Tetsuya Sakai and Xiao-Ming Wu. 2024. AI can be cognitively biased: An exploratory study on threshold priming in LLM-based batch relevance assessment. arXiv:2409.16022. Retrieved from https:\/\/arxiv.org\/abs\/2409.16022","DOI":"10.1145\/3673791.3698420"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-99739-7_8"},{"key":"e_1_3_3_22_2","unstructured":"Hyung Won Chung Le Hou Shayne Longpre Barret Zoph Yi Tay William Fedus Eric Li Xuezhi Wang Mostafa Dehghani Siddhartha Brahma et\u00a0al. 2022. Scaling instruction-finetuned language models. arXiv:2210.11416. Retrieved from https:\/\/arxiv.org\/abs\/2210.11416"},{"key":"e_1_3_3_23_2","volume-title":"TREC","author":"Craswell Nick","year":"2020","unstructured":"Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2020. Overview of the TREC 2020 deep learning track. In TREC."},{"key":"e_1_3_3_24_2","volume-title":"TREC","author":"Craswell Nick","year":"2019","unstructured":"Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2019. Overview of the TREC 2019 deep learning track. In TREC."},{"key":"e_1_3_3_25_2","volume-title":"TREC","author":"Craswell Nick","year":"2021","unstructured":"Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Fernando Campos, and Jimmy Lin. 2021. Overview of the TREC 2021 deep learning track. In TREC."},{"key":"e_1_3_3_26_2","volume-title":"TREC","author":"Craswell Nick","year":"2022","unstructured":"Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Fernando Campos, Jimmy Lin, Ellen M. Voorhees, and Ian Soboroff. 2022. Overview of the TREC 2022 deep learning track. In TREC."},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/564376.564429"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/2009916.2010063"},{"key":"e_1_3_3_29_2","volume-title":"TREC","author":"Dalton Jeffrey","year":"2020","unstructured":"Jeffrey Dalton, Chenyan Xiong, and Jamie Callan. 2020. CAsT 2020: The conversational assistance track overview. In TREC."},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401206"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3488560.3498491"},{"key":"e_1_3_3_32_2","first-page":"1","volume-title":"ACM Transactions on Information Systems","volume":"41","author":"Datta Suchana","year":"2022","unstructured":"Suchana Datta, Debasis Ganguly, Mandar Mitra, and Derek Greene. 2022. A relative information gain-based query performance prediction framework with generated query variants. ACM Transactions on Information Systems 41 (2022), 1\u201331."},{"key":"e_1_3_3_33_2","first-page":"2148","volume-title":"SIGIR","author":"Datta Suchana","year":"2022","unstructured":"Suchana Datta, Sean MacAvaney, Debasis Ganguly, and Derek Greene. 2022. A \u2018pointwise-query, listwise-document\u2019 based query performance prediction approach. In SIGIR, 2148\u20132153."},{"key":"e_1_3_3_34_2","unstructured":"Tim Dettmers Artidoro Pagnoni Ari Holtzman and Luke Zettlemoyer. 2023. QLoRA: Efficient finetuning of quantized LLMs. arXiv:2305.14314. Retrieved from https:\/\/arxiv.org\/abs\/2305.14314"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/2983323.2983894"},{"key":"e_1_3_3_36_2","first-page":"4171","volume-title":"NAACL","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 4171\u20134186."},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.3390\/app11199075"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/1277741.1277841"},{"key":"e_1_3_3_39_2","unstructured":"Qingxiu Dong Lei Li Damai Dai Ce Zheng Zhiyong Wu Baobao Chang Xu Sun Jingjing Xu and Zhifang Sui. 2022. A survey for in-context learning. arXiv:2301.00234. Retrieved from https:\/\/arxiv.org\/abs\/2301.00234"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.950"},{"key":"e_1_3_3_41_2","first-page":"39","volume-title":"ICTIR","author":"Faggioli Guglielmo","unstructured":"Guglielmo Faggioli, Laura Dietz, Charles L. A. Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, et\u00a0al. 2023. Perspectives on large language models for relevance judgment. In ICTIR, 39\u201350."},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3463090"},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3636341.3636356"},{"key":"e_1_3_3_44_2","first-page":"41","volume-title":"IIR","author":"Faggioli Guglielmo","year":"2023","unstructured":"Guglielmo Faggioli, Nicola Ferro, Cristina Muntean, Raffaele Perego, and Nicola Tonellotto. 2023. A spatial approach to predict performance of conversational search systems. In IIR, 41\u201346."},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591625"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3578337.3605142"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-28244-7_15"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-72113-8_8"},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-28244-7_20"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-99736-6_15"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3592046"},{"key":"e_1_3_3_52_2","unstructured":"Aryo Gema Luke Daines Pasquale Minervini and Beatrice\u00a0Alex. 2023. Parameter-efficient fine-tuning of LLaMA for the clinical domain. arXiv:2307.03042. Retrieved from https:\/\/arxiv.org\/abs\/2307.03042"},{"key":"e_1_3_3_53_2","doi-asserted-by":"crossref","unstructured":"Fabrizio Gilardi Meysam Alizadeh and Ma\u00ebl Kubli. 2023. ChatGPT outperforms crowd-workers for text-annotation tasks. arXiv:2303.15056. Retrieved from https:\/\/arxiv.org\/abs\/2303.15056","DOI":"10.1073\/pnas.2305016120"},{"key":"e_1_3_3_54_2","unstructured":"Yuxian Gu Li Dong Furu Wei and Minlie Huang. 2023. Knowledge distillation of large language models. arXiv:2306.08543. Retrieved from https:\/\/arxiv.org\/abs\/2306.08543"},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-15712-8_41"},{"key":"e_1_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3341981.3344249"},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/1458082.1458311"},{"key":"e_1_3_3_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462891"},{"key":"e_1_3_3_59_2","unstructured":"Yupeng Hou Junjie Zhang Zihan Lin Hongyu Lu Ruobing Xie Julian McAuley and Wayne Xin Zhao. 2023. Large language models are zero-shot rankers for recommender systems. arXiv:2305.08845. Retrieved from https:\/\/arxiv.org\/abs\/2305.08845"},{"key":"e_1_3_3_60_2","volume-title":"ICLR","author":"Hu Edward J.","year":"2021","unstructured":"Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2021. LoRA: Low-rank adaptation of large language models. In ICLR."},{"key":"e_1_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/582415.582418"},{"key":"e_1_3_3_62_2","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Arthur Mensch Chris Bamford Devendra Singh Chaplot Diego de las Casas Florian Bressand Gianna Lengyel Guillaume Lample Lucile Saulnier et\u00a0al. 2023. Mistral 7B. arXiv:2310.06825. Retrieved from https:\/\/arxiv.org\/abs\/2310.06825"},{"key":"e_1_3_3_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/2766462.2767824"},{"key":"e_1_3_3_64_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2023.119015"},{"key":"e_1_3_3_65_2","doi-asserted-by":"crossref","unstructured":"Ekaterina Khramtsova Shengyao Zhuang Mahsa Baktashmotlagh and Guido Zuccon. 2024. Leveraging LLMs for unsupervised dense retriever ranking. arXiv:2402.04853. Retrieved from https:\/\/arxiv.org\/abs\/2402.04853","DOI":"10.1145\/3626772.3657798"},{"key":"e_1_3_3_66_2","volume-title":"ICLR","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In ICLR."},{"key":"e_1_3_3_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/383952.383970"},{"key":"e_1_3_3_68_2","volume-title":"SIGIR","author":"Lassance Carlos","year":"2023","unstructured":"Carlos Lassance and St\u00e9phane Clinchant. 2023. The tale of two MSMARCO\u2014and their unfair comparisons. In SIGIR."},{"key":"e_1_3_3_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/383952.383972"},{"key":"e_1_3_3_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3463238"},{"key":"e_1_3_3_71_2","volume-title":"NeurIPS","author":"Liu Haokun","year":"2022","unstructured":"Haokun Liu, Derek Tam, Muqeeth Mohammed, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. 2022. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. In NeurIPS."},{"key":"e_1_3_3_72_2","unstructured":"Tiedong Liu and Bryan Kian Hsiang Low. 2023. Goat: Fine-tuned LLaMA outperforms GPT-4 on arithmetic tasks. arXiv:2305.14201. Retrieved from https:\/\/arxiv.org\/abs\/2305.14201"},{"key":"e_1_3_3_73_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-88708-6_25"},{"key":"e_1_3_3_74_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-016-9282-6"},{"key":"e_1_3_3_75_2","unstructured":"Yadong Lu Chunyuan Li Haotian Liu Jianwei Yang Jianfeng Gao and Yelong Shen. 2023. An empirical study of scaling instruct-tuned large multimodal models. arXiv:2309.09958. Retrieved from https:\/\/arxiv.org\/abs\/2309.09958"},{"key":"e_1_3_3_76_2","unstructured":"Shengjie Ma Chong Chen Qi Chu and Jiaxin Mao. 2024. Leveraging large language models for relevance judgments in legal case retrieval. arXiv:2403.18405. Retrieved from https:\/\/arxiv.org\/abs\/2403.18405"},{"key":"e_1_3_3_77_2","unstructured":"Xueguang Ma Liang Wang Nan Yang Furu Wei and Jimmy Lin. 2023. Fine-tuning LLaMA for multi-stage text retrieval. arXiv:2310.08319. Retrieved from https:\/\/arxiv.org\/abs\/2310.08319"},{"key":"e_1_3_3_78_2","unstructured":"Xueguang Ma Xinyu Zhang Ronak Pradeep and Jimmy Lin. 2023. Zero-shot listwise document reranking with a large language model. arXiv:2305.02156. Retrieved from https:\/\/arxiv.org\/abs\/2305.02156"},{"key":"e_1_3_3_79_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3463250"},{"key":"e_1_3_3_80_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3592032"},{"key":"e_1_3_3_81_2","doi-asserted-by":"publisher","DOI":"10.1145\/2396761.2398691"},{"key":"e_1_3_3_82_2","doi-asserted-by":"publisher","DOI":"10.1109\/DEXA.2017.38"},{"key":"e_1_3_3_83_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDIM.2016.7829763"},{"key":"e_1_3_3_84_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657658"},{"key":"e_1_3_3_85_2","first-page":"25","volume-title":"QPP++2023","author":"Meng Chuan","year":"2023","unstructured":"Chuan Meng, Mohammad Aliannejadi, and Maarten de Rijke. 2023. Performance prediction for conversational search using perplexities of query rewrites. In QPP++2023, 25\u201328."},{"key":"e_1_3_3_86_2","doi-asserted-by":"publisher","DOI":"10.1145\/3583780.3615070"},{"key":"e_1_3_3_87_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591919"},{"key":"e_1_3_3_88_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657864"},{"key":"e_1_3_3_89_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-88720-8_49"},{"key":"e_1_3_3_90_2","volume-title":"SIGIR","author":"Meng Chuan","year":"2025","unstructured":"Chuan Meng, Francesco Tonolini, Fengran Mo, Nikolaos Aletras, Emine Yilmaz, and Gabriella Kazai. 2025. Bridging the gap: From ad-hoc to proactive search in conversations. In SIGIR."},{"key":"e_1_3_3_91_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210146"},{"key":"e_1_3_3_92_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.135"},{"key":"e_1_3_3_93_2","unstructured":"Fengran Mo Kelong Mao Ziliang Zhao Hongjin Qian Haonan Chen Yiruo Cheng Xiaoxi Li Yutao Zhu Zhicheng Dou and Jian-Yun Nie. 2024. A survey of conversational search. arXiv:2410.15576. Retrieved from https:\/\/arxiv.org\/abs\/2410.15576"},{"key":"e_1_3_3_94_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.274"},{"key":"e_1_3_3_95_2","volume-title":"SIGIR","author":"Mo Fengran","year":"2025","unstructured":"Fengran Mo, Chuan Meng, Mohammad Aliannejadi, and Jian-Yun Nie. 2025. Conversational search: From fundamentals to frontiers in the LLM era. In SIGIR."},{"key":"e_1_3_3_96_2","doi-asserted-by":"publisher","DOI":"10.1145\/3580305.3599411"},{"key":"e_1_3_3_97_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2017.2754371"},{"key":"e_1_3_3_98_2","doi-asserted-by":"publisher","DOI":"10.1145\/860435.860510"},{"key":"e_1_3_3_99_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2005.03.023"},{"key":"e_1_3_3_100_2","first-page":"207","volume-title":"SPIRE","author":"P\u00e9rez-Iglesias Joaqu\u00edn","year":"2010","unstructured":"Joaqu\u00edn P\u00e9rez-Iglesias and Lourdes Araujo. 2010. Standard deviation as a query hardness estimator. In SPIRE, 207\u2013212."},{"key":"e_1_3_3_101_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591901"},{"key":"e_1_3_3_102_2","unstructured":"Ronak Pradeep Rodrigo Nogueira and Jimmy Lin. 2021. The expando-mono-duo design pattern for text ranking with pretrained sequence-to-sequence models. arXiv:2101.05667. Retrieved from https:\/\/arxiv.org\/abs\/2101.05667"},{"key":"e_1_3_3_103_2","unstructured":"Ronak Pradeep Sahel Sharifymoghaddam and Jimmy Lin. 2023. RankVicuna: Zero-shot listwise document reranking with open-source large language models. arXiv:2309.15088. Retrieved from https:\/\/arxiv.org\/abs\/2309.15088"},{"key":"e_1_3_3_104_2","unstructured":"Ronak Pradeep Sahel Sharifymoghaddam and Jimmy Lin. 2023. RankZephyr: Effective and robust zero-shot listwise reranking is a breeze! arXiv:2312.02724. Retrieved from https:\/\/arxiv.org\/abs\/2312.02724"},{"key":"e_1_3_3_105_2","doi-asserted-by":"crossref","unstructured":"Zhen Qin Rolf Jagerman Kai Hui Honglei Zhuang Junru Wu Jiaming Shen Tianqi Liu Jialu Liu Donald Metzler Xuanhui Wang et\u00a0al. 2023. Large language models are effective text rankers with pairwise ranking prompting. arXiv:2306.17563. Retrieved from https:\/\/arxiv.org\/abs\/2306.17563","DOI":"10.18653\/v1\/2024.findings-naacl.97"},{"key":"e_1_3_3_106_2","doi-asserted-by":"publisher","DOI":"10.1108\/AJIM-03-2015-0046"},{"key":"e_1_3_3_107_2","first-page":"109","article-title":"Okapi at TREC-3","volume":"109","author":"Robertson Stephen E.","year":"1995","unstructured":"Stephen E. Robertson, Steve Walker, Susan Jones, Micheline M. Hancock-Beaulieu, and Mike Gatford. 1995. Okapi at TREC-3. Nist Special Publication Sp 109 (1995), 109.","journal-title":"Nist Special Publication Sp"},{"key":"e_1_3_3_108_2","doi-asserted-by":"publisher","DOI":"10.1145\/3077136.3080665"},{"key":"e_1_3_3_109_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.249"},{"key":"e_1_3_3_110_2","doi-asserted-by":"crossref","unstructured":"Alireza Salemi and Hamed Zamani. 2024. Evaluating retrieval quality in retrieval-augmented generation. arXiv:2404.13781. Retrieved from https:\/\/arxiv.org\/abs\/2404.13781","DOI":"10.1145\/3626772.3657957"},{"key":"e_1_3_3_111_2","unstructured":"Mohammadreza Samadi and Davood Rafiei. 2023. Performance prediction for multi-hop questions. arXiv:2308.06431. Retrieved from https:\/\/arxiv.org\/abs\/2308.06431"},{"key":"e_1_3_3_112_2","unstructured":"Andrea Santilli and Emanuele Rodol\u00e0. 2023. Camoscio: An Italian instruction-tuned Llama. arXiv:2307.16456. Retrieved from https:\/\/arxiv.org\/abs\/2307.16456"},{"key":"e_1_3_3_113_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210078"},{"key":"e_1_3_3_114_2","doi-asserted-by":"publisher","DOI":"10.1145\/1835449.1835494"},{"key":"e_1_3_3_115_2","doi-asserted-by":"publisher","DOI":"10.1145\/2180868.2180873"},{"key":"e_1_3_3_116_2","first-page":"2486","volume-title":"SIGIR","author":"Singh Ashutosh","year":"2023","unstructured":"Ashutosh Singh, Debasis Ganguly, Suchana Datta, and Craig McDonald. 2023. Unsupervised query performance prediction for neural models utilising pairwise rank preferences. In SIGIR, 2486\u20132490."},{"key":"e_1_3_3_117_2","doi-asserted-by":"publisher","DOI":"10.1145\/383952.383961"},{"key":"e_1_3_3_118_2","unstructured":"Jiuding Sun Chantal Shaib and Byron C. Wallace. 2023. Evaluating the zero-shot robustness of instruction-tuned language models. arXiv:2306.11270. Retrieved from https:\/\/arxiv.org\/abs\/2306.11270"},{"key":"e_1_3_3_119_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462883"},{"key":"e_1_3_3_120_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.923"},{"key":"e_1_3_3_121_2","doi-asserted-by":"crossref","unstructured":"Rikiya Takehi Ellen M. Voorhees and Tetsuya Sakai. 2024. LLM-assisted relevance assessments: When should we ask LLMs for help? arXiv:2411.06877. Retrieved from https:\/\/arxiv.org\/abs\/2411.06877","DOI":"10.1145\/3726302.3729916"},{"key":"e_1_3_3_122_2","unstructured":"Raphael Tang Xinyu Zhang Xueguang Ma Jimmy Lin and Ferhan Ture. 2023. Found in the middle: Permutation self-consistency improves listwise ranking in large language models. arXiv:2310.07712. Retrieved from https:\/\/arxiv.org\/abs\/2310.07712"},{"key":"e_1_3_3_123_2","doi-asserted-by":"publisher","DOI":"10.1145\/2661829.2661906"},{"key":"e_1_3_3_124_2","volume-title":"ICLR","author":"Tay Yi","year":"2022","unstructured":"Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Steven Zheng, et\u00a0al. 2022. Ul2: Unifying language learning paradigms. In ICLR."},{"key":"e_1_3_3_125_2","doi-asserted-by":"publisher","DOI":"10.1145\/3166072.3166079"},{"key":"e_1_3_3_126_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657707"},{"key":"e_1_3_3_127_2","volume-title":"TREC","author":"Tomlinson Stephen","year":"2007","unstructured":"Stephen Tomlinson, Douglas W. Oard, Jason R. Baron, and Paul Thompson. 2007. Overview of the TREC 2007 legal track. In TREC."},{"key":"e_1_3_3_128_2","doi-asserted-by":"publisher","DOI":"10.1145\/2433396.2433407"},{"key":"e_1_3_3_129_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et\u00a0al. 2023. LLaMA: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_3_130_2","unstructured":"Shivani Upadhyay Ehsan Kamalloo and Jimmy Lin. 2024. LLMs can patch up missing relevance judgments in evaluation. arXiv:2405.04727. Retrieved from https:\/\/arxiv.org\/abs\/2405.04727"},{"key":"e_1_3_3_131_2","unstructured":"Shivani Upadhyay Ronak Pradeep Nandan Thakur Daniel Campos Nick Craswell Ian Soboroff Hoa Trang Dang and Jimmy Lin. 2024. A large-scale study of relevance assessments with large language models: An initial look. arXiv:2411.08275. Retrieved from https:\/\/arxiv.org\/abs\/2411.08275"},{"key":"e_1_3_3_132_2","unstructured":"Shivani Upadhyay Ronak Pradeep Nandan Thakur Nick Craswell and Jimmy Lin. 2024. UMBRELA: UMbrela is the (open-source reproduction of the) Bing RELevance Assessor. arXiv:2406.06519. Retrieved from https:\/\/arxiv.org\/abs\/2406.06519"},{"key":"e_1_3_3_133_2","unstructured":"Maria Vlachou and Craig Macdonald. 2023. On coherence-based predictors for dense query performance prediction. arXiv:2310.11405. Retrieved from https:\/\/arxiv.org\/abs\/2310.11405"},{"key":"e_1_3_3_134_2","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS, Vol. 35, 24824\u201324837.","journal-title":"NeurIPS"},{"key":"e_1_3_3_135_2","volume-title":"ICLR","author":"Xiong Lee","year":"2021","unstructured":"Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate nearest neighbor negative contrastive learning for dense text retrieval. In ICLR."},{"key":"e_1_3_3_136_2","unstructured":"Mingxue Xu Yao Lei Xu and Danilo P. Mandic. 2023. TensorGPT: Efficient compression of the embedding layer in LLMs based on the tensor-train decomposition. arXiv:2307.00526. Retrieved from https:\/\/arxiv.org\/abs\/2307.00526"},{"key":"e_1_3_3_137_2","doi-asserted-by":"crossref","unstructured":"Le Yan Zhen Qin Honglei Zhuang Rolf Jagerman Xuanhui Wang Michael Bendersky and Harrie Oosterhuis. 2024. Consolidating ranking and relevance predictions of large language models through post-processing. arXiv:2404.11791. Retrieved from https:\/\/arxiv.org\/abs\/2404.11791","DOI":"10.18653\/v1\/2024.emnlp-main.25"},{"key":"e_1_3_3_138_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462856"},{"key":"e_1_3_3_139_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210041"},{"key":"e_1_3_3_140_2","first-page":"37","volume-title":"QPP++2023","author":"Zendel Oleg","year":"2023","unstructured":"Oleg Zendel, Binsheng Liu, J. Shane Culpepper, and Falk Scholer. 2023. Entropy-based query performance prediction for neural information retrieval systems. In QPP++2023, 37\u201344."},{"key":"e_1_3_3_141_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657784"},{"key":"e_1_3_3_142_2","unstructured":"Shengyu Zhang Linfeng Dong Xiaoya Li Sen Zhang Xiaofei Sun Shuhe Wang Jiwei Li Runyi Hu Tianwei Zhang Fei Wu et\u00a0al. 2023. Instruction tuning for large language models: A survey. arXiv:2308.10792. Retrieved from https:\/\/arxiv.org\/abs\/2308.10792"},{"key":"e_1_3_3_143_2","unstructured":"Xinyu Zhang Sebastian Hofst\u00e4tter Patrick Lewis Raphael Tang and Jimmy Lin. 2023. Rank-without-GPT: Building GPT-independent listwise rerankers on open-source large language models. arXiv:2312.02969. Retrieved from https:\/\/arxiv.org\/abs\/2312.02969"},{"key":"e_1_3_3_144_2","unstructured":"Yue Zhang Leyang Cui Deng Cai Xinting Huang Tao Fang and Wei Bi. 2023. Multi-task instruction tuning of LLaMa for specific scenarios: A preliminary study on writing assistance. arXiv:2305.13225. Retrieved from https:\/\/arxiv.org\/abs\/2305.13225"},{"key":"e_1_3_3_145_2","article-title":"Judging LLM-as-a-judge with MT-bench and chatbot arena. In","author":"Zheng Lianmin","year":"2024","unstructured":"Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et\u00a0al. 2024. Judging LLM-as-a-judge with MT-bench and chatbot arena. In NeurIPS, Vol. 36.","journal-title":"NeurIPS"},{"key":"e_1_3_3_146_2","doi-asserted-by":"publisher","DOI":"10.1145\/1183614.1183696"},{"key":"e_1_3_3_147_2","doi-asserted-by":"publisher","DOI":"10.1145\/1277741.1277835"},{"key":"e_1_3_3_148_2","unstructured":"Yutao Zhu Huaying Yuan Shuting Wang Jiongnan Liu Wenhan Liu Chenlong Deng Zhicheng Dou and Ji-Rong Wen. 2023. Large language models for information retrieval: A survey. arXiv:2308.07107. Retrieved from https:\/\/arxiv.org\/abs\/2308.07107"},{"key":"e_1_3_3_149_2","doi-asserted-by":"crossref","unstructured":"Honglei Zhuang Zhen Qin Kai Hui Junru Wu Le Yan Xuanhui Wang and Michael Berdersky. 2023. Beyond yes and no: Improving zero-shot LLM rankers via scoring fine-grained relevance labels. arXiv:2310.14122. Retrieved from https:\/\/arxiv.org\/abs\/2310.14122","DOI":"10.18653\/v1\/2024.naacl-short.31"},{"key":"e_1_3_3_150_2","doi-asserted-by":"crossref","unstructured":"Shengyao Zhuang Bing Liu Bevan Koopman and Guido Zuccon. 2023. Open-source large language models are strong zero-shot query likelihood models for document ranking. arXiv:2310.13243. Retrieved from https:\/\/arxiv.org\/abs\/2310.13243","DOI":"10.18653\/v1\/2023.findings-emnlp.590"},{"key":"e_1_3_3_151_2","unstructured":"Shengyao Zhuang Honglei Zhuang Bevan Koopman and Guido Zuccon. 2023. A setwise approach for effective and highly efficient zero-shot ranking with large language models. arXiv:2310.09497. Retrieved from https:\/\/arxiv.org\/abs\/2310.09497"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3736402","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,17]],"date-time":"2025-07-17T10:29:31Z","timestamp":1752748171000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3736402"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,12]]},"references-count":150,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,7,31]]}},"alternative-id":["10.1145\/3736402"],"URL":"https:\/\/doi.org\/10.1145\/3736402","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,12]]},"assertion":[{"value":"2024-06-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-10","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}