{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T05:13:38Z","timestamp":1784178818118,"version":"3.55.0"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/501100000923","name":"Australian Research Council","doi-asserted-by":"crossref","award":["DP190101113"],"award-info":[{"award-number":["DP190101113"]}],"id":[{"id":"10.13039\/501100000923","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>\n                    Large Language Models (LLMs) are increasingly used to replace human judges to assess the relevance of information objects, raising concerns about circularity, bias, and whether simulated preferences can substitute for human judgement. This work presents experiments using multiple LLMs to label passages for relevance. It examines their gullibility\u2014how easily they are misled into labelling irrelevant passages as relevant. It also compares LLMs with human judges in ranking systems, analysing differences in discriminative power and whether some systems benefit under LLM-based evaluation. Results show that LLMs are influenced by the presence of query terms, even with irrelevant or random passages. Moreover, LLM-generated rankings are highly correlated with those of human judges, with strong agreement on which system is better in pairwise comparisons. However, LLMs may exhibit lower discriminative power, as seen in flatter ranking slopes and missed significance for meaningful improvements. Yet, there are no cases where capable LLMs and human judges reach opposing conclusions with significance. LLMs may boost traditional systems more than neural ones, adding a new concern of system bias. These findings highlight the strong potential of LLMs for relevance labelling, while also highlighting failure cases that call for careful adoption and further research to maintain evaluation integrity.\n                    <jats:xref ref-type=\"fn\">\n                      <jats:sup>1<\/jats:sup>\n                    <\/jats:xref>\n                  <\/jats:p>","DOI":"10.1145\/3788872","type":"journal-article","created":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T10:05:52Z","timestamp":1768817152000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["On the Use of LLMs for Relevance Labelling"],"prefix":"10.1145","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0008-8650","authenticated-orcid":false,"given":"Marwah","family":"Alaofi","sequence":"first","affiliation":[{"name":"Taibah University, Medina, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2425-3136","authenticated-orcid":false,"given":"Paul","family":"Thomas","sequence":"additional","affiliation":[{"name":"Microsoft, Adelaide, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9094-0810","authenticated-orcid":false,"given":"Falk","family":"Scholer","sequence":"additional","affiliation":[{"name":"RMIT University, Melbourne, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0487-9609","authenticated-orcid":false,"given":"Mark","family":"Sanderson","sequence":"additional","affiliation":[{"name":"RMIT University, Melbourne, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,4,17]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"Zahra Abbasiantaeb Chuan Meng Leif Azzopardi and Mohammad Aliannejadi. 2024. Can we use large language models to fill relevance judgment holes? arXiv:2405.05600. Retrieved from https:\/\/arxiv.org\/abs\/2405.05600"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591960"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/1390334.1390447"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3730348"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.20736\/0002002105"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1177\/001316446002000104"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/290941.291009"},{"key":"e_1_3_3_9_2","volume-title":"Proceedings of the 30th Text Retrieval Conference (TREC \u201921)","author":"Craswell Nick","year":"2021","unstructured":"Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Jimmy Lin. 2021. Overview of the TREC 2021 deep learning track. In Proceedings of the 30th Text Retrieval Conference (TREC \u201921), NIST Special Publication. Ian Soboroff and Angela Ellis (Eds.), National Institute of Standards and Technology (NIST). Retrieved from https:\/\/trec.nist.gov\/pubs\/trec30\/papers\/Overview-DL.pdf"},{"key":"e_1_3_3_10_2","volume-title":"Proceedings of the 31st Text Retrieval Conference (TREC \u201922)","volume":"500","author":"Craswell Nick","year":"2022","unstructured":"Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Jimmy Lin, Ellen M. Voorhees, and Ian Soboroff. 2022. Overview of the TREC 2022 deep learning track. In Proceedings of the 31st Text Retrieval Conference (TREC \u201922), NIST Special Publication500-338, Ian Soboroff and Angela Ellis (Eds.), National Institute of Standards and Technology (NIST). Retrieved from https:\/\/trec.nist.gov\/pubs\/trec31\/papers\/Overview_deep.pdf"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3077136.3080729"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3731120.3744588"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3578337.3605136"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3488560.3498406"},{"key":"e_1_3_3_15_2","first-page":"511","volume-title":"Nist Special Publication Sp","author":"Franz Martin","year":"1998","unstructured":"Martin Franz and Salim Roukos. 1998. Trec-6 ad-hoc retrieval. Nist Special Publication Sp., 511\u2013516."},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3336191.3371857"},{"key":"e_1_3_3_17_2","first-page":"192","volume-title":"Proceedings of the 17th Annual International ACM-SIGIR Conference on Research and Development (SIGIR \u201994)","author":"Hersh William","year":"1994","unstructured":"William Hersh, Chris Buckley, T. J. Leone, and David Hickam. 1994. OHSUMED: An interactive retrieval evaluation and new large test collection for research. In Proceedings of the 17th Annual International ACM-SIGIR Conference on Research and Development (SIGIR \u201994). Bruce W. Croft and C. J. van Rijsbergen (Eds.), Springer, 192\u2013201."},{"key":"e_1_3_3_18_2","unstructured":"Jared Kaplan Sam McCandlish Tom Henighan Tom B. Brown Benjamin Chess Rewon Child Scott Gray Alec Radford Jeffrey Wu and Dario Amodei. 2020. Scaling laws for neural language models. arXiv:2001.08361. Retrieved from https:\/\/arxiv.org\/abs\/2001.08361"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/30.1-2.81"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401075"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/1458082.1458160"},{"key":"e_1_3_3_22_2","unstructured":"Klaus Krippendorff. 2011. Computing Krippendorff\u2019s Alpha-Reliability. Retrieved from https:\/\/repository.upenn.edu\/handle\/20.500.14332\/2089"},{"key":"e_1_3_3_23_2","first-page":"71","article-title":"Computational analysis of present-day American English","author":"Ku\u010dera Henry","year":"1967","unstructured":"Henry Ku\u010dera, Winthrop Francis, William Freeman Twaddell, Mary Lois Marckworth, Laura M. Bell, and John Bissell Carroll. 1967. Computational analysis of present-day American English. International Journal of American Linguistics 35 (1967), 71\u201375.","journal-title":"International Journal of American Linguistics"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3271764"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3463238"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3592032"},{"key":"e_1_3_3_27_2","volume-title":"Proceedings of ICTIR 2020","author":"Macdonald Craig","year":"2020","unstructured":"Craig Macdonald and Nicola Tonellotto. 2020. Declarative experimentation in information retrieval using PyTerrier. In Proceedings of ICTIR 2020."},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591992"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/2407085.2407092"},{"key":"e_1_3_3_30_2","volume-title":"Proceedings of the Workshop on Cognitive Computation: Integrating Neural and Symbolic Approaches 2016 co-Located with the 30th Annual Conference on Neural Information Processing Systems (NIPS \u201916)","volume":"1773","author":"Nguyen Tri","year":"2016","unstructured":"Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A human generated machine reading comprehension dataset. In Proceedings of the Workshop on Cognitive Computation: Integrating Neural and Symbolic Approaches 2016 co-Located with the 30th Annual Conference on Neural Information Processing Systems (NIPS \u201916), Vol. 1773. Tarek Richard Besold, Antoine Bordes, Artur S. d\u2019Avila Garcez, and Greg Wayne (Eds.), Retrieved from CEUR-WS.org. https:\/\/ceur-ws.org\/Vol-1773\/CoCoNIPS_2016_paper9.pdf."},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.63"},{"key":"e_1_3_3_32_2","unstructured":"Rodrigo Frassetto Nogueira and Kyunghyun Cho. 2019. Passage re-ranking with BERT. arXiv:1901.04085. Retrieved from http:\/\/arxiv.org\/abs\/1901.04085"},{"key":"e_1_3_3_33_2","unstructured":"Rodrigo Frassetto Nogueira Wei Yang Jimmy Lin and Kyunghyun Cho. 2019. Document expansion by query prediction. arXiv:1904.08375. Retrieved from http:\/\/arxiv.org\/abs\/1904.08375"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3730221"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657942"},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3596511"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597201"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3600227"},{"key":"e_1_3_3_39_2","volume-title":"Proceeding of the Australasian Document Computing Symposium","author":"Sanderson Mark","year":"2010","unstructured":"Mark Sanderson, Falk Scholer, and Andrew Turpin. 2010. Relatively relevant: Assessor shift in document judgements. In Proceeding of the Australasian Document Computing Symposium. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:14426189"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/2009916.2010057"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.54195\/irrj.19625"},{"key":"e_1_3_3_42_2","volume-title":"Proceedings of the ACM SIGIR Conference on Human Information Interaction and Retrieval","author":"Thomas Paul","year":"2022","unstructured":"Paul Thomas, Gabriella Kazai, Ryen W. White, and Nick Craswell. 2022. The crowd is made of people: Observations from large-scale crowd labelling. In Proceedings of the ACM SIGIR Conference on Human Information Interaction and Retrieval."},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657707"},{"key":"e_1_3_3_44_2","unstructured":"Shivani Upadhyay Ehsan Kamalloo and Jimmy Lin. 2024. LLMs can patch up missing relevance judgments in evaluation. arXiv:2405.04727. Retrieved from https:\/\/arxiv.org\/abs\/2405.04727"},{"key":"e_1_3_3_45_2","unstructured":"Shivani Upadhyay Ronak Pradeep Nandan Thakur Nick Craswell and Jimmy Lin. 2024. UMBRELA: Umbrela is the (open-source reproduction of the) Bing relevance assessor. arXiv:2406.06519. Retrieved from https:\/\/arxiv.org\/abs\/2406.06519"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","unstructured":"Ellen Voorhees. 2005. Overview of the TREC 2004 robust retrieval track. In Proceedings of the 13th Text Retrieval Conference (TREC \u201904). DOI: 10.6028\/NIST.SP.500-261","DOI":"10.6028\/NIST.SP.500-261"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/290941.291017"},{"key":"e_1_3_3_48_2","unstructured":"Lee Xiong Chenyan Xiong Ye Li Kwok-Fung Tang Jialin Liu Paul Bennett Junaid Ahmed and Arnold Overwijk. 2021. Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. Retrieved from https:\/\/openreview.net\/forum?id=zeFrfgyZln"},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3411998"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3788872","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T12:00:12Z","timestamp":1776427212000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3788872"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,17]]},"references-count":48,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3788872"],"URL":"https:\/\/doi.org\/10.1145\/3788872","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,17]]},"assertion":[{"value":"2025-06-24","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-26","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-17","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}