{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T05:13:38Z","timestamp":1784178818076,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":55,"publisher":"ACM","funder":[{"DOI":"10.13039\/100018696","name":"HORIZON EUROPE Health","doi-asserted-by":"publisher","award":["101137074"],"award-info":[{"award-number":["101137074"]}],"id":[{"id":"10.13039\/100018696","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Ministero dell\u2019Istruzione, dell\u2019Universit\u00e0 e della Ricerca","award":["2022ZLL7MW"],"award-info":[{"award-number":["2022ZLL7MW"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,11,10]]},"DOI":"10.1145\/3746252.3761200","type":"proceedings-article","created":{"date-parts":[[2025,11,8]],"date-time":"2025-11-08T01:03:27Z","timestamp":1762563807000},"page":"2115-2126","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["A Cost-Effective Framework to Evaluate LLM-Generated Relevance Judgements"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-8003-4795","authenticated-orcid":false,"given":"Simone","family":"Merlo","sequence":"first","affiliation":[{"name":"University of Padua, Padua, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0362-5893","authenticated-orcid":false,"given":"Stefano","family":"Marchesin","sequence":"additional","affiliation":[{"name":"University of Padua, Padua, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5070-2049","authenticated-orcid":false,"given":"Guglielmo","family":"Faggioli","sequence":"additional","affiliation":[{"name":"University of Padua, Padua, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9219-6239","authenticated-orcid":false,"given":"Nicola","family":"Ferro","sequence":"additional","affiliation":[{"name":"University of Padua, Padua, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,11,10]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","unstructured":"Marah I Abdin Sam Ade Jacobs Ammar Ahmad Awan Jyoti Aneja Ahmed Awadallah Hany Awadalla Nguyen Bach Amit Bahree Arash Bakhtiari Harkirat S. Behl Alon Benhaim Misha Bilenko Johan Bjorck S\u00e9bastien Bubeck Martin Cai Caio C\u00e9sar Teodoro Mendes Weizhu Chen Vishrav Chaudhary Parul Chopra Allie Del Giorno Gustavo de Rosa Matthew Dixon Ronen Eldan Dan Iter Amit Garg Abhishek Goswami Suriya Gunasekar Emman Haider Junheng Hao Russell J. Hewett Jamie Huynh Mojan Javaheripi Xin Jin Piero Kauffmann Nikos Karampatziakis Dongwoo Kim Mahoud Khademi Lev Kurilenko James R. Lee Yin Tat Lee Yuanzhi Li Chen Liang Weishung Liu Eric Lin Zeqi Lin Piyush Madan Arindam Mitra Hardik Modi Anh Nguyen Brandon Norick Barun Patra Daniel Perez-Becker Thomas Portet Reid Pryzant Heyang Qin Marko Radmilac Corby Rosset Sambudha Roy Olatunji Ruwase Olli Saarikivi Amin Saied Adil Salim Michael Santacroce Shital Shah Ning Shang Hiteshi Sharma Xia Song Masahiro Tanaka Xin Wang Rachel Ward Guanhua Wang Philipp Witte Michael Wyatt Can Xu Jiahang Xu Sonali Yadav Fan Yang Ziyi Yang Donghan Yu Chengruidong Zhang Cyril Zhang Jianwen Zhang Li Lyna Zhang Yi Zhang Yue Zhang Yunan Zhang and Xiren Zhou. 2024. Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone. CoRR Vol. abs\/2404.14219 (2024). arXiv:2404.14219 doi:10.48550\/ARXIV.2404.14219","DOI":"10.48550\/ARXIV.2404.14219"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3673791.3698431"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.6028\/NIST.SP.500-324.core-overview"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1097-0258(20000315)19:5723::AID-SIM3793.0.CO;2-A"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.2307\/2532052"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3673791.3698420"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3589335.3648327"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.18653\/V1\/2023.ACL-LONG.870"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.IPM.2012.02.003"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2912574"},{"key":"e_1_3_2_1_11_1","volume-title":"Clarke and Laura Dietz","author":"Charles L.","year":"2024","unstructured":"Charles L. A. Clarke and Laura Dietz. 2024. LLM-based relevance assessment still can't replace human relevance assessment. arXiv:2412.17156 [cs.IR] https:\/\/arxiv.org\/abs\/2412.17156"},{"key":"e_1_3_2_1_12_1","volume-title":"The Cranfield tests on index language devices","author":"Cleverdon Cyril","unstructured":"Cyril Cleverdon. 1997. The Cranfield tests on index language devices. Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 47-59."},{"key":"e_1_3_2_1_13_1","volume-title":"Sampling Techniques","author":"Cochran William G.","unstructured":"William G. Cochran. 1977. Sampling Techniques, 3rd Edition. John Wiley.","edition":"3"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1177\/001316446002000104"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.6028\/NIST.SP.1266.deep-overview"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.6028\/NIST.SP.500-335.deep-overview"},{"key":"e_1_3_2_1_17_1","volume-title":"Proceedings of the Thirty-First Text REtrieval Conference, TREC 2022","author":"Craswell Nick","year":"2022","unstructured":"Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Jimmy Lin, Ellen M. Voorhees, and Ian Soboroff. 2022. Overview of the TREC 2022 Deep Learning Track. In Proceedings of the Thirty-First Text REtrieval Conference, TREC 2022, online, November 15-19, 2022 (NIST Special Publication, Vol. 500-338), Ian Soboroff and Angela Ellis (Eds.). National Institute of Standards and Technology (NIST). https:\/\/trec.nist.gov\/pubs\/trec31\/papers\/Overview_deep.pdf"},{"key":"e_1_3_2_1_18_1","volume-title":"Overview of the TREC 2019 deep learning track. CoRR","volume":"2003","author":"Craswell Nick","year":"2020","unstructured":"Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2020b. Overview of the TREC 2019 deep learning track. CoRR, Vol. abs\/2003.07820 (2020). arXiv:2003.07820 https:\/\/arxiv.org\/abs\/2003.07820"},{"key":"e_1_3_2_1_19_1","volume-title":"Overview of the TREC 2023 Deep Learning Track. In The Thirty-Second Text REtrieval Conference Proceedings (TREC 2023","author":"Craswell Nick","year":"2023","unstructured":"Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Hossein A. Rahmani, Daniel Campos, Jimmy Lin, Ellen M. Voorhees, and Ian Soboroff. 2023. Overview of the TREC 2023 Deep Learning Track. In The Thirty-Second Text REtrieval Conference Proceedings (TREC 2023), Gaithersburg, MD, USA, November 14-17, 2023 (NIST Special Publication, Vol. 500-xxx), Ian Soboroff and Angela Ellis (Eds.). National Institute of Standards and Technology (NIST). https:\/\/trec.nist.gov\/pubs\/trec32\/papers\/Overview_deep.pdf"},{"key":"e_1_3_2_1_20_1","volume-title":"Proceedings of The First Workshop on Large Language Models for Evaluation in Information Retrieval (LLM4Eval 2024) co-located with 10th International Conference on Online Publishing (SIGIR","author":"de Jesus Gabriel","year":"2024","unstructured":"Gabriel de Jesus and S\u00e9rgio Sobral Nunes. 2024. Exploring Large Language Models for Relevance Judgments in Tetun. In Proceedings of The First Workshop on Large Language Models for Evaluation in Information Retrieval (LLM4Eval 2024) co-located with 10th International Conference on Online Publishing (SIGIR 2024), Washington D.C., USA, July 18, 2024 (CEUR Workshop Proceedings, Vol. 3752), Clemencia Siro, Mohammad Aliannejadi, Hossein A. Rahmani, Nick Craswell, Charles L. A. Clarke, Guglielmo Faggioli, Bhaskar Mitra, Paul Thomas, and Emine Yilmaz (Eds.). CEUR-WS.org, 19-30. https:\/\/ceur-ws.org\/Vol-3752\/paper2.pdf"},{"key":"e_1_3_2_1_21_1","unstructured":"Laura Dietz Oleg Zendel Peter Bailey Charles Clarke Ellese Cotterill Jeff Dalton Faegheh Hasibi Mark Sanderson and Nick Craswell. 2025. LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations. arXiv:2504.19076 [cs.IR] https:\/\/arxiv.org\/abs\/2504.19076"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","unstructured":"Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman Akhil Mathur Alan Schelten Amy Yang Angela Fan Anirudh Goyal Anthony Hartshorn Aobo Yang Archi Mitra Archie Sravankumar Artem Korenev Arthur Hinsvark Arun Rao Aston Zhang Aur\u00e9lien Rodriguez Austen Gregerson Ava Spataru Baptiste Rozi\u00e8re Bethany Biron Binh Tang Bobbie Chern Charlotte Caucheteux Chaya Nayak Chloe Bi Chris Marra Chris McConnell Christian Keller Christophe Touret Chunyang Wu Corinne Wong Cristian Canton Ferrer Cyrus Nikolaidis Damien Allonsius Daniel Song Danielle Pintz Danny Livshits David Esiobu Dhruv Choudhary Dhruv Mahajan Diego Garcia-Olano Diego Perino Dieuwke Hupkes Egor Lakomkin Ehab AlBadawy Elina Lobanova Emily Dinan Eric Michael Smith Filip Radenovic Frank Zhang Gabriel Synnaeve Gabrielle Lee Georgia Lewis Anderson Graeme Nail Gr\u00e9goire Mialon Guan Pang Guillem Cucurell Hailey Nguyen Hannah Korevaar Hu Xu Hugo Touvron Iliyan Zarov Imanol Arrieta Ibarra Isabel M. Kloumann Ishan Misra Ivan Evtimov Jade Copet Jaewon Lee Jan Geffert Jana Vranes Jason Park Jay Mahadeokar Jeet Shah Jelmer van der Linde Jennifer Billock Jenny Hong Jenya Lee Jeremy Fu Jianfeng Chi Jianyu Huang Jiawen Liu Jie Wang Jiecao Yu Joanna Bitton Joe Spisak Jongsoo Park Joseph Rocca Joshua Johnstun Joshua Saxe Junteng Jia Kalyan Vasuden Alwala Kartikeya Upasani Kate Plawiak Ke Li Kenneth Heafield Kevin Stone and et al. 2024. The Llama 3 Herd of Models. CoRR Vol. abs\/2407.21783 (2024). arXiv:2407.21783 doi:10.48550\/ARXIV.2407.21783","DOI":"10.48550\/ARXIV.2407.21783"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3578337.3605136"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-22948-1"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1037\/h0028106"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342642"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1002\/sim.4780100512"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2303.15056"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.18653\/V1\/2024.NAACL-INDUSTRY.15"},{"key":"e_1_3_2_1_30_1","volume-title":"Tanis","author":"Hogg Robert V.","year":"2015","unstructured":"Robert V. Hogg, Dale L. Zimmerman, and Elliot A. Tanis. 2015. Probability and statistical inference (ninth edition. global edition ed.). Pearson, Boston. http:\/\/search.ebscohost.com\/login.aspx?direct=true&scope=site&db=nlebk&db=nlabk&AN=1419274"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Arthur Mensch Chris Bamford Devendra Singh Chaplot Diego de Las Casas Florian Bressand Gianna Lengyel Guillaume Lample Lucile Saulnier L\u00e9lio Renard Lavaud Marie-Anne Lachaux Pierre Stock Teven Le Scao Thibaut Lavril Thomas Wang Timoth\u00e9e Lacroix and William El Sayed. 2023. Mistral 7B. CoRR Vol. abs\/2310.06825 (2023). arXiv:2310.06825 doi:10.48550\/ARXIV.2310.06825","DOI":"10.48550\/ARXIV.2310.06825"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3592032"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.14778\/3665844.3665865"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-88708-6_19"},{"key":"e_1_3_2_1_35_1","unstructured":"Tri Nguyen Mir Rosenberg Xia Song Jianfeng Gao Saurabh Tiwary Rangan Majumder and Li Deng. 2016. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. In Proceedings of the Workshop on Cognitive Computation: Integrating neural and symbolic approaches 2016 co-located with the 30th Annual Conference on Neural Information Processing Systems (NIPS 2016) Barcelona Spain December 9 2016 (CEUR Workshop Proceedings Vol. 1773) Tarek Richard Besold Antoine Bordes Artur S. d'Avila Garcez and Greg Wayne (Eds.). CEUR-WS.org. https:\/\/ceur-ws.org\/Vol-1773\/CoCoNIPS_2016_paper9.pdf"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-88708-6_9"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657992"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2412.13268"},{"key":"e_1_3_2_1_39_1","volume-title":"Proceedings of The First Workshop on Large Language Models for Evaluation in Information Retrieval (LLM4Eval 2024) co-located with 10th International Conference on Online Publishing (SIGIR","author":"Rahmani Hossein A.","year":"2024","unstructured":"Hossein A. Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra, Paul Thomas, Charles L. A. Clarke, Mohammad Aliannejadi, Clemencia Siro, and Guglielmo Faggioli. 2024c. LLMJudge: LLMs for Relevance Judgments. In Proceedings of The First Workshop on Large Language Models for Evaluation in Information Retrieval (LLM4Eval 2024) co-located with 10th International Conference on Online Publishing (SIGIR 2024), Washington D.C., USA, July 18, 2024 (CEUR Workshop Proceedings, Vol. 3752), Clemencia Siro, Mohammad Aliannejadi, Hossein A. Rahmani, Nick Craswell, Charles L. A. Clarke, Guglielmo Faggioli, Bhaskar Mitra, Paul Thomas, and Emine Yilmaz (Eds.). CEUR-WS.org, 1-3. https:\/\/ceur-ws.org\/Vol-3752\/paper8.pdf"},{"key":"e_1_3_2_1_40_1","volume-title":"Evaluating information retrieval and access tasks: NTCIR's legacy of research impact","author":"Sakai Tetsuya","unstructured":"Tetsuya Sakai, Tetsuya. Sakai, Douglas W. Oard, and Noriko. Kando. 2021 - 2021. Evaluating information retrieval and access tasks: NTCIR's legacy of research impact. Springer Nature, Singapore."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3531766"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3701716.3715488"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1177\/0962280214552881"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2409.15133"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1162\/COLI.2006.32.4.563"},{"key":"e_1_3_2_1_46_1","volume-title":"Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges. arXiv:2504","author":"Thakur Nandan","year":"2025","unstructured":"Nandan Thakur, Ronak Pradeep, Shivani Upadhyay, Daniel Campos, Nick Craswell, and Jimmy Lin. 2025. Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges. arXiv:2504.15205 [cs.CL] https:\/\/arxiv.org\/abs\/2504.15205"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657707"},{"key":"e_1_3_2_1_48_1","unstructured":"Mehmet Deniz T\u00fcrkmen Mucahid Kutlu Bahadir Altun and Gokalp Cosgun. 2025. GenTREC: The First Test Collection Generated by Large Language Models for Evaluating Information Retrieval Systems. arXiv:2501.02408 [cs.IR] https:\/\/arxiv.org\/abs\/2501.02408"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2405.04727"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2406.06519"},{"key":"e_1_3_2_1_51_1","volume-title":"Overview of the TREC 2004 Robust Retrieval Track. In TREC.","author":"Voorhees Ellen","year":"2004","unstructured":"Ellen Voorhees. 2004. Overview of the TREC 2004 Robust Retrieval Track. In TREC."},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","unstructured":"Ellen M. Voorhees. 1996. NIST TREC Disks 4 and 5: Retrieval Test Collections Document Set. doi:10.18434\/t47g6m","DOI":"10.18434\/t47g6m"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/1390334.1390346"},{"key":"e_1_3_2_1_54_1","volume-title":"Proceedings of The First Workshop on Large Language Models for Evaluation in Information Retrieval (LLM4Eval 2024) co-located with 10th International Conference on Online Publishing (SIGIR","author":"Yang Jheng-Hong","year":"2024","unstructured":"Jheng-Hong Yang and Jimmy Lin. 2024. Toward Automatic Relevance Judgment using Vision-Language Models for Image-Text Retrieval Evaluation. In Proceedings of The First Workshop on Large Language Models for Evaluation in Information Retrieval (LLM4Eval 2024) co-located with 10th International Conference on Online Publishing (SIGIR 2024), Washington D.C., USA, July 18, 2024 (CEUR Workshop Proceedings, Vol. 3752), Clemencia Siro, Mohammad Aliannejadi, Hossein A. Rahmani, Nick Craswell, Charles L. A. Clarke, Guglielmo Faggioli, Bhaskar Mitra, Paul Thomas, and Emine Yilmaz (Eds.). CEUR-WS.org, 113-123. https:\/\/ceur-ws.org\/Vol-3752\/paper7.pdf"},{"key":"e_1_3_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2304.10145"}],"event":{"name":"CIKM '25: The 34th ACM International Conference on Information and Knowledge Management","location":"Seoul Republic of Korea","acronym":"CIKM '25","sponsor":["SIGIR ACM Special Interest Group on Information Retrieval","SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web"]},"container-title":["Proceedings of the 34th ACM International Conference on Information and Knowledge Management"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746252.3761200","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,12]],"date-time":"2025-12-12T02:46:33Z","timestamp":1765507593000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746252.3761200"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,10]]},"references-count":55,"alternative-id":["10.1145\/3746252.3761200","10.1145\/3746252"],"URL":"https:\/\/doi.org\/10.1145\/3746252.3761200","relation":{},"subject":[],"published":{"date-parts":[[2025,11,10]]},"assertion":[{"value":"2025-11-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}