{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T17:25:54Z","timestamp":1782840354193,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":38,"publisher":"ACM","license":[{"start":{"date-parts":[[2024,6,18]],"date-time":"2024-06-18T00:00:00Z","timestamp":1718668800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"DFG RESIRE","award":["509543643"],"award-info":[{"award-number":["509543643"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,6,18]]},"DOI":"10.1145\/3641525.3663619","type":"proceedings-article","created":{"date-parts":[[2024,7,11]],"date-time":"2024-07-11T19:02:19Z","timestamp":1720724539000},"page":"25-29","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Toward Evaluating the Reproducibility of Information Retrieval Systems with Simulated Users"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1765-2449","authenticated-orcid":false,"given":"Timo","family":"Breuer","sequence":"first","affiliation":[{"name":"TH K\u00f6ln (University of Applied Sciences), Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7001-4817","authenticated-orcid":false,"given":"Maria","family":"Maistro","sequence":"additional","affiliation":[{"name":"University of Copenhagen, Denmark"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,7,11]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Amig\u00f3 Enrique","year":"2022","unstructured":"Enrique Amig\u00f3, Pablo Castells, Julio Gonzalo, Ben Carterette, J.\u00a0Shane Culpepper, and Gabriella Kazai (Eds.). 2022. SIGIR \u201922: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain, July 11 - 15, 2022. ACM."},{"key":"e_1_3_2_1_2_1","volume-title":"SIGIR 2015 Workshop on Reproducibility, Inexplicability, and Generalizability of Results (RIGOR). In SIGIR. ACM, 1147\u20131148","author":"Arguello Jaime","year":"2015","unstructured":"Jaime Arguello, Fernando Diaz, Jimmy Lin, and Andrew Trotman. 2015. SIGIR 2015 Workshop on Reproducibility, Inexplicability, and Generalizability of Results (RIGOR). In SIGIR. ACM, 1147\u20131148."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature.2015.17433"},{"key":"e_1_3_2_1_4_1","volume-title":"User Simulation for Evaluating Information Access Systems. CoRR abs\/2306.08550","author":"Balog Krisztian","year":"2023","unstructured":"Krisztian Balog and ChengXiang Zhai. 2023. User Simulation for Evaluating Information Access Systems. CoRR abs\/2306.08550 (2023)."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723872.2723882"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Timo Breuer Nicola Ferro Norbert Fuhr Maria Maistro Tetsuya Sakai Philipp Schaer and Ian Soboroff. 2020. How to Measure the Reproducibility of System-oriented IR Experiments. In SIGIR. ACM 349\u2013358.","DOI":"10.1145\/3397271.3401036"},{"key":"e_1_3_2_1_7_1","volume-title":"ECIR (2)(Lecture Notes in Computer Science, Vol.\u00a012657)","author":"Breuer Timo","unstructured":"Timo Breuer, Nicola Ferro, Maria Maistro, and Philipp Schaer. 2021. repro_eval: A Python Interface to Reproducibility Measures of System-Oriented IR Experiments. In ECIR (2)(Lecture Notes in Computer Science, Vol.\u00a012657). Springer, 481\u2013486."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","unstructured":"Timo Breuer and Maria Maistro. 2024. Toward Evaluating the Reproducibility of Information Retrieval Systems with Simulated Users (Zenodo). https:\/\/doi.org\/10.5281\/zenodo.10931438","DOI":"10.5281\/zenodo.10931438"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"crossref","unstructured":"Olivier Chapelle Donald Metlzer Ya Zhang and Pierre Grinspan. 2009. Expected reciprocal rank for graded relevance. In CIKM. ACM 621\u2013630.","DOI":"10.1145\/1645953.1646033"},{"key":"e_1_3_2_1_10_1","volume-title":"The SIGIR 2019 Open-Source IR Replicability Challenge (OSIRRC","author":"Clancy Ryan","year":"2019","unstructured":"Ryan Clancy, Nicola Ferro, Claudia Hauff, Jimmy Lin, Tetsuya Sakai, and Ze\u00a0Zhong Wu. 2019. The SIGIR 2019 Open-Source IR Replicability Challenge (OSIRRC 2019). In SIGIR. ACM, 1432\u20131434."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Marco Ferrante Nicola Ferro and Maria Maistro. 2014. Injecting user models and time into precision via Markov chains. In SIGIR. ACM 597\u2013606.","DOI":"10.1145\/2600428.2609637"},{"key":"e_1_3_2_1_12_1","volume-title":"CLEF (Working Notes)(CEUR Workshop Proceedings, Vol.\u00a02380)","author":"Ferro Nicola","year":"2019","unstructured":"Nicola Ferro, Norbert Fuhr, Maria Maistro, Tetsuya Sakai, and Ian Soboroff. 2019. CENTRE@CLEF2019: Overview of the Replicability and Reproducibility Tasks. In CLEF (Working Notes)(CEUR Workshop Proceedings, Vol.\u00a02380). CEUR-WS.org."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3274784.3274786"},{"key":"e_1_3_2_1_14_1","unstructured":"Nicola Ferro Stefano Marchesin Alberto Purpura and Gianmaria Silvello. 2019. A Docker-Based Replicability Study of a Neural Information Retrieval Model. In OSIRRC@SIGIR(CEUR Workshop Proceedings Vol.\u00a02409). CEUR-WS.org 37\u201343."},{"key":"e_1_3_2_1_15_1","volume-title":"ECIR (2)(Lecture Notes in Computer Science, Vol.\u00a012036)","author":"Grand Adrien","unstructured":"Adrien Grand, Robert Muir, Jim Ferenczi, and Jimmy Lin. 2020. From MAXSCORE to Block-Max Wand: The Story of How Lucene Significantly Improved Query Evaluation Performance. In ECIR (2)(Lecture Notes in Computer Science, Vol.\u00a012036). Springer, 20\u201327."},{"key":"e_1_3_2_1_16_1","volume-title":"ECIR 2015, Vienna, Austria, March 29 - April 2, 2015. Proceedings. Lecture Notes in Computer Science, Vol.\u00a09022","author":"Hanbury Allan","year":"2015","unstructured":"Allan Hanbury, Gabriella Kazai, Andreas Rauber, and Norbert Fuhr (Eds.). 2015. Advances in Information Retrieval - 37th European Conference on IR Research, ECIR 2015, Vienna, Austria, March 29 - April 2, 2015. Proceedings. Lecture Notes in Computer Science, Vol.\u00a09022."},{"key":"e_1_3_2_1_17_1","volume-title":"Information Retrieval Evaluation","author":"Harman Donna","unstructured":"Donna Harman. 2011. Information Retrieval Evaluation. Morgan & Claypool Publishers."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"crossref","unstructured":"William\u00a0R. Hersh Andrew Turpin Susan Price Benjamin Chan Dale Kraemer Lynetta Sacherek and Daniel Olson. 2000. Do batch and user evaluation give the same results?. In SIGIR. ACM 17\u201324.","DOI":"10.1145\/345508.345539"},{"key":"e_1_3_2_1_19_1","volume-title":"Reproducibility in Scientific Computing. ACM Comput. Surv. 51, 3","author":"Ivie Peter","year":"2018","unstructured":"Peter Ivie and Douglas Thain. 2018. Reproducibility in Scientific Computing. ACM Comput. Surv. 51, 3 (2018), 63:1\u201363:36."},{"key":"e_1_3_2_1_20_1","volume-title":"ECIR (2)(Lecture Notes in Computer Science, Vol.\u00a012036)","author":"Kamphuis Chris","unstructured":"Chris Kamphuis, Arjen\u00a0P. de Vries, Leonid Boytsov, and Jimmy Lin. 2020. Which BM25 Do You Mean? A Large-Scale Reproducibility Study of Scoring Variants. In ECIR (2)(Lecture Notes in Computer Science, Vol.\u00a012036). Springer, 28\u201334."},{"key":"e_1_3_2_1_21_1","volume-title":"Rank Correlation Methods","author":"Kendall G.","unstructured":"M.\u00a0G. Kendall. 1948. Rank Correlation Methods. Griffin, Oxford, England."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Sahiti Labhishetty and Chengxiang Zhai. 2021. An Exploration of Tester-based Evaluation of User Simulators for Comparing Interactive Retrieval Systems. In SIGIR. ACM 1598\u20131602.","DOI":"10.1145\/3404835.3463091"},{"key":"e_1_3_2_1_23_1","volume-title":"RATE: A Reliability-Aware Tester-Based Evaluation Framework of User Simulators. In ECIR (1)(Lecture Notes in Computer Science, Vol.\u00a013185)","author":"Labhishetty Sahiti","year":"2022","unstructured":"Sahiti Labhishetty and ChengXiang Zhai. 2022. RATE: A Reliability-Aware Tester-Based Evaluation Framework of User Simulators. In ECIR (1)(Lecture Notes in Computer Science, Vol.\u00a013185). Springer, 336\u2013350."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"crossref","unstructured":"Ana Lucic Maurits J.\u00a0R. Bleeker Maarten de Rijke Koustuv Sinha Sami Jullien and Robert Stojnic. 2022. Towards Reproducible Machine Learning Research in Information Retrieval. In SIGIR. ACM 3459\u20133461.","DOI":"10.1145\/3477495.3532686"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Sean MacAvaney Andrew Yates Sergey Feldman Doug Downey Arman Cohan and Nazli Goharian. 2021. Simplified Data Wrangling with ir_datasets. In SIGIR. ACM 2429\u20132436.","DOI":"10.1145\/3404835.3463254"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"crossref","unstructured":"Craig Macdonald and Nicola Tonellotto. 2020. Declarative Experimentation in Information Retrieval using PyTerrier. In ICTIR. ACM 161\u2013168.","DOI":"10.1145\/3409256.3409829"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2023.103332"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"crossref","unstructured":"David Maxwell and Leif Azzopardi. 2016. Simulating Interactive Information Retrieval: SimIIR: A Framework for the Simulation of Interaction. In SIGIR. ACM 1141\u20131144.","DOI":"10.1145\/2911451.2911469"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1038\/nrd3439-c1"},{"key":"e_1_3_2_1_30_1","volume-title":"Robertson and Steve Walker","author":"E.","year":"1994","unstructured":"Stephen\u00a0E. Robertson and Steve Walker. 1994. Some Simple Effective Approximations to the 2-Poisson Model for Probabilistic Weighted Retrieval. In SIGIR. ACM\/Springer, 232\u2013241."},{"key":"e_1_3_2_1_31_1","volume-title":"TREC(NIST Special Publication, Vol.\u00a0500-225)","author":"Robertson E.","unstructured":"Stephen\u00a0E. Robertson, Steve Walker, Susan Jones, Micheline Hancock-Beaulieu, and Mike Gatford. 1994. Okapi at TREC-3. In TREC(NIST Special Publication, Vol.\u00a0500-225). National Institute of Standards and Technology (NIST), 109\u2013126."},{"key":"e_1_3_2_1_32_1","volume-title":"Virtual Event","author":"Santos Rodrygo","year":"2020","unstructured":"Rodrygo L.\u00a0T. Santos, Leandro\u00a0Balby Marinho, Elizabeth\u00a0M. Daly, Li Chen, Kim Falk, Noam Koenigstein, and Edleno\u00a0Silva de Moura (Eds.). 2020. RecSys 2020: Fourteenth ACM Conference on Recommender Systems, Virtual Event, Brazil, September 22-26, 2020. ACM."},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"crossref","unstructured":"Andrew Turpin and William\u00a0R. Hersh. 2001. Why Batch and User Evaluations Do Not Give the Same Results. In SIGIR. ACM 225\u2013231.","DOI":"10.1145\/383952.383992"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"crossref","unstructured":"Andrew Turpin and William\u00a0R. Hersh. 2002. User interface effects in past batch versus user experiments. In SIGIR. ACM 431\u2013432.","DOI":"10.1145\/564376.564479"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3451964.3451965"},{"key":"e_1_3_2_1_36_1","volume-title":"CORD-19: The Covid-19 Open Research Dataset. CoRR abs\/2004.10706","author":"Wang Lucy\u00a0Lu","year":"2020","unstructured":"Lucy\u00a0Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Darrin Eide, Kathryn Funk, Rodney Kinney, Ziyang Liu, William Merrill, Paul Mooney, Dewey\u00a0A. Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex\u00a0D. Wade, Kuansan Wang, Chris Wilhelm, Boya Xie, Douglas Raymond, Daniel\u00a0S. Weld, Oren Etzioni, and Sebastian Kohlmeier. 2020. CORD-19: The Covid-19 Open Research Dataset. CoRR abs\/2004.10706 (2020)."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"crossref","unstructured":"Saber Zerhoudi Sebastian G\u00fcnther Kim Plassmeier Timo Borst Christin Seifert Matthias Hagen and Michael Granitzer. 2022. The SimIIR 2.0 Framework: User Types Markov Model-Based Interaction Simulation and Advanced Query Generation. In CIKM. ACM 4661\u20134666.","DOI":"10.1145\/3511808.3557711"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3582524.3582540"}],"event":{"name":"ACM REP '24: ACM Conference on Reproducibility and Replicability","location":"Rennes France","acronym":"ACM REP '24","sponsor":["EIGREP Emerging Interest Group on Reproducibility and Replicability"]},"container-title":["Proceedings of the 2nd ACM Conference on Reproducibility and Replicability"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3641525.3663619","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3641525.3663619","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T02:04:41Z","timestamp":1755914681000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3641525.3663619"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,18]]},"references-count":38,"alternative-id":["10.1145\/3641525.3663619","10.1145\/3641525"],"URL":"https:\/\/doi.org\/10.1145\/3641525.3663619","relation":{},"subject":[],"published":{"date-parts":[[2024,6,18]]},"assertion":[{"value":"2024-07-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}