{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T02:27:57Z","timestamp":1783736877371,"version":"3.55.0"},"reference-count":46,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2022,6,1]],"date-time":"2022-06-01T00:00:00Z","timestamp":1654041600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGIR Forum"],"published-print":{"date-parts":[[2022,6]]},"abstract":"<jats:p>The discipline of information retrieval (IR) has a long history of examination of how best to measure performance. In particular, there is an extensive literature on the practice of assessing retrieval systems using batch experiments based on collections and relevance judgements. However, this literature has only rarely considered an underlying principle: that measured scores are inherently incomplete as a representation of human activity, that is, there is an innate gap between measured scores and the desired goal of human satisfaction. There are separate challenges such as poor experimental practices or the shortcomings of specific measures, but the issue considered here is more fundamental - straightforwardly, in batch experiments the human-machine gap cannot be closed. In other disciplines, the issue of the gap is well recognised and has been the subject of observations that provide valuable perspectives on the behaviour and effects of measures and the ways in which they can lead to unintended consequences, notably Goodhart's law and the Lucas critique. Here I describe these observations and argue that there is evidence that they apply to IR, thus showing that blind pursuit of performance gains based on optimisation of scores, and analysis based solely on aggregated measurements, can lead to misleading and unreliable outcomes.<\/jats:p>","DOI":"10.1145\/3582524.3582540","type":"journal-article","created":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T17:06:33Z","timestamp":1674839193000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["When Measurement Misleads"],"prefix":"10.1145","volume":"56","author":[{"given":"Justin","family":"Zobel","sequence":"first","affiliation":[{"name":"The University of Melbourne, Parkville, Victoria, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,1,27]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1645953.1646031"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766462.2767728"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/1012294.1012302"},{"key":"e_1_2_1_5_1","volume-title":"Does a proxy measure up? A framework to assess and convey proxy reliability. Climate Past, 16","author":"Boudinot F. G.","year":"2020","unstructured":"F. G. Boudinot and J. Wilson . Does a proxy measure up? A framework to assess and convey proxy reliability. Climate Past, 16 , 2020 . doi: 0.5194\/cp-16-1807-2020. F. G. Boudinot and J. Wilson. Does a proxy measure up? A framework to assess and convey proxy reliability. Climate Past, 16, 2020. doi: 0.5194\/cp-16-1807-2020."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1815663116"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.4337\/9781781950777.00022"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1148170.1148262"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1093\/gigascience\/giz053"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1177\/2515245920952393"},{"key":"e_1_2_1_12_1","volume-title":"Proc SIGIR Workshop: Redundancy, Diversity, and Interdependent Document Relevance","author":"Hawking D.","year":"2009","unstructured":"D. Hawking , T. Rowlands , and P. Thomas . C-TEST: Supporting novelty and diversity in test-files for search tuning . In Proc SIGIR Workshop: Redundancy, Diversity, and Interdependent Document Relevance , 2009 . D. Hawking, T. Rowlands, and P. Thomas. C-TEST: Supporting novelty and diversity in test-files for search tuning. In Proc SIGIR Workshop: Redundancy, Diversity, and Interdependent Document Relevance, 2009."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1277741.1277839"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/231880.231890"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1561\/1500000012"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-016-9282-6"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3077136.3080793"},{"key":"e_1_2_1_18_1","first-page":"19","volume-title":"Carnegie-Rochester Conference Series on Public Policy","volume":"1","author":"Lucas R. E.","year":"1976","unstructured":"R. E. Lucas , Jr. Econometric policy analysis: A critique . In Carnegie-Rochester Conference Series on Public Policy , volume 1 , 1976 . URL https:\/\/EconPapers.repec.org\/RePEc:eee:crcspp:v:1:y:1976:i::p: 19 - 46 . R. E. Lucas, Jr. Econometric policy analysis: A critique. In Carnegie-Rochester Conference Series on Public Policy, volume 1, 1976. URL https:\/\/EconPapers.repec.org\/RePEc:eee:crcspp:v:1:y:1976:i::p:19-46."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511809071"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.3389\/fpsyg.2013.00169"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1177\/0959354320930817"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-45068-6_1"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1416950.1416952"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2505515.2507665"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3052768"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3239572"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3415242"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3463040"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1093\/femsle\/fny059"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2911451.2911492"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3431813"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1561\/1500000009"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/215206.215351"},{"key":"e_1_2_1_34_1","volume-title":"University of D\u00fcsseldorf","author":"Sirotkin P.","year":"2012","unstructured":"P. Sirotkin . On Search Engine Evaluation Metrics. PhD thesis , University of D\u00fcsseldorf , 2012 . URL https:\/\/arxiv.org\/pdf\/1302.2318.pdf. P. Sirotkin. On Search Engine Evaluation Metrics. PhD thesis, University of D\u00fcsseldorf, 2012. URL https:\/\/arxiv.org\/pdf\/1302.2318.pdf."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1234-981X(199707)5:3(305::AID-EURO184)3.0.CO;2-4"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/383952.383992"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/1148170.1148176"},{"key":"e_1_2_1_38_1","volume-title":"Information Retrieval. Butterworths","author":"van Rijsbergen C. J.","year":"1979","unstructured":"C. J. van Rijsbergen . Information Retrieval. Butterworths , 2 nd edition, 1979 . C. J. van Rijsbergen. Information Retrieval. Butterworths, 2nd edition, 1979.","edition":"2"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-45691-0_34"},{"key":"e_1_2_1_40_1","volume-title":"TREC: Experiment and Evaluation in Information Retrieval","author":"Voorhees E. M.","year":"2005","unstructured":"E. M. Voorhees and D. K. Harman . TREC: Experiment and Evaluation in Information Retrieval . MIT Press , 2005 . E. M. Voorhees and D. K. Harman. TREC: Experiment and Evaluation in Information Retrieval. MIT Press, 2005."},{"key":"e_1_2_1_42_1","volume-title":"Proc. ACM SIGIR Int. Conf. on Research and Development in Information Retrieval","author":"Webber W.","year":"2008","unstructured":"W. Webber , A. Moffat , J. Zobel , and T. Sakai . Precision-at-ten considered redundant . In Proc. ACM SIGIR Int. Conf. on Research and Development in Information Retrieval , 2008 . NO DOI. W. Webber, A. Moffat, J. Zobel, and T. Sakai. Precision-at-ten considered redundant. In Proc. ACM SIGIR Int. Conf. on Research and Development in Information Retrieval, 2008. NO DOI."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/1852102.1852106"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3336191.3371799"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401162"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3190580.3190584"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/11610113_4"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3411998"},{"key":"e_1_2_1_49_1","first-page":"10","author":"Zobel J.","year":"2009","unstructured":"J. Zobel , A. Moffat , and L. Park . Against recall: Is it persistence, cardinality, density, coverage, or totality? SIGIR Forum , 2009 . doi: 10 .1.1.415.6729. J. Zobel, A. Moffat, and L. Park. Against recall: Is it persistence, cardinality, density, coverage, or totality? SIGIR Forum, 2009. doi: 10.1.1.415.6729.","journal-title":"SIGIR Forum"}],"container-title":["ACM SIGIR Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3582524.3582540","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3582524.3582540","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:47:16Z","timestamp":1750178836000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3582524.3582540"}},"subtitle":["The Limits of Batch Assessment of Retrieval Systems"],"short-title":[],"issued":{"date-parts":[[2022,6]]},"references-count":46,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,6]]}},"alternative-id":["10.1145\/3582524.3582540"],"URL":"https:\/\/doi.org\/10.1145\/3582524.3582540","relation":{},"ISSN":["0163-5840"],"issn-type":[{"value":"0163-5840","type":"print"}],"subject":[],"published":{"date-parts":[[2022,6]]},"assertion":[{"value":"2023-01-27","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}