{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,26]],"date-time":"2025-06-26T04:02:47Z","timestamp":1750910567568,"version":"3.41.0"},"reference-count":83,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2017,8,29]],"date-time":"2017-08-29T00:00:00Z","timestamp":1503964800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2018,4,30]]},"abstract":"<jats:p>We propose the<jats:italic>Assessor-driven Weighted Averages for Retrieval Evaluation (AWARE)<\/jats:italic>probabilistic framework, a novel methodology for dealing with multiple crowd assessors that may be contradictory and\/or noisy. By modeling relevance judgements and crowd assessors as sources of uncertainty, AWARE takes the expectation of a generic performance measure, like Average Precision, composed with these random variables. In this way, it approaches the problem of aggregating different crowd assessors from a new perspective, that is, directly combining the performance measures computed on the ground truth generated by the crowd assessors instead of adopting some classification technique to merge the labels produced by them. We propose several unsupervised estimators that instantiate the AWARE framework and we compare them with state-of-the-art approaches, that is,Majoriity Vote and Expectation Maximization, on TREC collections. We found that AWARE approaches improve in terms of their capability of correctly ranking systems and predicting their actual performance scores.<\/jats:p>","DOI":"10.1145\/3110217","type":"journal-article","created":{"date-parts":[[2017,8,29]],"date-time":"2017-08-29T17:49:18Z","timestamp":1504028958000},"page":"1-38","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["AWARE"],"prefix":"10.1145","volume":"36","author":[{"given":"Marco","family":"Ferrante","sequence":"first","affiliation":[{"name":"University of Padua, Padova, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nicola","family":"Ferro","sequence":"additional","affiliation":[{"name":"University of Padua, Padova, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Maria","family":"Maistro","sequence":"additional","affiliation":[{"name":"University of Padua, Padova, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,8,29]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-012-9204-1"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2012.01.004"},{"key":"e_1_2_2_3_1","doi-asserted-by":"crossref","unstructured":"P. Bailey N. Craswell I. Soboroff P. Thomas A. P. de Vries and E. Yilmaz. 2008. Relevance assessment: Are judges exchangeable and does it matter? See [12] 667--674. P. Bailey N. Craswell I. Soboroff P. Thomas A. P. de Vries and E. Yilmaz. 2008. Relevance assessment: Are judges exchangeable and does it matter? See [12] 667--674.","DOI":"10.1145\/1390334.1390447"},{"key":"e_1_2_2_4_1","doi-asserted-by":"crossref","unstructured":"M. Bashir J. Anderton J. Wu M. Ekstrand-Abueg P. B. Golbus V. Pavlu and J. A. Aslam. 2013. Northeastern university runs at the TREC12 crowdsourcing track. See [74]. M. Bashir J. Anderton J. Wu M. Ekstrand-Abueg P. B. Golbus V. Pavlu and J. A. Aslam. 2013. Northeastern university runs at the TREC12 crowdsourcing track. See [74].","DOI":"10.6028\/NIST.SP.500-298.crowd-NEU"},{"key":"e_1_2_2_5_1","doi-asserted-by":"crossref","unstructured":"R. Blanco H. Halpin D. M. Herzig P. Mika J. Pound and H. S. Thompson. 2011. Repeatable and reliable search system evaluation using crowdsourcing. See [50] 923--932. R. Blanco H. Halpin D. M. Herzig P. Mika J. Pound and H. S. Thompson. 2011. Repeatable and reliable search system evaluation using crowdsourcing. See [50] 923--932.","DOI":"10.1145\/2009916.2010039"},{"key":"e_1_2_2_6_1","unstructured":"C. Buckley and E. M. Voorhees. 2005. Retrieval system evaluation. In TREC. Experiment and Evaluation in Information Retrieval D. K. Harman and E. M. Voorhees (Eds.). MIT Press Cambridge MA 53--78. C. Buckley and E. M. Voorhees. 2005. Retrieval system evaluation. In TREC. Experiment and Evaluation in Information Retrieval D. K. Harman and E. M. Voorhees (Eds.). MIT Press Cambridge MA 53--78."},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1016\/0306-4573(92)90031-T"},{"key":"e_1_2_2_8_1","unstructured":"K. P. Burnham and D. R. Anderson. 2002. Model Selection and Multimodel Inference: A Practical Information-Theoretic Approach (2nd ed.). Springer-Verlag Heidelberg Germany. K. P. Burnham and D. R. Anderson. 2002. Model Selection and Multimodel Inference: A Practical Information-Theoretic Approach (2nd ed.). Springer-Verlag Heidelberg Germany."},{"key":"e_1_2_2_9_1","doi-asserted-by":"crossref","unstructured":"B. A. Carterette and I. Soboroff. 2010. The effect of assessor errors on IR system evaluation. See [15] 539--546. B. A. Carterette and I. Soboroff. 2010. The effect of assessor errors on IR system evaluation. See [15] 539--546.","DOI":"10.1145\/1835449.1835540"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1645953.1646033"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2396761"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1390334"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/MIC.2012.95"},{"volume-title":"TREC 2005 spam track overview. In Proceedings of the 14th Text REtrieval Conference (TREC\u201905)","author":"Cormack G.","key":"e_1_2_2_14_1"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1835449"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.2307\/2346806"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2348283.2348400"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2005.10.010"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-11382-1_3"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2808194.2809452"},{"key":"e_1_2_2_21_1","first-page":"2","article-title":"Reproducibility challenges in information retrieval evaluation","volume":"8","author":"Ferro N.","year":"2017","journal-title":"ACM J. Data Inf. Qual."},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964797.2964808"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1002\/asi.23416"},{"edition":"4","volume-title":"Matrix Computations","author":"Golub G. H.","key":"e_1_2_2_24_1"},{"key":"e_1_2_2_25_1","unstructured":"C. Grady and M. Lease. 2010. Crowdsourcing document relevance assessment with mechanical turk. In Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon\u2019s Mechanical Turk C. Callison-Burch and M. Dredze (Eds.). The Association for Computational Linguistics 172--179. C. Grady and M. Lease. 2010. Crowdsourcing document relevance assessment with mechanical turk. In Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon\u2019s Mechanical Turk C. Callison-Burch and M. Dredze (Eds.). The Association for Computational Linguistics 172--179."},{"key":"e_1_2_2_26_1","doi-asserted-by":"crossref","unstructured":"M. Halvey R. Villa and P. Clough. 2014. SIGIR 2014 workshop on gathering efficient assessments of relevance (GEAR). In Proceedigns of the 37th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201914) S. Geva A. Trotman P. Bruza C. L. A. Clarke and K. J\u00e4rvelin (Eds.). ACM Press New York NY 1293. M. Halvey R. Villa and P. Clough. 2014. SIGIR 2014 workshop on gathering efficient assessments of relevance (GEAR). In Proceedigns of the 37th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201914) S. Geva A. Trotman P. Bruza C. L. A. Clarke and K. J\u00e4rvelin (Eds.). ACM Press New York NY 1293.","DOI":"10.1145\/2600428.2600735"},{"key":"e_1_2_2_27_1","doi-asserted-by":"crossref","unstructured":"C. Harris and P. Srinivasan. 2013. Using hybrid methods for relevance assessment in TREC crowd\u201912. See [74]. C. Harris and P. Srinivasan. 2013. Using hybrid methods for relevance assessment in TREC crowd\u201912. See [74].","DOI":"10.6028\/NIST.SP.500-298.crowd-UIowaS"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1002\/9780470316672"},{"volume":"7224","volume-title":"Proceedings of the 32nd European Conference on IR Research: Advances in Information Retrieval (ECIR\u201912)","author":"Hosseini M.","key":"e_1_2_2_29_1"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2566486.2567988"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/582415.582418"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1609\/aimag.v35i2.2537"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-16354-3_17"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-20161-5_17"},{"key":"e_1_2_2_35_1","doi-asserted-by":"crossref","unstructured":"G. Kazai N. Craswell E. Yilmaz and S. S. M. Tahaghoghi. 2012. An analysis of systematic judging errors in information retrieval. See [11] 105--114. G. Kazai N. Craswell E. Yilmaz and S. S. M. Tahaghoghi. 2012. An analysis of systematic judging errors in information retrieval. See [11] 105--114.","DOI":"10.1145\/2396761.2396779"},{"key":"e_1_2_2_36_1","doi-asserted-by":"crossref","unstructured":"G. Kazai J. Kamps M. Koolen and N. Mili\u0107-Frayling. 2011. Crowdsourcing for book search evaluation: Impact of HIT design on comparative system ranking. See [50] 205--214. G. Kazai J. Kamps M. Koolen and N. Mili\u0107-Frayling. 2011. Crowdsourcing for book search evaluation: Impact of HIT design on comparative system ranking. See [50] 205--214.","DOI":"10.1145\/2009916.2009947"},{"volume-title":"Proceedings of the 20th International Conference on Information and Knowledge Management (CIKM\u201911)","author":"Kazai G.","key":"e_1_2_2_37_1"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-012-9205-0"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2505515.2505716"},{"volume-title":"Rank Correlation Methods","author":"Kendall M. G.","key":"e_1_2_2_40_1","doi-asserted-by":"crossref","DOI":"10.2307\/2333282"},{"key":"e_1_2_2_41_1","unstructured":"J. F. Kenney and E. S. Keeping. 1954. Mathematics of Statistics\u2014Part One (3rd ed.). D. Van Nostrand Company Princeton NJ. J. F. Kenney and E. S. Keeping. 1954. Mathematics of Statistics\u2014Part One (3rd ed.). D. Van Nostrand Company Princeton NJ."},{"key":"e_1_2_2_42_1","first-page":"4","article-title":"Special issue: Crowd in intelligent systems","volume":"7","author":"King I.","year":"2016","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/1458082.1458160"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177729694"},{"key":"e_1_2_2_45_1","doi-asserted-by":"crossref","unstructured":"E. Law P. N. Bennett and E. Horvitz. 2011. The effects of choice in routing relevance judgments. See [50] 1127--1128. E. Law P. N. Bennett and E. Horvitz. 2011. The effects of choice in routing relevance judgments. See [50] 1127--1128.","DOI":"10.1145\/2009916.2010082"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-013-9222-7"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1016\/0020-0271(68)90029-6"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.5555\/2964060.2964164"},{"volume-title":"Proceedings of the MediaEval 2013 Multimedia Benchmark Workshop, M. Larson, X. Anguera, T. Reuter, G. J. F. Jones, B. Ionescu, M. Schedl, T. Piatrik, C. Hauff, and M. Soleymani (Eds.).","author":"Loni B.","key":"e_1_2_2_49_1"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/2009916"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1561\/1900000044"},{"key":"e_1_2_2_52_1","doi-asserted-by":"crossref","unstructured":"S. Maxwell and H. D. Delaney. 2004. Designing Experiments and Analyzing Data. A Model Comparison Perspective (2nd ed.). Lawrence Erlbaum Associates Mahwah NJ. S. Maxwell and H. D. Delaney. 2004. Designing Experiments and Analyzing Data. A Model Comparison Perspective (2nd ed.). Lawrence Erlbaum Associates Mahwah NJ.","DOI":"10.4324\/9781410609243"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/2873063"},{"key":"e_1_2_2_54_1","unstructured":"K. R. Murphy B. Myors and A. Wolach. 2014. Statistical Power Analysis: A Simple and General Model for Traditional and Modern Hypothesis Tests (4th ed.). Routledge Taylor 8 Francis Group UK. K. R. Murphy B. Myors and A. Wolach. 2014. Statistical Power Analysis: A Simple and General Model for Traditional and Modern Hypothesis Tests (4th ed.). Routledge Taylor 8 Francis Group UK."},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1037\/1082-989X.8.4.434"},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2013.01.035"},{"key":"e_1_2_2_57_1","unstructured":"V. C. Raykar and S. Yu. 2012. Eliminating spammers and ranking annotators for crowdsourced labeling tasks. Journ. Mach. Learn. Res. 13 (February 2012) 491--518. V. C. Raykar and S. Yu. 2012. Eliminating spammers and ranking annotators for crowdsourced labeling tasks. Journ. Mach. Learn. Res. 13 (February 2012) 491--518."},{"volume-title":"Proceedings of the 26th Annual International Conference on Machine Learning (ICML\u201909)","author":"Raykar V. C.","key":"e_1_2_2_58_1"},{"key":"e_1_2_2_59_1","unstructured":"V. C. Raykar L. H. Zhao G. Hermosillo Valadez C. Florin L. Bogoni and L. Moy. 2010. Learning from crowds. J. Mach. Learn. Res. 11 (April 2010) 1297--1322. V. C. Raykar L. H. Zhao G. Hermosillo Valadez C. Florin L. Bogoni and L. Moy. 2010. Learning from crowds. J. Mach. Learn. Res. 11 (April 2010) 1297--1322."},{"key":"e_1_2_2_60_1","doi-asserted-by":"crossref","unstructured":"S. E. Robertson E. Kanoulas and E. Yilmaz. 2010. Extending average precision to graded relevance judgments See [15] 603--610. S. E. Robertson E. Kanoulas and E. Yilmaz. 2010. Extending average precision to graded relevance judgments See [15] 603--610.","DOI":"10.1145\/1835449.1835550"},{"edition":"2","volume-title":"A GLM Approach","author":"Rutherford A.","key":"e_1_2_2_61_1"},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1108\/JD-02-2014-0031"},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/2641383.2641385"},{"volume-title":"Proceedings of the 28th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201905)","author":"Sanderson M.","key":"e_1_2_2_64_1"},{"volume-title":"Proceedings of the SIGIR 2011 Workshop on Crowdsourcing for Information Retrieval.","author":"Smucker M. D.","key":"e_1_2_2_65_1"},{"key":"e_1_2_2_66_1","doi-asserted-by":"crossref","unstructured":"M. D. Smucker and C. P. Jethani. 2011. Measuring assessor accuracy: A comparison of nist assessors and user study participants. See [50] 1231--1232. M. D. Smucker and C. P. Jethani. 2011. Measuring assessor accuracy: A comparison of nist assessors and user study participants. See [50] 1231--1232.","DOI":"10.1145\/2009916.2010134"},{"volume-title":"Overview of the TREC 2012 crowdsourcing track. See [74].","author":"Smucker M. D.","key":"e_1_2_2_67_1"},{"volume-title":"Overview of the TREC 2013 crowdsourcing track. In Proceedings of the 22nd Text REtrieval Conference Proceedings (TREC","year":"2013","author":"Smucker M. D.","key":"e_1_2_2_68_1"},{"key":"e_1_2_2_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/383952.383961"},{"key":"e_1_2_2_70_1","doi-asserted-by":"publisher","DOI":"10.1126\/science.103.2684.677"},{"key":"e_1_2_2_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/290941.291017"},{"key":"e_1_2_2_72_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0306-4573(00)00010-8"},{"volume-title":"Overview of the TREC 2004 robust track. In Proceedings of the 13th Text REtrieval Conference Proceedings (TREC\u201904)","year":"2004","author":"Voorhees E. M.","key":"e_1_2_2_73_1"},{"volume-title":"In Proceedings of the 21st Text REtrieval Conference Proceedings (TREC\u201912)","year":"2013","author":"Voorhees E. M.","key":"e_1_2_2_74_1"},{"volume-title":"Proceedings of the 8th Text REtrieval Conference (TREC-8), E. M. Voorhees and D. K. Harman (Eds.). National Institute of Standards and Technology (NIST), Special Publication 500-246","author":"Voorhees E. M.","key":"e_1_2_2_75_1"},{"key":"e_1_2_2_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/MIC.2012.71"},{"key":"e_1_2_2_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/2854946.2854968"},{"key":"e_1_2_2_78_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4899-4493-1"},{"key":"e_1_2_2_79_1","doi-asserted-by":"crossref","unstructured":"W. Webber P. Chandar and B. A. Carterette. 2012. Alternative assessor disagreement and retrieval depth. See [11] 125--134. W. Webber P. Chandar and B. A. Carterette. 2012. Alternative assessor disagreement and retrieval depth. See [11] 125--134.","DOI":"10.1145\/2396761.2396781"},{"key":"e_1_2_2_80_1","doi-asserted-by":"publisher","DOI":"10.1145\/2484028.2484156"},{"key":"e_1_2_2_81_1","doi-asserted-by":"publisher","DOI":"10.1214\/aos\/1176346060"},{"key":"e_1_2_2_82_1","unstructured":"K. Yadati P. S. N. Shakthinathan C. Ayyanathan and M. Larson. 2014. Crowdsorting timed comments about music: Foundations for a new crowdsourcing task. In Proceedings of the MediaEval 2014 Multimedia Benchmark Workshop M. Larson B. Ionescu X. Anguera M. Eskevich P. Korshunov M. Schedl M. Soleymani P. Petkos R. Sutcliffe J. Choi and G. J. F. Jones (Eds.). CEUR Workshop Proceedings (CEUR-WS.org). K. Yadati P. S. N. Shakthinathan C. Ayyanathan and M. Larson. 2014. Crowdsorting timed comments about music: Foundations for a new crowdsourcing task. In Proceedings of the MediaEval 2014 Multimedia Benchmark Workshop M. Larson B. Ionescu X. Anguera M. Eskevich P. Korshunov M. Schedl M. Soleymani P. Petkos R. Sutcliffe J. Choi and G. J. F. Jones (Eds.). CEUR Workshop Proceedings (CEUR-WS.org)."},{"key":"e_1_2_2_83_1","doi-asserted-by":"crossref","unstructured":"E. Yilmaz J. A. Aslam and S. E. Robertson. 2008. A new rank correlation coefficient for information retrieval See [12] 587--594. E. Yilmaz J. A. Aslam and S. E. Robertson. 2008. A new rank correlation coefficient for information retrieval See [12] 587--594.","DOI":"10.1145\/1390334.1390435"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3110217","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3110217","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,25]],"date-time":"2025-06-25T05:56:23Z","timestamp":1750830983000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3110217"}},"subtitle":["Exploiting Evaluation Measures to Combine Multiple Assessors"],"short-title":[],"issued":{"date-parts":[[2017,8,29]]},"references-count":83,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2018,4,30]]}},"alternative-id":["10.1145\/3110217"],"URL":"https:\/\/doi.org\/10.1145\/3110217","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"type":"print","value":"1046-8188"},{"type":"electronic","value":"1558-2868"}],"subject":[],"published":{"date-parts":[[2017,8,29]]},"assertion":[{"value":"2016-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-08-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}