{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T04:40:41Z","timestamp":1780634441160,"version":"3.54.1"},"reference-count":71,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2023,8,18]],"date-time":"2023-08-18T00:00:00Z","timestamp":1692316800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2024,1,31]]},"abstract":"<jats:p>We present a simple and versatile framework for evaluating ranked lists in terms of Group Fairness and Relevance, in which the groups (i.e., possible attribute values) can be either nominal or ordinal in nature. First, we demonstrate that when our framework is applied to a binary hard group membership setting, our Group Fairness and Relevance (GFR) measures can easily quantify the overall polarity of each ranked list. Second, by utilising an existing diversified search test collection and treating each intent as an attribute value, we demonstrate that our framework can also handle soft group membership and that the GFR measures are highly correlated with a diversified information retrieval (IR) measure in this context as well. Third, using real data from a Japanese local search service, we demonstrate how our framework enables researchers to study intersectional group fairness based on multiple attribute sets. We also show that the similarity function for comparing the achieved and target distributions over the attribute values should be chosen carefully when the attribute values are ordinal. For such situations, our recommendation is to use multiple similarity functions with our framework: for example, one based on Jensen-Shannon Divergence (which disregards the ordinal nature of the groups) and another based on Root Normalised Order-aware Divergence (which has been designed specifically for handling ordinal groups). In addition, we highlight the fundamental differences between our framework and Attention-Weighted Rank Fairness (AWRF), a group fairness measure used at the TREC Fair Ranking Track.<\/jats:p>","DOI":"10.1145\/3589763","type":"journal-article","created":{"date-parts":[[2023,5,31]],"date-time":"2023-05-31T05:10:22Z","timestamp":1685509822000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["A Versatile Framework for Evaluating Ranked Lists in Terms of Group Fairness and Relevance"],"prefix":"10.1145","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6720-963X","authenticated-orcid":false,"given":"Tetsuya","family":"Sakai","sequence":"first","affiliation":[{"name":"Waseda University\/Naver Corporation, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-2678-3197","authenticated-orcid":false,"given":"Jin Young","family":"Kim","sequence":"additional","affiliation":[{"name":"Naver Corporation, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-8636-6880","authenticated-orcid":false,"given":"Inho","family":"Kang","sequence":"additional","affiliation":[{"name":"Naver Corporation, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,8,18]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"crossref","unstructured":"Rakesh Agrawal Gollapudi Sreenivas Alan Halverson and Samuel Leong. 2009. Diversifying search results. In Proceedings of the Second ACM International Conference on Web Search and Data Mining (WSDM\u201909 Barcelona Spain) . Association for Computing Machinery 5\u201314.","DOI":"10.1145\/1498759.1498766"},{"key":"e_1_3_3_3_2","doi-asserted-by":"crossref","unstructured":"Yamen Ajjour Henning Wachsmuth Johannes Kiesel Martin Potthast Matthias Hagen and Benno Stein. 2019. Data acquisition for argument search: The args.me corpus. In Advances in Artificial Intelligence (KI\u201919) (Lecture Notes in Computer Science 11793) Christoph Benzm\u00fcller and Heiner Stuckenschmidt (Eds.). Springer 48\u201359.","DOI":"10.1007\/978-3-030-30179-8_4"},{"key":"e_1_3_3_4_2","doi-asserted-by":"crossref","unstructured":"Enrique Amig\u00f3 Damiano Spina and Jorge Carrillo de Albornoz. 2018. An axiomatic analysis of diversity evaluation metrics: Introducing the rank-biased utility metric. In The 41st International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201918 Ann Arbor MI USA) . Association for Computing Machinery 625\u2013634.","DOI":"10.1145\/3209978.3210024"},{"key":"e_1_3_3_5_2","doi-asserted-by":"crossref","unstructured":"Vito Walter Anelli Tommaso Di Noia Eugenio Di Sciascio Claudio Pomo and Azzurra Ragone. 2019. On the discriminative power of hyper-parameters in cross-validation and how to choose them. In Proceedings of the 13th ACM Conference on Recommender Systems (RecSys\u201919 Copenhagen Denmark) . Association for Computing Machinery 447\u2013451.","DOI":"10.1145\/3298689.3347010"},{"key":"e_1_3_3_6_2","doi-asserted-by":"crossref","unstructured":"Azin Ashkan and Donald Metzler. 2019. Revisiting online personal search metrics with the user in mind. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201919 Paris France) . Association for Computing Machinery 625\u2013634.","DOI":"10.1145\/3331184.3331266"},{"key":"e_1_3_3_7_2","doi-asserted-by":"crossref","unstructured":"Alex Beutel Jilin Chen Tulsee Doshi Hai Qian Li Wei Yi Wu Lukasz Heldt Zhe Zhao Lichan Hong Ed H. Chi and Cristos Goodrow. 2019. Fairness in recommendation ranking through pairwise comparisons. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD\u201919 Anchorage AK USA) . Association for Computing Machinery 2212\u20132220.","DOI":"10.1145\/3292500.3330745"},{"key":"e_1_3_3_8_2","unstructured":"Asia J. Biega Fernando Diaz Michael D. Ekstrand Sergey Feldman and Sebastian Kohlmeier. 2021. Overview of the TREC 2020 fair ranking track. In The 29th Text REtrieval Conference (TREC\u201920) Proceedings . NIST. https:\/\/trec.nist.gov\/pubs\/trec29\/papers\/OVERVIEW.FR.pdf."},{"key":"e_1_3_3_9_2","unstructured":"Asia J. Biega Fernando Diaz Michael D. Ekstrand and Sebastian Kohlmeier. 2020. Overview of the TREC 2019 fair ranking track. In The 28th Text REtrieval Conference (TREC\u201919) Proceedings . NIST. https:\/\/trec.nist.gov\/pubs\/trec28\/papers\/OVERVIEW.FR.pdf."},{"key":"e_1_3_3_10_2","doi-asserted-by":"crossref","unstructured":"Asia J. Biega Krishna P. Gummadi and Gerhard Weikum. 2018. Equity of attention: Amortizing individual fairness in rankings. In The 41st International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201918 Ann Arbor MI USA) . Association for Computing Machinery 405\u2013414.","DOI":"10.1145\/3209978.3210063"},{"key":"e_1_3_3_11_2","doi-asserted-by":"crossref","unstructured":"Alexander Bondarenko Maik Fr\u00f6be Meriem Beloucif Lukas Gienapp Yamen Ajjour Alexander Panchenko Chris Biemann Benno Stein Henning Wachsmuth Martin Potthast and Matthias Hagen. 2020. Overview of touch\u00e9 2020: Argument retrieval. In Experimental IR Meets Multilinguality Multimodality and Interaction (Lecture Notes in Computer Science 12260) Avi Arampatzis Evangelos Kanoulas Theodora Tsikrika Stefanos Vrochidis Hideo Joho Christina Lioma Carsten Eickhoff Aur\u00e9lie N\u00e9v\u00e9ol Linda Cappellato and Nicola Ferro (Eds.). Springer 384\u2013395.","DOI":"10.1007\/978-3-030-58219-7_26"},{"key":"e_1_3_3_12_2","doi-asserted-by":"crossref","unstructured":"Alexander Bondarenko Lukas Gienapp Maik Fr\u00f6be Meriem Beloucif Yamen Ajjour Alexander Panchenko Chris Biemann Benno Stein Henning Wachsmuth Martin Potthast and Matthias Hagen. 2021. Overview of touch\u00e9 2020: Argument retrieval. In Experimental IR Meets Multilinguality Multimodality and Interaction (Lecture Notes in Computer Science 12880) K. Sel\u00e7uk Candan Bogdan Ionescu Lorraine Goeuriot Birger Larsen Henning M\u00fcller Alexis Joly Maria Maistro Florina Piroi Guglielmo Faggioli and Nicola Ferro (Eds.). Springer 450\u2013467.","DOI":"10.1007\/978-3-030-85251-1_28"},{"key":"e_1_3_3_13_2","volume-title":"Proceedings of Machine Learning Research","author":"Buolamwini Joy","year":"2018","unstructured":"Joy Buolamwini and Timnit Gebru. 2018. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Proceedings of Machine Learning Research, Sorelle A. Friedler and Christo Wilson (Eds.). Vol. 81. PMLR, 77\u201391."},{"key":"e_1_3_3_14_2","unstructured":"L. Elisa Celis Damian Straszak and Nisheeth K. Vishnoi. 2017. Ranking with fairness constraints. (2017). http:\/\/arxiv.org\/abs\/1704.06840."},{"issue":"6","key":"e_1_3_3_15_2","doi-asserted-by":"crossref","first-page":"572","DOI":"10.1007\/s10791-011-9167-7","article-title":"Intent-based diversification of web search results: Metrics and algorithms","volume":"14","author":"Chapelle Olivier","year":"2011","unstructured":"Olivier Chapelle, Shihao Ji, Ciya Liao, Emre Velipasaoglu, Larry Lai, and Su-Lin Wu. 2011. Intent-based diversification of web search results: Metrics and algorithms. Information Retrieval 14, 6 (2011), 572\u2013592.","journal-title":"Information Retrieval"},{"key":"e_1_3_3_16_2","doi-asserted-by":"crossref","unstructured":"Olivier Chapelle Donald Metzler Ya Zhang and Pierre Grinspan. 2009. Expected reciprocal rank for graded relevance. In Proceedings of the 18th ACM Conference on Information and Knowledge Management (CIKM\u201909 Hong Kong China) . Association for Computing Machinery 621\u2013630.","DOI":"10.1145\/1645953.1646033"},{"key":"e_1_3_3_17_2","doi-asserted-by":"crossref","unstructured":"Sachin Pathiyan Cherumanal Damiano Spina Falk Scholer and W. Bruce Croft. 2021. Evaluating fairness in argument retrieval. In Proceedings of the 30th ACM International Conference on Information and Knowledge Management (CIKM\u201921 Virtual Event Queensland Australia) . Association for Computing Machinery 3363\u20133367.","DOI":"10.1145\/3459637.3482099"},{"key":"e_1_3_3_18_2","doi-asserted-by":"crossref","unstructured":"Aleksandr Chuklin Pavel Serdyuov and Maarten de Rijke. 2013. Click model-based information retrieval metrics. In Proceedings of the 36th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201913 Dublin Ireland) . Association for Computing Machinery 493\u2013502.","DOI":"10.1145\/2484028.2484071"},{"key":"e_1_3_3_19_2","unstructured":"Charles L. A. Clarke Nick Craswell Ian Soboroff and Azin Ashkan. 2011. A comparative analysis of cascade measures for novelty and diversity. In Proceedings of the 4th ACM International Conference on Web Search and Data Mining (WSDM\u201911 Hong Kong China) . Association for Computing Machinery 75\u201384."},{"key":"e_1_3_3_20_2","unstructured":"Charles L. A. Clarke Nick Craswell and Ellen M. Voorhees. 2013. Overview of the TREC 2012 web track. In The 21st Text REtrieval Conference (TREC 2012) Proceedings . NIST. https:\/\/trec.nist.gov\/pubs\/trec21\/papers\/WEB12.overview.pdf."},{"key":"e_1_3_3_21_2","doi-asserted-by":"crossref","unstructured":"Charles L. A. Clarke Alexandra Vtyurina and Mark D. Smucker. 2020. Offline evaluation without gain. In Proceedings of the 2020 ACM SIGIR on International Conference on Theory of Information Retrieval (ICTIR\u201920 Virtual Event Norway) . Association for Computing Machinery 185\u2013192.","DOI":"10.1145\/3409256.3409816"},{"key":"e_1_3_3_22_2","doi-asserted-by":"crossref","unstructured":"Fernando Diaz Bhaskar Mitra Michael D. Ekstrand Asia J. Biega and Ben Carterette. 2020. Evaluating stochastic rankings with expected exposure. In Proceedings of the 29th ACM International Conference on Information and Knowledge Management (CIKM\u201920 Virtual Event Ireland) . Association for Computing Machinery 275\u2013284.","DOI":"10.1145\/3340531.3411962"},{"key":"e_1_3_3_23_2","doi-asserted-by":"crossref","first-page":"50","DOI":"10.1145\/3468507.3468515","article-title":"Assessing viewpoint diversity in search results using ranking fairness metrics","volume":"23","author":"Draws Tim","year":"2021","unstructured":"Tim Draws, Nava Tintarev, and Ujwal Gadiraju. 2021. Assessing viewpoint diversity in search results using ranking fairness metrics. ACM SIGKDD Explorations NewsLetter 23, 1 (2021), 50\u201358.","journal-title":"ACM SIGKDD Explorations NewsLetter"},{"key":"e_1_3_3_24_2","doi-asserted-by":"crossref","unstructured":"Michael D. Ekstrand Anubrata Das Robin Burke and Fernando Diaz. 2021. Fairness and discrimination in information access systems. (2021). https:\/\/arxiv.org\/abs\/2105.05779.","DOI":"10.1561\/9781638280415"},{"key":"e_1_3_3_25_2","doi-asserted-by":"crossref","unstructured":"Michael D. Ekstrand Graham McDonald and Amifa Raj. 2022. Overview of the TREC 2021 fair ranking track. In The 30th Text REtrieval Conference (TREC\u201921) Proceedings . NIST. https:\/\/trec.nist.gov\/pubs\/trec30\/papers\/Overview-F.pdf.","DOI":"10.6028\/NIST.SP.500-338.fair-overview"},{"key":"e_1_3_3_26_2","unstructured":"James R. Foulds Rashidul Islam Kamrun Naher Keya and Shimei Pan. 2019. An intersectional definition of fairness. (2019). http:\/\/arxiv.org\/abs\/1807.08362"},{"key":"e_1_3_3_27_2","doi-asserted-by":"crossref","unstructured":"Sahin Cem Geyik Stuart Ambler and Krishnaram Kenthapadi. 2019. Fairness-aware ranking in search & recommendation systems with application to LinkedIn talent search. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD\u201919 Anchorage AK USA) . Association for Computing Machinery 2221\u20132231.","DOI":"10.1145\/3292500.3330691"},{"key":"e_1_3_3_28_2","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1007\/s10791-020-09386-w","article-title":"Evaluation metrics for measuring bias in search engine results","volume":"24","author":"Gezici Gizem","year":"2021","unstructured":"Gizem Gezici, Aldo Lipani, Yucel Saygin, and Emine Yilmaz. 2021. Evaluation metrics for measuring bias in search engine results. Information Retrieval Journal 24 (2021), 85\u2013113.","journal-title":"Information Retrieval Journal"},{"key":"e_1_3_3_29_2","doi-asserted-by":"crossref","unstructured":"Avijit Ghosh Ritam Dutt and Christo Wilson. 2021. When fair ranking meets uncertain inference. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201921 Virtual Event Canada) . Association for Computing Machinery 1033\u20131043.","DOI":"10.1145\/3404835.3462850"},{"key":"e_1_3_3_30_2","article-title":"Characterizing intersectional group fairness with worst-case comparisons","volume":"2101","author":"Ghosh Avijit","year":"2021","unstructured":"Avijit Ghosh, Lea Genuit, and Mary Reagan. 2021. Characterizing intersectional group fairness with worst-case comparisons. CoRR abs\/2101.01673 (2021). https:\/\/arxiv.org\/abs\/2101.01673.","journal-title":"CoRR"},{"key":"e_1_3_3_31_2","doi-asserted-by":"crossref","first-page":"530","DOI":"10.1007\/s10791-012-9218-8","article-title":"Increasing evaluation sensitivity to diversity","volume":"16","author":"Golbus Peter B.","year":"2013","unstructured":"Peter B. Golbus, Javed A. Aslam, and Carles L. A. Clarke. 2013. Increasing evaluation sensitivity to diversity. Information Retrieval 16 (2013), 530\u2013555.","journal-title":"Information Retrieval"},{"key":"e_1_3_3_32_2","doi-asserted-by":"crossref","DOI":"10.4324\/9781315629049","volume-title":"What If There Were No Significance Tests?","author":"Harlow Lisa L.","year":"2016","unstructured":"Lisa L. Harlow, Stanley A. Mulaik, and James H. Steiger. 2016. What If There Were No Significance Tests? (Classic Edition). Routledge."},{"issue":"4","key":"e_1_3_3_33_2","doi-asserted-by":"crossref","first-page":"422","DOI":"10.1145\/582415.582418","article-title":"Cumulated gain-based evaluation of IR techniques","volume":"20","author":"J\u00e4rvelin Kalervo","year":"2002","unstructured":"Kalervo J\u00e4rvelin and Jaana Kek\u00e4l\u00e4inen. 2002. Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems 20, 4 (2002), 422\u2013446.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_3_34_2","doi-asserted-by":"crossref","unstructured":"Evangelos Kanoulas and Javed A. Aslam. 2009. Empirical justification of the gain and discount function for nDCG. In Proceedings of the 18th ACM Conference on Information and Knowledge Management (CIKM\u201909 Hong Kong China) . Association for Computing Machinery 611\u2013620.","DOI":"10.1145\/1645953.1646032"},{"key":"e_1_3_3_35_2","volume-title":"Rank Correlation Methods","author":"Kendall Maurice G.","year":"1962","unstructured":"Maurice G. Kendall. 1962. Rank Correlation Methods (3rd Edition). Charles Griffin and Company Limited."},{"key":"e_1_3_3_36_2","doi-asserted-by":"crossref","unstructured":"\u00d6mer Kirnap Fernando Diaz Asia Biega Michael Ekstrand Ben Carterette and Emine Yilmaz. 2021. Estimation of fair ranking metrics with incomplete judgments. In Proceedings of the Web Conference 2021 (Ljubljana Slovenia) . Association for Computing Machinery 1065\u20131075.","DOI":"10.1145\/3442381.3450080"},{"key":"e_1_3_3_37_2","doi-asserted-by":"crossref","unstructured":"Caitlin Kuhlman Walter Gerych and Elke Rundensteiner. 2021. Measuring group advantage: A comparative study of fair ranking metrics. In Proceedings of the 2021 AAAI\/ACM Conference on AI Ethics and Society (AIES\u201921 Virtual Event USA) . Association for Computing Machinery 674\u2013682.","DOI":"10.1145\/3461702.3462588"},{"key":"e_1_3_3_38_2","doi-asserted-by":"crossref","unstructured":"Caitlin Kuhlman MaryAnn VanValkenburg and Elke Rundensteiner. 2019. FARE: Diagnostics for fair ranking using pairwise error metrics. In The World Wide Web Conference (WWW\u201919 San Francisco CA USA) . Association for Computing Machinery 2936\u20132942.","DOI":"10.1145\/3308558.3313443"},{"key":"e_1_3_3_39_2","doi-asserted-by":"crossref","unstructured":"Juhi Kulshrestha Motahhare Eslami Johnnatan Messias Muhammad Bilal Zafar Saptarshi Ghosh Krishna P. Gunmadi and Karrie Karahalios. 2017. Quantifying search bias: Investigating sources of bias for political searches in social media. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing (CSCW\u201917 Portland Oregon USA) . Association for Computing Machinery 417\u2013432.","DOI":"10.1145\/2998181.2998321"},{"key":"e_1_3_3_40_2","doi-asserted-by":"crossref","unstructured":"Teerapong Leelanupab Guido Zuccon and Joemon M. Jose. 2012. A comprehensive analysis of parameter settings for novelty-biased cumulative gain. In Proceedings of the 21st ACM International Conference on Information and Knowledge Management (CIKM\u201912 Maui Hawaii USA) . Association for Computing Machinery 1950\u20131954.","DOI":"10.1145\/2396761.2398550"},{"key":"e_1_3_3_41_2","doi-asserted-by":"crossref","unstructured":"Elizaveta Levina and Peter Bickel. 2001. The earth mover\u2019s distance is the mallows distance: Some insights from statistics. In Proceedings 8th IEEE International Conference on Computer Vision (ICCV\u201901 Vancouver BC Canada) . IEEE 251\u2013256.","DOI":"10.1109\/ICCV.2001.937632"},{"issue":"1","key":"e_1_3_3_42_2","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1109\/18.61115","article-title":"Divergence measures based on the Shannon entropy","volume":"37","author":"Lin Jianhua","year":"1991","unstructured":"Jianhua Lin. 1991. Divergence measures based on the Shannon entropy. IEEE Transactions on Information Theory 37, 1 (1991), 145\u2013151.","journal-title":"IEEE Transactions on Information Theory"},{"key":"e_1_3_3_43_2","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1111\/j.2044-8317.1997.tb01100.x","article-title":"Confidence intervals for Kendall\u2019s tau","volume":"50","author":"Long Jeffrey D.","year":"1997","unstructured":"Jeffrey D. Long and Norman Cliff. 1997. Confidence intervals for Kendall\u2019s tau. British Journal of Mathematical and Statistical Psychology 50 (1997), 31\u201341.","journal-title":"British Journal of Mathematical and Statistical Psychology"},{"issue":"4","key":"e_1_3_3_44_2","doi-asserted-by":"crossref","first-page":"416","DOI":"10.1007\/s10791-016-9282-6","article-title":"The effect of pooling and evaluation depth on IR metrics","volume":"19","author":"Lu Xiaolu","year":"2016","unstructured":"Xiaolu Lu, Alistair Moffat, and J. Shane Culpepper. 2016. The effect of pooling and evaluation depth on IR metrics. Information Retrieval Journal 19, 4 (2016), 416\u2013445.","journal-title":"Information Retrieval Journal"},{"issue":"3","key":"e_1_3_3_45_2","first-page":"24","article-title":"Incorporating user expectations and behavior into the measurement of search effectiveness","volume":"35","author":"Moffat Alistair","year":"2017","unstructured":"Alistair Moffat, Peter Bailey, Falk Scholer, and Paul Thomas. 2017. Incorporating user expectations and behavior into the measurement of search effectiveness. ACM Transactions on Information Systems 35, 3, Article 24 (2017).","journal-title":"ACM Transactions on Information Systems"},{"issue":"1","key":"e_1_3_3_46_2","first-page":"2","article-title":"Rank-biased precision for measurement of retrieval effectiveness","volume":"27","author":"Moffat Alistair","year":"2008","unstructured":"Alistair Moffat and Justin Zobel. 2008. Rank-biased precision for measurement of retrieval effectiveness. ACM Transactions on Information Systems 27, 1, Article 2 (2008).","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_3_47_2","doi-asserted-by":"crossref","unstructured":"Harikrishna Narasimhan Andrew Cotter Maya Gupta and Serena Wang. 2020. Pairwise fairness for ranking and regression. In Proceedings of the AAAI Conference on Artificial Intelligence (New York USA) Vol. 34 AAAI 5248\u20135255.","DOI":"10.1609\/aaai.v34i04.5970"},{"key":"e_1_3_3_48_2","doi-asserted-by":"crossref","first-page":"16","DOI":"10.1145\/3186549.3186553","article-title":"On measuring bias in online information","volume":"46","author":"Pitoura Evaggelia","year":"2017","unstructured":"Evaggelia Pitoura, Panayiotis Tsaparas, Giorgos Flouris, Irini Fundulaki, Panagiotis Papadakos, Serge Abiteboul, and Gerhard Weikum. 2017. On measuring bias in online information. ACM SIGMOD Record 46, 4 (2017), 16\u201321.","journal-title":"ACM SIGMOD Record"},{"key":"e_1_3_3_49_2","doi-asserted-by":"crossref","unstructured":"Amifa Raj and Michael D. Ekstrand. 2022. Measuring fairness in ranked results. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201922 Madrid Spain) . Association for Computing Machinery 726\u2013736.","DOI":"10.1145\/3477495.3532018"},{"key":"e_1_3_3_50_2","volume-title":"Proceedings of the ACM on Human-Computer Interaction","volume":"2","author":"Robertson Ronald E.","year":"2018","unstructured":"Ronald E. Robertson, Shan Jiang, Kenneth Joseph, Lisa Friedland, David Lazer, and Christo Wilson. 2018. Auditing partisan audience bias within Google search. In Proceedings of the ACM on Human-Computer Interaction, Vol. 2."},{"key":"e_1_3_3_51_2","doi-asserted-by":"crossref","unstructured":"Stephen E. Robertson Evangelos Kanoulas and Emine Yilmaz. 2010. Extending average precision to graded relevance judgements. In Proceedings of the 33rd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201910 Geneva Switzerland) . Association for Computing Machinery 603\u2013610.","DOI":"10.1145\/1835449.1835550"},{"key":"e_1_3_3_52_2","doi-asserted-by":"crossref","unstructured":"Tetsuya Sakai. 2006. Evaluating evaluation metrics based on the bootstrap. In Proceedings of the 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201906 Seattle Washington USA) . Association for Computing Machinery 525\u2013532.","DOI":"10.1145\/1148170.1148261"},{"key":"e_1_3_3_53_2","doi-asserted-by":"crossref","unstructured":"Tetsuya Sakai. 2007. Alternatives to bpref. In Proceedings of the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201907 Amsterdam The Netherlands) . Association for Computing Machinery 71\u201378.","DOI":"10.1145\/1277741.1277756"},{"key":"e_1_3_3_54_2","doi-asserted-by":"crossref","unstructured":"Tetsuya Sakai. 2012. Evaluation with informational and navigational intents. In Proceedings of the 21st International Conference on World Wide Web (Lyon France) . Association for Computing Machinery 499\u2013508.","DOI":"10.1145\/2187836.2187904"},{"key":"e_1_3_3_55_2","doi-asserted-by":"crossref","unstructured":"Tetsuya Sakai. 2014. Metrics statistics tests. In Bridging Between Information Retrieval and Databases (Lecture Notes in Computer Science 8173) Nicola Ferro (Ed.). Springer 116\u2013163.","DOI":"10.1007\/978-3-642-54798-0_6"},{"key":"e_1_3_3_56_2","doi-asserted-by":"crossref","unstructured":"Tetsuya Sakai. 2018. Comparing two binned probability distributions for information access evaluation. In The 41st International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201918 Ann Arbor MI USA) . Association for Computing Machinery 1073\u20131076.","DOI":"10.1145\/3209978.3210073"},{"key":"e_1_3_3_57_2","doi-asserted-by":"crossref","unstructured":"Tetsuya Sakai. 2018. Laboratory experiments in information retrieval: Sample sizes effect sizes and statistical power. Springer.","DOI":"10.1007\/978-981-13-1199-4"},{"key":"e_1_3_3_58_2","doi-asserted-by":"crossref","unstructured":"Tetsuya Sakai. 2020. Evaluating evaluation measures for ordinal classification and ordinal quantification. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (Online) . Association for Computational Linguistics 2759\u20132769.","DOI":"10.18653\/v1\/2021.acl-long.214"},{"key":"e_1_3_3_59_2","unstructured":"Tetsuya Sakai. 2021. A closer look at evaluation measures for ordinal quantification. In Proceedings of the CIKM 2021 Workshops co-located with 30th ACM International Conference on Information and Knowledge Management (CIKM\u201921) . https:\/\/ceur-ws.org\/Vol-3052\/paper21.pdf."},{"key":"e_1_3_3_60_2","unstructured":"Tetsuya Sakai. 2021. On the instability of diminishing return IR measures. In Advances in Information Retrieval: 43rd European Conference on IR Research ECIR 2021 Part I (Lecture Notes in Computer Science 12656) Djoerd Hiemstra Marie-Francine Moens Josiane Mothe Raffaele Perego Martin Potthast and Fabrizio Sebastiani (Eds.). Springer 572\u2013586."},{"key":"e_1_3_3_61_2","doi-asserted-by":"crossref","unstructured":"Tetsuya Sakai Jin Young Kim and Inho Kang. 2022. A versatile framework for evaluating ranked lists in terms of group fairness and relevance. (2022). http:\/\/arxiv.org\/abs\/2204.00280.","DOI":"10.1145\/3589763"},{"key":"e_1_3_3_62_2","unstructured":"Tetsuya Sakai and Stephen E. Robertson. 2008. Modelling a user population for designing information retrieval metrics. In Proceedings of the 2nd International Workshop on Evaluating Information Access (EVIA\u201908) . 30\u201341. http:\/\/research.nii.ac.jp\/ntcir\/workshop\/OnlineProceedings7\/pdf\/EVIA2008\/07-EVIA2008-SakaiT.pdf."},{"key":"e_1_3_3_63_2","doi-asserted-by":"crossref","unstructured":"Tetsuya Sakai and Ruihua Song. 2011. Evaluating diversified search results using per-intent graded relevance. In Proceedings of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201911 Beijing China) . Association for Computing Machinery 1043\u20131052.","DOI":"10.1145\/2009916.2010055"},{"issue":"2","key":"e_1_3_3_64_2","first-page":"14","article-title":"Retrieval evaluation measures that agree with users\u2019 SERP preferences: Traditional, preference-based, and diversity measures","volume":"39","author":"Sakai Tetsuya","year":"2020","unstructured":"Tetsuya Sakai and Zhaohao Zeng. 2020. Retrieval evaluation measures that agree with users\u2019 SERP preferences: Traditional, preference-based, and diversity measures. ACM Transactions on Information Systems 39, 2, Article 14 (2020).","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_3_65_2","doi-asserted-by":"crossref","unstructured":"Piotr Sapiezynski Wesley Zeng and Ronald E. Robertson. 2019. Quantifying the impact of user attention on fair group representation in ranked lists. In Companion Proceedings of The 2019 World WideWeb Conference (San Francisco USA) . Association for Computing Machinery 553\u2013562.","DOI":"10.1145\/3308560.3317595"},{"key":"e_1_3_3_66_2","doi-asserted-by":"crossref","unstructured":"Ashudeep Singh and Thorsten Joachims. 2018. Fairness of exposure in rankings. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD\u201918 London United Kingdom) . Association for Computing Machinery 2219\u20132228.","DOI":"10.1145\/3219819.3220088"},{"key":"e_1_3_3_67_2","unstructured":"Ruihua Song Min Zhang Tetsuya Sakai Makoto P. Kato Yiqun Liu Miho Sugimoto Qinglei Wang and Naoki Orii. 2011. Overview of the NTCIR-9 INTENT task. In Proceedings of The 9th NTCIR Workshop Meeting (Tokyo Japan) . National Institute of Informatics 82\u2013105."},{"key":"e_1_3_3_68_2","doi-asserted-by":"crossref","first-page":"411","DOI":"10.1007\/s10791-020-09377-x","article-title":"Assessing ranking metrics in top-N recommendation","volume":"23","author":"Valcarce Daniel","year":"2020","unstructured":"Daniel Valcarce, Alejandro, Bellog\u00edn, Javier Parapar, and Pablo Castells. 2020. Assessing ranking metrics in top-N recommendation. Information Retrieval Journal 23 (2020), 411\u2013448.","journal-title":"Information Retrieval Journal"},{"key":"e_1_3_3_69_2","doi-asserted-by":"crossref","first-page":"328","DOI":"10.1016\/0734-189X(85)90055-6","article-title":"A distance metric for multidimensional histograms","volume":"32","author":"Werman Michael","year":"1985","unstructured":"Michael Werman, Shmuel Peleg, and Azriel Rosenfeld. 1985. A distance metric for multidimensional histograms. Computer Vision, Graphics, and Image Processing 32 (1985), 328\u2013336.","journal-title":"Computer Vision, Graphics, and Image Processing"},{"key":"e_1_3_3_70_2","doi-asserted-by":"crossref","unstructured":"Ke Yang and Julia Stoyanovich. 2017. Measuring fairness in ranked outputs. In Proceedings of the 29th International Conference on Scientific and Statistical Database Management (SSDBM\u201917: Chicago IL USA) . Association for Computing Machinery.","DOI":"10.1145\/3085504.3085526"},{"key":"e_1_3_3_71_2","doi-asserted-by":"crossref","unstructured":"Meike Zehlike Francesco Bonchi Carlos Castillo Sara Hajian Mohamed Megahed and Ricardo Baeza-Yates. 2017. FA*IR: A fair top-k ranking algorithm. In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management (CIKM\u201917 Singapore Singapore) . Association for Computing Machinery 1569\u20131578.","DOI":"10.1145\/3132847.3132938"},{"key":"e_1_3_3_72_2","doi-asserted-by":"crossref","unstructured":"ChengXiang Zhai William W. Cohen and John Lafferty. 2003. Beyond independent relevance: Methods and evaluation metrics for subtopic retrieval. In Proceedings of the 26th Annual International ACM SIGIR Conference on Research and Development in Informaion Retrieval (SIGIR\u201903 Toronto Canada) . Association for Computing Machinery 10\u201317.","DOI":"10.1145\/860435.860440"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589763","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3589763","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:22Z","timestamp":1750182562000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589763"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,18]]},"references-count":71,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,1,31]]}},"alternative-id":["10.1145\/3589763"],"URL":"https:\/\/doi.org\/10.1145\/3589763","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,18]]},"assertion":[{"value":"2022-04-06","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-03-14","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-08-18","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}