{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T19:50:57Z","timestamp":1779306657459,"version":"3.51.4"},"reference-count":41,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2024,5,31]],"date-time":"2024-05-31T00:00:00Z","timestamp":1717113600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,5,31]],"date-time":"2024-05-31T00:00:00Z","timestamp":1717113600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"European Union\u2019s Horizon 2020 for the project : \u201cNoBIAS - Artificial Intelligence without Bias\u201d","award":["860630"],"award-info":[{"award-number":["860630"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Data Min Knowl Disc"],"published-print":{"date-parts":[[2024,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Human feedback is often used, either directly or indirectly, as input to algorithmic decision making. However, humans are biased: if the algorithm that takes as input the human feedback does not control for potential biases, this might result in biased algorithmic decision making, which can have a tangible impact on people\u2019s lives. In this paper, we study how to detect and correct for evaluators\u2019 bias in the task of <jats:italic>ranking people (or items) from pairwise comparisons<\/jats:italic>. Specifically, we assume we are given pairwise comparisons of the items to be ranked produced by a set of evaluators. While the pairwise assessments of the evaluators should reflect to a certain extent the latent (unobservable) true quality scores of the items, they might be affected by each evaluator\u2019s own bias against, or in favor, of some groups of items. By detecting and amending evaluators\u2019 biases, we aim to produce a ranking of the items that is, as much as possible, in accordance with the ranking one would produce by having access to the latent quality scores. Our proposal is a novel method that extends the classic Bradley-Terry model by having a bias parameter for each evaluator which distorts the true quality score of each item, depending on the group the item belongs to. Thanks to the simplicity of the model, we are able to write explicitly its log-likelihood w.r.t. the parameters (i.e., items\u2019 latent scores and evaluators\u2019 bias) and optimize by means of the alternating approach. Our experiments on synthetic and real-world data confirm that our method is able to reconstruct the bias of each single evaluator extremely well and thus to outperform several non-trivial competitors in the task of producing a ranking which is as much as possible close to the unbiased ranking.<\/jats:p>","DOI":"10.1007\/s10618-024-01024-z","type":"journal-article","created":{"date-parts":[[2024,5,31]],"date-time":"2024-05-31T02:01:30Z","timestamp":1717120890000},"page":"2062-2086","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Bias-aware ranking from pairwise comparisons"],"prefix":"10.1007","volume":"38","author":[{"given":"Antonio","family":"Ferrara","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Francesco","family":"Bonchi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Francesco","family":"Fabbri","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fariba","family":"Karimi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Claudia","family":"Wagner","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,5,31]]},"reference":[{"issue":"2","key":"1024_CR1","doi-asserted-by":"publisher","first-page":"151","DOI":"10.1177\/1948550619837002","volume":"11","author":"A Almaatouq","year":"2020","unstructured":"Almaatouq A, Krafft P, Dunham Y, Rand DG, Pentland A (2020) Turkers of the world unite: multilevel in-group bias among crowdworkers on amazon mechanical Turk. Soc Psychol Personal Scince 11(2):151\u2013159","journal-title":"Soc Psychol Personal Scince"},{"key":"1024_CR2","doi-asserted-by":"crossref","unstructured":"Alvarez JM, Ruggieri S (2023) Counterfactual situation testing: Uncovering discrimination under fairness given the difference. Preprint arXiv:2302.11944","DOI":"10.1145\/3617694.3623222"},{"issue":"10","key":"1024_CR3","first-page":"923","volume":"4","author":"RJ Beaver","year":"1975","unstructured":"Beaver RJ, Gokhale D (1975) A model to incorporat within-pair order effects in paired comparisons. Commun Stat Theory Methods 4(10):923\u2013939","journal-title":"Commun Stat Theory Methods"},{"issue":"3\/4","key":"1024_CR4","doi-asserted-by":"publisher","first-page":"324","DOI":"10.2307\/2334029","volume":"39","author":"RA Bradley","year":"1952","unstructured":"Bradley RA, Terry ME (1952) Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika 39(3\/4):324\u2013345","journal-title":"Biometrika"},{"key":"1024_CR5","unstructured":"Bugakova N, Fedorova V, Gusev G, Drutsa A (2019) Aggregation of pairwise comparisons with reduction of biases. Preprint arXiv:1906.03711"},{"issue":"4","key":"1024_CR6","doi-asserted-by":"publisher","first-page":"835","DOI":"10.1017\/S0003055417000302","volume":"111","author":"D Carlson","year":"2017","unstructured":"Carlson D, Montgomery JM (2017) A pairwise comparison framework for fast, flexible, and reliable human coding of political texts. Am Polit Sci Rev 111(4):835\u2013843","journal-title":"Am Polit Sci Rev"},{"key":"1024_CR7","unstructured":"Celis LE, Straszak D, Vishnoi NK (2017) Ranking with fairness constraints. Preprint arXiv:1704.06840"},{"key":"1024_CR8","doi-asserted-by":"crossref","unstructured":"Chen X, Bennett PN, Collins-Thompson K, Horvitz E (2013) Pairwise ranking aggregation in a crowdsourced setting. In: Proceedings of the sixth ACM international conference on web search and data mining, pp 193\u2013202","DOI":"10.1145\/2433396.2433420"},{"key":"1024_CR9","unstructured":"Chen Y, Suh C (2015) Spectral MLE: Top-K rank aggregation from pairwise comparisons. In: International conference on machine learning, pp 371\u2013380"},{"key":"1024_CR10","unstructured":"Crenshaw K (2013) Demarginalizing the intersection of race and sex: a black feminist critique of antidiscrimination doctrine, feminist theory and antiracist politics. In: Feminist legal theories. Routledge, pp 23\u201351"},{"issue":"2","key":"1024_CR11","doi-asserted-by":"publisher","first-page":"432","DOI":"10.1093\/biomet\/74.2.432","volume":"74","author":"HA David","year":"1987","unstructured":"David HA (1987) Ranking from unbalanced paired-comparison data. Biometrika 74(2):432\u2013436","journal-title":"Biometrika"},{"key":"1024_CR12","doi-asserted-by":"crossref","unstructured":"Davidson RR, Beaver RJ (1977) On extending the Bradley-Terry model to incorporate within-pair order effects. Biometrics 693\u2013702","DOI":"10.2307\/2529467"},{"key":"1024_CR13","unstructured":"Fogel F, d\u2019Aspremont A, Vojnovic M (2014) Serialrank: spectral ranking using seriation. Adv Neural Inf Process Syst 27"},{"key":"1024_CR14","doi-asserted-by":"crossref","unstructured":"Garc\u00eda-Soriano D, Bonchi F (2021) Maxmin-fair ranking: individual fairness under group-fairness constraints. In: KDD \u201921: The 27th ACM SIGKDD conference on knowledge discovery and data mining, pp 436\u2013446","DOI":"10.1145\/3447548.3467349"},{"key":"1024_CR15","doi-asserted-by":"crossref","unstructured":"Geva M, Goldberg Y, Berant J (2019) Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets. In: Inui K, Jiang J, Ng V, Wan X (eds) Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing, EMNLP-IJCNLP, pp 1161\u20131166","DOI":"10.18653\/v1\/D19-1107"},{"key":"1024_CR16","doi-asserted-by":"crossref","unstructured":"Gohar U, Cheng L (2023) A survey on intersectional fairness in machine learning: Notions, mitigation, and challenges. Preprint arXiv:2305.06969","DOI":"10.24963\/ijcai.2023\/742"},{"key":"1024_CR17","unstructured":"He Y, Gan Q, Wipf D, Reinert GD, Yan J, Cucuringu M (2022) GNNRank: learning global rankings from pairwise comparisons via directed graph neural networks. In: International conference on machine learning, pp 8581\u20138612"},{"key":"1024_CR18","doi-asserted-by":"crossref","unstructured":"Hube C, Fetahu B, Gadiraju U (2019) Understanding and mitigating worker biases in the crowdsourced collection of subjective judgments. In: Proceedings of the 2019 CHI conference on human factors in computing systems","DOI":"10.1145\/3290605.3300637"},{"key":"1024_CR19","doi-asserted-by":"crossref","unstructured":"Kamar E, Kapoor A, Horvitz E (2015) Identifying and accounting for task-dependent bias in crowdsourcing. In: Proceedings of the AAAI conference on human computation and crowdsourcing, vol 3, no 1, pp 92\u2013101","DOI":"10.1609\/hcomp.v3i1.13238"},{"issue":"1\/2","key":"1024_CR20","doi-asserted-by":"publisher","first-page":"81","DOI":"10.2307\/2332226","volume":"30","author":"MG Kendall","year":"1938","unstructured":"Kendall MG (1938) A new measure of rank correlation. Biometrika 30(1\/2):81\u201393","journal-title":"Biometrika"},{"issue":"3","key":"1024_CR21","doi-asserted-by":"publisher","first-page":"239","DOI":"10.1093\/biomet\/33.3.239","volume":"33","author":"MG Kendall","year":"1945","unstructured":"Kendall MG (1945) The treatment of ties in ranking problems. Biometrika 33(3):239\u2013251","journal-title":"Biometrika"},{"issue":"3\/4","key":"1024_CR22","doi-asserted-by":"publisher","first-page":"324","DOI":"10.2307\/2332613","volume":"31","author":"MG Kendall","year":"1940","unstructured":"Kendall MG, Smith BB (1940) On the method of paired comparisons. Biometrika 31(3\/4):324\u2013345","journal-title":"Biometrika"},{"issue":"4","key":"1024_CR23","doi-asserted-by":"publisher","first-page":"53","DOI":"10.5430\/rwe.v11n4p53","volume":"11","author":"I Koshkalda","year":"2020","unstructured":"Koshkalda I, Kniaz O, Ryasnyanska A, Velieva V (2020) Motivation mechanism for stimulating the labor potential. Res World Econ 11(4):53\u201361","journal-title":"Res World Econ"},{"key":"1024_CR24","doi-asserted-by":"crossref","unstructured":"Kotturi Y, Kahng A, Procaccia A, Kulkarni C (2020) Hirepeer: impartial peer-assessed hiring at scale in expert crowdsourcing markets. In: Proceedings of the AAAI conference on artificial intelligence, vol 34, pp 2577\u20132584","DOI":"10.1609\/aaai.v34i03.5641"},{"key":"1024_CR25","unstructured":"Kuo T-S, Hankin M, Miranda L, Ying A, Wang C (2020) Assessing political bias using crowdsourced pairwise comparisons.  In: Proceedings of the AAAI conference on human computation and crowdsourcing"},{"key":"1024_CR26","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2021.106633","volume":"138","author":"SK Kuttal","year":"2021","unstructured":"Kuttal SK, Chen X, Wang Z, Balali S, Sarma A (2021) Visual resume: exploring developers\u2019s online contributions for hiring. Inf Softw Technol 138:106633","journal-title":"Inf Softw Technol"},{"key":"1024_CR27","doi-asserted-by":"crossref","unstructured":"Liu H, Thekinen J, Mollaoglu S, Tang D, Yang J, Cheng Y, Liu H, Tang J (2022) Toward annotator group bias in crowdsourcing. In: Proceedings of the 60th annual meeting of the association for computational linguistics","DOI":"10.18653\/v1\/2022.acl-long.126"},{"key":"1024_CR28","unstructured":"Negahban S, Oh S, Shah D (2012) Iterative ranking from pair-wise comparisons. Adv Neural Inf Process Syst 25"},{"key":"1024_CR29","doi-asserted-by":"crossref","unstructured":"Nocedal J, Wright SJ (1999) Numerical optimization","DOI":"10.1007\/b98874"},{"key":"1024_CR30","unstructured":"Pavlichenko N, Ustalov D (2021) IMDB-WIKI-SbS: an evaluation dataset for crowdsourced pairwise comparisons. Preprint arXiv:2110.14990"},{"key":"1024_CR31","volume-title":"Race gaps in sat scores highlight inequality and hinder upward mobility","author":"RV Reeves","year":"2017","unstructured":"Reeves RV, Halikias D (2017) Race gaps in sat scores highlight inequality and hinder upward mobility. Brookings Institute, Washington"},{"issue":"2\u20134","key":"1024_CR32","doi-asserted-by":"publisher","first-page":"144","DOI":"10.1007\/s11263-016-0940-3","volume":"126","author":"R Rothe","year":"2018","unstructured":"Rothe R, Timofte R, Gool LV (2018) Deep expectation of real and apparent age from a single image without facial landmarks. Int J Comput Vis 126(2\u20134):144\u2013157","journal-title":"Int J Comput Vis"},{"key":"1024_CR33","doi-asserted-by":"crossref","unstructured":"Sap M, Card D, Gabriel S, Choi Y, Smith NA (2019) The risk of racial bias in hate speech detection. In: Proceedings of the 57th annual meeting of the association for computational linguistics, pp 1668\u20131678","DOI":"10.18653\/v1\/P19-1163"},{"key":"1024_CR34","doi-asserted-by":"crossref","unstructured":"Sap M, Swayamdipta S, Vianna L, Zhou X, Choi Y, Smith NA (2021) Annotators with attitudes: how annotator beliefs and identities bias toxic language detection. Preprint arXiv:2111.07997","DOI":"10.18653\/v1\/2022.naacl-main.431"},{"key":"1024_CR35","doi-asserted-by":"crossref","unstructured":"Sarma A, Chen X, Kuttal S, Dabbish L, Wang Z (2016) Hiring in the global stage: profiles of online contributions. In: 2016 IEEE 11th international conference on global software engineering (ICGSE), pp 1\u201310","DOI":"10.1109\/ICGSE.2016.35"},{"key":"1024_CR36","doi-asserted-by":"crossref","unstructured":"Singh A, Joachims T (2018) Fairness of exposure in rankings. In: Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining, pp 2219\u20132228","DOI":"10.1145\/3219819.3220088"},{"issue":"4","key":"1024_CR37","doi-asserted-by":"publisher","first-page":"384","DOI":"10.1037\/h0065439","volume":"21","author":"LL Thurstone","year":"1927","unstructured":"Thurstone LL (1927) The method of paired comparisons for social values. J Abnorm Soc Psychol 21(4):384","journal-title":"J Abnorm Soc Psychol"},{"key":"1024_CR38","unstructured":"Wightman LF (1998) LSAC national longitudinal bar passage study. LSAC research report series"},{"issue":"1","key":"1024_CR39","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1007\/s10107-015-0892-3","volume":"151","author":"SJ Wright","year":"2015","unstructured":"Wright SJ (2015) Coordinate descent algorithms. Math Program 151(1):3\u201334","journal-title":"Math Program"},{"key":"1024_CR40","doi-asserted-by":"crossref","unstructured":"Zehlike M, Bonchi F, Castillo C, Hajian S, Megahed M, Baeza-Yates R (2017) FA*IR: a fair top-k ranking algorithm. In: Proceedings of the 2017 ACM on conference on information and knowledge management, pp 1569\u20131578","DOI":"10.1145\/3132847.3132938"},{"key":"1024_CR41","doi-asserted-by":"crossref","unstructured":"Zehlike M, Castillo C (2020) Reducing disparate exposure in ranking: a learning to rank approach. In: Proceedings of the web conference, pp 2849\u20132855","DOI":"10.1145\/3366424.3380048"}],"container-title":["Data Mining and Knowledge Discovery"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10618-024-01024-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10618-024-01024-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10618-024-01024-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,30]],"date-time":"2024-07-30T10:05:03Z","timestamp":1722333903000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10618-024-01024-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,31]]},"references-count":41,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,7]]}},"alternative-id":["1024"],"URL":"https:\/\/doi.org\/10.1007\/s10618-024-01024-z","relation":{},"ISSN":["1384-5810","1573-756X"],"issn-type":[{"value":"1384-5810","type":"print"},{"value":"1573-756X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5,31]]},"assertion":[{"value":"5 December 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 April 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"31 May 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"This paper focuses on bias and fairness issues in ranking from pairwise comparisons. Although the goal is to mitigate such issues, there are several ethical aspects to be considered. First, in order to tackle bias issues we need access to the sensitive attributes (for example gender or race) for which we want to mitigate biases. Furthermore, ranking from pairwise comparison methods, and hence BARP as such, can be used in decision making contexts, such as hiring or university admissions, that can affect human lives. Moreover, we considered biases against groups of people, however discriminatory behaviours against single individuals might still exist and special attention is required in such cases.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}}]}}