{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T14:04:32Z","timestamp":1787061872363,"version":"build-2736575974"},"reference-count":30,"publisher":"MIT Press","issue":"2","license":[{"start":{"date-parts":[[2024,1,8]],"date-time":"2024-01-08T00:00:00Z","timestamp":1704672000000},"content-version":"vor","delay-in-days":7,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,6,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Multiple measures, such as WEAT or MAC, attempt to quantify the magnitude of bias present in word embeddings in terms of a single-number metric. However, such metrics and the related statistical significance calculations rely on treating pre-averaged data as individual data points and utilizing bootstrapping techniques with low sample sizes. We show that similar results can be easily obtained using such methods even if the data are generated by a null model lacking the intended bias. Consequently, we argue that this approach generates false confidence. To address this issue, we propose a Bayesian alternative: hierarchical Bayesian modeling, which enables a more uncertainty-sensitive inspection of bias in word embeddings at different levels of granularity. To showcase our method, we apply it to Religion, Gender, and Race word lists from the original research, together with our control neutral word lists. We deploy the method using Google, GloVe, and Reddit embeddings. Further, we utilize our approach to evaluate a debiasing technique applied to the Reddit word embedding. Our findings reveal a more complex landscape than suggested by the proponents of single-number metrics. The datasets and source code for the paper are publicly available.1<\/jats:p>","DOI":"10.1162\/coli_a_00507","type":"journal-article","created":{"date-parts":[[2024,1,8]],"date-time":"2024-01-08T16:19:55Z","timestamp":1704730795000},"page":"563-617","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":4,"title":["A Bayesian Approach to Uncertainty in Word Embedding Bias\n                    Estimation"],"prefix":"10.1162","volume":"50","author":[{"given":"Alicja","family":"Dobrzeniecka","sequence":"first","affiliation":[{"name":"Vrije Universiteit Amsterdam The Netherlands. alizjonizm@gmail.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rafal","family":"Urbaniak","sequence":"additional","affiliation":[{"name":"Basis Research Institute, NYC, USA The University of Gdansk, Poland. rfl.urbaniak@gmail.com"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,6,1]]},"reference":[{"key":"2024070814304505200_bib1","first-page":"4356","article-title":"Man is to computer programmer as woman is to\n                        homemaker? Debiasing word embeddings","volume-title":"Proceedings\n                        of the 30th International Conference on Neural Information Processing\n                        Systems","author":"Bolukbasi","year":"2016"},{"issue":"6334","key":"2024070814304505200_bib2","doi-asserted-by":"publisher","first-page":"183","DOI":"10.1126\/science.aal4230","article-title":"Semantics derived automatically from\n                        language corpora contain human-like biases","volume":"356","author":"Caliskan","year":"2017","journal-title":"Science"},{"key":"2024070814304505200_bib3","doi-asserted-by":"publisher","first-page":"10012","DOI":"10.18653\/v1\/2021.emnlp-main.785","article-title":"Assessing the reliability of word\n                        embedding gender bias measures","volume-title":"Proceedings of\n                        the 2021 Conference on Empirical Methods in Natural Language\n                        Processing","author":"Du","year":"2021"},{"key":"2024070814304505200_bib4","doi-asserted-by":"publisher","first-page":"2914","DOI":"10.18653\/v1\/2020.acl-main.262","article-title":"Is your classifier actually biased?\n                        Measuring fairness under uncertainty with Bernstein bounds","volume-title":"Proceedings of the 58th Annual Meeting of the Association for\n                        Computational Linguistics","author":"Ethayarajh","year":"2020"},{"key":"2024070814304505200_bib5","doi-asserted-by":"publisher","first-page":"1696","DOI":"10.18653\/v1\/P19-1166","article-title":"Understanding undesirable word embedding\n                        associations","volume-title":"Proceedings of the 57th Annual\n                        Meeting of the Association for Computational Linguistics","author":"Ethayarajh","year":"2019"},{"issue":"16","key":"2024070814304505200_bib6","doi-asserted-by":"publisher","first-page":"E3635\u2013E3644","DOI":"10.1073\/pnas.1720347115","article-title":"Word embeddings quantify 100 years of gender and ethnic\n                        stereotypes","volume":"115","author":"Garg","year":"2018","journal-title":"Proceedings of the National Academy of\n                        Sciences"},{"key":"2024070814304505200_bib7","doi-asserted-by":"publisher","first-page":"1926","DOI":"10.18653\/v1\/2021.acl-long.150","article-title":"Intrinsic bias metrics do not correlate with\n                        application bias","volume-title":"Proceedings of the 59th Annual\n                        Meeting of the Association for Computational Linguistics and the 11th\n                        International Joint Conference on Natural Language Processing (Volume 1:\n                        Long Papers)","author":"Goldfarb-Tarrant","year":"2021"},{"key":"2024070814304505200_bib8","first-page":"609","article-title":"Lipstick on a pig: Debiasing methods cover\n                        up systematic gender biases in word embeddings but do not remove\n                        them","volume-title":"Proceedings of the 2019 Conference of the\n                        North American Chapter of the Association for Computational Linguistics:\n                        Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Gonen","year":"2019"},{"key":"2024070814304505200_bib9","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1145\/2509558.2509563","article-title":"Reporting bias and knowledge\n                            acquisition","volume-title":"Proceedings of the 2013\n                        Workshop on Automated Knowledge Base Construction","author":"Gordon","year":"2013"},{"key":"2024070814304505200_bib10","doi-asserted-by":"publisher","first-page":"122","DOI":"10.1145\/3461702.3462536","article-title":"Detecting emergent intersectional biases:\n                        Contextualized word embeddings contain a distribution of human-like\n                        biases","volume-title":"Proceedings of the 2021 AAAI\/ACM\n                        Conference on AI, Ethics, and Society","author":"Guo","year":"2021"},{"issue":"5","key":"2024070814304505200_bib11","doi-asserted-by":"publisher","first-page":"1157","DOI":"10.3758\/s13423-013-0572-3","article-title":"Robust misinterpretation of confidence\n                        intervals","volume":"21","author":"Hoekstra","year":"2014","journal-title":"Psychonomic Bulletin &\n                        Review"},{"key":"2024070814304505200_bib12","doi-asserted-by":"publisher","first-page":"4212","DOI":"10.18653\/v1\/2022.findings-emnlp.311","article-title":"Mind your bias: A critical review of bias\n                        detection methods for contextual language models","volume-title":"Findings of the Association for Computational Linguistics: EMNLP\n                        2022","author":"Husse","year":"2022"},{"key":"2024070814304505200_bib13","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4614-7138-7","volume-title":"An Introduction to Statistical Learning","author":"James","year":"2013"},{"key":"2024070814304505200_bib14","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1201\/9781003278290-6","article-title":"Are algorithms value-free? Feminist\n                        theoretical virtues in machine learning","volume":"1","author":"Johnson","year":"2023","journal-title":"Journal of\n                        Moral Philosophy"},{"key":"2024070814304505200_bib15","volume-title":"Doing Bayesian Data Analysis","author":"Kruschke","year":"2015","edition":"2nd ed."},{"key":"2024070814304505200_bib16","doi-asserted-by":"publisher","first-page":"85","DOI":"10.18653\/v1\/S19-1010","article-title":"Are we consistently biased?\n                        Multidimensional analysis of biases in distributional word\n                        vectors","volume-title":"Proceedings of the Eighth Joint\n                        Conference on Lexical and Computational Semantics (*SEM\n                        2019)","author":"Lauscher","year":"2019"},{"key":"2024070814304505200_bib17","doi-asserted-by":"publisher","first-page":"379","DOI":"10.1145\/3531146.3533105","article-title":"De-biasing \u201cbias\u201d\n                        measurement","volume-title":"Proceedings of the 2022 ACM\n                        Conference on Fairness, Accountability, and Transparency","author":"Lum","year":"2022"},{"key":"2024070814304505200_bib18","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1062","article-title":"Black is to criminal as caucasian is to\n                        police: Detecting and removing multiclass bias in word\n                        embeddings","author":"Manzini","year":"2019","journal-title":"arXiv preprint\n                    arXiv:1904.04047"},{"key":"2024070814304505200_bib19","doi-asserted-by":"publisher","first-page":"622","DOI":"10.18653\/v1\/N19-1063","article-title":"On measuring social biases in sentence\n                        encoders","volume-title":"Proceedings of the 2019 Conference of\n                        the North American Chapter of the Association for Computational Linguistics:\n                        Human Language Technologies, Volume 1 (Long and Short Papers)","author":"May","year":"2019"},{"key":"2024070814304505200_bib20","doi-asserted-by":"publisher","DOI":"10.1201\/9780429029608","volume-title":"Statistical Rethinking: A Bayesian Course with\n                        Examples in R and Stan","author":"McElreath","year":"2020","edition":"2nd ed."},{"key":"2024070814304505200_bib21","article-title":"Efficient estimation of word representations in vector\n                        space","volume-title":"1st International Conference on Learning\n                        Representations, ICLR 2013, Workshop Track\n                    Proceedings","author":"Mikolov","year":"2013"},{"key":"2024070814304505200_bib22","doi-asserted-by":"publisher","first-page":"103","DOI":"10.3758\/s13423-015-0947-8","article-title":"The fallacy of placing confidence in\n                        confidence intervals","volume":"23","author":"Morey","year":"2015","journal-title":"Psychonomic Bulletin &\n                        Review"},{"issue":"2","key":"2024070814304505200_bib23","doi-asserted-by":"publisher","first-page":"487","DOI":"10.1162\/coli_a_00379","article-title":"Fair is better than sensational: Man is to\n                        doctor as woman is to doctor","volume":"46","author":"Nissim","year":"2020","journal-title":"Computational\n                        Linguistics"},{"issue":"1","key":"2024070814304505200_bib24","doi-asserted-by":"publisher","first-page":"101","DOI":"10.1037\/\/1089-2699.6.1.101","article-title":"Harvesting implicit group attitudes and\n                        beliefs from a demonstration web site","volume":"6","author":"Nosek","year":"2002","journal-title":"Group\n                        Dynamics: Theory, Research, and Practice"},{"key":"2024070814304505200_bib25","doi-asserted-by":"publisher","first-page":"1532","DOI":"10.3115\/v1\/D14-1162","article-title":"GloVe: Global vectors for word\n                        representation","volume-title":"Proceedings of the 2014\n                        Conference on Empirical Methods in Natural Language Processing\n                        (EMNLP)","author":"Pennington","year":"2014"},{"key":"2024070814304505200_bib26","doi-asserted-by":"publisher","first-page":"329","DOI":"10.1162\/tacl_a_00024","article-title":"Native language cognate effects on second\n                        language lexical choice","volume":"6","author":"Rabinovich","year":"2018","journal-title":"Transactions of the\n                        Association for Computational Linguistics"},{"key":"2024070814304505200_bib27","article-title":"Evaluating metrics for bias in word\n                        embeddings","author":"Schr\u00f6der","year":"2021","journal-title":"arXiv preprint\n                    arXiv:2111.07864"},{"key":"2024070814304505200_bib28","doi-asserted-by":"publisher","first-page":"552","DOI":"10.24963\/ijcai.2021\/77","article-title":"Bias silhouette analysis: Towards\n                        assessing the quality of bias metrics for word embedding\n                        models","volume-title":"Proceedings of the Thirtieth\n                        International Joint Conference on Artificial Intelligence,\n                    IJCAI-21","author":"Splieth\u00f6ver","year":"2021"},{"issue":"01","key":"2024070814304505200_bib29","doi-asserted-by":"publisher","first-page":"7322","DOI":"10.1609\/aaai.v33i01.33017322","article-title":"Quantifying uncertainties in natural language processing\n                        tasks","volume":"33","author":"Xiao","year":"2019","journal-title":"Proceedings of the AAAI Conference on\n                        Artificial Intelligence"},{"key":"2024070814304505200_bib30","first-page":"759","article-title":"Robustness and reliability of gender bias\n                        assessment in word embeddings: The role of base pairs","volume-title":"Proceedings of the 1st Conference of the Asia-Pacific Chapter of the\n                        Association for Computational Linguistics and the 10th International Joint\n                        Conference on Natural Language Processing","author":"Zhang","year":"2020"}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/50\/2\/563\/2457607\/coli_a_00507.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/50\/2\/563\/2457607\/coli_a_00507.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,8]],"date-time":"2024-07-08T10:34:42Z","timestamp":1720434882000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/50\/2\/563\/118988\/A-Bayesian-Approach-to-Uncertainty-in-Word"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":30,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2024,6,1]]},"published-print":{"date-parts":[[2024,6,1]]}},"URL":"https:\/\/doi.org\/10.1162\/coli_a_00507","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"value":"0891-2017","type":"print"},{"value":"1530-9312","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}