{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,3]],"date-time":"2026-08-03T19:50:47Z","timestamp":1785786647306,"version":"3.56.0"},"reference-count":132,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,6,22]],"date-time":"2023-06-22T00:00:00Z","timestamp":1687392000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["J. Data and Information Quality"],"published-print":{"date-parts":[[2023,6,30]]},"abstract":"<jats:p>In this article, we introduce and discuss the pervasive issue of bias in the large language models that are currently at the core of mainstream approaches to Natural Language Processing (NLP). We first introduce data selection bias, that is, the bias caused by the choice of texts that make up a training corpus. Then, we survey the different types of social bias evidenced in the text generated by language models trained on such corpora, ranging from gender to age, from sexual orientation to ethnicity, and from religion to culture. We conclude with directions focused on measuring, reducing, and tackling the aforementioned types of bias.<\/jats:p>","DOI":"10.1145\/3597307","type":"journal-article","created":{"date-parts":[[2023,5,16]],"date-time":"2023-05-16T12:01:40Z","timestamp":1684238500000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":324,"title":["Biases in Large Language Models: Origins, Inventory, and Discussion"],"prefix":"10.1145","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3831-9706","authenticated-orcid":false,"given":"Roberto","family":"Navigli","sequence":"first","affiliation":[{"name":"Sapienza University of Rome, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6238-7816","authenticated-orcid":false,"given":"Simone","family":"Conia","sequence":"additional","affiliation":[{"name":"Sapienza University of Rome, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2717-3705","authenticated-orcid":false,"given":"Bj\u00f6rn","family":"Ross","sequence":"additional","affiliation":[{"name":"University of Edinburgh, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,6,22]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3461702.3462624"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.42"},{"key":"e_1_3_2_4_2","unstructured":"Julia Angwin Jeff Larson Lauren Kirchner and Surya Mattu. 2016. Machine bias. Retrieved from https:\/\/www.propublica.org\/article\/machine-bias-risk-assessments-in-criminal-sentencing."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.2307\/2524353"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/2908131.2908135"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.371"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.112"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.177"},{"key":"e_1_3_2_10_2","first-page":"1","volume-title":"Proceedings of the 2nd Workshop on Gender Bias in Natural Language Processing","author":"Bartl Marion","year":"2020","unstructured":"Marion Bartl, Malvina Nissim, and Albert Gatt. 2020. Unmasking contextual stereotypes: Measuring and mitigating BERT\u2019s gender bias. In Proceedings of the 2nd Workshop on Gender Bias in Natural Language Processing. Association for Computational Linguistics, 1\u201316. Retrieved from https:\/\/aclanthology.org\/2020.gebnlp-1.1."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1037\/emo0000444"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445922"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.463"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i14.17489"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.255"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2021\/593"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1257\/jel.20160995"},{"key":"e_1_3_2_18_2","article-title":"Language contamination explains the cross-lingual capabilities of English pretrained models","volume":"2204","author":"Blevins Terra","year":"2022","unstructured":"Terra Blevins and Luke Zettlemoyer. 2022. Language contamination explains the cross-lingual capabilities of English pretrained models. CoRR abs\/2204.08110 (2022).","journal-title":"CoRR"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2021\/521"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.485"},{"key":"e_1_3_2_21_2","article-title":"Language models are few-shot learners","volume":"2005","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. CoRR abs\/2005.14165 (2020).","journal-title":"CoRR"},{"key":"e_1_3_2_22_2","volume-title":"Proceedings of the 13th International Conference on English Language Research on Computerized Corpora","author":"Burnage Gavin","year":"1992","unstructured":"Gavin Burnage and Dominic Dunlop. 1992. Encoding the British National Corpus. In Proceedings of the 13th International Conference on English Language Research on Computerized Corpora."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2022.103437"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.2105\/AJPH.93.2.191"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.298"},{"key":"e_1_3_2_26_2","article-title":"Highly parallel autoregressive entity linking with discriminative correction","volume":"2109","author":"Cao Nicola De","year":"2021","unstructured":"Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021. Highly parallel autoregressive entity linking with discriminative correction. CoRR abs\/2109.03792 (2021).","journal-title":"CoRR"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.3115\/1706543.1706571"},{"key":"e_1_3_2_28_2","first-page":"arXiv:2103.1202","article-title":"Quality at a glance: An audit of web-crawled multilingual datasets","author":"Caswell Isaac","year":"2021","unstructured":"Isaac Caswell, Julia Kreutzer, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Beno\u00eet Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Javier Ortiz Su\u00e1rez, Iroro Orife, Kelechi Ogueji, Rubungo Andre Niyongabo, Toan Q. Nguyen, Mathias M\u00fcller, Andr\u00e9 M\u00fcller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine \u00c7abuk Ball\u00fd, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2021. Quality at a glance: An audit of web-crawled multilingual datasets. arXiv e-prints, Article arXiv:2103.12028 (March2021).","journal-title":"arXiv e-prints"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1057\/s41288-020-00166-7"},{"key":"e_1_3_2_30_2","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3\u20137, 2019 - Tutorial Abstracts","author":"Chang Kai-Wei","year":"2019","unstructured":"Kai-Wei Chang, Vinod Prabhakaran, and Vicente Ordonez. 2019. Bias and fairness in natural language processing. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3\u20137, 2019 - Tutorial Abstracts, Timothy Baldwin and Marine Carpuat (Eds.). Association for Computational Linguistics. Retrieved from https:\/\/aclanthology.org\/D19-2004\/."},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.3386\/w24904"},{"key":"e_1_3_2_32_2","article-title":"PaLM: Scaling language modeling with pathways","volume":"2204","author":"Chowdhery Aakanksha","year":"2022","unstructured":"Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. PaLM: Scaling language modeling with pathways. arXiv abs\/2204.02311 (2022).","journal-title":"arXiv"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.31"},{"key":"e_1_3_2_34_2","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Conia Simone","year":"2022","unstructured":"Simone Conia, Edoardo Barba, Alessandro Scir\u00e8, and Roberto Navigli. 2022. Semantic role labeling meets definition modeling: Using natural language to describe predicate-argument structures. In Findings of the Association for Computational Linguistics: EMNLP 2022. Association for Computational Linguistics."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.eacl-main.286"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.316"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.747"},{"key":"e_1_3_2_38_2","volume-title":"Proceedings of the 3rd Workshop on Gender Bias in Natural Language Processing","author":"Costa-juss\u00e0 Marta","year":"2021","unstructured":"Marta Costa-juss\u00e0, Hila Gonen, Christian Hardmeier, and Kellie Webster (Eds.). 2021. Proceedings of the 3rd Workshop on Gender Bias in Natural Language Processing. Association for Computational Linguistics. Retrieved from https:\/\/aclanthology.org\/2021.gebnlp-1.0."},{"key":"e_1_3_2_39_2","volume-title":"Proceedings of the 1st Workshop on Gender Bias in Natural Language Processing","author":"Costa-juss\u00e0 Marta R.","year":"2019","unstructured":"Marta R. Costa-juss\u00e0, Christian Hardmeier, Will Radford, and Kellie Webster (Eds.). 2019. In Proceedings of the 1st Workshop on Gender Bias in Natural Language Processing. Association for Computational Linguistics. Retrieved from https:\/\/aclanthology.org\/W19-3800."},{"key":"e_1_3_2_40_2","volume-title":"Proceedings of the 2nd Workshop on Gender Bias in Natural Language Processing","author":"Costa-juss\u00e0 Marta R.","year":"2020","unstructured":"Marta R. Costa-juss\u00e0, Christian Hardmeier, Will Radford, and Kellie Webster (Eds.). 2020. In Proceedings of the 2nd Workshop on Gender Bias in Natural Language Processing. Association for Computational Linguistics. Retrieved from https:\/\/aclanthology.org\/2020.gebnlp-1.0."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00425"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1080\/1369118X.2012.678878"},{"key":"e_1_3_2_43_2","unstructured":"Jeffrey Dastin. 2018. Amazon scraps secret AI recruiting tool that showed bias against women. Retrieved from https:\/\/www.reuters.com\/article\/us-amazon-com-jobs-automation-insight-idUSKCN1MK08G."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.522"},{"key":"e_1_3_2_45_2","first-page":"7659","volume-title":"Proceedings of the 34th AAAI Conference on Artificial Intelligence, AAAI 2020, the 32nd Innovative Applications of Artificial Intelligence Conference, IAAI 2020, the 10th AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, February 7\u201312, 2020","author":"Dev Sunipa","year":"2020","unstructured":"Sunipa Dev, Tao Li, Jeff M. Phillips, and Vivek Srikumar. 2020. On measuring and mitigating biased inferences of word embeddings. In Proceedings of the 34th AAAI Conference on Artificial Intelligence, AAAI 2020, the 32nd Innovative Applications of Artificial Intelligence Conference, IAAI 2020, the 10th AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, February 7\u201312, 2020. AAAI Press, 7659\u20137666. Retrieved from https:\/\/ojs.aaai.org\/index.php\/AAAI\/article\/view\/6267."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3173574.3173986"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2020.113679"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1037\/dev0000550"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00373"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.149"},{"key":"e_1_3_2_52_2","article-title":"The Pile: An 800GB dataset of diverse text for language modeling","author":"Gao Leo","year":"2020","unstructured":"Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The Pile: An 800GB dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027 (2020).","journal-title":"arXiv preprint arXiv:2101.00027"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.203"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1720347115"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10869-013-9290-0"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1162\/089120102760275983"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.150"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1061"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.3115\/1596409.1596411"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/MIS.2009.36"},{"key":"e_1_3_2_61_2","volume-title":"Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP)","author":"Hardmeier Christian","year":"2022","unstructured":"Christian Hardmeier, Christine Basta, Marta R. Costa-juss\u00e0, Gabriel Stanovsky, and Hila Gonen (Eds.). 2022. In Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP). Association for Computational Linguistics. Retrieved from https:\/\/aclanthology.org\/2022.gebnlp-1.0."},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.slpat-1.8"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.482"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1111\/lnc3.12432"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2012.10.002"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1031"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.7"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.487"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1017\/ssh.2020.18"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.831"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/3351095.3375671"},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.1080\/00224540903365414"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.560"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.1162\/089120103322711569"},{"key":"e_1_3_2_75_2","first-page":"79","volume-title":"Proceedings of Machine Translation Summit X: Papers","author":"Koehn Philipp","year":"2005","unstructured":"Philipp Koehn. 2005. EuroParl: A parallel corpus for statistical machine translation. In Proceedings of Machine Translation Summit X: Papers. 79\u201386. Retrieved from https:\/\/aclanthology.org\/2005.mtsummit-papers.11."},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.1177\/8755123315576212"},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W19-3823"},{"key":"e_1_3_2_78_2","unstructured":"Jeff Larson Surya Mattu Lauren Kirchner and Julia Angwin. 2016. How We Analyzed the COMPAS Recidivism Algorithm. Retrieved from https:\/\/www.propublica.org\/article\/how-we-analyzed-the-compas-recidivism-algorithm."},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1066"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.818"},{"key":"e_1_3_2_82_2","doi-asserted-by":"publisher","DOI":"10.1145\/3560815"},{"key":"e_1_3_2_83_2","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","volume":"1907","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv abs\/1907.11692 (2019).","journal-title":"arXiv"},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.nuse-1.5"},{"key":"e_1_3_2_85_2","volume-title":"Proceedings of the 2nd International Conference on Language Resources and Evaluation (LREC\u201900)","author":"Macleod Catherine","year":"2000","unstructured":"Catherine Macleod, Nancy Ide, and Ralph Grishman. 2000. The American national corpus: A standardized resource for American English. In Proceedings of the 2nd International Conference on Language Resources and Evaluation (LREC\u201900). European Language Resources Association (ELRA). Retrieved from http:\/\/www.lrec-conf.org\/proceedings\/lrec2000\/pdf\/196.pdf."},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.121"},{"key":"e_1_3_2_87_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.324"},{"key":"e_1_3_2_88_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S17-2090"},{"key":"e_1_3_2_89_2","doi-asserted-by":"publisher","DOI":"10.29366\/2018tlc.2.4.2"},{"key":"e_1_3_2_90_2","doi-asserted-by":"publisher","DOI":"10.3115\/1075671.1075742"},{"key":"e_1_3_2_91_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.104"},{"key":"e_1_3_2_92_2","doi-asserted-by":"publisher","DOI":"10.2139\/ssrn.3906300"},{"key":"e_1_3_2_93_2","doi-asserted-by":"publisher","DOI":"10.1145\/3461702.3462469"},{"key":"e_1_3_2_94_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.416"},{"key":"e_1_3_2_95_2","article-title":"WebGPT: Browser-assisted question-answering with human feedback","volume":"2112","author":"Nakano Reiichiro","year":"2021","unstructured":"Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021. WebGPT: Browser-assisted question-answering with human feedback. CoRR abs\/2112.09332 (2021).","journal-title":"CoRR"},{"key":"e_1_3_2_96_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.154"},{"key":"e_1_3_2_97_2","doi-asserted-by":"publisher","DOI":"10.1145\/1459352.1459355"},{"key":"e_1_3_2_98_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2021\/620"},{"key":"e_1_3_2_99_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2012.07.001"},{"key":"e_1_3_2_100_2","first-page":"137","article-title":"Socioeconomic bias in the judiciary","volume":"61","author":"Neitz Michele Benedetto","year":"2013","unstructured":"Michele Benedetto Neitz. 2013. Socioeconomic bias in the judiciary. Clevel. State Law Rev. 61 (2013), 137\u2013165. Retrieved from https:\/\/ssrn.com\/abstract=2149311.","journal-title":"Clevel. State Law Rev."},{"key":"e_1_3_2_101_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.583"},{"key":"e_1_3_2_102_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.computel-1.14"},{"key":"e_1_3_2_103_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.ltedi-1.4"},{"key":"e_1_3_2_104_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.aax2342"},{"key":"e_1_3_2_105_2","unstructured":"OpenAI. 2023. GPT-4 Technical Report. arxiv:2303.08774 [cs.CL]."},{"key":"e_1_3_2_106_2","volume-title":"Proceedings of the 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3\u20137, 2021","author":"Paolini Giovanni","year":"2021","unstructured":"Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Jie Ma, Alessandro Achille, Rishita Anubhai, C\u00edcero Nogueira dos Santos, Bing Xiang, and Stefano Soatto. 2021. Structured prediction as translation between augmented natural languages. In Proceedings of the 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3\u20137, 2021. OpenReview.net. Retrieved from https:\/\/openreview.net\/forum?id=US-TP-xnXI."},{"key":"e_1_3_2_107_2","volume-title":"Proceedings of the 1st Workshop on Trustworthy Natural Language Processing","author":"Pruksachatkun Yada","year":"2021","unstructured":"Yada Pruksachatkun, Anil Ramakrishna, Kai-Wei Chang, Satyapriya Krishna, Jwala Dhamala, Tanaya Guha, and Xiang Ren (Eds.). 2021. In Proceedings of the 1st Workshop on Trustworthy Natural Language Processing. Association for Computational Linguistics. Retrieved from https:\/\/aclanthology.org\/2021.trustnlp-1.0."},{"issue":"8","key":"e_1_3_2_108_2","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019). Retrieved from https:\/\/d4mucfpksywv.cloudfront.net\/better-language-models\/language-models.pdf.","journal-title":"OpenAI Blog"},{"key":"e_1_3_2_109_2","first-page":"140:1\u2013140:67","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21 (2020), 140:1\u2013140:67. Retrieved from http:\/\/jmlr.org\/papers\/v21\/20-074.html.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_2_110_2","unstructured":"Sebastian Ruder. 2021. Recent advances in language model fine-tuning. Retrieved from https:\/\/ruder.io\/recent-advances-lm-fine-tuning\/."},{"key":"e_1_3_2_111_2","unstructured":"Sebastian Ruder. 2022. Scaling NLP systems to the next 1000 languages. Retrieved from https:\/\/www.2022.aclweb.org\/invited-talks."},{"key":"e_1_3_2_112_2","article-title":"BLOOM: A 176B-parameter open-access multilingual language model","author":"Scao Teven Le","year":"2022","unstructured":"Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili\u0107, Daniel Hesslow, Roman Castagn\u00e9, Alexandra Sasha Luccioni, Fran\u00e7ois Yvon, Matthias Gall\u00e9 et\u00a0al. 2022. BLOOM: A 176B-parameter open-access multilingual language model. arXiv Preprint arXiv:2211.05100 (2022).","journal-title":"arXiv Preprint arXiv:2211.05100"},{"key":"e_1_3_2_113_2","doi-asserted-by":"publisher","DOI":"10.3233\/SW-222986"},{"key":"e_1_3_2_114_2","doi-asserted-by":"publisher","DOI":"10.1007\/s43681-020-00035-y"},{"key":"e_1_3_2_115_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1159"},{"key":"e_1_3_2_116_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ibusrev.2021.101969"},{"key":"e_1_3_2_117_2","first-page":"to appear","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Tedeschi Simone","year":"2023","unstructured":"Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Haji\u010d, Daniel Hershcovich, Eduard H. Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, and Roberto Navigli. 2023. What\u2019s the meaning of superhuman performance in today\u2019s NLU? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics (to appear)."},{"key":"e_1_3_2_118_2","article-title":"LaMDA: Language Models for Dialog Applications","volume":"2201","author":"Thoppilan Romal","year":"2022","unstructured":"Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Kathleen S. Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed H. Chi, and Quoc Le. 2022. LaMDA: Language Models for Dialog Applications. CoRR abs\/2201.08239 (2022).","journal-title":"CoRR"},{"key":"e_1_3_2_119_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10676-021-09583-1"},{"key":"e_1_3_2_120_2","article-title":"LLaMA: Open and efficient foundation language models","volume":"2302","author":"Touvron Hugo","year":"2023","unstructured":"Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth\u00e9e Lacroix, Baptiste Rozi\u00e8re, Naman Goyal, Eric Hambro, Faisal Azhar, Aur\u00e9lien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. LLaMA: Open and efficient foundation language models. CoRR abs\/2302.13971 (2023).","journal-title":"CoRR"},{"key":"e_1_3_2_121_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2210.14552"},{"key":"e_1_3_2_122_2","unstructured":"Starre Vartan. 2019. Racial bias found in a major health care risk algorithm. Retrieved from https:\/\/www.scientificamerican.com\/article\/racial-bias-found-in-a-major-health-care-risk-algorithm\/."},{"key":"e_1_3_2_123_2","article-title":"Nationality bias in text generation","volume":"2302","author":"Venkit Pranav Narayanan","year":"2023","unstructured":"Pranav Narayanan Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao K. Huang, and Shomir Wilson. 2023. Nationality bias in text generation. CoRR abs\/2302.02463 (2023).","journal-title":"CoRR"},{"key":"e_1_3_2_124_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.61"},{"key":"e_1_3_2_125_2","unstructured":"Wikipedia contributors. 2022. Who writes Wikipedia? Retrieved from https:\/\/en.wikipedia.org\/wiki\/Wikipedia:Who_writes_Wikipedia%3F."},{"key":"e_1_3_2_126_2","unstructured":"Wikipedia contributors. 2022. Wikipedians. Retrieved from https:\/\/en.wikipedia.org\/wiki\/Wikipedia:Wikipedians."},{"key":"e_1_3_2_127_2","doi-asserted-by":"publisher","DOI":"10.1080\/13557858.2019.1620176"},{"key":"e_1_3_2_128_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.41"},{"key":"e_1_3_2_129_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.45"},{"key":"e_1_3_2_130_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-acl.435"},{"key":"e_1_3_2_131_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3449830"},{"key":"e_1_3_2_132_2","article-title":"OPT: Open pre-trained transformer language models","volume":"2205","author":"Zhang Susan","year":"2022","unstructured":"Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. OPT: Open pre-trained transformer language models. arXiv abs\/2205.01068 (2022).","journal-title":"arXiv"},{"key":"e_1_3_2_133_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.11"}],"container-title":["Journal of Data and Information Quality"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3597307","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3597307","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:06Z","timestamp":1750182546000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3597307"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,22]]},"references-count":132,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,6,30]]}},"alternative-id":["10.1145\/3597307"],"URL":"https:\/\/doi.org\/10.1145\/3597307","relation":{},"ISSN":["1936-1955","1936-1963"],"issn-type":[{"value":"1936-1955","type":"print"},{"value":"1936-1963","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,22]]},"assertion":[{"value":"2022-12-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-02","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}