{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,3]],"date-time":"2026-08-03T19:51:37Z","timestamp":1785786697877,"version":"3.56.0"},"reference-count":100,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,12,12]],"date-time":"2024-12-12T00:00:00Z","timestamp":1733961600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,12,12]],"date-time":"2024-12-12T00:00:00Z","timestamp":1733961600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100000865","name":"Bill and Melinda Gates Foundation","doi-asserted-by":"publisher","award":["OPP114"],"award-info":[{"award-number":["OPP114"]}],"id":[{"id":"10.13039\/100000865","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000865","name":"Bill and Melinda Gates Foundation","doi-asserted-by":"publisher","award":["OPP114"],"award-info":[{"award-number":["OPP114"]}],"id":[{"id":"10.13039\/100000865","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Nat Comput Sci"],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Social identity biases, particularly the tendency to favor one\u2019s own group (ingroup solidarity) and derogate other groups (outgroup hostility), are deeply rooted in human psychology and social behavior. However, it is unknown if such biases are also present in artificial intelligence systems. Here we show that large language models (LLMs) exhibit patterns of social identity bias, similarly to humans. By administering sentence completion prompts to 77 different LLMs (for instance, \u2018We are\u2026\u2019), we demonstrate that nearly all base models and some instruction-tuned and preference-tuned models display clear ingroup favoritism and outgroup derogation. These biases manifest both in controlled experimental settings and in naturalistic human\u2013LLM conversations. However, we find that careful curation of training data and specialized fine-tuning can substantially reduce bias levels. These findings have important implications for developing more equitable artificial intelligence systems and highlight the urgent need to understand how human\u2013LLM interactions might reinforce existing social biases.<\/jats:p>","DOI":"10.1038\/s43588-024-00741-1","type":"journal-article","created":{"date-parts":[[2024,12,12]],"date-time":"2024-12-12T10:04:34Z","timestamp":1733997874000},"page":"65-75","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":84,"title":["Generative language models exhibit social identity biases"],"prefix":"10.1038","volume":"5","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-7354-1088","authenticated-orcid":false,"given":"Tiancheng","family":"Hu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0636-5046","authenticated-orcid":false,"given":"Yara","family":"Kyrychenko","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Steve","family":"Rathje","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nigel","family":"Collier","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sander","family":"van der Linden","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8150-9305","authenticated-orcid":false,"given":"Jon","family":"Roozenbeek","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,12,12]]},"reference":[{"key":"741_CR1","unstructured":"Milmo, D. ChatGPT reaches 100 million users two months after launch The Guardian (2 February 2023); https:\/\/www.theguardian.com\/technology\/2023\/feb\/02\/chatgpt-100-million-users-open-ai-fastest-growing-app"},{"key":"741_CR2","unstructured":"Microsoft. Global online safety survey results (Microsoft, 2024); https:\/\/www.microsoft.com\/en-us\/DigitalSafety\/research\/global-online-safety-survey"},{"key":"741_CR3","doi-asserted-by":"publisher","unstructured":"Bordia, S. & Bowman, S. R. Identifying and reducing gender bias in word-level language models. In Proc. 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop 7\u201315 (ACL, 2019); https:\/\/doi.org\/10.18653\/v1\/N19-3002","DOI":"10.18653\/v1\/N19-3002"},{"key":"741_CR4","doi-asserted-by":"publisher","unstructured":"Abid, A., Farooqi, M. & Zou, J. Persistent anti-muslim bias in large language models. In Proc. 2021 AAAI\/ACM Conference on AI, Ethics and Society 298\u2013306 (ACM, 2021); https:\/\/doi.org\/10.1145\/3461702.3462624","DOI":"10.1145\/3461702.3462624"},{"key":"741_CR5","doi-asserted-by":"crossref","unstructured":"Ahn, J. & Oh, A. Mitigating language-dependent ethnic bias in BERT. In Proc. 2021 Conference on Empirical Methods in Natural Language Processing 533\u2013549 (ACL, 2021); https:\/\/aclanthology.org\/2021.emnlp-main.42","DOI":"10.18653\/v1\/2021.emnlp-main.42"},{"key":"741_CR6","doi-asserted-by":"publisher","first-page":"147","DOI":"10.1038\/s41586-024-07856-5","volume":"633","author":"V Hofmann","year":"2024","unstructured":"Hofmann, V., Kalluri, P. R., Jurafsky, D. & King, S. AI generates covertly racist decisions about people based on their dialect. Nature 633, 147\u2013154 (2024).","journal-title":"Nature"},{"key":"741_CR7","doi-asserted-by":"publisher","first-page":"405","DOI":"10.1093\/poq\/nfs038","volume":"76","author":"S Iyengar","year":"2012","unstructured":"Iyengar, S., Sood, G. & Lelkes, Y. Affect, not ideology: a social identity perspective on polarization. Public Opinion Q. 76, 405\u2013431 (2012).","journal-title":"Public Opinion Q."},{"key":"741_CR8","doi-asserted-by":"publisher","first-page":"129","DOI":"10.1146\/annurev-polisci-051117-073034","volume":"22","author":"S Iyengar","year":"2019","unstructured":"Iyengar, S., Lelkes, Y., Levendusky, M., Malhotra, N. & Westwood, S. J. The origins and consequences of affective polarization in the United States. Annu. Rev. Political Sci. 22, 129\u2013146 (2019).","journal-title":"Annu. Rev. Political Sci."},{"key":"741_CR9","unstructured":"Tajfel, H., & Turner, J. C. An integrative theory of intergroup conflict. In The Social Psychology of Intergroup Relations (eds Austin, W. G. & Worchel, S.) 33\u201337 (Brooks\/Cole, 1979)."},{"key":"741_CR10","volume-title":"Rediscovering the Social Group","author":"JC Turner","year":"1987","unstructured":"Turner, J. C., Hogg, M. A., Oakes, P. J., Reicher, S. D. & Wetherell, M. S. Rediscovering the Social Group: A Self-Categorization Theory. (Basil Blackwell, 1987)."},{"key":"741_CR11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/bs.aesp.2018.03.001","volume":"58","author":"DM Mackie","year":"2018","unstructured":"Mackie, D. M. & Smith, E. R. Intergroup emotions theory: production, regulation and modification of group-based emotions. Adv. Exp. Soc. Psychol 58, 1\u201369 (2018).","journal-title":"Adv. Exp. Soc. Psychol"},{"key":"741_CR12","unstructured":"Hogg, M. A. & Abrams, D. Social Identifications: A Social Psychology of Intergroup Relations and Group Processes (Taylor & Francis, 1988)."},{"key":"741_CR13","doi-asserted-by":"publisher","first-page":"149","DOI":"10.1002\/ejsp.2420010202","volume":"1","author":"H Tajfel","year":"1971","unstructured":"Tajfel, H., Billig, M. G., Bundy, R. P. & Flament, C. Social categorization and intergroup behaviour. Eur. J. Soc. Psychol. 1, 149\u2013178 (1971).","journal-title":"Eur. J. Soc. Psychol."},{"key":"741_CR14","doi-asserted-by":"publisher","first-page":"81","DOI":"10.1177\/1368430210375251","volume":"14","author":"B Pinter","year":"2011","unstructured":"Pinter, B. & Greenwald, A. G. A comparison of minimal group induction procedures. Group Process. Intergr. Relat. 14, 81\u201398 (2011).","journal-title":"Group Process. Intergr. Relat."},{"key":"741_CR15","doi-asserted-by":"publisher","first-page":"981","DOI":"10.1037\/0022-3514.57.6.981","volume":"57","author":"A Maass","year":"1989","unstructured":"Maass, A., Salvi, D., Arcuri, L. & Semin, G. Language use in intergroup contexts: the linguistic intergroup bias. J. Pers. Soc. Psychol. 57, 981\u2013993 (1989).","journal-title":"J. Pers. Soc. Psychol."},{"key":"741_CR16","doi-asserted-by":"publisher","first-page":"753","DOI":"10.1521\/soco.2006.24.6.753","volume":"24","author":"GT Viki","year":"2006","unstructured":"Viki, G. T. et al. Beyond secondary emotions: the infrahumanization of outgroups using human-related and animal-related words. Soc. Cogn. 24, 753\u2013775 (2006).","journal-title":"Soc. Cogn."},{"key":"741_CR17","doi-asserted-by":"publisher","first-page":"685","DOI":"10.1007\/s13347-020-00415-6","volume":"33","author":"S Cave","year":"2020","unstructured":"Cave, S. & Dihal, K. The whiteness of AI. Philos. Technol. 33, 685\u2013703 (2020).","journal-title":"Philos. Technol."},{"key":"741_CR18","doi-asserted-by":"publisher","DOI":"10.18574\/nyu\/9781479833641.001.0001","volume-title":"Algorithms of Oppression","author":"SU Noble","year":"2018","unstructured":"Noble, S. U. Algorithms of Oppression: How Search Engines Reinforce Racism (New York Univ. Press, 2018)."},{"key":"741_CR19","doi-asserted-by":"crossref","unstructured":"Bender, E. M., Gebru, T., McMillan-Major, A. & Shmitchell, S. On the dangers of stochastic parrots: can language models be too big? In Proc. 2021 ACM Conference on Fairness, Accountability and Transparency 610\u2013623 (ACM, 2021); https:\/\/dl.acm.org\/doi\/10.1145\/3442188.3445922","DOI":"10.1145\/3442188.3445922"},{"key":"741_CR20","doi-asserted-by":"publisher","first-page":"183","DOI":"10.1126\/science.aal4230","volume":"356","author":"A Caliskan","year":"2017","unstructured":"Caliskan, A., Bryson, J. J. & Narayanan, A. Semantics derived automatically from language corpora contain human-like biases. Science 356, 183\u2013186 (2017).","journal-title":"Science"},{"key":"741_CR21","doi-asserted-by":"publisher","first-page":"1526","DOI":"10.1038\/s41562-023-01659-w","volume":"7","author":"T Webb","year":"2023","unstructured":"Webb, T., Holyoak, K. J. & Lu, H. Emergent analogical reasoning in large language models. Nat. Hum. Behav. 7, 1526\u20131541 (2023).","journal-title":"Nat. Hum. Behav."},{"key":"741_CR22","doi-asserted-by":"publisher","first-page":"e2405460121","DOI":"10.1073\/pnas.2405460121","volume":"121","author":"M Kosinski","year":"2024","unstructured":"Kosinski, M. Evaluating large language models in theory of mind tasks. Proc. Natl Acad. Sci. USA 121, e2405460121 (2024).","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"741_CR23","doi-asserted-by":"crossref","unstructured":"Caron, G. & Srivastava, S. Manipulating the perceived personality traits of language models. In Findings of the Association for Computational Linguistics: EMNLP 2023 (eds Bouamor, H. et al.) 2370\u20132386 (ACL, 2023); https:\/\/aclanthology.org\/2023.findings-emnlp.156","DOI":"10.18653\/v1\/2023.findings-emnlp.156"},{"key":"741_CR24","doi-asserted-by":"publisher","first-page":"337","DOI":"10.1017\/pan.2023.2","volume":"31","author":"LP Argyle","year":"2023","unstructured":"Argyle, L. P. et al. Out of one, many: using language models to simulate human samples. Political Anal 31, 337\u2013351 (2023).","journal-title":"Political Anal"},{"key":"741_CR25","doi-asserted-by":"crossref","unstructured":"Park, J. S. et al. Generative agents: interactive simulacra of human behavior. In Proc. 36th Annual ACM Symposium on User Interface Software and Technology (UIST \u201923) Vol. 2, 1\u201322 (ACM, 2023).","DOI":"10.1145\/3586183.3606763"},{"key":"741_CR26","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-024-53755-0","volume":"14","author":"S Matz","year":"2024","unstructured":"Matz, S. et al. The potential of generative AI for personalized persuasion at scale. Sci. Rep. 14, 4692 (2024).","journal-title":"Sci. Rep."},{"key":"741_CR27","doi-asserted-by":"crossref","unstructured":"Jakesch, M., Bhat, A., Buschek, D., Zalmanson, L. & Naaman, M. Co-writing with opinionated language models affects users\u2019 views. In Proc. 2023 CHI Conference on Human Factors in Computing Systems 1\u201315 (ACM, 2023); https:\/\/dl.acm.org\/doi\/10.1145\/3544548.3581196","DOI":"10.1145\/3544548.3581196"},{"key":"741_CR28","doi-asserted-by":"crossref","unstructured":"Bowman, S. R. & Dahl, G. E. What will it take to fix benchmarking in natural language understanding? In Proc. 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (eds Toutanova, K. et al.) 4843\u20134855 (ACL, 2021); https:\/\/aclanthology.org\/2021.naacl-main.385","DOI":"10.18653\/v1\/2021.naacl-main.385"},{"key":"741_CR29","unstructured":"Anwar, U. et al. Foundational challenges in assuring alignment and safety of large language models. Trans. Mach. Learn. Res. (2024); https:\/\/openreview.net\/forum?id=oVTkOs8Pka"},{"key":"741_CR30","doi-asserted-by":"crossref","unstructured":"Blodgett, S. L., Barocas, S., Daum\u00e9 III, H. & Wallach, H. Language (technology) is power: a critical survey of \u2018bias\u2019 in NLP. In Proc. 58th Annual Meeting of the Association for Computational Linguistics (eds Jurafsky, D. et al.) 5454\u20135476 (ACL, 2020); https:\/\/aclanthology.org\/2020.acl-main.485","DOI":"10.18653\/v1\/2020.acl-main.485"},{"key":"741_CR31","doi-asserted-by":"crossref","unstructured":"Parrish, A. et al. BBQ: a hand-built bias benchmark for question answering. In Findings of the Association for Computational Linguistics: ACL 2022 (eds Muresan, S. et al.) 2086\u20132105 (ACL, 2022); https:\/\/aclanthology.org\/2022.findings-acl.165","DOI":"10.18653\/v1\/2022.findings-acl.165"},{"key":"741_CR32","unstructured":"Ganguli, D., Schiefer, N., Favaro, M. & Clark, J. Challenges in evaluating AI systems https:\/\/www.anthropic.com\/index\/evaluating-ai-systems (Anthropic, 2023)."},{"key":"741_CR33","unstructured":"Santurkar, S. et al. Whose opinions do language models reflect? In Proc. 40th International Conference on Machine Learning 1244, 29971\u201330004 (ACM, 2023)."},{"key":"741_CR34","doi-asserted-by":"crossref","unstructured":"Blodgett, S. L., Lopez, G., Olteanu, A., Sim, R. & Wallach, H. Stereotyping Norwegian salmon: an inventory of pitfalls in fairness benchmark datasets. In Proc. 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (eds Zong, C. et al.) 1004\u20131015 (ACL, 2021); https:\/\/aclanthology.org\/2021.acl-long.81","DOI":"10.18653\/v1\/2021.acl-long.81"},{"key":"741_CR35","unstructured":"Zhao, W. et al. WildChat: 1M ChatGPT interaction logs in the wild. In Proc. Twelfth International Conference on Learning Representations (ICLR, 2024); https:\/\/openreview.net\/forum?id=Bl8u7ZRlbM"},{"key":"741_CR36","unstructured":"Zheng, L. et al. LMSYS-Chat-1M: a large-scale real-world LLM conversation dataset. In Proc. Twelfth International Conference on Learning Representations (ICLR, 2024); https:\/\/openreview.net\/forum?id=BOfDKxfwt0"},{"key":"741_CR37","unstructured":"Brown, T. B. et al. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020 (eds Larochelle, H. et al.) (Curran Associates, 2020); https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html"},{"key":"741_CR38","unstructured":"Touvron, H. et al. Llama 2: open foundation and fine-tuned chat models. Preprint at https:\/\/arxiv.org\/abs\/2307.09288 (2023)."},{"key":"741_CR39","unstructured":"Biderman, S. et al. Pythia: a suite for analyzing large language models across training and scaling. In Proc. 40th International Conference on Machine Learning Vol. 102 (eds Krause, A. et al.) 2397\u20132430 (PMLR, 2023); https:\/\/proceedings.mlr.press\/v202\/biderman23a.html"},{"key":"741_CR40","unstructured":"The Gemma Team et al. Gemma: open models based on gemini research and technology. Preprint at https:\/\/arxiv.org\/abs\/2403.08295 (2024)."},{"key":"741_CR41","unstructured":"Jiang, A. Q. et al. Mixtral of experts. Preprint at https:\/\/arxiv.org\/abs\/2401.04088 (2024)."},{"key":"741_CR42","unstructured":"OpenAI et al. GPT-4 technical report. Preprint at https:\/\/arxiv.org\/abs\/2303.08774 (2024)."},{"key":"741_CR43","unstructured":"OpenAI. Introducing ChatGPT https:\/\/openai.com\/index\/chatgpt\/ (2022)."},{"key":"741_CR44","unstructured":"Conover, M. et al. Free Dolly: introducing the world\u2019s first truly open instruction-tuned LLM https:\/\/www.databricks.com\/blog\/2023\/04\/12\/dolly-first-open-commercially-viable-instruction-tuned-llm (2023)."},{"key":"741_CR45","unstructured":"Taori, R. et al. Stanford alpaca: an instruction-following LLaMA model. GitHub https:\/\/github.com\/tatsu-lab\/stanford_alpaca (2023)."},{"key":"741_CR46","unstructured":"Wang, G. et al. OpenChat: advancing open-source language models with mixed-quality data. In Proc. Twelfth International Conference on Learning Representations (ICLR, 2024); https:\/\/openreview.net\/forum?id=AOJyfhWYHf"},{"key":"741_CR47","doi-asserted-by":"publisher","first-page":"475","DOI":"10.1037\/0022-3514.59.3.475","volume":"59","author":"CW Perdue","year":"1990","unstructured":"Perdue, C. W. et al. Us and them: social categorization and the process of intergroup bias. J. Pers. Soc. Psychol. 59, 475\u2013486 (1990).","journal-title":"J. Pers. Soc. Psychol."},{"key":"741_CR48","first-page":"5485","volume":"21","author":"C Raffel","year":"2020","unstructured":"Raffel, C. et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 5485\u20135551 (2020).","journal-title":"J. Mach. Learn. Res."},{"key":"741_CR49","unstructured":"Liu, Y. et al. RoBERTa: a robustly optimized BERT pretraining approach. Preprint at https:\/\/arxiv.org\/abs\/1907.11692 (2019)."},{"key":"741_CR50","doi-asserted-by":"crossref","unstructured":"Loureiro, D., Barbieri, F., Neves, L., Espinosa Anke, L. & Camacho-Collados, J. TimeLMs: diachronic language models from Twitter. In Proc. 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations 251\u2013260 (ACL, 2022); https:\/\/aclanthology.org\/2022.acl-demo.25","DOI":"10.18653\/v1\/2022.acl-demo.25"},{"key":"741_CR51","doi-asserted-by":"publisher","first-page":"121","DOI":"10.1080\/19312458.2020.1869198","volume":"15","author":"W Van Atteveldt","year":"2021","unstructured":"Van Atteveldt, W., Van der Velden, M. A. & Boukes, M. The validity of sentiment analysis: comparing manual annotation, crowd-coding, dictionary approaches and machine learning algorithms. Commun. Methods Measures 15, 121\u2013140 (2021).","journal-title":"Commun. Methods Measures"},{"key":"741_CR52","doi-asserted-by":"publisher","first-page":"5514","DOI":"10.1287\/mnsc.2021.4156","volume":"68","author":"R Frankel","year":"2022","unstructured":"Frankel, R., Jennings, J. & Lee, J. Disclosure sentiment: machine learning vs. dictionary methods. Manage. Sci. 68, 5514\u20135532 (2022).","journal-title":"Manage. Sci."},{"key":"741_CR53","doi-asserted-by":"publisher","first-page":"e2308950121","DOI":"10.1073\/pnas.2308950121","volume":"121","author":"S Rathje","year":"2024","unstructured":"Rathje, S. et al. GPT is an effective tool for multilingual psychological text analysis. Proc. Natl Acad. Sci. USA 121, e2308950121 (2024).","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"741_CR54","doi-asserted-by":"crossref","unstructured":"Templin, M. C.Certain Language Skills in Children; their Development and Interrelationships (Univ. Minnesota Press, 1957).","DOI":"10.5749\/j.ctttv2st"},{"key":"741_CR55","unstructured":"Gao, L. et al. The Pile: an 800\u2009GB dataset of diverse text for language modeling. Preprint at https:\/\/arxiv.org\/abs\/2101.00027 (2020)."},{"key":"741_CR56","unstructured":"Gokaslan, A. & Cohen, V. OpenWebText corpus. GitHub http:\/\/Skylion007.github.io\/OpenWebTextCorpus (2019)."},{"key":"741_CR57","unstructured":"Thrush, T., Ngo, H., Lambert, N. & Kiela, D. Online language modelling data pipeline. GitHub https:\/\/github.com\/huggingface\/olm-datasets (2022)."},{"key":"741_CR58","unstructured":"Kaplan, J. et al. Scaling laws for neural language models. Preprint at https:\/\/arxiv.org\/abs\/2001.08361 (2020)."},{"key":"741_CR59","first-page":"1","volume":"25","author":"HW Chung","year":"2024","unstructured":"Chung, H. W. et al. Scaling instruction-finetuned language models. J. Mach. Learn. Res. 25, 1\u201353 (2024).","journal-title":"J. Mach. Learn. Res."},{"key":"741_CR60","unstructured":"Jiang, H., Beeferman, D., Roy, B. & Roy, D. CommunityLM: probing partisan worldviews from language models. In Proc. 29th International Conference on Computational Linguistics 6818\u20136826 (ACL, 2022); https:\/\/aclanthology.org\/2022.coling-1.593"},{"key":"741_CR61","doi-asserted-by":"publisher","first-page":"e2024292118","DOI":"10.1073\/pnas.2024292118","volume":"118","author":"S Rathje","year":"2021","unstructured":"Rathje, S., Van Bavel, J. J. & van der Linden, S. Out-group animosity drives engagement on social media. Proc. Natl Acad. Sci. USA 118, e2024292118 (2021).","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"741_CR62","doi-asserted-by":"publisher","first-page":"12","DOI":"10.1016\/j.electstud.2015.11.001","volume":"41","author":"AI Abramowitz","year":"2016","unstructured":"Abramowitz, A. I. & Webster, S. The rise of negative partisanship and the nationalization of US elections in the 21st century. Elect. Stud. 41, 12\u201322 (2016).","journal-title":"Elect. Stud."},{"key":"741_CR63","doi-asserted-by":"publisher","first-page":"8127","DOI":"10.1038\/s41467-024-52179-8","volume":"15","author":"Y Kyrychenko","year":"2024","unstructured":"Kyrychenko, Y., Brik, T., van der Linden, S. & Roozenbeek, J. Social identity correlates of social media engagement before and after the 2022 Russian invasion of Ukraine. Nat. Commun. 15, 8127 (2024).","journal-title":"Nat. Commun."},{"key":"741_CR64","unstructured":"Bolukbasi, T., Chang, K.-W., Zou, J. Y., Saligrama, V. & Kalai, A. T. Man is to computer programmer as woman is to homemaker? Debiasing word embeddings. In Advances in Neural Information Processing Systems Vol. 29 (Curran Associates, 2016); https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2016\/hash\/a486cd07e4ac3d270571622f4f316ec5-Abstract.html"},{"key":"741_CR65","doi-asserted-by":"publisher","first-page":"E3635","DOI":"10.1073\/pnas.1720347115","volume":"115","author":"N Garg","year":"2018","unstructured":"Garg, N., Schiebinger, L., Jurafsky, D. & Zou, J. Word embeddings quantify 100 years of gender and ethnic stereotypes. Proc. Natl Acad. Sci. USA 115, E3635\u2013E3644 (2018).","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"741_CR66","doi-asserted-by":"crossref","unstructured":"Nadeem, M., Bethke, A. & Reddy, S. StereoSet: measuring stereotypical bias in pretrained language models. In Proc. 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (eds Zong, C. et al.) 5356\u20135371 (ACL, 2021); https:\/\/aclanthology.org\/2021.acl-long.416","DOI":"10.18653\/v1\/2021.acl-long.416"},{"key":"741_CR67","unstructured":"Liang, P. P., Wu, C., Morency, L.-P. & Salakhutdinov, R. Towards understanding and mitigating social biases in language models. In Proc. 38th International Conference on Machine Learning 139 (eds Meila, M. & Zhang, T.) 6565\u20136576 (PMLR, 2021); https:\/\/proceedings.mlr.press\/v139\/liang21a.html"},{"key":"741_CR68","doi-asserted-by":"publisher","first-page":"409","DOI":"10.1111\/j.1468-2958.1993.tb00308.x","volume":"19","author":"K Fiedler","year":"1993","unstructured":"Fiedler, K., Semin, G. R. & Finkenauer, C. The battle of words between gender groups: a language-based approach to intergroup processes. Hum. Commun. Res. 19, 409\u2013441 (1993).","journal-title":"Hum. Commun. Res."},{"key":"741_CR69","doi-asserted-by":"publisher","first-page":"116","DOI":"10.1037\/0022-3514.68.1.116","volume":"68","author":"A Maass","year":"1995","unstructured":"Maass, A., Milesi, A., Zabbini, S. & Stahlberg, D. Linguistic intergroup bias: differential expectancies or in-group protection? J. Pers. Soc. Psychol. 68, 116\u2013126 (1995).","journal-title":"J. Pers. Soc. Psychol."},{"key":"741_CR70","unstructured":"Bai, Y. et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. Preprint at https:\/\/arxiv.org\/abs\/2204.05862 (2022)."},{"key":"741_CR71","unstructured":"Sharma, M. et al. Towards understanding sycophancy in language models. In Proc. Twelfth International Conference on Learning Representations (ICLR, 2024); https:\/\/openreview.net\/forum?id=tvhaxkMKAn"},{"key":"741_CR72","unstructured":"Laban, P., Murakhovs\u2019ka, L., Xiong, C. & Wu, C.-S. Are you sure? Challenging LLMs leads to performance drops in the flipflop experiment. Preprint at https:\/\/arxiv.org\/abs\/2311.08596 (2024)."},{"key":"741_CR73","doi-asserted-by":"crossref","unstructured":"Feng, S., Park, C. Y., Liu, Y. & Tsvetkov, Y. From pretraining data to language models to downstream tasks: tracking the trails of political biases leading to unfair NLP models. In Proc. 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 11737\u201311762 (ACL, 2023); https:\/\/aclanthology.org\/2023.acl-long.656","DOI":"10.18653\/v1\/2023.acl-long.656"},{"key":"741_CR74","unstructured":"Chu, E., Andreas, J., Ansolabehere, S. & Roy, D. Language models trained on media diets can predict public opinion. Preprint at https:\/\/arxiv.org\/abs\/2303.16779 (2023)."},{"key":"741_CR75","doi-asserted-by":"publisher","first-page":"398","DOI":"10.1126\/science.abp9364","volume":"381","author":"AM Guess","year":"2023","unstructured":"Guess, A. M. et al. How do social media feed algorithms affect attitudes and behavior in an election campaign? Science 381, 398\u2013404 (2023).","journal-title":"Science"},{"key":"741_CR76","unstructured":"Anil, C. et al. Many-shot jailbreaking https:\/\/www.anthropic.com\/research\/many-shot-jailbreaking (Anthropic, 2024)."},{"key":"741_CR77","unstructured":"Radford, A. et al. Language models are unsupervised multitask learners https:\/\/openai.com\/research\/better-language-models (OpenAI, 2019)."},{"key":"741_CR78","unstructured":"Dey, N. et al. Cerebras-GPT: open compute-optimal language models trained on the Cerebras Wafer-Scale Cluster. Preprint at https:\/\/arxiv.org\/abs\/2304.03208 (2023)."},{"key":"741_CR79","unstructured":"BigScience Workshop et al. BLOOM: A 176B-parameter open-access multilingual language model. Preprint at https:\/\/arxiv.org\/abs\/2211.05100 (2023)."},{"key":"741_CR80","unstructured":"Touvron, H. et al. LLaMA: open and efficient foundation language models. Preprint at https:\/\/arxiv.org\/abs\/2302.13971 (2023)."},{"key":"741_CR81","unstructured":"Zhang, S. et al. OPT: open pre-trained transformer language models. Preprint at https:\/\/arxiv.org\/abs\/2205.01068 (2022)."},{"key":"741_CR82","unstructured":"Jiang, A. Q. et al. Mistral 7b. Preprint at https:\/\/arxiv.org\/abs\/2310.06825 (2023)."},{"key":"741_CR83","unstructured":"Groeneveld, D. et al. OLMo: Accelerating the science of language models. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds Ku, L.-W. et al.) 15789\u201315809 (ACL, 2024); https:\/\/aclanthology.org\/2024.acl-long.841"},{"key":"741_CR84","doi-asserted-by":"crossref","unstructured":"Muennighoff, N. et al. Crosslingual generalization through multitask finetuning. In Proc. 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 15991\u201316111 (Association for Computational Linguistics, 2023).","DOI":"10.18653\/v1\/2023.acl-long.891"},{"key":"741_CR85","unstructured":"Iyer, S. et al. OPT-IML: scaling language model instruction meta learning through the lens of generalization. Preprint at https:\/\/arxiv.org\/abs\/2212.12017 (2023)."},{"key":"741_CR86","unstructured":"Tay, Y. et al. UL2: unifying language learning paradigms. In The Eleventh International Conference on Learning Representations (ICLR, 2023); https:\/\/openreview.net\/forum?id=6ruVLB727MC"},{"key":"741_CR87","unstructured":"AI21studio. Announcing Jurassic-2 and task-specific APIs https:\/\/www.ai21.com\/blog\/introducing-j2 (AI21, 2023)."},{"key":"741_CR88","unstructured":"Ivison, H. et al. Camels in a changing climate: enhancing LM adaptation with TULU 2. Preprint at https:\/\/arxiv.org\/abs\/2311.10702 (2023)."},{"key":"741_CR89","unstructured":"Tunstall, L. et al. Zephyr: direct distillation of LM alignment. Preprint at https:\/\/arxiv.org\/abs\/2310.16944 (2023)."},{"key":"741_CR90","unstructured":"Zhu, B. et al. Starling-7B: improving helpfulness and harmlessness with RLAIF. In Proc. First Conference on Language Modeling https:\/\/openreview.net\/forum?id=GqDntYTTbk (2024)."},{"key":"741_CR91","unstructured":"Anil, R. et al. Palm 2 technical report. Preprint at https:\/\/arxiv.org\/abs\/2305.10403 (2023)."},{"key":"741_CR92","unstructured":"Wolf, T. et al. Transformers: state-of-the-art natural language processing. In Proc. 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations 38\u201345 (ACL, 2020); https:\/\/aclanthology.org\/2020.emnlp-demos.6"},{"key":"741_CR93","unstructured":"Holtzman, A., Buys, J., Du, L., Forbes, M. & Choi, Y. The curious case of neural text degeneration. In Proc. International Conference on Learning Representations (ICLR, 2020); https:\/\/openreview.net\/forum?id=rygGQyrFvH"},{"key":"741_CR94","unstructured":"Dettmers, T., Lewis, M., Belkada, Y. & Zettlemoyer, L. GPT3.int8(): 8-bit matrix multiplication for transformers at scale. In Proc. Advances in Neural Information Processing Systems 35 (eds Koyejo, S. et al.) 30318\u201330332 (Curran Associates, 2022); https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/file\/c3ba4962c05c49636d4c6206a97e9c8a-Paper-Conference.pdf"},{"key":"741_CR95","doi-asserted-by":"publisher","first-page":"216","DOI":"10.1609\/icwsm.v8i1.14550","volume":"8","author":"C Hutto","year":"2014","unstructured":"Hutto, C. & Gilbert, E. VADER: a parsimonious rule-based model for sentiment analysis of social media text. Proc. Int. AAAI Conf. Web Social Media 8, 216\u2013225 (2014).","journal-title":"Proc. Int. AAAI Conf. Web Social Media"},{"key":"741_CR96","unstructured":"\u212brup Nielsen, F. A new ANEW: evaluation of a word list for sentiment analysis in microblogs. In Proc. ESWC2011 Workshop on \u2018Making Sense of Microposts\u2019: Big Things Come in Small Packages, CEUR Workshop Proceedings Vol. 718 (eds Rowe, M. et al.) 93\u201398 (2011)."},{"key":"741_CR97","unstructured":"Loria, S. textblob Documentation, release 0.18.0.post0 edn https:\/\/readthedocs.org\/projects\/textblob\/downloads\/pdf\/dev\/ (Readthedocs, 2024)."},{"key":"741_CR98","doi-asserted-by":"crossref","unstructured":"Potts, C., Wu, Z., Geiger, A. & Kiela, D. DynaSent: A dynamic benchmark for sentiment analysis. Proc. 59th Annu. Meet. Assoc. Comput. Linguist. 11th Int. Jt. Conf. Nat. Lang. Process. Vol. 1: Long Pap. 2388\u20132404 (2021).","DOI":"10.18653\/v1\/2021.acl-long.186"},{"key":"741_CR99","volume-title":"The Development and Psychometric Properties of LIWC-22","author":"RL Boyd","year":"2022","unstructured":"Boyd, R. L., Ashokkumar, A., Seraj, S. & Pennebaker, J. W. The Development and Psychometric Properties of LIWC-22. (Univ. Texas at Austin, 2022)."},{"key":"741_CR100","doi-asserted-by":"publisher","unstructured":"Hu, T. et al. Generative language models exhibit social identity biases https:\/\/doi.org\/10.17605\/OSF.IO\/9HT32 (OSF, 2024).","DOI":"10.17605\/OSF.IO\/9HT32"}],"container-title":["Nature Computational Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.nature.com\/articles\/s43588-024-00741-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s43588-024-00741-1","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s43588-024-00741-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,28]],"date-time":"2025-01-28T23:08:39Z","timestamp":1738105719000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.nature.com\/articles\/s43588-024-00741-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,12]]},"references-count":100,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,1]]}},"alternative-id":["741"],"URL":"https:\/\/doi.org\/10.1038\/s43588-024-00741-1","relation":{},"ISSN":["2662-8457"],"issn-type":[{"value":"2662-8457","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,12]]},"assertion":[{"value":"22 May 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 November 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 December 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}]}}