{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T00:27:16Z","timestamp":1784766436936,"version":"3.55.0"},"reference-count":59,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,11,7]],"date-time":"2024-11-07T00:00:00Z","timestamp":1730937600000},"content-version":"vor","delay-in-days":311,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,11,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Due to the widespread use of large language models (LLMs), we need to understand whether they embed a specific \u201cworldview\u201d and what these views reflect. Recent studies report that, prompted with political questionnaires, LLMs show left-liberal leanings (Feng et al., 2023; Motoki et al., 2024). However, it is as yet unclear whether these leanings are reliable (robust to prompt variations) and whether the leaning is consistent across policies and political leaning. We propose a series of tests which assess the reliability and consistency of LLMs\u2019 stances on political statements based on a dataset of voting-advice questionnaires collected from seven EU countries and annotated for policy issues. We study LLMs ranging in size from 7B to 70B parameters and find that their reliability increases with parameter count. Larger models show overall stronger alignment with left-leaning parties but differ among policy programs: They show a (left-wing) positive stance towards environment protection, social welfare state, and liberal society but also (right-wing) law and order, with no consistent preferences in the areas of foreign policy and migration.<\/jats:p>","DOI":"10.1162\/tacl_a_00710","type":"journal-article","created":{"date-parts":[[2024,11,7]],"date-time":"2024-11-07T20:18:43Z","timestamp":1731010723000},"page":"1378-1400","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":15,"title":["Beyond Prompt Brittleness: Evaluating the Reliability and Consistency of Political Worldviews in LLMs"],"prefix":"10.1162","volume":"12","author":[{"given":"Tanise","family":"Ceron","sequence":"first","affiliation":[{"name":"Institute for Natural Language Processing, University of Stuttgart, Germany. tanise.ceron@ims.uni-stuttgart.de"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Neele","family":"Falk","sequence":"additional","affiliation":[{"name":"Institute for Natural Language Processing, University of Stuttgart, Germany. neele.falk@ims.uni-stuttgart.de"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ana","family":"Bari\u0107","sequence":"additional","affiliation":[{"name":"Faculty of Electrical Engineering and Computing, University of Zagreb, Croatia. ana.baric@fer.hr"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dmitry","family":"Nikolaev","sequence":"additional","affiliation":[{"name":"Department of Linguistics and English Language, University of Manchester, UK. dmitry.nikolaev@manchester.ac.uk"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sebastian","family":"Pad\u00f3","sequence":"additional","affiliation":[{"name":"Institute for Natural Language Processing, University of Stuttgart, Germany. pado@ims.uni-stuttgart.de"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,11,4]]},"reference":[{"key":"2024110720183700600_bib1","volume-title":"Standards for Educational and Psychological Testing","author":"American Educational Research Association","year":"1999"},{"key":"2024110720183700600_bib2","doi-asserted-by":"publisher","first-page":"114","DOI":"10.18653\/v1\/2023.c3nlp-1.12","article-title":"Probing pre-trained language models for cross-cultural differences in values","volume-title":"Proceedings of the First Workshop on Cross-Cultural Considerations in NLP (C3NLP)","author":"Arora","year":"2023"},{"key":"2024110720183700600_bib3","doi-asserted-by":"publisher","first-page":"325","DOI":"10.4324\/9781351154161-9","article-title":"Digital speech and democratic culture: A theory of freedom of expression for the information society","volume-title":"Law and Society Approaches to Cyberspace","author":"Balkin","year":"2017"},{"key":"2024110720183700600_bib4","article-title":"Simple linguistic inferences of large language models (LLMs): Blind spots and blinds","author":"Basmov","year":"2024","journal-title":"ArXiv"},{"key":"2024110720183700600_bib5","doi-asserted-by":"publisher","first-page":"610","DOI":"10.1145\/3442188.3445922","article-title":"On the dangers of stochastic parrots: Can language models be too big?","volume-title":"Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency","author":"Bender","year":"2021"},{"key":"2024110720183700600_bib6","doi-asserted-by":"publisher","DOI":"10.4324\/9780203028179","volume-title":"Party Policy in Modern Democracies","author":"Benoit","year":"2006"},{"key":"2024110720183700600_bib7","doi-asserted-by":"publisher","first-page":"5454","DOI":"10.18653\/v1\/2020.acl-main.485","article-title":"Language (technology) is power: A critical survey of \u201cbias\u201d in NLP","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Lin Blodgett","year":"2020"},{"key":"2024110720183700600_bib8","unstructured":"Ian\n              Budge\n            \n          . 2013. The standard right-left scale. Technical report, Comparative Manifesto Project."},{"key":"2024110720183700600_bib9","doi-asserted-by":"publisher","DOI":"10.1093\/oso\/9780199244003.001.0001","volume-title":"Mapping Policy Preferences: Estimates for Parties, Electors, and Governments 1945\u20131998","author":"Budge","year":"2001"},{"key":"2024110720183700600_bib10","doi-asserted-by":"publisher","first-page":"7874","DOI":"10.18653\/v1\/2023.findings-acl.499","article-title":"Additive manifesto decomposition: A policy domain aware method for understanding party positioning","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Ceron","year":"2023"},{"key":"2024110720183700600_bib11","first-page":"19","article-title":"Navigating the modern evaluation landscape: Considerations in benchmarks and frameworks for large language models (LLMs)","volume-title":"Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024): Tutorial Summaries","author":"Choshen","year":"2024"},{"issue":"70","key":"2024110720183700600_bib12","first-page":"1","article-title":"Scaling instruction-finetuned language models","volume":"25","author":"Chung","year":"2024","journal-title":"Journal of Machine Learning Research"},{"issue":"1","key":"2024110720183700600_bib13","doi-asserted-by":"publisher","DOI":"10.3384\/nejlt.2000-1533.2022.3505","article-title":"Bias identification and attribution in NLP models with regression and effect sizes","volume":"8","author":"Dayanik","year":"2022","journal-title":"Northern European Journal of Language Technology"},{"key":"2024110720183700600_bib14","article-title":"Questioning the survey responses of large language models","author":"Dominguez-Olmedo","year":"2023","journal-title":"ArXiv"},{"key":"2024110720183700600_bib15","article-title":"Towards measuring the representation of subjective global opinions in language models","author":"Durmus","year":"2023","journal-title":"ArXiv"},{"issue":"1822","key":"2024110720183700600_bib16","doi-asserted-by":"publisher","first-page":"20200145","DOI":"10.1098\/rstb.2020.0145","article-title":"Corrections of political misinformation: No evidence for an effect of partisan worldview in a US convenience sample","volume":"376","author":"Ecker","year":"2021","journal-title":"Philosophical Transactions of the Royal Society B: Biological Sciences"},{"key":"2024110720183700600_bib17","doi-asserted-by":"publisher","first-page":"3764","DOI":"10.18653\/v1\/2023.emnlp-main.230","article-title":"ROBBIE: Robust bias evaluation of large generative language models","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Esiobu","year":"2023"},{"issue":"1","key":"2024110720183700600_bib18","doi-asserted-by":"publisher","first-page":"93","DOI":"10.2307\/591118","article-title":"Measuring left-right and libertarian-authoritarian values in the British electorate","volume":"47","author":"Evans","year":"1996","journal-title":"The British Journal of Sociology"},{"key":"2024110720183700600_bib19","doi-asserted-by":"publisher","first-page":"11737","DOI":"10.18653\/v1\/2023.acl-long.656","article-title":"From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Feng","year":"2023"},{"key":"2024110720183700600_bib20","article-title":"The pile: An 800GB dataset of diverse text for language modeling","author":"Gao","year":"2020","journal-title":"ArXiv"},{"key":"2024110720183700600_bib21","doi-asserted-by":"publisher","first-page":"688","DOI":"10.18653\/v1\/E17-2109","article-title":"Unsupervised cross-lingual scaling of political texts","volume-title":"Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers","author":"Glava\u0161","year":"2017"},{"key":"2024110720183700600_bib22","doi-asserted-by":"publisher","first-page":"1862","DOI":"10.18653\/v1\/2023.emnlp-main.115","article-title":"\u201cFifty shades of bias\u201d: Normative ratings of gender bias in GPT generated English text","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Hada","year":"2023"},{"key":"2024110720183700600_bib23","doi-asserted-by":"publisher","DOI":"10.2139\/ssrn.4316084","article-title":"The political ideology of conversational AI: Converging evidence on ChatGPT\u2019s pro-environmental, left- libertarian orientation","author":"Hartmann","year":"2023","journal-title":"SSRN Electronic Journal"},{"issue":"4","key":"2024110720183700600_bib24","doi-asserted-by":"publisher","first-page":"39","DOI":"10.1002\/j.1662-6370.2001.tb00327.x","article-title":"Weltanschauung und ihre soziale Basis im Spiegel eidgen\u00f6ssischer Volksabstimmungen","volume":"7","author":"Hermann","year":"2001","journal-title":"Swiss Political Science Review"},{"key":"2024110720183700600_bib25","volume-title":"Atlas der politischen Landschaften: Ein weltanschauliches Portr\u00e4t der Schweiz","author":"Hermann","year":"2003"},{"key":"2024110720183700600_bib26","volume-title":"Political Ideologies: An Introduction","author":"Heywood","year":"2021"},{"key":"2024110720183700600_bib27","article-title":"The curious case of neural text degeneration","volume-title":"Proceedings of ICLR","author":"Holtzman","year":"2020"},{"issue":"8","key":"2024110720183700600_bib28","doi-asserted-by":"publisher","first-page":"e12432","DOI":"10.1111\/lnc3.12432","article-title":"Five sources of bias in natural language processing","volume":"15","author":"Hovy","year":"2021","journal-title":"Language and Linguistics Compass"},{"issue":"1","key":"2024110720183700600_bib29","doi-asserted-by":"publisher","first-page":"45","DOI":"10.2307\/2111335","article-title":"Political leadership and representation in West European democracies: A test of three models of voting","volume":"38","author":"Iversen","year":"1994","journal-title":"American Journal of Political Science"},{"key":"2024110720183700600_bib30","doi-asserted-by":"publisher","first-page":"308","DOI":"10.1057\/s41295-022-00305-5","article-title":"The changing relevance and meaning of left and right in 34 party systems from 1945 to 2020","volume":"21","author":"Jahn","year":"2023","journal-title":"Comparative European Politics"},{"key":"2024110720183700600_bib31","article-title":"Mistral 7b","author":"Jiang","year":"2023","journal-title":"ArXiv"},{"key":"2024110720183700600_bib32","doi-asserted-by":"publisher","first-page":"102420","DOI":"10.1016\/j.electstud.2021.102420","article-title":"Chapel Hill expert survey trend file, 1999\u20132019","volume":"75","author":"Jolly","year":"2022","journal-title":"Electoral Studies"},{"key":"2024110720183700600_bib33","doi-asserted-by":"publisher","first-page":"3631","DOI":"10.18653\/v1\/2022.naacl-main.266","article-title":"Prompt waywardness: The curious case of discretized interpretation of continuous prompts","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Khashabi","year":"2022"},{"key":"2024110720183700600_bib34","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511622014","volume-title":"The Transformation of European Social Democracy","author":"Kitschelt","year":"1994"},{"key":"2024110720183700600_bib35","doi-asserted-by":"publisher","DOI":"10.1145\/3582269.3615599","article-title":"Gender bias and stereotypes in large language models","volume-title":"Proceedings of The ACM Collective Intelligence Conference","author":"Kotek","year":"2023"},{"issue":"1","key":"2024110720183700600_bib36","doi-asserted-by":"publisher","DOI":"10.1145\/3369026","article-title":"Stance detection: A survey","volume":"53","author":"K\u00fc\u00e7\u00fck","year":"2020","journal-title":"ACM Computing Surveys"},{"issue":"2","key":"2024110720183700600_bib37","doi-asserted-by":"publisher","first-page":"311","DOI":"10.1017\/S0003055403000698","article-title":"Extracting policy positions from political texts using words as data","volume":"97","author":"Laver","year":"2003","journal-title":"American Political Science Review"},{"key":"2024110720183700600_bib38","doi-asserted-by":"publisher","DOI":"10.21203\/rs.3.rs-3996137\/v1","article-title":"Datasets for large language models: A comprehensive survey","author":"Liu","year":"2024","journal-title":"ArXiv"},{"issue":"6","key":"2024110720183700600_bib39","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3457607","article-title":"A survey on bias and fairness in machine learning","volume":"54","author":"Mehrabi","year":"2021","journal-title":"ACM Computing Surveys (CSUR)"},{"key":"2024110720183700600_bib40","doi-asserted-by":"publisher","first-page":"11048","DOI":"10.18653\/v1\/2022.emnlp-main.759","article-title":"Rethinking the role of demonstrations: What makes in-context learning work?","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Min","year":"2022"},{"key":"2024110720183700600_bib41","doi-asserted-by":"publisher","first-page":"933","DOI":"10.1162\/tacl_a_00681","article-title":"State of what art? A call for multi-prompt llm evaluation","volume":"12","author":"Mizrahi","year":"2024","journal-title":"Transactions of the Association for Computational Linguistics"},{"issue":"3","key":"2024110720183700600_bib42","doi-asserted-by":"publisher","first-page":"395","DOI":"10.1111\/j.1533-8525.2004.tb02296.x","article-title":"Structuring political opinions: Attitude consistency and democratic competence among the u.s. mass public","volume":"45","author":"Moskowitz","year":"2004","journal-title":"The Sociological Quarterly"},{"key":"2024110720183700600_bib43","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1007\/s11127-023-01097-2","article-title":"More human than human: Measuring ChatGPT political bias","volume":"198","author":"Motoki","year":"2024","journal-title":"Public Choice"},{"issue":"1","key":"2024110720183700600_bib44","doi-asserted-by":"publisher","DOI":"10.1038\/s41746-023-00939-z","article-title":"Large language models propagate race-based medicine","volume":"6","author":"Omiye","year":"2023","journal-title":"npj Digital Medicine"},{"issue":"3","key":"2024110720183700600_bib45","doi-asserted-by":"publisher","first-page":"511","DOI":"10.2307\/2111281","article-title":"The relationship between information, ideology, and voting behavior","volume":"31","author":"Palfrey","year":"1987","journal-title":"American Journal of Political Science"},{"issue":"1","key":"2024110720183700600_bib46","doi-asserted-by":"publisher","first-page":"7115633","DOI":"10.1155\/2024\/7115633","article-title":"The self-perception and political biases of chatgpt","volume":"2024","author":"Rutinowski","year":"2024","journal-title":"Human Behavior and Emerging Technologies"},{"key":"2024110720183700600_bib47","article-title":"Whose opinions do language models reflect?","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Santurkar","year":"2023"},{"key":"2024110720183700600_bib48","doi-asserted-by":"publisher","first-page":"5263","DOI":"10.18653\/v1\/2024.naacl-long.295","article-title":"You don\u2019t need a personality test to know these models are unreliable: Assessing the reliability of large language models on psychometric instruments","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Shu","year":"2024"},{"issue":"3","key":"2024110720183700600_bib49","doi-asserted-by":"publisher","first-page":"705","DOI":"10.1111\/j.1540-5907.2008.00338.x","article-title":"A scaling model for estimating time-series party positions from texts","volume":"52","author":"Slapin","year":"2008","journal-title":"American Journal of Political Science"},{"key":"2024110720183700600_bib50","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3465416.3483305","article-title":"A framework for understanding sources of harm throughout the machine learning life cycle","volume-title":"Proceedings of the 1st ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization","author":"Suresh","year":"2021"},{"issue":"1","key":"2024110720183700600_bib51","doi-asserted-by":"publisher","first-page":"151","DOI":"10.1111\/j.1540-5907.2007.00243.x","article-title":"Principle vs. pragmatism: Policy shifts and political competition","volume":"51","author":"Tavits","year":"2007","journal-title":"American Journal of Political Science"},{"key":"2024110720183700600_bib52","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00685","article-title":"Do llms exhibit human-like response biases? A case study in survey design","author":"Tjuatja","year":"2023","journal-title":"arXiv preprint arXiv:2311.04076"},{"key":"2024110720183700600_bib53","article-title":"Llama 2: Open foundation and fine-tuned chat models","author":"Touvron","year":"2023","journal-title":"ArXiv"},{"key":"2024110720183700600_bib54","first-page":"77013","article-title":"Evaluating Open-QA Evaluation","volume-title":"Advances in Neural Information Processing Systems","author":"Wang","year":"2023"},{"key":"2024110720183700600_bib55","article-title":"Not all countries celebrate Thanksgiving: On the cultural dominance in large language models","author":"Wang","year":"2023","journal-title":"ArXiv"},{"key":"2024110720183700600_bib56","doi-asserted-by":"publisher","first-page":"2300","DOI":"10.18653\/v1\/2022.naacl-main.167","article-title":"Do prompt-based models really understand the meaning of their prompts?","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Webson","year":"2022"},{"key":"2024110720183700600_bib57","doi-asserted-by":"publisher","first-page":"102821","DOI":"10.1016\/j.ijinfomgt.2024.102821","article-title":"Chatgpt usage in everyday life: A motivation-theoretic mixed-methods study","volume":"79","author":"Wolf","year":"2024","journal-title":"International Journal of Information Management"},{"key":"2024110720183700600_bib58","first-page":"12697","article-title":"Calibrate before use: Improving few-shot performance of language models","volume-title":"Proceedings of the 38th International Conference on Machine Learning","author":"Zhao","year":"2021"},{"key":"2024110720183700600_bib59","article-title":"Large language models are not robust multiple choice selectors","volume-title":"The Twelfth International Conference on Learning Representations","author":"Zheng","year":"2023"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00710\/2478626\/tacl_a_00710.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00710\/2478626\/tacl_a_00710.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,7]],"date-time":"2024-11-07T20:18:51Z","timestamp":1731010731000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00710\/125176\/Beyond-Prompt-Brittleness-Evaluating-the"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":59,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00710","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}