{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,20]],"date-time":"2026-08-20T16:29:26Z","timestamp":1787243366520,"version":"build-2736575974"},"reference-count":62,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,9,4]],"date-time":"2024-09-04T00:00:00Z","timestamp":1725408000000},"content-version":"vor","delay-in-days":247,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,9,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>One widely cited barrier to the adoption of LLMs as proxies for humans in subjective tasks is their sensitivity to prompt wording\u2014but interestingly, humans also display sensitivities to instruction changes in the form of response biases. We investigate the extent to which LLMs reflect human response biases, if at all. We look to survey design, where human response biases caused by changes in the wordings of \u201cprompts\u201d have been extensively explored in social psychology literature. Drawing from these works, we design a dataset and framework to evaluate whether LLMs exhibit human-like response biases in survey questionnaires. Our comprehensive evaluation of nine models shows that popular open and commercial LLMs generally fail to reflect human-like behavior, particularly in models that have undergone RLHF. Furthermore, even if a model shows a significant change in the same direction as humans, we find that they are sensitive to perturbations that do not elicit significant changes in humans. These results highlight the pitfalls of using LLMs as human proxies, and underscore the need for finer-grained characterizations of model behavior.1<\/jats:p>","DOI":"10.1162\/tacl_a_00685","type":"journal-article","created":{"date-parts":[[2024,9,4]],"date-time":"2024-09-04T09:35:08Z","timestamp":1725442508000},"page":"1011-1026","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":62,"title":["Do LLMs Exhibit Human-like Response Biases? A Case Study in Survey Design"],"prefix":"10.1162","volume":"12","author":[{"given":"Lindia","family":"Tjuatja","sequence":"first","affiliation":[{"name":"Carnegie Mellon University, USA. lindiat@andrew.cmu.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Valerie","family":"Chen","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, USA. vchen2@andrew.cmu.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tongshuang","family":"Wu","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ameet","family":"Talwalkwar","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Graham","family":"Neubig","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,9,4]]},"reference":[{"key":"2024090413342128400_bib1","first-page":"337","article-title":"Using large language models to simulate multiple humans and replicate human subject studies","volume-title":"International Conference on Machine Learning","author":"Aher","year":"2023"},{"issue":"1","key":"2024090413342128400_bib2","doi-asserted-by":"publisher","first-page":"3","DOI":"10.5001\/omj.2014.02","article-title":"Patient satisfaction survey as a tool towards quality improvement","volume":"29","author":"Al-Abri","year":"2014","journal-title":"Oman Medical Journal"},{"key":"2024090413342128400_bib3","first-page":"819","article-title":"Out of one, many: Using language models to simulate human samples","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Argyle","year":"2022"},{"issue":"2","key":"2024090413342128400_bib4","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1086\/269200","article-title":"Response effects in mail surveys","volume":"54","author":"Ayidiya","year":"1990","journal-title":"Public Opinion Quarterly"},{"key":"2024090413342128400_bib5","article-title":"Synthetic and natural noise both break neural machine translation","author":"Belinkov","year":"2017","journal-title":"arXiv preprint arXiv:1711.02173"},{"key":"2024090413342128400_bib6","volume-title":"Questionnaire Design: How to Plan, Structure and Write Survey Material for Effective Market Research","author":"Brace","year":"2018"},{"key":"2024090413342128400_bib7","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2024090413342128400_bib8","first-page":"1764","article-title":"Use-case-grounded simulations for explanation evaluation","volume-title":"Advances in Neural Information Processing Systems","author":"Chen","year":"2022"},{"issue":"1","key":"2024090413342128400_bib9","article-title":"Peer reviewed: A catalog of biases in questionnaires","volume":"2","author":"Choi","year":"2005","journal-title":"Preventing Chronic Disease"},{"key":"2024090413342128400_bib10","article-title":"Language models trained on media diets can predict public opinion","author":"Chu","year":"2023","journal-title":"arXiv preprint arXiv:2303.16779"},{"issue":"4","key":"2024090413342128400_bib11","doi-asserted-by":"publisher","first-page":"407","DOI":"10.1177\/002224378001700401","article-title":"The optimal number of response alternatives for a scale: A review","volume":"17","author":"Cox","year":"1980","journal-title":"Journal of Marketing Research"},{"key":"2024090413342128400_bib12","article-title":"Language models show human-like content effects on reasoning","author":"Dasgupta","year":"2022","journal-title":"arXiv preprint arXiv:2207.07051"},{"key":"2024090413342128400_bib13","doi-asserted-by":"publisher","DOI":"10.1016\/j.tics.2023.04.008","article-title":"Can AI language models replace human participants?","author":"Dillion","year":"2023","journal-title":"Trends in Cognitive Sciences"},{"key":"2024090413342128400_bib14","article-title":"Towards measuring the representation of subjective global opinions in language models","author":"Durmus","year":"2023","journal-title":"arXiv preprint arXiv:2306.16388"},{"key":"2024090413342128400_bib15","doi-asserted-by":"publisher","first-page":"1643","DOI":"10.1162\/tacl_a_00626","article-title":"Bridging the gap: A survey on integrating (human) feedback for natural language generation","volume":"11","author":"Fernandes","year":"2023","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024090413342128400_bib16","first-page":"3816","article-title":"Making pre-trained language models better few-shot learners","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Gao","year":"2021"},{"issue":"30","key":"2024090413342128400_bib17","doi-asserted-by":"publisher","first-page":"e2305016120","DOI":"10.1073\/pnas.2305016120","article-title":"ChatGPT outperforms crowd workers for text-annotation tasks","volume":"120","author":"Gilardi","year":"2023","journal-title":"Proceedings of the National Academy of Sciences of the United States of America"},{"issue":"1","key":"2024090413342128400_bib18","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1207\/s15328023top1401_11","article-title":"Social desirability bias: A demonstration and technique for its reduction","volume":"14","author":"Gordon","year":"1987","journal-title":"Teaching of Psychology"},{"issue":"2","key":"2024090413342128400_bib19","doi-asserted-by":"publisher","first-page":"278","DOI":"10.1287\/opre.28.2.278","article-title":"Intensity measures of consumer preference","volume":"28","author":"Hauser","year":"1980","journal-title":"Operations Research"},{"key":"2024090413342128400_bib20","doi-asserted-by":"publisher","first-page":"102","DOI":"10.1007\/978-1-4612-4798-2_6","article-title":"Response effects in surveys","volume-title":"Social information processing and survey methodology","author":"Hippler","year":"1987"},{"key":"2024090413342128400_bib21","doi-asserted-by":"crossref","unstructured":"John J.\n              Horton\n            \n          . 2023. Large language models as simulated economic agents: What can we learn from homo silicus?Working Paper 31122, National Bureau of Economic Research. 10.3386\/w31122","DOI":"10.3386\/w31122"},{"key":"2024090413342128400_bib22","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3544548.3580688","article-title":"Evaluating large language models in generating synthetic HCI research data: A case study","volume-title":"Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems","author":"H\u00e4m\u00e4l\u00e4inen","year":"2023"},{"key":"2024090413342128400_bib23","doi-asserted-by":"publisher","first-page":"423","DOI":"10.1162\/tacl_a_00324","article-title":"How can we know what language models know?","volume":"8","author":"Jiang","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024090413342128400_bib24","first-page":"11785","article-title":"Capturing failures of large language models via human cognitive biases","volume-title":"Advances in Neural Information Processing Systems","author":"Jones","year":"2022"},{"issue":"1","key":"2024090413342128400_bib25","doi-asserted-by":"publisher","first-page":"42","DOI":"10.2307\/2981421","article-title":"The effect of the question on survey responses: A review","volume":"145","author":"Kalton","year":"1982","journal-title":"Journal of the Royal Statistical Society Series A: Statistics in Society"},{"key":"2024090413342128400_bib26","article-title":"AI-augmented surveys: Leveraging large language models for opinion prediction in nationally representative surveys","author":"Kim","year":"2023","journal-title":"arXiv preprint arXiv: 2305.09620"},{"key":"2024090413342128400_bib27","doi-asserted-by":"publisher","first-page":"8086","DOI":"10.18653\/v1\/2022.acl-long.556","article-title":"Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Yao","year":"2022"},{"key":"2024090413342128400_bib28","article-title":"Black box adversarial prompting for foundation models","volume-title":"The Second Workshop on New Frontiers in Adversarial Machine Learning","author":"Maus","year":"2023"},{"issue":"1","key":"2024090413342128400_bib29","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1177\/0049124191020001003","article-title":"Acquiescence and recency response-order effects in interview surveys","volume":"20","author":"McClendon","year":"1991","journal-title":"Sociological Methods & Research"},{"issue":"2","key":"2024090413342128400_bib30","doi-asserted-by":"publisher","first-page":"208","DOI":"10.1086\/268651","article-title":"Effects of question order on survey responses","volume":"45","author":"McFarland","year":"1981","journal-title":"Public Opinion Quarterly"},{"key":"2024090413342128400_bib31","article-title":"Inverse scaling: When bigger isn\u2019t better","author":"McKenzie","year":"2023","journal-title":"Transactions on Machine Learning Research"},{"key":"2024090413342128400_bib32","doi-asserted-by":"publisher","first-page":"13","DOI":"10.18653\/v1\/2022.conll-1.2","article-title":"Collateral facilitation in humans and language models","volume-title":"Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL)","author":"Michaelov","year":"2022"},{"issue":"1","key":"2024090413342128400_bib33","doi-asserted-by":"publisher","first-page":"53","DOI":"10.1086\/209466","article-title":"Do polls reflect opinions or do opinions reflect polls? The impact of political polling on voters\u2019 expectations, preferences, and behavior","volume":"23","author":"Morwitz","year":"1996","journal-title":"Journal of Consumer Research"},{"issue":"3","key":"2024090413342128400_bib34","doi-asserted-by":"publisher","DOI":"10.29115\/SP-2014-0013","article-title":"Response order effects in the youth tobacco survey: Results of a split-ballot experiment","volume":"7","author":"O\u2019Halloran","year":"2014","journal-title":"Survey Practice"},{"key":"2024090413342128400_bib35","volume-title":"Middle Alternatives, Acquiescence, and the Quality of Questionnaire Data","author":"O\u2019Muircheartaigh","year":"2001"},{"key":"2024090413342128400_bib36","first-page":"27730","article-title":"Training language models to follow instructions with human feedback","volume":"35","author":"Ouyang","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2024090413342128400_bib37","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3586183.3606763","article-title":"Generative agents: Interactive simulacra of human behavior","volume-title":"Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology","author":"Park","year":"2023"},{"key":"2024090413342128400_bib38","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3526113.3545616","article-title":"Social simulacra: Creating populated prototypes for social computing systems","volume-title":"Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology","author":"Park","year":"2022"},{"key":"2024090413342128400_bib39","article-title":"Artificial intelligence in psychology research","author":"Park","year":"2023","journal-title":"arXiv preprint arXiv: 2302.07267"},{"key":"2024090413342128400_bib40","article-title":"Ignore previous prompt: Attack techniques for language models","author":"Perez","year":"2022","journal-title":"arXiv preprint arXiv:2211.09527"},{"key":"2024090413342128400_bib41","article-title":"Large language models sensitivity to the order of options in multiple-choice questions","author":"Pezeshkpour","year":"2023","journal-title":"arXiv preprint arXiv:2308.11483"},{"key":"2024090413342128400_bib42","article-title":"Combating adversarial misspellings with robust word recognition","author":"Pruthi","year":"2019","journal-title":"arXiv preprint arXiv:1905.11268"},{"issue":"1","key":"2024090413342128400_bib43","doi-asserted-by":"publisher","first-page":"26","DOI":"10.1109\/MAES.2007.327521","article-title":"The significance of letter position in word recognition","volume":"22","author":"Rawlinson","year":"2007","journal-title":"IEEE Aerospace and Electronic Systems Magazine"},{"key":"2024090413342128400_bib44","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.10970","article-title":"Robsut wrod reocginiton via semi-character recurrent neural network","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Sakaguchi","year":"2017"},{"key":"2024090413342128400_bib45","article-title":"Multitask prompted training enables zero-shot task generalization","volume-title":"International Conference on Learning Representations","author":"Sanh","year":"2022"},{"key":"2024090413342128400_bib46","article-title":"Whose opinions do language models reflect?","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Santurkar","year":"2023"},{"key":"2024090413342128400_bib47","article-title":"Evaluating the moral beliefs encoded in LLMs","volume-title":"Thirty-seventh Conference on Neural Information Processing Systems","author":"Scherrer","year":"2023"},{"key":"2024090413342128400_bib48","volume-title":"Questions and Answers in Attitude Surveys: Experiments on Question Form, Wording, and Context","author":"Schuman","year":"1996"},{"key":"2024090413342128400_bib49","doi-asserted-by":"publisher","first-page":"187","DOI":"10.1007\/978-1-4612-2848-6_13","article-title":"A cognitive model of response-order effects in survey measurement","volume-title":"Context Effects in Social and Psychological Research","author":"Schwarz","year":"1992"},{"key":"2024090413342128400_bib50","article-title":"Quantifying language models\u2019 sensitivity to spurious features in prompt design or: How I learned to start worrying about prompt formatting","author":"Sclar","year":"2023","journal-title":"arXiv preprint arXiv:2310 .11324"},{"key":"2024090413342128400_bib51","doi-asserted-by":"publisher","first-page":"1031","DOI":"10.1162\/tacl_a_00504","article-title":"Structural persistence in language models: Priming as a window into abstract language representations","volume":"10","author":"Sinclair","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024090413342128400_bib52","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.starsem-1.14","article-title":"Syntax and semantics meet in the \u201cmiddle\u201d: Probing the syntax-semantics interface of LMs through agentivity","volume-title":"STARSEM","author":"Tjuatja","year":"2023"},{"key":"2024090413342128400_bib53","article-title":"Llama 2: Open foundation and fine-tuned chat models","author":"Touvron","year":"2023"},{"key":"2024090413342128400_bib54","article-title":"ChatGPT-4 outperforms experts and crowd workers in annotating political twitter messages with zero-shot learning","author":"T\u00f6rnberg","year":"2023"},{"key":"2024090413342128400_bib55","doi-asserted-by":"publisher","first-page":"2153","DOI":"10.18653\/v1\/D19-1221","article-title":"Universal adversarial triggers for attacking and analyzing NLP","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Wallace","year":"2019"},{"key":"2024090413342128400_bib56","doi-asserted-by":"publisher","first-page":"7662","DOI":"10.18653\/v1\/2023.findings-emnlp.514","article-title":"Are language models worse than humans at following prompts? It\u2019s complicated","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Webson","year":"2023"},{"key":"2024090413342128400_bib57","doi-asserted-by":"publisher","first-page":"2300","DOI":"10.18653\/v1\/2022.naacl-main.167","article-title":"Do prompt-based models really understand the meaning of their prompts?","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Webson","year":"2022"},{"key":"2024090413342128400_bib58","article-title":"Finetuned language models are zero-shot learners","volume-title":"International Conference on Learning Representations","author":"Wei","year":"2022"},{"key":"2024090413342128400_bib59","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume-title":"Advances in Neural Information Processing Systems","author":"Wei","year":"2022"},{"key":"2024090413342128400_bib60","volume-title":"An Introduction to Survey Research, Polling, and Data Analysis","author":"Weisberg","year":"1996"},{"key":"2024090413342128400_bib61","article-title":"On large language models\u2019 selection bias in multi-choice questions","author":"Zheng","year":"2023","journal-title":"arXiv preprint arXiv:2309.03882"},{"key":"2024090413342128400_bib62","article-title":"Universal and transferable adversarial attacks on aligned language models","author":"Zou","year":"2023"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00685\/2468689\/tacl_a_00685.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00685\/2468689\/tacl_a_00685.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,4]],"date-time":"2024-09-04T09:35:33Z","timestamp":1725442533000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00685\/124261\/Do-LLMs-Exhibit-Human-like-Response-Biases-A-Case"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":62,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00685","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}