{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,27]],"date-time":"2026-08-27T06:07:14Z","timestamp":1787810834620,"version":"build-2784847793"},"reference-count":92,"publisher":"MIT Press","issue":"4","license":[{"start":{"date-parts":[[2024,7,30]],"date-time":"2024-07-30T00:00:00Z","timestamp":1722297600000},"content-version":"vor","delay-in-days":211,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,12,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>How should we compare the capabilities of language models (LMs) and humans? In this article, I draw inspiration from comparative psychology to highlight challenges in these comparisons. I focus on a case study: processing of recursively nested grammatical structures. Prior work suggests that LMs cannot process these structures as reliably as humans can. However, the humans were provided with instructions and substantial training, while the LMs were evaluated zero-shot. I therefore match the evaluation more closely. Providing large LMs with a simple prompt\u2014with substantially less content than the human training\u2014allows the LMs to consistently outperform the human results, even in more deeply nested conditions than were tested with humans. Furthermore, the effects of prompting are robust to the particular structures and vocabulary used in the prompt. Finally, reanalyzing the existing human data suggests that the humans may not perform above chance at the difficult structures initially. Thus, large LMs may indeed process recursively nested grammatical structures as reliably as humans, when evaluated comparably. This case study highlights how discrepancies in the evaluation methods can confound comparisons of language models and humans. I conclude by reflecting on the broader challenge of comparing human and model capabilities, and highlight an important difference between evaluating cognitive models and foundation models.<\/jats:p>","DOI":"10.1162\/coli_a_00525","type":"journal-article","created":{"date-parts":[[2024,7,30]],"date-time":"2024-07-30T15:02:55Z","timestamp":1722351775000},"page":"1441-1476","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":21,"title":["Can Language Models Handle Recursively Nested Grammatical Structures? A Case Study on Comparing Models and Humans"],"prefix":"10.1162","volume":"50","author":[{"given":"Andrew","family":"Lampinen","sequence":"first","affiliation":[{"name":"Google DeepMind. lampinen@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,12,1]]},"reference":[{"key":"2024122021045446000_bib1","article-title":"Neural machine translation by jointly learning to align and translate","author":"Bahdanau","year":"2014","journal-title":"arXiv preprint arXiv:1409.0473"},{"key":"2024122021045446000_bib2","doi-asserted-by":"publisher","first-page":"47","DOI":"10.4324\/9781410602367-7","article-title":"On the emergence of grammar from the lexicon","volume-title":"The Emergence of Language","author":"Bates","year":"2013"},{"issue":"11","key":"2024122021045446000_bib3","doi-asserted-by":"publisher","first-page":"e3000539","DOI":"10.1371\/journal.pbio.3000539","article-title":"All or nothing: No half-merge and the evolution of syntax","volume":"17","author":"Berwick","year":"2019","journal-title":"PLoS Biology"},{"issue":"6","key":"2024122021045446000_bib4","doi-asserted-by":"publisher","first-page":"e2218523120","DOI":"10.1073\/pnas.2218523120","article-title":"Using cognitive psychology to understand GPT-3","volume":"120","author":"Binz","year":"2023","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2024122021045446000_bib5","article-title":"On the opportunities and risks of foundation models","author":"Bommasani","year":"2021","journal-title":"arXiv preprint arXiv:2108.07258"},{"key":"2024122021045446000_bib6","doi-asserted-by":"publisher","first-page":"7484","DOI":"10.18653\/v1\/2022.acl-long.516","article-title":"The dangers of underclaiming: Reasons for caution when reporting how NLP systems fail","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Bowman","year":"2022"},{"key":"2024122021045446000_bib7","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"2","key":"2024122021045446000_bib8","doi-asserted-by":"publisher","first-page":"234","DOI":"10.1037\/0033-295X.113.2.234","article-title":"Becoming syntactic","volume":"113","author":"Chang","year":"2006","journal-title":"Psychological Review"},{"key":"2024122021045446000_bib9","first-page":"482","article-title":"Derivation by phase","volume-title":"An Annotated Syntax Reader","author":"Chomsky","year":"1999"},{"key":"2024122021045446000_bib10","volume-title":"Aspects of the Theory of Syntax","author":"Chomsky","year":"2014"},{"key":"2024122021045446000_bib11","doi-asserted-by":"publisher","DOI":"10.7551\/mitpress\/9780262527347.001.0001","volume-title":"The Minimalist Program","author":"Chomsky","year":"2014"},{"issue":"2","key":"2024122021045446000_bib12","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1207\/s15516709cog2302_2","article-title":"Toward a connectionist model of recursion in human linguistic performance","volume":"23","author":"Christiansen","year":"1999","journal-title":"Cognitive Science"},{"key":"2024122021045446000_bib13","doi-asserted-by":"publisher","first-page":"126","DOI":"10.1111\/j.1467-9922.2009.00538.x","article-title":"A usage-based approach to recursion in sentence processing","volume":"59","author":"Christiansen","year":"2009","journal-title":"Language Learning"},{"issue":"70","key":"2024122021045446000_bib14","first-page":"1","article-title":"Scaling instruction-finetuned language models","volume":"25","author":"Chung","year":"2024","journal-title":"Journal of Machine Learning Research"},{"issue":"3","key":"2024122021045446000_bib15","doi-asserted-by":"publisher","first-page":"972","DOI":"10.3758\/s13423-016-1161-z","article-title":"Beliefs and Bayesian reasoning","volume":"24","author":"Cohen","year":"2017","journal-title":"Psychonomic Bulletin & Review"},{"issue":"5","key":"2024122021045446000_bib16","doi-asserted-by":"publisher","first-page":"547","DOI":"10.1002\/wcs.131","article-title":"Recursion: What is it, who has it, and how did it evolve?","volume":"2","author":"Coolidge","year":"2011","journal-title":"Wiley Interdisciplinary Reviews: Cognitive Science"},{"key":"2024122021045446000_bib17","article-title":"Language models show human-like content effects on reasoning","author":"Dasgupta","year":"2022","journal-title":"arXiv preprint arXiv:2207.07051"},{"key":"2024122021045446000_bib18","doi-asserted-by":"publisher","first-page":"8698","DOI":"10.18653\/v1\/2024.naacl-long.482","article-title":"Investigating data contamination in modern benchmarks for large language models","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Deng","year":"2024"},{"issue":"3","key":"2024122021045446000_bib19","doi-asserted-by":"publisher","first-page":"831","DOI":"10.1037\/a0019634","article-title":"Assessing the belief bias effect with ROCs: It\u2019s a response bias effect","volume":"117","author":"Dube","year":"2010","journal-title":"Psychological Review"},{"key":"2024122021045446000_bib20","first-page":"199","article-title":"Recurrent neural network grammars","volume-title":"Proceedings of the North American Chapter of the Association for Computational Linguistics","author":"Dyer","year":"2016"},{"issue":"2","key":"2024122021045446000_bib21","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1007\/BF00114844","article-title":"Distributed representations, simple recurrent networks, and grammatical structure","volume":"7","author":"Elman","year":"1991","journal-title":"Machine Learning"},{"issue":"1","key":"2024122021045446000_bib22","doi-asserted-by":"publisher","first-page":"71","DOI":"10.1016\/0010-0277(93)90058-4","article-title":"Learning and development in neural networks: The importance of starting small","volume":"48","author":"Elman","year":"1993","journal-title":"Cognition"},{"key":"2024122021045446000_bib23","doi-asserted-by":"publisher","DOI":"10.7551\/mitpress\/5929.001.0001","volume-title":"Rethinking Innateness: A Connectionist Perspective on Development","author":"Elman","year":"1996"},{"key":"2024122021045446000_bib24","volume-title":"Bias in Human Reasoning: Causes and Consequences","author":"Evans","year":"1989"},{"issue":"9","key":"2024122021045446000_bib25","doi-asserted-by":"publisher","first-page":"1362","DOI":"10.1037\/xlm0000236","article-title":"The role of verb repetition in cumulative structural priming in comprehension","volume":"42","author":"Fine","year":"2016","journal-title":"Journal of Experimental Psychology: Learning, Memory, and Cognition"},{"issue":"10","key":"2024122021045446000_bib26","doi-asserted-by":"publisher","first-page":"e77661","DOI":"10.1371\/journal.pone.0077661","article-title":"Rapid expectation adaptation during syntactic comprehension","volume":"8","author":"Fine","year":"2013","journal-title":"PloS One"},{"issue":"43","key":"2024122021045446000_bib27","doi-asserted-by":"publisher","first-page":"26562","DOI":"10.1073\/pnas.1905334117","article-title":"Performance vs. competence in human\u2013machine comparisons","volume":"117","author":"Firestone","year":"2020","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2024122021045446000_bib28","volume-title":"Experimentology: An Open Science Approach to Experimental Psychology Methods","author":"Frank","year":"2024"},{"key":"2024122021045446000_bib29","first-page":"14035","article-title":"Interpretability illusions in the generalization of simplified models","volume-title":"Forty-first International Conference on Machine Learning","author":"Friedman","year":"2024"},{"key":"2024122021045446000_bib30","article-title":"RNNs as psycholinguistic subjects: Syntactic state and grammatical dependency","author":"Futrell","year":"2018","journal-title":"arXiv preprint arXiv:1809.01329"},{"key":"2024122021045446000_bib31","doi-asserted-by":"publisher","first-page":"32","DOI":"10.18653\/v1\/N19-1004","article-title":"Neural language models as psycholinguistic subjects: Representations of syntactic state","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Futrell","year":"2019"},{"issue":"43","key":"2024122021045446000_bib32","doi-asserted-by":"publisher","first-page":"e2122602119","DOI":"10.1073\/pnas.2122602119","article-title":"A resource-rational model of human processing of recursive linguistic structure","volume":"119","author":"Hahn","year":"2022","journal-title":"Proceedings of the National Academy of Sciences"},{"issue":"2\u20133","key":"2024122021045446000_bib33","doi-asserted-by":"publisher","first-page":"61","DOI":"10.1017\/S0140525X0999152X","article-title":"The weirdest people in the world?","volume":"33","author":"Henrich","year":"2010","journal-title":"Behavioral and Brain Sciences"},{"key":"2024122021045446000_bib34","first-page":"30016","article-title":"An empirical analysis of compute-optimal large language model training","volume":"35","author":"Hoffmann","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2024122021045446000_bib35","doi-asserted-by":"publisher","first-page":"7038","DOI":"10.18653\/v1\/2021.emnlp-main.564","article-title":"Surface form competition: Why the highest probability answer isn\u2019t always right","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Holtzman","year":"2021"},{"issue":"1","key":"2024122021045446000_bib36","doi-asserted-by":"publisher","first-page":"43","DOI":"10.1162\/nol_a_00137","article-title":"Artificial neural network language models align neurally and behaviorally with humans even after a developmentally realistic amount of training","volume":"5","author":"Hosseini","year":"2022","journal-title":"Neurobiology of Language"},{"key":"2024122021045446000_bib37","article-title":"Auxiliary task demands mask the capabilities of smaller language models","author":"Hu","year":"2024","journal-title":"arXiv preprint arXiv:2404.02418"},{"key":"2024122021045446000_bib38","doi-asserted-by":"publisher","first-page":"1725","DOI":"10.18653\/v1\/2020.acl-main.158","article-title":"A systematic assessment of syntactic generalization in neural language models","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Hu","year":"2020"},{"key":"2024122021045446000_bib39","doi-asserted-by":"publisher","first-page":"5040","DOI":"10.18653\/v1\/2023.emnlp-main.306","article-title":"Prompting is not a substitute for probability measurements in large language models","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Hu","year":"2023"},{"key":"2024122021045446000_bib40","doi-asserted-by":"publisher","first-page":"757","DOI":"10.1613\/jair.1.11674","article-title":"Compositionality decomposed: How do neural networks generalise?","volume":"67","author":"Hupkes","year":"2020","journal-title":"Journal of Artificial Intelligence Research"},{"issue":"2","key":"2024122021045446000_bib41","doi-asserted-by":"publisher","first-page":"365","DOI":"10.1017\/S0022226707004616","article-title":"Constraints on multiple center-embedding of clauses","volume":"43","author":"Karlsson","year":"2007","journal-title":"Journal of Linguistics"},{"key":"2024122021045446000_bib42","first-page":"22199","article-title":"Large language models are zero-shot reasoners","volume":"35","author":"Kojima","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2024122021045446000_bib43","doi-asserted-by":"publisher","first-page":"66","DOI":"10.18653\/v1\/D18-2012","article-title":"SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Kudo","year":"2018"},{"key":"2024122021045446000_bib44","doi-asserted-by":"publisher","first-page":"10421","DOI":"10.18653\/v1\/2022.emnlp-main.712","article-title":"Context limitations make neural language models more human-like","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Kuribayashi","year":"2022"},{"key":"2024122021045446000_bib45","first-page":"3226","article-title":"Can transformers process recursive nested constructions, like humans?","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Lakretz","year":"2022"},{"issue":"104699","key":"2024122021045446000_bib46","doi-asserted-by":"publisher","DOI":"10.1016\/j.cognition.2021.104699","article-title":"Mechanisms for handling nested dependencies in neural-network language models and humans","volume":"213","author":"Lakretz","year":"2021","journal-title":"Cognition"},{"key":"2024122021045446000_bib47","doi-asserted-by":"publisher","first-page":"537","DOI":"10.18653\/v1\/2022.findings-emnlp.38","article-title":"Can language models learn from explanations in context?","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Lampinen","year":"2022"},{"key":"2024122021045446000_bib48","doi-asserted-by":"publisher","first-page":"5210","DOI":"10.18653\/v1\/2020.acl-main.465","article-title":"How can we accelerate progress towards human-like linguistic generalization?","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Linzen","year":"2020"},{"key":"2024122021045446000_bib49","article-title":"Are emergent abilities in large language models just in-context learning?","author":"Lu","year":"2023","journal-title":"arXiv preprint arXiv:2309.01809"},{"key":"2024122021045446000_bib50","doi-asserted-by":"publisher","first-page":"8086","DOI":"10.18653\/v1\/2022.acl-long.556","article-title":"Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Lu","year":"2022"},{"issue":"4","key":"2024122021045446000_bib51","doi-asserted-by":"publisher","first-page":"676","DOI":"10.1037\/0033-295X.101.4.676","article-title":"The lexical nature of syntactic ambiguity resolution","volume":"101","author":"MacDonald","year":"1994","journal-title":"Psychological Review"},{"issue":"48","key":"2024122021045446000_bib52","doi-asserted-by":"publisher","first-page":"30046","DOI":"10.1073\/pnas.1907367117","article-title":"Emergent linguistic structure in artificial neural networks trained by self-supervision","volume":"117","author":"Manning","year":"2020","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2024122021045446000_bib53","article-title":"GPT-3, Bloviator: OpenAI\u2019s language generator has no idea what it\u2019s talking about","author":"Marcus","year":"2020","journal-title":"Technology Review"},{"issue":"8","key":"2024122021045446000_bib54","doi-asserted-by":"publisher","first-page":"348","DOI":"10.1016\/j.tics.2010.06.002","article-title":"Letting structure emerge: Connectionist and dynamical systems approaches to cognition","volume":"14","author":"McClelland","year":"2010","journal-title":"Trends in Cognitive Sciences"},{"issue":"42","key":"2024122021045446000_bib55","doi-asserted-by":"publisher","first-page":"25966","DOI":"10.1073\/pnas.1910416117","article-title":"Placing language in an integrated understanding system: Next steps toward human-level performance in neural language models","volume":"117","author":"McClelland","year":"2020","journal-title":"Proceedings of the National Academy of Sciences"},{"issue":"1","key":"2024122021045446000_bib56","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1080\/07268608508599333","article-title":"What a performance! Some problems with the competence-performance distinction","volume":"5","author":"Milroy","year":"1985","journal-title":"Australian Journal of Linguistics"},{"key":"2024122021045446000_bib57","article-title":"Experimental contexts can facilitate robust semantic property inference in language models, but inconsistently","author":"Misra","year":"2024","journal-title":"arXiv preprint arXiv:2401.06640"},{"key":"2024122021045446000_bib58","doi-asserted-by":"publisher","first-page":"2928","DOI":"10.18653\/v1\/2023.eacl-main.213","article-title":"COMPS: Conceptual minimal pair sentences for testing robust property knowledge and its inheritance in pre-trained language models","volume-title":"Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics","author":"Misra","year":"2023"},{"key":"2024122021045446000_bib59","doi-asserted-by":"publisher","first-page":"439","DOI":"10.18653\/v1\/2023.acl-short.38","article-title":"Grokking of hierarchical structure in vanilla transformers","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Murty","year":"2023"},{"key":"2024122021045446000_bib60","article-title":"Show your work: Scratchpads for intermediate computation with language models","author":"Nye","year":"2021","journal-title":"arXiv preprint arXiv:2112.00114"},{"key":"2024122021045446000_bib61","first-page":"27730","article-title":"Training language models to follow instructions with human feedback","volume":"35","author":"Ouyang","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2024122021045446000_bib62","first-page":"11054","article-title":"True few-shot learning with language models","volume-title":"Advances in Neural Information Processing Systems","author":"Perez","year":"2021"},{"issue":"4","key":"2024122021045446000_bib63","doi-asserted-by":"publisher","first-page":"136","DOI":"10.1016\/S1364-6613(99)01293-0","article-title":"Syntactic priming in language production","volume":"3","author":"Pickering","year":"1999","journal-title":"Trends in Cognitive Sciences"},{"issue":"3","key":"2024122021045446000_bib64","doi-asserted-by":"publisher","first-page":"427","DOI":"10.1037\/0033-2909.134.3.427","article-title":"Structural priming: A critical review","volume":"134","author":"Pickering","year":"2008","journal-title":"Psychological Bulletin"},{"key":"2024122021045446000_bib65","doi-asserted-by":"publisher","first-page":"66","DOI":"10.18653\/v1\/K19-1007","article-title":"Using priming to uncover the organization of syntactic representations in neural language models","volume-title":"Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL)","author":"Prasad","year":"2019"},{"issue":"8","key":"2024122021045446000_bib66","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford","year":"2019","journal-title":"OpenAI blog"},{"key":"2024122021045446000_bib67","article-title":"Scaling language models: Methods, analysis & insights from training gopher","author":"Rae","year":"2021","journal-title":"arXiv preprint arXiv:2112.11446"},{"issue":"140","key":"2024122021045446000_bib68","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel","year":"2020","journal-title":"Journal of Machine Learning Research"},{"key":"2024122021045446000_bib69","doi-asserted-by":"publisher","first-page":"10776","DOI":"10.18653\/v1\/2023.findings-emnlp.722","article-title":"NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Sainz","year":"2023"},{"key":"2024122021045446000_bib70","first-page":"43197","article-title":"Stay on topic with classifier-free guidance","volume-title":"Forty-first International Conference on Machine Learning","author":"Sanchez","year":"2024"},{"key":"2024122021045446000_bib71","first-page":"55565","article-title":"Are emergent abilities of large language models a mirage?","volume":"36","author":"Schaeffer","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"45","key":"2024122021045446000_bib72","doi-asserted-by":"publisher","first-page":"e2105646118","DOI":"10.1073\/pnas.2105646118","article-title":"The neural architecture of language: Integrative modeling converges on predictive processing","volume":"118","author":"Schrimpf","year":"2021","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2024122021045446000_bib73","doi-asserted-by":"publisher","first-page":"969","DOI":"10.18653\/v1\/2022.naacl-main.71","article-title":"When a sentence does not introduce a discourse entity, transformer-based models still sometimes refer to it","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Schuster","year":"2022"},{"key":"2024122021045446000_bib74","article-title":"Prompting GPT-3 to be reliable","volume-title":"The Eleventh International Conference on Learning Representations","author":"Si","year":"2023"},{"key":"2024122021045446000_bib75","doi-asserted-by":"publisher","first-page":"1031","DOI":"10.1162\/tacl_a_00504","article-title":"Structural persistence in language models: Priming as a window into abstract language representations","volume":"10","author":"Sinclair","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024122021045446000_bib76","doi-asserted-by":"publisher","first-page":"1","DOI":"10.46867\/ijcp.2018.31.01.12","article-title":"The importance of a truly comparative methodology for comparative psychology","volume":"31","author":"Smith","year":"2018","journal-title":"International Journal of Comparative Psychology"},{"key":"2024122021045446000_bib77","doi-asserted-by":"crossref","first-page":"1631","DOI":"10.18653\/v1\/D13-1170","article-title":"Recursive deep models for semantic compositionality over a sentiment treebank","volume-title":"Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing","author":"Socher","year":"2013"},{"key":"2024122021045446000_bib78","article-title":"Beyond the imitation game: Quantifying and extrapolating the capabilities of language models","author":"Srivastava","year":"2023","journal-title":"Transactions on Machine Learning Research"},{"key":"2024122021045446000_bib79","doi-asserted-by":"publisher","first-page":"13003","DOI":"10.18653\/v1\/2023.findings-acl.824","article-title":"Challenging BIG-bench tasks and whether chain-of-thought can solve them","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Suzgun","year":"2023"},{"key":"2024122021045446000_bib80","doi-asserted-by":"publisher","first-page":"1471","DOI":"10.18653\/v1\/2023.emnlp-main.91","article-title":"Transcending scaling laws with 0.1% extra compute","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Tay","year":"2023"},{"key":"2024122021045446000_bib81","doi-asserted-by":"publisher","DOI":"10.1037\/11294-000","author":"Underwood","year":"1949","journal-title":"Experimental Psychology: An Introduction"},{"key":"2024122021045446000_bib82","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2024122021045446000_bib83","first-page":"22964","article-title":"What language model architecture and pretraining objective works best for zero-shot generalization?","volume-title":"International Conference on Machine Learning","author":"Wang","year":"2022"},{"key":"2024122021045446000_bib84","volume-title":"Psychology of Reasoning: Structure and Content","author":"Wason","year":"1972"},{"key":"2024122021045446000_bib85","doi-asserted-by":"publisher","first-page":"7662","DOI":"10.18653\/v1\/2023.findings-emnlp.514","article-title":"Are language models worse than humans at following prompts? It\u2019s complicated","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Webson","year":"2023"},{"key":"2024122021045446000_bib86","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.167","article-title":"Do prompt-based models really understand the meaning of their prompts?","author":"Webson","year":"2021","journal-title":"arXiv preprint arXiv:2109.01247"},{"key":"2024122021045446000_bib87","first-page":"1","article-title":"Emergent abilities of large language models","author":"Wei","year":"2022","journal-title":"Transactions on Machine Learning Research"},{"key":"2024122021045446000_bib88","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2024122021045446000_bib89","doi-asserted-by":"publisher","first-page":"181","DOI":"10.18653\/v1\/W19-4819","article-title":"Hierarchical representation in neural language models: Suppression and recovery of expectations","volume-title":"Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Wilcox","year":"2019"},{"key":"2024122021045446000_bib90","article-title":"On the predictive power of neural language models for human real-time comprehension behavior","author":"Wilcox","year":"2020","journal-title":"arXiv preprint arXiv:2006.01912"},{"key":"2024122021045446000_bib91","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1017\/S0140525X21001758","article-title":"The generalizability crisis","volume":"45","author":"Yarkoni","year":"2022","journal-title":"Behavioral and Brain Sciences"},{"key":"2024122021045446000_bib92","article-title":"Least-to-most prompting enables complex reasoning in large language models","volume-title":"The Eleventh International Conference on Learning Representations","author":"Zhou","year":"2023"}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/50\/4\/1441\/2470042\/coli_a_00525.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/50\/4\/1441\/2470042\/coli_a_00525.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,12,20]],"date-time":"2024-12-20T21:05:19Z","timestamp":1734728719000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/50\/4\/1441\/123789\/Can-Language-Models-Handle-Recursively-Nested"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":92,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2024,12,1]]},"published-print":{"date-parts":[[2024,12,1]]}},"URL":"https:\/\/doi.org\/10.1162\/coli_a_00525","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"value":"0891-2017","type":"print"},{"value":"1530-9312","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}