{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,12,20]],"date-time":"2024-12-20T21:40:02Z","timestamp":1734730802714,"version":"3.32.0"},"reference-count":121,"publisher":"MIT Press","issue":"4","license":[{"start":{"date-parts":[[2024,7,30]],"date-time":"2024-07-30T00:00:00Z","timestamp":1722297600000},"content-version":"vor","delay-in-days":211,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,12,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>The staggering pace with which the capabilities of large language models (LLMs) are increasing, as measured by a range of commonly used natural language understanding (NLU) benchmarks, raises many questions regarding what \u201cunderstanding\u201d means for a language model and how it compares to human understanding. This is especially true since many LLMs are exclusively trained on text, casting doubt on whether their stellar benchmark performances are reflective of a true understanding of the problems represented by these benchmarks, or whether LLMs simply excel at uttering textual forms that correlate with what someone who understands the problem would say. In this philosophically inspired work, we aim to create some separation between form and meaning, with a series of tests that leverage the idea that world understanding should be consistent across presentational modes\u2014inspired by Fregean senses\u2014of the same meaning. Specifically, we focus on consistency across languages as well as paraphrases. Taking GPT-3.5 as our object of study, we evaluate multisense consistency across five different languages and various tasks. We start the evaluation in a controlled setting, asking the model for simple facts, and then proceed with an evaluation on four popular NLU benchmarks. We find that the model\u2019s multisense consistency is lacking and run several follow-up analyses to verify that this lack of consistency is due to a sense-dependent task understanding. We conclude that, in this aspect, the understanding of LLMs is still quite far from being consistent and human-like, and deliberate on how this impacts their utility in the context of learning about human language and understanding.<\/jats:p>","DOI":"10.1162\/coli_a_00529","type":"journal-article","created":{"date-parts":[[2024,7,30]],"date-time":"2024-07-30T15:07:44Z","timestamp":1722352064000},"page":"1507-1556","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":0,"title":["From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency"],"prefix":"10.1162","volume":"50","author":[{"given":"Xenia","family":"Ohmer","sequence":"first","affiliation":[{"name":"Osnabrueck University, Institute for Cognitive Science. xenia.ohmer@uni-osnabrueck.de"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elia","family":"Bruni","sequence":"additional","affiliation":[{"name":"Osnabrueck University, Institute for Cognitive Science. elia.bruni@uni-osnabrueck.de"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dieuwke","family":"Hupke","sequence":"additional","affiliation":[{"name":"Meta. dieuwkehupkes@meta.com"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2024,12,1]]},"reference":[{"key":"2024122021045224100_bib1","doi-asserted-by":"publisher","first-page":"109","DOI":"10.18653\/v1\/2021.conll-1.9","article-title":"Can language models encode perceptual structure without grounding? A case study in color","volume-title":"Proceedings of the 25th Conference on Computational Natural Language Learning","author":"AbdouStella","year":"2021"},{"key":"2024122021045224100_bib2","doi-asserted-by":"publisher","first-page":"191","DOI":"10.18653\/v1\/W19-4820","article-title":"Blackbox meets blackbox: Representational similarity & stability analysis of neural language models and brains","volume-title":"Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Abnar","year":"2019"},{"key":"2024122021045224100_bib3","doi-asserted-by":"publisher","first-page":"6168","DOI":"10.18653\/v1\/P19-1620","article-title":"Synthetic QA corpora generation with roundtrip consistency","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Alberti","year":"2019"},{"key":"2024122021045224100_bib4","doi-asserted-by":"publisher","first-page":"301","DOI":"10.18653\/v1\/2022.conll-1.20","article-title":"Syntactic surprisal from neural models predicts, but underestimates, human processing difficulty from syntactic ambiguities","volume-title":"Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL)","author":"Arehalli","year":"2022"},{"key":"2024122021045224100_bib5","doi-asserted-by":"publisher","first-page":"5642","DOI":"10.18653\/v1\/2020.acl-main.499","article-title":"Logic-guided data augmentation and regularization for consistent question answering","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Asai","year":"2020"},{"issue":"5","key":"2024122021045224100_bib6","doi-asserted-by":"publisher","first-page":"1474","DOI":"10.1111\/j.1467-8624.1990.tb02876.x","article-title":"The principle of mutual exclusivity in word learning: To honor or not to honor?","volume":"61","author":"Au","year":"1990","journal-title":"Child Development"},{"issue":"2","key":"2024122021045224100_bib7","doi-asserted-by":"publisher","first-page":"170","DOI":"10.1016\/j.tics.2017.11.005","article-title":"Frontal cortex and the hierarchical control of behavior","volume":"22","author":"Badre","year":"2018","journal-title":"Trends in Cognitive Sciences"},{"key":"2024122021045224100_bib8","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.44","article-title":"The belebele benchmark: A parallel reading comprehension dataset in 122 language variants","volume":"arXiv:2308.16884","author":"Bandarkar","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib9","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1201\/9781003205388-1","article-title":"On the proper role of linguistically oriented deep net analysis in linguistic theorising","volume-title":"Algebraic Structures in Natural Language","author":"Baroni","year":"2023"},{"key":"2024122021045224100_bib10","first-page":"389","article-title":"Abstraction as dynamic interpretation in perceptual symbol systems","volume-title":"Building Object Categories in Developmental Time","author":"Barsalou","year":"2005"},{"key":"2024122021045224100_bib11","article-title":"Worldsense: A synthetic benchmark for grounded reasoning in large language models","volume":"arXiv:2311.15930","author":"Benchekroun","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib12","doi-asserted-by":"publisher","first-page":"5185","DOI":"10.18653\/v1\/2020.acl-main.463","article-title":"Climbing towards NLU: On meaning, form, and understanding in the age of data","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Bender","year":"2020"},{"key":"2024122021045224100_bib13","article-title":"The reversal curse: LLMs trained on \u201cA is B\u201d fail to learn \u201cB is A\u201d","volume":"arXiv:2309.12288","author":"Berglund","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib14","doi-asserted-by":"publisher","first-page":"686","DOI":"10.1038\/d41586-023-02361-7","article-title":"ChatGPT broke the turing test\u2014the race is on for new ways to assess AI","volume":"619","author":"Biever","year":"2023","journal-title":"Nature (News Feature)"},{"key":"2024122021045224100_bib15","doi-asserted-by":"publisher","first-page":"8718","DOI":"10.18653\/v1\/2020.emnlp-main.703","article-title":"Experience grounds language","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Bisk","year":"2020"},{"key":"2024122021045224100_bib16","doi-asserted-by":"publisher","first-page":"5698","DOI":"10.18653\/v1\/2023.acl-long.313","article-title":"Zero-shot approach to overcome perturbation sensitivity of prompts","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Chakraborty","year":"2023"},{"key":"2024122021045224100_bib17","article-title":"Language model behavior: A comprehensive survey","volume":"arXiv:2303.11504","author":"Chang","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib18","doi-asserted-by":"publisher","first-page":"3841","DOI":"10.18653\/v1\/2021.findings-emnlp.324","article-title":"Can NLI models verify QA systems\u2019 predictions?","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2021","author":"Chen","year":"2021"},{"issue":"2","key":"2024122021045224100_bib19","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1207\/s15516709cog2302_2","article-title":"Toward a connectionist model of recursion in human linguistic performance","volume":"23","author":"Christiansen","year":"1999","journal-title":"Cognitive Science"},{"key":"2024122021045224100_bib20","doi-asserted-by":"publisher","first-page":"2475","DOI":"10.18653\/v1\/D18-1269","article-title":"XNLI: Evaluating cross-lingual sentence representations","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Conneau","year":"2018"},{"issue":"3","key":"2024122021045224100_bib21","doi-asserted-by":"publisher","first-page":"e13256","DOI":"10.1111\/cogs.13256","article-title":"Large language models demonstrate the potential of statistical learning in language","volume":"47","author":"Contreras Kallens","year":"2023","journal-title":"Cognitive Science"},{"key":"2024122021045224100_bib22","doi-asserted-by":"publisher","first-page":"8493","DOI":"10.18653\/v1\/2022.acl-long.581","article-title":"Knowledge neurons in pretrained transformers","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Dai","year":"2022"},{"key":"2024122021045224100_bib23","doi-asserted-by":"publisher","first-page":"94","DOI":"10.18653\/v1\/2021.conll-1.8","article-title":"Generalising to German plural noun classes, from the perspective of a recurrent neural network","volume-title":"Proceedings of the 25th Conference on Computational Natural Language Learning","author":"Dankers","year":"2021"},{"key":"2024122021045224100_bib24","doi-asserted-by":"publisher","DOI":"10.1145\/3596490","article-title":"Shortcut learning of large language models in natural language understanding","volume":"arXiv:2208.11857","author":"Du","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib25","doi-asserted-by":"publisher","first-page":"43","DOI":"10.1016\/j.cognition.2017.11.008","article-title":"Cognitive science in the era of artificial intelligence: A roadmap for reverse-engineering the infant language-learner","volume":"173","author":"Dupoux","year":"2018","journal-title":"Cognition"},{"key":"2024122021045224100_bib26","doi-asserted-by":"publisher","first-page":"1012","DOI":"10.1162\/tacl_a_00410","article-title":"Measuring and improving consistency in pretrained language models","volume":"9","author":"Elazar","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"issue":"2","key":"2024122021045224100_bib27","doi-asserted-by":"publisher","first-page":"179","DOI":"10.1016\/0364-0213(90)90002-E","article-title":"Finding structure in time","volume":"14","author":"Elman","year":"1990","journal-title":"Cognitive Science"},{"key":"2024122021045224100_bib28","doi-asserted-by":"publisher","first-page":"251","DOI":"10.1093\/oso\/9780195151770.003.0014","article-title":"Bilingual semantic and conceptual representation","author":"Francis","year":"2009"},{"issue":"6","key":"2024122021045224100_bib29","doi-asserted-by":"publisher","first-page":"829","DOI":"10.1177\/0956797611409589","article-title":"Insensitivity of the human sentence-processing system to hierarchical structure","volume":"22","author":"Frank","year":"2011","journal-title":"Psychological Science"},{"issue":"1","key":"2024122021045224100_bib30","first-page":"25","article-title":"\u00dcber Sinn und Bedeutung [\u201cOn sense and reference\u201d]","volume":"100","author":"Frege","year":"1892","journal-title":"Zeitschrift f\u00fcr Philosophie und philosophische Kritik"},{"key":"2024122021045224100_bib31","first-page":"58","article-title":"The thought: A logical inquiry [\u201cDer Gedanke. Eine logische Untersuchung\u201d]","volume":"2","author":"Frege","year":"1918\u20131919","journal-title":"Beitr\u00e4ge Zur Philosophie des Deutschen Idealismus"},{"key":"2024122021045224100_bib32","doi-asserted-by":"publisher","first-page":"688","DOI":"10.18653\/v1\/E17-1065","article-title":"Noisy-context surprisal as a human sentence processing cost model","volume-title":"Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers","author":"Futrell","year":"2017"},{"issue":"3","key":"2024122021045224100_bib33","doi-asserted-by":"publisher","first-page":"672","DOI":"10.1111\/tops.12278","article-title":"Analogy and abstraction","volume":"9","author":"Gentner","year":"2017","journal-title":"Topics in Cognitive Science"},{"key":"2024122021045224100_bib34","doi-asserted-by":"publisher","first-page":"1161","DOI":"10.18653\/v1\/D19-1107","article-title":"Are we modeling the task or the annotator? An investigation of annotator bias in natural language understanding datasets","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Geva","year":"2019"},{"key":"2024122021045224100_bib35","doi-asserted-by":"publisher","first-page":"5484","DOI":"10.18653\/v1\/2021.emnlp-main.446","article-title":"Transformer feed-forward layers are key-value memories","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Geva","year":"2021"},{"key":"2024122021045224100_bib36","doi-asserted-by":"publisher","first-page":"240","DOI":"10.18653\/v1\/W18-5426","article-title":"Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information","volume-title":"Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Giulianelli","year":"2018"},{"key":"2024122021045224100_bib37","doi-asserted-by":"publisher","first-page":"1195","DOI":"10.18653\/v1\/N18-1108","article-title":"Colorless green recurrent networks dream hierarchically","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers","author":"Gulordava","year":"2018"},{"key":"2024122021045224100_bib38","doi-asserted-by":"publisher","first-page":"107","DOI":"10.18653\/v1\/N18-2017","article-title":"Annotation artifacts in natural language inference data","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers)","author":"Gururangan","year":"2018"},{"key":"2024122021045224100_bib39","doi-asserted-by":"publisher","first-page":"5457","DOI":"10.18653\/v1\/2023.emnlp-main.332","article-title":"The effect of scaling, retrieval augmentation and form on the factual consistency of language models","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Hagstr\u00f6m","year":"2023"},{"key":"2024122021045224100_bib40","first-page":"1","volume-title":"Rethinking reasoning evaluation with theories of intelligence","author":"Heineman","year":"2023"},{"key":"2024122021045224100_bib41","first-page":"1","article-title":"Measuring massive multitask language understanding","author":"Hendrycks","year":"2021","journal-title":"International Conference on Learning Representations (ICLR)"},{"issue":"5","key":"2024122021045224100_bib42","doi-asserted-by":"publisher","first-page":"220","DOI":"10.1016\/j.tics.2005.03.003","article-title":"The emergence of competing modules in bilingualism","volume":"9","author":"Hernandez","year":"2005","journal-title":"Trends in Cognitive Sciences"},{"key":"2024122021045224100_bib43","doi-asserted-by":"publisher","first-page":"1301","DOI":"10.18653\/v1\/2021.naacl-main.102","article-title":"Understanding by understanding not: Modeling negation in language models","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Hosseini","year":"2021"},{"key":"2024122021045224100_bib44","first-page":"4411","article-title":"XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Hu","year":"2020"},{"key":"2024122021045224100_bib45","doi-asserted-by":"publisher","first-page":"z38u6","DOI":"10.31234\/osf.io\/z38u6","article-title":"Surprisal does not explain syntactic disambiguation difficulty: Evidence from a large-scale benchmark","author":"Huang","year":"2023","journal-title":"PsyArXiv Preprint"},{"key":"2024122021045224100_bib46","unstructured":"Hupkes, Dieuwke\n          . 2020. Hierarchy and Interpretability in Neural Models of Language Processing . Ph.D. thesis, University of Amsterdam."},{"issue":"10","key":"2024122021045224100_bib47","doi-asserted-by":"publisher","first-page":"1161","DOI":"10.1038\/s42256-023-00729-y","article-title":"A taxonomy and review of generalization research in NLP","volume":"5","author":"Hupkes","year":"2023","journal-title":"Nature Machine Intelligence"},{"issue":"251","key":"2024122021045224100_bib48","first-page":"1","article-title":"Atlas: Few-shot learning with retrieval augmented language models","volume":"24","author":"Izacard","year":"2023","journal-title":"Journal of Machine Learning Research"},{"key":"2024122021045224100_bib49","first-page":"3680","article-title":"BECEL: Benchmark for consistency evaluation of language models","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Jang","year":"2022"},{"key":"2024122021045224100_bib50","doi-asserted-by":"publisher","first-page":"15970","DOI":"10.18653\/v1\/2023.emnlp-main.991","article-title":"Consistency analysis of ChatGPT","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Jang","year":"2023"},{"key":"2024122021045224100_bib51","doi-asserted-by":"publisher","first-page":"1","DOI":"10.34133\/icomputing.0064","article-title":"What should replace the Turing Test?","volume":"2","author":"Johnson-Laird","year":"2023","journal-title":"Intelligent Computing"},{"key":"2024122021045224100_bib52","doi-asserted-by":"publisher","first-page":"4958","DOI":"10.18653\/v1\/2021.findings-acl.439","article-title":"Language models use monotonicity to assess NPI licensing","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Jumelet","year":"2021"},{"key":"2024122021045224100_bib53","doi-asserted-by":"publisher","first-page":"7811","DOI":"10.18653\/v1\/2020.acl-main.698","article-title":"Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Kassner","year":"2020"},{"key":"2024122021045224100_bib54","doi-asserted-by":"publisher","first-page":"8849","DOI":"10.18653\/v1\/2021.emnlp-main.697","article-title":"BeliefBank: Adding memory to a pre-trained language model for a systematic notion of belief","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Kassner","year":"2021"},{"key":"2024122021045224100_bib55","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2023.103971","article-title":"The defeat of the Winograd Schema Challenge","volume":"arXiv:2201.02387","author":"Kocijan","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib56","first-page":"169","article-title":"Lexical and conceptual memory in the bilingual: Mapping form to meaning in two languages","volume-title":"Tutorials in Bilingualism","author":"Kroll","year":"1997"},{"key":"2024122021045224100_bib57","first-page":"3226","article-title":"Can transformers process recursive nested constructions, like humans?","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Lakretz","year":"2022"},{"key":"2024122021045224100_bib58","doi-asserted-by":"publisher","first-page":"104699","DOI":"10.1016\/j.cognition.2021.104699","article-title":"Mechanisms for handling nested dependencies in neural-network language models and humans","volume":"213","author":"Lakretz","year":"2021","journal-title":"Cognition"},{"key":"2024122021045224100_bib59","doi-asserted-by":"publisher","first-page":"11","DOI":"10.18653\/v1\/N19-1002","article-title":"The emergence of number and syntax units in LSTM language models","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Lakretz","year":"2019"},{"key":"2024122021045224100_bib60","doi-asserted-by":"publisher","first-page":"3924","DOI":"10.18653\/v1\/D19-1405","article-title":"A logic-driven framework for consistency of neural models","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Li","year":"2019"},{"key":"2024122021045224100_bib61","first-page":"1","article-title":"Holistic evaluation of language models","author":"Liang","year":"2023","journal-title":"Transactions on Machine Learning Research"},{"key":"2024122021045224100_bib62","doi-asserted-by":"publisher","first-page":"6008","DOI":"10.18653\/v1\/2020.emnlp-main.484","article-title":"XGLUE: A new benchmark dataset for cross-lingual pre-training, understanding and generation","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Liang","year":"2020"},{"key":"2024122021045224100_bib63","first-page":"74","article-title":"ROUGE: A package for automatic evaluation of summaries","volume-title":"Text Summarization Branches Out","author":"Lin","year":"2004"},{"key":"2024122021045224100_bib64","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1146\/annurev-linguistics-032020-051035","article-title":"Syntactic structure from deep learning","volume":"7","author":"Linzen","year":"2021","journal-title":"Annual Review of Linguistics"},{"key":"2024122021045224100_bib65","doi-asserted-by":"publisher","first-page":"521","DOI":"10.1162\/tacl_a_00115","article-title":"Assessing the ability of LSTMs to learn syntax-sensitive dependencies","volume":"4","author":"Linzen","year":"2016","journal-title":"Transactions of the Association for Computational Linguistics"},{"issue":"3","key":"2024122021045224100_bib66","doi-asserted-by":"publisher","first-page":"640","DOI":"10.1016\/j.cell.2019.06.012","article-title":"Human replay spontaneously reorganizes experience","volume":"178","author":"Liu","year":"2019","journal-title":"Cell"},{"key":"2024122021045224100_bib67","doi-asserted-by":"publisher","DOI":"10.1016\/j.tics.2024.01.011","article-title":"Dissociating language and thought in large language models","volume":"arXiv:2301.06627","author":"Mahowald","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib68","doi-asserted-by":"publisher","first-page":"431","DOI":"10.1007\/s11525-017-9307-x","article-title":"Abstractive morphological learning with a recurrent neural network","volume":"27","author":"Malouf","year":"2017","journal-title":"Morphology"},{"key":"2024122021045224100_bib69","article-title":"Do language models refer?","volume":"arXiv:2308.05576","author":"Mandelkern","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib70","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.2322420121","volume-title":"Thomas Embers of autoregression: Understanding large language models through the problem they are trained to solve","author":"McCoy","year":"2023"},{"key":"2024122021045224100_bib71","doi-asserted-by":"publisher","first-page":"3428","DOI":"10.18653\/v1\/P19-1334","article-title":"Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"McCoy","year":"2019"},{"key":"2024122021045224100_bib72","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.182","volume-title":"Sources of hallucination by large language models on inference tasks","author":"McKenna","year":"2023"},{"issue":"1","key":"2024122021045224100_bib73","doi-asserted-by":"publisher","first-page":"202","DOI":"10.1016\/j.neuron.2014.05.019","article-title":"Hippocampal representation of related and opposing memories develop within distinct, hierarchically organized neural schemas","volume":"83","author":"McKenzie","year":"2014","journal-title":"Neuron"},{"key":"2024122021045224100_bib74","doi-asserted-by":"publisher","first-page":"65","DOI":"10.18653\/v1\/K18-1007","article-title":"Adversarially regularising neural NLI models to integrate logical background knowledge","volume-title":"Proceedings of the 22nd Conference on Computational Natural Language Learning","author":"Minervini","year":"2018"},{"key":"2024122021045224100_bib75","doi-asserted-by":"publisher","first-page":"1754","DOI":"10.18653\/v1\/2022.emnlp-main.115","article-title":"Enhancing self-consistency and performance of pre-trained language models through natural language inference","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Mitchell","year":"2022"},{"issue":"13","key":"2024122021045224100_bib76","doi-asserted-by":"publisher","first-page":"e2215907120","DOI":"10.1073\/pnas.2215907120","article-title":"The debate over understanding in AI\u2019s large language models","volume":"120","author":"Mitchell","year":"2023","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2024122021045224100_bib77","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00681","article-title":"State of what art? A call for multi-prompt LLM evaluation","volume":"arXiv:2401.00595","author":"Mizrahi","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib78","article-title":"The vector grounding problem","volume":"arXiv:2304:01481","author":"Mollo","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib79","doi-asserted-by":"publisher","first-page":"4885","DOI":"10.18653\/v1\/2020.acl-main.441","article-title":"Adversarial NLI: A new benchmark for natural language understanding","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Nie","year":"2020"},{"key":"2024122021045224100_bib80","doi-asserted-by":"publisher","first-page":"4658","DOI":"10.18653\/v1\/P19-1459","volume-title":"Probing neural network comprehension of natural language arguments","author":"Niven","year":"2019"},{"key":"2024122021045224100_bib81","first-page":"258","volume-title":"Separating form and meaning: Using self-consistency to quantify task understanding across multiple senses","author":"Ohmer","year":"2023"},{"key":"2024122021045224100_bib82","first-page":"27730","article-title":"Training language models to follow instructions with human feedback","volume-title":"Advances in Neural Information Processing Systems","author":"OuyangJeffrey","year":"2022"},{"key":"2024122021045224100_bib83","doi-asserted-by":"publisher","first-page":"311","DOI":"10.3115\/1073083.1073135","article-title":"BLEU: A method for automatic evaluation of machine translation","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni","year":"2002"},{"key":"2024122021045224100_bib84","first-page":"1","article-title":"Mapping language models to grounded conceptual spaces","volume-title":"International Conference on Learning Representations","author":"Patel","year":"2022"},{"issue":"2251","key":"2024122021045224100_bib85","doi-asserted-by":"publisher","first-page":"20220041","DOI":"10.1098\/rsta.2022.0041","article-title":"Symbols and grounding in large language models","volume":"381","author":"Pavlick","year":"2023","journal-title":"Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences"},{"key":"2024122021045224100_bib86","first-page":"7180","article-title":"Modern language models refute Chomsky\u2019s approach to language","author":"Piantadosi","year":"2023","journal-title":"Lingbuzz preprint"},{"key":"2024122021045224100_bib87","first-page":"1","article-title":"Meaning without reference in large language models","volume-title":"NeurIPS 2022 Workshop on Neuro Causal and Symbolic AI (nCSI)","author":"Piantadosi","year":"2022"},{"key":"2024122021045224100_bib88","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/IJCNN52387.2021.9534299","article-title":"How can the [mask] know? The sources and limitations of knowledge in BERT","volume-title":"2021 International Joint Conference on Neural Networks (IJCNN)","author":"Podkorytov","year":"2021"},{"key":"2024122021045224100_bib89","doi-asserted-by":"publisher","first-page":"2362","DOI":"10.18653\/v1\/2020.emnlp-main.185","article-title":"XCOPA: A multilingual dataset for causal commonsense reasoning","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Ponti","year":"2020"},{"key":"2024122021045224100_bib90","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.658","article-title":"Cross-lingual consistency of factual knowledge in multilingual language models","volume":"arXiv:2310:10478","author":"Qi","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib91","first-page":"1","article-title":"AI and the everything in the whole wide world benchmark","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1","author":"Raji","year":"2021"},{"key":"2024122021045224100_bib92","first-page":"78","article-title":"Machine reading, fast and slow: When do models \u201cunderstand\u201d language?","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Ray Choudhury","year":"2022"},{"key":"2024122021045224100_bib93","first-page":"578","volume-title":"COMET-22: Unbabel-IST 2022 submission for the metrics shared task","author":"Rei","year":"2022"},{"key":"2024122021045224100_bib94","first-page":"90","article-title":"Choice of plausible alternatives: An evaluation of commonsense causal reasoning","author":"Roemmele","year":"2011","journal-title":"Proceedings of the 2011 AAAI Spring Symposium Series"},{"key":"2024122021045224100_bib95","doi-asserted-by":"publisher","first-page":"10215","DOI":"10.18653\/v1\/2021.emnlp-main.802","volume-title":"XTREME-R: Towards more challenging and nuanced multilingual evaluation","author":"Ruder","year":"2021"},{"key":"2024122021045224100_bib96","doi-asserted-by":"publisher","first-page":"61","DOI":"10.18653\/v1\/2021.cmcl-1.6","volume-title":"Accounting for agreement phenomena in sentence comprehension with transformer language models: Effects of similarity-based interference on surprisal and attention","author":"Ryu","year":"2021"},{"key":"2024122021045224100_bib97","doi-asserted-by":"publisher","first-page":"3061","DOI":"10.18653\/v1\/2023.eacl-main.223","volume-title":"PECO: Examining single sentence label leakage in natural language inference datasets through progressive evaluation of cluster outliers","author":"Saxon","year":"2023"},{"key":"2024122021045224100_bib98","doi-asserted-by":"publisher","first-page":"2429","DOI":"10.18653\/v1\/2020.emnlp-main.190","volume-title":"What do models learn from question answering datasets?","author":"Sen","year":"2020"},{"key":"2024122021045224100_bib99","doi-asserted-by":"publisher","first-page":"274","DOI":"10.18653\/v1\/2023.conll-1.19","volume-title":"The validity of evaluation results: Assessing concurrence across compositionality benchmarks","author":"Sun","year":"2023"},{"key":"2024122021045224100_bib100","article-title":"General-purpose question-answering with Macaw","volume":"arXiv:2109.02593","author":"Tafjord","year":"2021","journal-title":"ArXiv preprint"},{"issue":"6022","key":"2024122021045224100_bib101","doi-asserted-by":"publisher","first-page":"1279","DOI":"10.1126\/science.1192788","article-title":"How to grow a mind: Statistics, structure, and abstraction","volume":"331","author":"Tenenbaum","year":"2011","journal-title":"Science"},{"key":"2024122021045224100_bib102","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.582","article-title":"A language model with limited memory capacity captures interference in human sentence processing","volume":"arXiv:2310:16142","author":"Timkey","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib103","article-title":"Llama: Open and efficient foundation language models","volume":"arXiv:2302:13971","author":"Touvron","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib104","doi-asserted-by":"publisher","first-page":"209","DOI":"10.18653\/v1\/W19-4324","article-title":"Assessing incrementality in sequence-to-sequence models","volume-title":"Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019)","author":"Ulmer","year":"2019"},{"key":"2024122021045224100_bib105","doi-asserted-by":"publisher","first-page":"e63226","DOI":"10.7554\/eLife.63226","article-title":"Neural representation of abstract task structure during generalization","volume":"10","author":"Vaidya","year":"2021","journal-title":"eLife"},{"key":"2024122021045224100_bib106","first-page":"2603","article-title":"Modeling garden path effects without explicit hierarchical syntax","author":"Van Schijndel","year":"2018","journal-title":"Proceedings of the 40th Annual Meeting of the Cognitive Science Society (CogSci)"},{"issue":"6","key":"2024122021045224100_bib107","doi-asserted-by":"publisher","first-page":"e12988","DOI":"10.1111\/cogs.12988","article-title":"Single-stage prediction models do not explain the magnitude of syntactic disambiguation difficulty","volume":"45","author":"Van Schijndel","year":"2021","journal-title":"Cognitive Science"},{"key":"2024122021045224100_bib108","first-page":"1","article-title":"SuperGLUE: A stickier benchmark for general-purpose language understanding systems","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems (NeurIPS)","author":"Wang","year":"2019"},{"key":"2024122021045224100_bib109","doi-asserted-by":"publisher","first-page":"353","DOI":"10.18653\/v1\/W18-5446","article-title":"GLUE: A multi-task benchmark and analysis platform for natural language understanding","volume-title":"Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Wang","year":"2018"},{"key":"2024122021045224100_bib110","first-page":"1","article-title":"Adversarial GLUE: A multi-task benchmark for robustness evaluation of language models","volume-title":"Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)","author":"Wang","year":"2021"},{"key":"2024122021045224100_bib111","doi-asserted-by":"publisher","first-page":"7136","DOI":"10.1609\/aaai.v33i01.33017136","article-title":"What if we simply swap the two text fragments? A straightforward yet effective way to test the robustness of methods to confounding signals in natural language inference tasks","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Wang","year":"2019"},{"key":"2024122021045224100_bib112","article-title":"Are large language models really robust to word-level perturbations?","volume":"arXiv:2309.11166","author":"Wang","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib113","doi-asserted-by":"publisher","first-page":"11132","DOI":"10.18653\/v1\/2022.emnlp-main.765","article-title":"Finding skill neurons in pre-trained transformer-based language models","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Wang","year":"2022"},{"key":"2024122021045224100_bib114","doi-asserted-by":"publisher","DOI":"10.1201\/9781003205388-2","article-title":"What artificial neural networks can tell us about human language acquisition","volume":"arXiv:2208.07998","author":"Warstadt","year":"2022","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib115","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.conll-babylm.1","article-title":"Call for papers \u2013 The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus","volume":"arXiv:2301.11796","author":"Warstadt","year":"2023","journal-title":"ArXiv preprint"},{"key":"2024122021045224100_bib116","doi-asserted-by":"publisher","first-page":"294","DOI":"10.18653\/v1\/2023.conll-1.20","article-title":"Mind the instructions: A holistic evaluation of consistency and interactions in prompt-based learning","volume-title":"Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL)","author":"Weber","year":"2023"},{"key":"2024122021045224100_bib117","article-title":"On the predictive power of neural language models for human real-time comprehension behavior","volume":"arXiv:2006.01912","author":"Wilcox","year":"2020","journal-title":"ArXiv preprint"},{"volume-title":"Philosophical investigations. Philosophische Untersuchungen","year":"1953","author":"Wittgenstein","key":"2024122021045224100_bib118"},{"key":"2024122021045224100_bib119","doi-asserted-by":"publisher","first-page":"3687","DOI":"10.18653\/v1\/D19-1382","article-title":"PAWS-X: A cross-lingual adversarial dataset for paraphrase identification","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Yang","year":"2019"},{"key":"2024122021045224100_bib120","doi-asserted-by":"publisher","first-page":"1112","DOI":"10.18653\/v1\/2021.acl-long.90","article-title":"When do you need billions of words of pretraining data?","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Zhang","year":"2021"},{"key":"2024122021045224100_bib121","doi-asserted-by":"publisher","first-page":"1298","DOI":"10.18653\/v1\/N19-1131","article-title":"PAWS: Paraphrase adversaries from word scrambling","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Zhang","year":"2019"}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/50\/4\/1507\/2480441\/coli_a_00529.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/50\/4\/1507\/2480441\/coli_a_00529.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,12,20]],"date-time":"2024-12-20T21:05:20Z","timestamp":1734728720000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/50\/4\/1507\/123794\/From-Form-s-to-Meaning-Probing-the-Semantic-Depths"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":121,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2024,12,1]]},"published-print":{"date-parts":[[2024,12,1]]}},"URL":"https:\/\/doi.org\/10.1162\/coli_a_00529","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"type":"print","value":"0891-2017"},{"type":"electronic","value":"1530-9312"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}