{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,30]],"date-time":"2026-08-30T09:02:06Z","timestamp":1788080526057,"version":"build-2784847793"},"reference-count":393,"publisher":"MIT Press","issue":"1","license":[{"start":{"date-parts":[[2024,4,24]],"date-time":"2024-04-24T00:00:00Z","timestamp":1713916800000},"content-version":"vor","delay-in-days":114,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,3,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Transformer language models have received widespread public attention, yet their generated text is often surprising even to NLP researchers. In this survey, we discuss over 250 recent studies of English language model behavior before task-specific fine-tuning. Language models possess basic capabilities in syntax, semantics, pragmatics, world knowledge, and reasoning, but these capabilities are sensitive to specific inputs and surface features. Despite dramatic increases in generated text quality as models scale to hundreds of billions of parameters, the models are still prone to unfactual responses, commonsense errors, memorized text, and social biases. Many of these weaknesses can be framed as over-generalizations or under-generalizations of learned patterns in text. We synthesize recent results to highlight what is currently known about large language model capabilities, thus providing a resource for applied work and for research in adjacent fields that use language models.<\/jats:p>","DOI":"10.1162\/coli_a_00492","type":"journal-article","created":{"date-parts":[[2023,11,15]],"date-time":"2023-11-15T14:26:05Z","timestamp":1700058365000},"page":"293-350","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":85,"title":["Language Model Behavior: A Comprehensive Survey"],"prefix":"10.1162","volume":"50","author":[{"given":"Tyler A.","family":"Chang","sequence":"first","affiliation":[{"name":"Department of Cognitive Science, UC San Diego. tachang@ucsd.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Benjamin K.","family":"Bergen","sequence":"additional","affiliation":[{"name":"Department of Cognitive Science, UC San Diego. bkbergen@ucsd.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,3,1]]},"reference":[{"key":"2024042419444642400_bib1","doi-asserted-by":"publisher","first-page":"6907","DOI":"10.18653\/v1\/2022.acl-long.476","article-title":"Word order does matter and shuffled language models know it","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Abdou","year":"2022"},{"key":"2024042419444642400_bib2","doi-asserted-by":"publisher","first-page":"298","DOI":"10.1145\/3461702.3462624","article-title":"Persistent anti-Muslim bias in large language models","volume-title":"The AAAI\/ACM Conference on AI, Ethics, and Society","author":"Abid","year":"2021"},{"key":"2024042419444642400_bib3","article-title":"How to query language models?","author":"Adolphs","year":"2021","journal-title":"ArXiv"},{"key":"2024042419444642400_bib4","article-title":"Using large language models to simulate multiple humans","author":"Aher","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib5","doi-asserted-by":"publisher","first-page":"42","DOI":"10.18653\/v1\/2021.blackboxnlp-1.4","article-title":"The language model understood the prompt was ambiguous: Probing syntactic uncertainty through generation","volume-title":"Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"Aina","year":"2021"},{"key":"2024042419444642400_bib6","doi-asserted-by":"publisher","first-page":"76","DOI":"10.18653\/v1\/2022.gebnlp-1.9","article-title":"Challenges in measuring bias via open-ended language generation","volume-title":"Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP)","author":"Aky\u00fcrek","year":"2022"},{"key":"2024042419444642400_bib7","doi-asserted-by":"publisher","first-page":"2824","DOI":"10.18653\/v1\/2022.naacl-main.203","article-title":"Using natural sentence prompts for understanding biases in language models","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Alnegheimish","year":"2022"},{"key":"2024042419444642400_bib8","doi-asserted-by":"publisher","first-page":"79","DOI":"10.18653\/v1\/2021.blackboxnlp-1.7","article-title":"ALL dolphins are intelligent and SOME are friendly: Probing BERT for nouns\u2019 semantic properties and their prototypicality","volume-title":"Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"Apidianaki","year":"2021"},{"key":"2024042419444642400_bib9","doi-asserted-by":"publisher","first-page":"1242","DOI":"10.18653\/v1\/2020.coling-main.107","article-title":"Always keep your target in mind: Studying semantics and improving performance of neural lexical substitution","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"Arefyev","year":"2020"},{"key":"2024042419444642400_bib10","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1017\/pan.2023.2","article-title":"Out of one, many: Using language models to simulate human samples","author":"Argyle","year":"2023","journal-title":"Political Analysis"},{"key":"2024042419444642400_bib11","doi-asserted-by":"publisher","first-page":"1778","DOI":"10.18653\/v1\/2021.findings-acl.155","article-title":"How reliable are model diagnostics?","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Aribandi","year":"2021"},{"key":"2024042419444642400_bib12","doi-asserted-by":"publisher","first-page":"405","DOI":"10.18653\/v1\/2022.conll-1.28","article-title":"Characterizing verbatim short-term memory in neural language models","volume-title":"Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL)","author":"Armeni","year":"2022"},{"key":"2024042419444642400_bib13","doi-asserted-by":"publisher","first-page":"4597","DOI":"10.18653\/v1\/2021.findings-acl.404","article-title":"PROST: Physical reasoning about objects through space and time","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Aroca-Ouellette","year":"2021"},{"key":"2024042419444642400_bib14","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.c3nlp-1.12","article-title":"Probing pre-trained language models for cross-cultural differences in values","author":"Arora","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib15","doi-asserted-by":"publisher","first-page":"3973","DOI":"10.18653\/v1\/2022.findings-emnlp.293","article-title":"On the role of bidirectionality in language model pre-training","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Artetxe","year":"2022"},{"key":"2024042419444642400_bib16","article-title":"Does BERT agree? Evaluating knowledge of structure dependence through agreement relations","author":"Bacon","year":"2019","journal-title":"ArXiv"},{"key":"2024042419444642400_bib17","doi-asserted-by":"publisher","first-page":"548","DOI":"10.18653\/v1\/2021.sigdial-1.57","article-title":"Assessing political prudence of open-domain chatbots","volume-title":"Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue","author":"Bang","year":"2021"},{"key":"2024042419444642400_bib18","first-page":"1","article-title":"Unmasking contextual stereotypes: Measuring and mitigating BERT\u2019s gender bias","volume-title":"Proceedings of the Second Workshop on Gender Bias in Natural Language Processing","author":"Bartl","year":"2020"},{"issue":"2","key":"2024042419444642400_bib19","doi-asserted-by":"publisher","first-page":"312","DOI":"10.1111\/tops.12141","article-title":"The non-redundant contributions of Marr\u2019s three levels of analysis for explaining information-processing mechanisms","volume":"7","author":"Bechtel","year":"2015","journal-title":"Topics in Cognitive Science"},{"issue":"1","key":"2024042419444642400_bib20","doi-asserted-by":"publisher","first-page":"207","DOI":"10.1162\/coli_a_00422","article-title":"Probing classifiers: Promises, shortcomings, and advances","volume":"48","author":"Belinkov","year":"2022","journal-title":"Computational Linguistics"},{"key":"2024042419444642400_bib21","doi-asserted-by":"publisher","first-page":"2554","DOI":"10.18653\/v1\/2021.findings-emnlp.218","article-title":"Probing pre-trained language models for semantic attributes and their values","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2021","author":"Beloucif","year":"2021"},{"key":"2024042419444642400_bib22","doi-asserted-by":"publisher","first-page":"610","DOI":"10.1145\/3442188.3445922","article-title":"On the dangers of stochastic parrots: Can language models be too big?","volume-title":"Proceedings of the ACM Conference on Fairness, Accountability, and Transparency","author":"Bender","year":"2021"},{"key":"2024042419444642400_bib23","doi-asserted-by":"publisher","first-page":"5185","DOI":"10.18653\/v1\/2020.acl-main.463","article-title":"Climbing towards NLU: On meaning, form, and understanding in the age of data","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Bender","year":"2020"},{"issue":"7","key":"2024042419444642400_bib24","doi-asserted-by":"publisher","first-page":"1207","DOI":"10.1111\/j.1551-6709.2011.01189.x","article-title":"Poverty of the stimulus revisited","volume":"35","author":"Berwick","year":"2011","journal-title":"Cognitive Science"},{"key":"2024042419444642400_bib25","article-title":"Thinking aloud: Dynamic context generation improves zero-shot reasoning performance of GPT-2","author":"Betz","year":"2021","journal-title":"ArXiv"},{"key":"2024042419444642400_bib26","doi-asserted-by":"publisher","first-page":"4164","DOI":"10.18653\/v1\/2021.naacl-main.328","article-title":"Is incoherence surprising? Targeted evaluation of coherence prediction from language models","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Beyer","year":"2021"},{"key":"2024042419444642400_bib27","doi-asserted-by":"publisher","first-page":"298","DOI":"10.18653\/v1\/2022.inlg-main.25","article-title":"Analogy generation by prompting large language models: A case study of instructGPT","volume-title":"Proceedings of the 15th International Conference on Natural Language Generation","author":"Bhavya","year":"2022"},{"issue":"6","key":"2024042419444642400_bib28","doi-asserted-by":"publisher","first-page":"e2218523120","DOI":"10.1073\/pnas.2218523120","article-title":"Using cognitive psychology to understand GPT-3","volume":"120","author":"Binz","year":"2023","journal-title":"Proceedings of the National Academy of Sciences of the United States of America"},{"key":"2024042419444642400_bib29","doi-asserted-by":"publisher","first-page":"5454","DOI":"10.18653\/v1\/2020.acl-main.485","article-title":"Language (technology) is power: A critical survey of \u201cbias\u201d in NLP","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Blodgett","year":"2020"},{"key":"2024042419444642400_bib30","article-title":"On the opportunities and risks of foundation models","author":"Bommasani","year":"2021","journal-title":"ArXiv"},{"key":"2024042419444642400_bib31","first-page":"2206","article-title":"Improving language models by retrieving from trillions of tokens","volume-title":"International Conference on Machine Learning","author":"Borgeaud","year":"2022"},{"key":"2024042419444642400_bib32","doi-asserted-by":"publisher","first-page":"7484","DOI":"10.18653\/v1\/2022.acl-long.516","article-title":"The dangers of underclaiming: Reasons for caution when reporting how NLP systems fail","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Bowman","year":"2022"},{"key":"2024042419444642400_bib33","doi-asserted-by":"publisher","first-page":"3624","DOI":"10.18653\/v1\/2022.naacl-main.265","article-title":"How conservative are language models? Adapting to the introduction of gender-neutral pronouns","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Brandl","year":"2022"},{"key":"2024042419444642400_bib34","doi-asserted-by":"publisher","first-page":"2280","DOI":"10.1145\/3531146.3534642","article-title":"What does it mean for a language model to preserve privacy?","volume-title":"Proceedings of the ACM Conference on Fairness, Accountability, and Transparency","author":"Brown","year":"2022"},{"key":"2024042419444642400_bib35","first-page":"1877","article-title":"Language models are few-shot learners","volume-title":"Advances in Neural Information Processing Systems","author":"Brown","year":"2020"},{"key":"2024042419444642400_bib36","article-title":"Build a medical sentence matching application using BERT and Amazon SageMaker","author":"Broyde","year":"2021","journal-title":"AWS Machine Learning Blog"},{"key":"2024042419444642400_bib37","article-title":"Isotropy in the contextual embedding space: Clusters and manifolds","volume-title":"International Conference on Learning Representations","author":"Cai","year":"2021"},{"key":"2024042419444642400_bib38","doi-asserted-by":"publisher","first-page":"5796","DOI":"10.18653\/v1\/2022.acl-long.398","article-title":"Can prompt probe pretrained language models? Understanding the invisible risks from a causal view","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Cao","year":"2022"},{"key":"2024042419444642400_bib39","doi-asserted-by":"publisher","first-page":"1860","DOI":"10.18653\/v1\/2021.acl-long.146","article-title":"Knowledgeable or educated guess? Revisiting language models as knowledge bases","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Cao","year":"2021"},{"key":"2024042419444642400_bib40","article-title":"Quantifying memorization across neural language models","volume-title":"International Conference on Learning Representations","author":"Carlini","year":"2023"},{"key":"2024042419444642400_bib41","first-page":"2633","article-title":"Extracting training data from large language models","volume-title":"USENIX Security Symposium","author":"Carlini","year":"2021"},{"key":"2024042419444642400_bib42","volume-title":"Syntax: A Generative Introduction","author":"Carnie","year":"2002"},{"key":"2024042419444642400_bib43","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.156","article-title":"Identifying and manipulating the personality traits of language models","author":"Caron","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib44","doi-asserted-by":"publisher","first-page":"119","DOI":"10.18653\/v1\/2022.emnlp-main.9","article-title":"The geometry of multilingual language model representations","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Chang","year":"2022"},{"key":"2024042419444642400_bib45","doi-asserted-by":"publisher","first-page":"4322","DOI":"10.18653\/v1\/2021.acl-long.333","article-title":"Convolutions and self-attention: Re-interpreting relative positions in pre-trained language models","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Chang","year":"2021"},{"key":"2024042419444642400_bib46","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1162\/tacl_a_00444","article-title":"Word acquisition in neural language models","volume":"10","author":"Chang","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024042419444642400_bib47","first-page":"28","article-title":"Look at that! BERT can be easily distracted from paying attention to morphosyntax","volume-title":"Proceedings of the Society for Computation in Linguistics 2021","author":"Chaves","year":"2021"},{"key":"2024042419444642400_bib48","article-title":"A critical appraisal of equity in conversational AI: Evidence from auditing GPT-3\u2019s dialogues with different publics on climate change and Black Lives Matter","volume":"arXiv:2209.13627","author":"Chen","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib49","article-title":"Evaluating large language models trained on code","author":"Chen","year":"2021","journal-title":"ArXiv"},{"key":"2024042419444642400_bib50","doi-asserted-by":"publisher","first-page":"6813","DOI":"10.18653\/v1\/2020.emnlp-main.553","article-title":"Pretrained language model embryology: The birth of ALBERT","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Chiang","year":"2020"},{"key":"2024042419444642400_bib51","doi-asserted-by":"publisher","first-page":"228","DOI":"10.18653\/v1\/2021.blackboxnlp-1.16","article-title":"Relating neural text degeneration to exposure bias","volume-title":"Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"Chiang","year":"2021"},{"key":"2024042419444642400_bib52","doi-asserted-by":"publisher","first-page":"2922","DOI":"10.18653\/v1\/2021.findings-acl.258","article-title":"Modeling the influence of verb aspect on the activation of typical event locations with BERT","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Cho","year":"2021"},{"key":"2024042419444642400_bib53","doi-asserted-by":"publisher","first-page":"1477","DOI":"10.18653\/v1\/2021.emnlp-main.111","article-title":"Stepmothers are mean and academics are pretentious: What do pretrained language models learn about you?","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Choenni","year":"2021"},{"key":"2024042419444642400_bib54","doi-asserted-by":"publisher","first-page":"8281","DOI":"10.18653\/v1\/2022.acl-long.568","article-title":"The grammar-learning trajectories of neural language models","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Choshen","year":"2022"},{"key":"2024042419444642400_bib55","doi-asserted-by":"publisher","first-page":"12710","DOI":"10.1609\/aaai.v35i14.17505","article-title":"How linguistically fair are multilingual pre-trained language models?","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Choudhury","year":"2021"},{"key":"2024042419444642400_bib56","article-title":"PaLM: Scaling language modeling with Pathways","author":"Chowdhery","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib57","doi-asserted-by":"publisher","first-page":"100","DOI":"10.18653\/v1\/2022.acl-short.12","article-title":"Buy Tesla, sell Ford: Assessing implicit stock market preference in pre-trained language models","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Chuang","year":"2022"},{"key":"2024042419444642400_bib58","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-short.92","article-title":"Black-box language model explanation by context length probing","author":"C\u00edfka","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib59","doi-asserted-by":"publisher","first-page":"7282","DOI":"10.18653\/v1\/2021.acl-long.565","article-title":"All that\u2019s \u2018human\u2019 is not gold: Evaluating human evaluation of generated text","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Clark","year":"2021"},{"key":"2024042419444642400_bib60","doi-asserted-by":"publisher","first-page":"276","DOI":"10.18653\/v1\/W19-4828","article-title":"What does BERT look at? An analysis of BERT\u2019s attention","volume-title":"Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Clark","year":"2019"},{"key":"2024042419444642400_bib61","article-title":"LaMDA: Language models for dialog applications","author":"Cohen","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib62","first-page":"373","article-title":"MiQA: A benchmark for inference on metaphorical questions","volume-title":"Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)","author":"Com\u0219a","year":"2022"},{"key":"2024042419444642400_bib63","doi-asserted-by":"publisher","first-page":"17","DOI":"10.18653\/v1\/2022.csrr-1.3","article-title":"Psycholinguistic diagnosis of language models\u2019 commonsense reasoning","volume-title":"Proceedings of the First Workshop on Commonsense Representation and Reasoning (CSRR 2022)","author":"Cong","year":"2022"},{"key":"2024042419444642400_bib64","doi-asserted-by":"publisher","first-page":"1249","DOI":"10.1162\/tacl_a_00425","article-title":"Quantifying social biases in NLP: A generalization and empirical comparison of extrinsic fairness metrics","volume":"9","author":"Czarnowska","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024042419444642400_bib65","doi-asserted-by":"publisher","first-page":"2094","DOI":"10.18653\/v1\/2022.findings-emnlp.153","article-title":"Scientific and creative analogies in pretrained language models","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Czinczoll","year":"2022"},{"key":"2024042419444642400_bib66","doi-asserted-by":"publisher","first-page":"852","DOI":"10.3389\/fpsyg.2015.00852","article-title":"What exactly is Universal Grammar, and has anyone seen it?","volume":"6","author":"Dabrowska","year":"2015","journal-title":"Frontiers in Psychology"},{"key":"2024042419444642400_bib67","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.247","article-title":"Why can GPT learn in-context? Language models secretly perform gradient descent as meta-optimizers","author":"Dai","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib68","doi-asserted-by":"publisher","first-page":"2978","DOI":"10.18653\/v1\/P19-1285","article-title":"Transformer-XL: Attentive language models beyond a fixed-length context","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Dai","year":"2019"},{"key":"2024042419444642400_bib69","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.893","article-title":"Analyzing Transformers in embedding space","author":"Dar","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib70","article-title":"Language models show human-like content effects on reasoning","author":"Dasgupta","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib71","doi-asserted-by":"publisher","first-page":"396","DOI":"10.18653\/v1\/2020.conll-1.32","article-title":"Discourse structure interacts with reference but not syntax in neural language models","volume-title":"Proceedings of the 24th Conference on Computational Natural Language Learning","author":"Davis","year":"2020"},{"key":"2024042419444642400_bib72","doi-asserted-by":"publisher","first-page":"1173","DOI":"10.18653\/v1\/D19-1109","article-title":"Commonsense knowledge mining from pretrained models","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Davison","year":"2019"},{"key":"2024042419444642400_bib73","doi-asserted-by":"publisher","first-page":"80","DOI":"10.18653\/v1\/2022.blackboxnlp-1.7","article-title":"Is it smaller than a tennis ball? Language models play the game of twenty questions","volume-title":"Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"De Bruyn","year":"2022"},{"key":"2024042419444642400_bib74","doi-asserted-by":"publisher","first-page":"2232","DOI":"10.18653\/v1\/2021.eacl-main.190","article-title":"Stereotype and skew: Quantifying gender bias in pre-trained and fine-tuned language models","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"de Vassimon Manela","year":"2021"},{"key":"2024042419444642400_bib75","doi-asserted-by":"publisher","first-page":"1968","DOI":"10.18653\/v1\/2021.emnlp-main.150","article-title":"Harms of gender exclusivity and challenges in non-binary representation in language technologies","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Dev","year":"2021"},{"key":"2024042419444642400_bib76","first-page":"246","article-title":"On measures of biases and harms in NLP","volume-title":"Findings of the Association for Computational Linguistics: AACL-IJCNLP 2022","author":"Dev","year":"2022"},{"key":"2024042419444642400_bib77","first-page":"4171","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2024042419444642400_bib78","doi-asserted-by":"publisher","first-page":"862","DOI":"10.1145\/3442188.3445924","article-title":"BOLD: Dataset and metrics for measuring biases in open-ended language generation","volume-title":"Proceedings of the ACM Conference on Fairness, Accountability, and Transparency","author":"Dhamala","year":"2021"},{"key":"2024042419444642400_bib79","doi-asserted-by":"publisher","first-page":"7250","DOI":"10.18653\/v1\/2022.acl-long.501","article-title":"Is GPT-3 text indistinguishable from human text? Scarecrow: A framework for scrutinizing machine text","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Dou","year":"2022"},{"key":"2024042419444642400_bib80","article-title":"Shortcut learning of large language models in natural language understanding: A survey","author":"Du","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib81","doi-asserted-by":"publisher","first-page":"5436","DOI":"10.24963\/ijcai.2022\/762","article-title":"A survey of vision-language pre-trained models","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence","author":"Du","year":"2022"},{"issue":"3","key":"2024042419444642400_bib82","doi-asserted-by":"publisher","first-page":"733","DOI":"10.1162\/coli_a_00445","article-title":"Position information in transformers: An overview","volume":"48","author":"Dufter","year":"2022","journal-title":"Computational Linguistics"},{"key":"2024042419444642400_bib83","doi-asserted-by":"publisher","first-page":"12763","DOI":"10.1609\/aaai.v37i11.26501","article-title":"Real or fake text?: Investigating human ability to detect boundaries between human-written and machine-generated text","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Dugan","year":"2023"},{"key":"2024042419444642400_bib84","article-title":"Measuring causal effects of data statistics on language model\u2019s \u2018factual\u2019 predictions","author":"Elazar","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib85","doi-asserted-by":"publisher","first-page":"1012","DOI":"10.1162\/tacl_a_00410","article-title":"Measuring and improving consistency in pretrained language models","volume":"9","author":"Elazar","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024042419444642400_bib86","doi-asserted-by":"publisher","first-page":"34","DOI":"10.1162\/tacl_a_00298","article-title":"What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models","volume":"8","author":"Ettinger","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024042419444642400_bib87","first-page":"1","article-title":"Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity","volume":"23","author":"Fedus","year":"2022","journal-title":"Journal of Machine Learning Research"},{"key":"2024042419444642400_bib88","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.507","article-title":"Towards WinoQueer: Developing a benchmark for anti-queer bias in large language models","volume-title":"Queer in AI Workshop","author":"Felkner","year":"2022"},{"key":"2024042419444642400_bib89","doi-asserted-by":"publisher","first-page":"1828","DOI":"10.18653\/v1\/2021.acl-long.144","article-title":"Causal analysis of syntactic agreement mechanisms in neural language models","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Finlayson","year":"2021"},{"issue":"6084","key":"2024042419444642400_bib90","doi-asserted-by":"publisher","first-page":"998","DOI":"10.1126\/science.1218633","article-title":"Predicting pragmatic reasoning in language games","volume":"336","author":"Frank","year":"2012","journal-title":"Science"},{"key":"2024042419444642400_bib91","doi-asserted-by":"publisher","first-page":"56","DOI":"10.18653\/v1\/W17-3207","article-title":"Beam search strategies for neural machine translation","volume-title":"Proceedings of the First Workshop on Neural Machine Translation","author":"Freitag","year":"2017"},{"issue":"1","key":"2024042419444642400_bib92","doi-asserted-by":"publisher","first-page":"145","DOI":"10.5195\/jmla.2018.280","article-title":"Semantic Scholar","volume":"106","author":"Fricke","year":"2018","journal-title":"Journal of the Medical Library Association"},{"key":"2024042419444642400_bib93","article-title":"Logical tasks for measuring extrapolation and rule comprehension","author":"Fujisawa","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib94","doi-asserted-by":"publisher","first-page":"1747","DOI":"10.1145\/3531146.3533229","article-title":"Predictability and surprise in large generative models","volume-title":"Proceedings of the ACM Conference on Fairness, Accountability, and Transparency","author":"Ganguli","year":"2022"},{"key":"2024042419444642400_bib95","article-title":"Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned","author":"Ganguli","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib96","doi-asserted-by":"publisher","first-page":"70","DOI":"10.18653\/v1\/2020.acl-demos.10","article-title":"SyntaxGym: An online platform for targeted evaluation of language models","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations","author":"Gauthier","year":"2020"},{"key":"2024042419444642400_bib97","doi-asserted-by":"publisher","DOI":"10.1093\/acrefore\/9780199384655.013.29","article-title":"Lexical semantics","author":"Geeraerts","year":"2017","journal-title":"Oxford Research Encyclopedia of Linguistics"},{"key":"2024042419444642400_bib98","doi-asserted-by":"publisher","first-page":"3356","DOI":"10.18653\/v1\/2020.findings-emnlp.301","article-title":"RealToxicityPrompts: Evaluating neural toxic degeneration in language models","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Gehman","year":"2020"},{"key":"2024042419444642400_bib99","doi-asserted-by":"publisher","first-page":"163","DOI":"10.18653\/v1\/2020.blackboxnlp-1.16","article-title":"Neural natural language inference models partially embed theories of lexical entailment and negation","volume-title":"Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"Geiger","year":"2020"},{"key":"2024042419444642400_bib100","doi-asserted-by":"publisher","first-page":"30","DOI":"10.18653\/v1\/2022.emnlp-main.3","article-title":"Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Geva","year":"2022"},{"key":"2024042419444642400_bib101","doi-asserted-by":"publisher","first-page":"5484","DOI":"10.18653\/v1\/2021.emnlp-main.446","article-title":"Transformer feed-forward layers are key-value memories","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Geva","year":"2021"},{"key":"2024042419444642400_bib102","article-title":"Assessing BERT\u2019s syntactic abilities","author":"Goldberg","year":"2019","journal-title":"ArXiv"},{"key":"2024042419444642400_bib103","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1163\/9789004368811_003","article-title":"Logic and conversation","author":"Grice","year":"1975","journal-title":"Syntax and Semantics: Vol. 3: Speech Acts"},{"key":"2024042419444642400_bib104","doi-asserted-by":"publisher","first-page":"173","DOI":"10.18653\/v1\/2022.flp-1.25","article-title":"On the cusp of comprehensibility: Can language models distinguish between metaphors and nonsense?","volume-title":"Proceedings of the 3rd Workshop on Figurative Language Processing (FLP)","author":"Grici\u016bt\u0117","year":"2022"},{"key":"2024042419444642400_bib105","doi-asserted-by":"publisher","first-page":"5877","DOI":"10.18653\/v1\/2020.emnlp-main.473","article-title":"Investigating African-American Vernacular English in transformer-based text generation","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Groenwold","year":"2020"},{"key":"2024042419444642400_bib106","doi-asserted-by":"publisher","first-page":"4602","DOI":"10.18653\/v1\/2022.acl-long.315","article-title":"Context matters: A pragmatic study of PLMs\u2019 negation understanding","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Gubelmann","year":"2022"},{"key":"2024042419444642400_bib107","first-page":"3929","article-title":"Retrieval augmented language model pre-training","volume-title":"International Conference on Machine Learning","author":"Guu","year":"2020"},{"key":"2024042419444642400_bib108","doi-asserted-by":"publisher","DOI":"10.1038\/s43588-023-00527-x","article-title":"Machine intuition: Uncovering human-like intuitive decision-making in GPT-3.5","author":"Hagendorff","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib109","doi-asserted-by":"publisher","first-page":"156","DOI":"10.1162\/tacl_a_00306","article-title":"Theoretical limitations of self-attention in neural sequence models","volume":"8","author":"Hahn","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024042419444642400_bib110","article-title":"FOLIO: Natural language reasoning with first-order logic","author":"Han","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib111","doi-asserted-by":"publisher","first-page":"275","DOI":"10.18653\/v1\/2021.blackboxnlp-1.20","article-title":"Analyzing BERT\u2019s knowledge of hypernymy via prompting","volume-title":"Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"Hanna","year":"2021"},{"key":"2024042419444642400_bib112","doi-asserted-by":"publisher","first-page":"3116","DOI":"10.18653\/v1\/2021.findings-emnlp.267","article-title":"Unpacking the interdependent systems of discrimination: Ableist bias in NLP systems through an intersectional lens","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2021","author":"Hassan","year":"2021"},{"key":"2024042419444642400_bib113","doi-asserted-by":"publisher","first-page":"1382","DOI":"10.18653\/v1\/2022.findings-emnlp.99","article-title":"Transformer language models without positional encodings still learn positional information","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Haviv","year":"2022"},{"key":"2024042419444642400_bib114","doi-asserted-by":"publisher","first-page":"4653","DOI":"10.18653\/v1\/2020.emnlp-main.376","article-title":"Investigating representations of verb bias in neural language models","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Hawkins","year":"2020"},{"key":"2024042419444642400_bib115","doi-asserted-by":"publisher","first-page":"7875","DOI":"10.18653\/v1\/2022.acl-long.543","article-title":"Can pre-trained language models interpret similes as smart as human?","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"He","year":"2022"},{"key":"2024042419444642400_bib116","doi-asserted-by":"publisher","first-page":"10758","DOI":"10.1609\/aaai.v36i10.21321","article-title":"Protecting intellectual property of language generation APIs with lexical watermark","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"He","year":"2022"},{"key":"2024042419444642400_bib117","doi-asserted-by":"publisher","first-page":"566","DOI":"10.1145\/3461702.3462578","article-title":"The Earth is flat and the Sun is not a star: The susceptibility of GPT-2 to universal adversarial triggers","volume-title":"Proceedings of the 2021 AAAI\/ACM Conference on AI, Ethics, and Society","author":"Heidenreich","year":"2021"},{"key":"2024042419444642400_bib118","article-title":"Measuring massive multitask language understanding","volume-title":"International Conference on Learning Representations","author":"Hendrycks","year":"2021"},{"key":"2024042419444642400_bib119","article-title":"Measuring mathematical problem solving with the MATH dataset","volume-title":"Advances in Neural Information Processing Systems Datasets and Benchmarks Track","author":"Hendrycks","year":"2021"},{"key":"2024042419444642400_bib120","article-title":"Scaling laws and interpretability of learning from repeated data","author":"Hernandez","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib121","doi-asserted-by":"publisher","first-page":"6997","DOI":"10.18653\/v1\/2022.acl-long.482","article-title":"Challenges and strategies in cross-cultural NLP","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Hershcovich","year":"2022"},{"key":"2024042419444642400_bib122","first-page":"30016","article-title":"Training compute-optimal large language models","volume-title":"Advances in Neural Information Processing Systems","author":"Hoffmann","year":"2022"},{"key":"2024042419444642400_bib123","article-title":"The curious case of neural text degeneration","volume-title":"International Conference on Learning Representations","author":"Holtzman","year":"2020"},{"key":"2024042419444642400_bib124","doi-asserted-by":"publisher","first-page":"716","DOI":"10.18653\/v1\/2022.acl-short.81","article-title":"An analysis of negation in natural language understanding corpora","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Hossain","year":"2022"},{"key":"2024042419444642400_bib125","doi-asserted-by":"publisher","first-page":"9106","DOI":"10.18653\/v1\/2020.emnlp-main.732","article-title":"An analysis of natural language inference benchmarks through the lens of negation","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Hossain","year":"2020"},{"key":"2024042419444642400_bib126","doi-asserted-by":"publisher","first-page":"272","DOI":"10.18653\/v1\/2022.blackboxnlp-1.22","article-title":"On the compositional generalization gap of in-context learning","volume-title":"Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"Hosseini","year":"2022"},{"key":"2024042419444642400_bib127","first-page":"323","article-title":"A closer look at the performance of neural language models on reflexive anaphor licensing","volume-title":"Proceedings of the Society for Computation in Linguistics 2020","author":"Hu","year":"2020"},{"key":"2024042419444642400_bib128","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.230","article-title":"A fine-grained comparison of pragmatic language understanding in humans and language models","author":"Hu","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib129","doi-asserted-by":"publisher","first-page":"1725","DOI":"10.18653\/v1\/2020.acl-main.158","article-title":"A systematic assessment of syntactic generalization in neural language models","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Hu","year":"2020"},{"key":"2024042419444642400_bib130","doi-asserted-by":"publisher","first-page":"2038","DOI":"10.18653\/v1\/2022.findings-emnlp.148","article-title":"Are large pre-trained language models leaking your personal information?","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Huang","year":"2022"},{"key":"2024042419444642400_bib131","doi-asserted-by":"publisher","first-page":"624","DOI":"10.18653\/v1\/2021.conll-1.49","article-title":"BabyBERTa: Learning more grammar with small-scale child-directed language","volume-title":"Proceedings of the 25th Conference on Computational Natural Language Learning","author":"Huebner","year":"2021"},{"key":"2024042419444642400_bib132","article-title":"State-of-the-art generalisation research in NLP: A taxonomy and review","author":"Hupkes","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib133","article-title":"Implicit causality in GPT-2: A case study","author":"Huynh","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib134","doi-asserted-by":"publisher","first-page":"1808","DOI":"10.18653\/v1\/2020.acl-main.164","article-title":"Automatic detection of generated text is easiest when humans are fooled","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Ippolito","year":"2020"},{"key":"2024042419444642400_bib135","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.inlg-main.3","article-title":"Preventing verbatim memorization in language models gives a false sense of privacy","author":"Ippolito","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib136","article-title":"OPT-IML: Scaling language model instruction meta learning through the lens of generalization","author":"Iyer","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib137","first-page":"3543","article-title":"Attention is not Explanation","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Jain","year":"2019"},{"issue":"11","key":"2024042419444642400_bib138","doi-asserted-by":"publisher","first-page":"e2208839120","DOI":"10.1073\/pnas.2208839120","article-title":"Human heuristics for AI-generated language are flawed","volume":"120","author":"Jakesch","year":"2023","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2024042419444642400_bib139","first-page":"52","article-title":"Can large language models truly understand prompts? A case study with negated prompts","volume-title":"Proceedings of the 1st Transfer Learning for Natural Language Processing Workshop","author":"Jang","year":"2022"},{"key":"2024042419444642400_bib140","doi-asserted-by":"publisher","first-page":"2296","DOI":"10.18653\/v1\/2020.coling-main.208","article-title":"Automatic detection of machine generated text: A critical survey","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"Jawahar","year":"2020"},{"key":"2024042419444642400_bib141","doi-asserted-by":"publisher","first-page":"2586","DOI":"10.18653\/v1\/2020.findings-emnlp.235","article-title":"Learning numeral embedding","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Jiang","year":"2020"},{"key":"2024042419444642400_bib142","article-title":"MPI: Evaluating and inducing personality in pre-trained language models","author":"Jiang","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib143","doi-asserted-by":"publisher","first-page":"6941","DOI":"10.18653\/v1\/2021.acl-long.540","article-title":"Learning prototypical functions for physical artifacts","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Jiang","year":"2021"},{"key":"2024042419444642400_bib144","doi-asserted-by":"publisher","first-page":"423","DOI":"10.1162\/tacl_a_00324","article-title":"How can we know what language models know?","volume":"8","author":"Jiang","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024042419444642400_bib145","article-title":"Perspective API","author":"Jigsaw","year":"2017","journal-title":"Google Jigsaw"},{"key":"2024042419444642400_bib146","first-page":"28458","article-title":"When to make exceptions: Exploring language models as accounts of human moral judgment","volume-title":"Advances in Neural Information Processing Systems","author":"Jin","year":"2022"},{"key":"2024042419444642400_bib147","doi-asserted-by":"publisher","first-page":"87","DOI":"10.18653\/v1\/2022.umios-1.10","article-title":"Probing script knowledge from pre-trained models","volume-title":"Proceedings of the Workshop on Unimodal and Multimodal Induction of Linguistic Structures (UM-IoS)","author":"Jin","year":"2022"},{"key":"2024042419444642400_bib148","article-title":"The ghost in the machine has an American accent: Value conflict in GPT-3","author":"Johnson","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib149","article-title":"A.I. is mastering language. Should we trust what it says?","author":"Johnson","year":"2022","journal-title":"The New York Times"},{"key":"2024042419444642400_bib150","first-page":"2876","article-title":"The role of physical inference in pronoun resolution","volume-title":"Proceedings of the Annual Meeting of the Cognitive Science Society","author":"Jones","year":"2021"},{"key":"2024042419444642400_bib151","first-page":"482","article-title":"Distributional semantics still can\u2019t account for affordances","volume-title":"Proceedings of the Annual Meeting of the Cognitive Science Society","author":"Jones","year":"2022"},{"key":"2024042419444642400_bib152","doi-asserted-by":"publisher","first-page":"6282","DOI":"10.18653\/v1\/2020.acl-main.560","article-title":"The state and fate of linguistic diversity and inclusion in the NLP world","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Joshi","year":"2020"},{"key":"2024042419444642400_bib153","first-page":"779","article-title":"Investigating the performance of transformer-based NLI models on presuppositional inferences","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Kabbara","year":"2022"},{"key":"2024042419444642400_bib154","article-title":"Language models (mostly) know what they know","author":"Kadavath","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib155","article-title":"KAMEL: Knowledge analysis with multitoken entities in language models","volume-title":"4th Conference on Automated Knowledge Base Construction","author":"Kalo","year":"2022"},{"key":"2024042419444642400_bib156","article-title":"Large language models struggle to learn long-tail knowledge","author":"Kandpal","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib157","first-page":"10697","article-title":"Deduplicating training data mitigates privacy risks in language models","volume-title":"International Conference on Machine Learning","author":"Kandpal","year":"2022"},{"key":"2024042419444642400_bib158","article-title":"Scaling laws for neural language models","author":"Kaplan","year":"2020","journal-title":"ArXiv"},{"key":"2024042419444642400_bib159","article-title":"MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning","author":"Karpas","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib160","doi-asserted-by":"publisher","first-page":"552","DOI":"10.18653\/v1\/2020.conll-1.45","article-title":"Are pretrained language models symbolic reasoners over knowledge?","volume-title":"Proceedings of the 24th Conference on Computational Natural Language Learning","author":"Kassner","year":"2020"},{"key":"2024042419444642400_bib161","doi-asserted-by":"publisher","first-page":"7811","DOI":"10.18653\/v1\/2020.acl-main.698","article-title":"Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Kassner","year":"2020"},{"key":"2024042419444642400_bib162","doi-asserted-by":"publisher","first-page":"2548","DOI":"10.18653\/v1\/2022.findings-emnlp.188","article-title":"Inferring implicit relations in complex questions with language models","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Katz","year":"2022"},{"key":"2024042419444642400_bib163","doi-asserted-by":"publisher","DOI":"10.1111\/cogs.13386","article-title":"Event knowledge in large language models: The gap between the impossible and the unlikely","author":"Kauf","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib164","first-page":"1105","article-title":"Balanced COPA: Countering superficial cues in causal reasoning","author":"Kavumba","year":"2020","journal-title":"Association for Natural Language Processing"},{"key":"2024042419444642400_bib165","doi-asserted-by":"publisher","first-page":"2333","DOI":"10.18653\/v1\/2022.acl-long.166","article-title":"Are prompt-based models clueless?","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Kavumba","year":"2022"},{"key":"2024042419444642400_bib166","doi-asserted-by":"publisher","first-page":"4859","DOI":"10.18653\/v1\/2021.findings-acl.429","article-title":"John praised Mary because _he_? Implicit causality bias and its interaction with explicit cues in LMs","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Kementchedjhieva","year":"2021"},{"key":"2024042419444642400_bib167","article-title":"Generalization through memorization: Nearest neighbor language models","volume-title":"International Conference on Learning Representations","author":"Khandelwal","year":"2020"},{"key":"2024042419444642400_bib168","article-title":"How BPE affects memorization in Transformers","author":"Kharitonov","year":"2021","journal-title":"ArXiv"},{"key":"2024042419444642400_bib169","first-page":"863","article-title":"\u201cno, they did not\u201d: Dialogue response dynamics in pre-trained language models","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Kim","year":"2022"},{"key":"2024042419444642400_bib170","article-title":"A watermark for large language models","author":"Kirchenbauer","year":"2023","journal-title":"ArXiv"},{"key":"2024042419444642400_bib171","first-page":"2611","article-title":"Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models","volume-title":"Advances in Neural Information Processing Systems","author":"Kirk","year":"2021"},{"key":"2024042419444642400_bib172","doi-asserted-by":"publisher","first-page":"52","DOI":"10.18653\/v1\/2020.inlg-1.8","article-title":"Assessing discourse relations in language generation from GPT-2","volume-title":"Proceedings of the 13th International Conference on Natural Language Generation","author":"Ko","year":"2020"},{"key":"2024042419444642400_bib173","first-page":"22199","article-title":"Large language models are zero-shot reasoners","volume-title":"Advances in Neural Information Processing Systems","author":"Kojima","year":"2022"},{"key":"2024042419444642400_bib174","doi-asserted-by":"publisher","first-page":"4365","DOI":"10.18653\/v1\/D19-1445","article-title":"Revealing the dark secrets of BERT","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Kovaleva","year":"2019"},{"key":"2024042419444642400_bib175","article-title":"Bard is getting better at logic and reasoning","author":"Krawczyk","year":"2023","journal-title":"The Keyword: Google Blog"},{"key":"2024042419444642400_bib176","doi-asserted-by":"publisher","first-page":"66","DOI":"10.18653\/v1\/P18-1007","article-title":"Subword regularization: Improving neural network translation models with multiple subword candidates","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Kudo","year":"2018"},{"key":"2024042419444642400_bib177","doi-asserted-by":"publisher","first-page":"166","DOI":"10.18653\/v1\/W19-3823","article-title":"Measuring bias in contextualized word representations","volume-title":"Proceedings of the First Workshop on Gender Bias in Natural Language Processing","author":"Kurita","year":"2019"},{"key":"2024042419444642400_bib178","article-title":"Why do masked neural language models still need common sense knowledge?","author":"Kwon","year":"2019","journal-title":"ArXiv"},{"key":"2024042419444642400_bib179","doi-asserted-by":"publisher","first-page":"1336","DOI":"10.1162\/tacl_a_00430","article-title":"On generative spoken language modeling from raw audio","volume":"9","author":"Lakhotia","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024042419444642400_bib180","first-page":"3226","article-title":"Can transformers process recursive nested constructions, like humans?","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Lakretz","year":"2022"},{"key":"2024042419444642400_bib181","doi-asserted-by":"publisher","first-page":"1204","DOI":"10.18653\/v1\/2022.emnlp-main.79","article-title":"Using commonsense knowledge to answer why-questions","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Lal","year":"2022"},{"key":"2024042419444642400_bib182","article-title":"Can language models handle recursively nested grammatical structures? A case study on comparing models and humans","author":"Lampinen","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib183","doi-asserted-by":"publisher","first-page":"537","DOI":"10.18653\/v1\/2022.findings-emnlp.38","article-title":"Can language models learn from explanations in context?","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Lampinen","year":"2022"},{"key":"2024042419444642400_bib184","doi-asserted-by":"publisher","first-page":"2309","DOI":"10.18653\/v1\/2022.findings-acl.181","article-title":"Does BERT really agree? Fine-grained analysis of lexical dependence on a syntactic task","volume-title":"Findings of the Association for Computational Linguistics: ACL 2022","author":"Lasri","year":"2022"},{"key":"2024042419444642400_bib185","doi-asserted-by":"publisher","first-page":"1808","DOI":"10.18653\/v1\/2022.emnlp-main.118","article-title":"Word order matters when you increase masking","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Lasri","year":"2022"},{"key":"2024042419444642400_bib186","first-page":"37","article-title":"Subject verb agreement error patterns in meaningless sentences: Humans vs. BERT","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Lasri","year":"2022"},{"key":"2024042419444642400_bib187","article-title":"What are large language models used for?","author":"Lee","year":"2023","journal-title":"NVIDIA Blog"},{"key":"2024042419444642400_bib188","doi-asserted-by":"publisher","first-page":"3637","DOI":"10.1145\/3543507.3583199","article-title":"Do language models plagiarize?","volume-title":"The ACM Web Conference","author":"Lee","year":"2023"},{"key":"2024042419444642400_bib189","doi-asserted-by":"publisher","first-page":"8424","DOI":"10.18653\/v1\/2022.acl-long.577","article-title":"Deduplicating training data makes language models better","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Lee","year":"2022"},{"key":"2024042419444642400_bib190","doi-asserted-by":"publisher","first-page":"1971","DOI":"10.18653\/v1\/2021.naacl-main.158","article-title":"Towards few-shot fact-checking via perplexity","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Lee","year":"2021"},{"key":"2024042419444642400_bib191","first-page":"206","article-title":"Can language models capture syntactic associations without surface cues? A case study of reflexive anaphor licensing in English control constructions","volume-title":"Proceedings of the Society for Computation in Linguistics 2022","author":"Lee","year":"2022"},{"key":"2024042419444642400_bib192","doi-asserted-by":"publisher","first-page":"3197","DOI":"10.1145\/3534678.3539147","article-title":"A new generation of perspective API: Efficient multilingual character-level Transformers","volume-title":"Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Lees","year":"2022"},{"key":"2024042419444642400_bib193","doi-asserted-by":"publisher","first-page":"946","DOI":"10.18653\/v1\/2021.naacl-main.73","article-title":"Does BERT pretrained on clinical notes reveal sensitive data?","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Lehman","year":"2021"},{"key":"2024042419444642400_bib194","doi-asserted-by":"publisher","first-page":"3045","DOI":"10.18653\/v1\/2021.emnlp-main.243","article-title":"The power of scale for parameter-efficient prompt tuning","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Lester","year":"2021"},{"key":"2024042419444642400_bib195","doi-asserted-by":"publisher","first-page":"2407","DOI":"10.18653\/v1\/2022.emnlp-main.154","article-title":"SafeText: A benchmark for exploring physical safety in language models","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Levy","year":"2022"},{"key":"2024042419444642400_bib196","doi-asserted-by":"publisher","first-page":"4718","DOI":"10.18653\/v1\/2021.findings-acl.416","article-title":"Investigating memorization of conspiracy theories in text generation","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Levy","year":"2021"},{"key":"2024042419444642400_bib197","article-title":"Counterfactual reasoning: Do language models need world knowledge for causal inference?","volume-title":"Workshop on Neuro Causal and Symbolic AI (nCSI)","author":"Li","year":"2022"},{"key":"2024042419444642400_bib198","doi-asserted-by":"publisher","first-page":"4492","DOI":"10.24963\/ijcai.2021\/612","article-title":"Pretrained language models for text generation: A survey","volume-title":"International Joint Conference on Artificial Intelligence","author":"Li","year":"2021"},{"key":"2024042419444642400_bib199","doi-asserted-by":"publisher","first-page":"11838","DOI":"10.18653\/v1\/2022.emnlp-main.812","article-title":"A systematic investigation of commonsense knowledge in large language models","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Li","year":"2022"},{"key":"2024042419444642400_bib200","article-title":"Is GPT-3 a psychopath? Evaluating large language models from a psychological perspective","author":"Li","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib201","article-title":"Jurassic-1: Technical details and evaluation","author":"Lieber","year":"2021","journal-title":"White Paper. AI21 Labs"},{"key":"2024042419444642400_bib202","doi-asserted-by":"publisher","first-page":"6862","DOI":"10.18653\/v1\/2020.emnlp-main.557","article-title":"Birds have four legs?! NumerSense: Probing numerical commonsense knowledge of pre-trained language models","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Lin","year":"2020"},{"key":"2024042419444642400_bib203","doi-asserted-by":"publisher","first-page":"3214","DOI":"10.18653\/v1\/2022.acl-long.229","article-title":"TruthfulQA: Measuring how models mimic human falsehoods","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Lin","year":"2022"},{"key":"2024042419444642400_bib204","doi-asserted-by":"publisher","first-page":"4437","DOI":"10.18653\/v1\/2022.naacl-main.330","article-title":"Testing the ability of language models to interpret figurative language","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Liu","year":"2022"},{"key":"2024042419444642400_bib205","first-page":"210","article-title":"Do ever larger octopi still amplify reporting biases? Evidence from judgments of typical colour","volume-title":"Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)","author":"Liu","year":"2022"},{"key":"2024042419444642400_bib206","doi-asserted-by":"publisher","first-page":"103654","DOI":"10.1016\/j.artint.2021.103654","article-title":"Quantifying and alleviating political bias in language models","volume":"304","author":"Liu","year":"2022","journal-title":"Artificial Intelligence"},{"key":"2024042419444642400_bib207","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","author":"Liu","year":"2019","journal-title":"ArXiv"},{"key":"2024042419444642400_bib208","doi-asserted-by":"publisher","first-page":"820","DOI":"10.18653\/v1\/2021.findings-emnlp.71","article-title":"Probing across time: What does RoBERTa know and when?","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2021","author":"Liu","year":"2021"},{"key":"2024042419444642400_bib209","article-title":"Intersectional bias in causal language models","author":"Magee","year":"2021","journal-title":"ArXiv"},{"key":"2024042419444642400_bib210","doi-asserted-by":"publisher","first-page":"265","DOI":"10.18653\/v1\/2023.eacl-main.20","article-title":"A discerning several thousand judgments: GPT-3 rates the article + adjective + numeral + noun construction","volume-title":"Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics","author":"Mahowald","year":"2023"},{"key":"2024042419444642400_bib211","article-title":"Experimentally measuring the redundancy of grammatical cues in transitive clauses","author":"Mahowald","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib212","doi-asserted-by":"publisher","first-page":"10351","DOI":"10.18653\/v1\/2021.emnlp-main.809","article-title":"Studying word order through iterative shuffling","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Malkin","year":"2021"},{"key":"2024042419444642400_bib213","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.546","article-title":"When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories","author":"Mallen","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib214","doi-asserted-by":"publisher","DOI":"10.7551\/mitpress\/9780262514620.001.0001","volume-title":"Vision: A Computational Investigation into the Human Representation and Processing of Visual Information","author":"Marr","year":"2010"},{"key":"2024042419444642400_bib215","doi-asserted-by":"publisher","first-page":"95","DOI":"10.18653\/v1\/2021.blackboxnlp-1.8","article-title":"ProSPer: Probing human and neural network language model understanding of spatial perspective","volume-title":"Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"Masis","year":"2021"},{"key":"2024042419444642400_bib216","doi-asserted-by":"publisher","first-page":"223","DOI":"10.18653\/v1\/2020.findings-emnlp.22","article-title":"How decoding strategies affect the verifiability of generated text","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Massarelli","year":"2020"},{"key":"2024042419444642400_bib217","article-title":"Understanding stereotypes in language models: Towards robust measurement and zero-shot debiasing","author":"Mattern","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib218","article-title":"How much do language models copy from their training data? Evaluating linguistic novelty in text generation using RAVEN","author":"McCoy","year":"2021","journal-title":"ArXiv"},{"key":"2024042419444642400_bib219","first-page":"2096","article-title":"Revisiting the poverty of the stimulus: Hierarchical generalization without a hierarchical bias in recurrent neural networks","volume-title":"Proceedings of the Annual Meeting of the Cognitive Science Society","author":"McCoy","year":"2018"},{"key":"2024042419444642400_bib220","doi-asserted-by":"publisher","first-page":"3428","DOI":"10.18653\/v1\/P19-1334","article-title":"Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"McCoy","year":"2019"},{"key":"2024042419444642400_bib221","doi-asserted-by":"publisher","first-page":"1878","DOI":"10.18653\/v1\/2022.acl-long.132","article-title":"An empirical survey of the effectiveness of debiasing techniques for pre-trained language models","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Meade","year":"2022"},{"key":"2024042419444642400_bib222","doi-asserted-by":"publisher","first-page":"2831","DOI":"10.18653\/v1\/2022.naacl-main.204","article-title":"Robust conversational agents against imperceptible toxicity triggers","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Mehrabi","year":"2022"},{"key":"2024042419444642400_bib223","doi-asserted-by":"publisher","first-page":"5328","DOI":"10.18653\/v1\/2021.acl-long.414","article-title":"Language model evaluation beyond perplexity","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Meister","year":"2021"},{"key":"2024042419444642400_bib224","first-page":"17359","article-title":"Locating and editing factual associations in GPT","volume-title":"Advances in Neural Information Processing Systems","author":"Meng","year":"2022"},{"key":"2024042419444642400_bib225","doi-asserted-by":"publisher","first-page":"745","DOI":"10.18653\/v1\/2020.coling-main.65","article-title":"Linguistic profiling of a neural language model","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"Miaschi","year":"2020"},{"key":"2024042419444642400_bib226","doi-asserted-by":"publisher","first-page":"13","DOI":"10.18653\/v1\/2022.conll-1.2","article-title":"Collateral facilitation in humans and language models","volume-title":"Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL)","author":"Michaelov","year":"2022"},{"key":"2024042419444642400_bib227","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.891","article-title":"\u2018Rarely\u2019 a problem? Language models exhibit inverse scaling in their predictions following \u2018few\u2019-type quantifiers","author":"Michaelov","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib228","doi-asserted-by":"publisher","first-page":"11048","DOI":"10.18653\/v1\/2022.emnlp-main.759","article-title":"Rethinking the role of demonstrations: What makes in-context learning work?","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Min","year":"2022"},{"key":"2024042419444642400_bib229","doi-asserted-by":"publisher","first-page":"218","DOI":"10.18653\/v1\/2022.nlpcss-1.24","article-title":"Who is GPT-3? An exploration of personality, values and demographics","volume-title":"Proceedings of the Fifth Workshop on Natural Language Processing and Computational Social Science (NLP+CSS)","author":"Miotto","year":"2022"},{"key":"2024042419444642400_bib230","article-title":"minicons: Enabling flexible behavioral and representational analyses of Transformer language models","author":"Misra","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib231","doi-asserted-by":"publisher","first-page":"4625","DOI":"10.18653\/v1\/2020.findings-emnlp.415","article-title":"Exploring BERT\u2019s sensitivity to lexical cues using tests from semantic priming","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Misra","year":"2020"},{"key":"2024042419444642400_bib232","first-page":"216","article-title":"Do language models learn typicality judgments from text?","volume-title":"Proceedings of the Annual Meeting of the Cognitive Science Society","author":"Misra","year":"2021"},{"key":"2024042419444642400_bib233","doi-asserted-by":"publisher","first-page":"2928","DOI":"10.18653\/v1\/2023.eacl-main.213","article-title":"COMPS: Conceptual minimal pair sentences for testing property knowledge and inheritance in pre-trained language models","volume-title":"Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics","author":"Misra","year":"2023"},{"key":"2024042419444642400_bib234","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.2215907120","article-title":"The debate over understanding in AI\u2019s large language models","author":"Mitchell","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib235","article-title":"Learning in the rational speech acts model","author":"Monroe","year":"2015","journal-title":"ArXiv"},{"key":"2024042419444642400_bib236","doi-asserted-by":"publisher","first-page":"2502","DOI":"10.18653\/v1\/2020.findings-emnlp.227","article-title":"On the interplay between fine-tuning and sentence-level probing for linguistic knowledge in pre-trained Transformers","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Mosbach","year":"2020"},{"key":"2024042419444642400_bib237","doi-asserted-by":"publisher","first-page":"5356","DOI":"10.18653\/v1\/2021.acl-long.416","article-title":"StereoSet: Measuring stereotypical bias in pretrained language models","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Nadeem","year":"2021"},{"key":"2024042419444642400_bib238","doi-asserted-by":"publisher","first-page":"1953","DOI":"10.18653\/v1\/2020.emnlp-main.154","article-title":"CrowS-Pairs: A challenge dataset for measuring social biases in masked language models","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Nangia","year":"2020"},{"key":"2024042419444642400_bib239","article-title":"Understanding searches better than ever before","author":"Nayak","year":"2019","journal-title":"The Keyword: Google Blog"},{"key":"2024042419444642400_bib240","doi-asserted-by":"publisher","first-page":"3710","DOI":"10.18653\/v1\/2021.naacl-main.290","article-title":"Refining targeted syntactic evaluation of language models","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Newman","year":"2021"},{"key":"2024042419444642400_bib241","doi-asserted-by":"publisher","first-page":"2398","DOI":"10.18653\/v1\/2021.naacl-main.191","article-title":"HONEST: Measuring hurtful sentence completion in language models","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Nozza","year":"2021"},{"key":"2024042419444642400_bib242","doi-asserted-by":"publisher","first-page":"26","DOI":"10.18653\/v1\/2022.ltedi-1.4","article-title":"Measuring harmful sentence completion in language models for LGBTQIA+ individuals","volume-title":"Proceedings of the Second Workshop on Language Technology for Equality, Diversity and Inclusion","author":"Nozza","year":"2022"},{"key":"2024042419444642400_bib243","doi-asserted-by":"publisher","first-page":"851","DOI":"10.18653\/v1\/2021.acl-long.70","article-title":"What context features can transformer language models use?","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"O\u2019Connor","year":"2021"},{"key":"2024042419444642400_bib244","article-title":"In-context learning and induction heads","author":"Olsson","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib245","article-title":"ChatGPT: Optimizing language models for dialogue","author":"OpenAI","year":"2022","journal-title":"OpenAI Blog"},{"key":"2024042419444642400_bib246","article-title":"GPT-4 technical report","author":"OpenAI","year":"2023","journal-title":"OpenAI"},{"key":"2024042419444642400_bib247","article-title":"Model index for researchers","author":"OpenAI","year":"2023","journal-title":"OpenAI"},{"key":"2024042419444642400_bib248","doi-asserted-by":"publisher","first-page":"4262","DOI":"10.18653\/v1\/2021.acl-long.329","article-title":"Probing toxic content in large pre-trained language models","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Ousidhoum","year":"2021"},{"key":"2024042419444642400_bib249","first-page":"27730","article-title":"Training language models to follow instructions with human feedback","volume-title":"Advances in Neural Information Processing Systems","author":"Ouyang","year":"2022"},{"key":"2024042419444642400_bib250","doi-asserted-by":"publisher","first-page":"823","DOI":"10.18653\/v1\/2021.emnlp-main.63","article-title":"The world of an octopus: How reporting bias influences a language model\u2019s perception of color","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Paik","year":"2021"},{"key":"2024042419444642400_bib251","doi-asserted-by":"publisher","first-page":"367","DOI":"10.18653\/v1\/2021.conll-1.29","article-title":"Pragmatic competence of pre-trained language models through the lens of discourse connectives","volume-title":"Proceedings of the 25th Conference on Computational Natural Language Learning","author":"Pandia","year":"2021"},{"key":"2024042419444642400_bib252","doi-asserted-by":"publisher","first-page":"1583","DOI":"10.18653\/v1\/2021.emnlp-main.119","article-title":"Sorting through the noise: Testing robustness of information processing in pre-trained language models","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Pandia","year":"2021"},{"key":"2024042419444642400_bib253","doi-asserted-by":"publisher","first-page":"4153","DOI":"10.18653\/v1\/2021.naacl-main.327","article-title":"Probing for bridging inference in transformer language models","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Pandit","year":"2021"},{"issue":"2","key":"2024042419444642400_bib254","first-page":"395","article-title":"Deep learning can contrast the minimal pairs of syntactic data","volume":"38","author":"Park","year":"2021","journal-title":"Linguistic Research"},{"key":"2024042419444642400_bib255","doi-asserted-by":"publisher","first-page":"10080","DOI":"10.18653\/v1\/2021.emnlp-main.790","article-title":"\u201cwas it \u201cstated\u201d or was it \u201cclaimed\u201d?: How linguistic bias affects generative language models","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Patel","year":"2021"},{"key":"2024042419444642400_bib256","doi-asserted-by":"publisher","first-page":"192","DOI":"10.18653\/v1\/2021.blackboxnlp-1.13","article-title":"A howling success or a working sea? Testing what BERT knows about metaphors","volume-title":"Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"Pedinotti","year":"2021"},{"key":"2024042419444642400_bib257","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18653\/v1\/2021.starsem-1.1","article-title":"Did the cat drink the coffee? Challenging transformers with generalized event knowledge","volume-title":"Proceedings of *SEM 2021: The Tenth Joint Conference on Lexical and Computational Semantics","author":"Pedinotti","year":"2021"},{"key":"2024042419444642400_bib258","doi-asserted-by":"publisher","first-page":"5015","DOI":"10.18653\/v1\/2022.emnlp-main.335","article-title":"COPEN: Probing conceptual knowledge in pre-trained language models","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Peng","year":"2022"},{"key":"2024042419444642400_bib259","doi-asserted-by":"publisher","first-page":"388","DOI":"10.1145\/3383313.3412249","article-title":"What does BERT know about books, movies and music? Probing BERT for conversational recommendation","volume-title":"Proceedings of the 14th ACM Conference on Recommender Systems","author":"Penha","year":"2020"},{"key":"2024042419444642400_bib260","doi-asserted-by":"publisher","first-page":"3419","DOI":"10.18653\/v1\/2022.emnlp-main.225","article-title":"Red teaming language models with language models","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Perez","year":"2022"},{"key":"2024042419444642400_bib261","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.847","article-title":"Discovering language model behaviors with model-written evaluations","author":"Perez","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib262","doi-asserted-by":"publisher","first-page":"1571","DOI":"10.18653\/v1\/2021.emnlp-main.118","article-title":"How much pretraining data do language models need to learn syntax?","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"P\u00e9rez-Mayos","year":"2021"},{"key":"2024042419444642400_bib263","doi-asserted-by":"publisher","first-page":"2463","DOI":"10.18653\/v1\/D19-1250","article-title":"Language models as knowledge bases?","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Petroni","year":"2019"},{"key":"2024042419444642400_bib264","article-title":"Transformers generalize linearly","author":"Petty","year":"2021","journal-title":"ArXiv"},{"issue":"2","key":"2024042419444642400_bib265","doi-asserted-by":"publisher","first-page":"141","DOI":"10.1093\/jole\/lzw013","article-title":"Infinitely productive language can arise from chance under communicative pressure","volume":"2","author":"Piantadosi","year":"2017","journal-title":"Journal of Language Evolution"},{"key":"2024042419444642400_bib266","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/IJCNN52387.2021.9534299","article-title":"How can the [MASK] know? The sources and limitations of knowledge in BERT","volume-title":"IEEE International Joint Conference on Neural Networks","author":"Podkorytov","year":"2021"},{"key":"2024042419444642400_bib267","article-title":"BERT is not a knowledge base (yet): Factual knowledge vs. name-based reasoning in unsupervised QA","author":"Poerner","year":"2019","journal-title":"ArXiv"},{"key":"2024042419444642400_bib268","doi-asserted-by":"publisher","first-page":"4550","DOI":"10.18653\/v1\/2022.naacl-main.337","article-title":"Does pre-training induce systematic inference? How masked language models acquire commonsense knowledge","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Porada","year":"2022"},{"key":"2024042419444642400_bib269","first-page":"663","article-title":"Poverty of the stimulus? A rational approach","volume-title":"Proceedings of the Annual Meeting of the Cognitive Science Society","author":"Prefors","year":"2006"},{"key":"2024042419444642400_bib270","article-title":"Train short, test long: Attention with linear biases enables input length extrapolation","volume-title":"International Conference on Learning Representations","author":"Press","year":"2022"},{"key":"2024042419444642400_bib271","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.378","article-title":"Measuring and narrowing the compositionality gap in language models","author":"Press","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib272","doi-asserted-by":"publisher","first-page":"7066","DOI":"10.18653\/v1\/2021.acl-long.549","article-title":"TIMEDIAL: Temporal commonsense reasoning in dialog","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Qin","year":"2021"},{"key":"2024042419444642400_bib273","doi-asserted-by":"publisher","first-page":"9157","DOI":"10.18653\/v1\/2022.emnlp-main.624","article-title":"Evaluating the impact of model scale for compositional generalization in semantic parsing","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Qiu","year":"2022"},{"key":"2024042419444642400_bib274","article-title":"Robust speech recognition via large-scale weak supervision","author":"Radford","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib275","article-title":"Improving language understanding by generative pre-training","author":"Radford","year":"2018","journal-title":"OpenAI"},{"key":"2024042419444642400_bib276","article-title":"Language models are unsupervised multitask learners","author":"Radford","year":"2019","journal-title":"OpenAI"},{"key":"2024042419444642400_bib277","article-title":"Scaling language models: Methods, analysis & insights from training gopher","author":"Rae","year":"2021","journal-title":"ArXiv"},{"issue":"140","key":"2024042419444642400_bib278","first-page":"5485","article-title":"Exploring the limits of transfer learning with a unified text-to-text Transformer","volume":"21","author":"Raffel","year":"2020","journal-title":"Journal of Machine Learning Research"},{"key":"2024042419444642400_bib279","article-title":"Measuring reliability of large language models through semantic consistency","volume-title":"NeurIPS ML Safety Workshop","author":"Raj","year":"2022"},{"key":"2024042419444642400_bib280","first-page":"88","article-title":"On the systematicity of probing contextualized word representations: The case of hypernymy in BERT","volume-title":"Proceedings of the Ninth Joint Conference on Lexical and Computational Semantics","author":"Ravichander","year":"2020"},{"key":"2024042419444642400_bib281","doi-asserted-by":"publisher","first-page":"840","DOI":"10.18653\/v1\/2022.findings-emnlp.59","article-title":"Impact of pretraining term frequencies on few-shot numerical reasoning","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Razeghi","year":"2022"},{"key":"2024042419444642400_bib282","doi-asserted-by":"publisher","first-page":"837","DOI":"10.18653\/v1\/2022.acl-short.94","article-title":"A recipe for arbitrary text style transfer with large language models","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Reif","year":"2022"},{"key":"2024042419444642400_bib283","first-page":"8594","article-title":"Visualizing and measuring the geometry of BERT","volume-title":"Advances in Neural Information Processing Systems","author":"Reif","year":"2019"},{"key":"2024042419444642400_bib284","doi-asserted-by":"publisher","first-page":"842","DOI":"10.1162\/tacl_a_00349","article-title":"A primer in BERTology: What we know about how BERT works","volume":"8","author":"Rogers","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024042419444642400_bib285","doi-asserted-by":"publisher","first-page":"10954","DOI":"10.18653\/v1\/2022.emnlp-main.752","article-title":"Do children texts hold the key to commonsense knowledge?","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Romero","year":"2022"},{"key":"2024042419444642400_bib286","article-title":"Large language models are not zero-shot communicators","author":"Ruis","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib287","doi-asserted-by":"publisher","first-page":"61","DOI":"10.18653\/v1\/2021.cmcl-1.6","article-title":"Accounting for agreement phenomena in sentence comprehension with transformer language models: Effects of similarity-based interference on surprisal and attention","volume-title":"Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics","author":"Ryu","year":"2021"},{"key":"2024042419444642400_bib288","article-title":"Unpacking large language models with conceptual consistency","author":"Sahu","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib289","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18653\/v1\/2022.starsem-1.1","article-title":"What do large language models learn about scripts?","volume-title":"Proceedings of the 11th Joint Conference on Lexical and Computational Semantics","author":"Sancheti","year":"2022"},{"key":"2024042419444642400_bib290","article-title":"DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter","volume-title":"Workshop on Energy Efficient Machine Learning and Cognitive Computing","author":"Sanh","year":"2019"},{"key":"2024042419444642400_bib291","doi-asserted-by":"publisher","first-page":"3762","DOI":"10.18653\/v1\/2022.emnlp-main.248","article-title":"Neural theory-of-mind? On the limits of social intelligence in large LMs","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Sap","year":"2022"},{"key":"2024042419444642400_bib292","article-title":"Language models are greedy reasoners: A systematic formal analysis of chain-of-thought","volume-title":"International Conference on Learning Representations","author":"Saparov","year":"2023"},{"key":"2024042419444642400_bib293","article-title":"Toolformer: Language models can teach themselves to use tools","author":"Schick","year":"2023","journal-title":"ArXiv"},{"key":"2024042419444642400_bib294","doi-asserted-by":"publisher","first-page":"969","DOI":"10.18653\/v1\/2022.naacl-main.71","article-title":"When a sentence does not introduce a discourse entity, transformer-based models still sometimes refer to it","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Schuster","year":"2022"},{"key":"2024042419444642400_bib295","doi-asserted-by":"publisher","first-page":"532","DOI":"10.18653\/v1\/2021.eacl-main.42","article-title":"Does she wink or does she nod? A challenging benchmark for evaluating word understanding of language models","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Senel","year":"2021"},{"key":"2024042419444642400_bib296","doi-asserted-by":"publisher","first-page":"1715","DOI":"10.18653\/v1\/P16-1162","article-title":"Neural machine translation of rare words with subword units","volume-title":"Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Sennrich","year":"2016"},{"key":"2024042419444642400_bib297","doi-asserted-by":"publisher","first-page":"2931","DOI":"10.18653\/v1\/P19-1282","article-title":"Is attention interpretable?","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Serrano","year":"2019"},{"key":"2024042419444642400_bib298","article-title":"Quantifying social biases using templates is unreliable","volume-title":"Workshop on Trustworthy and Socially Responsible Machine Learning","author":"Seshadri","year":"2022"},{"key":"2024042419444642400_bib299","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.244","article-title":"On second thought, let\u2019s not think step by step! Bias and toxicity in zero-shot reasoning","author":"Shaikh","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib300","article-title":"Deanthropomorphising NLP: Can a language model be conscious?","author":"Shardlow","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib301","doi-asserted-by":"publisher","first-page":"464","DOI":"10.18653\/v1\/N18-2074","article-title":"Self-attention with relative position representations","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers)","author":"Shaw","year":"2018"},{"key":"2024042419444642400_bib302","doi-asserted-by":"publisher","first-page":"3407","DOI":"10.18653\/v1\/D19-1339","article-title":"The woman worked as a babysitter: On biases in language generation","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Sheng","year":"2019"},{"key":"2024042419444642400_bib303","doi-asserted-by":"publisher","first-page":"750","DOI":"10.18653\/v1\/2021.naacl-main.60","article-title":"\u201cnice try, kiddo\u201d: Investigating ad hominems in dialogue responses","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Sheng","year":"2021"},{"key":"2024042419444642400_bib304","doi-asserted-by":"publisher","first-page":"4275","DOI":"10.18653\/v1\/2021.acl-long.330","article-title":"Societal biases in language generation: Progress and challenges","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Sheng","year":"2021"},{"key":"2024042419444642400_bib305","article-title":"Large language models can be easily distracted by irrelevant context","author":"Shi","year":"2023","journal-title":"ArXiv"},{"key":"2024042419444642400_bib306","first-page":"2218","article-title":"What Transformers might know about the physical world: T5 and the origins of knowledge","volume-title":"Proceedings of the Annual Meeting of the Cognitive Science Society","author":"Shi","year":"2021"},{"key":"2024042419444642400_bib307","doi-asserted-by":"publisher","first-page":"6863","DOI":"10.18653\/v1\/2020.coling-main.605","article-title":"Do neural language models overcome reporting bias?","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"Shwartz","year":"2020"},{"key":"2024042419444642400_bib308","doi-asserted-by":"publisher","first-page":"6850","DOI":"10.18653\/v1\/2020.emnlp-main.556","article-title":"\u201cyou are grounded!\u201d: Latent name artifacts in pre-trained language models","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Shwartz","year":"2020"},{"issue":"3","key":"2024042419444642400_bib309","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1371\/journal.pone.0248388","article-title":"Reevaluating pragmatic reasoning in language games","volume":"16","author":"Sikos","year":"2021","journal-title":"PLOS One"},{"key":"2024042419444642400_bib310","doi-asserted-by":"publisher","first-page":"2383","DOI":"10.18653\/v1\/2021.naacl-main.189","article-title":"Towards a comprehensive understanding and accurate evaluation of societal biases in pre-trained transformers","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Silva","year":"2021"},{"key":"2024042419444642400_bib311","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-srw.40","article-title":"Moral mimicry: Large language models produce moral rationalizations tailored to political identity","author":"Simmons","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib312","doi-asserted-by":"publisher","first-page":"1031","DOI":"10.1162\/tacl_a_00504","article-title":"Structural persistence in language models: Priming as a window into abstract language representations","volume":"10","author":"Sinclair","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024042419444642400_bib313","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.333","article-title":"Language model acceptability judgements are not always robust to context","author":"Sinha","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib314","doi-asserted-by":"publisher","first-page":"2888","DOI":"10.18653\/v1\/2021.emnlp-main.230","article-title":"Masked language modeling and the distributional hypothesis: Order word matters pre-training for little","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Sinha","year":"2021"},{"key":"2024042419444642400_bib315","doi-asserted-by":"publisher","first-page":"4449","DOI":"10.18653\/v1\/2022.findings-emnlp.326","article-title":"The curious case of absolute position embeddings","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2022","author":"Sinha","year":"2022"},{"key":"2024042419444642400_bib316","doi-asserted-by":"publisher","first-page":"9180","DOI":"10.18653\/v1\/2022.emnlp-main.625","article-title":"\u201cI\u2019m sorry to hear that\u201d: Finding new biases in language models with a holistic descriptor dataset","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Smith","year":"2022"},{"key":"2024042419444642400_bib317","article-title":"Using DeepSpeed and Megatron to train Megatron-Turing NLG 530B, a large-scale generative language model","author":"Smith","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib318","doi-asserted-by":"publisher","DOI":"10.1126\/sciadv.adh1850","article-title":"AI model GPT-3 (dis)informs us better than humans","author":"Spitale","year":"2023","journal-title":"ArXiv"},{"key":"2024042419444642400_bib319","article-title":"Beyond the imitation game: Quantifying and extrapolating the capabilities of language models","author":"Srivastava","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib320","doi-asserted-by":"publisher","first-page":"47","DOI":"10.18653\/v1\/2022.wnu-1.6","article-title":"Heroes, villains, and victims, and GPT-3: Automated extraction of character roles without training data","volume-title":"Proceedings of the 4th Workshop of Narrative Understanding (WNU2022)","author":"Stammbach","year":"2022"},{"key":"2024042419444642400_bib321","first-page":"164","article-title":"Putting GPT-3\u2019s creativity to the (alternative uses) test","volume-title":"International Conference on Computational Creativity","author":"Stevenson","year":"2022"},{"key":"2024042419444642400_bib322","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.32","article-title":"A causal framework to quantify the robustness of mathematical reasoning with language models","author":"Stolfo","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib323","doi-asserted-by":"publisher","first-page":"3645","DOI":"10.18653\/v1\/P19-1355","article-title":"Energy and policy considerations for deep learning in NLP","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Strubell","year":"2019"},{"key":"2024042419444642400_bib324","article-title":"RoFormer: Enhanced Transformer with rotary position embedding","author":"Su","year":"2021","journal-title":"ArXiv"},{"key":"2024042419444642400_bib325","doi-asserted-by":"publisher","first-page":"73","DOI":"10.18653\/v1\/2021.mrqa-1.7","article-title":"What can a generative language model answer about a passage?","volume-title":"Proceedings of the 3rd Workshop on Machine Reading for Question Answering","author":"Summers-Stay","year":"2021"},{"key":"2024042419444642400_bib326","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.824","article-title":"Challenging BIG-Bench tasks and whether chain-of-thought can solve them","author":"Suzgun","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib327","article-title":"Interpreting language models through knowledge graph extraction","volume-title":"Workshop on eXplainable AI Approaches for Debugging and Diagnosis","author":"Swamy","year":"2021"},{"key":"2024042419444642400_bib328","doi-asserted-by":"publisher","first-page":"112","DOI":"10.18653\/v1\/2022.gebnlp-1.13","article-title":"Fewer errors, but more stereotypes? The effect of model size on gender bias","volume-title":"Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP)","author":"Tal","year":"2022"},{"key":"2024042419444642400_bib329","doi-asserted-by":"publisher","first-page":"3878","DOI":"10.18653\/v1\/2020.acl-main.357","article-title":"Pre-training is (almost) all you need: An application to commonsense reasoning","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Tamborrino","year":"2020"},{"key":"2024042419444642400_bib330","article-title":"Gender biases unexpectedly fluctuate in the pre-training stage of masked language models","author":"Tang","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib331","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.825","article-title":"Scaling laws vs model architectures: How does inductive bias influence scaling?","author":"Tay","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib332","article-title":"Scale efficiently: Insights from pre-training and fine-tuning Transformers","volume-title":"International Conference on Learning Representations","author":"Tay","year":"2022"},{"key":"2024042419444642400_bib333","first-page":"47","article-title":"A study of BERT\u2019s processing of negations to determine sentiment","volume-title":"Benelux Conference on Artificial Intelligence and the Belgian Dutch Conference on Machine Learning","author":"Tejada","year":"2021"},{"key":"2024042419444642400_bib334","doi-asserted-by":"publisher","first-page":"4593","DOI":"10.18653\/v1\/P19-1452","article-title":"BERT rediscovers the classical NLP pipeline","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Tenney","year":"2019"},{"key":"2024042419444642400_bib335","article-title":"Bring structure to diverse documents with Amazon Textract and transformer-based models on Amazon SageMaker","author":"Thewsey","year":"2021","journal-title":"AWS Machine Learning Blog"},{"key":"2024042419444642400_bib336","first-page":"38274","article-title":"Memorization without overfitting: Analyzing the training dynamics of large language models","volume-title":"Advances in Neural Information Processing Systems","author":"Tirumala","year":"2022"},{"key":"2024042419444642400_bib337","first-page":"423","article-title":"Exploring the effects of negation and grammatical tense on bias probes","volume-title":"Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)","author":"Touileb","year":"2022"},{"key":"2024042419444642400_bib338","doi-asserted-by":"publisher","first-page":"158","DOI":"10.18653\/v1\/2021.acl-short.21","article-title":"AND does not mean OR: Using formal languages to study language models\u2019 representations","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)","author":"Traylor","year":"2021"},{"key":"2024042419444642400_bib339","article-title":"In cautious defense of LLM-ology","author":"Trott","year":"2023","journal-title":"Blog Post"},{"issue":"7","key":"2024042419444642400_bib340","doi-asserted-by":"publisher","first-page":"e13309","DOI":"10.1111\/cogs.13309","article-title":"Do large language models know what humans know?","volume":"47","author":"Trott","year":"2023","journal-title":"Cognitive Science"},{"key":"2024042419444642400_bib341","first-page":"883","article-title":"Not another negation benchmark: The NaN-NLI test suite for sub-clausal negation","volume-title":"Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Truong","year":"2022"},{"key":"2024042419444642400_bib342","doi-asserted-by":"publisher","first-page":"99","DOI":"10.18653\/v1\/2022.naacl-demo.11","article-title":"SentSpace: Large-scale benchmarking and evaluation of text using cognitively motivated lexical, syntactic, and semantic features","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: System Demonstrations","author":"Tuckute","year":"2022"},{"key":"2024042419444642400_bib343","doi-asserted-by":"publisher","first-page":"977","DOI":"10.18653\/v1\/2020.emnlp-main.70","article-title":"Predicting reference: What do language models learn about discourse models?","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Upadhye","year":"2020"},{"key":"2024042419444642400_bib344","doi-asserted-by":"publisher","first-page":"3609","DOI":"10.18653\/v1\/2021.acl-long.280","article-title":"BERT is to NLP what AlexNet is to CV: Can pre-trained language models identify analogies?","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Ushio","year":"2021"},{"key":"2024042419444642400_bib345","article-title":"Large language models still can\u2019t plan (a benchmark for LLMs on planning and reasoning about change)","volume-title":"Foundation Models for Decision Making Workshop","author":"Valmeekam","year":"2022"},{"key":"2024042419444642400_bib346","doi-asserted-by":"publisher","first-page":"5831","DOI":"10.18653\/v1\/D19-1592","article-title":"Quantity doesn\u2019t buy quality syntax with neural language models","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"van Schijndel","year":"2019"},{"key":"2024042419444642400_bib347","first-page":"5998","article-title":"Attention is all you need","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani","year":"2017"},{"key":"2024042419444642400_bib348","doi-asserted-by":"publisher","first-page":"63","DOI":"10.18653\/v1\/W19-4808","article-title":"Analyzing the structure of attention in a transformer language model","volume-title":"Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP","author":"Vig","year":"2019"},{"key":"2024042419444642400_bib349","first-page":"12388","article-title":"Investigating gender bias in language models using causal mediation analysis","volume-title":"Advances in Neural Information Processing Systems","author":"Vig","year":"2020"},{"key":"2024042419444642400_bib350","doi-asserted-by":"publisher","first-page":"952","DOI":"10.18653\/v1\/2022.emnlp-main.62","article-title":"How large language models are transforming machine-paraphrase plagiarism","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Wahle","year":"2022"},{"key":"2024042419444642400_bib351","doi-asserted-by":"publisher","first-page":"2153","DOI":"10.18653\/v1\/D19-1221","article-title":"Universal adversarial triggers for attacking and analyzing NLP","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Wallace","year":"2019"},{"key":"2024042419444642400_bib352","doi-asserted-by":"publisher","first-page":"5307","DOI":"10.18653\/v1\/D19-1534","article-title":"Do NLP models know numbers? Probing numeracy in embeddings","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Wallace","year":"2019"},{"key":"2024042419444642400_bib353","first-page":"3266","article-title":"SuperGLUE: A stickier benchmark for general-purpose language understanding systems","volume-title":"Advances in Neural Information Processing Systems","author":"Wang","year":"2019"},{"key":"2024042419444642400_bib354","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.153","article-title":"Towards understanding chain-of-thought prompting: An empirical study of what matters","author":"Wang","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib355","article-title":"On position embeddings in BERT","volume-title":"International Conference on Learning Representations","author":"Wang","year":"2021"},{"key":"2024042419444642400_bib356","doi-asserted-by":"publisher","first-page":"758","DOI":"10.1007\/978-3-030-88480-2_61","article-title":"Exploring generalization ability of pretrained language models on arithmetic and logical reasoning","volume-title":"Natural Language Processing and Chinese Computing","author":"Wang","year":"2021"},{"key":"2024042419444642400_bib357","doi-asserted-by":"publisher","first-page":"1719","DOI":"10.18653\/v1\/2022.findings-naacl.130","article-title":"Identifying and mitigating spurious correlations for improving robustness in NLP models","volume-title":"Findings of the Association for Computational Linguistics: NAACL 2022","author":"Wang","year":"2022"},{"key":"2024042419444642400_bib358","doi-asserted-by":"publisher","first-page":"2877","DOI":"10.18653\/v1\/D19-1286","article-title":"Investigating BERT\u2019s knowledge of language: Five analysis methods with NPIs","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Warstadt","year":"2019"},{"key":"2024042419444642400_bib359","doi-asserted-by":"publisher","first-page":"377","DOI":"10.1162\/tacl_a_00321","article-title":"BLiMP: The benchmark of linguistic minimal pairs for English","volume":"8","author":"Warstadt","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024042419444642400_bib360","doi-asserted-by":"publisher","DOI":"10.1038\/s41562-023-01659-w","article-title":"Emergent analogical reasoning in large language models","author":"Webb","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib361","article-title":"Finetuned language models are zero-shot learners","volume-title":"International Conference on Learning Representations","author":"Wei","year":"2022"},{"key":"2024042419444642400_bib362","doi-asserted-by":"publisher","first-page":"932","DOI":"10.18653\/v1\/2021.emnlp-main.72","article-title":"Frequency effects on syntactic rule learning in transformers","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Wei","year":"2021"},{"key":"2024042419444642400_bib363","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.963","article-title":"Inverse scaling can become U-shaped","author":"Wei","year":"2023","journal-title":"ArXiv"},{"key":"2024042419444642400_bib364","article-title":"Emergent abilities of large language models","author":"Wei","year":"2022","journal-title":"Transactions on Machine Learning Research"},{"key":"2024042419444642400_bib365","first-page":"24824","article-title":"Chain of thought prompting elicits reasoning in large language models","volume-title":"Advances in Neural Information Processing Systems","author":"Wei","year":"2022"},{"key":"2024042419444642400_bib366","article-title":"Ethical and social risks of harm from language models","author":"Weidinger","year":"2021","journal-title":"ArXiv"},{"key":"2024042419444642400_bib367","doi-asserted-by":"publisher","first-page":"214","DOI":"10.1145\/3531146.3533088","article-title":"Taxonomy of risks posed by language models","volume-title":"Proceedings of the ACM Conference on Fairness, Accountability, and Transparency","author":"Weidinger","year":"2022"},{"key":"2024042419444642400_bib368","first-page":"377","article-title":"Probing neural language models for human tacit assumptions","volume-title":"Annual Meeting of the Cognitive Science Society","author":"Weir","year":"2020"},{"key":"2024042419444642400_bib369","doi-asserted-by":"publisher","first-page":"10859","DOI":"10.18653\/v1\/2022.emnlp-main.746","article-title":"The better your syntax, the better your semantics? Probing pretrained language models for the English comparative correlative","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Weissweiler","year":"2022"},{"key":"2024042419444642400_bib370","doi-asserted-by":"publisher","first-page":"2447","DOI":"10.18653\/v1\/2021.findings-emnlp.210","article-title":"Challenges in detoxifying language models","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2021","author":"Welbl","year":"2021"},{"key":"2024042419444642400_bib371","doi-asserted-by":"publisher","first-page":"2985","DOI":"10.18653\/v1\/2023.eacl-main.217","article-title":"Should you mask 15% in masked language modeling?","volume-title":"Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics","author":"Wettig","year":"2023"},{"key":"2024042419444642400_bib372","doi-asserted-by":"publisher","first-page":"454","DOI":"10.18653\/v1\/2021.acl-long.38","article-title":"Examining the inductive bias of neural language models with artificial languages","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"White","year":"2021"},{"key":"2024042419444642400_bib373","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1162\/ling_a_00491","article-title":"Using computational models to test syntactic learnability","author":"Wilcox","year":"2022","journal-title":"Linguistic Inquiry"},{"key":"2024042419444642400_bib374","doi-asserted-by":"publisher","first-page":"1112","DOI":"10.18653\/v1\/N18-1101","article-title":"A broad-coverage challenge corpus for sentence understanding through inference","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)","author":"Williams","year":"2018"},{"key":"2024042419444642400_bib375","doi-asserted-by":"publisher","first-page":"120","DOI":"10.18653\/v1\/2020.repl4nlp-1.16","article-title":"Are all languages created equal in multilingual BERT?","volume-title":"Proceedings of the 5th Workshop on Representation Learning for NLP","author":"Wu","year":"2020"},{"key":"2024042419444642400_bib376","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.767","article-title":"Training trajectories of language models across scales","author":"Xia","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib377","doi-asserted-by":"publisher","first-page":"4040","DOI":"10.18653\/v1\/2020.emnlp-main.331","article-title":"Word frequency does not predict grammatical knowledge in language models","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Yu","year":"2020"},{"key":"2024042419444642400_bib378","doi-asserted-by":"publisher","first-page":"36","DOI":"10.1109\/EMC2-NIPS53020.2019.00016","article-title":"Q8BERT: Quantized 8bit BERT","volume-title":"Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing","author":"Zafrir","year":"2019"},{"key":"2024042419444642400_bib379","doi-asserted-by":"publisher","first-page":"4856","DOI":"10.18653\/v1\/2021.naacl-main.386","article-title":"TuringAdvice: A generative and dynamic evaluation of language use","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Zellers","year":"2021"},{"key":"2024042419444642400_bib380","doi-asserted-by":"publisher","DOI":"10.1145\/3617680","article-title":"A survey of controllable text generation using transformer-based pre-trained language models","author":"Zhang","year":"2023","journal-title":"ArXiv"},{"key":"2024042419444642400_bib381","doi-asserted-by":"publisher","first-page":"297","DOI":"10.18653\/v1\/2022.blackboxnlp-1.24","article-title":"Probing GPT-3\u2019s linguistic knowledge on semantic tasks","volume-title":"Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP","author":"Zhang","year":"2022"},{"key":"2024042419444642400_bib382","doi-asserted-by":"publisher","first-page":"415","DOI":"10.18653\/v1\/2023.findings-eacl.31","article-title":"Causal reasoning of entities and events in procedural texts","volume-title":"Findings of the Association for Computational Linguistics: EACL 2023","author":"Zhang","year":"2023"},{"key":"2024042419444642400_bib383","article-title":"OPT: Open pre-trained Transformer language models","author":"Zhang","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib384","doi-asserted-by":"publisher","first-page":"4581","DOI":"10.18653\/v1\/2021.emnlp-main.375","article-title":"Sociolectal analysis of pretrained language models","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Zhang","year":"2021"},{"key":"2024042419444642400_bib385","article-title":"BERTScore: Evaluating text generation with BERT","volume-title":"International Conference on Learning Representations","author":"Zhang","year":"2020"},{"key":"2024042419444642400_bib386","doi-asserted-by":"publisher","first-page":"1112","DOI":"10.18653\/v1\/2021.acl-long.90","article-title":"When do you need billions of words of pretraining data?","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Zhang","year":"2021"},{"key":"2024042419444642400_bib387","doi-asserted-by":"publisher","first-page":"1441","DOI":"10.18653\/v1\/P19-1139","article-title":"ERNIE: Enhanced language representation with informative entities","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Zhang","year":"2019"},{"key":"2024042419444642400_bib388","doi-asserted-by":"publisher","first-page":"72","DOI":"10.18653\/v1\/2021.conll-1.6","article-title":"Do pretrained transformers infer telicity like humans?","volume-title":"Proceedings of the 25th Conference on Computational Natural Language Learning","author":"Zhao","year":"2021"},{"key":"2024042419444642400_bib389","doi-asserted-by":"publisher","first-page":"500","DOI":"10.1145\/3442442.3452313","article-title":"A comparative study of using pre-trained language models for toxic comment classification","volume-title":"The ACM Web Conference","author":"Zhao","year":"2021"},{"key":"2024042419444642400_bib390","doi-asserted-by":"publisher","first-page":"2074","DOI":"10.18653\/v1\/2022.findings-acl.164","article-title":"Richer countries and richer representations","volume-title":"Findings of the Association for Computational Linguistics: ACL 2022","author":"Zhou","year":"2022"},{"key":"2024042419444642400_bib391","doi-asserted-by":"publisher","first-page":"7560","DOI":"10.18653\/v1\/2021.emnlp-main.598","article-title":"RICA: Evaluating robust inference capabilities based on commonsense axioms","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Zhou","year":"2021"},{"key":"2024042419444642400_bib392","article-title":"A survey on GPT-3","author":"Zong","year":"2022","journal-title":"ArXiv"},{"key":"2024042419444642400_bib393","doi-asserted-by":"publisher","first-page":"1028","DOI":"10.3758\/s13423-015-0864-x","article-title":"Situation models, mental simulations, and abstract concepts in discourse comprehension","volume":"23","author":"Zwaan","year":"2016","journal-title":"Psychonomic Bulletin & Review"}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/50\/1\/293\/2367117\/coli_a_00492.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/50\/1\/293\/2367117\/coli_a_00492.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,4,24]],"date-time":"2024-04-24T15:46:26Z","timestamp":1713973586000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/50\/1\/293\/118131\/Language-Model-Behavior-A-Comprehensive-Survey"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":393,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,3,1]]},"published-print":{"date-parts":[[2024,3,1]]}},"URL":"https:\/\/doi.org\/10.1162\/coli_a_00492","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"value":"0891-2017","type":"print"},{"value":"1530-9312","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}