{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T22:21:25Z","timestamp":1784240485847,"version":"3.55.0"},"reference-count":60,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2023,1,24]],"date-time":"2023-01-24T00:00:00Z","timestamp":1674518400000},"content-version":"vor","delay-in-days":23,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,1,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Today\u2019s probabilistic language generators fall short when it comes to producing coherent and fluent text despite the fact that the underlying models perform well under standard metrics (e.g., perplexity). This discrepancy has puzzled the language generation community for the last few years. In this work, we posit that the abstraction of natural language generation as a discrete stochastic process\u2014which allows for an information-theoretic analysis\u2014can provide new insights into the behavior of probabilistic language generators, for example, why high-probability texts can be dull or repetitive. Humans use language as a means of communicating information, aiming to do so in a simultaneously efficient and error-minimizing manner; in fact, psycholinguistics research suggests humans choose each word in a string with this subconscious goal in mind. We formally define the set of strings that meet this criterion: Those for which each word has an information content close to the expected information content, namely, the conditional entropy of our model. We then propose a simple and efficient procedure for enforcing this criterion when generating from probabilistic models, which we call locally typical sampling. Automatic and human evaluations show that, in comparison to nucleus and top-k sampling, locally typical sampling offers competitive performance (in both abstractive summarization and story generation) in terms of quality while consistently reducing degenerate repetitions.<\/jats:p>","DOI":"10.1162\/tacl_a_00536","type":"journal-article","created":{"date-parts":[[2023,1,24]],"date-time":"2023-01-24T16:28:16Z","timestamp":1674577696000},"page":"102-121","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":43,"title":["Locally Typical Sampling"],"prefix":"10.1162","volume":"11","author":[{"given":"Clara","family":"Meister","sequence":"first","affiliation":[{"name":"ETH Z\u00fcrich, Switzerland. clara.meister@inf.ethz.ch"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tiago","family":"Pimentel","sequence":"additional","affiliation":[{"name":"University of Cambridge, UK. tp472@cam.ac.uk"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gian","family":"Wiher","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich, Switzerland. gian.wiher@inf.ethz.ch"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ryan","family":"Cotterell","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich, Switzerland"},{"name":"University of Cambridge, UK. ryan.cotterell@inf.ethz.ch"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2023,1,12]]},"reference":[{"issue":"1","key":"2023012416103901800_bib1","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1177\/00238309040470010201","article-title":"The smooth signal redundancy hypothesis: A functional explanation for relationships between redundancy, prosodic prominence, and duration in spontaneous speech","volume":"47","author":"Aylett","year":"2004","journal-title":"Language and Speech"},{"key":"2023012416103901800_bib2","article-title":"Mirostat: A perplexity- controlled neural text decoding algorithm","volume-title":"Proceedings of the 9th International Conference on Learning Representations","author":"Basu","year":"2021"},{"issue":"5","key":"2023012416103901800_bib3","doi-asserted-by":"publisher","first-page":"442","DOI":"10.1109\/T-C.1973.223746","article-title":"Applying probability measures to abstract languages","volume":"C-22","author":"Booth","year":"1973","journal-title":"IEEE Transactions on Computers"},{"key":"2023012416103901800_bib4","first-page":"1089","article-title":"Calibration, entropy rates, and memory in language models","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Braverman","year":"2020"},{"issue":"3","key":"2023012416103901800_bib5","doi-asserted-by":"publisher","first-page":"809","DOI":"10.1214\/aoms\/1177706899","article-title":"The individual ergodic theorem of information theory","volume":"28","author":"Breiman","year":"1957","journal-title":"The Annals of Mathematical Statistics"},{"key":"2023012416103901800_bib6","first-page":"1877","article-title":"Language models are few- shot learners","volume-title":"Advances in Neural Information Processing Systems","author":"Brown","year":"2020"},{"key":"2023012416103901800_bib7","doi-asserted-by":"publisher","DOI":"10.1515\/9783112316009","volume-title":"Syntactic Structures","author":"Chomsky","year":"1957"},{"key":"2023012416103901800_bib8","volume-title":"The Minimalist Program","author":"Chomsky","year":"1995"},{"key":"2023012416103901800_bib9","article-title":"PaLM: Scaling language modeling with pathways","author":"Chowdhery","year":"2022","journal-title":"CoRR"},{"issue":"9","key":"2023012416103901800_bib10","doi-asserted-by":"publisher","DOI":"10.1126\/sciadv.aaw2594","article-title":"Different languages, similar encoding efficiency: Comparable information rates across the human communicative niche","volume":"5","author":"Coup\u00e9","year":"2019","journal-title":"Science Advances"},{"key":"2023012416103901800_bib11","volume-title":"Elements of Information Theory","author":"Cover","year":"2012"},{"key":"2023012416103901800_bib12","doi-asserted-by":"publisher","first-page":"166","DOI":"10.18653\/v1\/2021.gem-1.16","article-title":"Decoding methods for neural narrative generation","volume-title":"Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021)","author":"DeLucia","year":"2021"},{"key":"2023012416103901800_bib13","article-title":"Musings on typicality","author":"Dieleman","year":"2020"},{"key":"2023012416103901800_bib14","doi-asserted-by":"publisher","first-page":"489","DOI":"10.18653\/v1\/D18-1045","article-title":"Understanding back- translation at scale","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Edunov","year":"2018"},{"key":"2023012416103901800_bib15","doi-asserted-by":"publisher","first-page":"4506","DOI":"10.18653\/v1\/2020.coling-main.398","article-title":"Is MAP decoding all you need? The inadequacy of the mode in neural machine translation","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics, COLING","author":"Eikema","year":"2020"},{"key":"2023012416103901800_bib16","doi-asserted-by":"crossref","first-page":"889","DOI":"10.18653\/v1\/P18-1082","article-title":"Hierarchical neural story generation","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Fan","year":"2018"},{"issue":"3","key":"2023012416103901800_bib17","first-page":"400","article-title":"Konstanz im Kurzzeitged\u00e4chtnis-Konstanz im sprachlichen Informationsflu\u00df","volume":"27","author":"Fenk","year":"1980","journal-title":"Zeitschrift f\u00fcr experimentelle und angewandte Psychologie"},{"issue":"5","key":"2023012416103901800_bib18","doi-asserted-by":"publisher","first-page":"389","DOI":"10.1016\/j.tics.2019.02.003","article-title":"How efficiency shapes human language","volume":"23","author":"Gibson","year":"2019","journal-title":"Trends in Cognitive Sciences"},{"key":"2023012416103901800_bib19","first-page":"1724","article-title":"An empirical investigation of global and local normalization for recurrent neural sequence models using a continuous relaxation to beam search","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Goyal","year":"2019"},{"key":"2023012416103901800_bib20","doi-asserted-by":"crossref","first-page":"93","DOI":"10.1162\/tacl_a_00302","article-title":"A knowledge- enhanced pretraining model for commonsense story generation","volume":"8","author":"Guan","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2023012416103901800_bib21","doi-asserted-by":"publisher","DOI":"10.3115\/1073336.1073357","article-title":"A probabilistic Earley parser as a psycholinguistic model","volume-title":"Second Meeting of the North American Chapter of the Association for Computational Linguistics","author":"Hale","year":"2001"},{"key":"2023012416103901800_bib22","article-title":"Training compute-optimal large language models","author":"Hoffmann","year":"2022","journal-title":"CoRR"},{"key":"2023012416103901800_bib23","article-title":"The curious case of neural text degeneration","volume-title":"Proceedings of the 8th International Conference on Learning Representations","author":"Holtzman","year":"2020"},{"key":"2023012416103901800_bib24","doi-asserted-by":"publisher","DOI":"10.1515\/9783110219258.43","volume-title":"3. Syntactic recursion and iteration","author":"Karlsson","year":"2010"},{"key":"2023012416103901800_bib25","doi-asserted-by":"publisher","first-page":"284","DOI":"10.18653\/v1\/P18-1027","article-title":"Sharp nearby, fuzzy far away: How neural language models use context","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Khandelwal","year":"2018"},{"key":"2023012416103901800_bib26","volume-title":"Theory and Application of Infinite Series","author":"Knopp","year":"1954"},{"issue":"5","key":"2023012416103901800_bib27","doi-asserted-by":"publisher","first-page":"1202","DOI":"10.1111\/cogs.12414","article-title":"Grammaticality, acceptability, and probability: A probabilistic view of linguistic knowledge","volume":"41","author":"Lau","year":"2017","journal-title":"Cognitive Science"},{"key":"2023012416103901800_bib28","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/7503.003.0111","article-title":"Speakers optimize information density through syntactic reduction","volume-title":"Advances in Neural Information Processing Systems","author":"Levy","year":"2007"},{"key":"2023012416103901800_bib29","doi-asserted-by":"publisher","first-page":"7871","DOI":"10.18653\/v1\/2020.acl-main.703","article-title":"BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Lewis","year":"2020"},{"key":"2023012416103901800_bib30","first-page":"110","article-title":"A diversity-promoting objective function for neural conversation models","volume-title":"Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Li","year":"2016"},{"key":"2023012416103901800_bib31","doi-asserted-by":"crossref","first-page":"4715","DOI":"10.18653\/v1\/2020.acl-main.428","article-title":"Don\u2019t say that! Making inconsistent dialogue unlikely with unlikelihood training","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Li","year":"2020"},{"issue":"2","key":"2023012416103901800_bib32","doi-asserted-by":"publisher","first-page":"313","DOI":"10.1016\/j.cognition.2012.09.010","article-title":"Info\/information theory: Speakers choose shorter words in predictive contexts","volume":"126","author":"Mahowald","year":"2013","journal-title":"Cognition"},{"issue":"2","key":"2023012416103901800_bib33","doi-asserted-by":"publisher","first-page":"196","DOI":"10.1214\/aoms\/1177729028","article-title":"The basic theorems of information theory","volume":"24","author":"McMillan","year":"1953","journal-title":"The Annals of Mathematical Statistics"},{"key":"2023012416103901800_bib34","doi-asserted-by":"publisher","first-page":"6870","DOI":"10.18653\/v1\/2020.acl-main.615","article-title":"Generalized entropy regularization or: There\u2019s nothing special about label smoothing","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Meister","year":"2020"},{"key":"2023012416103901800_bib35","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.170","article-title":"If beam search is the answer, what was the question?","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing","author":"Meister","year":"2020"},{"key":"2023012416103901800_bib36","doi-asserted-by":"publisher","first-page":"36","DOI":"10.18653\/v1\/2022.acl-short.5","article-title":"On the probability\u2013 quality paradox in language generation","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Meister","year":"2022"},{"key":"2023012416103901800_bib37","article-title":"Pointer sentinel mixture models","volume-title":"Proceedings of the 5th International Conference on Learning Representations","author":"Merity","year":"2017"},{"key":"2023012416103901800_bib38","first-page":"334","article-title":"A systematic characterization of sampling algorithms for open- ended language generation","volume-title":"Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing","author":"Nadeem","year":"2020"},{"key":"2023012416103901800_bib39","doi-asserted-by":"publisher","first-page":"280","DOI":"10.18653\/v1\/K16-1028","article-title":"Abstractive text summarization using sequence-to-sequence RNNs and beyond","volume-title":"Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning","author":"Nallapati","year":"2016"},{"key":"2023012416103901800_bib40","doi-asserted-by":"crossref","first-page":"314","DOI":"10.18653\/v1\/W19-5333","article-title":"Facebook FAIR\u2019s WMT19 news translation task submission","volume-title":"Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1)","author":"Ng","year":"2019"},{"key":"2023012416103901800_bib41","article-title":"Regularizing neural networks by penalizing confident output distributions","volume-title":"Proceedings of the 5th International Conference on Learning Representations","author":"Pereyra","year":"2017"},{"issue":"9","key":"2023012416103901800_bib42","doi-asserted-by":"publisher","first-page":"3526","DOI":"10.1073\/pnas.1012551108","article-title":"Word lengths are optimized for efficient communication","volume":"108","author":"Piantadosi","year":"2011","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2023012416103901800_bib43","first-page":"4816","article-title":"MAUVE: Measuring the gap between neural text and human text using divergence frontiers","volume-title":"Advances in Neural Information Processing Systems","author":"Pillutla","year":"2021"},{"key":"2023012416103901800_bib44","article-title":"Cluster-based evaluation of automatically generated text","author":"Pimentel","year":"2022","journal-title":"arXiv preprint arXiv:2205.16001"},{"key":"2023012416103901800_bib45","doi-asserted-by":"publisher","first-page":"949","DOI":"10.18653\/v1\/2021.emnlp-main.73","article-title":"A surprisal\u2013duration trade-off across and within the world\u2019s languages","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Pimentel","year":"2021"},{"key":"2023012416103901800_bib46","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1162\/tacl_a_00296","article-title":"Phonotactic complexity and its trade-offs","volume":"8","author":"Pimentel","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2023012416103901800_bib47","article-title":"Language models are unsupervised multitask learners","author":"Radford","year":"2019"},{"issue":"4","key":"2023012416103901800_bib48","doi-asserted-by":"publisher","first-page":"831","DOI":"10.2307\/412337","article-title":"The finiteness of natural language","volume":"45","author":"Reich","year":"1969","journal-title":"Language"},{"key":"2023012416103901800_bib49","doi-asserted-by":"publisher","DOI":"10.26530\/OAPEN_603356","volume-title":"The empirical base of linguistics: Grammaticality judgments and linguistic methodology","author":"Sch\u00fctze","year":"2016"},{"key":"2023012416103901800_bib50","doi-asserted-by":"crossref","first-page":"843","DOI":"10.18653\/v1\/K19-1079","article-title":"Do massively pretrained language models make better storytellers?","volume-title":"Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL)","author":"See","year":"2019"},{"key":"2023012416103901800_bib51","doi-asserted-by":"crossref","first-page":"623","DOI":"10.1002\/j.1538-7305.1948.tb00917.x","article-title":"A mathematical theory of communication","volume":"27","author":"Shannon","year":"1948","journal-title":"Bell System Technical Journal"},{"issue":"1","key":"2023012416103901800_bib52","doi-asserted-by":"publisher","first-page":"50","DOI":"10.1002\/j.1538-7305.1951.tb01366.x","article-title":"Prediction and entropy of printed English","volume":"30","author":"Shannon","year":"1951","journal-title":"Bell System Technical Journal"},{"key":"2023012416103901800_bib53","doi-asserted-by":"publisher","first-page":"3356","DOI":"10.18653\/v1\/D19-1331","article-title":"On NMT search errors and model errors: Cat got your tongue?","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Stahlberg","year":"2019"},{"key":"2023012416103901800_bib54","doi-asserted-by":"publisher","first-page":"355","DOI":"10.18653\/v1\/W19-8643","article-title":"Best practices for the human evaluation of automatically generated text","volume-title":"Proceedings of the 12th International Conference on Natural Language Generation","author":"van der Lee","year":"2019"},{"key":"2023012416103901800_bib55","article-title":"Neural text generation with unlikelihood training","volume-title":"Proceedings of the 8th International Conference on Learning Representations","author":"Welleck","year":"2020"},{"key":"2023012416103901800_bib56","doi-asserted-by":"publisher","first-page":"38","DOI":"10.18653\/v1\/2020.emnlp-demos.6","article-title":"Transformers: State-of-the-art natural language processing","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Wolf","year":"2020"},{"key":"2023012416103901800_bib57","article-title":"Google\u2019s neural machine translation system: Bridging the gap between human and machine translation","author":"Yonghui","year":"2016","journal-title":"CoRR"},{"issue":"31","key":"2023012416103901800_bib58","doi-asserted-by":"publisher","first-page":"7937","DOI":"10.1073\/pnas.1800521115","article-title":"Efficient compression in color naming and its evolution","volume":"115","author":"Zaslavsky","year":"2018","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2023012416103901800_bib59","first-page":"25","article-title":"Trading off diversity and quality in natural language generation","volume-title":"Proceedings of the Workshop on Human Evaluation of NLP Systems (HumEval)","author":"Zhang","year":"2021"},{"key":"2023012416103901800_bib60","volume-title":"Human Behavior and the Principle of Least Effort","author":"Zipf","year":"1949"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00536\/2067865\/tacl_a_00536.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00536\/2067865\/tacl_a_00536.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T14:31:46Z","timestamp":1701786706000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00536\/114593\/Locally-Typical-Sampling"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023]]},"references-count":60,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00536","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023]]},"published":{"date-parts":[[2023]]}}}