{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T10:46:49Z","timestamp":1781779609485,"version":"3.54.5"},"reference-count":51,"publisher":"MIT Press - Journals","license":[{"start":{"date-parts":[[2021,12,23]],"date-time":"2021-12-23T00:00:00Z","timestamp":1640217600000},"content-version":"vor","delay-in-days":356,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,12,17]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>We introduce a simple but flexible mechanism to learn an intermediate plan to ground the generation of abstractive summaries. Specifically, we prepend (or prompt) target summaries with entity chains\u2014ordered sequences of entities mentioned in the summary. Transformer-based sequence-to-sequence models are then trained to generate the entity chain and then continue generating the summary conditioned on the entity chain and the input. We experimented with both pretraining and finetuning with this content planning objective. When evaluated on CNN\/DailyMail, XSum, SAMSum, and BillSum, we demonstrate empirically that the grounded generation with the planning objective improves entity specificity and planning in summaries for all datasets, and achieves state-of-the-art performance on XSum and SAMSum in terms of rouge. Moreover, we demonstrate empirically that planning with entity chains provides a mechanism to control hallucinations in abstractive summaries. By prompting the decoder with a modified content plan that drops hallucinated entities, we outperform state-of-the-art approaches for faithfulness when evaluated automatically and by humans.<\/jats:p>","DOI":"10.1162\/tacl_a_00438","type":"journal-article","created":{"date-parts":[[2021,12,24]],"date-time":"2021-12-24T05:58:18Z","timestamp":1640325498000},"page":"1475-1492","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":45,"title":["Planning with Learned Entity Prompts for Abstractive Summarization"],"prefix":"10.1162","volume":"9","author":[{"given":"Shashi","family":"Narayan","sequence":"first","affiliation":[{"name":"Google Research. shashinarayan@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yao","family":"Zhao","sequence":"additional","affiliation":[{"name":"Google Brain. yaozhaoyz@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Joshua","family":"Maynez","sequence":"additional","affiliation":[{"name":"Google Research. joshuahm@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gon\u00e7alo","family":"Sim\u00f5es","sequence":"additional","affiliation":[{"name":"Google Research. gsimoes@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Vitaly","family":"Nikolaev","sequence":"additional","affiliation":[{"name":"Google Research. vitalyn@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ryan","family":"McDonald","sequence":"additional","affiliation":[{"name":"ASAPP. ryanmcd@asapp.com"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2021,12,17]]},"reference":[{"key":"2021122316260765800_bib1","doi-asserted-by":"publisher","DOI":"10.3115\/1608810.1608825","article-title":"Using coreference chains for text summarization","volume-title":"Coreference and Its Applications","author":"Azzam","year":"1999"},{"key":"2021122316260765800_bib2","article-title":"Neural machine translation by jointly learning to align and translate","author":"Bahdanau","year":"2015","journal-title":"CoRR"},{"key":"2021122316260765800_bib3","article-title":"Using lexical chains for text summarization","volume-title":"Intelligent Scalable Text Summarization","author":"Barzilay","year":"1997"},{"key":"2021122316260765800_bib4","first-page":"4171","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2021122316260765800_bib5","doi-asserted-by":"publisher","first-page":"4884","DOI":"10.18653\/v1\/P19-1483","article-title":"Handling divergent reference texts when evaluating table-to-text generation","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Dhingra","year":"2019"},{"key":"2021122316260765800_bib6","first-page":"13042","article-title":"Unified language model pre-training for natural language understanding and generation","volume-title":"Advances in Neural Information Processing Systems 32","author":"Li","year":"2019"},{"key":"2021122316260765800_bib7","doi-asserted-by":"publisher","first-page":"4830","DOI":"10.18653\/v1\/2021.naacl-main.384","article-title":"GSum: A general framework for guided neural abstractive summarization","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Dou","year":"2021"},{"key":"2021122316260765800_bib8","doi-asserted-by":"publisher","first-page":"5055","DOI":"10.18653\/v1\/2020.acl-main.454","article-title":"FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Durmus","year":"2020"},{"key":"2021122316260765800_bib9","doi-asserted-by":"publisher","first-page":"2214","DOI":"10.18653\/v1\/P19-1213","article-title":"Ranking generated summaries by correctness: An interesting but challenging application for natural language inference","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Falke","year":"2019"},{"key":"2021122316260765800_bib10","doi-asserted-by":"crossref","first-page":"45","DOI":"10.18653\/v1\/W18-2706","article-title":"Controllable abstractive summarization","volume-title":"Proceedings of the 2nd Workshop on Neural Machine Translation and Generation","author":"Fan","year":"2018"},{"key":"2021122316260765800_bib11","doi-asserted-by":"publisher","first-page":"478","DOI":"10.18653\/v1\/2021.findings-acl.42","article-title":"GO FIGURE: A meta evaluation of factuality in summarization","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Gabriel","year":"2021"},{"key":"2021122316260765800_bib12","doi-asserted-by":"publisher","first-page":"70","DOI":"10.18653\/v1\/D19-5409","article-title":"SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization","volume-title":"Proceedings of the 2nd Workshop on New Frontiers in Summarization","author":"Gliwa","year":"2019"},{"key":"2021122316260765800_bib13","article-title":"REALM: Retrieval-augmented language model pre-training","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Guu","year":"2020"},{"key":"2021122316260765800_bib14","volume-title":"Cohesion in English","author":"Halliday","year":"1976"},{"key":"2021122316260765800_bib15","article-title":"CTRLsum: Towards generic controllable text summarization","author":"He","year":"2020","journal-title":"CoRR"},{"key":"2021122316260765800_bib16","first-page":"1693","article-title":"Teaching machines to read and comprehend","volume-title":"Advances in Neural Information Processing Systems 28","author":"Hermann","year":"2015"},{"issue":"8","key":"2021122316260765800_bib17","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Computation"},{"key":"2021122316260765800_bib18","doi-asserted-by":"publisher","first-page":"781","DOI":"10.18653\/v1\/2020.emnlp-main.57","article-title":"PAIR: Planning and iterative refinement in pre-trained transformers for long text generation","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Hua","year":"2020"},{"key":"2021122316260765800_bib19","first-page":"9","article-title":"What might be in a summary?","author":"Jones","year":"1993","journal-title":"Information Retrieval"},{"key":"2021122316260765800_bib20","article-title":"Sample efficient text summarization using a single pre-trained transformer","author":"Khandelwal","year":"2019","journal-title":"CoRR"},{"key":"2021122316260765800_bib21","doi-asserted-by":"publisher","first-page":"329","DOI":"10.18653\/v1\/D16-1032","article-title":"Globally coherent text generation with neural checklist models","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing","author":"Kiddon","year":"2016"},{"key":"2021122316260765800_bib22","first-page":"48","article-title":"BillSum: A corpus for automatic summarization of US legislation","volume-title":"Proceedings of the 2nd Workshop on New Frontiers in Summarization","author":"Kornilova","year":"2019"},{"key":"2021122316260765800_bib23","doi-asserted-by":"publisher","first-page":"9332","DOI":"10.18653\/v1\/2020.emnlp-main.750","article-title":"Evaluating the factual consistency of abstractive text summarization","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Kryscinski","year":"2020"},{"key":"2021122316260765800_bib24","doi-asserted-by":"publisher","first-page":"7871","DOI":"10.18653\/v1\/2020.acl-main.703","article-title":"BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Lewis","year":"2020"},{"key":"2021122316260765800_bib25","first-page":"150","article-title":"Automatic evaluation of summaries using n-gram co-occurrence statistics","volume-title":"Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics","author":"Lin","year":"2003"},{"issue":"2","key":"2021122316260765800_bib26","doi-asserted-by":"publisher","first-page":"159","DOI":"10.1147\/rd.22.0159","article-title":"The automatic creation of literature abstracts","volume":"2","author":"Luhn","year":"1958","journal-title":"IBM Journal of Research and Development"},{"issue":"3","key":"2021122316260765800_bib27","first-page":"243","article-title":"Rhetorical Structure Theory: Toward a functional theory of text organization","volume":"8","author":"Mann","year":"1988","journal-title":"Text"},{"key":"2021122316260765800_bib28","article-title":"Constrained abstractive summarization: Preserving factual consistency with constrained generation","author":"Mao","year":"2020","journal-title":"CoRR"},{"key":"2021122316260765800_bib29","first-page":"868","article-title":"Event representations for automated story generation with deep neural nets","volume-title":"Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18), New Orleans, Louisiana, USA, February 2\u20137, 2018","author":"Martin","year":"2018"},{"key":"2021122316260765800_bib30","doi-asserted-by":"publisher","first-page":"1906","DOI":"10.18653\/v1\/2020.acl-main.173","article-title":"On faithfulness and factuality in abstractive summarization","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Maynez","year":"2020"},{"key":"2021122316260765800_bib31","doi-asserted-by":"publisher","first-page":"74","DOI":"10.1145\/215206.215334","article-title":"Generating summaries of multiple news articles","volume-title":"Proceedings of the 18th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"McKeown","year":"1995"},{"key":"2021122316260765800_bib32","first-page":"404","article-title":"TextRank: Bringing order into text","volume-title":"Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing","author":"Mihalcea","year":"2004"},{"key":"2021122316260765800_bib33","doi-asserted-by":"publisher","first-page":"1797","DOI":"10.18653\/v1\/D18-1206","article-title":"Don\u2019t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Narayan","year":"2018"},{"key":"2021122316260765800_bib34","article-title":"QURIOUS: Question generation pretraining for text generation","author":"Narayan","year":"2020","journal-title":"CoRR"},{"key":"2021122316260765800_bib35","first-page":"646","article-title":"Multi-reward reinforced summarization with saliency and entailment","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers)","author":"Pasunuru","year":"2018"},{"issue":"1","key":"2021122316260765800_bib36","doi-asserted-by":"publisher","first-page":"6908","DOI":"10.1609\/aaai.v33i01.33016908","article-title":"Data-to-text generation with content selection and planning","volume":"33","author":"Puduppully","year":"2019","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2021122316260765800_bib37","doi-asserted-by":"crossref","first-page":"2401","DOI":"10.18653\/v1\/2020.findings-emnlp.217","article-title":"ProphetNet: Predicting future n-gram for sequence-to-SequencePre-training","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Qi","year":"2020"},{"key":"2021122316260765800_bib38","unstructured":"Alec\n              Radford\n            , KarthikNarasimhan, TimSalimans, and IlyaSutskever. 2018. Improving language understanding by generative pre-training. Technical report, OpenAI."},{"key":"2021122316260765800_bib39","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","author":"Raffel","year":"2019","journal-title":"CoRR"},{"issue":"1","key":"2021122316260765800_bib40","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1017\/S1351324997001502","article-title":"Building applied natural language generation systems","volume":"3","author":"Reiter","year":"1997","journal-title":"Natural Language Engineering"},{"key":"2021122316260765800_bib41","doi-asserted-by":"publisher","first-page":"264","DOI":"10.1017\/S1351324997001502","article-title":"Leveraging pre-trained checkpoints for sequence generation tasks","volume":"8","author":"Rothe","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2021122316260765800_bib42","doi-asserted-by":"crossref","first-page":"1073","DOI":"10.18653\/v1\/P17-1099","article-title":"Get to the point: Summarization with pointer-generator networks","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"See","year":"2017"},{"key":"2021122316260765800_bib43","first-page":"4596","article-title":"Adafactor: Adaptive learning rates with sublinear memory cost","volume-title":"International Conference on Machine Learning","author":"Shazeer","year":"2018"},{"key":"2021122316260765800_bib44","first-page":"5926","article-title":"MASS: Masked sequence to sequence pre-training for language generation","volume-title":"Proceedings of the 36th International Conference on Machine Learning","author":"Song","year":"2019"},{"key":"2021122316260765800_bib45","first-page":"3104","article-title":"Sequence to sequence learning with neural networks","volume-title":"Advances in Neural Information Processing Systems 27","author":"Sutskever","year":"2014"},{"key":"2021122316260765800_bib46","doi-asserted-by":"publisher","first-page":"142","DOI":"10.3115\/1119176.1119195","article-title":"Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition","volume-title":"Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003","author":"Tjong Kim Sang","year":"2003"},{"key":"2021122316260765800_bib47","first-page":"5998","article-title":"Attention is all you need","volume-title":"Advances in Neural Information Processing Systems 30","author":"Vaswani","year":"2017"},{"key":"2021122316260765800_bib48","doi-asserted-by":"publisher","first-page":"1112","DOI":"10.18653\/v1\/N18-1101","article-title":"A broad-coverage challenge corpus for sentence understanding through inference","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)","author":"Williams","year":"2018"},{"key":"2021122316260765800_bib49","doi-asserted-by":"publisher","first-page":"3174","DOI":"10.18653\/v1\/D18-1356","article-title":"Learning neural templates for text generation","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Wiseman","year":"2018"},{"key":"2021122316260765800_bib50","doi-asserted-by":"publisher","first-page":"7378","DOI":"10.1609\/aaai.v33i01.33017378","article-title":"Plan-and-write: Towards better automatic storytelling","volume-title":"The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 \u2013 February 1, 2019","author":"Yao","year":"2019"},{"key":"2021122316260765800_bib51","article-title":"PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Zhang","year":"2020"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00438\/1979348\/tacl_a_00438.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00438\/1979348\/tacl_a_00438.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,12,24]],"date-time":"2021-12-24T06:00:09Z","timestamp":1640325609000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00438\/108867\/Planning-with-Learned-Entity-Prompts-for"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021]]},"references-count":51,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00438","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2021]]},"published":{"date-parts":[[2021]]}}}