{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T16:36:37Z","timestamp":1780677397435,"version":"3.54.1"},"reference-count":76,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2022,9,21]],"date-time":"2022-09-21T00:00:00Z","timestamp":1663718400000},"content-version":"vor","delay-in-days":263,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,9,19]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Knowledge-grounded dialogue systems powered by large language models often generate responses that, while fluent, are not attributable to a relevant source of information. Progress towards models that do not exhibit this issue requires evaluation metrics that can quantify its prevalence. To this end, we introduce the Benchmark for Evaluation of Grounded INteraction (Begin), comprising 12k dialogue turns generated by neural dialogue systems trained on three knowledge-grounded dialogue corpora. We collect human annotations assessing the extent to which the models\u2019 responses can be attributed to the given background information. We then use Begin to analyze eight evaluation metrics. We find that these metrics rely on spurious correlations, do not reliably distinguish attributable abstractive responses from unattributable ones, and perform substantially worse when the knowledge source is longer. Our findings underscore the need for more sophisticated and robust evaluation metrics for knowledge-grounded dialogue. We make Begin publicly available at https:\/\/github.com\/google\/BEGIN-dataset.<\/jats:p>","DOI":"10.1162\/tacl_a_00506","type":"journal-article","created":{"date-parts":[[2022,9,21]],"date-time":"2022-09-21T18:03:04Z","timestamp":1663783384000},"page":"1066-1083","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":25,"title":["Evaluating Attribution in Dialogue Systems: The BEGIN Benchmark"],"prefix":"10.1162","volume":"10","author":[{"given":"Nouha","family":"Dziri","sequence":"first","affiliation":[{"name":"University of Alberta, Canada. dziri@cs.ualberta.ca"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hannah","family":"Rashkin","sequence":"additional","affiliation":[{"name":"Google Research, USA. hrashkin@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tal","family":"Linzen","sequence":"additional","affiliation":[{"name":"Google Research, USA. linzen@google.com"},{"name":"New York University, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"David","family":"Reitter","sequence":"additional","affiliation":[{"name":"Google Research, USA. reitter@google.com"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2022,9,19]]},"reference":[{"key":"2022092118021074700_bib1","article-title":"Towards a human-like open-domain chatbot","author":"Adiwardana","year":"2020","journal-title":"CoRR (arXiv preprint)"},{"key":"2022092118021074700_bib2","article-title":"A neural probabilistic language model","volume-title":"Advances in Neural Information Processing Systems","author":"Bengio","year":"2000"},{"key":"2022092118021074700_bib3","doi-asserted-by":"publisher","first-page":"9347","DOI":"10.18653\/v1\/2020.emnlp-main.751","article-title":"Re-evaluating evaluation in text summarization","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Bhandari","year":"2020"},{"key":"2022092118021074700_bib4","first-page":"549","article-title":"The ISO standard for dialogue act annotation, second edition","volume-title":"Proceedings of the 12th Language Resources and Evaluation Conference","author":"Bunt","year":"2020"},{"issue":"1","key":"2022092118021074700_bib5","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11912","article-title":"Faithful to the original: Fact aware neural abstractive summarization","volume":"32","author":"Cao","year":"2018","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2022092118021074700_bib6","doi-asserted-by":"publisher","first-page":"2174","DOI":"10.18653\/v1\/D18-1241","article-title":"QuAC: Question answering in context","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Choi","year":"2018"},{"key":"2022092118021074700_bib7","doi-asserted-by":"publisher","first-page":"4884","DOI":"10.18653\/v1\/P19-1483","article-title":"Handling divergent reference texts when evaluating table-to-text generation","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Dhingra","year":"2019"},{"key":"2022092118021074700_bib8","article-title":"Wizard of wikipedia: Knowledge-powered conversational agents","volume-title":"7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6\u20139, 2019","author":"Dinan","year":"2019"},{"key":"2022092118021074700_bib9","doi-asserted-by":"publisher","first-page":"5055","DOI":"10.18653\/v1\/2020.acl-main.454","article-title":"FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Durmus","year":"2020"},{"key":"2022092118021074700_bib10","doi-asserted-by":"publisher","first-page":"1443","DOI":"10.18653\/v1\/2022.acl-long.102","article-title":"Spurious correlations in reference-free evaluation of text generation","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Durmus","year":"2022"},{"key":"2022092118021074700_bib11","doi-asserted-by":"publisher","first-page":"3806","DOI":"10.18653\/v1\/N19-1381","article-title":"Evaluating coherence in dialogue systems using entailment","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Dziri","year":"2019"},{"key":"2022092118021074700_bib12","article-title":"Faithdial: A faithful benchmark for information-seeking dialogue","author":"Dziri","year":"2022","journal-title":"CoRR (arXiv preprint)"},{"key":"2022092118021074700_bib13","doi-asserted-by":"publisher","first-page":"2197","DOI":"10.18653\/v1\/2021.emnlp-main.168","article-title":"Neural path hunter: Reducing hallucination in dialogue systems via path grounding","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Dziri","year":"2021"},{"key":"2022092118021074700_bib14","article-title":"On the origin of hallucinations in conversational models: Is it the datasets or the models?","author":"Dziri","year":"2022","journal-title":"CoRR (arXiv preprint)"},{"key":"2022092118021074700_bib15","doi-asserted-by":"publisher","first-page":"391","DOI":"10.1162\/tacl_a_00373","article-title":"SummEval: Re-evaluating summarization evaluation","volume":"9","author":"Fabbri","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2022092118021074700_bib16","doi-asserted-by":"publisher","first-page":"2214","DOI":"10.18653\/v1\/P19-1213","article-title":"Ranking generated summaries by correctness: An interesting but challenging application for natural language inference","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Falke","year":"2019"},{"key":"2022092118021074700_bib17","article-title":"ERG semantic documentation","author":"Flickinger","year":"2014"},{"key":"2022092118021074700_bib18","doi-asserted-by":"publisher","first-page":"1460","DOI":"10.1162\/tacl_a_00437","article-title":"Experts, errors, and context: A large-scale study of human evaluation for machine translation","volume":"9","author":"Freitag","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2022092118021074700_bib19","doi-asserted-by":"publisher","first-page":"478","DOI":"10.18653\/v1\/2021.findings-acl.42","article-title":"GO FIGURE: A meta evaluation of factuality in summarization","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Gabriel","year":"2021"},{"key":"2022092118021074700_bib20","doi-asserted-by":"publisher","first-page":"96","DOI":"10.18653\/v1\/2021.gem-1.10","article-title":"The GEM benchmark: Natural language generation, its evaluation and metrics","volume-title":"Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021)","author":"Gehrmann","year":"2021"},{"key":"2022092118021074700_bib21","article-title":"Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text","author":"Gehrmann","year":"2022","journal-title":"CoRR (arXiv preprint)"},{"key":"2022092118021074700_bib22","doi-asserted-by":"publisher","first-page":"1891","DOI":"10.21437\/Interspeech.2019-3079","article-title":"Topical-Chat: Towards knowledge-grounded open-domain conversations","volume-title":"Proceedings of Interspeech 2019","author":"Gopalakrishnan","year":"2019"},{"key":"2022092118021074700_bib23","volume-title":"Studies in the Way of Words","author":"Grice","year":"1989"},{"key":"2022092118021074700_bib24","doi-asserted-by":"publisher","first-page":"708","DOI":"10.18653\/v1\/N18-1065","article-title":"Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)","author":"Grusky","year":"2018"},{"key":"2022092118021074700_bib25","doi-asserted-by":"publisher","first-page":"3785","DOI":"10.18653\/v1\/2022.acl-long.263","article-title":"DialFact: A benchmark for fact-checking in dialogue","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Gupta","year":"2022"},{"key":"2022092118021074700_bib26","article-title":"DeBERTa: Decoding- enhanced BERT with disentangled attention","volume-title":"International Conference on Learning Representations","author":"He","year":"2021"},{"key":"2022092118021074700_bib27","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.1212303","article-title":"SpaCy: Industrial-strength natural language processing in Python","author":"Honnibal","year":"2020"},{"key":"2022092118021074700_bib28","doi-asserted-by":"publisher","first-page":"7856","DOI":"10.18653\/v1\/2021.emnlp-main.619","article-title":"Q2: Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Or","year":"2021"},{"issue":"3","key":"2022092118021074700_bib29","doi-asserted-by":"publisher","first-page":"44","DOI":"10.1109\/MIC.2020.3037151","article-title":"Chatbots as conversational healthcare services","volume":"25","author":"Jovanovi\u0107","year":"2021","journal-title":"IEEE Internet Computing"},{"key":"2022092118021074700_bib30","doi-asserted-by":"publisher","first-page":"718","DOI":"10.18653\/v1\/2020.acl-main.66","article-title":"Improved natural language generation via loss truncation","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Kang","year":"2020"},{"key":"2022092118021074700_bib31","article-title":"CTRL: A conditional transformer language model for controllable generation","author":"Keskar","year":"2019","journal-title":"CoRR (arXiv preprint)"},{"key":"2022092118021074700_bib32","article-title":"Adam: A method for stochastic optimization","volume-title":"3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7\u20139, 2015, Conference Track Proceedings","author":"Kingma","year":"2015"},{"key":"2022092118021074700_bib33","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s40593-021-00267-x","article-title":"Automated data-driven generation of personalized pedagogical interventions in intelligent tutoring systems","author":"Kochmar","year":"2021","journal-title":"International Journal of Artificial Intelligence in Education"},{"key":"2022092118021074700_bib34","doi-asserted-by":"publisher","first-page":"9332","DOI":"10.18653\/v1\/2020.emnlp-main.750","article-title":"Evaluating the factual consistency of abstractive text summarization","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Kryscinski","year":"2020"},{"issue":"9","key":"2022092118021074700_bib35","doi-asserted-by":"publisher","first-page":"1248","DOI":"10.1093\/jamia\/ocy072","article-title":"Conversational agents in healthcare: A systematic review","volume":"25","author":"Laranjo","year":"2018","journal-title":"Journal of the American Medical Informatics Association"},{"key":"2022092118021074700_bib36","doi-asserted-by":"publisher","first-page":"7871","DOI":"10.18653\/v1\/2020.acl-main.703","article-title":"BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Lewis","year":"2020"},{"key":"2022092118021074700_bib37","first-page":"110","article-title":"A diversity- promoting objective function for neural conversation models","volume-title":"Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Li","year":"2016"},{"key":"2022092118021074700_bib38","first-page":"74","article-title":"ROUGE: A package for automatic evaluation of summaries","volume-title":"Text Summarization Branches Out","author":"Lin","year":"2004"},{"key":"2022092118021074700_bib39","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","author":"Liu","year":"2019","journal-title":"CoRR (arXiv preprint)"},{"key":"2022092118021074700_bib40","doi-asserted-by":"publisher","first-page":"4984","DOI":"10.18653\/v1\/2020.acl-main.448","article-title":"Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Mathur","year":"2020"},{"key":"2022092118021074700_bib41","doi-asserted-by":"publisher","first-page":"1906","DOI":"10.18653\/v1\/2020.acl-main.173","article-title":"On faithfulness and factuality in abstractive summarization","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Maynez","year":"2020"},{"key":"2022092118021074700_bib42","doi-asserted-by":"publisher","first-page":"1906","DOI":"10.18653\/v1\/2020.acl-main.173","article-title":"On faithfulness and factuality in abstractive summarization","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Maynez","year":"2020"},{"key":"2022092118021074700_bib43","doi-asserted-by":"publisher","first-page":"3428","DOI":"10.18653\/v1\/P19-1334","article-title":"Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"McCoy","year":"2019"},{"key":"2022092118021074700_bib44","doi-asserted-by":"publisher","DOI":"10.3115\/1075812.1075938","article-title":"WordNet: A lexical database for English","volume-title":"Human Language Technology: Proceedings of a Workshop held at Plainsboro, New Jersey, March 8\u201311, 1994","author":"Miller","year":"1994"},{"key":"2022092118021074700_bib45","doi-asserted-by":"crossref","first-page":"2339","DOI":"10.18653\/v1\/2020.acl-main.212","article-title":"Syntactic data augmentation increases robustness to inference heuristics","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Min","year":"2020"},{"key":"2022092118021074700_bib46","doi-asserted-by":"publisher","first-page":"6881","DOI":"10.18653\/v1\/2021.acl-long.536","article-title":"Improving factual consistency of abstractive summarization via question answering","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Nan","year":"2021"},{"key":"2022092118021074700_bib47","doi-asserted-by":"publisher","first-page":"4812","DOI":"10.18653\/v1\/2021.naacl-main.383","article-title":"Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Pagnoni","year":"2021"},{"key":"2022092118021074700_bib48","doi-asserted-by":"publisher","first-page":"311","DOI":"10.3115\/1073083.1073135","article-title":"BLEU: A method for automatic evaluation of machine translation","volume-title":"Proceedings of 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni","year":"2002"},{"key":"2022092118021074700_bib49","doi-asserted-by":"publisher","first-page":"4274","DOI":"10.18653\/v1\/2021.naacl-main.338","article-title":"Focused attention improves document-grounded generation","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Prabhumoye","year":"2021"},{"key":"2022092118021074700_bib50","doi-asserted-by":"publisher","first-page":"751","DOI":"10.18653\/v1\/2021.emnlp-main.58","article-title":"Learning compact metrics for MT","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Amy","year":"2021"},{"issue":"8","key":"2022092118021074700_bib51","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford","year":"2019","journal-title":"OpenAI Blog"},{"issue":"140","key":"2022092118021074700_bib52","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel","year":"2020","journal-title":"Journal of Machine Learning Research"},{"key":"2022092118021074700_bib53","article-title":"Measuring attribution in natural language generation models","author":"Rashkin","year":"2021","journal-title":"CoRR (arXiv preprint)"},{"key":"2022092118021074700_bib54","doi-asserted-by":"publisher","first-page":"704","DOI":"10.18653\/v1\/2021.acl-long.58","article-title":"Increasing faithfulness in knowledge-grounded dialogue with controllable features","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Rashkin","year":"2021"},{"key":"2022092118021074700_bib55","doi-asserted-by":"publisher","first-page":"249","DOI":"10.1162\/tacl_a_00266","article-title":"CoQA: A conversational question answering challenge","volume":"7","author":"Reddy","year":"2019","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2022092118021074700_bib56","doi-asserted-by":"publisher","first-page":"300","DOI":"10.18653\/v1\/2021.eacl-main.24","article-title":"Recipes for building an open-domain chatbot","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Roller","year":"2021"},{"key":"2022092118021074700_bib57","first-page":"1702","article-title":"What makes a good conversation? How controllable attributes affect human judgments","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"See","year":"2019"},{"key":"2022092118021074700_bib58","doi-asserted-by":"publisher","first-page":"7881","DOI":"10.18653\/v1\/2020.acl-main.704","article-title":"BLEURT: Learning robust metrics for text generation","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Sellam","year":"2020"},{"key":"2022092118021074700_bib59","doi-asserted-by":"publisher","first-page":"3784","DOI":"10.18653\/v1\/2021.findings-emnlp.320","article-title":"Retrieval augmentation reduces hallucination in conversation","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2021","author":"Shuster","year":"2021"},{"issue":"56","key":"2022092118021074700_bib60","first-page":"1929","article-title":"Dropout: A simple way to prevent neural networks from overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"Journal of Machine Learning Research"},{"key":"2022092118021074700_bib61","volume-title":"Describing Talk: A Taxonomy of Verbal Response Modes","author":"Stiles","year":"1992"},{"key":"2022092118021074700_bib62","article-title":"Sticking to the facts: Confident decoding for faithful data-to-text generation","author":"Tian","year":"2019","journal-title":"CoRR (arXiv preprint)"},{"key":"2022092118021074700_bib63","article-title":"Attention is all you need","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani","year":"2017"},{"key":"2022092118021074700_bib64","doi-asserted-by":"publisher","first-page":"5008","DOI":"10.18653\/v1\/2020.acl-main.450","article-title":"Asking and answering questions to evaluate the factual consistency of summaries","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Wang","year":"2020"},{"key":"2022092118021074700_bib65","doi-asserted-by":"publisher","first-page":"3731","DOI":"10.18653\/v1\/P19-1363","article-title":"Dialogue natural language inference","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Welleck","year":"2019"},{"key":"2022092118021074700_bib66","doi-asserted-by":"publisher","first-page":"3731","DOI":"10.18653\/v1\/P19-1363","article-title":"Dialogue natural language inference","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Welleck","year":"2019"},{"key":"2022092118021074700_bib67","doi-asserted-by":"publisher","first-page":"1112","DOI":"10.18653\/v1\/N18-1101","article-title":"A broad-coverage challenge corpus for sentence understanding through inference","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)","author":"Williams","year":"2018"},{"key":"2022092118021074700_bib68","doi-asserted-by":"publisher","first-page":"38","DOI":"10.18653\/v1\/2020.emnlp-demos.6","article-title":"Transformers: State- of-the-art natural language processing","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Wolf","year":"2020"},{"key":"2022092118021074700_bib69","article-title":"Transfertransfo: A transfer learning approach for neural network based conversational agents","author":"Wolf","year":"2019","journal-title":"CoRR (arXiv preprint)"},{"key":"2022092118021074700_bib70","doi-asserted-by":"publisher","first-page":"79","DOI":"10.1145\/3371647.3371659","article-title":"Opportunities and challenges in using AI chatbots in higher education","volume-title":"Proceedings of the 2019 3rd International Conference on Education and E-Learning","author":"Yang","year":"2019"},{"key":"2022092118021074700_bib71","doi-asserted-by":"crossref","first-page":"15","DOI":"10.18653\/v1\/2021.eancs-1.3","article-title":"A comprehensive assessment of dialog evaluation metrics","volume-title":"The First Workshop on Evaluations and Assessments of Neural Conversation Systems","author":"Yeh","year":"2021"},{"key":"2022092118021074700_bib72","first-page":"27263","article-title":"Bartscore: Evaluating generated text as text generation","volume-title":"Advances in Neural Information Processing Systems","author":"Yuan","year":"2021"},{"key":"2022092118021074700_bib73","doi-asserted-by":"publisher","first-page":"2204","DOI":"10.18653\/v1\/P18-1205","article-title":"Personalizing dialogue agents: I have a dog, do you have pets too?","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhang","year":"2018"},{"key":"2022092118021074700_bib74","article-title":"Bertscore: Evaluating text generation with bert","volume-title":"International Conference on Learning Representations","author":"Zhang","year":"2020"},{"key":"2022092118021074700_bib75","doi-asserted-by":"publisher","first-page":"270","DOI":"10.18653\/v1\/2020.acl-demos.30","article-title":"DIALOGPT : Large-scale generative pre-training for conversational response generation","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations","author":"Zhang","year":"2020"},{"key":"2022092118021074700_bib76","doi-asserted-by":"publisher","first-page":"708","DOI":"10.18653\/v1\/D18-1076","article-title":"A dataset for document grounded conversations","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Zhou","year":"2018"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00506\/2043760\/tacl_a_00506.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00506\/2043760\/tacl_a_00506.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,9,21]],"date-time":"2022-09-21T18:03:56Z","timestamp":1663783436000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00506\/113023\/Evaluating-Attribution-in-Dialogue-Systems-The"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022]]},"references-count":76,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00506","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022]]},"published":{"date-parts":[[2022]]}}}