{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,11]],"date-time":"2026-01-11T02:33:18Z","timestamp":1768098798557,"version":"3.49.0"},"reference-count":64,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2023,5,30]],"date-time":"2023-05-30T00:00:00Z","timestamp":1685404800000},"content-version":"vor","delay-in-days":4,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,5,26]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>In this paper, we conduct the first study on spurious correlations for open-domain response generation models based on a corpus CGDialog curated by ourselves. The current models indeed suffer from spurious correlations and have a tendency to generate irrelevant and generic responses. Inspired by causal discovery algorithms, we propose a novel model-agnostic method for training and inference using a conditional independence classifier. The classifier is trained by a constrained self-training method, coined ConSTrain, to overcome data sparsity. The experimental results based on both human and automatic evaluation show that our method significantly outperforms the competitive baselines in terms of relevance, informativeness, and fluency.<\/jats:p>","DOI":"10.1162\/tacl_a_00561","type":"journal-article","created":{"date-parts":[[2023,5,30]],"date-time":"2023-05-30T17:27:01Z","timestamp":1685467621000},"page":"511-530","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":3,"title":["Less is More: Mitigate Spurious Correlations for Open-Domain Dialogue Response Generation Models by Causal Discovery"],"prefix":"10.1162","volume":"11","author":[{"given":"Tao","family":"Feng","sequence":"first","affiliation":[{"name":"Faculty of Information Technology, Monash University, Australia. tao.feng@monash.edu"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lizhen","family":"Qu","sequence":"additional","affiliation":[{"name":"Faculty of Information Technology, Monash University, Australia. lizhen.qu@monash.edu"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gholamreza","family":"Haffari","sequence":"additional","affiliation":[{"name":"Faculty of Information Technology, Monash University, Australia. gholamreza.haffari@monash.edu"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2023,5,26]]},"reference":[{"key":"2023053017265150300_bib1","doi-asserted-by":"publisher","first-page":"pages 941\u2013pages 958","DOI":"10.18653\/v1\/2021.eacl-main.229","article-title":"Filtering noisy dialogue corpora by connectivity and content relatedness","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Akama","year":"2020"},{"key":"2023053017265150300_bib2","first-page":"2662","article-title":"Informative and controllable opinion summarization","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Amplayo","year":"2021"},{"key":"2023053017265150300_bib3","doi-asserted-by":"publisher","first-page":"65","DOI":"10.18653\/v1\/2020.acl-main.9","article-title":"METEOR: An automatic metric for MT evaluation with improved correlation with human judgments","volume-title":"Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization","author":"Banerjee","year":"2005"},{"key":"2023053017265150300_bib4","doi-asserted-by":"crossref","first-page":"85","DOI":"10.18653\/v1\/2020.acl-main.9","article-title":"PLATO: Pre-trained dialogue generation model with discrete latent variable","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Bao","year":"2020"},{"key":"2023053017265150300_bib5","doi-asserted-by":"publisher","first-page":"830","DOI":"10.1609\/icwsm.v14i1.7347","article-title":"The pushshift Reddit dataset","volume-title":"Proceedings of the international AAAI conference on web and social media","author":"Baumgartner","year":"2020"},{"key":"2023053017265150300_bib6","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1907.04068","article-title":"Conditional independence testing using generative adversarial networks","author":"Bellot","year":"2019"},{"key":"2023053017265150300_bib7","article-title":"Comparing rating scales and preference judgements in language evaluation","volume-title":"Proceedings of the 6th International Natural Language Generation Conference","author":"Belz","year":"2010"},{"key":"2023053017265150300_bib8","doi-asserted-by":"publisher","first-page":"131","DOI":"10.18653\/v1\/W16-2301","article-title":"Findings of the 2016 conference on machine translation","volume-title":"Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers","author":"Bojar","year":"2016"},{"issue":"1","key":"2023053017265150300_bib9","doi-asserted-by":"publisher","DOI":"10.3390\/info13010041","article-title":"A literature survey of recent advances in chatbots","volume":"13","author":"Caldarini","year":"2022","journal-title":"Information"},{"key":"2023053017265150300_bib10","doi-asserted-by":"publisher","first-page":"136","DOI":"10.3115\/1626355.1626373","article-title":"(meta-) evaluation of machine translation","volume-title":"Proceedings of the Second Workshop on Statistical Machine Translation","author":"Callison-Burch","year":"2007"},{"key":"2023053017265150300_bib11","doi-asserted-by":"publisher","first-page":"5650","DOI":"10.18653\/v1\/P19-1567","article-title":"Improving neural conversational models with entropy-based data filtering","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Cs\u00e1ky","year":"2019"},{"key":"2023053017265150300_bib12","article-title":"Wizard of Wikipedia: Knowledge-powered conversational agents","author":"Dinan","year":"2019"},{"key":"2023053017265150300_bib13","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-02174-9","volume-title":"Statistical Significance Testing for Natural Language Processing","author":"Dror","year":"2020"},{"key":"2023053017265150300_bib14","doi-asserted-by":"publisher","first-page":"874","DOI":"10.18653\/v1\/2021.eacl-main.74","article-title":"Leveraging passage retrieval with generative models for open domain question answering","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Izacard","year":"2021"},{"key":"2023053017265150300_bib15","article-title":"Adam: A method for stochastic optimization","author":"Kingma","year":"2015"},{"key":"2023053017265150300_bib16","doi-asserted-by":"publisher","first-page":"465","DOI":"10.18653\/v1\/P17-2074","article-title":"Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Kiritchenko","year":"2017"},{"key":"2023053017265150300_bib17","doi-asserted-by":"crossref","first-page":"395","DOI":"10.1109\/BigComp54360.2022.00089","article-title":"Toward robust response selection model for cross negative sampling condition","volume-title":"2022 IEEE International Conference on Big Data and Smart Computing (BigComp)","author":"Lee","year":"2022"},{"key":"2023053017265150300_bib18","first-page":"9459","article-title":"Retrieval-augmented generation for knowledge-intensive nlp tasks","volume":"33","author":"Lewis","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2023053017265150300_bib19","first-page":"110","article-title":"A diversity-promoting objective function for neural conversation models","volume-title":"Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Li","year":"2016"},{"key":"2023053017265150300_bib20","first-page":"986","article-title":"DailyDialog: A manually labelled multi-turn dialogue dataset","volume-title":"Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Li","year":"2017"},{"key":"2023053017265150300_bib21","doi-asserted-by":"publisher","first-page":"128","DOI":"10.18653\/v1\/2021.acl-long.11","article-title":"Conversations are not flat: Modeling the dynamic information flow across dialogue utterances","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Li","year":"2021"},{"key":"2023053017265150300_bib22","doi-asserted-by":"crossref","first-page":"pages 2122\u2013pages 2132","DOI":"10.18653\/v1\/D16-1230","article-title":"How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing","author":"Liu","year":"2016"},{"key":"2023053017265150300_bib23","doi-asserted-by":"publisher","first-page":"3469","DOI":"10.18653\/v1\/2021.acl-long.269","article-title":"Towards emotional support dialog systems","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Liu","year":"2021"},{"key":"2023053017265150300_bib24","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1907.11692","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","author":"Liu","year":"2019"},{"key":"2023053017265150300_bib25","article-title":"Revisiting classifier two-sample tests","volume-title":"International Conference on Learning Representations","author":"Lopez-Paz","year":"2017"},{"key":"2023053017265150300_bib26","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9781107337855","volume-title":"Best-Worst Scaling: Theory, Methods and Applications","author":"Louviere","year":"2015"},{"key":"2023053017265150300_bib27","volume-title":"Introduction to Causal Inference from a Machine Learning Perspective","author":"Neal","year":"2020"},{"key":"2023053017265150300_bib28","doi-asserted-by":"crossref","first-page":"486","DOI":"10.18653\/v1\/K18-1047","article-title":"Adversarial over-sensitivity and over-stability strategies for dialogue models","volume-title":"Proceedings of the 22nd Conference on Computational Natural Language Learning","author":"Niu","year":"2018"},{"issue":"3","key":"2023053017265150300_bib29","doi-asserted-by":"publisher","first-page":"203","DOI":"10.3934\/jdg.20210085","article-title":"Causal discovery in machine learning: Theories and applications","volume":"8","author":"Nogueira","year":"2021","journal-title":"Journal of Dynamics and Games"},{"key":"2023053017265150300_bib30","doi-asserted-by":"publisher","first-page":"72","DOI":"10.18653\/v1\/N18-2012","article-title":"RankME: Reliable human ratings for natural language generation","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers)","author":"Novikova","year":"2018"},{"key":"2023053017265150300_bib31","doi-asserted-by":"publisher","first-page":"311","DOI":"10.3115\/1073083.1073135","article-title":"BLEU: A method for automatic evaluation of machine translation","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni","year":"2002"},{"key":"2023053017265150300_bib32","first-page":"8026","article-title":"PyTorch: An imperative style, high-performance deep learning library","author":"Paszke","year":"2019"},{"key":"2023053017265150300_bib33","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511803161","volume-title":"Causality","author":"Pearl","year":"2009"},{"key":"2023053017265150300_bib34","first-page":"441","article-title":"A theory of inferred causation","volume-title":"Proceedings of the Second International Conference on Principles of Knowledge Representation and Reasoning","author":"Pearl","year":"1991"},{"key":"2023053017265150300_bib35","first-page":"4816","article-title":"Mauve: Measuring the gap between neural text and human text using divergence frontiers","volume":"34","author":"Pillutla","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"5","key":"2023053017265150300_bib36","doi-asserted-by":"publisher","first-page":"1317","DOI":"10.1007\/s12559-021-09925-7","article-title":"Recognizing emotion cause in conversations","volume":"13","author":"Poria","year":"2021","journal-title":"Cognitive Computation"},{"key":"2023053017265150300_bib37","doi-asserted-by":"publisher","first-page":"510","DOI":"10.1162\/tacl_a_00381","article-title":"Data-to-text generation with macro planning","volume":"9","author":"Puduppully","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2023053017265150300_bib38","doi-asserted-by":"publisher","first-page":"pages 3826\u2013pages 3835","DOI":"10.18653\/v1\/P19-1372","article-title":"Are training samples correlated? Learning to generate dialogue responses with multiple references","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Qiu","year":"2019"},{"key":"2023053017265150300_bib39","doi-asserted-by":"publisher","first-page":"5835","DOI":"10.18653\/v1\/2021.naacl-main.466","article-title":"RocketQA: An optimized training approach to dense passage retrieval for open-domain question answering","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Yingqi","year":"2021"},{"key":"2023053017265150300_bib40","doi-asserted-by":"publisher","first-page":"2383","DOI":"10.18653\/v1\/D16-1264","article-title":"SQuAD: 100,000+ questions for machine comprehension of text","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing","author":"Rajpurkar","year":"2016"},{"key":"2023053017265150300_bib41","doi-asserted-by":"publisher","first-page":"5370","DOI":"10.18653\/v1\/P19-1534","article-title":"Towards empathetic open-domain conversation models: A new benchmark and dataset","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Rashkin","year":"2019"},{"key":"2023053017265150300_bib42","doi-asserted-by":"publisher","first-page":"300","DOI":"10.18653\/v1\/2021.eacl-main.24","article-title":"Recipes for building an open-domain chatbot","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Roller","year":"2021"},{"key":"2023053017265150300_bib43","first-page":"8346","article-title":"An investigation of why overparameterization exacerbates spurious correlations","volume-title":"International Conference on Machine Learning","author":"Sagawa","year":"2020"},{"key":"2023053017265150300_bib44","doi-asserted-by":"publisher","first-page":"32","DOI":"10.18653\/v1\/P19-1004","article-title":"Do neural dialog systems use the conversation history effectively? An empirical study","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Sankar","year":"2019"},{"key":"2023053017265150300_bib45","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1806.09708","article-title":"Mimic and classify: A meta-algorithm for conditional independence testing","author":"Sen","year":"2018"},{"key":"2023053017265150300_bib46","article-title":"Model-powered conditional independence test","volume":"30","author":"Sen","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"267","key":"2023053017265150300_bib47","doi-asserted-by":"publisher","first-page":"467","DOI":"10.1080\/01621459.1954.10483515","article-title":"Spurious correlation: A causal interpretation","volume":"49","author":"Simon","year":"1954","journal-title":"Journal of the American Statistical Association"},{"key":"2023053017265150300_bib48","volume-title":"Causation, Prediction, and Search","author":"Spirtes","year":"2000"},{"key":"2023053017265150300_bib49","doi-asserted-by":"publisher","first-page":"1861","DOI":"10.18653\/v1\/2021.eacl-main.160","article-title":"How to evaluate a summarizer: Study design and statistical analysis for manual linguistic quality evaluation","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Steen","year":"2021"},{"key":"2023053017265150300_bib50","article-title":"Unsupervised domain adaptation through self-supervision","author":"Sun","year":"2019"},{"key":"2023053017265150300_bib51","doi-asserted-by":"publisher","first-page":"5680","DOI":"10.18653\/v1\/2021.eacl-main.160","article-title":"Investigating crowdsourcing protocols for evaluating the factual consistency of summaries","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Tang","year":"2022"},{"key":"2023053017265150300_bib52","article-title":"Learning robust models by countering spurious correlations","author":"Wang","year":"2021"},{"key":"2023053017265150300_bib53","doi-asserted-by":"publisher","first-page":"14041","DOI":"10.1609\/aaai.v35i16.17653","volume-title":"Do response selection models really know what\u2019s next? Utterance manipulation strategies for multi-turn response selection","author":"Whang","year":"2021"},{"key":"2023053017265150300_bib54","doi-asserted-by":"publisher","first-page":"38","DOI":"10.18653\/v1\/2020.emnlp-demos.6","article-title":"Transformers: State-of-the-art natural language processing","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Wolf","year":"2020"},{"key":"2023053017265150300_bib55","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1901.08149","article-title":"TransferTransfo: A transfer learning approach for neural network based conversational agents","author":"Wolf","year":"2019"},{"key":"2023053017265150300_bib56","doi-asserted-by":"publisher","first-page":"5180","DOI":"10.18653\/v1\/2022.acl-long.356","article-title":"Beyond goldfish memory: Long-term open-domain conversation","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Jing","year":"2022"},{"issue":"11","key":"2023053017265150300_bib57","doi-asserted-by":"publisher","first-page":"e19735","DOI":"10.2196\/19735","article-title":"Measurement of semantic textual similarity in clinical texts: Comparison of transformer-based models","volume":"8","author":"Xi","year":"2020","journal-title":"JMIR Medical Informatics"},{"key":"2023053017265150300_bib58","doi-asserted-by":"publisher","first-page":"2204","DOI":"10.18653\/v1\/P18-1205","article-title":"Personalizing dialogue agents: I have a dog, do you have pets too?","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhang","year":"2018"},{"key":"2023053017265150300_bib59","article-title":"Bertscore: Evaluating text generation with bert","author":"Zhang*","year":"2020"},{"key":"2023053017265150300_bib60","doi-asserted-by":"publisher","first-page":"270","DOI":"10.18653\/v1\/2020.acl-demos.30","article-title":"DIALOGPT: Large-scale generative pre-training for conversational response generation","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations","author":"Zhang","year":"2020"},{"key":"2023053017265150300_bib61","doi-asserted-by":"publisher","first-page":"813","DOI":"10.18653\/v1\/2021.findings-acl.72","article-title":"CoMAE: A multi-factor hierarchical framework for empathetic response generation","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Zheng","year":"2021"},{"key":"2023053017265150300_bib62","doi-asserted-by":"publisher","first-page":"5808","DOI":"10.18653\/v1\/2022.naacl-main.426","article-title":"Less is more: Learning to refine dialogue history for personalized dialogue generation","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Zhong","year":"2022"},{"key":"2023053017265150300_bib63","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11325","article-title":"Emotional chatting machine: Emotional conversation generation with internal and external memory","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Zhou","year":"2018"},{"key":"2023053017265150300_bib64","doi-asserted-by":"publisher","first-page":"5982","DOI":"10.1109\/ICCV.2019.00608","article-title":"Confidence regularized self-training","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zou","year":"2019"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00561\/2110596\/tacl_a_00561.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00561\/2110596\/tacl_a_00561.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,30]],"date-time":"2023-05-30T17:27:31Z","timestamp":1685467651000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00561\/116161\/Less-is-More-Mitigate-Spurious-Correlations-for"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,26]]},"references-count":64,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00561","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,26]]}}}