{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,19]],"date-time":"2026-02-19T03:18:17Z","timestamp":1771471097264,"version":"3.50.1"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2020,11,13]],"date-time":"2020-11-13T00:00:00Z","timestamp":1605225600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["No. 61772076, 61751201 and 61602197"],"award-info":[{"award-number":["No. 61772076, 61751201 and 61602197"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Major Project of Zhijiang Lab","award":["No. 2019DH0ZX01"],"award-info":[{"award-number":["No. 2019DH0ZX01"]}]},{"name":"NSFB","award":["No. Z181100008918002"],"award-info":[{"award-number":["No. Z181100008918002"]}]},{"name":"National Key R&D Plan","award":["No. 2018YFB1005100"],"award-info":[{"award-number":["No. 2018YFB1005100"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2021,1,31]]},"abstract":"<jats:p>\n            Open-domain generative dialogue systems have attracted considerable attention over the past few years. Currently, how to automatically evaluate them is still a big challenge. As far as we know, there are three kinds of automatic evaluations for open-domain generative dialogue systems: (1) Word-overlap-based metrics; (2) Embedding-based metrics; (3) Learning-based metrics. Due to the lack of systematic comparison, it is not clear which kind of metrics is more effective. In this article, we first measure systematically all kinds of metrics to check which kind is best. Extensive experiments demonstrate that learning-based metrics are the most effective evaluation metrics for open-domain generative dialogue systems. Moreover, we observe that nearly all learning-based metrics depend on the negative sampling mechanism, which obtains extremely imbalanced and low-quality samples to train a score model. To address this issue, we propose a novel learning-based metric that significantly improves the correlation with human judgments by using augmented\n            <jats:bold>PO<\/jats:bold>\n            sitive samples and valuable\n            <jats:bold>NE<\/jats:bold>\n            gative samples, called PONE. Extensive experiments demonstrate that PONE significantly outperforms the state-of-the-art learning-based evaluation method. Besides, we have publicly released the codes of our proposed metric and state-of-the-art baselines.\n            <jats:sup>1<\/jats:sup>\n          <\/jats:p>","DOI":"10.1145\/3423168","type":"journal-article","created":{"date-parts":[[2020,11,24]],"date-time":"2020-11-24T21:12:30Z","timestamp":1606252350000},"page":"1-37","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["PONE"],"prefix":"10.1145","volume":"39","author":[{"given":"Tian","family":"Lan","sequence":"first","affiliation":[{"name":"Beijing Institute of Technology, Haidian, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xian-Ling","family":"Mao","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Haidian, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Wei","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoyan","family":"Gao","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Haidian, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Heyan","family":"Huang","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Haidian, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,11,13]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Le","author":"Freitas Adiwardana Daniel De","year":"2020","unstructured":"Daniel De Freitas Adiwardana , Minh-Thang Luong , David R. So , Jamie Hall , Noah Fiedel , Romal Thoppilan , Zi Yang , Apoorv Kulshreshtha , Gaurav Nemade , Yifeng Lu , and Quoc V . Le . 2020 . Towards a human-like open-domain chatbot. ArXiv abs\/2001.09977 (2020). Daniel De Freitas Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. 2020. Towards a human-like open-domain chatbot. ArXiv abs\/2001.09977 (2020)."},{"key":"e_1_2_1_2_1","volume-title":"Neural machine translation by jointly learning to align and translate. CoRR abs\/1409.0473","author":"Bahdanau Dzmitry","year":"2014","unstructured":"Dzmitry Bahdanau , Kyunghyun Cho , and Yoshua Bengio . 2014. Neural machine translation by jointly learning to align and translate. CoRR abs\/1409.0473 ( 2014 ). Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. CoRR abs\/1409.0473 (2014)."},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the IEEvaluation@ACL Conference.","author":"Banerjee Satanjeev","year":"2005","unstructured":"Satanjeev Banerjee and Alon Lavie . 2005 . METEOR: An automatic metric for MT evaluation with improved correlation with human judgments . In Proceedings of the IEEvaluation@ACL Conference. Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the IEEvaluation@ACL Conference."},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201917)","author":"Cai Liwei","year":"2017","unstructured":"Liwei Cai and William Yang Wang . 2017 . KBGAN: Adversarial learning for knowledge graph embeddings . In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201917) . Liwei Cai and William Yang Wang. 2017. KBGAN: Adversarial learning for knowledge graph embeddings. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201917)."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1179"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics (ACL\u201911)","author":"Danescu-Niculescu-Mizil Cristian","year":"2011","unstructured":"Cristian Danescu-Niculescu-Mizil and Lillian Lee . 2011 . Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs . In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics (ACL\u201911) . Cristian Danescu-Niculescu-Mizil and Lillian Lee. 2011. Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs. In Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics (ACL\u201911)."},{"key":"e_1_2_1_7_1","first-page":"19","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","volume":"1","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . BERT: Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , Volume 1 (Long and Short Papers). Association for Computational Linguistics, 4171--4186. DOI:https:\/\/doi.org\/10. 18653\/v1\/N 19 - 1423 Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics, 4171--4186. DOI:https:\/\/doi.org\/10.18653\/v1\/N19-1423"},{"key":"e_1_2_1_8_1","volume-title":"Modern Machine Learning and Natural Language Processing Workshop","volume":"2","author":"Forgues Gabriel","year":"2014","unstructured":"Gabriel Forgues , Joelle Pineau , Jean-Marie Larchev\u00eaque , and R\u00e9al Tremblay . 2014 . Bootstrapping dialog systems with word embeddings. In Nips , Modern Machine Learning and Natural Language Processing Workshop , Vol. 2 . Gabriel Forgues, Joelle Pineau, Jean-Marie Larchev\u00eaque, and R\u00e9al Tremblay. 2014. Bootstrapping dialog systems with word embeddings. In Nips, Modern Machine Learning and Natural Language Processing Workshop, Vol. 2."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1613\/jair.5477"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W19-2310"},{"key":"e_1_2_1_11_1","volume-title":"Predictive engagement: An efficient metric for automatic evaluation of open-domain dialogue systems. arXiv preprint arXiv:1911.01456","author":"Ghazarian Sarik","year":"2019","unstructured":"Sarik Ghazarian , Ralph Weischedel , Aram Galstyan , and Nanyun Peng . 2019. Predictive engagement: An efficient metric for automatic evaluation of open-domain dialogue systems. arXiv preprint arXiv:1911.01456 ( 2019 ). Sarik Ghazarian, Ralph Weischedel, Aram Galstyan, and Nanyun Peng. 2019. Predictive engagement: An efficient metric for automatic evaluation of open-domain dialogue systems. arXiv preprint arXiv:1911.01456 (2019)."},{"key":"e_1_2_1_12_1","first-page":"16","volume-title":"Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 1631--1640","author":"Gu Jiatao","year":"1865","unstructured":"Jiatao Gu , Zhengdong Lu , Hang Li , and Victor O. K. Li . 2016. Incorporating copying mechanism in sequence-to-sequence learning . In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 1631--1640 . DOI:https:\/\/doi.org\/10. 1865 3\/v1\/P 16 - 1154 Jiatao Gu, Zhengdong Lu, Hang Li, and Victor O. K. Li. 2016. Incorporating copying mechanism in sequence-to-sequence learning. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 1631--1640. DOI:https:\/\/doi.org\/10.18653\/v1\/P16-1154"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2014-38"},{"key":"e_1_2_1_14_1","volume-title":"The curious case of neural text degeneration. ArXiv abs\/1904.09751","author":"Holtzman Ari","year":"2020","unstructured":"Ari Holtzman , Jan Buys , Maxwell Forbes , and Yejin Choi . 2020. The curious case of neural text degeneration. ArXiv abs\/1904.09751 ( 2020 ). Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. ArXiv abs\/1904.09751 (2020)."},{"key":"e_1_2_1_15_1","volume-title":"Sequence-to-sequence data augmentation for dialogue language understanding. ArXiv abs\/1807.01554","author":"Hou Yutai","year":"2018","unstructured":"Yutai Hou , Yijia Liu , Wanxiang Che , and Ting Liu . 2018. Sequence-to-sequence data augmentation for dialogue language understanding. ArXiv abs\/1807.01554 ( 2018 ). Yutai Hou, Yijia Liu, Wanxiang Che, and Ting Liu. 2018. Sequence-to-sequence data augmentation for dialogue language understanding. ArXiv abs\/1807.01554 (2018)."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3383123"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-3529"},{"key":"e_1_2_1_18_1","volume-title":"Sentence-level fluency evaluation: References help, but can be spared! ArXiv abs\/1809.08731","author":"Kann Katharina","year":"2018","unstructured":"Katharina Kann , Sascha Rothe , and Katja Filippova . 2018. Sentence-level fluency evaluation: References help, but can be spared! ArXiv abs\/1809.08731 ( 2018 ). Katharina Kann, Sascha Rothe, and Katja Filippova. 2018. Sentence-level fluency evaluation: References help, but can be spared! ArXiv abs\/1809.08731 (2018)."},{"key":"e_1_2_1_19_1","volume-title":"Adversarial evaluation of dialogue models. ArXiv abs\/1701.08198","author":"Kannan Anjuli","year":"2016","unstructured":"Anjuli Kannan and Oriol Vinyals . 2016. Adversarial evaluation of dialogue models. ArXiv abs\/1701.08198 ( 2016 ). Anjuli Kannan and Oriol Vinyals. 2016. Adversarial evaluation of dialogue models. ArXiv abs\/1701.08198 (2016)."},{"key":"e_1_2_1_20_1","volume-title":"Rush","author":"Klein Guillaume","year":"2017","unstructured":"Guillaume Klein , Yoon Kim , Yuntian Deng , Jean Senellart , and Alexander M . Rush . 2017 . OpenNMT: Open-source toolkit for neural machine translation. In Proceedings of the Association for Computational Linguistics Conference (ACL\u2019 17). Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander M. Rush. 2017. OpenNMT: Open-source toolkit for neural machine translation. In Proceedings of the Association for Computational Linguistics Conference (ACL\u201917)."},{"key":"e_1_2_1_21_1","volume-title":"A manually annotated Chinese corpus for non-task-oriented dialogue systems. ArXiv abs\/1805.05542","author":"Li Jing","year":"2018","unstructured":"Jing Li , Yan Song , Haisong Zhang , and Shuming Shi . 2018. A manually annotated Chinese corpus for non-task-oriented dialogue systems. ArXiv abs\/1805.05542 ( 2018 ). Jing Li, Yan Song, Haisong Zhang, and Shuming Shi. 2018. A manually annotated Chinese corpus for non-task-oriented dialogue systems. ArXiv abs\/1805.05542 (2018)."},{"key":"e_1_2_1_22_1","volume-title":"Dice loss for data-imbalanced NLP tasks. ArXiv abs\/1911.02855","author":"Li Xiaoya","year":"2019","unstructured":"Xiaoya Li , Xiaofei Sun , Yuxian Meng , Junjun Liang , Fei Wu , and Jiwei Li. 2019. Dice loss for data-imbalanced NLP tasks. ArXiv abs\/1911.02855 ( 2019 ). Xiaoya Li, Xiaofei Sun, Yuxian Meng, Junjun Liang, Fei Wu, and Jiwei Li. 2019. Dice loss for data-imbalanced NLP tasks. ArXiv abs\/1911.02855 (2019)."},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the International Joint Conference on Natural Language Processing.","author":"Li Yanran","year":"2017","unstructured":"Yanran Li , Hui Su , Xiaoyu Shen , Wenjie Li , Ziqiang Cao , and Shuzi Niu . 2017 . DailyDialog: A manually labelled multi-turn dialogue dataset . In Proceedings of the International Joint Conference on Natural Language Processing. Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. DailyDialog: A manually labelled multi-turn dialogue dataset. In Proceedings of the International Joint Conference on Natural Language Processing."},{"key":"e_1_2_1_24_1","volume-title":"Proceedings of the Association for Computational Linguistics Conference (ACL\u201904)","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin . 2004 . ROUGE: A package for automatic evaluation of summaries . In Proceedings of the Association for Computational Linguistics Conference (ACL\u201904) . Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Proceedings of the Association for Computational Linguistics Conference (ACL\u201904)."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1230"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1103"},{"key":"e_1_2_1_27_1","volume-title":"Efficient estimation of word representations in vector space. CoRR abs\/1301.3781","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov , Kai Chen , Gregory S. Corrado , and Jeffrey Dean . 2013. Efficient estimation of word representations in vector space. CoRR abs\/1301.3781 ( 2013 ). Tomas Mikolov, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. CoRR abs\/1301.3781 (2013)."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201914)","author":"Pennington Jeffrey","unstructured":"Jeffrey Pennington , Richard Socher , and Christopher D. Manning . 2014. GloVe: Global vectors for word representation . In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201914) . Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201914)."},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201911)","author":"Ritter Alan","unstructured":"Alan Ritter , Colin Cherry , and William B. Dolan . 2011. Data-driven response generation in social media . In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201911) . Alan Ritter, Colin Cherry, and William B. Dolan. 2011. Data-driven response generation in social media. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201911)."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1561\/1500000019"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1128"},{"key":"e_1_2_1_33_1","volume-title":"Lintean","author":"Rus Vasile","year":"2012","unstructured":"Vasile Rus and Mihai C . Lintean . 2012 . A comparison of greedy and optimal assessment of natural language student input using word-to-word similarity metrics. In Proceedings of the Building Educational Applications Workshop at the Conference of the North American Chapter of the Association for Computational Linguistics : Human Language Technologies (BEA@NAACL-HLT\u201912). Vasile Rus and Mihai C. Lintean. 2012. A comparison of greedy and optimal assessment of natural language student input using word-to-word similarity metrics. In Proceedings of the Building Educational Applications Workshop at the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (BEA@NAACL-HLT\u201912)."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.55"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-1009"},{"key":"e_1_2_1_36_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201915)","author":"Serban Iulian","year":"2015","unstructured":"Iulian Serban , Alessandro Sordoni , Yoshua Bengio , Aaron C. Courville , and Joelle Pineau . 2015 . Building end-to-end dialogue systems using generative hierarchical neural network models . In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201915) . Iulian Serban, Alessandro Sordoni, Yoshua Bengio, Aaron C. Courville, and Joelle Pineau. 2015. Building end-to-end dialogue systems using generative hierarchical neural network models. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201915)."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11321"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2018\/616"},{"key":"e_1_2_1_39_1","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"van der Maaten Laurens","year":"2008","unstructured":"Laurens van der Maaten and Geoffrey E. Hinton . 2008 . Visualizing data using t-SNE . Journal of Machine Learning Research 9 (2008), 2579 -- 2605 . Laurens van der Maaten and Geoffrey E. Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9 (2008), 2579--2605.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_1_40_1","volume-title":"Proceedings of the Conference on Neural Information Processing Systems.","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N. Gomez , Lukasz Kaiser , and Illia Polosukhin . 2017 . Attention is all you need . In Proceedings of the Conference on Neural Information Processing Systems. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Conference on Neural Information Processing Systems."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210061"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1670"},{"key":"e_1_2_1_43_1","volume-title":"Domain adaptive training BERT for response selection. ArXiv abs\/1908.04812","author":"Whang Taesun","year":"2019","unstructured":"Taesun Whang , Dongyub Lee , Chanhee Lee , Kisu Yang , Dongsuk Oh , and Heuiseok Lim . 2019. Domain adaptive training BERT for response selection. ArXiv abs\/1908.04812 ( 2019 ). Taesun Whang, Dongyub Lee, Chanhee Lee, Kisu Yang, Dongsuk Oh, and Heuiseok Lim. 2019. Domain adaptive training BERT for response selection. ArXiv abs\/1908.04812 (2019)."},{"key":"e_1_2_1_44_1","volume-title":"Towards universal paraphrastic sentence embeddings. CoRR abs\/1511.08198","author":"Wieting John","year":"2015","unstructured":"John Wieting , Mohit Bansal , Kevin Gimpel , and Karen Livescu . 2015. Towards universal paraphrastic sentence embeddings. CoRR abs\/1511.08198 ( 2015 ). John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu. 2015. Towards universal paraphrastic sentence embeddings. CoRR abs\/1511.08198 (2015)."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00345"},{"key":"e_1_2_1_46_1","unstructured":"Han Xiao. 2018. Bert-as-service. Retrieved from https:\/\/github.com\/hanxiao\/bert-as-service.  Han Xiao. 2018. Bert-as-service. Retrieved from https:\/\/github.com\/hanxiao\/bert-as-service."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1233"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1362"},{"key":"e_1_2_1_49_1","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Zhang Tianyi","year":"2020","unstructured":"Tianyi Zhang , Varsha Kishore , Felix Wu , Kilian Q. Weinberger , and Yoav Artzi . 2020 . BERTScore: Evaluating text generation with BERT . In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id&equals;SkeHuCVFDr. Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating text generation with BERT. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id&equals;SkeHuCVFDr."},{"key":"e_1_2_1_50_1","volume-title":"Proceedings of the Conference on Neural Information Processing Systems (NIPS\u201915)","author":"Zhang Xiang","year":"2015","unstructured":"Xiang Zhang , Junbo Jake Zhao , and Yann LeCun . 2015 . Character-level convolutional networks for text classification . In Proceedings of the Conference on Neural Information Processing Systems (NIPS\u201915) . Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Proceedings of the Conference on Neural Information Processing Systems (NIPS\u201915)."},{"key":"e_1_2_1_51_1","volume-title":"Dolan","author":"Zhang Yizhe","year":"2019","unstructured":"Yizhe Zhang , Siqi Sun , Michel Galley , Yen-Chun Chen , Chris Brockett , Xiang Gao , Jianfeng Gao , Jingjing Liu , and William B . Dolan . 2019 . DialoGPT: Large-scale generative pre-training for conversational response generation. ArXiv abs\/1911.00536 (2019). Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and William B. Dolan. 2019. DialoGPT: Large-scale generative pre-training for conversational response generation. ArXiv abs\/1911.00536 (2019)."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2018\/643"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3423168","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3423168","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:24:57Z","timestamp":1750195497000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3423168"}},"subtitle":["A Novel Automatic Evaluation Metric for Open-domain Generative Dialogue Systems"],"short-title":[],"issued":{"date-parts":[[2020,11,13]]},"references-count":52,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,1,31]]}},"alternative-id":["10.1145\/3423168"],"URL":"https:\/\/doi.org\/10.1145\/3423168","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,11,13]]},"assertion":[{"value":"2020-05-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-11-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}