{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T10:08:37Z","timestamp":1783591717963,"version":"3.55.0"},"reference-count":95,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2025,2,26]],"date-time":"2025-02-26T00:00:00Z","timestamp":1740528000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Science Foundation of China","doi-asserted-by":"crossref","award":["62276029"],"award-info":[{"award-number":["62276029"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Technical Field Foundation","award":["2023-JCJQ-JJ-0747"],"award-info":[{"award-number":["2023-JCJQ-JJ-0747"]}]},{"name":"CCF-Zhipu.AI Large Model","award":["202217"],"award-info":[{"award-number":["202217"]}]},{"DOI":"10.13039\/501100012236","name":"Beijing Institute of Technology Research Fund Program for Young Scholars","doi-asserted-by":"crossref","award":["6120220261"],"award-info":[{"award-number":["6120220261"]}],"id":[{"id":"10.13039\/501100012236","id-type":"DOI","asserted-by":"crossref"}]},{"name":"CIPSC-SMP-Zhipu Large Model Cross-Disciplinary"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2025,5,31]]},"abstract":"<jats:p>\n            Incorporating explicit personas into dialogue models is critical for generating responses that fulfill specific user needs and preferences, creating a more personalized and engaging interaction. Early works on persona-based dialogue generation directly concatenate the persona descriptions and dialogue history into relatively small pre-trained language models (PLMs) for response generation, which leads to uninformative and inferior results due to the sparse persona information and the limited model generation capabilities. Recently, large language models (LLMs) have shown their surprising capabilities in language generation. Prompting the LLMs with the persona descriptions for role-playing dialogue generation has also achieved promising results. However, deploying LLMs is challenging for practical applications due to their large scale, spurring efforts to distill the generation capabilities into more concise and compact models through teacher-student learning. In this article, we propose an efficient compact\n            <jats:bold>K<\/jats:bold>\n            nowledge-grounded\n            <jats:bold>P<\/jats:bold>\n            ersona-based\n            <jats:bold>D<\/jats:bold>\n            ialogue model enhanced by LLM\n            <jats:bold>D<\/jats:bold>\n            istillation (KPDD). Specifically, first, we propose to enrich the annotated persona descriptions by integrating external knowledge graphs (KGs) with a mixed encoding network, coupled with a mixture of experts (MoE) module for both informative and diverse response generation. The mixed encoding network contains multiple layers of modality interaction operations, enabling information from both modalities propagates to the other. Second, to fully exploit the generation capabilities of LLMs, we turn to the distillation technique to improve the generation capabilities of our model, facilitated by a natural language inference (NLI)-based filtering mechanism to extract high-quality information from LLMs. In addition, we employ a curriculum learning strategy to train our model on the high-quality filtered distilled data and progressively on the relatively noisy original data, enhancing its adaptability and performance. Extensive experiments show that KPDD outperforms state-of-the-art baselines in terms of both automatic and human evaluation.\n          <\/jats:p>","DOI":"10.1145\/3711857","type":"journal-article","created":{"date-parts":[[2025,1,10]],"date-time":"2025-01-10T13:32:07Z","timestamp":1736515927000},"page":"1-29","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["Efficient and Effective Role Player: A Compact Knowledge-grounded Persona-based Dialogue Model Enhanced by LLM Distillation"],"prefix":"10.1145","volume":"43","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0993-8495","authenticated-orcid":false,"given":"Linmei","family":"Hu","sequence":"first","affiliation":[{"name":"School of Computer Science, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4504-7576","authenticated-orcid":false,"given":"Xinyu","family":"Zhang","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7239-6900","authenticated-orcid":false,"given":"Dandan","family":"Song","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3631-3976","authenticated-orcid":false,"given":"Changzhi","family":"Zhou","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7474-5848","authenticated-orcid":false,"given":"Hongyu","family":"He","sequence":"additional","affiliation":[{"name":"Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1476-0273","authenticated-orcid":false,"given":"Liqiang","family":"Nie","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,2,26]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"2787","volume-title":"Proceedings of the 27th Annual Conference on Neural Information Processing Systems","author":"Bordes Antoine","year":"2013","unstructured":"Antoine Bordes, Nicolas Usunier, Alberto Garc\u00eda-Dur\u00e1n, Jason Weston, and Oksana Yakhnenko. 2013. Translating embeddings for modeling multi-relational data. In Proceedings of the 27th Annual Conference on Neural Information Processing Systems, 2787\u20132795. Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2013\/hash\/1cecc7a77928ca8133fa24680a88d2f9-Abstract.html"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6244"},{"key":"e_1_3_2_4_2","first-page":"7984","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics","author":"Cao Yu","year":"2022","unstructured":"Yu Cao, Wei Bi, Meng Fang, Shuming Shi, and Dacheng Tao. 2022. A model-agnostic data manipulation method for persona-based dialogue generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 7984\u20138002. Retrieved from https:\/\/aclanthology.org\/2022.acl-long.550"},{"key":"e_1_3_2_5_2","doi-asserted-by":"crossref","unstructured":"Yu Cao Wei Bi Meng Fang and Dacheng Tao. 2020. Pretrained language models for dialogue generation with multiple input sources. In Findings of the Association for Computational Linguistics (EMNLP \u201920) 909\u2013917. Retrieved from https:\/\/aclanthology.org\/2020.findings-emnlp.81","DOI":"10.18653\/v1\/2020.findings-emnlp.81"},{"key":"e_1_3_2_6_2","unstructured":"Huajun Chen. 2023. Large knowledge model: Perspectives and challenges. arXiv:2312.02706. Retrieved from https:\/\/doi.org\/10.48550\/ARXIV.2312.02706"},{"key":"e_1_3_2_7_2","doi-asserted-by":"crossref","unstructured":"Liang Chen Hongru Wang Yang Deng Wai Chung Kwan Zezhong Wang and Kam-Fai Wong. 2023. Towards robust personalized dialogue generation via order-insensitive representation regularization. In Findings of the Association for Computational Linguistics (ACL \u201923) 7337\u20137345. Retrieved from https:\/\/aclanthology.org\/2023.findings-acl.462","DOI":"10.18653\/v1\/2023.findings-acl.462"},{"key":"e_1_3_2_8_2","first-page":"951","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Chen Maximillian","year":"2023","unstructured":"Maximillian Chen, Xiao Yu, Weiyan Shi, Urvi Awasthi, and Zhou Yu. 2023. Controllable mixed-initiative dialogue generation through prompting. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 951\u2013966. Retrieved from https:\/\/aclanthology.org\/2023.acl-short.82"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i11.26489"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3606368"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1308"},{"key":"e_1_3_2_12_2","unstructured":"Aakanksha Chowdhery Sharan Narang Jacob Devlin Maarten Bosma Gaurav Mishra Adam Roberts Paul Barham Hyung Won Chung Charles Sutton Sebastian Gehrmann et al. 2023. PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research (2023) 240:1\u2013240:113. Retrieved from http:\/\/jmlr.org\/papers\/v24\/22-1144.html"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1111\/j.2517-6161.1977.tb01600.x"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3507782"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3457533"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","unstructured":"Emily Dinan Varvara Logacheva Valentin Malykh Alexander Miller Kurt Shuster Jack Urbanek Douwe Kiela Arthur Szlam Iulian Serban Ryan Lowe et al. 2020. The second conversational intelligence challenge (convai2). In The NeurIPS\u201918 Competition 187\u2013208. Retrieved from https:\/\/link.springer.com\/chapter\/10.1007\/978-3-030-29135-8_7","DOI":"10.1007\/978-3-030-29135-8_7"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.18653\/V1\/2024.ACL-LONG.106"},{"key":"e_1_3_2_18_2","first-page":"146","volume-title":"Proceedings of the 2019 Workshop on Widening NLP","author":"Dziri Nouha","year":"2019","unstructured":"Nouha Dziri, Ehsan Kamalloo, Kory Mathewson, and Osmar Zaiane. 2019. Evaluating coherence in dialogue systems using entailment. In Proceedings of the 2019 Workshop on Widening NLP, 146\u2013148. Retrieved from https:\/\/aclanthology.org\/W19-3646"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.18653\/V1\/2020.EMNLP-MAIN.99"},{"key":"e_1_3_2_20_2","first-page":"10421","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Fu Yao","year":"2023","unstructured":"Yao Fu, Hao Peng, Litu Ou, Ashish Sabharwal, and Tushar Khot. 2023. Specializing smaller language models towards multi-step reasoning. In Proceedings of the International Conference on Machine Learning, 10421\u201310430. Retrieved from https:\/\/proceedings.mlr.press\/v202\/fu23d.html"},{"key":"e_1_3_2_21_2","first-page":"15387","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Gao Jingsheng","year":"2023","unstructured":"Jingsheng Gao, Yixin Lian, Ziyi Zhou, Yuzhuo Fu, and Baoyuan Wang. 2023. LiveChat: A large-scale personalized dialogue dataset automatically constructed from live streaming. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 15387\u201315405. Retrieved from https:\/\/aclanthology.org\/2023.acl-long.858"},{"key":"e_1_3_2_22_2","unstructured":"Geoffrey Hinton Oriol Vinyals and Jeff Dean. 2015. Distilling the knowledge in a neural network. arXiv:1503.02531. Retrieved from https:\/\/arxiv.org\/abs\/1503.02531"},{"key":"e_1_3_2_23_2","unstructured":"Jordan Hoffmann Sebastian Borgeaud Arthur Mensch Elena Buchatskaya Trevor Cai Eliza Rutherford Diego de Las Casas Lisa Anne Hendricks Johannes Welbl Aidan Clark et al. 2022. Training compute-optimal large language models. arXiv:2203.15556. Retrieved from https:\/\/arxiv.org\/abs\/2203.15556"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","unstructured":"Cheng-Yu Hsieh Chun-Liang Li Chih-kuan Yeh Hootan Nakhost Yasuhisa Fujii Alex Ratner Ranjay Krishna Chen-Yu Lee and Tomas Pfister. 2023. Distilling step-by-step! Outperforming larger language models with less training data and smaller model sizes. In Findings of the Association for Computational Linguistics (ACL \u201923) 8003\u20138017. Retrieved from https:\/\/aclanthology.org\/2023.findings-acl.507","DOI":"10.18653\/v1\/2023.findings-acl.507"},{"key":"e_1_3_2_25_2","volume-title":"The Tenth International Conference on Learning Representations (ICLR 2022)","author":"Hu Edward J.","year":"2022","unstructured":"Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Representations (ICLR 2022). OpenReview.net. Retrieved from https:\/\/openreview.net\/forum?id=nZeVKeeFYf9"},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","first-page":"102142","DOI":"10.1016\/j.ipm.2019.102142","article-title":"Graph neural news recommendation with long-term and short-term interest modeling","volume":"57","author":"Hu Linmei","year":"2020","unstructured":"Linmei Hu, Chen Li, Chuan Shi, Cheng Yang, and Chao Shao. 2020. Graph neural news recommendation with long-term and short-term interest modeling. Information Processing and Management 57 (2020), 102142.","journal-title":"Information Processing and Management"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2023.3310002"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1162\/dint_a_00243"},{"key":"e_1_3_2_29_2","doi-asserted-by":"crossref","first-page":"2523","DOI":"10.18653\/v1\/2023.emnlp-main.154","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Huang Qiushi","year":"2023","unstructured":"Qiushi Huang, Shuai Fu, Xubo Liu, Wenwu Wang, Tom Ko, Yu Zhang, and Lilian Tang. 2023. Learning retrieval augmentation for personalized dialogue generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2523\u20132540. Retrieved from https:\/\/aclanthology.org\/2023.emnlp-main.154"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.54"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","unstructured":"Xiaoqi Jiao Yichun Yin Lifeng Shang Xin Jiang Xiao Chen Linlin Li Fang Wang and Qun Liu. 2020. TinyBERT: Distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics 4163\u20134174. Retrieved from https:\/\/aclanthology.org\/2020.findings-emnlp.372","DOI":"10.18653\/v1\/2020.findings-emnlp.372"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.243"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/n16-1014"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3453183"},{"key":"e_1_3_2_36_2","first-page":"2665","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Li Liunian Harold","year":"2023","unstructured":"Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, and Yejin Choi. 2023. Symbolic chain-of-thought distillation: Small models can also \u201cthink\u201d step-by-step. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 2665\u20132679. Retrieved from https:\/\/aclanthology.org\/2023.acl-long.150"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i10.21347"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3498557"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.18653\/V1\/D19-1282"},{"key":"e_1_3_2_40_2","unstructured":"Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out 74\u201381. Retrieved from https:\/\/aclanthology.org\/W04-1013"},{"key":"e_1_3_2_41_2","first-page":"1001","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics","author":"Liu Chang","year":"2022","unstructured":"Chang Liu, Chongyang Tao, Jiazhan Feng, and Dongyan Zhao. 2022. Multi-granularity structural knowledge distillation for language model compression. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 1001\u20131011. Retrieved from https:\/\/aclanthology.org\/2022.acl-long.71"},{"key":"e_1_3_2_42_2","first-page":"8404","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Liu Shuai","year":"2023","unstructured":"Shuai Liu, Hyundong Cho, Marjorie Freedman, Xuezhe Ma, and Jonathan May. 2023. RECAP: Retrieval-enhanced context-aware prefix encoder for personalized dialogue response generation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 8404\u20138419. Retrieved from https:\/\/aclanthology.org\/2023.acl-long.468"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-short.8"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462828"},{"key":"e_1_3_2_45_2","doi-asserted-by":"crossref","first-page":"5454","DOI":"10.18653\/v1\/P19-1542","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Madotto Andrea","year":"2019","unstructured":"Andrea Madotto, Zhaojiang Lin, Chien-Sheng Wu, and Pascale Fung. 2019. Personalizing dialogue agents via meta-learning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 5454\u20135459. Retrieved from https:\/\/aclanthology.org\/P19-1542"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2019.03.010"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.173"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.18653\/V1\/P18-1076"},{"key":"e_1_3_2_49_2","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311\u2013318. Retrieved from https:\/\/aclanthology.org\/P02-1040"},{"key":"e_1_3_2_50_2","doi-asserted-by":"crossref","first-page":"364","DOI":"10.18653\/v1\/2021.emnlp-main.30","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Park Geondo","year":"2021","unstructured":"Geondo Park, Gyeongman Kim, and Eunho Yang. 2021. Distilling linguistic context for language model compression. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 364\u2013378. Retrieved from https:\/\/aclanthology.org\/2021.emnlp-main.30"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/n19-1119"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2018\/595"},{"key":"e_1_3_2_53_2","article-title":"Language models are unsupervised multitask learners","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog (2019).","journal-title":"OpenAI Blog"},{"key":"e_1_3_2_54_2","doi-asserted-by":"crossref","first-page":"2541","DOI":"10.18653\/v1\/2023.emnlp-main.155","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Rawte Vipula","year":"2023","unstructured":"Vipula Rawte, Swagata Chakraborty, Agnibh Pathak, Anubhav Sarkar, S. M. Towhidul Islam Tonmoy, Aman Chadha, Amit Sheth, and Amitava Das. 2023. The troubling emergence of hallucination in large language models - An extensive definition, quantification, and prescriptive remediations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2541\u20132573. Retrieved from https:\/\/aclanthology.org\/2023.emnlp-main.155"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3532077"},{"key":"e_1_3_2_56_2","unstructured":"Victor Sanh Lysandre Debut Julien Chaumond and Thomas Wolf. 2019. DistilBERT a distilled version of BERT: Smaller faster cheaper and lighter. arXiv:1910.01108. Retrieved from https:\/\/arxiv.org\/abs\/1910.01108"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-93417-4_38"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.814"},{"key":"e_1_3_2_59_2","first-page":"5719","volume-title":"Proceedings of the 36th International Conference on Machine Learning","author":"Shen Tianxiao","year":"2019","unstructured":"Tianxiao Shen, Myle Ott, Michael Auli, and Marc\u2019Aurelio Ranzato. 2019. Mixture models for diverse machine translation: Tricks of the trade. In Proceedings of the 36th International Conference on Machine Learning, 5719\u20135728. Retrieved from http:\/\/proceedings.mlr.press\/v97\/shen19c.html"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.14"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.516"},{"key":"e_1_3_2_62_2","first-page":"6651","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing","author":"Song Haoyu","year":"2020","unstructured":"Haoyu Song, Yan Wang, Wei-Nan Zhang, Zhengyu Zhao, Ting Liu, and Xiaojiang Liu. 2020. Profile consistency identification for open-domain dialogue agents. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 6651\u20136662. Retrieved from https:\/\/aclanthology.org\/2020.emnlp-main.539"},{"key":"e_1_3_2_63_2","volume-title":"Proceedings of the 28h International Joint Conference on Artificial Intelligence","author":"Song Haoyu","year":"2019","unstructured":"Haoyu Song, Weinan Zhang, Yiming Cui, Dong Wang, and Ting Liu. 2019. Exploiting persona information for diverse generation of conversational responses. In Proceedings of the 28h International Joint Conference on Artificial Intelligence. Retrieved from https:\/\/www.ijcai.org\/Proceedings\/2019\/0721.pdf"},{"key":"e_1_3_2_64_2","first-page":"8878","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Song Haoyu","year":"2020","unstructured":"Haoyu Song, Wei-Nan Zhang, Jingwen Hu, and Ting Liu. 2020. Generating persona consistent dialogues by exploiting natural language inference. In Proceedings of the AAAI Conference on Artificial Intelligence, 8878\u20138885."},{"key":"e_1_3_2_65_2","first-page":"4323","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing","author":"Sun Siqi","year":"2019","unstructured":"Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019. Patient knowledge distillation for BERT model compression. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, 4323\u20134332. Retrieved from https:\/\/aclanthology.org\/D19-1441"},{"key":"e_1_3_2_66_2","first-page":"498","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing","author":"Sun Siqi","year":"2020","unstructured":"Siqi Sun, Zhe Gan, Yuwei Fang, Yu Cheng, Shuohang Wang, and Jingjing Liu. 2020. Contrastive distillation on intermediate representations for language model compression. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 498\u2013508. Retrieved from https:\/\/aclanthology.org\/2020.emnlp-main.36"},{"key":"e_1_3_2_67_2","doi-asserted-by":"crossref","unstructured":"Chen Tang Chenghua Lin Henglin Huang Frank Guerin and Zhihao Zhang. 2022. EtriCA: Event-triggered context-aware story generation augmented by cross attention. arXiv:2210.12463. Retrieved from https:\/\/arxiv.org\/abs\/2210.12463","DOI":"10.18653\/v1\/2022.findings-emnlp.403"},{"key":"e_1_3_2_68_2","first-page":"4604","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Tang Chen","year":"2023","unstructured":"Chen Tang, Hongbo Zhang, Tyler Loakman, Chenghua Lin, and Frank Guerin. 2023. Enhancing dialogue generation via dynamic graph knowledge aggregation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 4604\u20134616. Retrieved from https:\/\/aclanthology.org\/2023.acl-long.253"},{"key":"e_1_3_2_69_2","first-page":"5456","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Tang Yihong","year":"2023","unstructured":"Yihong Tang, Bo Wang, Miao Fang, Dongming Zhao, Kun Huang, Ruifang He, and Yuexian Hou. 2023. Enhancing personalized dialogue generation with contrastive latent variables: Combining sparse and dense persona. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 5456\u20135468. Retrieved from https:\/\/aclanthology.org\/2023.acl-long.299"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1486"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3531830"},{"key":"e_1_3_2_72_2","unstructured":"Iulia Turc Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. Well-read students learn better: On the importance of pre-training compact models. arXiv:1908.08962. Retrieved from https:\/\/arxiv.org\/abs\/1908.08962"},{"key":"e_1_3_2_73_2","volume-title":"Proceedings of the 8th International Conference on Learning Representations","author":"Vashishth Shikhar","year":"2020","unstructured":"Shikhar Vashishth, Soumya Sanyal, Vikram Nitin, and Partha P. Talukdar. 2020. Composition-based multi-relational graph convolutional networks. In Proceedings of the 8th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=BylA_C4tPr"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.304"},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3539382"},{"key":"e_1_3_2_76_2","unstructured":"Zekun Moore Wang Zhongyuan Peng Haoran Que Jiaheng Liu Wangchunshu Zhou Yuhan Wu Hongcheng Guo Ruitong Gan Zehao Ni Man Zhang et al. 2023. RoleLLM: Benchmarking eliciting and enhancing role-playing abilities of large language models. arXiv.2310.00746. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2310.00746"},{"key":"e_1_3_2_77_2","unstructured":"Jason Wei Yi Tay Rishi Bommasani Colin Raffel Barret Zoph Sebastian Borgeaud Dani Yogatama Maarten Bosma Denny Zhou Donald Metzler et al. 2022. Emergent abilities of large language models. Transactions on Machine Learning Research (2022). Retrieved from https:\/\/openreview.net\/forum?id=yzkSU5zdwD"},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1363"},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.7"},{"key":"e_1_3_2_80_2","first-page":"1956","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics","author":"Wu Yuwei","year":"2021","unstructured":"Yuwei Wu, Xuezhe Ma, and Diyi Yang. 2021. Personalized response generation via generative split memory network. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics, 1956\u20131970. Retrieved from https:\/\/aclanthology.org\/2021.naacl-main.157"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.542"},{"key":"e_1_3_2_82_2","first-page":"32:1","article-title":"HGAT: Heterogeneous graph attention networks for semi-supervised short text classification","volume":"39","author":"Yang Tianchi","year":"2021","unstructured":"Tianchi Yang, Linmei Hu, Chuan Shi, Houye Ji, Xiaoli Li, and Liqiang Nie. 2021. HGAT: Heterogeneous graph attention networks for semi-supervised short text classification. ACM Transactions on Information Systems 39 (2021), 32:1\u201332:29.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_2_83_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSC.2023.3269396"},{"key":"e_1_3_2_84_2","unstructured":"Yizhe Yang Heyan Huang Palakorn Achananuparp Jing Jiang and Ee-Peng Lim. 2024. Speaker verification in agent-generated conversations. arXiv:2405.10150. Retrieved from https:\/\/arxiv.org\/abs\/2405.10150"},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","DOI":"10.18653\/V1\/2022.FINDINGS-ACL.149"},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.1145\/3580305.3599832"},{"key":"e_1_3_2_87_2","doi-asserted-by":"crossref","unstructured":"Qiang Zhang Jason Naradowsky and Yusuke Miyao. 2023. Ask an expert: Leveraging language models to improve strategic reasoning in goal-oriented dialogue models. In Findings of the Association for Computational Linguistics 6665\u20136694. Retrieved from https:\/\/aclanthology.org\/2023.findings-acl.417","DOI":"10.18653\/v1\/2023.findings-acl.417"},{"key":"e_1_3_2_88_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1205"},{"key":"e_1_3_2_89_2","first-page":"1815","volume-title":"Annual Conference on Neural Information Processing Systems","author":"Zhang Yizhe","year":"2018","unstructured":"Yizhe Zhang, Michel Galley, Jianfeng Gao, Zhe Gan, Xiujun Li, Chris Brockett, and Bill Dolan. 2018. Generating informative and diverse conversational responses via adversarial information maximization. In Annual Conference on Neural Information Processing Systems, 1815\u20131825. Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2018\/hash\/23ce1851341ec1fa9e0c259de10bf87c-Abstract.html"},{"key":"e_1_3_2_90_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6518"},{"key":"e_1_3_2_91_2","first-page":"5808","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics","author":"Zhong Hanxun","year":"2022","unstructured":"Hanxun Zhong, Zhicheng Dou, Yutao Zhu, Hongjin Qian, and Ji-Rong Wen. 2022. Less is more: Learning to refine dialogue history for personalized dialogue generation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics, 5808\u20135820. Retrieved from https:\/\/aclanthology.org\/2022.naacl-main.426"},{"key":"e_1_3_2_92_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i5.16602"},{"key":"e_1_3_2_93_2","doi-asserted-by":"crossref","unstructured":"Jinfeng Zhou Zhuang Chen Dazhen Wan Bosi Wen Yi Song Jifan Yu Yongkang Huang Libiao Peng Jiaming Yang Xiyao Xiao et al. 2023. CharacterGLM: Customizing Chinese conversational AI characters with large language models. arXiv:2311.16832. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2311.16832","DOI":"10.18653\/v1\/2024.emnlp-industry.107"},{"key":"e_1_3_2_94_2","first-page":"9945","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Zhou Junkai","year":"2023","unstructured":"Junkai Zhou, Liang Pang, Huawei Shen, and Xueqi Cheng. 2023. SimOAP: Improve coherence and consistency in persona-based dialogue generation via over-sampling and post-evaluation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 9945\u20139959. Retrieved from https:\/\/aclanthology.org\/2023.acl-long.553"},{"key":"e_1_3_2_95_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210080"},{"key":"e_1_3_2_96_2","first-page":"1132","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Zhuang Yimeng","year":"2023","unstructured":"Yimeng Zhuang and Mei Tu. 2023. Pretrained bidirectional distillation for machine translation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 1132\u20131145. Retrieved from https:\/\/aclanthology.org\/2023.acl-long.63"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3711857","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3711857","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:19:15Z","timestamp":1750295955000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3711857"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,26]]},"references-count":95,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,5,31]]}},"alternative-id":["10.1145\/3711857"],"URL":"https:\/\/doi.org\/10.1145\/3711857","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,26]]},"assertion":[{"value":"2024-06-04","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-10","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-26","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}