{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T05:10:46Z","timestamp":1784178646390,"version":"3.55.0"},"reference-count":71,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2023,8,18]],"date-time":"2023-08-18T00:00:00Z","timestamp":1692316800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2024,1,31]]},"abstract":"<jats:p>Task-oriented dialogue systems (TDSs) are assessed mainly in an offline setting or through human evaluation. The evaluation is often limited to single-turn or is very time-intensive. As an alternative, user simulators that mimic user behavior allow us to consider a broad set of user goals to generate human-like conversations for simulated evaluation. Employing existing user simulators to evaluate TDSs is challenging as user simulators are primarily designed to optimize dialogue policies for TDSs and have limited evaluation capabilities. Moreover, the evaluation of user simulators is an open challenge.<\/jats:p><jats:p>In this work, we propose a metaphorical user simulator for end-to-end TDS evaluation, where we define a simulator to be metaphorical if it simulates a user\u2019s analogical thinking in interactions with systems. We also propose a tester-based evaluation framework to generate variants, i.e., dialogue systems with different capabilities. Our user simulator constructs a metaphorical user model that assists the simulator in reasoning by referring to prior knowledge when encountering new items. We estimate the quality of simulators by checking the simulated interactions between simulators and variants. Our experiments are conducted using three TDS datasets. The proposed user simulator demonstrates better consistency with manual evaluation than an agenda-based simulator and a seq2seq model on three datasets; our tester framework demonstrates efficiency and has been tested on multiple tasks, such as conversational recommendation and e-commerce dialogues.<\/jats:p>","DOI":"10.1145\/3596510","type":"journal-article","created":{"date-parts":[[2023,5,22]],"date-time":"2023-05-22T12:03:43Z","timestamp":1684757023000},"page":"1-29","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Metaphorical User Simulators for Evaluating Task-oriented Dialogue Systems"],"prefix":"10.1145","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4817-9500","authenticated-orcid":false,"given":"Weiwei","family":"Sun","sequence":"first","affiliation":[{"name":"Shandong University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-8596-9747","authenticated-orcid":false,"given":"Shuyu","family":"Guo","sequence":"additional","affiliation":[{"name":"Shandong University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3179-4125","authenticated-orcid":false,"given":"Shuo","family":"Zhang","sequence":"additional","affiliation":[{"name":"Bloomberg, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2964-6422","authenticated-orcid":false,"given":"Pengjie","family":"Ren","sequence":"additional","affiliation":[{"name":"Shandong University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4592-4074","authenticated-orcid":false,"given":"Zhumin","family":"Chen","sequence":"additional","affiliation":[{"name":"Shandong University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1086-0202","authenticated-orcid":false,"given":"Maarten","family":"de Rijke","sequence":"additional","affiliation":[{"name":"University of Amsterdam, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9076-6565","authenticated-orcid":false,"given":"Zhaochun","family":"Ren","sequence":"additional","affiliation":[{"name":"Shandong University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,8,18]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-76298-0_52"},{"key":"e_1_3_2_3_2","unstructured":"Krisztian Balog. 2021. Conversational AI from an information retrieval perspective: Remaining challenges and a case for user simulation. In Proceedings of the Second International Conference on Design of Experimental Search & Information REtrieval Systems (DESIRES\u201921) Vol. 2950. 80\u201390."},{"key":"e_1_3_2_4_2","doi-asserted-by":"crossref","unstructured":"Krisztian Balog David Maxwell Paul Thomas and Shuo Zhang. 2021. Sim4IR: The SIGIR 2021 workshop on simulation for information retrieval evaluation. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201921) . 2697\u20132698.","DOI":"10.1145\/3404835.3462821"},{"key":"e_1_3_2_5_2","unstructured":"Alan W. Black Susanne Burger Alistair Conkie H. Hastie Simon Keizer Oliver Lemon Nicolas Merigaud Gabriel Parent Gabriel Schubiner Blaise Thomson J. Williams Kai Yu Steve J. Young and Maxine Esk\u00e9nazi. 2011. Spoken dialog challenge 2010: Comparison of live and control test results. In Proceedings of the SIGDIAL 2011 Conference (SIGDIAL\u201911) . 2\u20137."},{"key":"e_1_3_2_6_2","unstructured":"Dimitrios Bountouridis Jaron Harambam Mykola Makhortykh M\u00f3nica Marrero Nava Tintarev and Claudia Hauff. 2019. SIREN: A simulation framework for understanding the effects of recommender systems in online news environments. In Proceedings of the Conference on Fairness Accountability and Transparency (FAT*\u201919) . 150\u2013159."},{"key":"e_1_3_2_7_2","doi-asserted-by":"crossref","unstructured":"Ben Carterette Evangelos Kanoulas and Emine Yilmaz. 2011. Simulating simple user behavior for system effectiveness evaluation. In Proceedings of the 20th ACM International Conference on Information and Knowledge Management (CIKM\u201911) . 611\u2013620.","DOI":"10.1145\/2063576.2063668"},{"key":"e_1_3_2_8_2","unstructured":"Meng Chen Ruixue Liu Lei Shen Shaozu Yuan Jingyan Zhou Youzheng Wu Xiaodong He and Bowen Zhou. 2020. The JDDC Corpus: A large-scale multi-turn chinese dialogue dataset for e-commerce customer service. In Proceedings of the Twelfth Language Resources and Evaluation Conference (LREC\u201920) . 459\u2013466."},{"key":"e_1_3_2_9_2","first-page":"290","article-title":"Human-computer dialogue simulation using hidden markov models","author":"Cuay\u00e1huitl Heriberto","year":"2005","unstructured":"Heriberto Cuay\u00e1huitl, Steve Renals, Oliver Lemon, and Hiroshi Shimodaira. 2005. Human-computer dialogue simulation using hidden markov models. In IEEE Workshop on ASRU 2005, 290\u2013295.","journal-title":"IEEE Workshop on ASRU 2005"},{"key":"e_1_3_2_10_2","volume-title":"Proceedings of the 1st Workshop on Semiparametric Methods in NLP: Decoupling Logic from Knowledge","author":"Das Rajarshi","year":"2022","unstructured":"Rajarshi Das, Patrick Lewis, Sewon Min, June Thai, and Manzil Zaheer (Eds.). 2022. Proceedings of the 1st Workshop on Semiparametric Methods in NLP: Decoupling Logic from Knowledge."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-020-09866-x"},{"key":"e_1_3_2_12_2","first-page":"80","volume-title":"IEEE Workshop on ASRU 1997","author":"Eckert Wieland","year":"1997","unstructured":"Wieland Eckert, Esther Levin, and Roberto Pieraccini. 1997. User modeling for spoken dialogue system evaluation. In IEEE Workshop on ASRU 1997. 80\u201387."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2008.04.002"},{"key":"e_1_3_2_14_2","unstructured":"Mihail Eric Rahul Goel Shachi Paul Adarsh Kumar Abhishek Sethi Anuj Kumar Goyal Peter Ku Sanchit Agarwal and Shuyang Gao. 2020. MultiWOZ 2.1: A consolidated multi-domain dialogue dataset with state corrections and state tracking baselines. In Proceedings of the Twelfth Language Resources and Evaluation Conference (LREC\u201920) . 422\u2013428."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.aiopen.2021.06.002"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","unstructured":"Kallirroi Georgila James Henderson and Oliver Lemon. 2005. Learning user simulations for information state update dialogue systems. In INTERSPEECH 2005 . 893\u2013896.","DOI":"10.21437\/Interspeech.2005-401"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","unstructured":"Ting Han Ximing Liu Ryuichi Takanobu Yixin Lian Chongxuan Huang Dazhen Wan Wei Peng and Minlie Huang. 2021. MultiWOZ 2.3: A multi-domain task-oriented dialogue dataset enhanced with annotation corrections and co-reference annotation. In Natural Language Processing and Chinese Computing: 10th CCF International Conference (NLPCC\u201921) . 206\u2013218.","DOI":"10.1007\/978-3-030-88483-3_16"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","unstructured":"Wanwei He Yinpei Dai Yinhe Zheng Yuchuan Wu Zheng Cao Dermot Liu Peng Jiang Min Yang Fei Huang Luo Si Jian Sun and Yongbin Li. 2022. GALAXY: A generative pre-trained model for task-oriented dialog with semi-supervised learning and explicit policy injection. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201922) Vol. 36. 10749\u201310757.","DOI":"10.1609\/aaai.v36i10.21320"},{"key":"e_1_3_2_19_2","article-title":"A simple language model for task-oriented dialogue","volume":"2005","author":"Hosseini-Asl Ehsan","year":"2020","unstructured":"Ehsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz, and Richard Socher. 2020. A simple language model for task-oriented dialogue. ArXiv abs\/2005.00796 (2020).","journal-title":"ArXiv"},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","unstructured":"Jin Huang Harrie Oosterhuis Maarten de Rijke and Herke van Hoof. 2020. Keeping dataset biases out of the simulation: A debiased simulator for reinforcement learning based recommender systems. In Proceedings of the 14th ACM Conference on Recommender Systems (RecSys\u201920) . 190\u2013199.","DOI":"10.1145\/3383313.3412252"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3453154"},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","unstructured":"Xisen Jin Wenqiang Lei Zhaochun Ren Hongshen Chen Shangsong Liang Yihong Eric Zhao and Dawei Yin. 2018. Explicit state tracking with semi-supervisionfor neural dialogue generation. In Proceedings of the 27th ACM International Conference on Information and Knowledge Management (CIKM\u201918) . 1403\u20131412.","DOI":"10.1145\/3269206.3271683"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.5751\/ES-03802-160146"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","unstructured":"Filip Jurc\u00edcek Simon Keizer Milica Gasic Fran\u00e7ois Mairesse Blaise Thomson Kai Yu and Steve J. Young. 2011. Real user evaluation of spoken dialogue systems using amazon mechanical turk. In INTERSPEECH 2011 . 3061\u20133064.","DOI":"10.21437\/Interspeech.2011-766"},{"key":"e_1_3_2_25_2","volume-title":"Metaphor in Conversation","author":"Kaal Anna","year":"2012","unstructured":"Anna Kaal. 2012. Metaphor in Conversation. Ph.D. Dissertation. Vrije Universiteit Amsterdam."},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","unstructured":"Sahiti Labhishetty and ChengXiang Zhai. 2021. An exploration of tester-based evaluation of user simulators for comparing interactive retrieval systems. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201921) . 1598\u20131602.","DOI":"10.1145\/3404835.3463091"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-6393(99)00067-9"},{"key":"e_1_3_2_28_2","unstructured":"Wenqiang Lei Xiangnan He Yisong Miao Qingyun Wu Richang Hong Min-Yen Kan and Tat-Seng Chua. 2020. Estimation-action-reflection: Towards deep interaction between conversational and recommender systems. In Proceedings of the 13th International Conference on Web Search and Data Mining (WSDM\u201920) . 304\u2013312."},{"key":"e_1_3_2_29_2","unstructured":"Wenqiang Lei Xisen Jin Min-Yen Kan Zhaochun Ren Xiangnan He and Dawei Yin. 2018. Sequicity: Simplifying task-oriented dialogue systems with single sequence-to-sequence architectures. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL\u201918) . 1437\u20131447."},{"key":"e_1_3_2_30_2","unstructured":"Jiwei Li Michel Galley Chris Brockett Jianfeng Gao and Bill Dolan. 2016. A diversity-promoting objective function for neural conversation models. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL\u201916) . 110\u2013119."},{"key":"e_1_3_2_31_2","article-title":"Towards deep conversational recommendations","volume":"1812","author":"Li Raymond","year":"2018","unstructured":"Raymond Li, Samira Ebrahimi Kahou, Hannes Schulz, Vincent Michalski, Laurent Charlin, and Christopher Joseph Pal. 2018. Towards deep conversational recommendations. ArXiv abs\/1812.07617 (2018).","journal-title":"ArXiv"},{"key":"e_1_3_2_32_2","doi-asserted-by":"crossref","unstructured":"Weixin Liang Youzhi Tian Chengcai Chen and Zhou Yu. 2020. MOSS: End-to-end dialog system framework with modular supervision. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201920) Vol. 34 8327\u20138335.","DOI":"10.1609\/aaai.v34i05.6349"},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","unstructured":"Zhaojiang Lin Bing Liu Seungwhan Moon Paul A. Crook Zhenpeng Zhou Zhiguang Wang Zhou Yu Andrea Madotto Eunjoon Cho and Rajen Subba. 2021. Leveraging slot descriptions for zero-shot cross-domain dialogue StateTracking. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL\u201921) . 5640\u20135648.","DOI":"10.18653\/v1\/2021.naacl-main.448"},{"key":"e_1_3_2_34_2","unstructured":"Wenchang Ma Ryuichi Takanobu and Minlie Huang. 2021. CR-walker: Tree-structured graph reasoning and dialog acts for conversational recommendation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201921) . 1839\u20131851."},{"key":"e_1_3_2_35_2","article-title":"Language models as few-shot learner for task-oriented dialogue systems","volume":"2008","author":"Madotto Andrea","year":"2020","unstructured":"Andrea Madotto and Zihan Liu. 2020. Language models as few-shot learner for task-oriented dialogue systems. ArXiv abs\/2008.06239 (2020).","journal-title":"ArXiv"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","unstructured":"David Maxwell and Leif Azzopardi. 2016. Agents simulated users and humans: An analysis of performance and behaviour. In Proceedings of the 25th ACM International on Conference on Information and Knowledge Management (CIKM\u201916) . 731\u2013740.","DOI":"10.1145\/2983323.2983805"},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","unstructured":"Nikola Mrksic Diarmuid \u00d3. S\u00e9aghdha Blaise Thomson Milica Gasic Pei hao Su David Vandyke Tsung-Hsien Wen and Steve J. Young. 2015. Multi-domain dialog state tracking using recurrent neural networks. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (ACL\u201915) . 794\u2013799.","DOI":"10.3115\/v1\/P15-2130"},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","unstructured":"Nikola Mrksic Diarmuid \u00d3. S\u00e9aghdha Tsung-Hsien Wen Blaise Thomson and Steve J. Young. 2017. Neural belief tracker: Data-driven dialogue state tracking. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL\u201917) . 1777\u20131788.","DOI":"10.18653\/v1\/P17-1163"},{"key":"e_1_3_2_39_2","first-page":"34","volume-title":"Workshop on GEM 2021","author":"Nekvinda Tom\u00e1s","year":"2021","unstructured":"Tom\u00e1s Nekvinda and Ondrej Dusek. 2021. Shades of BLEU, flavours of success: The case of MultiWOZ. In Workshop on GEM 2021. 34\u201346."},{"key":"e_1_3_2_40_2","doi-asserted-by":"crossref","unstructured":"Rodrigo Nogueira Zhiying Jiang Ronak Pradeep and Jimmy J. Lin. 2020. Document ranking with a pretrained sequence-to-sequence model. In Findings of the Association for Computational Linguistics: (EMNLP\u201920) . 708\u2013718.","DOI":"10.18653\/v1\/2020.findings-emnlp.63"},{"key":"e_1_3_2_41_2","doi-asserted-by":"crossref","unstructured":"Alexandros Papangelis Yi-Chia Wang Piero Molino and G\u00f6khan T\u00fcr. 2019. Collaborative multi-agent dialogue model training via reinforcement learning. In Proceedings of the 20th Annual SIGdial Meeting on Discourse and Dialogue (SIGdial\u201919) . 92\u2013102.","DOI":"10.18653\/v1\/W19-5912"},{"key":"e_1_3_2_42_2","doi-asserted-by":"crossref","unstructured":"Kishore Papineni Salim Roukos Todd Ward and Wei-Jing Zhu. 2002. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL\u201902) . 311\u2013318.","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_2_43_2","doi-asserted-by":"crossref","unstructured":"Surendra Pathak Sheikh Ariful Islam Honglu Jiang Lei Xu and Emmett Tomai. 2022. A survey on security analysis of amazon echo devices. High-Confidence Computing 2 (2022) 100087.","DOI":"10.1016\/j.hcc.2022.100087"},{"key":"e_1_3_2_44_2","doi-asserted-by":"crossref","unstructured":"Baolin Peng Chunyuan Li Jinchao Li Shahin Shayandeh Lars Lid\u00e9n and Jianfeng Gao. 2021. SOLOIST: Building task bots at scale with transfer learning and machine teaching. Transactions of the Association for Computational Linguistics 9 (2021) 807\u2013824.","DOI":"10.1162\/tacl_a_00399"},{"key":"e_1_3_2_45_2","doi-asserted-by":"crossref","unstructured":"Fabio Petroni Aleksandra Piktus Angela Fan Patrick Lewis Majid Yazdani Nicola De Cao James Thorne Yacine Jernite Vassilis Plachouras Tim Rocktaschel and Sebastian Riedel. 2021. KILT: A benchmark for knowledge intensive language tasks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL\u201921) . 2523\u20132544.","DOI":"10.18653\/v1\/2021.naacl-main.200"},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","unstructured":"Olivier Pietquin and Helen Hastie. 2012. A survey on metrics for the evaluation of user simulations. The Knowledge Engineering Review 28 (2012) 59\u201373.","DOI":"10.1017\/S0269888912000343"},{"key":"e_1_3_2_47_2","first-page":"9","volume-title":"OpenAI blog","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. In OpenAI blog. 9."},{"key":"e_1_3_2_48_2","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR 21 (2020), 1\u201367.","journal-title":"JMLR"},{"key":"e_1_3_2_49_2","doi-asserted-by":"crossref","unstructured":"Siva Reddy Danqi Chen and Christopher D. Manning. 2019. CoQA: A conversational question answering challenge. Transactions of the Association for Computational Linguistics 7 (2019) 249\u2013266.","DOI":"10.1162\/tacl_a_00266"},{"key":"e_1_3_2_50_2","unstructured":"Pengjie Ren Zhongkun Liu Xiaomeng Song Hongtao Tian Zhumin Chen Zhaochun Ren and Maarten de Rijke. 2021. Wizard of search engine: Access to information through conversations with search engines. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201921) . 533\u2013543."},{"key":"e_1_3_2_51_2","unstructured":"Zhaochun Ren Zhi Tian Dongdong Li Pengjie Ren Liu Yang Xin Xin Huasheng Liang Maarten de Rijke and Zhumin Chen. 2022. Variational reasoning about user preferences for conversational recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201922) . 165\u2013175."},{"key":"e_1_3_2_52_2","unstructured":"Lina Maria Rojas-Barahona Milica Gasi\u0107 Nikola Mrksic Pei hao Su Stefan Ultes Tsung-Hsien Wen Steve J. Young and David Vandyke. 2017. A network-based end-to-end trainable task-oriented dialogue system. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1 Long Papers (EACL\u201917) . 438\u2013449."},{"key":"e_1_3_2_53_2","doi-asserted-by":"crossref","unstructured":"Jost Schatzmann Blaise Thomson Karl Weilhammer Hui Ye and Steve J. Young. 2007. Agenda-based user simulation for bootstrapping a POMDP dialogue system. In Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; Companion Volume Short Papers (NAACL\u201907) . 149\u2013152.","DOI":"10.3115\/1614108.1614146"},{"key":"e_1_3_2_54_2","unstructured":"Weiyan Shi Kun Qian Xuewei Wang and Zhou Yu. 2019. How to build user simulators to train RL-based dialog systems. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP\u201919) . 1990\u20132000."},{"key":"e_1_3_2_55_2","article-title":"Human evaluation of conversations is an open problem: Comparing the sensitivity of various methods for evaluating dialogue agents","volume":"2201","author":"Smith Eric Michael","year":"2022","unstructured":"Eric Michael Smith, Orion Hsu, Rebecca Qian, Stephen Roller, Y-Lan Boureau, and Jason Weston. 2022. Human evaluation of conversations is an open problem: Comparing the sensitivity of various methods for evaluating dialogue agents. ArXiv abs\/2201.04723 (2022).","journal-title":"ArXiv"},{"key":"e_1_3_2_56_2","article-title":"Multi-task pre-training for plug-and-play task-oriented dialogue system","volume":"2109","author":"Su Yixuan","year":"2021","unstructured":"Yixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta, Deng Cai, Yi-An Lai, and Yi Zhang. 2021. Multi-task pre-training for plug-and-play task-oriented dialogue system. ArXiv abs\/2109.14739 (2021).","journal-title":"ArXiv"},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","unstructured":"Weiwei Sun Shuo Zhang Krisztian Balog Zhaochun Ren Pengjie Ren Zhumin Chen and Maarten de Rijke. 2021. Simulating user satisfaction for the evaluation of task-oriented dialogue systems. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201921) . 2499\u20132506.","DOI":"10.1145\/3404835.3463241"},{"key":"e_1_3_2_58_2","unstructured":"Yueming Sun and Yi Zhang. 2018. Conversational recommender system. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval (SIGIR\u201918) . 235\u2013244."},{"key":"e_1_3_2_59_2","article-title":"Transferable dialogue systems and user simulators","volume":"2107","author":"Tseng Bo-Hsiang","year":"2021","unstructured":"Bo-Hsiang Tseng, Yinpei Dai, Florian Kreyssig, and Bill Byrne. 2021. Transferable dialogue systems and user simulators. ArXiv abs\/2107.11904 (2021).","journal-title":"ArXiv"},{"key":"e_1_3_2_60_2","article-title":"A neural conversational model","volume":"1506","author":"Vinyals Oriol","year":"2015","unstructured":"Oriol Vinyals and Quoc Le. 2015. A neural conversational model. ArXiv abs\/1506.05869 (2015).","journal-title":"ArXiv"},{"key":"e_1_3_2_61_2","doi-asserted-by":"crossref","unstructured":"Marilyn A. Walker Diane J. Litman Candace A. Kamm and Alicia Abella. 1997. PARADISE: A framework for evaluating spoken dialogue agents. In Proceedings of the 35th Annual Meeting of the Association for Computational Linguistics and Eighth Conference of the European Chapter of the Association for Computational Linguistics (ACL\u201997) . 271\u2013280.","DOI":"10.3115\/976909.979652"},{"key":"e_1_3_2_62_2","doi-asserted-by":"crossref","unstructured":"Yunyi Yang Yunhao Li and Xiaojun Quan. 2021. UBAR: Towards fully end-to-end task-oriented dialog systems with GPT-2. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201921) Vol. 35. 14230\u201314238.","DOI":"10.1609\/aaai.v35i16.17674"},{"key":"e_1_3_2_63_2","first-page":"1160","volume-title":"Proceedings of the IEEE","author":"Young Steve","year":"2013","unstructured":"Steve Young, Milica Ga\u0161i\u0107, Blaise Thomson, and Jason D. Williams. 2013. POMDP-based statistical spoken dialog systems: A review. In Proceedings of the IEEE, Vol. 101. 1160\u20131179."},{"key":"e_1_3_2_64_2","doi-asserted-by":"crossref","unstructured":"Shuo Zhang and Krisztian Balog. 2020. Evaluating conversational recommender systems via user simulation. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD\u201920) . 1512\u20131520.","DOI":"10.1145\/3394486.3403202"},{"key":"e_1_3_2_65_2","doi-asserted-by":"crossref","unstructured":"Shuo Zhang Mu-Chun Wang and Krisztian Balog. 2022. Analyzing and simulating user utterance reformulation in conversational recommender systems. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201922) . 133\u2013143.","DOI":"10.1145\/3477495.3531936"},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","unstructured":"Yongfeng Zhang Xu Chen Qingyao Ai Liu Yang and W. Bruce Croft. 2018. Towards conversational search and recommendation: System ask user respond. In Proceedings of the 27th ACM International Conference on Information and Knowledge Management (CIKM\u201918) . 177\u2013186.","DOI":"10.1145\/3269206.3271776"},{"key":"e_1_3_2_67_2","doi-asserted-by":"crossref","unstructured":"Yinan Zhang Xueqing Liu and ChengXiang Zhai. 2017. Information retrieval evaluation as search simulation: A general formal framework for IR evaluation. In Proceedings of the ACM SIGIR International Conference on Theory of Information Retrieval (ICTIR\u201917) . 193\u2013200.","DOI":"10.1145\/3121050.3121070"},{"key":"e_1_3_2_68_2","doi-asserted-by":"crossref","unstructured":"Yichi Zhang Zhijian Ou and Zhou Yu. 2020. Task-oriented dialog systems that consider multiple appropriate responses under the same context. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201920) Vol. 34. 9604\u20139611.","DOI":"10.1609\/aaai.v34i05.6507"},{"key":"e_1_3_2_69_2","article-title":"Recent advances and challenges in task-oriented dialog system","volume":"2003","author":"Zhang Zheng","year":"2020","unstructured":"Zheng Zhang, Ryuichi Takanobu, Minlie Huang, and Xiaoyan Zhu. 2020. Recent advances and challenges in task-oriented dialog system. ArXiv abs\/2003.07490 (2020).","journal-title":"ArXiv"},{"key":"e_1_3_2_70_2","article-title":"Mengzi: Towards lightweight yet ingenious pre-trained models for chinese","volume":"2110","author":"Zhang Zhuosheng","year":"2021","unstructured":"Zhuosheng Zhang, Hanqing Zhang, Keming Chen, Yuhang Guo, Jingyun Hua, Yulong Wang, and Ming Zhou. 2021. Mengzi: Towards lightweight yet ingenious pre-trained models for chinese. ArXiv abs\/2110.06696 (2021).","journal-title":"ArXiv"},{"key":"e_1_3_2_71_2","doi-asserted-by":"crossref","unstructured":"Qi Zhu Zheng Zhang Yan Fang Xiang Li Ryuichi Takanobu Jinchao Li Baolin Peng Jianfeng Gao Xiaoyan Zhu and Minlie Huang. 2020. ConvLab-2: An open-source toolkit for building evaluating and diagnosing dialogue systems. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (ACL\u201920) . 142\u2013149.","DOI":"10.18653\/v1\/2020.acl-demos.19"},{"key":"e_1_3_2_72_2","doi-asserted-by":"crossref","unstructured":"Lixin Zou Long Xia Yulong Gu Xiangyu Zhao Weidong Liu Xiangji Huang and Dawei Yin. 2020. Neural interactive collaborative filtering. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR\u201920) . 749\u2013758.","DOI":"10.1145\/3397271.3401181"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3596510","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3596510","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:48:00Z","timestamp":1750178880000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3596510"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,18]]},"references-count":71,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,1,31]]}},"alternative-id":["10.1145\/3596510"],"URL":"https:\/\/doi.org\/10.1145\/3596510","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,18]]},"assertion":[{"value":"2022-06-14","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-04-12","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-08-18","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}