{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T02:46:49Z","timestamp":1784256409519,"version":"3.55.0"},"reference-count":274,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2026,8,31]]},"abstract":"<jats:p>\n                    This survey examines evaluation methods for large language model (LLM)-based agents in multi-turn conversational settings. Using a PRISMA-inspired framework, we systematically reviewed nearly 250 scholarly sources, capturing the state-of-the-art from various venues of publication, and establishing a solid foundation for our analysis. Our study offers a structured approach by developing two interrelated taxonomy systems: one that defines\n                    <jats:italic toggle=\"yes\">what to evaluate<\/jats:italic>\n                    and another that explains\n                    <jats:italic toggle=\"yes\">how to evaluate<\/jats:italic>\n                    . The first taxonomy identifies key components of LLM-based agents for multi-turn conversations and their evaluation dimensions, including task completion, response quality, user experience, memory and context retention, as well as planning and tool integration. These components ensure that the performance of conversational agents is assessed in a holistic and meaningful manner. The second taxonomy system focuses on the evaluation methodologies. It categorizes approaches into annotation-based evaluations, automated metrics, hybrid strategies that combine human assessments with quantitative measures, and self-judging methods utilizing LLMs. This framework not only captures traditional metrics derived from language understanding, such as BLEU and ROUGE scores, but also incorporates advanced techniques that reflect the dynamic, interactive nature of multi-turn dialogues. Together, these frameworks summarize the current status quo, expose limitations in traditional practices, and provide a structured blueprint for improvement. Based on the summarization of existing studies, we identify several challenges and propose future directions, including the development of scalable, real-time evaluation pipelines, enhanced privacy-preserving mechanisms, and robust metrics that capture dynamic multi-turn interactions. Our contributions bridge historical insights with modern practices, paving the way for next-generation, reliably evaluated conversational AI systems and offering a comprehensive guide for researchers and practitioners.\n                  <\/jats:p>","DOI":"10.1145\/3793671","type":"journal-article","created":{"date-parts":[[2026,2,14]],"date-time":"2026-02-14T14:27:55Z","timestamp":1771079275000},"page":"1-40","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Evaluating LLM-based Agents for Multi-turn Conversations: A Survey"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-8321-6365","authenticated-orcid":false,"given":"Shengyue","family":"Guan","sequence":"first","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4833-0880","authenticated-orcid":false,"given":"Jindong","family":"Wang","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6337-9375","authenticated-orcid":false,"given":"Jiang","family":"Bian","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3571-7808","authenticated-orcid":false,"given":"Bin","family":"Zhu","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8496-033X","authenticated-orcid":false,"given":"Jian-Guang","family":"Lou","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5451-3253","authenticated-orcid":false,"given":"Haoyi","family":"Xiong","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,4,28]]},"reference":[{"key":"e_1_3_2_1_2","first-page":"229","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Student Research Workshop","author":"A Sujan Reddy","year":"2022","unstructured":"Sujan Reddy A. 2022. Automating human evaluation of dialogue systems. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Student Research Workshop. Daphne Ippolito, Liunian Harold Li, Maria Leonor Pacheco, Danqi Chen, and Nianwen Xue (Eds.), Association for Computational Linguistics, 229\u2013234. DOI: 10.18653\/v1\/2022.naacl-srw.29"},{"key":"e_1_3_2_2_2","first-page":"5073","volume-title":"Proceedings of 12th Language Resources and Evaluation Conference","author":"Abercrombie Gavin","year":"2020","unstructured":"Gavin Abercrombie and Riza Batista-Navarro. 2020. ParlVote: A corpus for sentiment analysis of political debates. In Proceedings of 12th Language Resources and Evaluation Conference. Nicoletta Calzolari, Fr\u00e9d\u00e9ric B\u00e9chet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, et al. (Eds.), European Language Resources Association, 5073\u20135078. Retrieved from https:\/\/aclanthology.org\/2020.lrec-1.624"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/s42001-019-00060-w"},{"key":"e_1_3_2_4_2","first-page":"5161","volume-title":"Proceedings of the 32nd ACM International Conference on Information and Knowledge Management (Birmingham, United Kingdom) (CIKM \u201923)","author":"Acharya Praveen","year":"2023","unstructured":"Praveen Acharya. 2023. Towards effective modeling and exploitation of search and user context in conversational information retrieval. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management (Birmingham, United Kingdom) (CIKM \u201923). ACM, 5161\u20135164. DOI: 10.1145\/3583780.3616005"},{"key":"e_1_3_2_5_2","first-page":"1255","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track","author":"Agarwal Divyansh","year":"2024","unstructured":"Divyansh Agarwal, Alexander Fabbri, Ben Risher, Philippe Laban, Shafiq Joty, and Chien-Sheng Wu. 2024. Prompt leakage effect and mitigation strategies for multi-turn LLM applications. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track. Franck Dernoncourt, Daniel Preo\u0163iuc-Pietro, and Anastasia Shimorina (Eds.), Association for Computational Linguistics, 1255\u20131275. DOI: 10.18653\/v1\/2024.emnlp-industry.94"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.dim.2022.100025"},{"key":"e_1_3_2_7_2","unstructured":"Lize Alberts Geoff Keeling and Amanda McCroskery. 2024. Should agentic conversational AI change how we think about ethics? Characterising an interactional ethics centred on respect. arXiv:2401.09082. Retrieved from https:\/\/arxiv.org\/abs\/2401.09082"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-023-09322-1"},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of the 2020 Conference on Human Information Interaction and Retrieval (CHIIR \u201920)","author":"Aliannejadi Mohammad","year":"2020","unstructured":"Mohammad Aliannejadi, Manajit Chakraborty, Esteban Andr\u00e9s R\u00edssola, and Fabio Crestani. 2020. Harnessing evolution of multi-turn conversations for effective answer retrieval. In Proceedings of the 2020 Conference on Human Information Interaction and Retrieval (CHIIR \u201920). ACM. DOI: 10.1145\/3343413.3377968"},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","unstructured":"Samuel Arcadinho David Aparicio and Mariana Almeida. 2024. Automated test generation to evaluate tool-augmented LLMs as conversational AI agents. arXiv:2409.15934. Retrieved from https:\/\/arxiv.org\/abs\/2409.15934","DOI":"10.18653\/v1\/2024.genbench-1.4"},{"key":"e_1_3_2_11_2","unstructured":"Suket Arora Kamaljeet Batra and Sarabjit Singh. 2013. Dialogue system: A brief review. arXiv:1306.4134. Retrieved from https:\/\/arxiv.org\/abs\/1306.4134"},{"key":"e_1_3_2_12_2","doi-asserted-by":"crossref","first-page":"7352","DOI":"10.18653\/v1\/2020.acl-main.656","volume-title":"Proceedings of 58th Annual Meeting of the Association for Computational Linguistics","author":"Atanasova Pepa","year":"2020","unstructured":"Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020. Generating fact checking explanations. In Proceedings of 58th Annual Meeting of the Association for Computational Linguistics. Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.), Association for Computational Linguistics, 7352\u20137364. DOI: 10.18653\/v1\/2020.acl-main.656"},{"key":"e_1_3_2_13_2","unstructured":"Sanghwan Bae Donghyun Kwak Soyoung Kang Min Young Lee Sungdong Kim Yuin Jeong Hyeri Kim Sang-Woo Lee Woomyoung Park and Nako Sung. 2022. Keep me updated! Memory management in long-term conversations. arXiv:2210.08750. Retrieved from https:\/\/arxiv.org\/abs\/2210.08750"},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","first-page":"7421","DOI":"10.18653\/v1\/2024.acl-long.401","volume-title":"Proceedings of 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Bai Ge","year":"2024","unstructured":"Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, et al. 2024. MT-Bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues. In Proceedings of 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 7421\u20137454. DOI: 10.18653\/v1\/2024.acl-long.401"},{"key":"e_1_3_2_15_2","volume-title":"Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023NeurIPS \u201923)","author":"Bai Yushi","year":"2023","unstructured":"Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, et al. 2023. Benchmarking foundation models with language-model-as-an-examiner. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023 (NeurIPS \u201923). Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.), Retrieved from http:\/\/papers.nips.cc\/paper_files\/paper\/2023\/hash\/f64e55d03e2fe61aa4114e49cb654acb-Abstract-Datasets_and_Benchmarks.html"},{"key":"e_1_3_2_16_2","unstructured":"Debarag Banerjee Pooja Singh Arjun Avadhanam and Saksham Srivastava. 2023. Benchmarking LLM powered chatbots: Methods and metrics. arXiv:2308.04624. Retrieved from https:\/\/arxiv.org\/abs\/2308.04624"},{"key":"e_1_3_2_17_2","first-page":"65","volume-title":"Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization","author":"Banerjee Satanjeev","year":"2005","unstructured":"Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization. Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare Voss (Eds.), Association for Computational Linguistics, 65\u201372. Retrieved from https:\/\/aclanthology.org\/W05-0909\/"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","first-page":"18393","DOI":"10.18653\/v1\/2024.emnlp-main.1022","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Bassani Elias","year":"2024","unstructured":"Elias Bassani and Ignacio Sanchez. 2024. GuardBench: A large-scale benchmark for guardrail models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 18393\u201318409. DOI: 10.18653\/v1\/2024.emnlp-main.1022"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i16.29720"},{"key":"e_1_3_2_20_2","unstructured":"Tom B. Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. arXiv:2005.14165. Retrieved from https:\/\/arxiv.org\/abs\/2005.14165"},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","first-page":"5016","DOI":"10.18653\/v1\/D18-1547","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Budzianowski Pawe\u0142","year":"2018","unstructured":"Pawe\u0142 Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, I\u00f1igo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Ga\u0161i\u0107. 2018. MultiWOZ\u2014A large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun\u2019ichi Tsujii (Eds.), Association for Computational Linguistics, 5016\u20135026. DOI: 10.18653\/v1\/D18-1547"},{"key":"e_1_3_2_22_2","unstructured":"Pawe\u0142 Budzianowski Tsung-Hsien Wen Bo-Hsiang Tseng I\u00f1igo Casanueva Stefan Ultes Osman Ramadan and Milica Ga\u0161i\u0107. 2020. MultiWOZ\u2014A large-scale multi-domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling. arXiv:1810.00278. Retrieved from https:\/\/arxiv.org\/abs\/1810.00278"},{"key":"e_1_3_2_23_2","first-page":"4516","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Byrne Bill","year":"2019","unstructured":"Bill Byrne, Karthik Krishnamoorthi, Chinnadhurai Sankar, Arvind Neelakantan, Ben Goodrich, Daniel Duckworth, Semih Yavuz, Amit Dubey, Kyu-Young Kim, and Andy Cedilnik. 2019. Taskmaster-1: Toward a realistic and diverse dialog dataset. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.), Association for Computational Linguistics, 4516\u20134525. DOI: 10.18653\/v1\/D19-1459"},{"key":"e_1_3_2_24_2","volume-title":"Proceedings of the 31st ACM International Conference on Information & Knowledge Management","author":"Cai Xiaoyu","year":"2022","unstructured":"Xiaoyu Cai, Yao Fu, Hong Zhao, Weihao Jiang, and Shi Pu. 2022. Memory graph with message rehearsal for multi-turn dialogue generation. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:252904675"},{"key":"e_1_3_2_25_2","unstructured":"Lang Cao. 2024. DiagGPT: An LLM-based and multi-agent dialogue system with automatic topic management for flexible task-oriented dialogue. arXiv:2308.08043. Retrieved from https:\/\/arxiv.org\/abs\/2308.08043"},{"key":"e_1_3_2_26_2","first-page":"1","volume-title":"Proceedings of 2nd Workshop on Natural Language Reasoning and Structured Explanations (@ACL \u201924)","author":"Cao Lang","year":"2024","unstructured":"Lang Cao. 2024. GraphReason: Enhancing reasoning capabilities of large language models through a graph-based verification approach. In Proceedings of 2nd Workshop on Natural Language Reasoning and Structured Explanations (@ACL \u201924). Bhavana Dalvi Mishra, Greg Durrett, Peter Jansen, Ben Lipkin, Danilo Neves Ribeiro, Lionel Wong, Xi Ye, and Wenting Zhao (Eds.), Association for Computational Linguistics, 1\u201312. Retrieved from https:\/\/aclanthology.org\/2024.n,lrse-1.1\/"},{"key":"e_1_3_2_27_2","first-page":"3628","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Cao Lang","year":"2024","unstructured":"Lang Cao. 2024. Learn to refuse: Making large language models more controllable and reliable through knowledge scope limitation and refusal mechanism. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 3628\u20133646. DOI: https:\/\/doi.org\/10.18653\/v1\/2024.emnlp-main.212"},{"key":"e_1_3_2_28_2","doi-asserted-by":"crossref","unstructured":"David Castillo-Bolado Joseph Davidson Finlay Gray and Marek Rosa. 2024. Beyond prompts: Dynamic conversational benchmarking of large language models. arXiv:2409.20222. Retrieved from https:\/\/arxiv.org\/abs\/2409.20222","DOI":"10.52202\/079017-1347"},{"key":"e_1_3_2_29_2","doi-asserted-by":"crossref","first-page":"5606","DOI":"10.18653\/v1\/2023.emnlp-main.342","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Chae Hyungjoo","year":"2023","unstructured":"Hyungjoo Chae, Yongho Song, Kai Ong, Taeyoon Kwon, Minjin Kim, Youngjae Yu, Dongha Lee, Dongyeop Kang, and Jinyoung Yeo. 2023. Dialogue chain-of-thought distillation for commonsense-aware conversational agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Houda Bouamor, Juan Pino, and Kalika Bali (Eds.), Association for Computational Linguistics, 5606\u20135632. DOI: 10.18653\/v1\/2023.emnlp-main.342"},{"key":"e_1_3_2_30_2","unstructured":"Chi-Min Chan Weize Chen Yusheng Su Jianxuan Yu Wei Xue Shanghang Zhang Jie Fu and Zhiyuan Liu. 2023. ChatEval: Towards better LLM-based evaluators through multi-agent debate. arXiv:2308.07201. Retrieved from https:\/\/arxiv.org\/abs\/2308.07201"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","unstructured":"Yupeng Chang Xu Wang Jindong Wang Yuan Wu Linyi Yang Kaijie Zhu Hao Chen Xiaoyuan Yi Cunxiang Wang Yidong Wang et al. 2024. A survey on evaluation of large language models. 15 3 Article 39 (March 2024) 45. DOI: 10.1145\/3641289","DOI":"10.1145\/3641289"},{"key":"e_1_3_2_32_2","unstructured":"Harrison Chase. 2022. LangChain. Retrieved from https:\/\/github.com\/langchain-ai\/langchain"},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","first-page":"7100","DOI":"10.18653\/v1\/2022.emnlp-main.478","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Chaudhury Subhajit","year":"2022","unstructured":"Subhajit Chaudhury, Sarathkrishna Swaminathan, Chulaka Gunasekara, Maxwell Crouse, Srinivas Ravishankar, Daiki Kimura, Keerthiram Murugesan, Ram\u00f3n Fernandez Astudillo, Tahira Naseem, Pavan Kapanipathi, et al. 2022. X-FACTOR: A cross-metric evaluation of factual correctness in abstractive summarization. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.), Association for Computational Linguistics, 7100\u20137110. DOI: 10.18653\/v1\/2022.emnlp-main.478"},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","first-page":"2108","DOI":"10.18653\/v1\/2024.findings-acl.125","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024","author":"Chen Hongzhan","year":"2024","unstructured":"Hongzhan Chen, Hehong Chen, Ming Yan, Wenshen Xu, Gao Xing, Weizhou Shen, Xiaojun Quan, Chenliang Li, Ji Zhang, and Fei Huang. 2024. SocialBench: Sociality evaluation of role-playing conversational agents. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024. Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.), Association for Computational Linguistics, 2108\u20132126. DOI: 10.18653\/v1\/2024.findings-acl.125"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3166054.3166058"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","first-page":"9057","DOI":"10.18653\/v1\/2024.findings-emnlp.529","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024","author":"Chen Kedi","year":"2024","unstructured":"Kedi Chen, Qin Chen, Jie Zhou, He Yishen, and Liang He. 2024. DiaHalu: A dialogue-level hallucination evaluation benchmark for large language models. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024. Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 9057\u20139079. DOI: 10.18653\/v1\/2024.findings-emnlp.529"},{"key":"e_1_3_2_37_2","unstructured":"Wenhu Chen Xueguang Ma Xinyi Wang and William W. Cohen. 2023. Program of thoughts prompting: disentangling computation from reasoning for numerical reasoning tasks. arXiv:2211.12588. Retrieved from https:\/\/arxiv.org\/abs\/2211.12588"},{"key":"e_1_3_2_38_2","unstructured":"Yanbing Chen Lin Li Xiaohui Tao and Dong Zhou. 2024. Persona-centric metamorphic relation guided robustness evaluation for multi-turn dialogue modelling. arXiv:2401.12483. Retrieved from https:\/\/arxiv.org\/abs\/2401.12483"},{"key":"e_1_3_2_39_2","doi-asserted-by":"crossref","first-page":"3014","DOI":"10.18653\/v1\/2022.emnlp-main.195","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Cheng Yi","year":"2022","unstructured":"Yi Cheng, Wenge Liu, Wenjie Li, Jiashuo Wang, Ruihui Zhao, Bang Liu, Xiaodan Liang, and Yefeng Zheng. 2022. Improving multi-turn emotional support dialogue generation with lookahead strategy planning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.), Association for Computational Linguistics, 3014\u20133026. DOI: 10.18653\/v1\/2022.emnlp-main.195"},{"key":"e_1_3_2_40_2","doi-asserted-by":"crossref","first-page":"15607","DOI":"10.18653\/v1\/2023.acl-long.870","volume-title":"Proceedings of 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Chiang Cheng-Han","year":"2023","unstructured":"Cheng-Han Chiang and Hung-Yi Lee. 2023. Can large language models be an alternative to human evaluations? In Proceedings of 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.), Association for Computational Linguistics, 15607\u201315631. DOI: 10.18653\/v1\/2023.acl-long.870"},{"key":"e_1_3_2_41_2","doi-asserted-by":"crossref","first-page":"8928","DOI":"10.18653\/v1\/2023.findings-emnlp.599","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201923","author":"Chiang Cheng-Han","year":"2023","unstructured":"Cheng-Han Chiang and Hung-Yi Lee. 2023. A closer look into using large language models for automatic evaluation. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201923). Houda Bouamor, Juan Pino, and Kalika Bali (Eds.), Association for Computational Linguistics, 8928\u20138942. DOI: 10.18653\/v1\/2023.findings-emnlp.599"},{"key":"e_1_3_2_42_2","first-page":"11346","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Cho Young Min","unstructured":"Young Min Cho, Sunny Rai, Lyle Ungar, Jo\u00e3o Sedoc, and Sharath Guntuku. 2023. An integrative survey on mental health conversational agents to bridge computer science and medical perspectives. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Houda Bouamor, Juan Pino, and Kalika Bali (Eds.), Association for Computational Linguistics, 11346\u201311369. DOI: 10.18653\/v1\/2023.emnlp-main.698"},{"key":"e_1_3_2_43_2","first-page":"40","volume-title":"Proceedings of 3rd Workshop on Bridging Human\u2013Computer Interaction and Natural Language Processing","author":"Choi Jason Ingyu","year":"2024","unstructured":"Jason Ingyu Choi, Marcus Collins, Eugene Agichtein, Oleg Rokhlenko, and Shervin Malmasi. 2024. Combining multiple metrics for evaluating retrieval-augmented conversations. In Proceedings of 3rd Workshop on Bridging Human\u2013Computer Interaction and Natural Language Processing. Su Lin Blodgett, Amanda Cercas Curry, Sunipa Dev, Michael Madaio, Ani Nenkova, Diyi Yang, and Ziang Xiao (Eds.), Association for Computational Linguistics, 40\u201350. DOI: 10.18653\/v1\/2024.hcinlp-1.4"},{"key":"e_1_3_2_44_2","unstructured":"Victor Costan and Srinivas Devadas. 2016. Intel SGX explained. Cryptology ePrint Archive. Retrieved from https:\/\/eprint.iacr.org\/2016\/086"},{"key":"e_1_3_2_45_2","unstructured":"Gautier Dagan Frank Keller and Alex Lascarides. 2023. Dynamic planning with a LLM. arXiv:2308.06391. Retrieved from https:\/\/arxiv.org\/abs\/2308.06391"},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","first-page":"8795","DOI":"10.18653\/v1\/2024.acl-long.477","volume-title":"Proceedings of 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Deng Yang","year":"2024","unstructured":"Yang Deng, Xuan Zhang, Wenxuan Zhang, Yifei Yuan, See-Kiong Ng, and Tat-Seng Chua. 2024. On the multi-turn instruction following for conversational web agents. In Proceedings of 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.), Association for Computational Linguistics, 8795\u20138812. DOI: 10.18653\/v1\/2024.acl-long.477"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-020-09866-x"},{"key":"e_1_3_2_48_2","unstructured":"Zican Dong Tianyi Tang Junyi Li Wayne Xin Zhao and Ji-Rong Wen. 2024. BAMBOO: A comprehensive benchmark for evaluating long text modeling capacities of large language models. arXiv:2309.13345. Retrieved from https:\/\/arxiv.org\/abs\/2309.13345"},{"key":"e_1_3_2_49_2","doi-asserted-by":"crossref","first-page":"6734","DOI":"10.18653\/v1\/2024.naacl-long.375","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Dong Zhichen","year":"2024","unstructured":"Zhichen Dong, Zhanhui Zhou, Chao Yang, Jing Shao, and Yu Qiao. 2024. Attacks, defenses and evaluations for LLM conversation safety: A survey. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Kevin Duh, Helena Gomez, and Steven Bethard (Eds.), Association for Computational Linguistics, 6734\u20136747. DOI: 10.18653\/v1\/2024.naacl-long.375"},{"key":"e_1_3_2_50_2","unstructured":"Haodong Duan Jueqi Wei Chonghua Wang Hongwei Liu Yixiao Fang Songyang Zhang Dahua Lin and Kai Chen. 2023. BotChat: Evaluating LLMs\u2019 capabilities of having multi-turn dialogues. arXiv:2310.13650. Retrieved from https:\/\/arxiv.org\/abs\/2310.13650"},{"key":"e_1_3_2_51_2","first-page":"16916","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924)","author":"Fan Zhiyuan","year":"2024","unstructured":"Zhiyuan Fan, Weinong Wang, Xing W, and Debing Zhang. 2024. SedarEval: Automated evaluation using self-adaptive rubrics. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 16916\u201316930. DOI: 10.18653\/v1\/2024.findings-emnlp.984"},{"key":"e_1_3_2_52_2","doi-asserted-by":"crossref","first-page":"7285","DOI":"10.18653\/v1\/2022.findings-emnlp.539","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201922)","author":"Feng Jiazhan","year":"2022","unstructured":"Jiazhan Feng, Chongyang Tao, Chang Liu, Rui Yan, and Dongyan Zhao. 2022. How to represent context better? An empirical study on context modeling for multi-turn response selection. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201922). Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.), Association for Computational Linguistics, 7285\u20137298. DOI: 10.18653\/v1\/2022.findings-emnlp.539"},{"key":"e_1_3_2_53_2","first-page":"2078","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (\u201923","author":"Ferron Amila","year":"2023","unstructured":"Amila Ferron, Amber Shore, Ekata Mitra, and Ameeta Agrawal. 2023. MEEP: Is this engaging? Prompting large language models for dialogue evaluation in multilingual settings. In Proceedings of the Findings of the Association for Computational Linguistics (\u201923). Houda Bouamor, Juan Pino, and Kalika Bali (Eds.), Association for Computational Linguistics, 2078\u20132100. DOI: 10.18653\/v1\/2023.findings-emnlp.137"},{"key":"e_1_3_2_54_2","unstructured":"Zafeirios Fountas Martin A. Benfeghoul Adnan Oomerjee Fenia Christopoulou Gerasimos Lampouras Haitham Bou-Ammar and Jun Wang. 2024. Human-like episodic memory for infinite context LLMs. arXiv:2407.09450. Retrieved from https:\/\/arxiv.org\/abs\/2407.09450"},{"key":"e_1_3_2_55_2","first-page":"643","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Fu Dayuan","year":"2024","unstructured":"Dayuan Fu, Biqing Qi, Yihuai Gao, Che Jiang, Guanting Dong, and Bowen Zhou. 2024. MSI-Agent: Incorporating multi-scale insight into embodied agents for superior planning and decision-making. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 643\u2013659. DOI: 10.18653\/v1\/2024.emnlp-main.38"},{"key":"e_1_3_2_56_2","first-page":"495","volume-title":"Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Fu Tingchen","year":"2023","unstructured":"Tingchen Fu, Xueliang Zhao, and Rui Yan. 2023. Delving into global dialogue structures: Structure planning augmented response selection for multi-turn conversations. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 495\u2013505. DOI: 10.1145\/3580305.3599304"},{"key":"e_1_3_2_57_2","unstructured":"Luyu Gao Aman Madaan Shuyan Zhou Uri Alon Pengfei Liu Yiming Yang Jamie Callan and Graham Neubig. 2023. PAL: Program-aided language models. arXiv:2211.10435. Retrieved from https:\/\/arxiv.org\/abs\/2211.10435"},{"key":"e_1_3_2_58_2","doi-asserted-by":"crossref","unstructured":"Xiang Gao Yizhe Zhang Michel Galley Chris Brockett and Bill Dolan. 2020. Dialogue response ranking training with large-scale human feedback data. arXiv:2009.06978. Retrieved from https:\/\/arxiv.org\/abs\/2009.06978","DOI":"10.18653\/v1\/2020.emnlp-main.28"},{"key":"e_1_3_2_59_2","doi-asserted-by":"crossref","first-page":"4194","DOI":"10.18653\/v1\/2022.findings-acl.331","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (ACL \u201922)","author":"Ghazarian Sarik","year":"2022","unstructured":"Sarik Ghazarian, Behnam Hedayatnia, Alexandros Papangelis, Yang Liu, and Dilek Hakkani-Tur. 2022. What is wrong with you? Leveraging user sentiment for automatic dialog evaluation. In Proceedings of the Findings of the Association for Computational Linguistics (ACL \u201922). Smaranda Muresan, Preslav Nakov, and Aline Villavicencio(Eds.), Association for Computational Linguistics, 4194\u20134204. DOI: 10.18653\/v1\/2022.findings-acl.331"},{"key":"e_1_3_2_60_2","first-page":"354","volume-title":"Proceedings of the Intelligent Information and Database Systems","author":"Goh Ong Sing","year":"2016","unstructured":"Ong Sing Goh, Yogan Jaya Kumar, Ngo Hea Choon, Pui Huang Leong, and Mohammad Safar. 2016. An Evaluation of the Conversation Agent System. In Proceedings of the Intelligent Information and Database Systems. Ngoc Thanh Nguyen, Bogdan Trawi\u0144ski, Hamido Fujita, and Tzung-Pei Hong (Eds.), Springer Berlin Heidelberg, 354\u2013365."},{"key":"e_1_3_2_61_2","doi-asserted-by":"crossref","first-page":"1908","DOI":"10.18653\/v1\/2024.findings-emnlp.106","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924)","author":"Golany Lotem","year":"2024","unstructured":"Lotem Golany, Filippo Galgani, Maya Mamo, Nimrod Parasol, Omer Vandsburger, Nadav Bar, and Ido Dagan. 2024. Efficient data generation for source-grounded information-seeking dialogs: A use case for meeting transcripts. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen(Eds.), Association for Computational Linguistics, 1908\u20131925. DOI: 10.18653\/v1\/2024.findings-emnlp.106"},{"key":"e_1_3_2_62_2","doi-asserted-by":"crossref","first-page":"1891","DOI":"10.21437\/Interspeech.2019-3079","volume-title":"Proceedings of the Interspeech 2019","author":"Gopalakrishnan Karthik","year":"2019","unstructured":"Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-T\u00fcr. 2019. Topical-Chat: towards knowledge-grounded open-domain conversations. In Proceedings of the Interspeech 2019, 1891\u20131895. DOI: 10.21437\/Interspeech.2019-3079"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1093\/oso\/9780198849063.003.0005"},{"key":"e_1_3_2_64_2","unstructured":"Milan Gritta Gerasimos Lampouras and Ignacio Iacobacci. 2024. HumanRankEval: Automatic evaluation of LMs as conversational assistants. arXiv:2405.09186. Retrieved from https:\/\/arxiv.org\/abs\/2405.09186"},{"key":"e_1_3_2_65_2","first-page":"86","volume-title":"Proceedings of 21st International Conference Service-Oriented Computing (ICSOC \u201923), Part I","author":"Gu Yang","year":"2023","unstructured":"Yang Gu, Jian Cao, Yuan Guo, Shiyou Qian, and Wei Guan. 2023. Plan, generate and match: Scientific workflow recommendation with large language models. In Proceedings of 21st International Conference Service-Oriented Computing (ICSOC \u201923), Part I. Springer-Verlag, 86\u2013102. DOI: 10.1007\/978-3-031-48421-6_7"},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","first-page":"7646","DOI":"10.18653\/v1\/2024.emnlp-main.436","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Gu Yu","year":"2024","unstructured":"Yu Gu, Yiheng Shu, Hao Yu, Xiao Liu, Yuxiao Dong, Jie Tang, Jayanth Srinivasa, Hugo Latapie, and Yu Su. 2024. Middleware for LLMs: Tools are instrumental for language agents in complex environments. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 7646\u20137663. DOI: 10.18653\/v1\/2024.emnlp-main.436"},{"key":"e_1_3_2_67_2","unstructured":"Ece Gumusel Kyrie Zhixuan Zhou and Madelyn Rose Sanfilippo. 2024. User privacy harms and risks in conversational AI: A proposed framework. arXiv:2402.09716. Retrieved from https:\/\/arxiv.org\/abs\/2402.09716"},{"key":"e_1_3_2_68_2","first-page":"15711","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (ACL \u201924)","author":"Guo Zishan","year":"2024","unstructured":"Zishan Guo, Yufei Huang, and Deyi Xiong. 2024. CToolEval: A Chinese benchmark for LLM-powered agent evaluation in real-world API interactions. In Proceedings of the Findings of the Association for Computational Linguistics (ACL \u201924). Lun-Wei Ku, Andre Martins, and Vivek Srikumar(Eds.), Association for Computational Linguistics, 15711\u201315724. DOI: 10.18653\/v1\/2024.findings-acl.928"},{"key":"e_1_3_2_69_2","doi-asserted-by":"crossref","first-page":"3785","DOI":"10.18653\/v1\/2022.acl-long.263","volume-title":"Proceedings of 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Gupta Prakhar","year":"2022","unstructured":"Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022. DialFact: A benchmark for fact-checking in dialogue. In Proceedings of 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.), Association for Computational Linguistics, 3785\u20133801. DOI: 10.18653\/v1\/2022.acl-long.263"},{"key":"e_1_3_2_70_2","unstructured":"Ilya Gusev. 2024. PingPong: A benchmark for role-playing language models with user emulation and multi-model evaluation. arXiv:2409.06820. Retrieved from https:\/\/arxiv.org\/abs\/2409.06820"},{"issue":"2","key":"e_1_3_2_71_2","first-page":"161","article-title":"Dialogues with colorful \u201cpersonalities\u201d of early AI","volume":"4","author":"G\u00fczeldere G\u00fcven","year":"1995","unstructured":"G\u00fcven G\u00fczeldere and Stefano Franchi. 1995. Dialogues with colorful \u201cpersonalities\u201d of early AI. Stanford Humanities Review 4, 2 (July 1995), 161\u2013169.","journal-title":"Stanford Humanities Review"},{"key":"e_1_3_2_72_2","doi-asserted-by":"crossref","first-page":"6988","DOI":"10.18653\/v1\/2023.emnlp-main.432","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processingand","author":"Han Xue","year":"2023","unstructured":"Xue Han, Yitong Wang, Qian Hu, Pengwei Hu, Chao Deng, and Junlan Feng. 2023. Log-FGAER: Logic-guided fine-grained address entity recognition from multi-turn spoken dialogue. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Houda Bouamor, Juan Pino, and Kalika Bali (Eds.), Association for Computational Linguistics, 6988\u20136997. DOI: 10.18653\/v1\/2023.emnlp-main.432"},{"key":"e_1_3_2_73_2","unstructured":"Zengguang Hao Jie Zhang Binxia Xu Yafang Wang Gerard de Melo and Xiaolong Li. 2023. IntentDial: An intent graph based multi-turn dialogue system with reasoning path visualization. arXiv:2310.11818. Retrieved from https:\/\/arxiv.org\/abs\/2310.11818"},{"key":"e_1_3_2_74_2","doi-asserted-by":"crossref","first-page":"75","DOI":"10.18653\/v1\/2024.scichat-1.8","volume-title":"Proceedings of 1st Workshop on Simulating Conversational Intelligence in Chat (SCI-CHAT \u201924)","author":"Hassan Islam A.","year":"2024","unstructured":"Islam A. Hassan, Yvette Graham, and Rameez Qureshi. 2024. Advancing open-domain conversational agents\u2014Designing an engaging system for natural multi-turn dialogue. In Proceedings of 1st Workshop on Simulating Conversational Intelligence in Chat (SCI-CHAT \u201924). Yvette Graham, Qun Liu, Gerasimos Lampouras, Ignacio Iacobacci, Sinead Madden, and Haider Khalid (Eds.), Association for Computational Linguistics, 75\u201379. Retrieved from https:\/\/aclanthology.org\/2024.scichat-1.8\/"},{"key":"e_1_3_2_75_2","doi-asserted-by":"crossref","DOI":"10.1609\/aaaiss.v2i1.27688","article-title":"Memory matters: The need to improve long-term memory in LLM-agents","author":"Hatalis Kostas","year":"2024","unstructured":"Kostas Hatalis, Despina Christou, Joshua Myers, Steven Jones, Keith Lambert, Adam Amos-Binks, Zohreh Dannenhauer, and Dustin Dannenhauer. 2024. Memory matters: The need to improve long-term memory in LLM-agents. In Proceedings of the AAAI Symposium Series. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:267224915","journal-title":"Proceedings of the AAAI Symposium Series"},{"key":"e_1_3_2_76_2","unstructured":"Ji He Jianshu Chen Xiaodong He Jianfeng Gao Lihong Li Li Deng and Mari Ostendorf. 2016. Deep reinforcement learning with a natural language action space. arXiv:1511.04636. Retrieved from https:\/\/arxiv.org\/abs\/1511.04636"},{"key":"e_1_3_2_77_2","unstructured":"Dan Hendrycks Collin Burns Steven Basart Andy Zou Mantas Mazeika Dawn Xiaodong Song and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv:2009.03300. Retrieved from https:\/\/arxiv.org\/abs\/2009.03300"},{"key":"e_1_3_2_78_2","volume-title":"Advances in Neural Information Processing Systems","author":"Herbrich Ralf","year":"2006","unstructured":"Ralf Herbrich, Tom Minka, and Thore Graepel. 2006. TrueSkill\u2122: A Bayesian skill rating system. In Advances in Neural Information Processing Systems. B. Sch\u00f6lkopf, J. Platt, and T. Hoffman (Eds.), Vol. 19, MIT Press. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2006\/file\/f44ee263952e65b3610b8ba51229d1f9-Paper.pdf"},{"key":"e_1_3_2_79_2","article-title":"Blockchain and AI-driven framework for measuring the digital economy in GCC","author":"Hoxha Julian","year":"2024","unstructured":"Julian Hoxha and Marsela Thanasi-Bo\u00e7e. 2024. Blockchain and AI-driven framework for measuring the digital economy in GCC. Emerging Science Journal. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:272139752","journal-title":"Emerging Science Journal"},{"key":"e_1_3_2_80_2","unstructured":"Sihao Hu Tiansheng Huang Fatih Ilhan Selim Tekin Gaowen Liu Ramana Kompella and Ling Liu. 2024. A survey on large language model-based game agents. arXiv:2404.02039. Retrieved from https:\/\/arxiv.org\/abs\/2404.02039"},{"key":"e_1_3_2_81_2","doi-asserted-by":"crossref","first-page":"4363","DOI":"10.18653\/v1\/2024.findings-acl.259","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (ACL \u201924)","author":"Huang Shijue","year":"2024","unstructured":"Shijue Huang, Wanjun Zhong, Jianqiao Lu, Qi Zhu, Jiahui Gao, Weiwen Liu, Yutai Hou, Xingshan Zeng, Yasheng Wang, Lifeng Shang, et al. 2024. Planning, creation, usage: Benchmarking LLMs for comprehensive tool utilization in real-world complex scenarios. In Proceedings of the Findings of the Association for Computational Linguistics (ACL \u201924). Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.), Association for Computational Linguistics, 4363\u20134400. DOI: 10.18653\/v1\/2024.findings-acl.259"},{"key":"e_1_3_2_82_2","doi-asserted-by":"crossref","unstructured":"Shih-Hong Huang Ya-Fang Lin Zeyu He Chieh-Yang Huang and Ting-Hao \u2018Kenneth\u2019 Huang. 2024. How does conversation length impact user\u2019s satisfaction? A case study of length-controlled conversations with LLM-powered chatbots. arXiv:2404.17025. Retrieved from https:\/\/arxiv.org\/abs\/2404.17025","DOI":"10.1145\/3613905.3650823"},{"key":"e_1_3_2_83_2","first-page":"1093","volume-title":"Proceedings of 26th International Joint Conference on Artificial Intelligence (IJCAI \u201917)","author":"Huang Xiao","year":"2017","unstructured":"Xiao Huang, Biqing Fang, Hai Wan, and Yongmei Liu. 2017. A general multi-agent epistemic planner based on higher-order belief change. In Proceedings of 26th International Joint Conference on Artificial Intelligence (IJCAI \u201917). International Joint Conferences on Artificial Intelligence Organization, 1093\u20131101. DOI: 10.24963\/ijcai.2017\/152"},{"key":"e_1_3_2_84_2","unstructured":"Xu Huang Weiwen Liu Xiaolong Chen Xingmei Wang Hao Wang Defu Lian Yasheng Wang Ruiming Tang and Enhong Chen. 2024. Understanding the planning of LLM agents: A survey. arXiv:2402.02716. Retrieved from https:\/\/arxiv.org\/abs\/2402.02716"},{"key":"e_1_3_2_85_2","unstructured":"Yue Huang Jiawen Shi Yuan Li Chenrui Fan Siyuan Wu Qihui Zhang Yixin Liu Pan Zhou Yao Wan Neil Zhenqiang Gong and Lichao Sun. 2024. MetaTool benchmark for large language models: Deciding whether to use tools and which to use. arXiv:2310.03128. Retrieved from https:\/\/arxiv.org\/abs\/2310.03128"},{"key":"e_1_3_2_86_2","doi-asserted-by":"crossref","unstructured":"Ziheng Huang Sebastian Gutierrez Hemanth Kamana and Stephen MacNeil. 2023. Memory sandbox: Transparent and interactive memory management for conversational agents. arXiv:2308.01542. Retrieved from https:\/\/arxiv.org\/abs\/2308.01542","DOI":"10.1145\/3586182.3615796"},{"key":"e_1_3_2_87_2","doi-asserted-by":"publisher","DOI":"10.1214\/aos\/1079120141"},{"key":"e_1_3_2_88_2","unstructured":"Kai Tzu Iunn Ong Namyoung Kim Minju Gwak Hyungjoo Chae Taeyoon Kwon Yohan Jo Seung Won Hwang Dongha Lee and Jinyoung Yeo. 2024. Towards lifelong dialogue agents via relation-aware memory construction and timeline-augmented response generation. arXiv:2406.10996. Retrieved from https:\/\/arxiv.org\/abs\/2406.10996"},{"key":"e_1_3_2_89_2","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Jang Jihyoung","year":"2023","unstructured":"Jihyoung Jang, MinSeong Boo, and Hyounghun Kim. 2023. Conversation chronicles: Towards diverse temporal and relational dynamics in multi-session conversations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Retrieved from https:\/\/openreview.net\/forum?id=9LPJK81xy1"},{"issue":"2","key":"e_1_3_2_90_2","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1109\/MSEC.2019.2947124","article-title":"Trusted execution environments: Properties, applications, and challenges","volume":"18","author":"Jauernig Patrick","year":"2020","unstructured":"Patrick Jauernig, Ahmad-Reza Sadeghi, and Emmanuel Stapf. 2020. Trusted execution environments: Properties, applications, and challenges. IEEE Security & Privacy 18, 2 (2020), 56\u201360.","journal-title":"IEEE Security & Privacy"},{"key":"e_1_3_2_91_2","article-title":"Intent detection for task\u2010oriented conversational agents: A comparative study of recurrent neural networks and transformer models","author":"Jbene Mourad","year":"2024","unstructured":"Mourad Jbene, Abdellah Chehri, Rachid Saadane, Smail Tigani, and Gwanggil Jeon. 2024. Intent detection for task\u2010oriented conversational agents: A comparative study of recurrent neural networks and transformer models. Expert Systems. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:272030065","journal-title":"Expert Systems"},{"key":"e_1_3_2_92_2","doi-asserted-by":"publisher","DOI":"10.1121\/1.2016299"},{"key":"e_1_3_2_93_2","unstructured":"Qi Jia Yizhu Liu Siyu Ren Kenny Q. Zhu and Haifeng Tang. 2023. Multi-turn response selection using dialogue dependency relations. arXiv:2010.01502. Retrieved from https:\/\/arxiv.org\/abs\/2010.01502"},{"key":"e_1_3_2_94_2","doi-asserted-by":"crossref","first-page":"11091","DOI":"10.18653\/v1\/2022.emnlp-main.762","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Jiang Zhihua","year":"2022","unstructured":"Zhihua Jiang, Guanghui Ye, Dongning Rao, Di Wang, and Xin Miao. 2022. IM2: An interpretable and multi-category integrated metric framework for automatic dialogue evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.), Association for Computational Linguistics, 11091\u201311103. DOI: 10.18653\/v1\/2022.emnlp-main.762"},{"key":"e_1_3_2_95_2","unstructured":"Jeff Johnson Matthijs Douze and Herv\u00e9 J\u00e9gou. 2017. Billion-scale similarity search with GPUs. arXiv:1702.08734. Retrieved from https:\/\/arxiv.org\/abs\/1702.08734"},{"key":"e_1_3_2_96_2","doi-asserted-by":"crossref","first-page":"147","DOI":"10.18653\/v1\/2024.naacl-short.14","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers)","author":"Jones Jaylen","year":"2024","unstructured":"Jaylen Jones, Lingbo Mo, Eric Fosler-Lussier, and Huan Sun. 2024. A multi-aspect framework for counter narrative evaluation using large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers). Kevin Duh, Helena Gomez, and Steven Bethard (Eds.), Association for Computational Linguistics, 147\u2013168. DOI: 10.18653\/v1\/2024.naacl-short.14"},{"key":"e_1_3_2_97_2","doi-asserted-by":"crossref","first-page":"1305","DOI":"10.18653\/v1\/2022.acl-long.93","volume-title":"Proceedings of 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Eddine Moussa Kamal","year":"2022","unstructured":"Moussa Kamal Eddine, Guokan Shang, Antoine Tixier, and Michalis Vazirgiannis. 2022. FrugalScore: Learning cheaper, lighter and faster evaluation metrics for automatic text generation. In Proceedings of 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.), Association for Computational Linguistics, 1305\u20131318. DOI: 10.18653\/v1\/2022.acl-long.93"},{"key":"e_1_3_2_98_2","doi-asserted-by":"crossref","first-page":"9475","DOI":"10.18653\/v1\/2024.findings-emnlp.553","volume-title":"Findings of the Association for Computational Linguistics (EMNLP \u201924)","author":"Kargupta Priyanka","year":"2024","unstructured":"Priyanka Kargupta, Ishika Agarwal, Dilek Hakkani Tur, and Jiawei Han. 2024. Instruct, not assist: LLM-based Multi-Turn planning and hierarchical questioning for socratic code debugging. In Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 9475\u20139495. DOI: 10.18653\/v1\/2024.findings-emnlp.553"},{"key":"e_1_3_2_99_2","unstructured":"Alex Kim Keonwoo Kim and Sangwon Yoon. 2024. DEBATE: Devil\u2019s advocate-based assessment and text evaluation. arXiv:2405.09935. Retrieved from https:\/\/arxiv.org\/abs\/2405.09935"},{"key":"e_1_3_2_100_2","doi-asserted-by":"crossref","unstructured":"Abishek Komma Nagesh Panyam Chandrasekarasastry Timothy Leffel Anuj Goyal Angeliki Metallinou Spyros Matsoukas and Aram Galstyan. 2023. Toward more accurate and generalizable evaluation metrics for task-oriented dialogs. arXiv:2306.03984. Retrieved from https:\/\/arxiv.org\/abs\/2306.03984","DOI":"10.18653\/v1\/2023.acl-industry.19"},{"key":"e_1_3_2_101_2","doi-asserted-by":"crossref","first-page":"17020","DOI":"10.18653\/v1\/2024.emnlp-main.946","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Kurtic Eldar","year":"2024","unstructured":"Eldar Kurtic, Amir Moeini, and Dan Alistarh. 2024. Mathador-LM: A dynamic benchmark for mathematical reasoning on large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 17020\u201317027. DOI: 10.18653\/v1\/2024.emnlp-main.946"},{"key":"e_1_3_2_102_2","unstructured":"Wai-Chung Kwan Xingshan Zeng Yuxin Jiang Yufei Wang Liangyou Li Lifeng Shang Xin Jiang Qun Liu and Kam-Fai Wong. 2024. MT-Eval: A multi-turn capabilities evaluation benchmark for large language models. arXiv:2401.16745. Retrieved from https:\/\/arxiv.org\/abs\/2401.16745"},{"key":"e_1_3_2_103_2","doi-asserted-by":"crossref","first-page":"707","DOI":"10.18653\/v1\/2023.acl-industry.68","volume-title":"Proceedings of 61st Annual Meeting of the Association for Computational Linguistics (Volume 5: Industry Track)","author":"Kwon Deuksin","year":"2023","unstructured":"Deuksin Kwon, Sunwoo Lee, Ki Hyun Kim, Seojin Lee, Taeyoon Kim, and Eric Davis. 2023. What, when, and how to ground: Designing user persona-aware conversational agents for engaging dialogue. In Proceedings of 61st Annual Meeting of the Association for Computational Linguistics (Volume 5: Industry Track). Sunayana Sitaram, Beata Beigman Klebanov, and Jason D. Williams (Eds.), Association for Computational Linguistics, 707\u2013719. DOI: 10.18653\/v1\/2023.acl-industry.68"},{"key":"e_1_3_2_104_2","unstructured":"Minseo Kwon Yaesol Kim and Young J. Kim. 2024. Fast and accurate task planning using neuro-symbolic language models and multi-level goal decomposition. arXiv:2409.19250. Retrieved from https:\/\/arxiv.org\/abs\/2409.19250"},{"key":"e_1_3_2_105_2","first-page":"228","volume-title":"Proceedings of the 2nd Workshop on Statistical Machine Translation (StatMT \u201907)","author":"Lavie Alon","unstructured":"Alon Lavie and Abhaya Agarwal. 2007. Meteor: An automatic metric for MT evaluation with high levels of correlation with human judgments. In Proceedings of the 2nd Workshop on Statistical Machine Translation (StatMT \u201907), Association for Computational Linguistics, 228\u2013231."},{"key":"e_1_3_2_106_2","unstructured":"Seolhwa Lee Heuiseok Lim and Jo\u00e3o Sedoc. 2020. An evaluation protocol for generative conversational systems. arXiv:2010.12741. Retrieved from https:\/\/arxiv.org\/abs\/2010.12741"},{"key":"e_1_3_2_107_2","doi-asserted-by":"crossref","unstructured":"Fangyu Lei Qian Liu Yiming Huang Shizhu He Jun Zhao and Kang Liu. 2024. S3Eval: A synthetic scalable systematic evaluation suite for large language models. arXiv:2310.15147. Retrieved from https:\/\/arxiv.org\/abs\/2310.15147","DOI":"10.18653\/v1\/2024.naacl-long.69"},{"key":"e_1_3_2_108_2","unstructured":"Quinn Leng Jacob Portes Sam Havens Matei Zaharia and Michael Carbin. 2024. Long context RAG performance of large language models. arXiv:2411.03538. Retrieved from https:\/\/arxiv.org\/abs\/2411.03538"},{"key":"e_1_3_2_109_2","first-page":"1774","volume-title":"Findings of the Association for Computational Linguistics (ACL \u201923)","author":"Li Daliang","year":"2023","unstructured":"Daliang Li, Ankit Singh Rawat, Manzil Zaheer, Xin Wang, Michal Lukasik, Andreas Veit, Felix Yu, and Sanjiv Kumar. 2023. Large language models with controllable working memory. In Findings of the Association for Computational Linguistics (ACL \u201923). Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.), Association for Computational Linguistics, 1774\u20131793. DOI: 10.18653\/v1\/2023.findings-acl.112"},{"key":"e_1_3_2_110_2","volume-title":"NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following","author":"Li Dacheng","year":"2023","unstructured":"Dacheng Li, Rulin Shao, Anze Xie, Ying Sheng, Lianmin Zheng, Joseph Gonzalez, Ion Stoica, Xuezhe Ma, and Hao Zhang. 2023. How long can context length of open-source LLMs truly promise? In NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following. Retrieved from https:\/\/openreview.net\/forum?id=LywifFNXV5"},{"key":"e_1_3_2_111_2","unstructured":"Kenneth Li Yiming Wang Fernanda Vi\u00e9gas and Martin Wattenberg. 2024. Dialogue action tokens: Steering language models in goal-directed dialogue with a multi-turn planner. arXiv:2406.11978. Retrieved from https:\/\/arxiv.org\/abs\/2406.11978"},{"key":"e_1_3_2_112_2","unstructured":"Margaret Li Jason Weston and Stephen Roller. 2019. ACUTE-EVAL: Improved dialogue evaluation with optimized questions and multi-turn comparisons. arXiv:1909.03087. Retrieved from https:\/\/arxiv.org\/abs\/1909.03087"},{"key":"e_1_3_2_113_2","first-page":"3102","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Li Minghao","year":"2023","unstructured":"Minghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song, Hangyu Li, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li. 2023. API-Bank: A comprehensive benchmark for tool-augmented LLMs. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Houda Bouamor, Juan Pino, and Kalika Bali (Eds.), Association for Computational Linguistics, 3102\u20133116. DOI: 10.18653\/v1\/2023.emnlp-main.187"},{"key":"e_1_3_2_114_2","unstructured":"Xinzhe Li. 2024. A review of prominent paradigms for LLM-based agents: Tool use (including RAG) planning and feedback learning. arXiv:2406.05804. Retrieved from https:\/\/arxiv.org\/abs\/2406.05804"},{"key":"e_1_3_2_115_2","unstructured":"Yanran Li Hui Su Xiaoyu Shen Wenjie Li Ziqiang Cao and Shuzi Niu. 2017. DailyDialog: A manually labelled multi-turn dialogue dataset. arXiv:1710.03957. Retrieved from https:\/\/arxiv.org\/abs\/1710.03957"},{"key":"e_1_3_2_116_2","first-page":"1652","volume-title":"Proceedings of 28th International Conference on Computational Linguistics","author":"Li Yixuan","year":"2020","unstructured":"Yixuan Li, Fangzhen Wu, Yanan Jiang, Yongbin Xie, and Xiaojiang Xu. 2020. Improve the response diversity of multi-turn dialogue system with bidirectional distillation. In Proceedings of 28th International Conference on Computational Linguistics. International Committee on Computational Linguistics, 1652\u20131662. DOI: 10.18653\/v1\/2020.coling-main.145"},{"key":"e_1_3_2_117_2","unstructured":"Youquan Li Miao Zheng Fan Yang Guosheng Dong Bin Cui Weipeng Chen Zenan Zhou and Wentao Zhang. 2024. FB-Bench: A fine-grained multi-task benchmark for evaluating LLMs\u2019 responsiveness to human feedback. arXiv:2410.09412. Retrieved from https:\/\/arxiv.org\/abs\/2410.09412"},{"key":"e_1_3_2_118_2","unstructured":"Xinnian Liang Bing Wang Hui Huang Shuangzhi Wu Peihao Wu Lu Lu Zejun Ma and Zhoujun Li. 2023. Unleashing infinite-length input capacity for large-scale language models with self-controlled memory system. arXiv:2304.13343. Retrieved from https:\/\/arxiv.org\/2304.13343"},{"key":"e_1_3_2_119_2","unstructured":"Jonathan Light Min Cai Weiqin Chen Guanzhi Wang Xiusi Chen Wei Cheng Yisong Yue and Ziniu Hu. 2024. Strategist: Learning strategic skills by LLMs via bi-level tree search. arXiv:2408.10635. Retrieved from https:\/\/arxiv.org\/abs\/2408.10635"},{"key":"e_1_3_2_120_2","unstructured":"Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out. Association for Computational Linguistics 74\u201381. Retrieved from https:\/\/aclanthology.org\/W04-1013"},{"key":"e_1_3_2_121_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2977471"},{"key":"e_1_3_2_122_2","first-page":"47","volume-title":"Proceedings of 5th Workshop on NLP for Conversational AI (NLP4ConvAI \u201923)","author":"Lin Yen-Ting","year":"2023","unstructured":"Yen-Ting Lin and Yun-Nung Chen. 2023. LLM-Eval: Unified multi-dimensional automatic evaluation for open-domain conversations with large language models. In Proceedings of 5th Workshop on NLP for Conversational AI (NLP4ConvAI \u201923). Yun-Nung Chen and Abhinav Rastogi (Eds.), Association for Computational Linguistics, 47\u201358. DOI: 10.18653\/v1\/2023.nlp4convai-1.5"},{"key":"e_1_3_2_123_2","first-page":"259","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (ACL \u201923)","author":"Liu Anqi","year":"2023","unstructured":"Anqi Liu, Bo Wang, Yue Tan, Dongming Zhao, Kun Huang, Ruifang He, and Yuexian Hou. 2023. MTGP: Multi-turn target-oriented dialogue guided by generative global path with flexible turns. In Proceedings of the Findings of the Association for Computational Linguistics (ACL \u201923). Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.), Association for Computational Linguistics, 259\u2013271. DOI: 10.18653\/v1\/2023.findings-acl.18"},{"key":"e_1_3_2_124_2","first-page":"4649","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924)","author":"Liu Bing","year":"2024","unstructured":"Bing Liu, Zhou Jianxiang, Dan Meng, and Haonan Lu. 2024. An evaluation mechanism of LLM-based agents on manipulating APIs. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 4649\u20134662. DOI: 10.18653\/v1\/2024.findings-emnlp.267"},{"key":"e_1_3_2_125_2","unstructured":"Hang Liu Meng Chen Youzheng Wu Xiaodong He and Bowen Zhou. 2021. Conversational query rewriting with self-supervised learning. arXiv:2102.04708. Retrieved from https:\/\/arxiv.org\/abs\/2102.04708"},{"key":"e_1_3_2_126_2","doi-asserted-by":"crossref","first-page":"16781","DOI":"10.18653\/v1\/2024.findings-emnlp.978","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924)","author":"Liu Hao","year":"2024","unstructured":"Hao Liu, Zi-Yi Dou, Yixin Wang, Nanyun Peng, and Yisong Yue. 2024. Uncertainty calibration for tool-using language agents. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 16781\u201316805. DOI: 10.18653\/v1\/2024.findings-emnlp.978"},{"key":"e_1_3_2_127_2","first-page":"1096","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track","author":"Liu Junhua","year":"2024","unstructured":"Junhua Liu, Tan Yong Keat, Bin Fu, and Kwan Hui Lim. 2024. Lara: Linguistic-adaptive retrieval-augmentation for multi-turn intent classification. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track. Franck Dernoncourt, Daniel Preo\u0163iuc-Pietro, and Anastasia Shimorina (Eds.), Association for Computational Linguistics, 1096\u20131106. DOI: 10.18653\/v1\/2024.emnlp-industry.82"},{"key":"e_1_3_2_128_2","unstructured":"Lei Liu Xiaoyan Yang Yue Shen Binbin Hu Zhiqiang Zhang Jinjie Gu and Guannan Zhang. 2023. Think-in-Memory: Recalling and post-thinking enable LLMs with long-term memory. arXiv:2311.08719. Retrieved from https:\/\/arxiv.org\/abs\/2311.08719"},{"key":"e_1_3_2_129_2","unstructured":"Na Liu Liangyu Chen Xiaoyu Tian Wei Zou Kaijiang Chen and Ming Cui. 2024. From LLM to conversational agent: A memory enhanced architecture with fine-tuning of large language models. arXiv:2401.02777. Retrieved from https:\/\/arxiv.org\/abs\/2401.02777"},{"key":"e_1_3_2_130_2","unstructured":"Shuo Liu Kaining Ying Hao Zhang Yue Yang Yuqi Lin Tianle Zhang Chuanhao Li Yu Qiao Ping Luo Wenqi Shao and Kaipeng Zhang. 2024. ConvBench: A multi-turn conversation evaluation benchmark with hierarchical capability for large vision-language models. arXiv:2403.20194. Retrieved from https:\/\/arxiv.org\/abs\/2403.20194"},{"key":"e_1_3_2_131_2","first-page":"470","volume-title":"Proceedings of62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Liu Xinyi","year":"2024","unstructured":"Xinyi Liu, Pinxin Liu, and Hangfeng He. 2024. An empirical analysis on large language models in debate evaluation. In Proceedings of62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.), Association for Computational Linguistics, 470\u2013487. DOI: 10.18653\/v1\/2024.acl-short.44"},{"key":"e_1_3_2_132_2","unstructured":"Xiao Liu Hao Yu Hanchen Zhang Yifan Xu Xuanyu Lei Hanyu Lai Yu Gu Hangliang Ding Kaiwen Men Kejuan Yang et al. 2023. AgentBench: Evaluating LLMs as agents. arXiv:2308.03688. Retrieved from https:\/\/arxiv.org\/abs\/2308.03688"},{"key":"e_1_3_2_133_2","doi-asserted-by":"publisher","unstructured":"Yizhu Liu Qi Jia and Kenny Zhu. 2022. Reference-free summarization evaluation via semantic correlation and compression ratio. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Marine Carpuat Marie-Catherine de Marneffe and Ivan Vladimir Meza Ruiz (Eds.) Association for Computational Linguistics 2109\u20132115. DOI: 10.18653\/v1\/2022.naacl-main.153","DOI":"10.18653\/v1\/2022.naacl-main.153"},{"key":"e_1_3_2_134_2","unstructured":"Ziyu Liu Tao Chu Yuhang Zang Xilin Wei Xiaoyi Dong Pan Zhang Zijian Liang Yuanjun Xiong Yu Qiao Dahua Lin et al. 2024. MMDU: A multi-turn multi-image dialog understanding benchmark and instruction-tuning dataset for LVLMs. arXiv:2406.11833. Retrieved from https:\/\/arxiv.org\/abs\/2406.11833"},{"key":"e_1_3_2_135_2","unstructured":"Zijun Liu Yanzhe Zhang Peng Li Yang Liu and Diyi Yang. 2024. A dynamic LLM-powered agent network for task-oriented agent collaboration. arXiv:2310.02170. Retrieved from https:\/\/arxiv.org\/abs\/2310.02170"},{"key":"e_1_3_2_136_2","unstructured":"Zhiling Luo Qiankun Shi Sha Zhao Wei Zhou Haiqing Chen Yuankai Ma and Haitao Leng. 2022. AliCHI: A large-scale multi-modal dataset and automated evaluation tool for human-like dialogue systems. arXiv:2212.05489. Retrieved from https:\/\/arxiv.org\/abs\/2212.05489"},{"key":"e_1_3_2_137_2","unstructured":"Daoming Lyu Bo Liu and Jianshu Chen. 2022. PRIMA: Planner-reasoner inside a multi-task reasoning agent. arXiv:2202.00531. Retrieved from https:\/\/arxiv.org\/abs\/2202.00531"},{"key":"e_1_3_2_138_2","unstructured":"Aman Madaan Niket Tandon Prakhar Gupta Skyler Hallinan Luyu Gao Sarah Wiegreffe Uri Alon Nouha Dziri Shrimai Prabhumoye Yiming Yang et al. 2023. Self-Refine: Iterative refinement with self-feedback. arXiv:2303.17651. Retrieved from https:\/\/arxiv.org\/abs\/2303.17651"},{"key":"e_1_3_2_139_2","doi-asserted-by":"crossref","unstructured":"Adyasha Maharana Dong-Ho Lee Sergey Tulyakov Mohit Bansal Francesco Barbieri and Yuwei Fang. 2024. Evaluating very long-term conversational memory of LLM agents. arXiv:2402.17753. Retrieved from https:\/\/arxiv.org\/abs\/2402.17753","DOI":"10.18653\/v1\/2024.acl-long.747"},{"key":"e_1_3_2_140_2","unstructured":"Shengyu Mao Xiaohan Wang Mengru Wang Yong Jiang Pengjun Xie Fei Huang and Ningyu Zhang. 2024. Editing personality for large language models. arXiv:2310.02168. Retrieved from https:\/\/arxiv.org\/abs\/2310.02168"},{"key":"e_1_3_2_141_2","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Mehnaz Laiba","year":"2021","unstructured":"Laiba Mehnaz, Debanjan Mahata, Rakesh Gosangi, Uma Sushmitha Gunturi, Riya Jain, Gauri Gupta, Amardeep Kumar, Isabelle G. Lee, Anish Acharya, and Rajiv Ratn Shah. 2021. GupShup: Summarizing open-domain code-switched conversations. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:243865218"},{"key":"e_1_3_2_142_2","first-page":"2831","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Mehrabi Ninareh","year":"2022","unstructured":"Ninareh Mehrabi, Ahmad Beirami, Fred Morstatter, and Aram Galstyan. 2022. Robust conversational agents against imperceptible toxicity triggers. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz (Eds.), Association for Computational Linguistics, 2831\u20132847. DOI: 10.18653\/v1\/2022.naacl-main.204"},{"key":"e_1_3_2_143_2","unstructured":"Shikib Mehri Mihail Eric and Dilek Hakkani-Tur. 2020. DialoGLUE: A natural language understanding benchmark for task-oriented dialogue. arXiv:2009.13570. Retrieved from https:\/\/arxiv.org\/abs\/2009.13570"},{"key":"e_1_3_2_144_2","doi-asserted-by":"crossref","DOI":"10.3390\/app10030762","article-title":"Human annotated dialogues dataset for natural conversational agents","author":"Merdivan Erinc","year":"2020","unstructured":"Erinc Merdivan, Deepika Singh, Sten Hanke, Johannes Kropf, Andreas Holzinger, and Matthieu Geist. 2020. Human annotated dialogues dataset for natural conversational agents. Applied Sciences. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:210874219","journal-title":"Applied Sciences"},{"key":"e_1_3_2_145_2","doi-asserted-by":"publisher","DOI":"10.3390\/app10030762"},{"key":"e_1_3_2_146_2","volume-title":"Proceedings of 12th International Conference on Learning Representations (ICLR \u201924)","author":"Mialon Gr\u00e9goire","year":"2024","unstructured":"Gr\u00e9goire Mialon, Cl\u00e9mentine Fourrier, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2024. GAIA: A benchmark for general AI assistants. In Proceedings of 12th International Conference on Learning Representations (ICLR \u201924). Retrieved from OpenReview.net. https:\/\/openreview.net\/forum?id=fibxvahvs3"},{"key":"e_1_3_2_147_2","doi-asserted-by":"crossref","first-page":"14420","DOI":"10.18653\/v1\/2024.findings-emnlp.843","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924)","author":"Miehling Erik","year":"2024","unstructured":"Erik Miehling, Manish Nagireddy, Prasanna Sattigeri, Elizabeth M. Daly, David Piorkowski, and John T. Richards. 2024. Language models in dialogue: Conversational maxims for human-AI interactions. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 14420\u201314437. DOI: 10.18653\/v1\/2024.findings-emnlp.843"},{"key":"e_1_3_2_148_2","unstructured":"Eric Mitchell Charles Lin Antoine Bosselut Christopher D. Manning and Chelsea Finn. 2022. Memory-based model editing at scale. arXiv:2206.06520. Retrieved from https:\/\/arxiv.org\/abs\/2206.06520"},{"key":"e_1_3_2_149_2","unstructured":"Fengran Mo Chen Qu Kelong Mao Yihong Wu Zhan Su Kaiyu Huang and Jian-Yun Nie. 2024. Aligning query representation with rewritten query and relevance judgments in conversational search. arXiv:2407.20189. Retrieved from https:\/\/arxiv.org\/abs\/2407.20189"},{"key":"e_1_3_2_150_2","unstructured":"Ali Modarressi Ayyoob Imani Mohsen Fayyaz and Hinrich Sch\u00fctze. 2024. RET-LLM: Towards a general read-write memory for large language models. arXiv:2305.14322. Retrieved from https:\/\/arxiv.org\/abs\/2305.14322"},{"key":"e_1_3_2_151_2","unstructured":"Behrad Moniri Hamed Hassani and Edgar Dobriban. 2024. Evaluating the performance of large language models via debates. arXiv:2406.11044. Retrieved from https:\/\/arxiv.org\/abs\/2406.11044"},{"key":"e_1_3_2_152_2","unstructured":"Christian Muise Tathagata Chakraborti Shubham Agarwal Ondrej Bajgar Arunima Chaudhary Luis A. Lastras-Montano Josef Ondrej Miroslav Vodolan and Charlie Wiecha. 2019. Planning for goal-oriented dialogue systems. arXiv:1910.08137. Retrieved from https:\/\/arxiv.org\/abs\/1910.08137"},{"key":"e_1_3_2_153_2","doi-asserted-by":"crossref","first-page":"6909","DOI":"10.18653\/v1\/2023.findings-emnlp.461","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Muthusamy Vinod","year":"2023","unstructured":"Vinod Muthusamy, Yara Rizk, Kiran Kate, Praveen Venkateswaran, Vatche Isahagian, Ashu Gulati, and Parijat Dube. 2023. Towards large language model-based personal agents in the enterprise: Current trends and open problems. In Findings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.), Association for Computational Linguistics, 6909\u20136921. DOI: 10.18653\/v1\/2023.findings-emnlp.461"},{"key":"e_1_3_2_154_2","first-page":"4556","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (NAACL \u201924)","author":"Nan Linyong","year":"2024","unstructured":"Linyong Nan, Ellen Zhang, Weijin Zou, Yilun Zhao, Wenfei Zhou, and Arman Cohan. 2024. On evaluating the integration of reasoning and action in LLM agents with database question answering. In Proceedings of the Findings of the Association for Computational Linguistics (NAACL \u201924). Kevin Duh, Helena Gomez, and Steven Bethard (Eds.), Association for Computational Linguistics, 4556\u20134579. DOI: 10.18653\/v1\/2024.findings-naacl.284"},{"key":"e_1_3_2_155_2","unstructured":"Jinjie Ni Tom Young Vlad Pandelea Fuzhao Xue and Erik Cambria. 2022. Recent advances in deep learning based dialogue systems: A systematic survey. arXiv:2105.04387. Retrieved from https:\/\/arxiv.org\/abs\/2105.04387"},{"key":"e_1_3_2_156_2","volume-title":"Proceedings of 11th International Conference on Learning Representations","author":"Nijkamp Erik","year":"2023","unstructured":"Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023. CodeGen: An open large language model for code with multi-turn program synthesis. In Proceedings of 11th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=iaYcJKpY2B_"},{"key":"e_1_3_2_157_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3249757"},{"key":"e_1_3_2_158_2","doi-asserted-by":"crossref","first-page":"121","DOI":"10.18653\/v1\/W19-4114","volume-title":"Proceedings of 1st Workshop on NLP for Conversational AI","author":"Olabiyi Oluwatobi","year":"2019","unstructured":"Oluwatobi Olabiyi, Alan O. Salimov, Anish Khazane, Erik, and Mueller, Su. 2019. Multi-turn dialogue response generation in an adversarial learning framework. In Proceedings of 1st Workshop on NLP for Conversational AI. Yun-Nung Chen, Tania Bedrax-Weiss, Dilek Hakkani-Tur, Anuj Kumar, Mike Lewis, Thang-Minh Luong, Pei-Hao and Tsung-Hsien Wen (Eds.), Association for Computational Linguistics, 121\u2013132. DOI: 10.18653\/v1\/W19-4114"},{"key":"e_1_3_2_159_2","unstructured":"OpenAI. 2023. ChatGPT. Retrieved November 19 2024 from https:\/\/chat.openai.com"},{"key":"e_1_3_2_160_2","unstructured":"OpenAI. 2024. ChatGPT Plugins. Retrieved November 21 2024 from https:\/\/openai.com\/index\/chatgpt-plugins\/"},{"key":"e_1_3_2_161_2","unstructured":"OpenAI. 2024. GPT-4 Technical Report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_2_162_2","unstructured":"OpenAI. 2025. Learning to Reason with LLMs. Retrieved January 6 from https:\/\/openai.com\/index\/learning-to-reason-with-llms\/"},{"key":"e_1_3_2_163_2","unstructured":"Long Ouyang Jeff Wu Xu Jiang Diogo Almeida Carroll L. Wainwright Pamela Mishkin Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray et al. 2022. Training language models to follow instructions with human feedback. arXiv.02155. Retrieved from https:\/\/arxiv.org\/abs\/2203.02155"},{"key":"e_1_3_2_164_2","doi-asserted-by":"crossref","first-page":"n71","DOI":"10.1136\/bmj.n71","article-title":"The PRISMA 2020 statement: An updated guideline for reporting systematic reviews","volume":"372","author":"Page Matthew J.","year":"2021","unstructured":"Matthew J. Page, Joanne E. McKenzie, Patrick M. Bossuyt, Isabelle Boutron, Tammy C. Hoffmann, Cynthia D. Mulrow, Larissa Shamseer, Jennifer M. Tetzlaff, Elie A. Akl, Sue E. Brennan, et al. 2021. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 372 (2021), n71.","journal-title":"BMJ"},{"key":"e_1_3_2_165_2","unstructured":"Arka Pal Deep Karkhanis Manley Roberts Samuel Dooley Arvind Sundararajan and Siddartha Naidu. 2023. Giraffe: Adventures in expanding context lengths in LLMs. arXiv:2308.10882. Retrieved from https:\/\/arxiv.org\/abs\/2308.10882"},{"key":"e_1_3_2_166_2","first-page":"311","volume-title":"Proceedings of 40th Annual Meeting on Association for Computational Linguistics (ACL \u201902)","author":"Papineni Kishore","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of 40th Annual Meeting on Association for Computational Linguistics (ACL \u201902). Association for Computational Linguistics, 311\u2013318. DOI: 10.3115\/1073083.1073135"},{"key":"e_1_3_2_167_2","unstructured":"ChaeHun Park Minseok Choi Dohyun Lee and Jaegul Choo. 2024. PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison. arXiv:2404.01015. Retrieved from https:\/\/arxiv.org\/abs\/2404.01015"},{"key":"e_1_3_2_168_2","unstructured":"Joon Sung Park Joseph C. O\u2019Brien Carrie J. Cai Meredith Ringel Morris Percy Liang and Michael S. Bernstein. 2023. Generative agents: Interactive ##. arXiv:2304.03442. Retrieved from https:\/\/arxiv.org\/abs\/2304.03442"},{"key":"e_1_3_2_169_2","unstructured":"Shishir G. Patil Tianjun Zhang Xin Wang and Joseph E. Gonzalez. 2023. Gorilla: Large language model connected with massive APIs. arXiv:2305.15334. Retrieved from https:\/\/arxiv.org\/abs\/2305.15334"},{"key":"e_1_3_2_170_2","doi-asserted-by":"crossref","unstructured":"Vitou Phy Yang Zhao and Akiko Aizawa. 2020. Deconstruct to reconstruct a configurable evaluation metric for open-domain dialogue systems. arXiv:2011.00483. Retrieved from https:\/\/arxiv.org\/abs\/2011.00483","DOI":"10.18653\/v1\/2020.coling-main.368"},{"key":"e_1_3_2_171_2","doi-asserted-by":"crossref","unstructured":"Matt Post. 2018. A call for clarity in reporting BLEU scores. arXiv:1804.08771. Retrieved from https:\/\/arxiv.org\/abs\/1804.08771","DOI":"10.18653\/v1\/W18-6319"},{"key":"e_1_3_2_172_2","doi-asserted-by":"crossref","first-page":"4226","DOI":"10.18653\/v1\/2024.findings-naacl.264","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (NAACL \u201924)","author":"Prasad Archiki","year":"2024","unstructured":"Archiki Prasad, Alexander Koller, Mareike Hartmann, Peter Clark, Ashish Sabharwal, Mohit Bansal, and Tushar Khot. 2024. ADaPT: As-needed decomposition and planning with language models. In Proceedings of the Findings of the Association for Computational Linguistics (NAACL \u201924). Kevin Duh, Helena Gomez, and Steven Bethard (Eds.), Association for Computational Linguistics, 4226\u20134252. DOI: 10.18653\/v1\/2024.findings-naacl.264"},{"key":"e_1_3_2_173_2","doi-asserted-by":"crossref","unstructured":"Ofir Press Muru Zhang Sewon Min Ludwig Schmidt Noah A. Smith and Mike Lewis. 2023. Measuring and narrowing the compositionality gap in language models. arXiv:2210.03350. Retrieved from https:\/\/arxiv.org\/abs\/2210.03350","DOI":"10.18653\/v1\/2023.findings-emnlp.378"},{"key":"e_1_3_2_174_2","doi-asserted-by":"crossref","first-page":"2333","DOI":"10.18653\/v1\/2024.findings-naacl.151","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (NAACL \u201924)","author":"Qiao Shanbao","year":"2024","unstructured":"Shanbao Qiao, Xuebing Liu, and Seung-Hoon Na. 2024. COMEM: In-context retrieval-augmented mass-editing memory in large language models. In Proceedings of the Findings of the Association for Computational Linguistics (NAACL \u201924). Kevin Duh, Helena Gomez, and Steven Bethard (Eds.), Association for Computational Linguistics, 2333\u20132347. DOI: 10.18653\/v1\/2024.findings-naacl.151"},{"key":"e_1_3_2_175_2","unstructured":"Yujia Qin Shihao Liang Yining Ye Kunlun Zhu Lan Yan Yaxi Lu Yankai Lin Xin Cong Xiangru Tang Bill Qian et al. 2023. ToolLLM: Facilitating large language models to master 16000+ real-world APIs. arXiv:2307.16789. Retrieved from https:\/\/arxiv.org\/abs\/2307.16789"},{"key":"e_1_3_2_176_2","first-page":"615","volume-title":"Proceedings of the Findings of the Association for Computational LinguisticsEMNLP \u201924)","author":"Qiu Huachuan","year":"2024","unstructured":"Huachuan Qiu, Hongliang He, Shuai Zhang, Anqi Li, and Zhenzhong Lan. 2024. SMILE: Single-turn to multi-turn inclusive language expansion via ChatGPT for mental health support. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 615\u2013636. DOI: 10.18653\/v1\/2024.findings-emnlp.34"},{"key":"e_1_3_2_177_2","doi-asserted-by":"publisher","DOI":"10.1109\/MASSP.1986.1165342"},{"key":"e_1_3_2_178_2","unstructured":"Alec Radford and Karthik Narasimhan. 2018. Improving Language Understanding By Generative Pre-training. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:49313245"},{"key":"e_1_3_2_179_2","unstructured":"Alec Radford Jeff Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language Models Are Unsupervised Multitask Learners. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:160025533"},{"key":"e_1_3_2_180_2","volume-title":"Proceedings of 23rd ACM International Conference on Intelligent Virtual Agents","author":"Reimann Merle M.","year":"2023","unstructured":"Merle M. Reimann, Catharine Oertel, Florian A. Kunneman, and Koen V. Hindriks. 2023. Predicting interaction quality aspects using level-based scores for conversational agents. In Proceedings of 23rd ACM International Conference on Intelligent Virtual Agents. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:266437223"},{"key":"e_1_3_2_181_2","unstructured":"Liliang Ren Mankeerat Sidhu Qi Zeng Revanth Gangi Reddy Heng Ji and ChengXiang Zhai. 2023. C-PMI: Conditional pointwise mutual information for turn-level dialogue evaluation. arXiv:2306.15245. Retrieved from https:\/\/arxiv.org\/abs\/2306.15245"},{"key":"e_1_3_2_182_2","doi-asserted-by":"crossref","first-page":"503","DOI":"10.1108\/00220410410560582","article-title":"Understanding inverse document frequency: On theoretical arguments for IDF","volume":"60","author":"Robertson Stephen E.","year":"2004","unstructured":"Stephen E. Robertson. 2004. Understanding inverse document frequency: On theoretical arguments for IDF. Journal of Documentation 60 (2004), 503\u2013520. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:8864928","journal-title":"Journal of Documentation"},{"key":"e_1_3_2_183_2","doi-asserted-by":"crossref","first-page":"6030","DOI":"10.18653\/v1\/2023.emnlp-main.368","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Ruiz-Dolz Ramon","year":"2023","unstructured":"Ramon Ruiz-Dolz, Stella Heras, and Ana Garcia. 2023. Automatic debate evaluation with argumentation semantics and natural language argument graph networks. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Houda Bouamor, Juan Pino, and Kalika Bali (Eds.), Association for Computational Linguistics, 6030\u20136040. DOI: 10.18653\/v1\/2023.emnlp-main.368"},{"key":"e_1_3_2_184_2","doi-asserted-by":"crossref","first-page":"3468","DOI":"10.18653\/v1\/2023.findings-emnlp.226","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201923)","author":"Sarch Gabriel","year":"2023","unstructured":"Gabriel Sarch, Yue Wu, Michael Tarr, and Katerina Fragkiadaki. 2023. Open-ended instructable embodied agents with memory-augmented large language models. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201923). Houda Bouamor, Juan Pino, and Kalika Bali (Eds.), Association for Computational Linguistics, 3468\u20133500. DOI: 10.18653\/v1\/2023.findings-emnlp.226"},{"key":"e_1_3_2_185_2","first-page":"6874","volume-title":"Proceedings of 12Language Resources and Evaluation Conference","author":"Sathe Aalok","year":"2020","unstructured":"Aalok Sathe, Salar Ather, Tuan Manh Le, Nathan Perry, and Joonsuk Park. 2020. Automated fact-checking of claims from Wikipedia. In Proceedings of 12th Language Resources and Evaluation Conference. Nicoletta Calzolari, Fr\u00e9d\u00e9ric B\u00e9chet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, et al. (Eds.), European Language Resources Association, 6874\u20136882. Retrieved from https:\/\/aclanthology.org\/2020.lrec-1.849\/"},{"key":"e_1_3_2_186_2","doi-asserted-by":"crossref","first-page":"101","DOI":"10.18653\/v1\/2021.fever-1.11","volume-title":"Proceedings of 4th Workshop on Fact Extraction and VERification (FEVER)","author":"Sathe Aalok","year":"2021","unstructured":"Aalok Sathe and Joonsuk Park. 2021. Automatic fact-checking with document-level annotations using BERT and multiple instance learning. In Proceedings of 4th Workshop on Fact Extraction and VERification (FEVER). Rami Aly, Christos Christodoulopoulos, Oana Cocarascu, Zhijiang Guo, Arpit Mittal, Michael Schlichtkrull, James Thorne, and Andreas Vlachos (Eds.), Association for Computational Linguistics, 101\u2013107. DOI: 10.18653\/v1\/2021.fever-1.11"},{"key":"e_1_3_2_187_2","first-page":"4036","volume-title":"Proceedings of 33rd ACM International Conference on Information and Knowledge Management (CIKM \u201924)","author":"Setty Ritvik","year":"2024","unstructured":"Ritvik Setty and Vinay Setty. 2024. QuestGen: Effectiveness of question generation methods for fact-checking applications. In Proceedings of 33rd ACM International Conference on Information and Knowledge Management (CIKM \u201924). ACM, 4036\u20134040. DOI: 10.1145\/3627673.3679985"},{"key":"e_1_3_2_188_2","volume-title":"Proceedings of 27th International Conference on Intelligent User Interfaces","author":"Shen Junxiao","year":"2022","unstructured":"Junxiao Shen, Boyin Yang, John J. Dudley, and Per Ola Kristensson. 2022. KWickChat: A multi-turn dialogue system for AAC using context-aware sentence generation by bag-of-keywords. In Proceedings of 27th International Conference on Intelligent User Interfaces. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:247585177"},{"key":"e_1_3_2_189_2","doi-asserted-by":"crossref","unstructured":"Yongliang Shen Kaitao Song Xu Tan Dongsheng Li Weiming Lu and Yueting Zhuang. 2023. HuggingGPT: Solving AI tasks with ChatGPT and its friends in hugging face. arXiv:2303.17580. Retrieved from https:\/\/arxiv.org\/abs\/2303.17580","DOI":"10.52202\/075280-1657"},{"key":"e_1_3_2_190_2","first-page":"774","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track","author":"Shen Yuanhao","year":"2024","unstructured":"Yuanhao Shen, Xiaodan Zhu, and Lei Chen. 2024. SMARTCAL: An approach to self-aware tool-use evaluation and calibration. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track. Franck Dernoncourt, Daniel Preo\u0163iuc-Pietro, and Anastasia Shimorina (Eds.), Association for Computational Linguistics, 774\u2013789. DOI: 10.18653\/v1\/2024.emnlp-industry.59"},{"key":"e_1_3_2_191_2","first-page":"2312","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Shi Wentao","year":"2024","unstructured":"Wentao Shi, Mengqi Yuan, Junkang Wu, Qifan Wang, and Fuli Feng. 2024. Direct multi-turn preference optimization for language agents. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 2312\u20132324. DOI: 10.18653\/v1\/2024.emnlp-main.138"},{"key":"e_1_3_2_192_2","first-page":"10642","volume-title":"Proceedings of the Findings of the Association for Computational LinguisticsEMNLP \u201924)","author":"Shi Zhengliang","year":"2024","unstructured":"Zhengliang Shi, Shen Gao, Xiuyi Chen, Yue Feng, Lingyong Yan, Haibo Shi, Dawei Yin, Pengjie Ren, Suzan Verberne, and Zhaochun Ren. 2024. Learning to use tools via cooperative and interactive agents. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 10642\u201310657. DOI: 10.18653\/v1\/2024.findings-emnlp.624"},{"key":"e_1_3_2_193_2","first-page":"12856","volume-title":"Proceedings of 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Shi Zhengliang","year":"2023","unstructured":"Zhengliang Shi, Weiwei Sun, Shuo Zhang, Zhen Zhang, Pengjie Ren, and Zhaochun Ren. 2023. RADE: Reference-assisted dialogue evaluation for open-domain dialogue. In Proceedings of 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.), Association for Computational Linguistics, 12856\u201312875. DOI: 10.18653\/v1\/2023.acl-long.719"},{"key":"e_1_3_2_194_2","volume-title":"Proceedings of 13th International Conference on Learning Representations","author":"Shim Jeonghoon","year":"2025","unstructured":"Jeonghoon Shim, Gyuhyeon Seo, Cheongsu Lim, and Yohan Jo. 2025. ToolDial: Multi-turn dialogue generation method for tool-augmented language models. In Proceedings of 13th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=J1J5eGJsKZ"},{"key":"e_1_3_2_195_2","first-page":"486","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track","author":"Singh Harmanpreet","year":"2024","unstructured":"Harmanpreet Singh, Nikhil Verma, Yixiao Wang, Manasa Bharadwaj, Homa Fashandi, Kevin Ferreira, and Chul Lee. 2024. Personal large language model agents: A case study on tailored travel planning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track. Franck Dernoncourt, Daniel Preo\u0163iuc-Pietro, and Anastasia Shimorina (Eds.), Association for Computational Linguistics, 486\u2013514. DOI: 10.18653\/v1\/2024.emnlp-industry.37"},{"key":"e_1_3_2_196_2","doi-asserted-by":"crossref","unstructured":"Ishika Singh Valts Blukis Arsalan Mousavian Ankit Goyal Danfei Xu Jonathan Tremblay Dieter Fox Jesse Thomason and Animesh Garg. 2022. ProgPrompt: Generating situated robot task plans using large language models. arXiv:2209.11302. Retrieved from https:\/\/arxiv.org\/abs\/2209.11302","DOI":"10.1109\/ICRA48891.2023.10161317"},{"key":"e_1_3_2_197_2","unstructured":"Ishika Singh David Traum and Jesse Thomason. 2024. TwoStep: Multi-agent task planning using classical planners and large language models. arXiv:2403.17246. Retrieved from https:\/\/arxiv.org\/abs\/2403.17246"},{"key":"e_1_3_2_198_2","unstructured":"Ved Sirdeshmukh Kaustubh Deshpande Johannes Mols Lifeng Jin Ed-Yeremai Cardona Dean Lee Jeremy Kritz Willow Primack Summer Yue and Chen Xing. 2025. MultiChallenge: A realistic multi-turn conversation evaluation benchmark challenging to frontier LLMs. arXiv:2501.17399. Retrieved from https:\/\/arxiv.org\/abs\/2501.17399"},{"key":"e_1_3_2_199_2","doi-asserted-by":"crossref","first-page":"77","DOI":"10.18653\/v1\/2022.nlp4convai-1.8","volume-title":"Proceedings of 4th Workshop on NLP for Conversational AI","author":"Smith Eric","year":"2022","unstructured":"Eric Smith, Orion Hsu, Rebecca Qian, Stephen Roller, Y-Lan Boureau, and Jason Weston. 2022. Human evaluation of conversations is an open problem: Comparing the sensitivity of various methods for evaluating dialogue agents. In Proceedings of 4th Workshop on NLP for Conversational AI. Bing Liu, Alexandros Papangelis, Stefan Ultes, Abhinav Rastogi, Yun-Nung Chen, Georgios Spithourakis, Elnaz Nouri, and Weiyan Shi (Eds.), Association for Computational Linguistics, 77\u201397. DOI: 10.18653\/v1\/2022.nlp4convai-1.8"},{"key":"e_1_3_2_200_2","unstructured":"Guijin Son Hyunwoo Ko Hoyoung Lee Yewon Kim and Seunghyeok Hong. 2024. LLM-as-a-judge & reward model: What they can and cannot do. arXiv:2409.11239. Retrieved from https:\/\/arxiv.org\/abs\/2409.11239"},{"key":"e_1_3_2_201_2","first-page":"196","volume-title":"Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Sordoni Alessandro","year":"2015","unstructured":"Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, and Bill Dolan. 2015. A neural network approach to context-sensitive generation of conversational responses. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Rada Mihalcea, Joyce Chai, and Anoop Sarkar (Eds.), Association for Computational Linguistics, 196\u2013205. DOI: 10.3115\/v1\/N15-1020"},{"key":"e_1_3_2_202_2","unstructured":"Akshaya Kesarimangalam Srinivasan Shambhavi Singh Geordan Gutow Howie Choset and Bhaskar Vundurthy. 2023. Multi-agent collective construction using 3D Decomposition. arXiv:2309.00985. Retrieved from https:\/\/arxiv.org\/abs\/2309.00985"},{"key":"e_1_3_2_203_2","first-page":"3008","volume-title":"Advances in Neural Information Processing Systems","author":"Stiennon Nisan","year":"2020","unstructured":"Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano. 2020. Learning to summarize with human feedback. In Advances in Neural Information Processing Systems. H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33, Curran Associates, Inc., 3008\u20133021. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2020\/file\/1f89885d556929e98d3ef9b86448f951-Paper.pdf"},{"key":"e_1_3_2_204_2","first-page":"22","volume-title":"Proceedings of 57th Annual Meeting of the Association for Computational Linguistics","author":"Su Hui","unstructured":"Hui Su, Xiaoyu Shen, Rongzhi Zhang, Fei Sun, Pengwei Hu, Cheng Niu, and Jie Zhou. 2019. Improving multi-turn dialogue modelling with utterance rewriter. In Proceedings of 57th Annual Meeting of the Association for Computational Linguistics. Anna Korhonen, David Traum, and Llu\u00eds M\u00e0rquez (Eds.), Association for Computational Linguistics, 22\u201331. DOI: 10.18653\/v1\/P19-1003"},{"key":"e_1_3_2_205_2","unstructured":"Haotian Sun Yuchen Zhuang Lingkai Kong Bo Dai and Chao Zhang. 2023. AdaPlanner: Adaptive planning from feedback with language models. arXiv:2305.16653. Retrieved from https:\/\/arxiv.org\/abs\/2305.16653"},{"key":"e_1_3_2_206_2","doi-asserted-by":"crossref","first-page":"1358","DOI":"10.18653\/v1\/2024.findings-emnlp.73","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924)","author":"Sun Kai","year":"2024","unstructured":"Kai Sun, Yushi Bai, Ji Qi, Lei Hou, and Juanzi Li. 2024. MM-MATH: Advancing multimodal math evaluation with process evaluation and fine-grained classification. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 1358\u20131375. DOI: 10.18653\/v1\/2024.findings-emnlp.73"},{"key":"e_1_3_2_207_2","unstructured":"Jihoon Tack Jaehyung Kim Eric Mitchell Jinwoo Shin Yee Whye Teh and Jonathan Richard Schwarz. 2024. Online adaptation of language models with a memory of amortized contexts. arXiv:2403.04317. Retrieved from https:\/\/arxiv.org\/abs\/2403.04317"},{"key":"e_1_3_2_208_2","first-page":"1","volume-title":"Proceedings of the 2022 30th Signal Processing and Communications Applications Conference (SIU)","author":"Talha Selamet Ekrem","year":"2022","unstructured":"Ekrem Talha Selamet and Borahan T\u00fcmer. 2022. Context detection and identification in multi-Agent reinforcement learning with non-stationary environment. In Proceedings of the 2022 30th Signal Processing and Communications Applications Conference (SIU), 1\u20134. DOI: 10.1109\/SIU55565.2022.9864802"},{"key":"e_1_3_2_209_2","first-page":"6476","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Tan Zhaoxuan","year":"2024","unstructured":"Zhaoxuan Tan, Qingkai Zeng, Yijun Tian, Zheyuan Liu, Bing Yin, and Meng Jiang. 2024. Democratizing large language models via personalized parameter-efficient fine-tuning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 6476\u20136491. DOI: 10.18653\/v1\/2024.emnlp-main.372"},{"key":"e_1_3_2_210_2","volume-title":"Proceedings of 34th International Conference on Automated Planning and Scheduling","author":"Caballero Test\u00f3n Javier","year":"2024","unstructured":"Javier Caballero Test\u00f3n and MariaD R. Moreno. 2024. Multi-agent temporal task solving and plan optimization. In Proceedings of 34th International Conference on Automated Planning and Scheduling. Retrieved from https:\/\/openreview.net\/forum?id=sPSw73rhQB"},{"key":"e_1_3_2_211_2","unstructured":"Jacob-Junqi Tian Hao Yu Yury Orlovskiy Tyler Vergho Mauricio Rivera Mayank Goel Zachary Yang Jean-Francois Godbout Reihaneh Rabbany and Kellin Pelrine. 2024. Web retrieval agents for evidence-based misinformation detection. arXiv:2409.00009. Retrieved from https:\/\/arxiv.org\/abs\/2409.00009"},{"key":"e_1_3_2_212_2","doi-asserted-by":"crossref","first-page":"12833","DOI":"10.18653\/v1\/2024.findings-emnlp.750","volume-title":"Proceedings of the Findings of the Association for Computational LinguisticsEMNLP \u201924)","author":"Tong Terry","year":"2024","unstructured":"Terry Tong, Qin Liu, Jiashu Xu, and Muhao Chen. 2024. Securing multi-turn conversational language models from distributed backdoor attacks. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 12833\u201312846. DOI: 10.18653\/v1\/2024.findings-emnlp.750"},{"key":"e_1_3_2_213_2","first-page":"11","volume-title":"Computing Machinery and Intelligence","author":"Turing A. M.","year":"1995","unstructured":"A. M. Turing. 1995. Computing Machinery and Intelligence. MIT Press, 11\u201335."},{"key":"e_1_3_2_214_2","unstructured":"Szymon Tworkowski Konrad Staniszewski Miko\u0142aj Pacek Yuhuai Wu Henryk Michalewski and Piotr Mi\u0142o\u015b. 2023. Focused transformer: Contrastive training for context scaling. arXiv:2307.03170. Retrieved from https:\/\/arxiv.org\/abs\/2307.03170"},{"key":"e_1_3_2_215_2","first-page":"351","volume-title":"Proceedings of 9th International Symposium on Information and Communication Technology (SoICT \u201918)","author":"Van Thao Nguyen","unstructured":"Thao Nguyen Van, Nugroho Fredivianus, Huu Tam Tran, Kurt Geihs, and Thi Thanh Binh Huynh. 2018. Formal verification of ALICA multi-agent plans using model checking. In Proceedings of 9th International Symposium on Information and Communication Technology (SoICT \u201918). ACM, 351\u2013358. DOI: 10.1145\/3287921.3287947"},{"key":"e_1_3_2_216_2","first-page":"6000","volume-title":"Proceedings of 31st International Conference on Neural Information Processing Systems (NIPS \u201917)","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of 31st International Conference on Neural Information Processing Systems (NIPS \u201917). Curran Associates Inc., Red Hook, NY, 6000\u20136010."},{"key":"e_1_3_2_217_2","unstructured":"Oriol Vinyals and Quoc Le. 2015. A neural conversational model. arXiv:1506.05869. Retrieved from https:\/\/arxiv.org\/abs\/1506.05869"},{"key":"e_1_3_2_218_2","unstructured":"Evan Wang Federico Cassano Catherine Wu Yunfeng Bai Will Song Vaskar Nath Ziwen Han Sean Hendryx Summer Yue and Hugh Zhang. 2024. Planning in Natural language improves LLM search for code generation. arXiv:2409.03733. Retrieved from https:\/\/arxiv.org\/abs\/2409.03733"},{"key":"e_1_3_2_219_2","doi-asserted-by":"crossref","first-page":"10582","DOI":"10.18653\/v1\/2024.findings-emnlp.620","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924)","author":"Wang Haoxiang","year":"2024","unstructured":"Haoxiang Wang, Wei Xiong, Tengyang Xie, Han Zhao, and Tong Zhang. 2024. Interpretable preferences via multi-objective reward modeling and mixture-of-experts. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 10582\u201310592. DOI: 10.18653\/v1\/2024.findings-emnlp.620"},{"key":"e_1_3_2_220_2","volume-title":"Proceedings of the International Conference on Computational Linguistics","author":"Wang Jian","year":"2020","unstructured":"Jian Wang, Junhao Liu, Wei Bi, Xiaojiang Liu, Kejing He, Ruifeng Xu, and Min Yang. 2020. Dual dynamic memory network for end-to-end multi-turn task-oriented dialog systems. In Proceedings of the International Conference on Computational Linguistics. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:227230549"},{"key":"e_1_3_2_221_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11704-024-40231-1"},{"key":"e_1_3_2_222_2","doi-asserted-by":"crossref","unstructured":"Lei Wang Wanyu Xu Yihuai Lan Zhiqiang Hu Yunshi Lan Roy Ka-Wei Lee and Ee-Peng Lim. 2023. Plan-and-solve prompting: improving zero-shot chain-of-thought reasoning by large language models. arXiv:2305.04091. Retrieved from https:\/\/arxiv.org\/abs\/2305.04091","DOI":"10.18653\/v1\/2023.acl-long.147"},{"key":"e_1_3_2_223_2","unstructured":"Lei Wang Jingsen Zhang Hao Yang Zhiyuan Chen Jiakai Tang Zeyu Zhang Xu Chen Yankai Lin Ruihua Song Wayne Xin Zhao et al. 2024. User behavior simulation with large language model based agents. arXiv:2306.02552. Retrieved from https:\/\/arxiv.org\/abs\/2306.02552"},{"key":"e_1_3_2_224_2","unstructured":"Pei Wang Yanan Wu Zekun Wang Jiaheng Liu Xiaoshuai Song Zhongyuan Peng Ken Deng Chenchen Zhang Jiakai Wang Junran Peng et al. 2024. MTU-Bench: A multi-granularity tool-use benchmark for large language models. arXiv:2410.11710. Retrieved from https:\/\/arxiv.org\/abs\/2410.11710"},{"key":"e_1_3_2_225_2","unstructured":"Xingyao Wang Zihan Wang Jiateng Liu Yangyi Chen Lifan Yuan Hao Peng and Heng Ji. 2024. MINT: Evaluating LLMs in multi-turn interaction with tools and language feedback. arXiv:2309.10691. Retrieved from https:\/\/arxiv.org\/abs\/2309.10691"},{"key":"e_1_3_2_226_2","doi-asserted-by":"crossref","first-page":"4351","DOI":"10.18653\/v1\/2024.findings-naacl.271","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (NAACL \u201924)","author":"Wang Yancheng","year":"2024","unstructured":"Yancheng Wang, Ziyan Jiang, Zheng Chen, Fan Yang, Yingxue Zhou, Eunah Cho, Xing Fan, Yanbin Lu, Xiaojiang Huang, and Yingzhen Yang. 2024. RecMind: Large language model powered agent for recommendation. In Proceedings of the Findings of the Association for Computational Linguistics (NAACL \u201924). Kevin Duh, Helena Gomez, and Steven Bethard (Eds.), Association for Computational Linguistics, 4351\u20134364. DOI: 10.18653\/v1\/2024.findings-naacl.271"},{"key":"e_1_3_2_227_2","doi-asserted-by":"crossref","unstructured":"Yuxia Wang Revanth Gangi Reddy Zain Muhammad Mujahid Arnav Arora Aleksandr Rubashevskii Jiahui Geng Osama Mohammed Afzal Liangming Pan Nadav Borenstein Aditya Pillai et al. 2024. Factcheck-Bench: Fine-grained evaluation benchmark for automatic fact-checkers. arXiv:2311.09000. Retrieved from https:\/\/arxiv.org\/abs\/2311.09000","DOI":"10.18653\/v1\/2024.findings-emnlp.830"},{"key":"e_1_3_2_228_2","unstructured":"Yaoxiang Wang Zhiyong Wu Junfeng Yao and Jinsong Su. 2024. TDAG: A multi-agent framework based on dynamic task decomposition and agent generation. arXiv:2402.10178. Retrieved from https:\/\/arxiv.org\/abs\/2402.10178"},{"key":"e_1_3_2_229_2","unstructured":"Zihao Wang Eugene Agichtein and Jinho Choi. 2023. FCC: Fusing conversation history and candidate provenance for contextual response ranking in dialogue systems. arXiv:2304.00180. Retrieved from https:\/\/arxiv.org\/abs\/2304.00180"},{"key":"e_1_3_2_230_2","unstructured":"Zhi Wang Li Zhang Wenhao Wu Yuanheng Zhu Dongbin Zhao and Chunlin Chen. 2024. Meta-DT: Offline meta-RL as conditional sequence modeling with world model disentanglement. arXiv:2410.11448. Retrieved from https:\/\/arxiv.org\/abs\/2410.11448"},{"key":"e_1_3_2_231_2","doi-asserted-by":"publisher","DOI":"10.1145\/365153.365168"},{"key":"e_1_3_2_232_2","unstructured":"Shiwei Wu Chen Zhang Yan Gao Qimeng Wang Tong Xu Yao Hu and Enhong Chen. 2024. Benchmarking large language models for conversational question answering in multi-instructional documents. arXiv:2410.00526. Retrieved from https:\/\/arxiv.org\/abs\/2410.00526"},{"key":"e_1_3_2_233_2","unstructured":"Yike Wu Jiatao Zhang Nan Hu LanLing Tang Guilin Qi Jun Shao Jie Ren and Wei Song. 2024. MLDT: Multi-level decomposition for complex long-horizon robotic task planning with open-source large language model. arXiv:2403.18760. Retrieved from https:\/\/arxiv.org\/abs\/2403.18760"},{"key":"e_1_3_2_234_2","article-title":"Research and application of large model-based intelligent customer service system","author":"Xi Yuan","year":"2024","unstructured":"Yuan Xi. 2024. Research and application of large model-based intelligent customer service system. International Journal of Emerging Technologies and Advanced Applications 1, 3 (2024). Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:269366511","journal-title":"International Journal of Emerging Technologies and Advanced Applications"},{"key":"e_1_3_2_235_2","unstructured":"Hengjia Xiao and Peng Wang. 2024. LLM A*: Human in the loop large language models enabled A* search for robotics. arXiv:2312.01797. Retrieved from https:\/\/arxiv.org\/abs\/2312.01797"},{"key":"e_1_3_2_236_2","doi-asserted-by":"crossref","first-page":"10883","DOI":"10.18653\/v1\/2024.findings-emnlp.638","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924)","author":"Xiao Ruixuan","year":"2024","unstructured":"Ruixuan Xiao, Wentao Ma, Ke Wang, Yuchuan Wu, Junbo Zhao, Haobo Wang, Fei Huang, and Yongbin Li. 2024. FlowBench: Revisiting and benchmarking workflow-guided planning for LLM-based agents. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 10883\u201310900. DOI: 10.18653\/v1\/2024.findings-emnlp.638"},{"key":"e_1_3_2_237_2","doi-asserted-by":"crossref","unstructured":"Yujie Xing and Jon Atle Gulla. 2022. Evaluating and improving context attention distribution on multi-turn response generation using self-contained distractions. arXiv:2211.04943. Retrieved from https:\/\/arxiv.org\/abs\/2211.04943","DOI":"10.5121\/csit.2023.130210"},{"key":"e_1_3_2_238_2","first-page":"4884","volume-title":"Proceedings of the Findings of the Association for Computational LinguisticsEMNLP \u201922)","author":"Xu Guangxuan","year":"2022","unstructured":"Guangxuan Xu, Ruibo Liu, Fabrice Harel-Canada, Nischal Reddy Chandra, and Nanyun Peng. 2022. EnDex: Evaluation of dialogue engagingness at scale. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201922). Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.), Association for Computational Linguistics, 4884\u20134893. DOI: 10.18653\/v1\/2022.findings-emnlp.359"},{"key":"e_1_3_2_239_2","first-page":"5180","volume-title":"Proceedings of 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Xu Jing","year":"2022","unstructured":"Jing Xu, Arthur Szlam, and Jason Weston. 2022. Beyond goldfish memory: Long-term open-domain conversation. In Proceedings of 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.), Association for Computational Linguistics, 5180\u20135197. DOI: 10.18653\/v1\/2022.acl-long.356"},{"key":"e_1_3_2_240_2","unstructured":"Qiantong Xu Fenglu Hong Bo Li Changran Hu Zhengyu Chen and Jian Zhang. 2023. On the tool manipulation capability of open-source large language models. arXiv:2305.16504. Retrieved from https:\/\/arxivorg\/abs\/2305.16504"},{"key":"e_1_3_2_241_2","doi-asserted-by":"publisher","DOI":"10.3390\/knowledge2010004"},{"key":"e_1_3_2_242_2","doi-asserted-by":"crossref","unstructured":"Yunyi Yang Yunhao Li and Xiaojun Quan. 2021. UBAR: Towards fully end-to-end task-oriented dialog systems with GPT-2. arXiv:2012.03539. Retrieved from https:\/\/arxiv.org\/abs\/2012.03539","DOI":"10.1609\/aaai.v35i16.17674"},{"key":"e_1_3_2_243_2","unstructured":"Shunyu Yao Dian Yu Jeffrey Zhao Izhak Shafran Thomas L. Griffiths Yuan Cao and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. arXiv:2305.10601. Retrieved from https:\/\/arxiv.org\/abs\/2305.10601"},{"key":"e_1_3_2_244_2","unstructured":"Shunyu Yao Jeffrey Zhao Dian Yu Nan Du Izhak Shafran Karthik Narasimhan and Yuan Cao. 2022. ReAct: Synergizing reasoning and acting in language models. arXiv:2210.03629. Retrieved from https:\/\/arxiv.org\/abs\/2210.03629"},{"key":"e_1_3_2_245_2","unstructured":"Zihao Yi Jiarui Ouyang Zhe Xu Yuwen Liu Tianhao Liao Haohao Luo and Ying Shen. 2024. A survey on recent advances in LLM-Based multi-turn dialogue systems. arXiv:2402.18013. Retrieved from https:\/\/arxiv.org\/abs\/2402.18013"},{"key":"e_1_3_2_246_2","unstructured":"Erxin Yu Jing Li Ming Liao Siqi Wang Zuchen Gao Fei Mi and Lanqing Hong. 2024. CoSafe: Evaluating large language model safety in multi-turn dialogue coreference. arXiv:2406.17626. Retrieved from https:\/\/arxiv.org\/abs\/2406.17626"},{"key":"e_1_3_2_247_2","doi-asserted-by":"crossref","first-page":"15992","DOI":"10.18653\/v1\/2024.emnlp-main.894","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Zeng Li","year":"2024","unstructured":"Li Zeng, Yingyu Shan, Zeming Liu, Jiashu Yao, and Yuhang Guo. 2024. FAME: Towards factual multi-task model editing. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 15992\u201316011. DOI: 10.18653\/v1\/2024.emnlp-main.894"},{"key":"e_1_3_2_248_2","unstructured":"Chen Zhang Xinyi Dai Yaxiong Wu Qu Yang Yasheng Wang Ruiming Tang and Yong Liu. 2025. A survey on multi-turn interaction capabilities of large language models. arXiv:2501.09959. Retrieved from https:\/\/arxiv.org\/abs\/2501.09959"},{"key":"e_1_3_2_249_2","unstructured":"Chen Zhang Luis Fernando D\u2019Haro Yiming Chen Malu Zhang and Haizhou Li. 2024. A comprehensive analysis of the effectiveness of large language models as automatic dialogue evaluators. arXiv:2312.15407. Retrieved from https:\/\/arxiv.org\/abs\/2312.15407"},{"key":"e_1_3_2_250_2","doi-asserted-by":"crossref","first-page":"1990","DOI":"10.1145\/3442381.3449902","volume-title":"Proceedings of the Web Conference 2021 (WWW \u201921)","author":"Zhang Chen","year":"2021","unstructured":"Chen Zhang, Hao Wang, Feijun Jiang, and Hongzhi Yin. 2021. Adapting to context-aware knowledge in natural conversation for multi-turn response selection. In Proceedings of the Web Conference 2021 (WWW \u201921). ACM, 1990\u20132001. DOI: 10.1145\/3442381.3449902"},{"key":"e_1_3_2_251_2","first-page":"190","volume-title":"Proceedings of the 2020 19th International Symposium on Distributed Computing and Applications for Business Engineering and Science (DCABES)","author":"Zhang Guodong","year":"2020","unstructured":"Guodong Zhang, Li Mao, and Jun Sun. 2020. Last utterance-context attention model for multi-turn response generation. In Proceedings of the 2020 19th International Symposium on Distributed Computing and Applications for Business Engineering and Science (DCABES), 190\u2013193. DOI: 10.1109\/DCABES50732.2020.00057"},{"key":"e_1_3_2_252_2","unstructured":"Kai Zhang Lizhi Qing Yangyang Kang and Xiaozhong Liu. 2024. Personalized LLM response generation with parameterized memory injection. arXiv:2404.03565. Retrieved from https:\/\/arxiv.org\/abs\/2404.03565"},{"key":"e_1_3_2_253_2","doi-asserted-by":"publisher","DOI":"10.1145\/3545570"},{"key":"e_1_3_2_254_2","doi-asserted-by":"crossref","unstructured":"Saizheng Zhang Emily Dinan Jack Urbanek Arthur Szlam Douwe Kiela and Jason Weston. 2018. Personalizing dialogue agents: I have a dog do you have pets too? arXiv:1801.07243. Retrieved from https:\/\/arxiv.org\/abs\/1801.07243","DOI":"10.18653\/v1\/P18-1205"},{"key":"e_1_3_2_255_2","doi-asserted-by":"crossref","first-page":"1304","DOI":"10.18653\/v1\/2024.findings-emnlp.70","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924)","author":"Zhang Shaoqing","year":"2024","unstructured":"Shaoqing Zhang, Zhuosheng Zhang, Kehai Chen, Xinbei Ma, Muyun Yang, Tiejun Zhao, and Min Zhang. 2024. Dynamic planning for LLM-based graphical user interface automation. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 1304\u20131320. DOI: 10.18653\/v1\/2024.findings-emnlp.70"},{"key":"e_1_3_2_256_2","unstructured":"Tianyi Zhang Varsha Kishore Felix Wu Kilian Q. Weinberger and Yoav Artzi. 2020. BERTScore: Evaluating text generation with BERT. arXiv:1904.09675. Retrieved from https:\/\/arxiv.org\/abs\/1904.09675"},{"key":"e_1_3_2_257_2","doi-asserted-by":"crossref","first-page":"3395","DOI":"10.18653\/v1\/2022.findings-emnlp.247","volume-title":"Proceedings of the Findings of the Association for Computational LinguisticsEMNLP \u201922","author":"Zhang Tong","year":"2022","unstructured":"Tong Zhang, Yong Liu, Boyang Li, Zhiwei Zeng, Pengwei Wang, Yuan You, Chunyan Miao, and Lizhen Cui. 2022. History-aware hierarchical transformer for multi-session open-domain dialogue system. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201922). Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.), Association for Computational Linguistics, 3395\u20133407. DOI: 10.18653\/v1\/2022.findings-emnlp.247"},{"key":"e_1_3_2_258_2","unstructured":"Tong Zhang Yong Liu Boyang Li Zhiwei Zeng Pengwei Wang Yuan You Chunyan Miao and Lizhen Cui. 2023. History-aware hierarchical transformer for multi-session open-domain dialogue system. arXiv:2302.00907. Retrieved from https:\/\/arxiv.org\/abs\/2302.00907"},{"key":"e_1_3_2_259_2","unstructured":"XiuYu Zhang and Zening Luo. 2024. Advancing conversational psychotherapy: Integrating privacy dual-memory and domain expertise with large language models. arXiv:2412.02987. Retrieved from https:\/\/arxiv.org\/abs\/2412.02987"},{"key":"e_1_3_2_260_2","doi-asserted-by":"crossref","unstructured":"Xiaopan Zhang Hao Qin Fuquan Wang Yue Dong and Jiachen Li. 2024. LaMMA-P: Generalizable multi-agent long-horizon task allocation and planning with LM-driven PDDL planner. arXiv:2409.20560. Retrieved from https:\/\/arxiv.org\/abs\/2409.20560","DOI":"10.1109\/ICRA55743.2025.11127951"},{"key":"e_1_3_2_261_2","doi-asserted-by":"crossref","unstructured":"Yuxiang Zhang Jing Chen Junjie Wang Yaxin Liu Cheng Yang Chufan Shi Xinyu Zhu Zihao Lin Hanwen Wan Yujiu Yang et al. 2024. ToolBeHonest: A multi-level hallucination diagnostic benchmark for tool-augmented large language models. arXiv:2406.20015. Retrieved from https:\/\/arxiv.org\/abs\/2406.20015","DOI":"10.18653\/v1\/2024.emnlp-main.637"},{"key":"e_1_3_2_262_2","doi-asserted-by":"crossref","unstructured":"Yizhe Zhang Jiarui Lu and Navdeep Jaitly. 2024. Probing the multi-turn planning capabilities of LLMs via 20 question games. arXiv:2310.01468. Retrieved from https:\/\/arxiv.org\/abs\/2310.01468","DOI":"10.18653\/v1\/2024.acl-long.82"},{"key":"e_1_3_2_263_2","unstructured":"Zeyu Zhang Xiaohe Bo Chen Ma Rui Li Xu Chen Quanyu Dai Jieming Zhu Zhenhua Dong and Ji-Rong Wen. 2024. A survey on the memory mechanism of large language model based agents. arXiv:2404.13501. Retrieved from https:\/\/arxiv.org\/abs\/2404.13501"},{"key":"e_1_3_2_264_2","unstructured":"Zeyu Zhang Quanyu Dai Luyu Chen Zeren Jiang Rui Li Jieming Zhu Xu Chen Yi Xie Zhenhua Dong and Ji-Rong Wen. 2024. MemSim: A Bayesian simulator for evaluating memory of LLM-based personal assistants. arXiv:2409.20163. Retrieved from https:\/\/arxiv.org\/abs\/2409.20163"},{"key":"e_1_3_2_265_2","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/1757\/1\/012023"},{"key":"e_1_3_2_266_2","doi-asserted-by":"crossref","first-page":"13186","DOI":"10.18653\/v1\/2024.findings-emnlp.771","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924)","author":"Zhao Xin","year":"2024","unstructured":"Xin Zhao, Naoki Yoshinaga, and Daisuke Oba. 2024. What matters in memorizing and recalling facts? Multifaceted benchmarks for knowledge probing in language models. In Proceedings of the Findings of the Association for Computational Linguistics (EMNLP \u201924). Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.), Association for Computational Linguistics, 13186\u201313214. DOI: 10.18653\/v1\/2024.findings-emnlp.771"},{"key":"e_1_3_2_267_2","unstructured":"Zirui Zhao Wee Sun Lee and David Hsu. 2023. Large language models as Commonsense knowledge for large-scale task planning. arXiv:2305.14078. Retrieved from https:\/\/arxiv.org\/abs\/2305.14078"},{"key":"e_1_3_2_268_2","unstructured":"Lianmin Zheng Wei-Lin Chiang Ying Sheng Tianle Li Siyuan Zhuang Zhanghao Wu Yonghao Zhuang Zhuohan Li Zi Lin Eric P. Xing et al. 2024. LMSYS-Chat-1M: A large-scale real-world LLM conversation dataset. arXiv:2309.11998. Retrieved from https:\/\/arxiv.org\/abs\/2309.11998"},{"key":"e_1_3_2_269_2","unstructured":"Lianmin Zheng Wei Lin Chiang Ying Sheng Siyuan Zhuang Zhanghao Wu Yonghao Zhuang Zi Lin Zhuohan Li Dacheng Li Eric P. Xing et al. 2023. Judging LLM-as-a-judge with MT-bench and chatbot arena. arXiv:2306.05685. Retrieved from https:\/\/arxiv.org\/abs\/2306.05685"},{"key":"e_1_3_2_270_2","doi-asserted-by":"crossref","unstructured":"Xin Zheng Jie Lou Boxi Cao Xueru Wen Yuqiu Ji Hongyu Lin Yaojie Lu Xianpei Han Debing Zhang and Le Sun. 2024. Critic-CoT: Boosting the reasoning abilities of large language model via chain-of-thoughts critic. arXiv:2408.16326. Retrieved from https:\/\/arxiv.org\/abs\/2408.16326","DOI":"10.18653\/v1\/2025.findings-acl.89"},{"key":"e_1_3_2_271_2","unstructured":"Wanjun Zhong Lianghong Guo Qiqi Gao He Ye and Yanlin Wang. 2023. MemoryBank: Enhancing large language models with long-term memory. arXiv:2305.10250. Retrieved from https:\/\/arxiv.org\/abs\/2305.10250"},{"key":"e_1_3_2_272_2","doi-asserted-by":"crossref","first-page":"9945","DOI":"10.18653\/v1\/2023.acl-long.553","volume-title":"Proceedings of 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhou Junkai","year":"2023","unstructured":"Junkai Zhou, Liang Pang, Huawei Shen, and Xueqi Cheng. 2023. SimOAP: Improve coherence and consistency in persona-based dialogue generation via over-sampling and post-evaluation. In Proceedings of 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.), Association for Computational Linguistics, 9945\u20139959. DOI: 10.18653\/v1\/2023.acl-long.553"},{"key":"e_1_3_2_273_2","unstructured":"Yuchen Zhuang Yue Yu Kuan Wang Haotian Sun and Chao Zhang. 2023. ToolQA: A dataset for LLM question answering with external tools. arXiv:2306.13304. Retrieved from https:\/\/arxiv.org\/abs\/2306.13304"},{"key":"e_1_3_2_274_2","unstructured":"Mingchen Zhuge Changsheng Zhao Dylan Ashley Wenyi Wang Dmitrii Khizbullin Yunyang Xiong Zechun Liu Ernie Chang Raghuraman Krishnamoorthi Yuandong Tian et al. 2024. Agent-as-a-judge: Evaluate agents with agents. arXiv:2410.10934. Retrieved from https:\/\/arxiv.org\/abs\/2410.10934"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3793671","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T14:44:33Z","timestamp":1777387473000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3793671"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,28]]},"references-count":274,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,8,31]]}},"alternative-id":["10.1145\/3793671"],"URL":"https:\/\/doi.org\/10.1145\/3793671","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,28]]},"assertion":[{"value":"2025-03-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-21","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-28","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}