{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T16:46:16Z","timestamp":1784738776322,"version":"3.55.0"},"reference-count":221,"publisher":"Association for Computing Machinery (ACM)","issue":"9","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62477011, 62207013"],"award-info":[{"award-number":["62477011, 62207013"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>With the rapid development of Large Language Models (LLMs), LLM-based agents have been widely adopted in various fields, becoming essential for autonomous decision-making and interactive tasks. However, current work typically relies on prompt design or fine-tuning strategies applied to vanilla LLMs, which often leads to limited effectiveness in complex agent-related environments. Although numerous recent studies have explored various strategies to optimize LLM-based agents for complex agent tasks, a systematic review summarizing and comparing these methods from a holistic perspective remains lacking. In this survey, we provide a comprehensive review of LLM-based agent optimization approaches, categorizing them into parameter-driven and parameter-free methods. We first focus on parameter-driven optimization, covering fine-tuning-based optimization, reinforcement learning-based optimization, and hybrid strategies, analyzing key aspects such as trajectory data construction, reward function design, and optimization algorithms. Additionally, we briefly discuss parameter-free strategies that optimize agent behavior through prompt engineering and external knowledge retrieval. Finally, we summarize the evaluation for agents, review key applications of LLM-based agents, and discuss the major challenges and promising future directions. A curated collection of the surveyed works is provided at https:\/\/github.com\/YoungDubbyDu\/LLM-Agent-Optimization.<\/jats:p>","DOI":"10.1145\/3789261","type":"journal-article","created":{"date-parts":[[2026,1,24]],"date-time":"2026-01-24T18:31:18Z","timestamp":1769279478000},"page":"1-37","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["A Survey on the Optimization of Large Language Model-based Agents"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5792-631X","authenticated-orcid":false,"given":"Shangheng","family":"Du","sequence":"first","affiliation":[{"name":"Shanghai Institute of Artificial Intelligence for Education, East China Normal University","place":["Shanghai, China"]},{"name":"School of Computer Science and Technology, East China Normal University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0691-5741","authenticated-orcid":false,"given":"Jiabao","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Donghua University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-8898-6030","authenticated-orcid":false,"given":"Jinxin","family":"Shi","sequence":"additional","affiliation":[{"name":"Shanghai Institute of Artificial Intelligence for Education, East China Normal University","place":["Shanghai, China"]},{"name":"School of Computer Science and Technology, East China Normal University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-3488-6982","authenticated-orcid":false,"given":"Zhentao","family":"Xie","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, East China Normal University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-2883-019X","authenticated-orcid":false,"given":"Xin","family":"Jiang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, East China Normal University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4707-1765","authenticated-orcid":false,"given":"Yanhong","family":"Bai","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, East China Normal University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4723-5486","authenticated-orcid":false,"given":"Liang","family":"He","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, East China Normal University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,2,11]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Michael Ahn Debidatta Dwibedi Chelsea Finn Montse Gonzalez Arenas Keerthana Gopalakrishnan Karol Hausman Brian Ichter Alex Irpan Nikhil Joshi Ryan Julian et\u00a0al. 2024. Autort: Embodied foundation models for large scale orchestration of robotic agents. arXiv:2401.12963. Retrieved from https:\/\/arxiv.org\/abs\/2401.12963"},{"key":"e_1_3_2_3_2","unstructured":"Renat Aksitov Sobhan Miryoosefi Zonglin Li Daliang Li Sheila Babayan Kavya Kopparapu Zachary Fisher Ruiqi Guo Sushant Prakash Pranesh Srinivasan et\u00a0al. 2023. Rest meets react: Self-improvement for multi-step reasoning llm agent. arXiv:2312.10003. Retrieved from https:\/\/arxiv.org\/abs\/2312.10003"},{"key":"e_1_3_2_4_2","unstructured":"Siyu An Qin Li Junru Lu Di Yin and Xing Sun. 2024. FinVerse: An autonomous agent system for versatile financial analysis. arXiv:2406.06379. Retrieved from https:\/\/arxiv.org\/abs\/2406.06379"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1039\/D4DD00252K"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_2_7_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Asai Akari","year":"2024","unstructured":"Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-RAG: Learning to retrieve, generate, and critique through self-reflection. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_2_8_2","unstructured":"Zhijie Bao Wei Chen Shengze Xiao Kuang Ren Jiaao Wu Cheng Zhong Jiajie Peng Xuanjing Huang and Zhongyu Wei. 2023. Disc-medllm: Bridging general large language models and real-world medical consultation. arXiv:2308.14346. Retrieved from https:\/\/arxiv.org\/abs\/2308.14346"},{"key":"e_1_3_2_9_2","first-page":"138595","article-title":"Reflective multi-agent collaboration based on large language models","volume":"37","author":"Bo Xiaohe","year":"2025","unstructured":"Xiaohe Bo, Zeyu Zhang, Quanyu Dai, Xueyang Feng, Lei Wang, Rui Li, Xu Chen, and Ji-Rong Wen. 2025. Reflective multi-agent collaboration based on large language models. Advances in Neural Information Processing Systems 37 (2025), 138595\u2013138631.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_10_2","first-page":"287","volume-title":"Proceedings of the Conference on Robot Learning","author":"Brohan Anthony","year":"2023","unstructured":"Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, Ryan Julian, et\u00a0al. 2023. Do as i can, not as i say: Grounding language in robotic affordances. In Proceedings of the Conference on Robot Learning. PMLR, 287\u2013318."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.3390\/app11114948"},{"key":"e_1_3_2_12_2","unstructured":"Mert Cemri Melissa Z. Pan Shuyi Yang Lakshya A. Agrawal Bhavya Chopra Rishabh Tiwari Kurt Keutzer Aditya Parameswaran Dan Klein Kannan Ramchandran et\u00a0al. 2025. Why do multi-agent llm systems fail? arXiv:2503.13657. Retrieved from https:\/\/arxiv.org\/abs\/2503.13657"},{"key":"e_1_3_2_13_2","unstructured":"Jun Shern Chan Neil Chowdhury Oliver Jaffe James Aung Dane Sherburn Evan Mays Giulio Starace Kevin Liu Leon Maksin Tejal Patwardhan et\u00a0al. 2024. Mle-bench: Evaluating machine learning agents on machine learning engineering. arXiv:2410.07095. Retrieved from https:\/\/arxiv.org\/abs\/2410.07095"},{"key":"e_1_3_2_14_2","unstructured":"Baian Chen Chang Shu Ehsan Shareghi Nigel Collier Karthik Narasimhan and Shunyu Yao. 2023. Fireact: Toward language agent fine-tuning. arXiv:2310.05915. Retrieved from https:\/\/arxiv.org\/abs\/2310.05915"},{"key":"e_1_3_2_15_2","unstructured":"Junying Chen Zhenyang Cai Ke Ji Xidong Wang Wanlong Liu Rongsheng Wang Jianye Hou and Benyou Wang. 2024. Huatuogpt-o1 towards medical complex reasoning with llms. arXiv:2412.18925. Retrieved from https:\/\/arxiv.org\/abs\/2412.18925"},{"key":"e_1_3_2_16_2","first-page":"589","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Chen Minghao","year":"2024","unstructured":"Minghao Chen, Yihang Li, Yanting Yang, Shiyu Yu, Binbin Lin, and Xiaofei He. 2024. AutoManual: Constructing instruction manuals by LLM agents via interactive environmental learning. In Proceedings of the Advances in Neural Information Processing Systems. 589\u2013631."},{"key":"e_1_3_2_17_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde De Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman et\u00a0al. 2021. Evaluating large language models trained on code. arXiv:2107.03374. Retrieved from https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_3_2_18_2","volume-title":"Proceedings of the ICLR","author":"Chen Weize","year":"2024","unstructured":"Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, et\u00a0al. 2024. AgentVerse: Facilitating multi-agent collaboration and exploring emergent behaviors. In Proceedings of the ICLR."},{"key":"e_1_3_2_19_2","unstructured":"Weize Chen Jiarui Yuan Chen Qian Cheng Yang Zhiyuan Liu and Maosong Sun. 2024. Optima: Optimizing effectiveness and efficiency for llm-based multi-agent system. arXiv:2410.08115. Retrieved from https:\/\/arxiv.org\/abs\/2410.08115"},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","first-page":"9510","DOI":"10.18653\/v1\/2024.acl-long.515","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Chen Zehui","year":"2024","unstructured":"Zehui Chen, Weihua Du, Wenwei Zhang, Kuikun Liu, Jiangning Liu, Miao Zheng, Jingming Zhuo, Songyang Zhang, Dahua Lin, Kai Chen, et\u00a0al. 2024. T-eval: Evaluating the tool utilization capability of large language models step by step. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 9510\u20139529."},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","unstructured":"Zhixun Chen Ming Li Yuxuan Huang Yali Du Meng Fang and Tianyi Zhou. 2025. ATLaS: Agent tuning via learning critical steps. arXiv:2503.02197. Retrieved from https:\/\/arxiv.org\/abs\/2503.02197","DOI":"10.18653\/v1\/2025.findings-acl.1299"},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","unstructured":"Zehui Chen Kuikun Liu Qiuchen Wang Wenwei Zhang Jiangning Liu Dahua Lin Kai Chen and Feng Zhao. 2024. Agent-FLAN: Designing data and methods of effective agent tuning for large language models. arXiv:2403.12881. Retrieved from https:\/\/arxiv.org\/abs\/2403.12881","DOI":"10.18653\/v1\/2024.findings-acl.557"},{"key":"e_1_3_2_23_2","volume-title":"Proceedings of the Machine Learning for Healthcare Conference","author":"Choi Jihye","year":"2024","unstructured":"Jihye Choi, Nils Palumbo, Prasad Chalasani, Matthew M. Engelhard, Somesh Jha, Anivarya Kumar, and David Page. 2024. MALADE: Orchestration of LLM-powered agents with retrieval augmented generation for pharmacovigilance. In Proceedings of the Machine Learning for Healthcare Conference. PMLR."},{"key":"e_1_3_2_24_2","unstructured":"Peter Clark Isaac Cowhey Oren Etzioni Tushar Khot Ashish Sabharwal Carissa Schoenick and Oyvind Tafjord. 2018. Think you have solved question answering? try arc the ai2 reasoning challenge. arXiv:1803.05457. Retrieved from https:\/\/arxiv.org\/abs\/1803.05457"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1146\/annurev-statistics-031219-041220"},{"key":"e_1_3_2_26_2","unstructured":"Karl Cobbe Vineet Kosaraju Mohammad Bavarian Mark Chen Heewoo Jun Lukasz Kaiser Matthias Plappert Jerry Tworek Jacob Hilton Reiichiro Nakano et\u00a0al. 2021. Training verifiers to solve math word problems. arXiv:2110.14168. Retrieved from https:\/\/arxiv.org\/abs\/2110.14168"},{"key":"e_1_3_2_27_2","unstructured":"crewAI Team. 2025. crewAI. Retrieved December 2 2025 from https:\/\/github.com\/crewAIInc\/crewAI"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.365"},{"key":"e_1_3_2_29_2","first-page":"28091","article-title":"Mind2web: Towards a generalist agent for the web","volume":"36","author":"Deng Xiang","year":"2023","unstructured":"Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023. Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems 36 (2023), 28091\u201328114.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_30_2","unstructured":"Zhirui Deng Zhicheng Dou Yutao Zhu Ji-Rong Wen Ruibin Xiong Mang Wang and Weipeng Chen. 2024. From novice to expert: LLM agent policy optimization via step-wise reinforcement learning. arXiv:2411.03817. Retrieved from https:\/\/arxiv.org\/abs\/2411.03817"},{"key":"e_1_3_2_31_2","article-title":"Qlora: Efficient finetuning of quantized llms","author":"Dettmers Tim","year":"2023","unstructured":"Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. Qlora: Efficient finetuning of quantized llms. Advances in Neural Information Processing Systems 36 (2023), 10088\u201310115.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41406.2024.00013"},{"key":"e_1_3_2_33_2","first-page":"15394","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Dou Zi-Yi","year":"2024","unstructured":"Zi-Yi Dou, Cheng-Fu Yang, Xueqing Wu, Kai-Wei Chang, and Nanyun Peng. 2024. Re-rest: Reflection-reinforced self-training for language agents. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 15394\u201315411."},{"key":"e_1_3_2_34_2","unstructured":"Shangheng Du Xiangchao Yan Dengyang Jiang Jiakang Yuan Yusong Hu Xin Li Liang He Bo Zhang and Lei Bai. 2025. AutoMLGen: Navigating fine-grained optimization for coding agents. arXiv:2510.08511. Retrieved from https:\/\/arxiv.org\/abs\/2510.08511"},{"key":"e_1_3_2_35_2","first-page":"111","article-title":"Introduction to reinforcement learning","author":"Ernst Damien","year":"2024","unstructured":"Damien Ernst and Arthur Louette. 2024. Introduction to reinforcement learning. Feuerriegel, S., Hartmann, J., Janiesch, C., and Zschech, P (2024), 111\u2013126.","journal-title":"Feuerriegel, S., Hartmann, J., Janiesch, C., and Zschech, P"},{"key":"e_1_3_2_36_2","first-page":"75","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Fan Yue","year":"2024","unstructured":"Yue Fan, Xiaojian Ma, Rujie Wu, Yuntao Du, Jiaqi Li, Zhi Gao, and Qing Li. 2024. Videoagent: A memory-augmented multimodal agent for video understanding. In Proceedings of the European Conference on Computer Vision. Springer, 75\u201392."},{"key":"e_1_3_2_37_2","first-page":"10183","volume-title":"Proceedings of the 31st International Conference on Computational Linguistics","author":"Fan Zhihao","year":"2025","unstructured":"Zhihao Fan, Lai Wei, Jialong Tang, Wei Chen, Wang Siyuan, Zhongyu Wei, and Fei Huang. 2025. AI hospital: Benchmarking large language models in a multi-agent medical interaction simulator. In Proceedings of the 31st International Conference on Computational Linguistics. 10183\u201310213."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3677052.3698688"},{"key":"e_1_3_2_39_2","volume-title":"Proceedings of the 38th Annual Conference on Neural Information Processing Systems","author":"Feng Peiyuan","year":"2024","unstructured":"Peiyuan Feng, Yichen He, Guanhua Huang, Yuan Lin, Hanchong Zhang, Yuchen Zhang, and Hang Li. 2024. AGILE: A novel reinforcement learning framework of LLM agents. In Proceedings of the 38th Annual Conference on Neural Information Processing Systems."},{"key":"e_1_3_2_40_2","unstructured":"Xidong Feng Ziyu Wan Haotian Fu Bo Liu Mengyue Yang Girish A. Koushik Zhiyuan Hu Ying Wen and Jun Wang. 2024. Natural language reinforcement learning. arXiv:2411.14251. Retrieved from https:\/\/arxiv.org\/abs\/2411.14251"},{"key":"e_1_3_2_41_2","first-page":"643","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Fu Dayuan","year":"2024","unstructured":"Dayuan Fu, Biqing Qi, Yihuai Gao, Che Jiang, Guanting Dong, and Bowen Zhou. 2024. MSI-Agent: Incorporating multi-scale insight into embodied agents for superior planning and decision-making. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 643\u2013659."},{"key":"e_1_3_2_42_2","volume-title":"Proceedings of the 38th Annual Conference on Neural Information Processing Systems","author":"Fu Yao","year":"2024","unstructured":"Yao Fu, Dong-Ki Kim, Jaekyeom Kim, Sungryull Sohn, Lajanugen Logeswaran, Kyunghoon Bae, and Honglak Lee. 2024. Autoguide: Automated generation and selection of context-aware guidelines for large language model agents. In Proceedings of the 38th Annual Conference on Neural Information Processing Systems."},{"key":"e_1_3_2_43_2","unstructured":"Dawei Gao Zitao Li Xuchen Pan Weirui Kuang Zhijian Ma Bingchen Qian Fei Wei Wenhao Zhang Yuexiang Xie Daoyuan Chen et\u00a0al. 2024. Agentscope: A flexible yet robust multi-agent platform. arXiv:2402.14034. Retrieved from https:\/\/arxiv.org\/abs\/2402.14034"},{"key":"e_1_3_2_44_2","unstructured":"Shen Gao Yuntao Wen Minghang Zhu Jianing Wei Yuhan Cheng Qunzi Zhang and Shuo Shang. 2024. Simulating financial market via large language model based agents. arXiv:2406.19966. Retrieved from https:\/\/arxiv.org\/abs\/2406.19966"},{"issue":"7","key":"e_1_3_2_45_2","doi-asserted-by":"crossref","first-page":"1389","DOI":"10.1039\/D4DD00013G","article-title":"ProtAgents: Protein discovery via large language model multi-agent collaborations combining physics and machine learning","volume":"3","author":"Ghafarollahi Alireza","year":"2024","unstructured":"Alireza Ghafarollahi and Markus J. Buehler. 2024. ProtAgents: Protein discovery via large language model multi-agent collaborations combining physics and machine learning. Digital Discovery 3, 7 (2024), 1389\u20131409.","journal-title":"Digital Discovery"},{"key":"e_1_3_2_46_2","first-page":"3397","article-title":"An operator view of policy gradient methods","volume":"33","author":"Ghosh Dibya","year":"2020","unstructured":"Dibya Ghosh, Marlos C. Machado, and Nicolas Le Roux. 2020. An operator view of policy gradient methods. Advances in Neural Information Processing Systems 33 (2020), 3397\u20133406.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_47_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Gou Zhibin","year":"2024","unstructured":"Zhibin Gou, Zhihong Shao, Yeyun Gong, yelong shen, Yujiu Yang, Minlie Huang, Nan Duan, and Weizhu Chen. 2024. ToRA: A tool-integrated reasoning agent for mathematical problem solving. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_2_48_2","doi-asserted-by":"crossref","first-page":"7646","DOI":"10.18653\/v1\/2024.emnlp-main.436","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Gu Yu","year":"2024","unstructured":"Yu Gu, Yiheng Shu, Hao Yu, Xiao Liu, Yuxiao Dong, Jie Tang, Jayanth Srinivasa, Hugo Latapie, and Yu Su. 2024. Middleware for LLMs: Tools are instrumental for language agents in complex environments. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 7646\u20137663."},{"key":"e_1_3_2_49_2","first-page":"126118","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Guan Jian","year":"2024","unstructured":"Jian Guan, Wei Wu, zujie wen, Peng Xu, Hongning Wang, and Minlie Huang. 2024. AMOR: A recipe for building adaptable modular knowledge agents through process feedback. In Proceedings of the Advances in Neural Information Processing Systems. 126118\u2013126148."},{"key":"e_1_3_2_50_2","unstructured":"Daya Guo Dejian Yang Haowei Zhang Junxiao Song Ruoyu Zhang Runxin Xu Qihao Zhu Shirong Ma Peiyi Wang Xiao Bi et\u00a0al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv:2501.12948. Retrieved from https:\/\/arxiv.org\/abs\/2501.12948"},{"key":"e_1_3_2_51_2","doi-asserted-by":"crossref","first-page":"8369","DOI":"10.18653\/v1\/2024.emnlp-main.477","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Gupta Priyanshu","year":"2024","unstructured":"Priyanshu Gupta, Shashank Kirtania, Ananya Singha, Sumit Gulwani, Arjun Radhakrishna, Gustavo Soares, and Sherry Shi. 2024. MetaReflection: Learning instructions for language agents using past reflections. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 8369\u20138385."},{"key":"e_1_3_2_52_2","article-title":"Measuring massive multitask language understanding","author":"Hendrycks Dan","year":"2021","unstructured":"Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (2021).","journal-title":"Proceedings of the International Conference on Learning Representations"},{"key":"e_1_3_2_53_2","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks","author":"Hendrycks Dan","year":"2021","unstructured":"Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the MATH dataset. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks."},{"key":"e_1_3_2_54_2","volume-title":"Proceedings of the ICLR","author":"Hong Sirui","year":"2024","unstructured":"Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, et\u00a0al. 2024. MetaGPT: Meta programming for a multi-agent collaborative framework. In Proceedings of the ICLR."},{"key":"e_1_3_2_55_2","first-page":"26406","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Hong Yining","year":"2024","unstructured":"Yining Hong, Zishuo Zheng, Peihao Chen, Yian Wang, Junyan Li, and Chuang Gan. 2024. Multiply: A multisensory object-centric embodied large language model in 3d world. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 26406\u201326416."},{"key":"e_1_3_2_56_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Hu Edward J","year":"2022","unstructured":"Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_57_2","unstructured":"Yusong Hu Runmin Ma Yue Fan Jinxin Shi Zongsheng Cao Yuhao Zhou Jiakang Yuan Xiangchao Yan Wenlong Zhang Lei Bai et\u00a0al. 2025. FlowSearch: Advancing deep research with dynamic structured knowledge flow. arXiv:2510.08521. Retrieved from https:\/\/arxiv.org\/abs\/2510.08521"},{"key":"e_1_3_2_58_2","unstructured":"Kaixuan Huang Yuanhao Qu Henry Cousins William A Johnson Di Yin Mihir Shah Denny Zhou Russ Altman Mengdi Wang and Le Cong. 2024. Crispr-gpt: An llm agent for automated design of gene-editing experiments. arXiv:2404.18021. Retrieved from https:\/\/arxiv.org\/abs\/2404.18021"},{"key":"e_1_3_2_59_2","doi-asserted-by":"crossref","first-page":"5014","DOI":"10.18653\/v1\/2024.acl-long.274","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Huang Xiang","year":"2024","unstructured":"Xiang Huang, Sitao Cheng, Shanshan Huang, Jiayu Shen, Yong Xu, Chaoyun Zhang, and Yuzhong Qu. 2024. QueryAgent: A reliable and efficient reasoning framework with environmental feedback based self-correction. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 5014\u20135035."},{"key":"e_1_3_2_60_2","unstructured":"Xu Huang Jianxun Lian Yuxuan Lei Jing Yao Defu Lian and Xing Xie. 2023. Recommender ai agent: Integrating large language models for interactive recommendations. arXiv:2308.16505. Retrieved from https:\/\/arxiv.org\/abs\/2308.16505"},{"key":"e_1_3_2_61_2","unstructured":"Xu Huang Weiwen Liu Xiaolong Chen Xingmei Wang Hao Wang Defu Lian Yasheng Wang Ruiming Tang and Enhong Chen. 2024. Understanding the planning of LLM agents: A survey. arXiv:2402.02716. Retrieved from https:\/\/arxiv.org\/abs\/2402.02716"},{"key":"e_1_3_2_62_2","unstructured":"Haolin Jin Linghan Huang Haipeng Cai Jun Yan Bo Li and Huaming Chen. 2024. From llms to llm-based agents for software engineering: A survey of current challenges and future. arXiv:2408.02479. Retrieved from https:\/\/arxiv.org\/abs\/2408.02479"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1259"},{"key":"e_1_3_2_64_2","doi-asserted-by":"crossref","unstructured":"Qiao Jin Zhizheng Wang Yifan Yang Qingqing Zhu Donald Wright Thomas Huang W. John Wilbur Zhe He Andrew Taylor Qingyu Chen et\u00a0al. 2024. Agentmd: Empowering language agents for risk prediction with large-scale clinical tool learning. arXiv:2402.13225. Retrieved from https:\/\/arxiv.org\/abs\/2402.13225","DOI":"10.1038\/s41467-025-64430-x"},{"key":"e_1_3_2_65_2","unstructured":"Dongkyu Kim Byoungwook Kim Donggeon Han and Matou\u0161 Eibich. 2024. AutoRAG: Automated framework for optimization of retrieval augmented generation pipeline. arXiv:2410.20878. Retrieved from https:\/\/arxiv.org\/abs\/2410.20878"},{"key":"e_1_3_2_66_2","first-page":"13511","volume-title":"Findings of the Association for Computational Linguistics ACL 2024","author":"Kim Minsoo","year":"2024","unstructured":"Minsoo Kim, Victor Bursztyn, Eunyee Koh, Shunan Guo, and Seung-won Hwang. 2024. Rada: Retrieval-augmented web agent planning with llms. In Findings of the Association for Computational Linguistics ACL 2024. 13511\u201313525."},{"key":"e_1_3_2_67_2","first-page":"79410","article-title":"Mdagents: An adaptive collaboration of llms for medical decision-making","volume":"37","author":"Kim Yubin","year":"2025","unstructured":"Yubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan, Xuhai Xu, Daniel McDuff, Hyeonhoon Lee, Marzyeh Ghassemi, Cynthia Breazeal, Hae Park, et\u00a0al. 2025. Mdagents: An adaptive collaboration of llms for medical decision-making. Advances in Neural Information Processing Systems 37 (2025), 79410\u201379452.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1145\/3589334.3645611"},{"key":"e_1_3_2_69_2","unstructured":"Jakub L\u00e1la Odhran O\u2019Donoghue Aleksandar Shtedritski Sam Cox Samuel G. Rodriques and Andrew D. White. 2023. Paperqa: Retrieval-augmented generative agent for scientific research. arXiv:2312.07559. Retrieved from https:\/\/arxiv.org\/abs\/2312.07559"},{"key":"e_1_3_2_70_2","unstructured":"Marc Lanctot Edward Lockhart Jean-Baptiste Lespiau Vinicius Zambaldi Satyaki Upadhyay Julien P\u00e9rolat Sriram Srinivasan Finbarr Timbers Karl Tuyls Shayegan Omidshafiei et\u00a0al. 2019. OpenSpiel: A framework for reinforcement learning in games. arXiv:1908.09453. Retrieved from https:\/\/arxiv.org\/abs\/1908.09453"},{"key":"e_1_3_2_71_2","first-page":"15737","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Lee Dong Won","year":"2024","unstructured":"Dong Won Lee, Hae Park, Yoon Kim, Cynthia Breazeal, and Louis-Philippe Morency. 2024. Global reward to local rewards: Multimodal-guided decomposition for improving dialogue agents. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 15737\u201315762."},{"key":"e_1_3_2_72_2","first-page":"8745","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2024","author":"Li Binxu","year":"2024","unstructured":"Binxu Li, Tiankai Yan, Yuanting Pan, Jie Luo, Ruiyang Ji, Jiayuan Ding, Zhe Xu, Shilong Liu, Haoyu Dong, Zihao Lin, et\u00a0al. 2024. MMedAgent: Learning to use medical tools with multi-modal agent. In Findings of the Association for Computational Linguistics: EMNLP 2024. 8745\u20138760."},{"key":"e_1_3_2_73_2","first-page":"2757","volume-title":"Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing","author":"Li Dawei","year":"2025","unstructured":"Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, et\u00a0al. 2025. From generation to judgment: Opportunities and challenges of llm-as-a-judge. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2757\u20132791."},{"key":"e_1_3_2_74_2","unstructured":"Dawei Li Zhen Tan Peijia Qian Yifan Li Kumar Satvik Chaudhary Lijie Hu and Jiayi Shen. 2024. Smoa: Improving multi-agent large language models with sparse mixture-of-agents. arXiv:2411.03284. Retrieved from https:\/\/arxiv.org\/abs\/2411.03284"},{"key":"e_1_3_2_75_2","first-page":"51991","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Li Guohao","year":"2023","unstructured":"Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. CAMEL: Communicative agents for \u201cmind\u201d exploration of large language model society. In Proceedings of the 37th International Conference on Neural Information Processing Systems. 51991\u201352008."},{"key":"e_1_3_2_76_2","article-title":"Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls","author":"Li Jinyang","year":"2023","unstructured":"Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Ruiying Geng, Nan Huo, et\u00a0al. 2023. Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls. Advances in Neural Information Processing Systems 36 (2023), 42330\u201342357.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_77_2","unstructured":"Junkai Li Yunghwei Lai Weitao Li Jingyi Ren Meng Zhang Xinhui Kang Siyu Wang Peng Li Ya-Qin Zhang Weizhi Ma et\u00a0al. 2024. Agent hospital: A simulacrum of hospital with evolvable medical agents. arXiv:2405.02957. Retrieved from https:\/\/arxiv.org\/abs\/2405.02957"},{"key":"e_1_3_2_78_2","first-page":"7595","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Li Ming","year":"2024","unstructured":"Ming Li, Yong Zhang, Zhitao Li, Jiuhai Chen, Lichang Chen, Ning Cheng, Jianzong Wang, Tianyi Zhou, and Jing Xiao. 2024. From quantity to quality: Boosting LLM performance with self-guided data selection for instruction tuning. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 7595\u20137628."},{"key":"e_1_3_2_79_2","first-page":"4703","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Li Renhao","year":"2024","unstructured":"Renhao Li, Minghuan Tan, Derek Wong, and Min Yang. 2024. CoEvol: Constructing better responses for instruction finetuning through multi-agent cooperation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 4703\u20134721."},{"key":"e_1_3_2_80_2","unstructured":"Xiaoxi Li Jiajie Jin Guanting Dong Hongjin Qian Yongkang Wu Ji-Rong Wen Yutao Zhu and Zhicheng Dou. 2025. Webthinker: Empowering large reasoning models with deep research capability. arXiv:2504.21776. Retrieved from https:\/\/arxiv.org\/abs\/2504.21776"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.1007\/s44336-024-00009-2"},{"key":"e_1_3_2_82_2","unstructured":"Yang Li Yangyang Yu Haohang Li Zhi Chen and Khaldoun Khashanah. 2023. Tradinggpt: Multi-agent system with layered memory and distinct characters for enhanced financial trading performance. arXiv:2309.03736. Retrieved from https:\/\/arxiv.org\/abs\/2309.03736"},{"key":"e_1_3_2_83_2","first-page":"49881","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"37","author":"Li Zaijing","year":"2024","unstructured":"Zaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen, Dongmei Jiang, and Liqiang Nie. 2024. Optimus-1: Hybrid multimodal memory empowered agents excel in long-horizon tasks. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 37. 49881\u201349913."},{"key":"e_1_3_2_84_2","doi-asserted-by":"crossref","first-page":"17889","DOI":"10.18653\/v1\/2024.emnlp-main.992","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Liang Tian","year":"2024","unstructured":"Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu. 2024. Encouraging divergent thinking in large language models through multi-agent debate. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 17889\u201317904."},{"key":"e_1_3_2_85_2","unstructured":"Xuechen Liang Meiling Tao Yinghui Xia Tianyu Shi Jun Wang and JingSong Yang. 2024. Cmat: A multi-agent collaboration tuning framework for enhancing small language models. arXiv:2404.01663. Retrieved from https:\/\/arxiv.org\/abs\/2404.01663"},{"key":"e_1_3_2_86_2","doi-asserted-by":"crossref","unstructured":"Xuechen Liang Meiling Tao Yinghui Xia Tianyu Shi Jun Wang and JingSong Yang. 2024. Self-evolving agents with reflective and memory-augmented abilities. arXiv:2409.00872. Retrieved from https:\/\/arxiv.org\/abs\/2409.00872","DOI":"10.2139\/ssrn.5182425"},{"key":"e_1_3_2_87_2","article-title":"Swiftsage: A generative agent with fast and slow thinking for complex interactive tasks","author":"Lin Bill Yuchen","year":"2023","unstructured":"Bill Yuchen Lin, Yicheng Fu, Karina Yang, Faeze Brahman, Shiyu Huang, Chandra Bhagavatula, Prithviraj Ammanabrolu, Yejin Choi, and Xiang Ren. 2023. Swiftsage: A generative agent with fast and slow thinking for complex interactive tasks. Advances in Neural Information Processing Systems 36 (2023), 23813\u201323825.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_88_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Lin Bill Yuchen","year":"2024","unstructured":"Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. 2024. The unlocking spell on base LLMs: Rethinking alignment via in-context learning. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_2_89_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.229"},{"key":"e_1_3_2_90_2","unstructured":"Evan Zheran Liu Kelvin Guu Panupong Pasupat Tianlin Shi and Percy Liang. 2018. Reinforcement learning on web interfaces using workflow-guided exploration. arXiv:1802.08802. Retrieved from https:\/\/arxiv.org\/abs\/1802.08802"},{"key":"e_1_3_2_91_2","article-title":"Visual instruction tuning","author":"Liu Haotian","year":"2023","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. Advances in neural information processing systems 36 (2023), 34892\u201334916.","journal-title":"Advances in neural information processing systems"},{"key":"e_1_3_2_92_2","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.1234"},{"key":"e_1_3_2_93_2","volume-title":"Proceedings of the 13th International Conference on Learning Representations","author":"Liu Jie","year":"2025","unstructured":"Jie Liu, Pan Zhou, Yingjun Du, Ah-Hwee Tan, Cees G. M. Snoek, Jan-Jakob Sonke, and Efstratios Gavves. 2025. CaPo: Cooperative plan optimization for efficient embodied multi-agent cooperation. In Proceedings of the 13th International Conference on Learning Representations."},{"key":"e_1_3_2_94_2","unstructured":"Sizhe Liu Yizhou Lu Siyu Chen Xiyang Hu Jieyu Zhao Tianfan Fu and Yue Zhao. 2024. Drugagent: Automating ai-aided drug discovery programming through llm multi-agent collaboration. arXiv:2411.15692. Retrieved from https:\/\/arxiv.org\/abs\/2411.15692"},{"key":"e_1_3_2_95_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-short.8"},{"key":"e_1_3_2_96_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Liu Xiao","year":"2024","unstructured":"Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et\u00a0al. 2024. AgentBench: Evaluating LLMs as agents. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_2_97_2","first-page":"598","volume-title":"Proceedings of the 2020 Chinese Control And Decision Conference","author":"Liu Yun-ting","year":"2020","unstructured":"Yun-ting Liu, Jia-ming Yang, Liang Chen, Ting Guo, and Yu Jiang. 2020. Overview of reinforcement learning based on value and policy. In Proceedings of the 2020 Chinese Control And Decision Conference. IEEE, 598\u2013603."},{"key":"e_1_3_2_98_2","volume-title":"Proceedings of the 1st Conference on Language Modeling","author":"Liu Zijun","year":"2024","unstructured":"Zijun Liu, Yanzhe Zhang, Peng Li, Yang Liu, and Diyi Yang. 2024. A dynamic LLM-powered agent network for task-oriented agent collaboration. In Proceedings of the 1st Conference on Language Modeling."},{"key":"e_1_3_2_99_2","first-page":"2507","article-title":"Learn to explain: Multimodal reasoning via thought chains for science question answering","volume":"35","author":"Lu Pan","year":"2022","unstructured":"Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems 35 (2022), 2507\u20132521.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_100_2","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-024-00832-8"},{"key":"e_1_3_2_101_2","volume-title":"Proceedings of the 38th Conference on Neural Information Processing Systems Datasets and Benchmarks Track","author":"Ma Chang","year":"2024","unstructured":"Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, Yaohui Jin, Zhenzhong Lan, Lingpeng Kong, and Junxian He. 2024. AgentBoard: An analytical evaluation board of multi-turn LLM agents. In Proceedings of the 38th Conference on Neural Information Processing Systems Datasets and Benchmarks Track."},{"key":"e_1_3_2_102_2","unstructured":"Hao Ma Tianyi Hu Zhiqiang Pu Boyin Liu Xiaolin Ai Yanyan Liang and Min Chen. 2024. Coevolving with the other you: Fine-tuning LLM with sequential cooperative multi-agent reinforcement learning. arXiv:2410.06101. Retrieved from https:\/\/arxiv.org\/abs\/2410.06101"},{"key":"e_1_3_2_103_2","first-page":"15701","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Ma Yubo","year":"2024","unstructured":"Yubo Ma, Zhibin Gou, Junheng Hao, Ruochen Xu, Shuohang Wang, Liangming Pan, Yujiu Yang, Yixin Cao, and Aixin Sun. 2024. SciAgent: Tool-augmented language models for scientific reasoning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 15701\u201315736."},{"key":"e_1_3_2_104_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA57147.2024.10610855"},{"key":"e_1_3_2_105_2","article-title":"Egoschema: A diagnostic benchmark for very long-form video language understanding","author":"Mangalam Karttikeya","year":"2023","unstructured":"Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. 2023. Egoschema: A diagnostic benchmark for very long-form video language understanding. Advances in Neural Information Processing Systems 36 (2023), 46212\u201346244.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"7","key":"e_1_3_2_106_2","doi-asserted-by":"crossref","first-page":"197605","DOI":"10.1007\/s11704-024-40663-9","article-title":"A survey on lora of large language models","volume":"19","author":"Mao Yuren","year":"2025","unstructured":"Yuren Mao, Yuhang Ge, Yijiang Fan, Wenyi Xu, Yu Mi, Zhonghao Hu, and Yunjun Gao. 2025. A survey on lora of large language models. Frontiers of Computer Science 19, 7 (2025), 197605.","journal-title":"Frontiers of Computer Science"},{"key":"e_1_3_2_107_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Mialon Gr\u00e9goire","year":"2024","unstructured":"Gr\u00e9goire Mialon, Cl\u00e9mentine Fourrier, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2024. Gaia: A benchmark for general ai assistants. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_2_108_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.92"},{"key":"e_1_3_2_109_2","doi-asserted-by":"crossref","first-page":"6129","DOI":"10.1145\/3711896.3736570","volume-title":"Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2","author":"Mohammadi Mahmoud","year":"2025","unstructured":"Mahmoud Mohammadi, Yipeng Li, Jane Lo, and Wendy Yip. 2025. Evaluation and benchmarking of llm agents: A survey. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 6129\u20136139."},{"key":"e_1_3_2_110_2","article-title":"Bridging the gap between value and policy based reinforcement learning","author":"Nachum Ofir","year":"2017","unstructured":"Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans. 2017. Bridging the gap between value and policy based reinforcement learning. Advances in Neural Information Processing Systems 30 (2017), 2772\u20132782.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_111_2","unstructured":"King Han Naman Jain Alex Gu Wen-Ding Li Fanjia Yan Tianjun Zhang Sida Wang Armando Solar-Lezama Koushik Sen and Ion Stoica. 2024. Livecodebench: Holistic and contamination free evaluation of large language models for code. arXiv:2403.07974. Retrieved from https:\/\/arxiv.org\/abs\/2403.07974"},{"key":"e_1_3_2_112_2","unstructured":"Alexander Novikov Ng\u00e2n V\u0169 Marvin Eisenberger Emilien Dupont Po-Sen Huang Adam Zsolt Wagner Sergey Shirobokov Borislav Kozlovskii Francisco J. R. Ruiz Abbas Mehrabian et\u00a0al. 2025. AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv:2506.13131. Retrieved from https:\/\/arxiv.org\/abs\/2506.13131"},{"key":"e_1_3_2_113_2","unstructured":"OpenAI. 2024. GPT-4 Technical Report. arxiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_2_114_2","first-page":"27730","article-title":"Training language models to follow instructions with human feedback","volume":"35","author":"Ouyang Long","year":"2022","unstructured":"Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et\u00a0al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35 (2022), 27730\u201327744.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_115_2","first-page":"126620","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Pang Jing-Cheng","year":"2024","unstructured":"Jing-Cheng Pang, Si-Hang Yang, Kaiyuan Li, Jiaji Zhang, Xiong-Hui Chen, Nan Tang, and Yang Yu. 2024. KALM: Knowledgeable agents by offline reinforcement learning from large language model rollouts. In Proceedings of the Advances in Neural Information Processing Systems. 126620\u2013126652."},{"key":"e_1_3_2_116_2","unstructured":"Venkatesh Balavadhani Parthasarathy Ahtsham Zafar Aafaq Khan and Arsalan Shahid. 2024. The ultimate guide to fine-tuning llms from basics to breakthroughs: An exhaustive review of technologies research best practices applied research challenges and opportunities. arXiv:2408.13296. Retrieved from https:\/\/arxiv.org\/abs\/2408.13296"},{"key":"e_1_3_2_117_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.168"},{"key":"e_1_3_2_118_2","first-page":"2219","volume-title":"Proceedings of the 2006 IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"Peters Jan","year":"2006","unstructured":"Jan Peters and Stefan Schaal. 2006. Policy gradient methods for robotics. In Proceedings of the 2006 IEEE\/RSJ International Conference on Intelligent Robots and Systems. IEEE, 2219\u20132225."},{"key":"e_1_3_2_119_2","unstructured":"Pranav Putta Edmund Mills Naman Garg Sumeet Motwani Chelsea Finn Divyansh Garg and Rafael Rafailov. 2024. Agent q: Advanced reasoning and learning for autonomous ai agents. arXiv:2408.07199. Retrieved from https:\/\/arxiv.org\/abs\/2408.07199"},{"key":"e_1_3_2_120_2","unstructured":"Zehan Qi Xiao Liu Iat Long Iong Hanyu Lai Xueqiao Sun Wenyi Zhao Yu Yang Xinyue Yang Jiadai Sun Shuntian Yao et\u00a0al. 2024. WebRL: Training LLM web agents via self-evolving online curriculum reinforcement learning. arXiv:2411.02337. Retrieved from https:\/\/arxiv.org\/abs\/2411.02337"},{"key":"e_1_3_2_121_2","doi-asserted-by":"crossref","first-page":"5628","DOI":"10.18653\/v1\/2024.acl-long.305","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Qian Chen","year":"2024","unstructured":"Chen Qian, Yufan Dang, Jiahao Li, Wei Liu, Zihao Xie, YiFei Wang, Weize Chen, Cheng Yang, Xin Cong, Xiaoyin Che, Zhiyuan Liu, and Maosong Sun. 2024. Experiential co-learning of software-developing agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 5628\u20135640."},{"key":"e_1_3_2_122_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.810"},{"key":"e_1_3_2_123_2","unstructured":"Chen Qian Zihao Xie Yifei Wang Wei Liu Yufan Dang Zhuoyun Du Weize Chen Cheng Yang Zhiyuan Liu and Maosong Sun. 2024. Scaling large-language-model-based multi-agent collaboration. arXiv:2406.07155. Retrieved from https:\/\/arxiv.org\/abs\/2406.07155"},{"key":"e_1_3_2_124_2","first-page":"114843","article-title":"Agent planning with world knowledge model","volume":"37","author":"Qiao Shuofei","year":"2024","unstructured":"Shuofei Qiao, Runnan Fang, Ningyu Zhang, Yuqi Zhu, Xiang Chen, Shumin Deng, Yong Jiang, Pengjun Xie, Fei Huang, and Huajun Chen. 2024. Agent planning with world knowledge model. Advances in Neural Information Processing Systems 37 (2024), 114843\u2013114871.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_125_2","doi-asserted-by":"crossref","first-page":"3003","DOI":"10.18653\/v1\/2024.acl-long.165","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Qiao Shuofei","year":"2024","unstructured":"Shuofei Qiao, Ningyu Zhang, Runnan Fang, Yujie Luo, Wangchunshu Zhou, Yuchen Jiang, Chengfei Lv, and Huajun Chen. 2024. AutoAct: Automatic agent learning from scratch for QA via self-planning. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 3003\u20133021."},{"key":"e_1_3_2_126_2","unstructured":"Yujia Qin Shihao Liang Yining Ye Kunlun Zhu Lan Yan Yaxi Lu Yankai Lin Xin Cong Xiangru Tang Bill Qian et\u00a0al. 2023. Toolllm: Facilitating large language models to master 16000+ real-world apis. arXiv:2307.16789. Retrieved from https:\/\/arxiv.org\/abs\/2307.16789"},{"key":"e_1_3_2_127_2","first-page":"55249","article-title":"Recursive introspection: Teaching language model agents how to self-improve","volume":"37","author":"Qu Yuxiao","year":"2024","unstructured":"Yuxiao Qu, Tianjun Zhang, Naman Garg, and Aviral Kumar. 2024. Recursive introspection: Teaching language model agents how to self-improve. Advances in Neural Information Processing Systems 37 (2024), 55249\u201355285.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_128_2","first-page":"53728","article-title":"Direct preference optimization: Your language model is secretly a reward model","volume":"36","author":"Rafailov Rafael","year":"2023","unstructured":"Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems 36 (2023), 53728\u201353741.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_129_2","unstructured":"Matthew Renze and Erhan Guven. 2024. Self-reflection in LLM agents: Effects on problem-solving performance. arXiv:2405.06682. Retrieved from https:\/\/arxiv.org\/abs\/2405.06682"},{"key":"e_1_3_2_130_2","unstructured":"Yusuf Roohani Andrew Lee Qian Huang Jian Vora Zachary Steinhart Kexin Huang Alexander Marson Percy Liang and Jure Leskovec. 2024. Biodiscoveryagent: An ai agent for designing genetic perturbation experiments. arXiv:2405.17631. Retrieved from https:\/\/arxiv.org\/abs\/2405.17631"},{"key":"e_1_3_2_131_2","unstructured":"Jingqing Ruan Yihong Chen Bin Zhang Zhiwei Xu Tianpeng Bao Guoqing Du Shiwei Shi Hangyu Mao Ziyue Li Xingyu Zeng et\u00a0al. 2023. TPTU: Large language model-based AI agents for task planning and tool usage. arXiv:2308.03427. Retrieved from https:\/\/arxiv.org\/abs\/2308.03427"},{"key":"e_1_3_2_132_2","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv:1707.06347. Retrieved from https:\/\/arxiv.org\/abs\/1707.06347"},{"key":"e_1_3_2_133_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20074-8_9"},{"key":"e_1_3_2_134_2","unstructured":"Jinxin Shi Zongsheng Cao Runmin Ma Yusong Hu Jie Zhou Xin Li Lei Bai Liang He and Bo Zhang. 2025. DualResearch: Entropy-gated dual-graph retrieval for answer reconstruction. arXiv:2510.08959. Retrieved from https:\/\/arxiv.org\/abs\/2510.08959"},{"key":"e_1_3_2_135_2","unstructured":"Wentao Shi Zichun Yu Fuli Feng Xiangnan He and Chenyan Xiong. 2025. Efficient multi-agent system training with data influence-oriented tree search. arXiv:2502.00955. Retrieved from https:\/\/arxiv.org\/abs\/2502.00955"},{"key":"e_1_3_2_136_2","first-page":"2312","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Shi Wentao","year":"2024","unstructured":"Wentao Shi, Mengqi Yuan, Junkang Wu, Qifan Wang, and Fuli Feng. 2024. Direct multi-turn preference optimization for language agents. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2312\u20132324."},{"key":"e_1_3_2_137_2","article-title":"Reflexion: Language agents with verbal reinforcement learning","author":"Shinn Noah","year":"2023","unstructured":"Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems 36 (2023), 8634\u20138652.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_138_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01075"},{"key":"e_1_3_2_139_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Shridhar Mohit","year":"2021","unstructured":"Mohit Shridhar, Xingdi Yuan, Marc-Alexandre C\u00f4t\u00e9, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2021. ALFWorld: Aligning text and embodied environments for interactive learning. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_140_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-023-06291-2"},{"key":"e_1_3_2_141_2","doi-asserted-by":"crossref","first-page":"2124","DOI":"10.18653\/v1\/2024.findings-emnlp.116","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2024","author":"Song Yifan","year":"2024","unstructured":"Yifan Song, Weimin Xiong, Xiutian Zhao, Dawei Zhu, Wenhao Wu, Ke Wang, Cheng Li, Wei Peng, and Sujian Li. 2024. AgentBank: Towards generalized LLM agents via fine-tuning on 50000+ interaction trajectories. In Findings of the Association for Computational Linguistics: EMNLP 2024. 2124\u20132141."},{"key":"e_1_3_2_142_2","doi-asserted-by":"crossref","first-page":"7584","DOI":"10.18653\/v1\/2024.acl-long.409","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Song Yifan","year":"2024","unstructured":"Yifan Song, Da Yin, Xiang Yue, Jie Huang, Sujian Li, and Bill Yuchen Lin. 2024. Trial and error: Exploration-based trajectory optimization of LLM agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 7584\u20137600."},{"key":"e_1_3_2_143_2","unstructured":"Haoyang Su Renqi Chen Shixiang Tang Zhenfei Yin Xinzhe Zheng Jinzhe Li Biqing Qi Qi Wu Hui Li Wanli Ouyang Philip Torr Bowen Zhou and Nanqing Dong. 2024. Many heads are better than one: Improved scientific idea generation by a LLM-based multi-agent system. arXiv:2410.09403. Retrieved from https:\/\/arxiv.org\/abs\/2410.09403"},{"key":"e_1_3_2_144_2","unstructured":"Vighnesh Subramaniam Yilun Du Joshua B. Tenenbaum Antonio Torralba Shuang Li and Igor Mordatch. 2025. Multiagent finetuning: Self improvement with diverse reasoning chains. arXiv:2501.05707. Retrieved from https:\/\/arxiv.org\/abs\/2501.05707"},{"key":"e_1_3_2_145_2","doi-asserted-by":"crossref","first-page":"8052","DOI":"10.18653\/v1\/2024.emnlp-main.458","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Sun Hao","year":"2024","unstructured":"Hao Sun, Jiayi Wu, Hengyi Cai, Xiaochi Wei, Yue Feng, Bo Wang, Shuaiqiang Wang, Yan Zhang, and Dawei Yin. 2024. AdaSwitch: Adaptive switching between small and large agents for effective cloud-local collaborative learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 8052\u20138062."},{"key":"e_1_3_2_146_2","article-title":"Policy gradient methods for reinforcement learning with function approximation","author":"Sutton Richard S.","year":"1999","unstructured":"Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour. 1999. Policy gradient methods for reinforcement learning with function approximation. Advances in Neural Information Processing Systems 12 (1999), 1057\u20131063.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_147_2","unstructured":"Zhengwei Tao Ting-En Lin Xiancai Chen Hangyu Li Yuchuan Wu Yongbin Li Zhi Jin Fei Huang Dacheng Tao and Jingren Zhou. 2024. A survey on self-evolution of large language models. arXiv:2404.14387. Retrieved from https:\/\/arxiv.org\/abs\/2404.14387"},{"key":"e_1_3_2_148_2","first-page":"arXiv\u20132505","article-title":"InternAgent: When agent becomes the scientist\u2013building closed-loop system from hypothesis to verification","author":"Team InternAgent","year":"2025","unstructured":"InternAgent Team, Bo Zhang, Shiyang Feng, Xiangchao Yan, Jiakang Yuan, Runmin Ma, Yusong Hu, Zhiyin Yu, Xiaohan He, Songtao Huang, et\u00a0al. 2025. InternAgent: When agent becomes the scientist\u2013building closed-loop system from hypothesis to verification. arXiv e-prints (2025), arXiv\u20132505.","journal-title":"arXiv e-prints"},{"key":"e_1_3_2_149_2","unstructured":"Microsoft Team. 2025. Semantic Kernel. Retrieved December 2 2025 from https:\/\/github.com\/microsoft\/semantic-kernel"},{"key":"e_1_3_2_150_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00475"},{"key":"e_1_3_2_151_2","doi-asserted-by":"crossref","first-page":"7601","DOI":"10.18653\/v1\/2024.acl-long.410","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Trung Luong","year":"2024","unstructured":"Luong Trung, Xinbo Zhang, Zhanming Jie, Peng Sun, Xiaoran Jin, and Hang Li. 2024. Reft: Reasoning with reinforced fine-tuning. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 7601\u20137614."},{"key":"e_1_3_2_152_2","doi-asserted-by":"crossref","unstructured":"Dennis Ulmer Elman Mansimov Kaixiang Lin Justin Sun Xibin Gao and Yi Zhang. 2024. Bootstrapping llm-based task-oriented dialogue agents via self-talk. arXiv:2401.05033. Retrieved from https:\/\/arxiv.org\/abs\/2401.05033","DOI":"10.18653\/v1\/2024.findings-acl.566"},{"key":"e_1_3_2_153_2","doi-asserted-by":"crossref","first-page":"10583","DOI":"10.18653\/v1\/2024.acl-long.570","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wang Boshi","year":"2024","unstructured":"Boshi Wang, Hao Fang, Jason Eisner, Benjamin Van Durme, and Yu Su. 2024. LLMs in the imaginarium: Tool learning through simulated trial and error. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 10583\u201310604."},{"key":"e_1_3_2_154_2","unstructured":"Guanzhi Wang Yuqi Xie Yunfan Jiang Ajay Mandlekar Chaowei Xiao Yuke Zhu Linxi Fan and Anima Anandkumar. 2023. Voyager: An open-ended embodied agent with large language models. arXiv:2305.16291. Retrieved from https:\/\/arxiv.org\/abs\/2305.16291"},{"key":"e_1_3_2_155_2","doi-asserted-by":"crossref","first-page":"7626","DOI":"10.18653\/v1\/2024.findings-emnlp.448","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2024","author":"Wang Hanlin","year":"2024","unstructured":"Hanlin Wang, Chak Tou Leong, Jian Wang, and Wenjie Li. 2024. E2CL: Exploration-based error correction learning for embodied agents. In Findings of the Association for Computational Linguistics: EMNLP 2024. 7626\u20137639."},{"key":"e_1_3_2_156_2","doi-asserted-by":"crossref","unstructured":"Hanlin Wang Jian Wang Chak Tou Leong and Wenjie Li. 2025. Steca: Step-level trajectory calibration for llm agent learning. arXiv:2502.14276. Retrieved from https:\/\/arxiv.org\/abs\/2502.14276","DOI":"10.18653\/v1\/2025.findings-acl.604"},{"key":"e_1_3_2_157_2","unstructured":"Pei Wang Yanan Wu Zekun Wang Jiaheng Liu Xiaoshuai Song Zhongyuan Peng Ken Deng Chenchen Zhang Jiakai Wang Junran Peng et\u00a0al. 2024. Mtu-bench: A multi-granularity tool-use benchmark for large language models. arXiv:2410.11710. Retrieved from https:\/\/arxiv.org\/abs\/2410.11710"},{"key":"e_1_3_2_158_2","doi-asserted-by":"crossref","first-page":"11279","DOI":"10.18653\/v1\/2022.emnlp-main.775","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Wang Ruoyao","year":"2022","unstructured":"Ruoyao Wang, Peter Jansen, Marc-Alexandre C\u00f4t\u00e9, and Prithviraj Ammanabrolu. 2022. ScienceWorld: Is your agent smarter than a 5th grader?. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 11279\u201311298."},{"key":"e_1_3_2_159_2","unstructured":"Renxi Wang Haonan Li Xudong Han Yixuan Zhang and Timothy Baldwin. 2024. Learning from failure: Integrating negative examples when fine-tuning large language models as agents. arXiv:2402.11651. Retrieved from https:\/\/arxiv.org\/abs\/2402.11651"},{"key":"e_1_3_2_160_2","first-page":"1","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Wang Xingyao","year":"2024","unstructured":"Xingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen, Lifan Yuan, Hao Peng, and Heng Ji. 2024. MINT: Evaluating LLMs in multi-turn interaction with tools and language feedback. In Proceedings of the 12th International Conference on Learning Representations. 1\u201335."},{"key":"e_1_3_2_161_2","doi-asserted-by":"crossref","first-page":"4891","DOI":"10.18653\/v1\/2024.emnlp-main.281","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Wang Zheng","year":"2024","unstructured":"Zheng Wang, Zhongyang Li, Zeren Jiang, Dandan Tu, and Wei Shi. 2024. Crafting personalized agents through retrieval-augmented generation on editable memory graphs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 4891\u20134906."},{"key":"e_1_3_2_162_2","article-title":"Q-learning","author":"Watkins Christopher J. C. H.","year":"1992","unstructured":"Christopher J. C. H. Watkins and Peter Dayan. 1992. Q-learning. Machine Learning 8, 3\u20134 (1992), 279\u2013292.","journal-title":"Machine Learning"},{"key":"e_1_3_2_163_2","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, Denny Zhou, et\u00a0al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (2022), 24824\u201324837.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_164_2","unstructured":"Muning Wen Ziyu Wan Weinan Zhang Jun Wang and Ying Wen. 2024. Reinforcing language agents via policy optimization with action decomposition. arXiv:2405.15821. Retrieved from https:\/\/arxiv.org\/abs\/2405.15821"},{"key":"e_1_3_2_165_2","volume-title":"Proceedings of the 38th Conference on Neural Information Processing Systems Datasets and Benchmarks Track","author":"Wu Cheng-Kuang","year":"2024","unstructured":"Cheng-Kuang Wu, Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, and Hung-yi Lee. 2024. StreamBench: Towards benchmarking continuous improvement of language agents. In Proceedings of the 38th Conference on Neural Information Processing Systems Datasets and Benchmarks Track."},{"key":"e_1_3_2_166_2","first-page":"68082","article-title":"ivideogpt: Interactive videogpts are scalable world models","volume":"37","author":"Wu Jialong","year":"2025","unstructured":"Jialong Wu, Shaofeng Yin, Ningya Feng, Xu He, Dong Li, Jianye Hao, and Mingsheng Long. 2025. ivideogpt: Interactive videogpts are scalable world models. Advances in Neural Information Processing Systems 37 (2025), 68082\u201368119.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_167_2","unstructured":"Qingyun Wu Gagan Bansal Jieyu Zhang Yiran Wu Beibin Li Erkang Zhu Li Jiang Xiaoyun Zhang Shaokun Zhang Jiale Liu et\u00a0al. 2023. Autogen: Enabling next-gen llm applications via multi-agent conversation. arXiv:2308.08155. Retrieved from https:\/\/arxiv.org\/abs\/2308.08155"},{"key":"e_1_3_2_168_2","first-page":"25981","article-title":"AvaTaR: Optimizing LLM agents for tool usage via contrastive reasoning","volume":"37","author":"Wu Shirley","year":"2025","unstructured":"Shirley Wu, Shiyu Zhao, Qian Huang, Kexin Huang, Michihiro Yasunaga, Kaidi Cao, Vassilis Ioannidis, Karthik Subbian, Jure Leskovec, and James Y. Zou. 2025. AvaTaR: Optimizing LLM agents for tool usage via contrastive reasoning. Advances in Neural Information Processing Systems 37 (2025), 25981\u201326010.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_169_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-024-4222-0"},{"key":"e_1_3_2_170_2","unstructured":"Zhiheng Xi Yiwen Ding Wenxiang Chen Boyang Hong Honglin Guo Junzhe Wang Dingwen Yang Chenyang Liao Xin Guo Wei He et\u00a0al. 2024. AgentGym: Evolving large language model-based agents across diverse environments. arXiv:2406.04151. Retrieved from https:\/\/arxiv.org\/abs\/2406.04151"},{"key":"e_1_3_2_171_2","unstructured":"Zhiheng Xi Jixuan Huang Chenyang Liao Baodai Huang Honglin Guo Jiaqi Liu Rui Zheng Junjie Ye Jiazheng Zhang Wenxiang Chen et\u00a0al. 2025. Agentgym-rl: Training llm agents for long-horizon decision making through multi-turn reinforcement learning. arXiv:2509.08755. Retrieved from https:\/\/arxiv.org\/abs\/2509.08755"},{"key":"e_1_3_2_172_2","doi-asserted-by":"crossref","first-page":"4650","DOI":"10.18653\/v1\/2024.emnlp-main.268","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Xiang Yufei","year":"2024","unstructured":"Yufei Xiang, Yiqun Shen, Yeqin Zhang, and Nguyen Cam-Tu. 2024. Retrospex: Language agent meets offline reinforcement learning critic. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 4650\u20134666."},{"key":"e_1_3_2_173_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00965"},{"key":"e_1_3_2_174_2","doi-asserted-by":"crossref","unstructured":"Yihang Xiao Jinyi Liu Yan Zheng Xiaohan Xie Jianye Hao Mingzhi Li Ruitao Wang Fei Ni Yuxiao Li Jintian Luo et\u00a0al. 2024. Cellagent: An llm-driven multi-agent framework for automated single-cell data analysis. arXiv:2407.09811. Retrieved from https:\/\/arxiv.org\/abs\/2407.09811","DOI":"10.1101\/2024.05.13.593861"},{"key":"e_1_3_2_175_2","unstructured":"Yijia Xiao Edward Sun Di Luo and Wei Wang. 2024. TradingAgents: Multi-agents LLM financial trading framework. arXiv:2412.20138. Retrieved from https:\/\/arxiv.org\/abs\/2412.20138"},{"key":"e_1_3_2_176_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Xiao Ziyang","year":"2023","unstructured":"Ziyang Xiao, Dongxiang Zhang, Yangjun Wu, Lilin Xu, Yuan Jessica Wang, Xiongwei Han, Xiaojin Fu, Tao Zhong, Jia Zeng, Mingli Song, et\u00a0al. 2023. Chain-of-experts: When llms meet complex operations research problems. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_2_177_2","doi-asserted-by":"crossref","unstructured":"Weimin Xiong Yifan Song Qingxiu Dong Bingchan Zhao Feifan Song Xun Wang and Sujian Li. 2025. MPO: Boosting LLM agents with meta plan optimization. arXiv:2503.02682. Retrieved from https:\/\/arxiv.org\/abs\/2503.02682","DOI":"10.18653\/v1\/2025.findings-emnlp.210"},{"key":"e_1_3_2_178_2","doi-asserted-by":"crossref","first-page":"1556","DOI":"10.18653\/v1\/2024.emnlp-main.93","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Xiong Weimin","year":"2024","unstructured":"Weimin Xiong, Yifan Song, Xiutian Zhao, Wenhao Wu, Xun Wang, Ke Wang, Cheng Li, Wei Peng, and Sujian Li. 2024. Watch every step! LLM agent learning via iterative step-level process refinement. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 1556\u20131572."},{"key":"e_1_3_2_179_2","unstructured":"Fangzhi Xu Qiushi Sun Kanzhi Cheng Jun Liu Yu Qiao and Zhiyong Wu. 2024. Interactive evolution: A neural-symbolic self-training framework for large language models. arXiv:2406.11736. Retrieved from https:\/\/arxiv.org\/abs\/2406.11736"},{"key":"e_1_3_2_180_2","unstructured":"Renjun Xu and Jingwen Peng. 2025. A comprehensive survey of deep research: Systems methodologies and applications. arXiv:2506.12594. Retrieved from https:\/\/arxiv.org\/abs\/2506.12594"},{"key":"e_1_3_2_181_2","first-page":"5985","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Xu Tianyang","year":"2024","unstructured":"Tianyang Xu, Shujin Wu, Shizhe Diao, Xiaoze Liu, Xingyao Wang, Yangyi Chen, and Jing Gao. 2024. SaySelf: Teaching LLMs to express confidence with self-reflective rationales. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 5985\u20135998."},{"key":"e_1_3_2_182_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations, ICLR 2024","author":"Yang Chengrun","year":"2024","unstructured":"Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou, and Xinyun Chen. 2024. Large language models as optimizers. In Proceedings of the 12th International Conference on Learning Representations, ICLR 2024."},{"key":"e_1_3_2_183_2","doi-asserted-by":"crossref","unstructured":"Hongyang Yang Boyu Zhang Neng Wang Cheng Guo Xiaoli Zhang Likun Lin Junlin Wang Tianyu Zhou Mao Guan Runjia Zhang et\u00a0al. 2024. FinRobot: An open-source AI agent platform for financial applications using large language models. arXiv:2405.14767. Retrieved from https:\/\/arxiv.org\/abs\/2405.14767","DOI":"10.2139\/ssrn.4841493"},{"key":"e_1_3_2_184_2","volume-title":"Proceedings of the 13th International Conference on Learning Representations","author":"Yang John","year":"2025","unstructured":"John Yang, Carlos E. Jimenez, Alex L. Zhang, Kilian Lieret, Joyce Yang, Xindi Wu, Ori Press, Niklas Muennighoff, Gabriel Synnaeve, Karthik R. Narasimhan, et\u00a0al. 2025. SWE-bench multimodal: Do AI systems generalize to visual software domains?. In Proceedings of the 13th International Conference on Learning Representations."},{"key":"e_1_3_2_185_2","article-title":"Intercode: Standardizing and benchmarking interactive coding with execution feedback","author":"Yang John","year":"2023","unstructured":"John Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao. 2023. Intercode: Standardizing and benchmarking interactive coding with execution feedback. Advances in Neural Information Processing Systems 36 (2023), 23826\u201323854.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_186_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02482"},{"key":"e_1_3_2_187_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1259"},{"key":"e_1_3_2_188_2","first-page":"20744","article-title":"Webshop: Towards scalable real-world web interaction with grounded language agents","volume":"35","author":"Yao Shunyu","year":"2022","unstructured":"Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems 35 (2022), 20744\u201320757.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_189_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Yao Shunyu","year":"2023","unstructured":"Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing reasoning and acting in language models. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_190_2","unstructured":"Weiran Yao Shelby Heinecke Juan Carlos Niebles Zhiwei Liu Yihao Feng Le Xue Rithesh Murthy Zeyuan Chen Jianguo Zhang Devansh Arpit et\u00a0al. 2023. Retroformer: Retrospective large language agents with policy gradient optimization. arXiv:2308.02151. Retrieved from https:\/\/arxiv.org\/abs\/2308.02151"},{"key":"e_1_3_2_191_2","doi-asserted-by":"publisher","DOI":"10.1093\/bib\/bbae693"},{"key":"e_1_3_2_192_2","doi-asserted-by":"crossref","first-page":"12380","DOI":"10.18653\/v1\/2024.acl-long.670","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Yin Da","year":"2024","unstructured":"Da Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Chandu, Kai-Wei Chang, Yejin Choi, and Bill Yuchen Lin. 2024. Agent lumos: Unified and modular training for open-source language agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 12380\u201312403."},{"key":"e_1_3_2_193_2","first-page":"595","volume-title":"Proceedings of the AAAI Symposium Series","author":"Yu Yangyang","year":"2024","unstructured":"Yangyang Yu, Haohang Li, Zhi Chen, Yuechen Jiang, Yang Li, Denghui Zhang, Rong Liu, Jordan W. Suchow, and Khaldoun Khashanah. 2024. FinMem: A performance-enhanced LLM trading agent with layered memory and character design. In Proceedings of the AAAI Symposium Series. 595\u2013597."},{"key":"e_1_3_2_194_2","first-page":"137010","article-title":"Fincon: A synthesized llm multi-agent system with conceptual verbal reinforcement for enhanced financial decision making","volume":"37","author":"Yu Yangyang","year":"2025","unstructured":"Yangyang Yu, Zhiyuan Yao, Haohang Li, Zhiyang Deng, Yuechen Jiang, Yupeng Cao, Zhi Chen, Jordan Suchow, Zhenyu Cui, Rong Liu, et\u00a0al. 2025. Fincon: A synthesized llm multi-agent system with conceptual verbal reinforcement for enhanced financial decision making. Advances in Neural Information Processing Systems 37 (2025), 137010\u2013137045.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_195_2","first-page":"1","volume-title":"Proceedings of the 15th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics","author":"Yue Ling","year":"2024","unstructured":"Ling Yue, Sixue Xing, Jintai Chen, and Tianfan Fu. 2024. Clinicalagent: Clinical trial multi-agent system with large language model-based reasoning. In Proceedings of the 15th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics. 1\u201310."},{"key":"e_1_3_2_196_2","unstructured":"Shengbin Yue Siyuan Wang Wei Chen Xuanjing Huang and Zhongyu Wei. 2024. Synergistic multi-agent framework with trajectory learning for knowledge-intensive tasks. arXiv:2407.09893. Retrieved from https:\/\/arxiv.org\/abs\/2407.09893"},{"key":"e_1_3_2_197_2","unstructured":"Kamer Ali Yuksel and Hassan Sawaf. 2024. A Multi-AI Agent System for Autonomous Optimization of Agentic AI Solutions via Iterative Refinement and LLM-Driven Feedback Loops. arXiv:2412.17149. Retrieved from https:\/\/arxiv.org\/abs\/2412.17149"},{"key":"e_1_3_2_198_2","unstructured":"Aohan Zeng Mingdao Liu Rui Lu Bowen Wang Xiao Liu Yuxiao Dong and Jie Tang. 2023. Agenttuning: Enabling generalized agent abilities for llms. arXiv:2310.12823. Retrieved from https:\/\/arxiv.org\/abs\/2310.12823"},{"key":"e_1_3_2_199_2","unstructured":"Yifan Zeng Yiran Wu Xiao Zhang Huazheng Wang and Qingyun Wu. 2024. Autodefense: Multi-agent llm defense against jailbreak attacks. arXiv:2403.04783. Retrieved from https:\/\/arxiv.org\/abs\/2403.04783"},{"key":"e_1_3_2_200_2","unstructured":"Daochen Zha Kwei-Herng Lai Yuanpu Cao Songyi Huang Ruzhe Wei Junyu Guo and Xia Hu. 2019. Rlcard: A toolkit for reinforcement learning in card games. arXiv:1910.04376. Retrieved from https:\/\/arxiv.org\/abs\/1910.04376"},{"key":"e_1_3_2_201_2","first-page":"110935","article-title":"Fine-tuning large vision-language models as decision-making agents via reinforcement learning","volume":"37","author":"Zhai Simon","year":"2024","unstructured":"Simon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan, Peter Tong, Yifei Zhou, Alane Suhr, Saining Xie, Yann LeCun, Yi Ma, et\u00a0al. 2024. Fine-tuning large vision-language models as decision-making agents via reinforcement learning. Advances in Neural Information Processing Systems 37 (2024), 110935\u2013110971.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_202_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Zhang Hongxin","year":"2024","unstructured":"Hongxin Zhang, Weihua Du, Jiaming Shan, Qinhong Zhou, Yilun Du, Joshua B. Tenenbaum, Tianmin Shu, and Chuang Gan. 2024. Building cooperative embodied agents modularly with large language models. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_203_2","unstructured":"Jianguo Zhang Tian Lan Rithesh Murthy Zhiwei Liu Weiran Yao Ming Zhu Juntao Tan Thai Hoang Zuxin Liu Liangwei Yang et\u00a0al. 2024. Agentohana: Design unified data and training pipeline for effective agent learning. arXiv:2402.15506. Retrieved from https:\/\/arxiv.org\/abs\/2402.15506"},{"key":"e_1_3_2_204_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-60990-0_12"},{"key":"e_1_3_2_205_2","unstructured":"Qizheng Zhang Changran Hu Shubhangi Upasani Boyuan Ma Fenglu Hong Vamsidhar Kamanuru Jay Rainton Chen Wu Mengmeng Ji Hanchen Li et\u00a0al. 2025. Agentic context engineering: Evolving contexts for self-improving language models. arXiv:2510.04618. Retrieved from https:\/\/arxiv.org\/abs\/2510.04618"},{"key":"e_1_3_2_206_2","unstructured":"Shengyu Zhang Linfeng Dong Xiaoya Li Sen Zhang Xiaofei Sun Shuhe Wang Jiwei Li Runyi Hu Tianwei Zhang Fei Wu et\u00a0al. 2023. Instruction tuning for large language models: A survey. arXiv:2308.10792. Retrieved from https:\/\/arxiv.org\/abs\/2308.10792"},{"key":"e_1_3_2_207_2","first-page":"60315","article-title":"Offline training of language model agents with functions as learnable weights","volume":"235","author":"Zhang Shaokun","year":"2024","unstructured":"Shaokun Zhang, Jieyu Zhang, Jiale Liu, Linxin Song, Chi Wang, Ranjay Krishna, and Qingyun Wu. 2024. Offline training of language model agents with functions as learnable weights. Proceedings of Machine Learning Research 235 (2024), 60315\u201360335.","journal-title":"Proceedings of Machine Learning Research"},{"key":"e_1_3_2_208_2","doi-asserted-by":"crossref","first-page":"5348","DOI":"10.18653\/v1\/2024.acl-long.292","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhang Wenqi","year":"2024","unstructured":"Wenqi Zhang, Ke Tang, Hai Wu, Mengna Wang, Yongliang Shen, Guiyang Hou, Zeqi Tan, Peng Li, Yueting Zhuang, and Weiming Lu. 2024. Agent-Pro: Learning to evolve via policy-level reflection and optimization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 5348\u20135375."},{"key":"e_1_3_2_209_2","doi-asserted-by":"crossref","first-page":"4314","DOI":"10.1145\/3637528.3671801","volume-title":"Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Zhang Wentao","year":"2024","unstructured":"Wentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun, Jiaze Sun, Molei Qin, Xinyi Li, Yuqing Zhao, Yilei Zhao, Xinyu Cai, et\u00a0al. 2024. A multimodal foundation agent for financial trading: Tool-augmented, diversified, and generalist. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 4314\u20134325."},{"key":"e_1_3_2_210_2","unstructured":"Yiming Zhang Zheng Chang Wentao Cai MengXing Ren Kang Yuan Yining Sun and Zenghui Ding. 2025. IIMedGPT: Promoting large language model capabilities of medical tasks by efficient human preference alignment. arXiv:2501.02869. Retrieved from https:\/\/arxiv.org\/abs\/2501.02869"},{"key":"e_1_3_2_211_2","unstructured":"Zeyu Zhang Xiaohe Bo Chen Ma Rui Li Xu Chen Quanyu Dai Jieming Zhu Zhenhua Dong and Ji-Rong Wen. 2024. A survey on the memory mechanism of large language model based agents. arXiv:2404.13501. Retrieved from https:\/\/arxiv.org\/abs\/2404.13501"},{"key":"e_1_3_2_212_2","first-page":"19632","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Zhao Andrew","year":"2024","unstructured":"Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, and Gao Huang. 2024. Expel: LLM agents are experiential learners. In Proceedings of the AAAI Conference on Artificial Intelligence. 19632\u201319642."},{"key":"e_1_3_2_213_2","doi-asserted-by":"crossref","first-page":"6401","DOI":"10.18653\/v1\/2024.emnlp-main.367","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Zhao Qi","year":"2024","unstructured":"Qi Zhao, Haotian Fu, Chen Sun, and George Konidaris. 2024. EPO: Hierarchical LLM agents with environment preference optimization. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 6401\u20136415."},{"key":"e_1_3_2_214_2","article-title":"Lyra: Orchestrating dual correction in automated theorem proving","author":"Zheng Chuanyang","year":"2024","unstructured":"Chuanyang Zheng, Haiming Wang, Enze Xie, Zhengying Liu, Jiankai Sun, Huajian Xin, Jianhao Shen, Zhenguo Li, and Yu Li. 2024. Lyra: Orchestrating dual correction in automated theorem proving. Transactions on Machine Learning Research (2024).","journal-title":"Transactions on Machine Learning Research"},{"key":"e_1_3_2_215_2","doi-asserted-by":"crossref","unstructured":"Yuxiang Zheng Dayuan Fu Xiangkun Hu Xiaojie Cai Lyumanshan Ye Pengrui Lu and Pengfei Liu. 2025. Deepresearcher: Scaling deep research via reinforcement learning in real-world environments. arXiv:2504.03160. Retrieved from https:\/\/arxiv.org\/abs\/2504.03160","DOI":"10.18653\/v1\/2025.emnlp-main.22"},{"key":"e_1_3_2_216_2","first-page":"4575","article-title":"Star-agents: Automatic data optimization with LLM agents for instruction tuning","volume":"37","author":"Zhou Hang","year":"2024","unstructured":"Hang Zhou, Yehui Tang, Haochen Qin, Yujie Yang, Renren Jin, Deyi Xiong, Kai Han, and Yunhe Wang. 2024. Star-agents: Automatic data optimization with LLM agents for instruction tuning. Advances in Neural Information Processing Systems 37 (2024), 4575\u20134597.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_217_2","unstructured":"Han Zhou Xingchen Wan Ruoxi Sun Hamid Palangi Shariq Iqbal Ivan Vuli\u0107 Anna Korhonen and Sercan \u00d6 Ar\u0131k. 2025. Multi-agent design: Optimizing agents with better prompts and topologies. arXiv:2502.02533. Retrieved from https:\/\/arxiv.org\/abs\/2502.02533"},{"key":"e_1_3_2_218_2","doi-asserted-by":"crossref","first-page":"2922","DOI":"10.18653\/v1\/2024.findings-naacl.184","volume-title":"Findings of the Association for Computational Linguistics: NAACL 2024","author":"Zhou Qinhao","year":"2024","unstructured":"Qinhao Zhou, Zihan Zhang, Xiang Xiang, Ke Wang, Yuchuan Wu, and Yongbin Li. 2024. Enhancing the general agent capabilities of low-paramter LLMs through tuning and multi-branch reasoning. In Findings of the Association for Computational Linguistics: NAACL 2024. 2922\u20132931."},{"key":"e_1_3_2_219_2","unstructured":"Shuyan Zhou Frank F. Xu Hao Zhu Xuhui Zhou Robert Lo Abishek Sridhar Xianyi Cheng Tianyue Ou Yonatan Bisk Daniel Fried et\u00a0al. 2023. Webarena: A realistic web environment for building autonomous agents. arXiv:2307.13854. Retrieved from https:\/\/arxiv.org\/abs\/2307.13854"},{"key":"e_1_3_2_220_2","unstructured":"Wangchunshu Zhou Yixin Ou Shengwei Ding Long Li Jialong Wu Tiannan Wang Jiamin Chen Shuai Wang Xiaohua Xu Ningyu Zhang et\u00a0al. 2024. Symbolic learning enables self-evolving agents. arXiv:2406.18532. Retrieved from https:\/\/arxiv.org\/abs\/2406.18532"},{"key":"e_1_3_2_221_2","first-page":"10902","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Zhu Junda","year":"2024","unstructured":"Junda Zhu, Lingyong Yan, Haibo Shi, Dawei Yin, and Lei Sha. 2024. ATM: Adversarial tuning multi-agent system makes a robust retrieval-augmented generator. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 10902\u201310919."},{"key":"e_1_3_2_222_2","unstructured":"Kaiwen Zuo Yirui Jiang Fan Mo and Pietro Lio. 2024. KG4Diagnosis: A hierarchical multi-agent LLM framework with knowledge graph enhancement for medical diagnosis. arXiv:2412.16833. Retrieved from https:\/\/arxiv.org\/abs\/2412.16833"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3789261","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,11]],"date-time":"2026-02-11T23:21:10Z","timestamp":1770852070000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3789261"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,11]]},"references-count":221,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3789261"],"URL":"https:\/\/doi.org\/10.1145\/3789261","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,11]]},"assertion":[{"value":"2025-04-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-08","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}