{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T21:42:49Z","timestamp":1775857369497,"version":"3.50.1"},"reference-count":144,"publisher":"Association for Computing Machinery (ACM)","issue":"10","license":[{"start":{"date-parts":[[2026,4,2]],"date-time":"2026-04-02T00:00:00Z","timestamp":1775088000000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"NSF","award":["MSPA -2434666"],"award-info":[{"award-number":["MSPA -2434666"]}]},{"name":"NSF","award":["IIS-2211526"],"award-info":[{"award-number":["IIS-2211526"]}]},{"name":"Learning Engineering Virtual Institute","award":["P0362134 SUB00004699"],"award-info":[{"award-number":["P0362134 SUB00004699"]}]},{"name":"NSF Graduate Research Fellowship Progra"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>\n                    Large Language Model (LLM)-based agents have recently emerged as a new paradigm that extends the capabilities of LLMs beyond text generation to dynamic interaction with external environments. A critical challenge lies in ensuring their\n                    <jats:italic toggle=\"yes\">g<\/jats:italic>\n                    eneralizability \u2013 the ability to maintain consistently high performance across varied instructions, tasks, environments, and domains, especially those different from the agent\u2019s fine-tuning data. Despite growing interest, the concept of generalizability in LLM-based agents remains underdefined, and systematic approaches to measure and improve it are lacking. We provide the first comprehensive review of generalizability in LLM-based agents. We begin by clarifying the definition and boundaries of agent generalizability. We then review existing benchmarks. Next, we categorize strategies for improving generalizability into three groups: methods targeting the backbone LLM, targeting agent components, and targeting their interactions. Furthermore, we introduce the distinction between\n                    <jats:italic toggle=\"yes\">g<\/jats:italic>\n                    eneralizable frameworks and\n                    <jats:italic toggle=\"yes\">g<\/jats:italic>\n                    eneralizable agents and outline how generalizable frameworks can be translated into agent-level generalizability. Finally, we identify future directions, including the development of standardized evaluation frameworks, variance- and cost-based metrics, and hybrid approaches that integrate methodological innovations with agent architecture-level designs. We aim to establish a foundation for principled research on building LLM-based agents that generalize reliably across diverse real-world applications.\n                  <\/jats:p>","DOI":"10.1145\/3794858","type":"journal-article","created":{"date-parts":[[2026,2,7]],"date-time":"2026-02-07T20:24:18Z","timestamp":1770495858000},"page":"1-44","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Generalizability of Large Language Model-Based Agents: A Comprehensive Survey"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5533-8514","authenticated-orcid":false,"given":"Minxing","family":"Zhang","sequence":"first","affiliation":[{"name":"Computer Science, Duke University","place":["Durham, United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-9957-2780","authenticated-orcid":false,"given":"Yi","family":"Yang","sequence":"additional","affiliation":[{"name":"Computer Science, Duke University","place":["Durham, United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-1096-1820","authenticated-orcid":false,"given":"Roy","family":"Xie","sequence":"additional","affiliation":[{"name":"Computer Science, Duke University","place":["Durham, United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6874-9515","authenticated-orcid":false,"given":"Bhuwan","family":"Dhingra","sequence":"additional","affiliation":[{"name":"Computer Science, Duke University","place":["Durham, United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-7270-6194","authenticated-orcid":false,"given":"Shuyan","family":"Zhou","sequence":"additional","affiliation":[{"name":"Computer Science, Duke University","place":["Durham, United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2200-8711","authenticated-orcid":false,"given":"Jian","family":"Pei","sequence":"additional","affiliation":[{"name":"Duke University","place":["Durham, United States"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,2]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"2024. Retrieved February 22 2026 from https:\/\/www.pingcap.com\/article\/common-issues-in-implementing-llm-agents\/"},{"key":"e_1_3_3_3_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et\u00a0al. 2023. Gpt-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_3_4_2","unstructured":"Renat Aksitov Sobhan Miryoosefi Zonglin Li Daliang Li Sheila Babayan Kavya Kopparapu Zachary Fisher Ruiqi Guo Sushant Prakash Pranesh Srinivasan et\u00a0al. 2023. Rest meets react: Self-improvement for multi-step reasoning LLM agent. arXiv:2312.10003. Retrieved from https:\/\/arxiv.org\/abs\/2312.10003"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","unstructured":"Sacha Alanoca Shira Gur-Arieh Tom Zick and Kevin Klyman. 2025. Comparing apples to oranges: A taxonomy for navigating the global landscape of AI regulation. In Proceedings of the 2025 ACM Conference on Fairness Accountability and Transparency (FAccT \u201925). Association for Computing Machinery New York NY USA. 914\u2013937. DOI:10.1145\/3715275.3732059","DOI":"10.1145\/3715275.3732059"},{"key":"e_1_3_3_6_2","unstructured":"Alon Albalak Yanai Elazar Sang Michael Xie Shayne Longpre Nathan Lambert Xinyi Wang Niklas Muennighoff Bairu Hou Liangming Pan Haewon Jeong et\u00a0al. 2024. A survey on data selection for language models. arXiv:2402.16827. Retrieved from https:\/\/arxiv.org\/abs\/2402.16827"},{"key":"e_1_3_3_7_2","unstructured":"Anthropic. 2024. Model Context Protocol (MCP). Retrieved February 22 2026 from https:\/\/modelcontextprotocol.io"},{"key":"e_1_3_3_8_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Asai Akari","year":"2023","unstructured":"Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2023. Self-rag: Learning to retrieve, generate, and critique through self-reflection. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_3_9_2","doi-asserted-by":"crossref","unstructured":"Ge Bai Jie Liu Xingyuan Bu Yancheng He Jiaheng Liu Zhanhui Zhou Zhuoran Lin Wenbo Su Tiezheng Ge Bo Zheng and Wanli Ouyang. 2024. Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. Association for Computational Bangkok Thailand. 1 (2024) 7421\u20137454.","DOI":"10.18653\/v1\/2024.acl-long.401"},{"key":"e_1_3_3_10_2","unstructured":"Yuntao Bai Andy Jones Kamal Ndousse Amanda Askell Anna Chen Nova DasSarma Dawn Drain Stanislav Fort Deep Ganguli Tom Henighan et\u00a0al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv:2204.05862. Retrieved from https:\/\/arxiv.org\/abs\/2204.05862"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1177\/0890334420906850"},{"key":"e_1_3_3_12_2","doi-asserted-by":"crossref","unstructured":"Angana Borah and Rada Mihalcea. 2024. Towards implicit bias detection and mitigation in multi-agent LLM interactions. In Findings of the Association for Computational Linguistics: EMNLP\u201924. Association for Computational Linguistics Miami Florida USA. 306\u20139326.","DOI":"10.18653\/v1\/2024.findings-emnlp.545"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2023.3288409"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-7091-2668-4_10"},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCIAIG.2012.2186810"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0031-9406(05)61211-4"},{"key":"e_1_3_3_17_2","unstructured":"Stephen Casper Xander Davies Claudia Shi Thomas Krendl Gilbert J\u00e9r\u00e9my Scheurer Javier Rando Rachel Freedman Tomasz Korbak David Lindner Pedro Freire et\u00a0al. 2023. Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv:2307.15217. Retrieved from https:\/\/arxiv.org\/abs\/2307.15217"},{"key":"e_1_3_3_18_2","volume-title":"LangChain: Build Context-Aware Reasoning Applications","author":"Chase Harrison","year":"2022","unstructured":"Harrison Chase and Ankush Gola. 2022. LangChain: Build Context-Aware Reasoning Applications. Retrieved February 22, 2026 from https:\/\/www.langchain.com"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.226"},{"key":"e_1_3_3_20_2","unstructured":"Yuxing Chen Weijie Wang Sylvain Lobry and Camille Kurtz. 2024. An LLM agent for automatic geospatial data analysis. arXiv:2410.18792. Retrieved from https:\/\/arxiv.org\/abs\/2410.18792"},{"key":"e_1_3_3_21_2","doi-asserted-by":"crossref","unstructured":"Yu-Neng Chuang Tianwei Xing Chia-Yuan Chang Zirui Liu Xun Chen and Xia Hu. 2024. Learning to compress prompt in natural language formats. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics Mexico City Mexico. 1 (2024) 7756\u20137767.","DOI":"10.18653\/v1\/2024.naacl-long.429"},{"key":"e_1_3_3_22_2","unstructured":"Christopher Clarke Karthik Krishnamurthy Walter Talamonti Yiping Kang Lingjia Tang and Jason Mars. 2024. One agent too many: User perspectives on approaches to multi-agent conversational AI. arXiv:2401.07123. Retrieved from https:\/\/arxiv.org\/abs\/2401.07123"},{"key":"e_1_3_3_23_2","first-page":"41","volume-title":"Proceedings of the Workshop on Computer Games","author":"C\u00f4t\u00e9 Marc-Alexandre","year":"2018","unstructured":"Marc-Alexandre C\u00f4t\u00e9, Akos K\u00e1d\u00e1r, Xingdi Yuan, Ben Kybartas, Tavian Barnes, Emery Fine, James Moore, Matthew Hausknecht, Layla El Asri, Mahmoud Adada, et\u00a0al. 2018. Textworld: A learning environment for text-based games. In Proceedings of the Workshop on Computer Games. Springer, 41\u201375."},{"key":"e_1_3_3_24_2","article-title":"Why AI alignment could be hard with modern deep learning","author":"Cotra Ajeya","year":"2021","unstructured":"Ajeya Cotra. 2021. Why AI alignment could be hard with modern deep learning. Cold Takes (2021). Retrieved February 22, 2026 from https:\/\/www.cold-takes.com\/why-ai-alignment-could-be-hard-with-modern-deep-learning\/","journal-title":"Cold Takes"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00323"},{"key":"e_1_3_3_26_2","article-title":"Mind2web: Towards a generalist agent for the web","volume":"36","author":"Deng Xiang","year":"2024","unstructured":"Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2024. Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems 36 (2024), 28091\u201328114.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_27_2","unstructured":"Yu Du Fangyun Wei and Hongyang Zhang. 2024. AnyTool: Self-reflective hierarchical agents for large-scale API calls. In Proceedings of the 41st International Conference on Machine Learning (ICML\u201924). JMLR.org. 235 Article 470 (2024) 11812\u201311829."},{"key":"e_1_3_3_28_2","volume-title":"Optimal Learning: Computational Procedures for Bayes-Adaptive Markov Decision Processes","author":"Duff Michael O\u2019Gordon","year":"2002","unstructured":"Michael O\u2019Gordon Duff. 2002. Optimal Learning: Computational Procedures for Bayes-Adaptive Markov Decision Processes. University of Massachusetts Amherst."},{"key":"e_1_3_3_29_2","doi-asserted-by":"crossref","unstructured":"Lutfi Eren Erdogan Nicholas Lee Siddharth Jha Sehoon Kim Ryan Tabrizi Suhong Moon Coleman Richard Charles Hooper Gopala Anumanchipalli Kurt Keutzer and Amir Gholami.2024. Tinyagent: Function calling at the edge. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for Computational Linguistics Miami Florida USA. 80\u201388.","DOI":"10.18653\/v1\/2024.emnlp-demo.9"},{"key":"e_1_3_3_30_2","unstructured":"Lutfi Eren Erdogan Nicholas Lee Sehoon Kim Suhong Moon Hiroki Furuta Gopala Anumanchipalli Kurt Keutzer and Amir Gholami. 2025. PLAN-AND-ACT: Improving planning of agents for long-horizon tasks. arXiv:2503.09572. Retrieved from https:\/\/arxiv.org\/abs\/2503.09572"},{"key":"e_1_3_3_31_2","doi-asserted-by":"crossref","unstructured":"Haishuo Fang Xiaodan Zhu and Iryna Gurevych. 2025. Preemptive detection and correction of misaligned actions in LLM agents. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics Suzhou China. 222\u2013244.","DOI":"10.18653\/v1\/2025.emnlp-main.12"},{"key":"e_1_3_3_32_2","unstructured":"Dayuan Fu Keqing He Yejie Wang Wentao Hong Zhuoma Gongque Weihao Zeng Wei Wang Jingang Wang Xunliang Cai and Weiran Xu. 2025. Agentrefine: Enhancing agent generalization through refinement tuning. arXiv:2501.01702. Retrieved from https:\/\/arxiv.org\/abs\/2501.01702"},{"key":"e_1_3_3_33_2","unstructured":"Leo Gao Stella Biderman Sid Black Laurence Golding Travis Hoppe Charles Foster Jason Phang Horace He Anish Thite Noa Nabeshima et\u00a0al. 2020. The pile: An 800gb dataset of diverse text for language modeling. arXiv:2101.00027. Retrieved from https:\/\/arxiv.org\/abs\/2101.00027"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.naacl-long.366"},{"key":"e_1_3_3_35_2","first-page":"49","volume-title":"The Colonization of Unfamiliar Landscapes","author":"Golledge Reginald G.","year":"2003","unstructured":"Reginald G. Golledge. 2003. Human wayfinding and cognitive maps. In The Colonization of Unfamiliar Landscapes. Marcy Rockman and James Steele (Eds.). Routledge, 49\u201354."},{"key":"e_1_3_3_36_2","doi-asserted-by":"crossref","unstructured":"Carlos G\u00f3mez-Rodr\u00edguez and Paul Williams. 2023. A confederacy of models: A comprehensive evaluation of LLMs on creative writing. In Findings of the Association for Computational Linguistics: EMNLP 2023. Association for Computational Linguistics Singapore. 14504\u201314528.","DOI":"10.18653\/v1\/2023.findings-emnlp.966"},{"key":"e_1_3_3_37_2","unstructured":"Arnav Gudibande Eric Wallace Charlie Snell Xinyang Geng Hao Liu Pieter Abbeel Sergey Levine and Dawn Song. 2023. The false promise of imitating proprietary LLMs. arXiv:2305.15717. Retrieved from https:\/\/arxiv.org\/abs\/2305.15717"},{"key":"e_1_3_3_38_2","unstructured":"Caglar Gulcehre Tom Le Paine Srivatsan Srinivasan Ksenia Konyushkova Lotte Weerts Abhishek Sharma Aditya Siddhant Alex Ahern Miaosen Wang Chenjie Gu et\u00a0al. 2023. Reinforced self-training (rest) for language modeling. arXiv:2308.08998. Retrieved from https:\/\/arxiv.org\/abs\/2308.08998"},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","unstructured":"Taicheng Guo Xiuying Chen Yaqi Wang Ruidi Chang Shichao Pei Nitesh V. Chawla Olaf Wiest and Xiangliang Zhang. 2024. Large language model based multi-agents: A survey of progress and challenges. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI \u201924) Article 890 (2024) 8048\u20138057. DOI:10.24963\/ijcai.2024\/890","DOI":"10.24963\/ijcai.2024\/890"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3123"},{"key":"e_1_3_3_41_2","doi-asserted-by":"crossref","unstructured":"Akash Gupta Ivaxi Sheth Vyas Raina Mark Gales and Mario Fritz. 2025. LLM task interference: An initial study on the impact of task-switch in conversational history. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2024.emnlp-main.811"},{"key":"e_1_3_3_42_2","unstructured":"Amine Ben Hassouna Hana Chaari and Ines Belhaj. 2024. LLM-Agent-UMF: LLM-based agent unified modeling framework for seamless integration of multi active\/passive core-agents. arxiv:2409.11393 [cs.SE]. Retrieved from https:\/\/arxiv.org\/abs\/2409.11393"},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","unstructured":"Feng He Tianqing Zhu Dayong Ye Bo Liu Wanlei Zhou and Philip S. Yu. 2025. The emerged security and privacy of LLM agent: A survey with case studies. ACM Comput. Surv. 58 6 Article 162 (April 2026) 36. DOI:10.1145\/3773080","DOI":"10.1145\/3773080"},{"key":"e_1_3_3_44_2","doi-asserted-by":"crossref","unstructured":"Mengkang Hu Tianxing Chen Qiguang Chen Yao Mu Wenqi Shao and Ping Luo. 2025. HiAgent: Hierarchical working memory management for solving long-horizon agent tasks with large language model. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics Vienna Austria. 1 (2025) 32779\u201332798.","DOI":"10.18653\/v1\/2025.acl-long.1575"},{"key":"e_1_3_3_45_2","doi-asserted-by":"crossref","unstructured":"Xiang Huang Sitao Cheng Shanshan Huang Jiayu Shen Yong Xu Chaoyun Zhang and Yuzhong Qu. 2024. Queryagent: A reliable and efficient reasoning framework with environmental feedback based self-correction. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics Bangkok Thailand. 1 (2024) 5014\u20135035.","DOI":"10.18653\/v1\/2024.acl-long.274"},{"key":"e_1_3_3_46_2","unstructured":"Xu Huang Weiwen Liu Xiaolong Chen Xingmei Wang Hao Wang Defu Lian Yasheng Wang Ruiming Tang and Enhong Chen. 2024. Understanding the planning of LLM agents: A survey. arXiv:2402.02716. Retrieved from https:\/\/arxiv.org\/abs\/2402.02716"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4842-8844-3_4"},{"key":"e_1_3_3_48_2","doi-asserted-by":"crossref","unstructured":"Jiarui Ji Runlin Lei Jialing Bi Zhewei Wei Yankai Lin Xuchen Pan Yaliang Li and Bolin Ding. 2025. LLM-based multi-agent systems are scalable graph generative models. In Findings of the Association for Computational Linguistics: ACL 2025. Association for Computational Linguistics Vienna Austria. 1492\u20131523.","DOI":"10.18653\/v1\/2025.findings-acl.78"},{"key":"e_1_3_3_49_2","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Antoine Roux Arthur Mensch Blanche Savary Chris Bamford Devendra Singh Chaplot Diego de las Casas Emma Bou Hanna Florian Bressand et\u00a0al. 2024. Mixtral of experts. arXiv:2401.04088. Retrieved from https:\/\/arxiv.org\/abs\/2401.04088"},{"key":"e_1_3_3_50_2","volume-title":"Proceedings of the ICML 2024 Workshop on LLMs and Cognition","author":"Jiang Bowen","year":"2024","unstructured":"Bowen Jiang, Yangxinyu Xie, Xiaomeng Wang, Weijie J. Su, Camillo Jose Taylor, and Tanwi Mallick. 2024. Multi-modal and multi-agent systems meet rationality: A survey. In Proceedings of the ICML 2024 Workshop on LLMs and Cognition."},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1631\/FITEE.1800514"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3678717.3691269"},{"key":"e_1_3_3_53_2","doi-asserted-by":"crossref","unstructured":"Douwe Kiela Max Bartolo Yixin Nie Divyansh Kaushik Atticus Geiger Zhengxuan Wu Bertie Vidgen Grusha Prasad Amanpreet Singh Pratik Ringshia et\u00a0al. 2021. Dynabench: Rethinking benchmarking in NLP. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics. 4110\u20134124.","DOI":"10.18653\/v1\/2021.naacl-main.324"},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-018-9646-y"},{"key":"e_1_3_3_55_2","first-page":"51991","article-title":"Camel: Communicative agents for \u201cmind\u201d exploration of large language model society","volume":"36","author":"Li Guohao","year":"2023","unstructured":"Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. Camel: Communicative agents for \u201cmind\u201d exploration of large language model society. Advances in Neural Information Processing Systems 36 (2023), 51991\u201352008.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_56_2","unstructured":"Yuanchun Li Hao Wen Weijun Wang Xiangyu Li Yizhen Yuan Guohong Liu Jiacheng Liu Wenxing Xu Xiang Wang Yi Sun et\u00a0al. 2024. Personal LLM agents: Insights and survey about the capability efficiency and security. arXiv:2401.05459. Retrieved from https:\/\/arxiv.org\/abs\/2401.05459"},{"key":"e_1_3_3_57_2","article-title":"Visual instruction tuning","volume":"36","author":"Liu Haotian","year":"2024","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. Advances in Neural Information Processing Systems 36 (2024), 34892\u201334916.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_58_2","unstructured":"Ruibo Liu Ruixin Yang Chenyan Jia Ge Zhang Denny Zhou Andrew M. Dai Diyi Yang and Soroush Vosoughi. 2023. Training socially aligned language models on simulated social interactions. arXiv:2305.16960. Retrieved from https:\/\/arxiv.org\/abs\/2305.16960"},{"key":"e_1_3_3_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3680616"},{"key":"e_1_3_3_60_2","doi-asserted-by":"publisher","unstructured":"Xiaoou Liu Tiejin Chen Longchao Da Chacha Chen Zhen Lin and Hua Wei. 2025. Uncertainty quantification and confidence calibration in large language models: A survey. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD \u201925). Association for Computing Machinery New York NY USA. 6107\u20136117. DOI:10.1145\/3711896.3736569","DOI":"10.1145\/3711896.3736569"},{"key":"e_1_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41591-024-03416-6"},{"key":"e_1_3_3_62_2","unstructured":"Xiao Liu Hao Yu Hanchen Zhang Yifan Xu Xuanyu Lei Hanyu Lai Yu Gu Hangliang Ding Kaiwen Men Kejuan Yang et\u00a0al. 2023. Agentbench: Evaluating LLMs as agents. arXiv:2308.03688. Retrieved from https:\/\/arxiv.org\/abs\/2308.03688"},{"key":"e_1_3_3_63_2","unstructured":"Zhihan Liu Hao Hu Shenao Zhang Hongyi Guo Shuqi Ke Boyi Liu and Zhaoran Wang. 2024. Reason for future act for now: A principled framework for autonomous LLM agents. In Proceedings of the 41st International Conference on Machine Learning (ICML\u201924). JMLR.org. 235 Article 1261 (2024) 31186\u201331261."},{"key":"e_1_3_3_64_2","volume-title":"Proceedings of the 1st Conference on Language Modeling","author":"Liu Zijun","year":"2024","unstructured":"Zijun Liu, Yanzhe Zhang, Peng Li, Yang Liu, and Diyi Yang. 2024. A dynamic LLM-powered agent network for task-oriented agent collaboration. In Proceedings of the 1st Conference on Language Modeling."},{"key":"e_1_3_3_65_2","unstructured":"J. Luo W. Zhang Y. Yuan et\u00a0al. 2025. Large language model agent: A survey on methodology applications and challenges. arXiv:2503.21460. Retrieved from https:\/\/arxiv.org\/abs\/2503.21460"},{"key":"e_1_3_3_66_2","doi-asserted-by":"crossref","unstructured":"Xinbei Ma Yiting Wang Yao Yao Tongxin Yuan Aston Zhang Zhuosheng Zhang and Hai Zhao. 2025. Caution for the environment: Multimodal agents are susceptible to environmental distractions. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics Vienna Austria. 1 (2025) 22324\u201322339.","DOI":"10.18653\/v1\/2025.acl-long.1087"},{"key":"e_1_3_3_67_2","first-page":"46534","article-title":"Self-refine: Iterative refinement with self-feedback","volume":"36","author":"Madaan Aman","year":"2023","unstructured":"Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et\u00a0al. 2023. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems 36 (2023), 46534\u201346594.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_68_2","unstructured":"Dayou Mao Yuhao Chen Yifan Wu Maximilian Gilles and Alexander Wong. 2024. Robust analysis of multi-task learning efficiency: New benchmarks on light-weighed backbones and effective measurement of multi-task learning challenges by feature disentanglement. arXiv:2402.03557. Retrieved from https:\/\/arxiv.org\/abs\/2402.03557"},{"key":"e_1_3_3_69_2","unstructured":"Tula Masterman Sandi Besen Mason Sawtell and Alex Chao. 2024. The landscape of emerging AI agent architectures for reasoning planning and tool calling: A survey. arXiv:2404.11584. Retrieved from https:\/\/arxiv.org\/abs\/2404.11584"},{"key":"e_1_3_3_70_2","unstructured":"Ning Miao Yee Whye Teh and Tom Rainforth. 2023. Selfcheck: Using LLMs to zero-shot check their own step-by-step reasoning. arXiv:2308.00436. Retrieved from https:\/\/arxiv.org\/abs\/2308.00436"},{"key":"e_1_3_3_71_2","doi-asserted-by":"crossref","unstructured":"Vincent Micheli and Fran\u00e7ois Fleuret. 2021. Language models are few-shot butlers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics. Online and Punta Cana Dominican Republic. 9312\u20139318.","DOI":"10.18653\/v1\/2021.emnlp-main.734"},{"key":"e_1_3_3_72_2","unstructured":"Dany Moshkovich Hadar Mulian Sergey Zeltyn Natti Eder Inna Skarbovsky and Roy Abitbol. 2025. Beyond black-box benchmarking: Observability analytics and optimization of agentic systems. arXiv:2503.06745. Retrieved from https:\/\/arxiv.org\/abs\/2503.06745"},{"key":"e_1_3_3_73_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639187"},{"key":"e_1_3_3_74_2","unstructured":"Minh Nguyen and Ehsan Shareghi. 2024. One STEP at a time: Language agents are stepwise planners. arXiv:2411.08432. Retrieved from https:\/\/arxiv.org\/abs\/2411.08432"},{"key":"e_1_3_3_75_2","first-page":"27730","article-title":"Training language models to follow instructions with human feedback","volume":"35","author":"Ouyang Long","year":"2022","unstructured":"Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et\u00a0al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35 (2022), 27730\u201327744.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_76_2","unstructured":"Aaron Parisi Yao Zhao and Noah Fiedel. 2022. Talm: Tool augmented language models. arXiv:2205.12255. Retrieved from https:\/\/arxiv.org\/abs\/2205.12255"},{"key":"e_1_3_3_77_2","doi-asserted-by":"publisher","DOI":"10.1006\/jevp.1999.0160"},{"key":"e_1_3_3_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00886"},{"key":"e_1_3_3_79_2","unstructured":"Zehan Qi Xiao Liu Iat Long Iong Hanyu Lai Xueqiao Sun Wenyi Zhao Yu Yang Xinyue Yang Jiadai Sun Shuntian Yao et\u00a0al. 2024. WebRL: Training LLM web agents via self-evolving online curriculum reinforcement learning. arXiv:2411.02337. Retrieved from https:\/\/arxiv.org\/abs\/2411.02337"},{"key":"e_1_3_3_80_2","unstructured":"Yujia Qin Shihao Liang Yining Ye Kunlun Zhu Lan Yan Yaxi Lu Yankai Lin Xin Cong Xiangru Tang Bill Qian et\u00a0al. 2023. ToolLLM: Facilitating large language models to master 16000+ real-world APIs. arXiv:2307.16789. Retrieved from https:\/\/arxiv.org\/abs\/2307.16789"},{"key":"e_1_3_3_81_2","first-page":"28492","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2023","unstructured":"Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 28492\u201328518."},{"issue":"140","key":"e_1_3_3_82_2","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21, 140 (2020), 1\u201367.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_3_83_2","unstructured":"Allen Z. Ren Anushri Dixit Alexandra Bodrova Sumeet Singh Stephen Tu Noah Brown Peng Xu Leila Takayama Fei Xia Jake Varley et\u00a0al. 2023. Robots that ask for help: Uncertainty alignment for large language model planners. arXiv:2307.01928. Retrieved from https:\/\/arxiv.org\/abs\/2307.01928"},{"key":"e_1_3_3_84_2","unstructured":"Yangjun Ruan Honghua Dong Andrew Wang Silviu Pitis Yongchao Zhou Jimmy Ba Yann Dubois Chris J. Maddison and Tatsunori Hashimoto. 2023. Identifying the risks of LM agents with an LM-emulated sandbox. arXiv:2309.15817. Retrieved from https:\/\/arxiv.org\/abs\/2309.15817"},{"key":"e_1_3_3_85_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(97)00026-X"},{"key":"e_1_3_3_86_2","doi-asserted-by":"crossref","unstructured":"Tanmana Sadhu Ali Pesaranghader Yanan Chen and Dong Hoon Yi. 2024. Athena: Safe autonomous agents with verbal contrastive learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track. Association for Computational Linguistics. Miami Florida US. 1121\u20131130.","DOI":"10.18653\/v1\/2024.emnlp-industry.84"},{"key":"e_1_3_3_87_2","first-page":"29971","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Santurkar Shibani","year":"2023","unstructured":"Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect?. In Proceedings of the International Conference on Machine Learning. PMLR, 29971\u201330004."},{"key":"e_1_3_3_88_2","first-page":"68539","article-title":"Toolformer: Language models can teach themselves to use tools","volume":"36","author":"Schick Timo","year":"2023","unstructured":"Timo Schick, Jane Dwivedi-Yu, Roberto Dess\u00ec, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems 36 (2023), 68539\u201368551.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_89_2","unstructured":"Jiawen Shi Zenghui Yuan Guiyao Tie Pan Zhou Neil Zhenqiang Gong and Lichao Sun. 2025. Prompt injection attack to tool selection in LLM agents. arXiv:2504.19793. Retrieved from https:\/\/arxiv.org\/abs\/2504.19793"},{"key":"e_1_3_3_90_2","doi-asserted-by":"publisher","DOI":"10.1145\/3696410.3714825"},{"key":"e_1_3_3_91_2","article-title":"Reflexion: Language agents with verbal reinforcement learning","volume":"36","author":"Shinn Noah","year":"2024","unstructured":"Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems 36 (2024), 8634\u20138652.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_92_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01075"},{"key":"e_1_3_3_93_2","unstructured":"Mohit Shridhar Xingdi Yuan Marc-Alexandre C\u00f4t\u00e9 Yonatan Bisk Adam Trischler and Matthew Hausknecht. 2020. Alfworld: Aligning text and embodied environments for interactive learning. arXiv:2010.03768. Retrieved from https:\/\/arxiv.org\/abs\/2010.03768"},{"key":"e_1_3_3_94_2","unstructured":"Ruoyu Song Muslum Ozgur Ozmen Hyungsub Kim Antonio Bianchi and Z. Berkay Celik. 2024. Enhancing LLM-based autonomous driving agents to mitigate perception attacks. arXiv:2409.14488. Retrieved from https:\/\/arxiv.org\/abs\/2409.14488"},{"key":"e_1_3_3_95_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-022-01611-x"},{"issue":"1","key":"e_1_3_3_96_2","first-page":"26","article-title":"Expanding approaches for research: Understanding and using trustworthiness in qualitative research","volume":"44","author":"Stahl Norman A.","year":"2020","unstructured":"Norman A. Stahl and James R. King. 2020. Expanding approaches for research: Understanding and using trustworthiness in qualitative research. Journal of Developmental Education 44, 1 (2020), 26\u201328.","journal-title":"Journal of Developmental Education"},{"key":"e_1_3_3_97_2","first-page":"3008","article-title":"Learning to summarize with human feedback","volume":"33","author":"Stiennon Nisan","year":"2020","unstructured":"Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems 33 (2020), 3008\u20133021.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_98_2","doi-asserted-by":"crossref","unstructured":"Zhe Su Xuhui Zhou Sanketh Rangreji Anubha Kabra Julia Mendelsohn Faeze Brahman and Maarten Sap. 2025. AI-LieDar: Examine the trade-off between utility and truthfulness in LLM agents. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics. Albuquerque New Mexico. 1 (2025) 11867\u201311894.","DOI":"10.18653\/v1\/2025.naacl-long.595"},{"key":"e_1_3_3_99_2","unstructured":"Theodore R. Sumers Shunyu Yao Karthik Narasimhan and Thomas L. Griffiths. 2023. Cognitive architectures for language agents. arXiv:2309.02427. Retrieved from https:\/\/arxiv.org\/abs\/2309.02427"},{"key":"e_1_3_3_100_2","unstructured":"Xingpeng Sun Haoming Meng Souradip Chakraborty Amrit Singh Bedi and Aniket Bera. 2024. Beyond Text: Utilizing vocal cues to improve decision making in LLMs for robot navigation tasks. arXiv:2402.03494. Retrieved from https:\/\/arxiv.org\/abs\/2402.03494"},{"key":"e_1_3_3_101_2","unstructured":"Xingpeng Sun Yiran Zhang Xindi Tang Amrit Singh Bedi and Aniket Bera. 2024. TrustNavGPT: Modeling uncertainty to improve trustworthiness of audio-guided LLM-based robot navigation. arXiv:2408.01867. Retrieved from https:\/\/arxiv.org\/abs\/2408.01867"},{"key":"e_1_3_3_102_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-022-10228-y"},{"key":"e_1_3_3_103_2","unstructured":"Alex Tamkin Kunal Handa Avash Shrestha and Noah Goodman. 2022. Task ambiguity in humans and language models. arXiv:2212.10711. Retrieved from https:\/\/arxiv.org\/abs\/2212.10711"},{"key":"e_1_3_3_104_2","doi-asserted-by":"publisher","unstructured":"Xiangru Tang Qiao Jin Kunlun Zhu Tongxin Yuan Yichi Zhang Wangchunshu Zhou Meng Qu Yilun Zhao Jian Tang Zhuosheng Zhang et\u00a0al. 2024. Prioritizing safeguarding over autonomy: Nature Communications 16 Article 8317 (2025). DOI:10.1038\/s41467-025-63913-1","DOI":"10.1038\/s41467-025-63913-1"},{"key":"e_1_3_3_105_2","unstructured":"Elizaveta Tennant Stephen Hailes and Mirco Musolesi. 2024. Moral alignment for LLM agents. arXiv:2410.01639. Retrieved from https:\/\/arxiv.org\/abs\/2410.01639"},{"key":"e_1_3_3_106_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et\u00a0al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_3_107_2","doi-asserted-by":"crossref","unstructured":"Harsh Trivedi Tushar Khot Mareike Hartmann Ruskin Manku Vinty Dong Edward Li Shashank Gupta Ashish Sabharwal and Niranjan Balasubramanian. 2024. Appworld: A controllable world of apps and people for benchmarking interactive coding agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics Bangkok Thailand. 1 (2024) 16022\u201316076.","DOI":"10.18653\/v1\/2024.acl-long.850"},{"key":"e_1_3_3_108_2","volume-title":"North American Industry Classification System (NAICS)","author":"Bureau U.S. Census","year":"2022","unstructured":"U.S. Census Bureau. 2022. North American Industry Classification System (NAICS). U.S. Department of Commerce."},{"key":"e_1_3_3_109_2","doi-asserted-by":"publisher","DOI":"10.1057\/s41599-025-04850-8"},{"key":"e_1_3_3_110_2","unstructured":"Jiayin Wang Weizhi Ma Peijie Sun Min Zhang and Jian-Yun Nie. 2024. Understanding user experience in large language model interactions. arXiv:2401.08329. Retrieved from https:\/\/arxiv.org\/abs\/2401.08329"},{"key":"e_1_3_3_111_2","unstructured":"Junlin Wang Jue Wang Ben Athiwaratkun Ce Zhang and James Zou. 2024. Mixture-of-agents enhances large language model capabilities. arXiv:2406.04692. Retrieved from https:\/\/arxiv.org\/abs\/2406.04692"},{"key":"e_1_3_3_112_2","volume-title":"Proceedings of the ICML 2024 Workshop on Foundation Models in the Wild","author":"Wang Kuan","year":"2024","unstructured":"Kuan Wang, Yadong Lu, Michael Santacroce, Yeyun Gong, Chao Zhang, et\u00a0al. 2024. Adapting LLM agents with universal feedback in communication. In Proceedings of the ICML 2024 Workshop on Foundation Models in the Wild."},{"key":"e_1_3_3_113_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3367329"},{"key":"e_1_3_3_114_2","first-page":"74530","article-title":"Augmenting language models with long-term memory","volume":"36","author":"Wang Weizhi","year":"2023","unstructured":"Weizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu, Xifeng Yan, Jianfeng Gao, and Furu Wei. 2023. Augmenting language models with long-term memory. Advances in Neural Information Processing Systems 36 (2023), 74530\u201374543.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_115_2","unstructured":"Xingyao Wang Yangyi Chen Lifan Yuan Yizhe Zhang Yunzhu Li Hao Peng and Heng Ji. 2024. Executable code actions elicit better LLM agents. In Proceedings of the 41st International Conference on Machine Learning (ICML\u201924). JMLR.org. 235 Article 2054 (2024) 50208\u201350232."},{"key":"e_1_3_3_116_2","unstructured":"Zhiruo Wang Zhoujun Cheng Hao Zhu Daniel Fried and Graham Neubig. 2024. What are tools anyway? A survey from the language model perspective. arXiv:2403.15452. Retrieved from https:\/\/arxiv.org\/abs\/2403.15452"},{"key":"e_1_3_3_117_2","unstructured":"Ziyan Wang Meng Fang Tristan Tomilin Fei Fang and Yali Du. 2024. Safe multi-agent reinforcement learning with natural language constraints. arXiv:2405.20018. Retrieved from https:\/\/arxiv.org\/abs\/2405.20018"},{"key":"e_1_3_3_118_2","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, Denny Zhou, et\u00a0al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (2022), 24824\u201324837.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_119_2","article-title":"LLM-powered autonomous agents","author":"Weng Lilian","year":"2023","unstructured":"Lilian Weng. 2023. LLM-powered autonomous agents. lilianweng.github.io (Jun22023). Retrieved February 22, 2026 from https:\/\/lilianweng.github.io\/posts\/2023-06-23-agent\/","journal-title":"lilianweng.github.io"},{"key":"e_1_3_3_120_2","doi-asserted-by":"publisher","DOI":"10.1001\/jamanetworkopen.2025.8052"},{"key":"e_1_3_3_121_2","unstructured":"Yue Wu Xuan Tang Tom M. Mitchell and Yuanzhi Li. 2023. Smartplay: A benchmark for LLMs as intelligent agents. arXiv:2310.01557. Retrieved from https:\/\/arxiv.org\/abs\/2310.01557"},{"key":"e_1_3_3_122_2","unstructured":"Zhiheng Xi Wenxiang Chen Xin Guo Wei He Yiwen Ding Boyang Hong Ming Zhang Junzhe Wang Senjie Jin Enyu Zhou et\u00a0al. 2023. The rise and potential of large language model based agents: A survey. arXiv:2309.07864. Retrieved from https:\/\/arxiv.org\/abs\/2309.07864"},{"key":"e_1_3_3_123_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-024-4222-0"},{"key":"e_1_3_3_124_2","unstructured":"Zhen Xiang Linzhi Zheng Yanjie Li Junyuan Hong Qinbin Li Han Xie Jiawei Zhang Zidi Xiong Chulin Xie Carl Yang et\u00a0al. 2024. GuardAgent: Safeguard LLM agents by a guard agent via knowledge-enabled reasoning. arXiv:2406.09187. Retrieved from https:\/\/arxiv.org\/abs\/2406.09187"},{"key":"e_1_3_3_125_2","doi-asserted-by":"crossref","unstructured":"Chengxing Xie Canyu Chen Feiran Jia Ziyu Ye Shiyang Lai Kai Shu Jindong Gu Adel Bibi Ziniu Hu David Jurgens James Evans Philip H. S. Torr Bernard Ghanem and Guohao Li.2024. Can large language model agents simulate human trust behaviors? In Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS \u201924). Curran Associates Inc. Red Hook NY USA. 37 Article 501 (2024) 15674\u201315729.","DOI":"10.52202\/079017-0501"},{"key":"e_1_3_3_126_2","doi-asserted-by":"crossref","unstructured":"Roy Xie Junlin Wang Ruomin Huang Minxing Zhang Rong Ge Jian Pei Neil Zhenqiang Gong and Bhuwan Dhingra. 2024. Recall: Membership inference via relative conditional log-likelihoods. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics Miami Florida USA. 8671\u20138689.","DOI":"10.18653\/v1\/2024.emnlp-main.493"},{"key":"e_1_3_3_127_2","first-page":"52040","article-title":"Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments","volume":"37","author":"Xie Tianbao","year":"2024","unstructured":"Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh J. Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et\u00a0al. 2024. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems 37 (2024), 52040\u201352094.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_128_2","doi-asserted-by":"crossref","unstructured":"Weimin Xiong Yifan Song Xiutian Zhao Wenhao Wu Xun Wang Ke Wang Cheng Li Wei Peng and Sujian Li. 2024. Watch every step! LLM agent learning via iterative step-level process refinement. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics. Miami Florida USA. 1556\u20131572.","DOI":"10.18653\/v1\/2024.emnlp-main.93"},{"key":"e_1_3_3_129_2","unstructured":"Ke Yang Yao Liu Sapana Chaudhary Rasool Fakoor Pratik Chaudhari George Karypis and Huzefa Rangwala. 2024. AgentOccam: A simple yet strong baseline for LLM-based web agents. arXiv:2410.13825. Retrieved from https:\/\/arxiv.org\/abs\/2410.13825"},{"key":"e_1_3_3_130_2","unstructured":"Zonghan Yang An Liu Zijun Liu Kaiming Liu Fangzhou Xiong Yile Wang Zeyuan Yang Qingyuan Hu Xinrui Chen Zhenhe Zhang Fuwen Luo Zhicheng Guo Peng Li and Yang Liu.2024. Position: Towards unified alignment between agents humans and environment. In Proceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research). PMLR. 235 (2024) 56251\u201356275."},{"key":"e_1_3_3_131_2","first-page":"20744","article-title":"Webshop: Towards scalable real-world web interaction with grounded language agents","volume":"35","author":"Yao Shunyu","year":"2022","unstructured":"Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems 35 (2022), 20744\u201320757.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_132_2","unstructured":"Shunyu Yao Jeffrey Zhao Dian Yu Nan Du Izhak Shafran Karthik Narasimhan and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. arXiv:2210.03629. Retrieved from https:\/\/arxiv.org\/abs\/2210.03629"},{"key":"e_1_3_3_133_2","doi-asserted-by":"crossref","unstructured":"Aohan Zeng Mingdao Liu Rui Lu Bowen Wang Xiao Liu Yuxiao Dong and Jie Tang. 2024. AgentTuning: Enabling generalized agent abilities for LLMs. In Findings of the Association for Computational Linguistics: ACL 2024. Association for Computational Linguistics Bangkok Thailand. 3053\u20133077.","DOI":"10.18653\/v1\/2024.findings-acl.181"},{"key":"e_1_3_3_134_2","doi-asserted-by":"crossref","unstructured":"Boyang Zhang Yicong Tan Yun Shen Ahmed Salem Michael Backes Savvas Zannettou and Yang Zhang. 2025. Breaking agents: Compromising autonomous LLM agents through malfunction amplification. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics Suzhou China. 34964\u201334976.","DOI":"10.18653\/v1\/2025.emnlp-main.1771"},{"key":"e_1_3_3_135_2","unstructured":"Hanrong Zhang Jingyuan Huang Kai Mei Yifei Yao Zhenting Wang Chenlu Zhan Hongwei Wang and Yongfeng Zhang. 2024. Agent security bench (ASB): Formalizing and benchmarking attacks and defenses in LLM-based agents. arXiv:2410.02644. Retrieved from https:\/\/arxiv.org\/abs\/2410.02644"},{"key":"e_1_3_3_136_2","unstructured":"Hanchen Zhang Xiao Liu Bowen Lv Xueqiao Sun Bohao Jing Iat Long Iong Zhenyu Hou Zehan Qi Hanyu Lai Yifan Xu et\u00a0al. 2025. AgentRL: Scaling agentic reinforcement learning with a multi-turn multi-task framework. arXiv:2510.04206. Retrieved from https:\/\/arxiv.org\/abs\/2510.04206"},{"key":"e_1_3_3_137_2","doi-asserted-by":"crossref","unstructured":"Michael JQ Zhang and Eunsol Choi. 2025. Clarify when necessary: Resolving ambiguity through interaction with LMS. In Findings of the Association for Computational Linguistics: NAACL 2025. Association for Computational Linguistics Albuquerque New Mexico. 5541\u20135558.","DOI":"10.18653\/v1\/2025.findings-naacl.306"},{"key":"e_1_3_3_138_2","unstructured":"Michael J. Q. Zhang W. Bradley Knox and Eunsol Choi. 2024. Modeling future conversation turns to teach LLMs to ask clarifying questions. arXiv:2410.13788. Retrieved from https:\/\/arxiv.org\/abs\/2410.13788"},{"key":"e_1_3_3_139_2","doi-asserted-by":"publisher","unstructured":"Shengyu Zhang Linfeng Dong Xiaoya Li Sen Zhang Xiaofei Sun Shuhe Wang Jiwei Li Runyi Hu Tianwei Zhang Guoyin Wang and Fei Wu. 2026. Instruction tuning for large language models: A survey. ACM Comput. Surv. 58 7 Article 169 (May 2026) 36. DOI:10.1145\/3777411","DOI":"10.1145\/3777411"},{"key":"e_1_3_3_140_2","doi-asserted-by":"crossref","unstructured":"Xiaopan Zhang Hao Qin Fuquan Wang Yue Dong and Jiachen Li. 2024. LaMMA-P: Generalizable multi-agent long-horizon task allocation and planning with LM-driven PDDL planner. arXiv:2409.20560. Retrieved from https:\/\/arxiv.org\/abs\/2409.20560","DOI":"10.1109\/ICRA55743.2025.11127951"},{"key":"e_1_3_3_141_2","unstructured":"Boyuan Zheng Boyu Gou Jihyung Kil Huan Sun and Yu Su. 2024. Gpt-4v (ision) is a generalist web agent if grounded. In Proceedings of the 41st International Conference on Machine Learning (ICML\u201924). JMLR.org. 235 Article 2538 (2024) 61349\u201361385."},{"key":"e_1_3_3_142_2","unstructured":"Zhi Zheng Qian Feng Hang Li Alois Knoll and Jianxiang Feng. 2024. Evaluating uncertainty-based failure detection for closed-loop LLM planners. arXiv:2406.00430. Retrieved from https:\/\/arxiv.org\/abs\/2406.00430"},{"key":"e_1_3_3_143_2","unstructured":"Qinhong Zhou Sunli Chen Yisong Wang Haozhe Xu Weihua Du Hongxin Zhang Yilun Du Joshua B. Tenenbaum and Chuang Gan. 2024. HAZARD challenge: Embodied decision making in dynamically changing environments. arXiv:2401.12975. Retrieved from https:\/\/arxiv.org\/abs\/2401.12975"},{"key":"e_1_3_3_144_2","unstructured":"Shuyan Zhou Frank F. Xu Hao Zhu Xuhui Zhou Robert Lo Abishek Sridhar Xianyi Cheng Tianyue Ou Yonatan Bisk Daniel Fried et\u00a0al. 2023. Webarena: A realistic web environment for building autonomous agents. arXiv:2307.13854. Retrieved from https:\/\/arxiv.org\/abs\/2307.13854"},{"key":"e_1_3_3_145_2","doi-asserted-by":"publisher","DOI":"10.1145\/3589335.3651955"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3794858","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3794858","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T20:43:13Z","timestamp":1775853793000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3794858"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,2]]},"references-count":144,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3794858"],"URL":"https:\/\/doi.org\/10.1145\/3794858","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,2]]},"assertion":[{"value":"2025-09-24","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-20","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-02","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}