{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T13:43:11Z","timestamp":1785332591524,"version":"3.55.0"},"reference-count":147,"publisher":"Association for Computing Machinery (ACM)","issue":"13","license":[{"start":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T00:00:00Z","timestamp":1782259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,10,31]]},"abstract":"<jats:p>The integration of Agentic AI into healthcare marks a new era in medical automation and decision support. Leveraging autonomous reasoning, planning, and collaboration, these AI systems are transforming clinical workflows, diagnostics, and patient management. Recent advancements have led Agentic AI to evolve from single-agent decision-making to multi-agent systems, enabling more sophisticated problem-solving in complex medical environments. In this survey, we provide an in-depth discussion on the core aspects and challenges of Agentic AI in healthcare, including its applications in electronic health records (EHR) interactions, clinical triage, medical question-answering, and disease diagnosis. We categorize existing single-agent and multi-agent frameworks, explore key evaluation metrics, implementation strategies, and commonly used benchmarks, and examine the communication, reasoning, and decision-making processes of these AI agents. Additionally, we address critical challenges such as model reliability, interoperability, ethical considerations, and regulatory constraints. To support further research, we summarize relevant datasets and benchmarks and we outline future research directions that emphasize human-AI collaboration, transparency, and the safe deployment of AI-driven medical systems.<\/jats:p>","DOI":"10.1145\/3809164","type":"journal-article","created":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T09:05:58Z","timestamp":1779181558000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Agentic AI in Healthcare: Opportunities, Challenges, and Future Directions"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0770-679X","authenticated-orcid":false,"given":"Mourad","family":"Gridach","sequence":"first","affiliation":[{"name":"Real World Solutions, IQVIA","place":["Cambridge, United Kingdom of Great Britain and Northern Ireland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-0122-8357","authenticated-orcid":false,"given":"Jay","family":"Nanavati","sequence":"additional","affiliation":[{"name":"Real World Solutions, IQVIA","place":["Cambridge, United Kingdom of Great Britain and Northern Ireland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-7301-4034","authenticated-orcid":false,"given":"Khaldoun","family":"Zine El Abidine","sequence":"additional","affiliation":[{"name":"Real World Solutions, IQVIA","place":["Ottawa, Canada"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6203-3954","authenticated-orcid":false,"given":"Calum","family":"Yacoubian","sequence":"additional","affiliation":[{"name":"Real World Solutions, IQVIA","place":["Cambridge, United Kingdom of Great Britain and Northern Ireland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3495-3796","authenticated-orcid":false,"given":"Christina","family":"Mack","sequence":"additional","affiliation":[{"name":"Real World Solutions, IQVIA","place":["Durham, United States of America"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41746-024-01074-z"},{"key":"e_1_3_2_3_2","doi-asserted-by":"crossref","unstructured":"Gavin Abercrombie Amanda Cercas Curry Tanvi Dinkar Verena Rieser and Zeerak Talat. 2023. Mirages: On anthropomorphism in dialogue systems. arXiv:2305.09800. Retrieved from https:\/\/arxiv.org\/abs\/2305.09800 (2023).","DOI":"10.18653\/v1\/2023.emnlp-main.290"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/BIBE60311.2023.00071"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383652.3423900"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.3390\/su15086655"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.envsoft.2023.105713"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.7466396"},{"key":"e_1_3_2_9_2","first-page":"12449","article-title":"wav2vec 2.0: A framework for self-supervised learning of speech representations","volume":"33","author":"Baevski Alexei","year":"2020","unstructured":"Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems 33 (2020), 12449\u201312460.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_10_2","unstructured":"Zhijie Bao Qingyun Liu Ying Guo Zhengqiang Ye Jun Shen Shirong Xie Jiajie Peng Xuanjing Huang and Zhongyu Wei. 2024. Piors: Personalized intelligent outpatient reception based on large language model with multi-agents medical scenario simulation. arXiv:2411.13902. Retrieved from https:\/\/arxiv.org\/abs\/2411.13902 (2024)."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1038\/s44387-025-00003-z"},{"key":"e_1_3_2_12_2","unstructured":"Mert Cemri Melissa Z. Pan Shuyi Yang Lakshya A. Agrawal Bhavya Chopra Rishabh Tiwari Kurt Keutzer Aditya Parameswaran Dan Klein Kannan Ramchandran et\u00a0al. 2025. Why do multi-agent llm systems fail? arXiv:2503.13657. Retrieved from https:\/\/arxiv.org\/abs\/2503.13657 (2025)."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3658948"},{"key":"e_1_3_2_14_2","unstructured":"Chi-Min Chan Jianxuan Yu Weize Chen Chunyang Jiang Xinyu Liu Weijie Shi Zhiyuan Liu Wei Xue and Yike Guo. 2024. Agentmonitor: A plug-and-play framework for predictive and secure multi-agent systems. arXiv:2408.14972. Retrieved from https:\/\/arxiv.org\/abs\/2408.14972 (2024)."},{"key":"e_1_3_2_15_2","doi-asserted-by":"crossref","unstructured":"Chaoran Chen Bingsheng Yao Ruishi Zou Wenyue Hua Weimin Lyu Toby Jia-Jun Li and Dakuo Wang. 2025. Towards a design guideline for RPA evaluation: A survey of large language model-based role-playing agents. arXiv:2502.13012. Retrieved from https:\/\/arxiv.org\/abs\/2502.13012 (2025).","DOI":"10.18653\/v1\/2025.findings-acl.938"},{"key":"e_1_3_2_16_2","unstructured":"Hanjie Chen Zhouxiang Fang Yash Singla and Mark Dredze. 2024. Benchmarking large language models on answering and explaining challenging medical questions. arXiv:2402.18060. Retrieved from https:\/\/arxiv.org\/abs\/2402.18060 (2024)."},{"key":"e_1_3_2_17_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Chen Weize","year":"2024","unstructured":"Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, et\u00a0al. 2024. AgentVerse: Facilitating multi-agent collaboration and exploring emergent behaviors. In Proceedings of the 12th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=EHg5GDnyq1"},{"key":"e_1_3_2_18_2","unstructured":"Xuanzhong Chen Ye Jin Xiaohao Mao Lun Wang Shuyang Zhang and Ting Chen. 2024. RareAgents: Autonomous multi-disciplinary team for rare disease diagnosis and treatment. arXiv:2412.12475. Retrieved from https:\/\/arxiv.org\/abs\/2412.12475 (2024)."},{"key":"e_1_3_2_19_2","unstructured":"Yuheng Cheng Ceyao Zhang Zhengwen Zhang Xiangrui Meng Sirui Hong Wenhao Li Zihao Wang Zekai Wang Feng Yin Junhua Zhao et\u00a0al. 2024. Exploring large language model based intelligent agents: Definitions methods and prospects. arXiv:2401.03428. Retrieved from https:\/\/arxiv.org\/abs\/2401.03428 (2024)."},{"key":"e_1_3_2_20_2","article-title":"Mind2web: Towards a generalist agent for the web","volume":"36","author":"Deng Xiang","year":"2024","unstructured":"Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2024. Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems 36 (2024).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_21_2","unstructured":"Han Ding Yinheng Li Junhao Wang and Hang Chen. 2024. Large language model agent in financial trading: A survey. arXiv:2408.06361. Retrieved from https:\/\/arxiv.org\/abs\/2408.06361 (2024)."},{"key":"e_1_3_2_22_2","unstructured":"Zane Durante Qiuyuan Huang Naoki Wake Ran Gong Jae Sung Park Bidipta Sarkar Rohan Taori Yusuke Noda Demetri Terzopoulos Yejin Choi et\u00a0al. 2024. Agent AI: Surveying the horizons of multimodal interaction. arXiv:2401.03568. Retrieved from https:\/\/arxiv.org\/abs\/2401.03568 (2024)."},{"key":"e_1_3_2_23_2","volume-title":"Proceedings of the 31st International Conference on Computational Linguistics","author":"Fan Zhihao","year":"2025","unstructured":"Zhihao Fan, Lai Wei, Jialong Tang, Wei Chen, Wang Siyuan, Zhongyu Wei, and Fei Huang. 2025. AI hospital: Benchmarking large language models in a multi-agent medical interaction simulator. In Proceedings of the 31st International Conference on Computational Linguistics, Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert (Eds.)."},{"key":"e_1_3_2_24_2","first-page":"1605","volume-title":"Proceedings of the 29th USENIX Security Symposium (USENIX Security 20)","author":"Fang Minghong","year":"2020","unstructured":"Minghong Fang, Xiaoyu Cao, Jinyuan Jia, and Neil Gong. 2020. Local model poisoning attacks to \\(\\lbrace\\) Byzantine-Robust \\(\\rbrace\\) federated learning. In Proceedings of the 29th USENIX Security Symposium (USENIX Security 20). 1605\u20131622."},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-2270"},{"key":"e_1_3_2_26_2","unstructured":"K. J. Feng David W. McDonald and Amy X. Zhang. 2025. Levels of autonomy for AI agents. arXiv:2506.12469. Retrieved from https:\/\/arxiv.org\/abs\/2506.12469 (2025)."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.76"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/191246.191322"},{"key":"e_1_3_2_29_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Furuta Hiroki","year":"2024","unstructured":"Hiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo, Aleksandra Faust, Shixiang Shane Gu, and Izzeddin Gur. 2024. Multimodal web navigation with instruction-finetuned foundation models. In Proceedings of the 12th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=efFmBWioSc"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1039\/D4DD00013G"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1002\/adma.202413523"},{"key":"e_1_3_2_32_2","doi-asserted-by":"crossref","unstructured":"Fatemeh Ghezloo Mehmet Saygin Seyfioglu Rustin Soraki Wisdom O. Ikezogwo Beibin Li Tejoram Vivekanandan Joann G. Elmore Ranjay Krishna and Linda Shapiro. 2025. PathFinder: A multi-modal multi-agent system for medical diagnostic decision-making applied to histopathology. arXiv:2502.08916. Retrieved from https:\/\/arxiv.org\/abs\/2502.08916 (2025).","DOI":"10.1109\/ICCV51701.2025.02175"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1101\/2024.04.07.24305462"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2024\/890"},{"key":"e_1_3_2_35_2","unstructured":"Taicheng Guo Xiuying Chen Yaqi Wang Ruidi Chang Shichao Pei Nitesh V. Chawla Olaf Wiest and Xiangliang Zhang. 2024. Large language model based multi-agents: A survey of progress and challenges. arXiv:2402.01680. Retrieved from https:\/\/arxiv.org\/abs\/2402.01680 (2024)."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41597-023-02036-y"},{"key":"e_1_3_2_37_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Gur Izzeddin","year":"2024","unstructured":"Izzeddin Gur, Hiroki Furuta, Austin V Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust. 2024. A real-world webagent with planning, long context understanding, and program synthesis. In Proceedings of the 12th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=9JQtrumvg8"},{"key":"e_1_3_2_38_2","unstructured":"Lewis Hammond Alan Chan Jesse Clifton Jason Hoelscher-Obermaier Akbir Khan Euan McLean Chandler Smith Wolfram Barfuss Jakob Foerster Tom\u2019a\u0161 Gaven\u010diak et\u00a0al. 2025. Multi-agent risks from advanced ai. arXiv:2502.14143. Retrieved from https:\/\/arxiv.org\/abs\/2502.14143 (2025)."},{"key":"e_1_3_2_39_2","unstructured":"Shanshan Han Qifan Zhang Yuhang Yao Weizhao Jin and Zhaozhuo Xu. 2024. LLM multi-agent systems: Challenges and open problems. arXiv:2402.03578. Retrieved from https:\/\/arxiv.org\/abs\/2402.03578 (2024)."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.4324\/9780429493744-5"},{"key":"e_1_3_2_41_2","unstructured":"Xuehai He Yichen Zhang Luntian Mou Eric Xing and Pengtao Xie. 2020. Pathvqa: 30000+ questions for medical visual question answering. arXiv:2003.10286. Retrieved from https:\/\/arxiv.org\/abs\/2003.10286 (2020)."},{"key":"e_1_3_2_42_2","unstructured":"Yifeng He Ethan Wang Yuyang Rong Zifei Cheng and Hao Chen. 2024. Security of ai agents. arXiv:2406.08689. Retrieved from https:\/\/arxiv.org\/abs\/2406.08689 (2024)."},{"key":"e_1_3_2_43_2","unstructured":"Axel H\u00f8jmark Govind Pimpale Arjun Panickssery Marius Hobbhahn and J\u00e9r\u00e9my Scheurer. 2024. Analyzing probabilistic methods for evaluating agent capabilities. arXiv:2409.16125. Retrieved from https:\/\/arxiv.org\/abs\/2409.16125 (2024)."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/BIBM62325.2024.10822109"},{"key":"e_1_3_2_45_2","unstructured":"Sirui Hong Xiawu Zheng Jonathan Chen Yuheng Cheng Jinlin Wang Ceyao Zhang Zili Wang Steven Ka Shing Yau Zijuan Lin Liyang Zhou et\u00a0al. 2023. Metagpt: Meta programming for multi-agent collaborative framework. In The Twelfth International Conference on Learning Representations 3 4 (2023) 6."},{"key":"e_1_3_2_46_2","unstructured":"Sihao Hu Tiansheng Huang Fatih Ilhan Selim Tekin Gaowen Liu Ramana Kompella and Ling Liu. 2024. A survey on large language model-based game agents. arXiv:2404.02039. Retrieved from https:\/\/arxiv.org\/abs\/2404.02039 (2024)."},{"key":"e_1_3_2_47_2","unstructured":"Shengran Hu Cong Lu and Jeff Clune. 2024. Automated design of agentic systems. arXiv:2408.08435. Retrieved from https:\/\/arxiv.org\/abs\/2408.08435 (2024)."},{"key":"e_1_3_2_48_2","unstructured":"Qiuyuan Huang Naoki Wake Bidipta Sarkar Zane Durante Ran Gong Rohan Taori Yusuke Noda Demetri Terzopoulos Noboru Kuno Ade Famoti et\u00a0al. 2024. Position paper: Agent AI towards a holistic intelligence. arXiv:2403.00833. Retrieved from https:\/\/arxiv.org\/abs\/2403.00833 (2024)."},{"key":"e_1_3_2_49_2","unstructured":"Carlos E Jimenez John Yang Alexander Wettig Shunyu Yao Kexin Pei Ofir Press and Karthik Narasimhan. 2023. Swe-bench: Can language models resolve real-world github issues? arXiv:2310.06770. Retrieved from https:\/\/arxiv.org\/abs\/2310.06770 (2023)."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.3390\/app11146421"},{"key":"e_1_3_2_51_2","doi-asserted-by":"crossref","unstructured":"Qiao Jin Bhuwan Dhingra Zhengping Liu William W. Cohen and Xinghua Lu. 2019. Pubmedqa: A dataset for biomedical research question answering. arXiv:1909.06146. Retrieved from https:\/\/arxiv.org\/abs\/1909.06146 (2019).","DOI":"10.18653\/v1\/D19-1259"},{"key":"e_1_3_2_52_2","doi-asserted-by":"crossref","unstructured":"Alistair E. W. Johnson Tom J. Pollard Nathaniel R. Greenbaum Matthew P. Lungren Chih-ying Deng Yifan Peng Zhiyong Lu Roger G. Mark Seth J. Berkowitz and Steven Horng. 2019. MIMIC-CXR-JPG a large publicly available database of labeled chest radiographs. arXiv:1901.07042. Retrieved from https:\/\/arxiv.org\/abs\/1901.07042 (2019).","DOI":"10.1038\/s41597-019-0322-0"},{"key":"e_1_3_2_53_2","unstructured":"Sayash Kapoor Benedikt Stroebl Zachary S Siegel Nitya Nadgir and Arvind Narayanan. 2024. Ai agents that matter. arXiv:2407.01502. Retrieved from https:\/\/arxiv.org\/abs\/2407.01502 (2024)."},{"key":"e_1_3_2_54_2","unstructured":"Raihan Khan Sayak Sarkar Sainik Kumar Mahata and Edwin Jose. 2024. Security threats in agentic AI system. arXiv:2410.14728. Retrieved from https:\/\/arxiv.org\/abs\/2410.14728 (2024)."},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2522"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1007\/s13735-024-00334-8"},{"key":"e_1_3_2_57_2","unstructured":"Thomas Kwa Ben West Joel Becker Amy Deng Katharyn Garcia Max Hasin Sami Jawhar Megan Kinniment Nate Rush Sydney Von Arx et\u00a0al. 2025. Measuring ai ability to complete long tasks. arXiv:2503.14499. Retrieved from https:\/\/arxiv.org\/abs\/2503.14499 (2025)."},{"key":"e_1_3_2_58_2","unstructured":"Donghyun Lee and Mo Tiwari. 2024. Prompt infection: Llm-to-llm prompt injection within multi-agent systems. arXiv:2410.07283. Retrieved from https:\/\/arxiv.org\/abs\/2410.07283 (2024)."},{"key":"e_1_3_2_59_2","unstructured":"Binbin Li Tianxin Meng Xiaoming Shi Jie Zhai and Tong Ruan. 2023. MedDM: LLM-executable clinical guidance tree for clinical decision-making. arXiv:2312.02441. Retrieved from https:\/\/arxiv.org\/abs\/2312.02441 (2023)."},{"key":"e_1_3_2_60_2","article-title":"Agent hospital: A simulacrum of hospital with evolvable medical agents","author":"Li Junkai","year":"2024","unstructured":"Junkai Li, Yunghwei Lai, Weitao Li, Jingyi Ren, Meng Zhang, Xinhui Kang, Siyu Wang, Peng Li, Ya-Qin Zhang, Weizhi Ma, et\u00a0al. 2024. Agent hospital: A simulacrum of hospital with evolvable medical agents. arXiv:2405.02957. Retrieved from https:\/\/arxiv.org\/abs\/2405.02957 (2024).","journal-title":"a"},{"key":"e_1_3_2_61_2","unstructured":"Yubo Li Xiaobin Shen Xinyu Yao Xueying Ding Yidi Miao Ramayya Krishnan and Rema Padman. 2025. Beyond single-turn: A survey on multi-turn interactions with large language models. arXiv:2504.04717. Retrieved from https:\/\/arxiv.org\/abs\/2504.04717 (2025)."},{"key":"e_1_3_2_62_2","unstructured":"Yusheng Liao Shuyang Jiang Yanfeng Wang and Yu Wang. 2024. ReflecTool: Towards reflection-aware tool-augmented clinical agents. arXiv:2410.17657. Retrieved from https:\/\/arxiv.org\/abs\/2410.17657 (2024)."},{"key":"e_1_3_2_63_2","unstructured":"Bill Yuchen Lin Yuntian Deng Khyathi Chandu Faeze Brahman Abhilasha Ravichander Valentina Pyatkin Nouha Dziri Ronan Le Bras and Yejin Choi. 2024. Wildbench: Benchmarking llms with challenging tasks from real users in the wild. arXiv:2406.04770. Retrieved from https:\/\/arxiv.org\/abs\/2406.04770 (2024)."},{"key":"e_1_3_2_64_2","unstructured":"Bo Liu Yuqian Jiang Xiaohan Zhang Qiang Liu Shiqi Zhang Joydeep Biswas and Peter Stone. 2023. Llm+ p: Empowering large language models with optimal planning proficiency. arXiv:2304.11477. Retrieved from https:\/\/arxiv.org\/abs\/2304.11477 (2023)."},{"key":"e_1_3_2_65_2","unstructured":"Jie Liu Wenxuan Wang Zizhan Ma Guolin Huang Yihang SU Kao-Jung Chang Wenting Chen Haoliang Li Linlin Shen and Michael Lyu. 2024. Medchain: Bridging the gap between LLM agents and clinical practice through interactive sequential benchmarking. arXiv:2412.01605. Retrieved from https:\/\/arxiv.org\/abs\/2412.01605 (2024)."},{"key":"e_1_3_2_66_2","unstructured":"Xiao Liu Hao Yu Hanchen Zhang Yifan Xu Xuanyu Lei Hanyu Lai Yu Gu Hangliang Ding Kaiwen Men Kejuan Yang et\u00a0al. 2023. Agentbench: Evaluating llms as agents. arXiv:2308.03688. Retrieved from https:\/\/arxiv.org\/abs\/2308.03688 (2023)."},{"key":"e_1_3_2_67_2","unstructured":"Xiaogeng Liu Zhiyuan Yu Yizhe Zhang Ning Zhang and Chaowei Xiao. [2024]. Automatic and universal prompt injection attacks against large language models. arXiv:2403.04957. Retrieved from https:\/\/arxiv.org\/abs\/2403.04957 (2024)."},{"key":"e_1_3_2_68_2","unstructured":"Yi Liu Gelei Deng Yuekang Li Kailong Wang Zihao Wang Xiaofeng Wang Tianwei Zhang Yepang Liu Haoyu Wang Yan Zheng et\u00a0al. 2023. Prompt injection attack against llm-integrated applications. arXiv:2306.05499. Retrieved from https:\/\/arxiv.org\/abs\/2306.05499 (2023)."},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.5555\/3698900.3699003"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-emnlp.329"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA57147.2024.10610855"},{"key":"e_1_3_2_72_2","unstructured":"Tula Masterman Sandi Besen Mason Sawtell and Alex Chao. 2024. The landscape of emerging ai agent architectures for reasoning planning and tool calling: A survey. arXiv:2404.11584. Retrieved from https:\/\/arxiv.org\/abs\/2404.11584 (2024)."},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.1143"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4612-5108-8_8"},{"key":"e_1_3_2_75_2","volume-title":"Society of mind","author":"Minsky Marvin","year":"1988","unstructured":"Marvin Minsky. 1988. Society of mind. Simon and Schuster."},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2019.03.054"},{"key":"e_1_3_2_77_2","unstructured":"Reiichiro Nakano Jacob Hilton Suchir Balaji Jeff Wu Long Ouyang Christina Kim Christopher Hesse Shantanu Jain Vineet Kosaraju William Saunders et\u00a0al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv:2112.09332. Retrieved from https:\/\/arxiv.org\/abs\/2112.09332 (2021)."},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.5555\/13437.13438"},{"key":"e_1_3_2_79_2","volume-title":"Proceedings of the Workshop on Reincarnating Reinforcement Learning at ICLR 2023","author":"Nottingham Kolby","year":"2023","unstructured":"Kolby Nottingham, Prithviraj Ammanabrolu, Alane Suhr, Yejin Choi, Hannaneh Hajishirzi, Sameer Singh, and Roy Fox. 2023. Do embodied agents dream of pixelated sheep?: Embodied decision making using language guided world modelling. In Proceedings of the Workshop on Reincarnating Reinforcement Learning at ICLR 2023. Retrieved from https:\/\/openreview.net\/forum?id=Z_qiOvqvnBl"},{"key":"e_1_3_2_80_2","unstructured":"Maxime Oquab Timoth\u00e9e Darcet Th\u00e9o Moutakanni Huy Vo Marc Szafraniec Vasil Khalidov Pierre Fernandez Daniel Haziza Francisco Massa Alaaeldin El-Nouby et\u00a0al. 2023. Dinov2: Learning robust visual features without supervision. arXiv:2304.07193. Retrieved from https:\/\/arxiv.org\/abs\/2304.07193 (2023)."},{"key":"e_1_3_2_81_2","unstructured":"Charles Packer Vivian Fang Shishir_G Patil Kevin Lin Sarah Wooders and Joseph_E Gonzalez. 2023. MemGPT: Towards LLMs as operating systems. (2023)."},{"key":"e_1_3_2_82_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.conll-1.21"},{"key":"e_1_3_2_83_2","doi-asserted-by":"publisher","DOI":"10.1109\/MNET.2024.3442880"},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.89"},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","DOI":"10.1145\/3586183.3606763"},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.1145\/3526113.3545616"},{"key":"e_1_3_2_87_2","unstructured":"M. Phuong M. Aitchison E. Catt S. Cogan A. Kaskasoli V. Krakovna D. Lindner M. Rahtz Y. Assael S. Hodkinson et\u00a0al. 2024. Evaluating frontier models for dangerous capabilities. arXiv:2403.13793. Retrieved from https:\/\/arxiv.org\/abs\/2403.13793 (2024)."},{"key":"e_1_3_2_88_2","first-page":"111715","article-title":"Cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents","volume":"37","author":"Piatti Giorgio","year":"2025","unstructured":"Giorgio Piatti, Zhijing Jin, Max Kleiman-Weiner, Bernhard Sch\u00f6lkopf, Mrinmaya Sachan, and Rada Mihalcea. 2025. Cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents. Advances in Neural Information Processing Systems 37 (2025), 111715\u2013111759.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_89_2","doi-asserted-by":"crossref","unstructured":"Harry E. Pople. 1984. CADUCEUS: An experimental expert system for medical diagnosis. (1984).","DOI":"10.7551\/mitpress\/1165.003.0007"},{"key":"e_1_3_2_90_2","doi-asserted-by":"crossref","unstructured":"Chen Qian Wei Liu Hongzhang Liu Nuo Chen Yufan Dang Jiahao Li Cheng Yang Weize Chen Yusheng Su Xin Cong Juyuan Xu Dahai Li Zhiyuan Liu and Maosong Sun. 2024. Chatdev: Communicative agents for software development. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 15174\u201315186.","DOI":"10.18653\/v1\/2024.acl-long.810"},{"key":"e_1_3_2_91_2","unstructured":"Chen Qian Wei Liu Hongzhang Liu Nuo Chen Yufan Dang Jiahao Li Cheng Yang Weize Chen Yusheng Su Xin Cong et\u00a0al. 2023. Chatdev: Communicative agents for software development. arXiv:2307.07924. Retrieved from https:\/\/arxiv.org\/abs\/2307.07924 (2023)."},{"key":"e_1_3_2_92_2","unstructured":"Chen Qian Zihao Xie Yifei Wang Wei Liu Yufan Dang Zhuoyun Du Weize Chen Cheng Yang Zhiyuan Liu and Maosong Sun. 2024. Scaling large-language-model-based multi-agent collaboration. arXiv:2406.07155. Retrieved from https:\/\/arxiv.org\/abs\/2406.07155 (2024)."},{"key":"e_1_3_2_93_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et\u00a0al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PmLR, 8748\u20138763."},{"key":"e_1_3_2_94_2","first-page":"28492","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2023","unstructured":"Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 28492\u201328518."},{"issue":"10","key":"e_1_3_2_95_2","first-page":"19","article-title":"The media equation: How people treat computers, television, and new media like real people","volume":"10","author":"Reeves Byron","year":"1996","unstructured":"Byron Reeves and Clifford Nass. 1996. The media equation: How people treat computers, television, and new media like real people. Cambridge, UK 10, 10 (1996), 19\u201336.","journal-title":"Cambridge, UK"},{"key":"e_1_3_2_96_2","article-title":"A conversation with Bing\u2019s chatbot left me deeply unsettled","volume":"16","author":"Roose Kevin","year":"2023","unstructured":"Kevin Roose. 2023. A conversation with Bing\u2019s chatbot left me deeply unsettled. The New York Times 16 (2023).","journal-title":"The New York Times"},{"key":"e_1_3_2_97_2","unstructured":"Daniel Rose Chia-Chien Hung Marco Lepri Israa Alqassem Kiril Gashteovski and Carolin Lawrence. 2025. MEDDxAgent: A unified modular agent framework for explainable automatic differential diagnosis. arXiv:2502.19175. Retrieved from https:\/\/arxiv.org\/abs\/2502.19175 (2025)."},{"issue":"3","key":"e_1_3_2_98_2","first-page":"e80494","article-title":"Bioethics artificial intelligence advisory (BAIA): An agentic artificial intelligence (AI) framework for bioethical clinical decision support","volume":"17","author":"Roy Taposh P. Dutta","year":"2025","unstructured":"Taposh P. Dutta Roy. 2025. Bioethics artificial intelligence advisory (BAIA): An agentic artificial intelligence (AI) framework for bioethical clinical decision support. Cureus 17, 3 (2025), e80494.","journal-title":"Cureus"},{"key":"e_1_3_2_99_2","unstructured":"Rob Royce Marcel Kaufmann Jonathan Becktor Sangwoo Moon Kalind Carpenter Kai Pak Amanda Towler Rohan Thakker and Shehryar Khattak. 2024. Enabling novel mission operations and interactions with ROSA: The robot operating system agent. arXiv:2410.06472. Retrieved from https:\/\/arxiv.org\/abs\/2410.06472 (2024)."},{"key":"e_1_3_2_100_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Ruan Yangjun","year":"2024","unstructured":"Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J. Maddison, and Tatsunori Hashimoto. 2024. Identifying the risks of LM agents with an LM-emulated sandbox. In Proceedings of the 12th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=GEcwtMk1uA"},{"key":"e_1_3_2_101_2","doi-asserted-by":"publisher","DOI":"10.1163\/9789004409965_004"},{"key":"e_1_3_2_102_2","unstructured":"Samuel Schmidgall Rojin Ziaei Carl Harris Eduardo Reis Jeffrey Jopling and Michael Moor. 2024. AgentClinic: A multimodal agent benchmark to evaluate AI in simulated clinical environments. arXiv:2405.07960. Retrieved from https:\/\/arxiv.org\/abs\/2405.07960 (2024)."},{"key":"e_1_3_2_103_2","first-page":"1559","volume-title":"Proceedings of the 30th USENIX Security Symposium (USENIX Security 21)","author":"Schuster Roei","year":"2021","unstructured":"Roei Schuster, Congzheng Song, Eran Tromer, and Vitaly Shmatikov. 2021. You autocomplete me: Poisoning vulnerabilities in neural code completion. In Proceedings of the 30th USENIX Security Symposium (USENIX Security 21). 1559\u20131575."},{"key":"e_1_3_2_104_2","unstructured":"Leo Schwinn David Dobre Stephan G\u00fcnnemann and Gauthier Gidel. 2023. Adversarial attacks and defenses in large language models: Old and new threats. In PMLR. 103\u2013117."},{"key":"e_1_3_2_105_2","unstructured":"Erfan Shayegani Md Abdullah Al Mamun Yu Fu Pedram Zaree Yue Dong and Nael Abu-Ghazaleh. 2023. Survey of vulnerabilities in large language models revealed by adversarial attacks. arXiv:2310.10844. Retrieved from https:\/\/arxiv.org\/abs\/2310.10844 (2023)."},{"key":"e_1_3_2_106_2","doi-asserted-by":"crossref","unstructured":"Wenqi Shi Ran Xu Yuchen Zhuang Yue Yu Jieyu Zhang Hang Wu Yuanda Zhu Joyce Ho Carl Yang and May D. Wang. 2024. Ehragent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records. arXiv:2401.07128. Retrieved from https:\/\/arxiv.org\/abs\/2401.07128 (2024).","DOI":"10.18653\/v1\/2024.emnlp-main.1245"},{"key":"e_1_3_2_107_2","doi-asserted-by":"publisher","DOI":"10.1145\/1408800.1408906"},{"key":"e_1_3_2_108_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-023-06291-2"},{"key":"e_1_3_2_109_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41591-024-03423-7"},{"key":"e_1_3_2_110_2","volume-title":"Intelligent systems for engineering: a knowledge-based approach","author":"Sriram Ram D.","year":"2012","unstructured":"Ram D. Sriram. 2012. Intelligent systems for engineering: a knowledge-based approach. Springer Science & Business Media."},{"key":"e_1_3_2_111_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cmpb.2023.107525"},{"key":"e_1_3_2_112_2","unstructured":"Theodore Sumers Shunyu Yao Karthik Narasimhan and Thomas Griffiths. 2023. Cognitive architectures for language agents. Transactions on Machine Learning Research (TMLR) (2023)."},{"key":"e_1_3_2_113_2","unstructured":"Yiyou Sun Shawn Hu Georgia Zhou Ken Zheng Hannaneh Hajishirzi Nouha Dziri and Dawn Song. 2025. OMEGA: Can LLMs reason outside the box in math? evaluating exploratory compositional and transformative generalization. arXiv:2506.18880. Retrieved from https:\/\/arxiv.org\/abs\/2506.18880 (2025)."},{"key":"e_1_3_2_114_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1009669824615"},{"key":"e_1_3_2_115_2","unstructured":"Xiangru Tang Qiao Jin Kunlun Zhu Tongxin Yuan Yichi Zhang Wangchunshu Zhou Meng Qu Yilun Zhao Jian Tang Zhuosheng Zhang et\u00a0al. 2024. Prioritizing safeguarding over autonomy: Risks of llm agents for science. arXiv:2402.04247. Retrieved from https:\/\/arxiv.org\/abs\/2402.04247 (2024)."},{"key":"e_1_3_2_116_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.33"},{"key":"e_1_3_2_117_2","first-page":"35413","volume-title":"International Conference on Machine Learning","author":"Wan Alexander","year":"2023","unstructured":"Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023. Poisoning language models during instruction tuning. In International Conference on Machine Learning. PMLR, 35413\u201335425."},{"key":"e_1_3_2_118_2","unstructured":"Cunxiang Wang Ruoxi Ning Boqi Pan Tonghui Wu Qipeng Guo Cheng Deng Guangsheng Bao Qian Wang and Yue Zhang. 2024. Novelqa: A benchmark for long-range novel question answering. arXiv.2403.12766v1. Retrieved from https:\/\/arxiv.org\/abs\/2403.12766v1 (2024)."},{"key":"e_1_3_2_119_2","article-title":"Voyager: An open-ended embodied agent with large language models","author":"Wang Guanzhi","year":"2024","unstructured":"Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2024. Voyager: An open-ended embodied agent with large language models. Transactions on Machine Learning Research (2024). Retrieved from https:\/\/openreview.net\/forum?id=ehfRiF0R3a","journal-title":"Transactions on Machine Learning Research"},{"key":"e_1_3_2_120_2","doi-asserted-by":"crossref","unstructured":"Wenxuan Wang Zizhan Ma Zheng Wang Chenghan Wu Wenting Chen Xiang Li and Yixuan Yuan. 2025. A survey of LLM-based agents in medicine: How far are we from Baymax? arXiv:2502.11211. Retrieved from https:\/\/arxiv.org\/abs\/2502.11211 (2025).","DOI":"10.18653\/v1\/2025.findings-acl.539"},{"key":"e_1_3_2_121_2","unstructured":"Yuntao Wang Yanghe Pan Quan Zhao Yi Deng Zhou Su Linkang Du and Tom H. Luan. 2024. Large model agents: State-of-the-art cooperation paradigms security and privacy and future trends. arXiv:2409.14457. Retrieved from https:\/\/arxiv.org\/abs\/2409.14457 (2024)."},{"key":"e_1_3_2_122_2","doi-asserted-by":"publisher","DOI":"10.52202\/075280-1480"},{"key":"e_1_3_2_123_2","doi-asserted-by":"publisher","DOI":"10.1001\/jama.2024.21451"},{"key":"e_1_3_2_124_2","unstructured":"Jinjie Wei Dingkang Yang Yanshu Li Qingyao Xu Zhaoyu Chen Mingcheng Li Yue Jiang Xiaolu Hou and Lihua Zhang. 2024. Medaide: Towards an omni medical aide via specialized llm-based multi-agent collaboration. arXiv:2410.12532. Retrieved from https:\/\/arxiv.org\/abs\/2410.12532 (2024)."},{"key":"e_1_3_2_125_2","volume-title":"KDD\u201924 Workshop: Artificial Intelligence and Data Science for Healthcare: Bridging Data-Centric AI and People-Centric Healthcare","author":"Wu Hao","year":"2024","unstructured":"Hao Wu, Yinghao Zhu, Zixiang Wang, Xiaochen Zheng, Ling Wang, Wen Tang, Yasha Wang, Chengwei Pan, Ewen M. Harrison, Junyi Gao, et\u00a0al. 2024. EHRFlow: A large language model-driven iterative multi-agent electronic health record data analysis workflow. In KDD\u201924 Workshop: Artificial Intelligence and Data Science for Healthcare: Bridging Data-Centric AI and People-Centric Healthcare."},{"key":"e_1_3_2_126_2","unstructured":"Qingyun Wu Gagan Bansal Jieyu Zhang Yiran Wu Beibin Li Erkang Zhu Li Jiang Xiaoyun Zhang Shaokun Zhang Jiale Liu et\u00a0al. 2023. Autogen: Enabling next-gen llm applications via multi-agent conversation. arXiv:2308.08155. Retrieved from https:\/\/arxiv.org\/abs\/2308.08155 (2023)."},{"key":"e_1_3_2_127_2","unstructured":"Yue Wu Xuan Tang Tom M. Mitchell and Yuanzhi Li. 2023. Smartplay: A benchmark for llms as intelligent agents. arXiv:2310.01557. Retrieved from https:\/\/arxiv.org\/abs\/2310.01557 (2023)."},{"key":"e_1_3_2_128_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-024-4222-0"},{"key":"e_1_3_2_129_2","unstructured":"Ancheng Xu Di Yang Renhao Li Jingwei Zhu Minghuan Tan Min Yang Wanxin Qiu Mingchen Ma Haihong Wu Bingyu Li et\u00a0al. 2025. AutoCBT: An autonomous multi-agent framework for cognitive behavioral therapy in psychological counseling. arXiv:2501.09426. Retrieved from https:\/\/arxiv.org\/abs\/2501.09426 (2025)."},{"key":"e_1_3_2_130_2","doi-asserted-by":"crossref","unstructured":"Ran Xu Wenqi Shi Jonathan Wang Jasmine Zhou and Carl Yang. 2025. MedAssist: LLM-empowered medical assistant for assisting the scrutinization and comprehension of electronic health records. In Companion Proceedings of the ACM on Web Conference. 2931\u20132934.","DOI":"10.1145\/3701716.3715186"},{"key":"e_1_3_2_131_2","unstructured":"Xinrun Xu Yuxin Wang Chaoyi Xu Ziluo Ding Jiechuan Jiang Zhiming Ding and B\u00f6rje F. Karlsson. 2024. A survey on game playing agents and large models: Methods applications and challenges. arXiv:2403.10249. Retrieved from https:\/\/arxiv.org\/abs\/2403.10249 (2024)."},{"key":"e_1_3_2_132_2","unstructured":"Bingyu Yan Xiaoming Zhang Litian Zhang Lian Zhang Ziyi Zhou Dezhuang Miao and Chaozhuo Li. 2025. Beyond self-talk: A communication-centric survey of LLM-based multi-agent systems. arXiv:2502.14321. Retrieved from https:\/\/arxiv.org\/abs\/2502.14321 (2025)."},{"key":"e_1_3_2_133_2","unstructured":"Hang Yang Hao Chen Hui Guo Yineng Chen Ching-Sheng Lin Shu Hu Jinrong Hu Xi Wu and Xin Wang. 2024. LLM-MedQA: Enhancing medical question answering through case studies in large language models. arXiv:2501.05464. Retrieved from https:\/\/arxiv.org\/abs\/2501.05464 (2024)."},{"key":"e_1_3_2_134_2","doi-asserted-by":"crossref","unstructured":"Jianwei Yang Reuben Tan Qianhui Wu Ruijie Zheng Baolin Peng Yongyuan Liang Yu Gu Mu Cai Seonghyeon Ye Joel Jang et\u00a0al. 2025. Magma: A foundation model for multimodal AI agents. arXiv:2502.13130. Retrieved from https:\/\/arxiv.org\/abs\/2502.13130 (2025).","DOI":"10.1109\/CVPR52734.2025.01325"},{"key":"e_1_3_2_135_2","doi-asserted-by":"crossref","unstructured":"Zhilin Yang Peng Qi Saizheng Zhang Yoshua Bengio William W. Cohen Ruslan Salakhutdinov and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse explainable multi-hop question answering. arXiv:1809.09600. Retrieved from https:\/\/arxiv.org\/abs\/1809.09600 (2018).","DOI":"10.18653\/v1\/D18-1259"},{"key":"e_1_3_2_136_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1508"},{"key":"e_1_3_2_137_2","unstructured":"Haoqi Yuan Chi Zhang Hongcheng Wang Feiyang Xie Penglin Cai Hao Dong and Zongqing Lu. 2023. Skill reinforcement learning and planning for open-world long-horizon tasks. arXiv:2303.16563. Retrieved from https:\/\/arxiv.org\/abs\/2303.16563 (2023)."},{"key":"e_1_3_2_138_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-emnlp.79"},{"key":"e_1_3_2_139_2","article-title":"The shift from models to compound ai systems","author":"Zaharia Matei","year":"2024","unstructured":"Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, et\u00a0al. 2024. The shift from models to compound ai systems. Berkeley Artificial Intelligence Research Lab. Available online at: https:\/\/bair.berkeley.edu\/blog\/2024\/02\/18\/compound-ai-systems\/ (accessed February 27, 2024) (2024).","journal-title":"Berkeley Artificial Intelligence Research Lab. Available online at:"},{"key":"e_1_3_2_140_2","unstructured":"Boyang Zhang Yicong Tan Yun Shen Ahmed Salem Michael Backes Savvas Zannettou and Yang Zhang. 2024. Breaking agents: Compromising autonomous llm agents through malfunction amplification. arXiv:2407.20859. Retrieved from https:\/\/arxiv.org\/abs\/2407.20859 (2024)."},{"key":"e_1_3_2_141_2","unstructured":"Hongxin Zhang Weihua Du Jiaming Shan Qinhong Zhou Yilun Du Joshua B. Tenenbaum Tianmin Shu and Chuang Gan. 2023. Building cooperative embodied agents modularly with large language models. arXiv:2307.02485. Retrieved from https:\/\/arxiv.org\/abs\/2307.02485 (2023)."},{"key":"e_1_3_2_142_2","unstructured":"Xiaoman Zhang Chaoyi Wu Ziheng Zhao Weixiong Lin Ya Zhang Yanfeng Wang and Weidi Xie. 2023. Pmc-vqa: Visual instruction tuning for medical visual question answering. arXiv:2305.10415. Retrieved from https:\/\/arxiv.org\/abs\/2305.10415 (2023)."},{"key":"e_1_3_2_143_2","unstructured":"Yiqun Zhang Xiaocui Yang Xiaobai Li Siyuan Yu Yi Luan Shi Feng Daling Wang and Yifei Zhang. 2024. PsyDraw: A multi-agent multimodal system for mental health screening in left-behind children. arXiv:2412.14769. Retrieved from https:\/\/arxiv.org\/abs\/2412.14769 (2024)."},{"key":"e_1_3_2_144_2","doi-asserted-by":"publisher","DOI":"10.1145\/3711896.3736573"},{"key":"e_1_3_2_145_2","unstructured":"Shuyan Zhou Frank F. Xu Hao Zhu Xuhui Zhou Robert Lo Abishek Sridhar Xianyi Cheng Tianyue Ou Yonatan Bisk Daniel Fried et\u00a0al. 2023. Webarena: A realistic web environment for building autonomous agents. arXiv:2307.13854. Retrieved from https:\/\/arxiv.org\/abs\/2307.13854 (2023)."},{"key":"e_1_3_2_146_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Zhou Shuyan","year":"2024","unstructured":"Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et\u00a0al. 2024. WebArena: A realistic web environment for building autonomous agents. In Proceedings of the 12th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=oKn9c6ytLx"},{"key":"e_1_3_2_147_2","unstructured":"Yuan Zhou Peng Zhang Mengya Song Alice Zheng Yiwen Lu Zhiheng Liu Yong Chen and Zhaohan Xi. 2024. Zodiac: A cardiologist-level LLM framework for multi-agent diagnostics. arXiv:2410.02026. Retrieved from https:\/\/arxiv.org\/abs\/2410.02026 (2024)."},{"key":"e_1_3_2_148_2","first-page":"3827","volume-title":"Proceedings of the 34th USENIX Security Symposium (USENIX Security 25)","author":"Zou Wei","year":"2025","unstructured":"Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2025. \\(\\lbrace\\) PoisonedRAG \\(\\rbrace\\) : Knowledge corruption attacks to \\(\\lbrace\\) retrieval-augmented \\(\\rbrace\\) generation of large language models. In Proceedings of the 34th USENIX Security Symposium (USENIX Security 25). 3827\u20133844."}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3809164","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T13:04:42Z","timestamp":1782306282000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3809164"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":147,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2026,10,31]]}},"alternative-id":["10.1145\/3809164"],"URL":"https:\/\/doi.org\/10.1145\/3809164","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-03-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-06","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}