{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T15:41:25Z","timestamp":1782574885498,"version":"3.54.5"},"reference-count":202,"publisher":"Association for Computing Machinery (ACM)","issue":"8","license":[{"start":{"date-parts":[[2025,3,22]],"date-time":"2025-03-22T00:00:00Z","timestamp":1742601600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62406188"],"award-info":[{"award-number":["62406188"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"Joint Funds of the National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U21B2020"],"award-info":[{"award-number":["U21B2020"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Natural Science Foundation of Shanghai","award":["24ZR1440300"],"award-info":[{"award-number":["24ZR1440300"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2025,8,31]]},"abstract":"<jats:p>\n            Large language models (LLMs) have dramatically enhanced the field of language intelligence, as demonstrably evidenced by their formidable empirical performance across a spectrum of complex reasoning tasks. Additionally, theoretical proofs have illuminated their emergent reasoning capabilities, providing a compelling showcase of their advanced cognitive abilities in linguistic contexts. Critical to their remarkable efficacy in handling complex reasoning tasks, LLMs leverage the intriguing chain-of-thought (CoT) reasoning techniques, obliging them to formulate intermediate steps en route to deriving an answer. The CoT reasoning approach has not only exhibited proficiency in amplifying reasoning performance but also in enhancing interpretability, controllability, and flexibility. In light of these merits, recent research endeavors have extended CoT reasoning methodologies to nurture the development of autonomous language agents, which adeptly adhere to language instructions and execute actions within varied environments. This survey article orchestrates a thorough discourse, penetrating vital research dimensions, encompassing (i) the foundational mechanics of CoT techniques, with a focus on elucidating the circumstances and justification behind its efficacy; (ii) the paradigm shift in CoT; and (iii) the burgeoning of language agents fortified by CoT approaches. Prospective research avenues envelop explorations into generalization, efficiency, customization, scaling, and safety. A repository for the related papers is available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/Zoeyyao27\/CoT-Igniting-Agent\">https:\/\/github.com\/Zoeyyao27\/CoT-Igniting-Agent<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3719341","type":"journal-article","created":{"date-parts":[[2025,2,25]],"date-time":"2025-02-25T09:42:47Z","timestamp":1740476567000},"page":"1-39","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":37,"title":["Igniting Language Intelligence: The Hitchhiker\u2019s Guide from Chain-of-Thought Reasoning to Language Agents"],"prefix":"10.1145","volume":"57","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4183-3645","authenticated-orcid":false,"given":"Zhuosheng","family":"Zhang","sequence":"first","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6055-2546","authenticated-orcid":false,"given":"Yao","family":"Yao","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0890-0524","authenticated-orcid":false,"given":"Aston","family":"Zhang","sequence":"additional","affiliation":[{"name":"Amazon Web Services, Seattle, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2700-4513","authenticated-orcid":false,"given":"Xiangru","family":"Tang","sequence":"additional","affiliation":[{"name":"Yale University, New Haven, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1505-8603","authenticated-orcid":false,"given":"Xinbei","family":"Ma","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4807-0062","authenticated-orcid":false,"given":"Zhiwei","family":"He","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5821-8895","authenticated-orcid":false,"given":"Yiming","family":"Wang","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9746-3719","authenticated-orcid":false,"given":"Mark","family":"Gerstein","sequence":"additional","affiliation":[{"name":"Yale University, New Haven, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8007-2503","authenticated-orcid":false,"given":"Rui","family":"Wang","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5194-1570","authenticated-orcid":false,"given":"Gongshen","family":"Liu","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7290-0487","authenticated-orcid":false,"given":"Hai","family":"Zhao","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,3,22]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"Adept. 2022. ACT-1: Transformer for Actions. Retrieved February 26 2025 from https:\/\/www.adept.ai\/act"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.5555\/3618408.3618425"},{"key":"e_1_3_3_4_2","article-title":"Flamingo: A visual language model for few-shot learning","author":"Alayrac Jean-Baptiste","year":"2022","unstructured":"Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et\u00a0al. 2022. Flamingo: A visual language model for few-shot learning. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS\u201922). 23716\u201323736.","journal-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS\u201922)."},{"key":"e_1_3_3_5_2","article-title":"Palm 2 technical report","author":"Anil Rohan","year":"2023","unstructured":"Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et\u00a0al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403 (2023).","journal-title":"arXiv preprint arXiv:2305.10403"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.34133\/icomputing.0025"},{"key":"e_1_3_3_7_2","article-title":"Qwen-VL: A frontier large vision-language model with versatile abilities","volume":"2308","author":"Bai Jinze","year":"2023","unstructured":"Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-VL: A frontier large vision-language model with versatile abilities. arXiv preprint abs\/2308.12966 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3643757"},{"key":"e_1_3_3_9_2","unstructured":"Rohan Bavishi Erich Elsen Curtis Hawthorne Maxwell Nye Augustus Odena Arushi Somani and Sa\u011fnak Ta\u015f\u0131rlar. 2023. Fuyu-8B: A Multimodal Architecture for AI Agents. Retrieved February 26 2025 from https:\/\/www.adept.ai\/blog\/fuyu-8b"},{"key":"e_1_3_3_10_2","volume-title":"Proceedings of the 38th AAAI Conference on Artificial Intelligence, the 36th Conference on Innovative Applications of Artificial Intelligence, and the 14th Symposium on Educational Advances in Artificial Intelligence (AAAI\u201924\/IAAI\u201924\/EAAI\u201924)","author":"Besta Maciej","year":"2024","unstructured":"Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et\u00a0al. 2024. Graph of thoughts: Solving elaborate problems with large language models. In Proceedings of the 38th AAAI Conference on Artificial Intelligence, the 36th Conference on Innovative Applications of Artificial Intelligence, and the 14th Symposium on Educational Advances in Artificial Intelligence (AAAI\u201924\/IAAI\u201924\/EAAI\u201924). 17682\u201317690."},{"key":"e_1_3_3_11_2","first-page":"17691","volume-title":"Proceedings of the 38th AAAI Conference on Artificial Intelligence, the 36th Conference on Innovative Applications of Artificial Intelligence, and the 14th Symposium on Educational Advances in Artificial Intelligence (AAAI\u201924\/IAAI\u201924\/EAAI\u201924)","author":"Bi Zhen","year":"2024","unstructured":"Zhen Bi, Ningyu Zhang, Yinuo Jiang, Shumin Deng, Guozhou Zheng, and Huajun Chen. 2024. When do program-of-thought works for reasoning? In Proceedings of the 38th AAAI Conference on Artificial Intelligence, the 36th Conference on Innovative Applications of Artificial Intelligence, and the 14th Symposium on Educational Advances in Artificial Intelligence (AAAI\u201924\/IAAI\u201924\/EAAI\u201924). 17691\u201317699."},{"key":"e_1_3_3_12_2","article-title":"Emergent autonomous scientific research capabilities of large language models","volume":"2304","author":"Boiko Daniil A.","year":"2023","unstructured":"Daniil A. Boiko, Robert MacKnight, and Gabe Gomes. 2023. Emergent autonomous scientific research capabilities of large language models. arXiv preprint abs\/2304.05332 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_13_2","article-title":"Language models are few-shot learners","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et\u00a0al. 2020. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS\u201920). 1877\u20131901.","journal-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS\u201920). 1877\u20131901."},{"key":"e_1_3_3_14_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Cai Tianle","year":"2024","unstructured":"Tianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen, and Denny Zhou. 2024. Large language models as tool makers. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3641289"},{"key":"e_1_3_3_16_2","article-title":"FireAct: Toward language agent fine-tuning","volume":"2310","author":"Chen Baian","year":"2023","unstructured":"Baian Chen, Chang Shu, Ehsan Shareghi, Nigel Collier, Karthik Narasimhan, and Shunyu Yao. 2023. FireAct: Toward language agent fine-tuning. arXiv preprint abs\/2310.05915 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_17_2","article-title":"Walking down the memory maze: Beyond context limit through interactive reading","volume":"2310","author":"Chen Howard","year":"2023","unstructured":"Howard Chen, Ramakanth Pasunuru, Jason Weston, and Asli Celikyilmaz. 2023. Walking down the memory maze: Beyond context limit through interactive reading. arXiv preprint abs\/2310.05029 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_18_2","first-page":"45767","article-title":"Are more LLM calls all you need? Towards the scaling properties of compound AI systems","volume":"37","author":"Chen Lingjiao","year":"2025","unstructured":"Lingjiao Chen, Jared Quincy Davis, Boris Hanin, Peter Bailis, Ion Stoica, Matei A Zaharia, and James Y. Zou. 2025. Are more LLM calls all you need? Towards the scaling properties of compound AI systems. Advances in Neural Information Processing Systems 37 (2025), 45767\u201345790.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_19_2","article-title":"Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks","author":"Chen Wenhu","year":"2023","unstructured":"Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. 2023. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. Transactions on Machine Learning Research. Published Online, November 1, 2023.","journal-title":"Transactions on Machine Learning Research."},{"key":"e_1_3_3_20_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Chen Xinyun","year":"2024","unstructured":"Xinyun Chen, Maxwell Lin, Nathanael Sch\u00e4rli, and Denny Zhou. 2024. Teaching large language models to self-debug. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_21_2","article-title":"Navigate through enigmatic labyrinth\u2014A survey of chain of thought reasoning: Advances, frontiers and future","author":"Chu Zheng","year":"2024","unstructured":"Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He, Haotian Wang, Weihua Peng, Ming Liu, Bing Qin, and Ting Liu. 2024. Navigate through enigmatic labyrinth\u2014A survey of chain of thought reasoning: Advances, frontiers and future. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1173\u20131203.","journal-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1173\u20131203"},{"issue":"70","key":"e_1_3_3_22_2","first-page":"1","article-title":"Scaling instruction-finetuned language models","volume":"25","author":"Chung Hyung Won","year":"2024","unstructured":"Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et\u00a0al. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research 25, 70 (2024), 1\u201353.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_3_23_2","article-title":"Training verifiers to solve math word problems","volume":"2110","author":"Cobbe Karl","year":"2021","unstructured":"Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et\u00a0al. 2021. Training verifiers to solve math word problems. arXiv preprint abs\/2110.14168 (2021).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_24_2","volume-title":"Proceedings of the 31st Annual Network and Distributed System Security Symposium (NDSS\u201924)","author":"Deng Gelei","year":"2024","unstructured":"Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2024. MASTERKEY: Automated jailbreaking of large language model chatbots. In Proceedings of the 31st Annual Network and Distributed System Security Symposium (NDSS\u201924)."},{"key":"e_1_3_3_25_2","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/2024.acl-long.73","article-title":"Active prompting with chain-of-thought for large language models","author":"Diao Shizhe","year":"2024","unstructured":"Shizhe Diao, Pengcheng Wang, Yong Lin, Rui Pan, Xiang Liu, and Tong Zhang. 2024. Active prompting with chain-of-thought for large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1330\u20131350.","journal-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1330\u20131350."},{"key":"e_1_3_3_26_2","first-page":"8469","volume-title":"Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)","author":"Driess Danny","year":"2023","unstructured":"Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et\u00a0al. 2023. PaLM-E: An embodied multimodal language model. In Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923). 8469\u20138488."},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.5555\/3692070.3692537"},{"key":"e_1_3_3_28_2","article-title":"The Llama 3 herd of models","author":"Dubey Abhimanyu","year":"2024","unstructured":"Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et\u00a0al. 2024. The Llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024).","journal-title":"arXiv preprint arXiv:2407.21783"},{"key":"e_1_3_3_29_2","article-title":"Misusing tools in large language models with visual adversarial examples","author":"Fu Xiaohan","year":"2023","unstructured":"Xiaohan Fu, Zihan Wang, Shuheng Li, Rajesh K. Gupta, Niloofar Mireshghallah, Taylor Berg-Kirkpatrick, and Earlence Fernandes. 2023. Misusing tools in large language models with visual adversarial examples. arXiv preprint arXiv:2310.03185 (2023).","journal-title":"arXiv preprint arXiv:2310.03185"},{"key":"e_1_3_3_30_2","article-title":"OpenAGI: When LLM meets domain experts","author":"Ge Yingqiang","year":"2023","unstructured":"Yingqiang Ge, Wenyue Hua, Kai Mei, Jianchao Ji, Juntao Tan, Shuyuan Xu, Zelong Li, and Yongfeng Zhang. 2023. OpenAGI: When LLM meets domain experts. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 5539\u20135568.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 5539\u20135568."},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00370"},{"key":"e_1_3_3_32_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Gou Zhibin","year":"2024","unstructured":"Zhibin Gou, Zhihong Shao, Yeyun Gong, yelong shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2024. CRITIC: Large language models can self-correct with tool-interactive critiquing. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_33_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Gou Zhibin","year":"2024","unstructured":"Zhibin Gou, Zhihong Shao, Yeyun Gong, yelong shen, Yujiu Yang, Minlie Huang, Nan Duan, and Weizhu Chen. 2024. ToRA: A tool-integrated reasoning agent for mathematical problem solving. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_34_2","unstructured":"Lin Guan Karthik Valmeekam Sarath Sreedharan and Subbarao Kambhampati. 2023. Leveraging pre-trained large language models to construct and utilize world models for model-based task planning. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 79081\u201379094."},{"key":"e_1_3_3_35_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Gur Izzeddin","year":"2024","unstructured":"Izzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust. 2024. A real-world webagent with planning, long context understanding, and program synthesis. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_36_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Gurnee Wes","year":"2024","unstructured":"Wes Gurnee and Max Tegmark. 2024. Language models represent space and time. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_37_2","first-page":"2455","article-title":"Recurrent world models facilitate policy evolution","author":"Ha David","year":"2018","unstructured":"David Ha and J\u00fcrgen Schmidhuber. 2018. Recurrent world models facilitate policy evolution. In Proceedings of the 32nd International Conference on Neural Information Processing Systems (NIPS\u201918).2455\u20132467.","journal-title":"Proceedings of the 32nd International Conference on Neural Information Processing Systems (NIPS\u201918)."},{"key":"e_1_3_3_38_2","first-page":"8154","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201923)","author":"Hao Shibo","year":"2023","unstructured":"Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. 2023. Reasoning with language model is planning with world model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201923). 8154\u20138173."},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00642"},{"key":"e_1_3_3_40_2","article-title":"Is there an intelligent agent in your future?","volume":"11","author":"Hendler James","year":"1999","unstructured":"James Hendler. 1999. Is there an intelligent agent in your future? Nature Web Matters 11 (1999), 1\u20138.","journal-title":"Nature Web Matters"},{"key":"e_1_3_3_41_2","volume-title":"Proceedings of the 9th International Conference on Learning Representations (ICLR\u201921)","author":"Hendrycks Dan","year":"2021","unstructured":"Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In Proceedings of the 9th International Conference on Learning Representations (ICLR\u201921)."},{"key":"e_1_3_3_42_2","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks","author":"Hendrycks Dan","year":"2021","unstructured":"Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the MATH dataset. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks."},{"key":"e_1_3_3_43_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Hong Sirui","year":"2024","unstructured":"Sirui Hong, Mingchen Zhuge, Jiaqi Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Jinlin Wang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, et\u00a0al. 2024. MetaGPT: Meta programming for a multi-agent collaborative framework. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_44_2","article-title":"ChatDB: Augmenting LLMs with databases as their symbolic memory","volume":"2306","author":"Hu Chenxu","year":"2023","unstructured":"Chenxu Hu, Jie Fu, Chenzhuang Du, Simian Luo, Junbo Zhao, and Hang Zhao. 2023. ChatDB: Augmenting LLMs with databases as their symbolic memory. arXiv preprint abs\/2306.03901 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_45_2","doi-asserted-by":"crossref","first-page":"1049","DOI":"10.18653\/v1\/2023.findings-acl.67","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Huang Jie","year":"2023","unstructured":"Jie Huang and Kevin Chen-Chuan Chang. 2023. Towards reasoning in large language models: A survey. In Findings of the Association for Computational Linguistics: ACL 2023. 1049\u20131065."},{"key":"e_1_3_3_46_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Huang Jie","year":"2024","unstructured":"Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2024. Large language models cannot self-correct reasoning yet. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_47_2","first-page":"72096","article-title":"Language is not all you need: Aligning perception with language models","author":"Huang Shaohan","year":"2023","unstructured":"Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, et\u00a0al. 2023. Language is not all you need: Aligning perception with language models. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923).72096\u201372109.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923)."},{"key":"e_1_3_3_48_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)","author":"Huang Zeyu","year":"2023","unstructured":"Zeyu Huang, Yikang Shen, Xiaofeng Zhang, Jie Zhou, Wenge Rong, and Zhang Xiong. 2023. Transformer-Patcher: One mistake worth one neuron. In Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)."},{"key":"e_1_3_3_49_2","volume-title":"Proceedings of the NeurIPS 2022 Foundation Models for Decision Making Workshop","author":"Jiang Yunfan","year":"2022","unstructured":"Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan. 2022. VIMA: General robot manipulation with multimodal prompts. In Proceedings of the NeurIPS 2022 Foundation Models for Decision Making Workshop."},{"key":"e_1_3_3_50_2","article-title":"ChatMOF: An autonomous AI system for predicting and generating metal-organic frameworks","volume":"2308","author":"Kang Yeonghun","year":"2023","unstructured":"Yeonghun Kang and Jihan Kim. 2023. ChatMOF: An autonomous AI system for predicting and generating metal-organic frameworks. arXiv preprint abs\/2308.01423 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_51_2","article-title":"Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP","volume":"2212","author":"Khattab Omar","year":"2022","unstructured":"Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia. 2022. Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP. arXiv preprint abs\/2212.14024 (2022).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_52_2","article-title":"Language models can solve computer tasks","author":"Kim Geunwoo","year":"2023","unstructured":"Geunwoo Kim, Pierre Baldi, and Stephen McAleer. 2023. Language models can solve computer tasks. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 39648\u201339677.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 39648\u201339677."},{"key":"e_1_3_3_53_2","article-title":"ProPILE: Probing privacy leakage in large language models","author":"Kim Siwon","year":"2023","unstructured":"Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh. 2023. ProPILE: Probing privacy leakage in large language models. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 20750\u201320762.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 20750\u201320762."},{"key":"e_1_3_3_54_2","article-title":"Evaluating language-model agents on realistic autonomous tasks","author":"Kinniment Megan","year":"2023","unstructured":"Megan Kinniment, Lucas Jun Koba Sato, Haoxing Du, Brian Goodrich, Max Hasin, Lawrence Chan, Luke Harold Miles, Tao R. Lin, Hjalmar Wijk, Joel Burget, et\u00a0al. 2023. Evaluating language-model agents on realistic autonomous tasks. arXiv preprint arXiv:2312.11671 (2023).","journal-title":"arXiv preprint arXiv:2312.11671"},{"key":"e_1_3_3_55_2","first-page":"22199","article-title":"Large language models are zero-shot reasoners","author":"Kojima Takeshi","year":"2022","unstructured":"Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS\u201922).22199\u201322213.","journal-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS\u201922)."},{"key":"e_1_3_3_56_2","first-page":"623","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Lee Soochan","year":"2023","unstructured":"Soochan Lee and Gunhee Kim. 2023. Recursion of thought: A divide-and-conquer approach to multi-context reasoning with language models. In Findings of the Association for Computational Linguistics: ACL 2023. 623\u2013658."},{"key":"e_1_3_3_57_2","first-page":"51991","article-title":"CAMEL: Communicative agents for \u201cmind\u201d exploration of large language model society","author":"Li Guohao","year":"2023","unstructured":"Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. CAMEL: Communicative agents for \u201cmind\u201d exploration of large language model society. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923).51991\u201352008.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923)."},{"key":"e_1_3_3_58_2","first-page":"9053","volume-title":"Findings of the Association for Computational Linguistics: ACL 2024","author":"Li Huayang","year":"2024","unstructured":"Huayang Li, Siheng Li, Deng Cai, Longyue Wang, Lemao Liu, Taro Watanabe, Yujiu Yang, and Shuming Shi. 2024. TextBind: Multi-turn interleaved multimodal instruction-following in the wild. In Findings of the Association for Computational Linguistics: ACL 2024. 9053\u20139076."},{"key":"e_1_3_3_59_2","volume-title":"Proceedings of the 40th International Conference on Machine Learning (ICML\u201923)","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning (ICML\u201923)."},{"key":"e_1_3_3_60_2","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 296\u2013310.","author":"Li Junlong","year":"2024","unstructured":"Junlong Li, Jinyuan Wang, Zhuosheng Zhang, and Hai Zhao. 2024. Self-prompting large language models for zero-shot open-domain QA. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 296\u2013310."},{"key":"e_1_3_3_61_2","first-page":"3102","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201923)","author":"Li Minghao","year":"2023","unstructured":"Minghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song, Hangyu Li, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li. 2023. API-Bank: A comprehensive benchmark for tool-augmented LLMs. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201923). 3102\u20133116."},{"key":"e_1_3_3_62_2","first-page":"1","volume-title":"Proceedings of the 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU\u201923)","author":"Li Yuang","year":"2023","unstructured":"Yuang Li, Yu Wu, Jinyu Li, and Shujie Liu. 2023. Prompting large language models for zero-shot domain adaptation in speech recognition. In Proceedings of the 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU\u201923). 1\u20138."},{"key":"e_1_3_3_63_2","article-title":"MetaAgents: Simulating interactions of human behaviors for LLM-based task-oriented coordination via collaborative generative agents","volume":"2310","author":"Li Yuan","year":"2023","unstructured":"Yuan Li, Yixuan Zhang, and Lichao Sun. 2023. MetaAgents: Simulating interactions of human behaviors for LLM-based task-oriented coordination via collaborative generative agents. arXiv preprint abs\/2310.06500 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_64_2","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201924)","author":"Liang Tian","year":"2024","unstructured":"Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi. 2024. Encouraging divergent thinking in large language models through multi-agent debate. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201924)."},{"key":"e_1_3_3_65_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2410.05318"},{"key":"e_1_3_3_66_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Lightman Hunter","year":"2024","unstructured":"Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2024. Let\u2019s verify step by step. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_67_2","article-title":"SwiftSage: A generative agent with fast and slow thinking for complex interactive tasks","author":"Lin Bill Yuchen","year":"2023","unstructured":"Bill Yuchen Lin, Yicheng Fu, Karina Yang, Faeze Brahman, Shiyu Huang, Chandra Bhagavatula, Prithviraj Ammanabrolu, Yejin Choi, and Xiang Ren. 2023. SwiftSage: A generative agent with fast and slow thinking for complex interactive tasks. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 23813\u201323825.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 23813\u201323825."},{"key":"e_1_3_3_68_2","article-title":"AgentSims: An open-source sandbox for large language model evaluation","volume":"2308","author":"Lin Jiaju","year":"2023","unstructured":"Jiaju Lin, Haoran Zhao, Aochi Zhang, Yiting Wu, Huqiuyue Ping, and Qin Chen. 2023. AgentSims: An open-source sandbox for large language model evaluation. arXiv preprint abs\/2308.04026 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_69_2","doi-asserted-by":"crossref","first-page":"158","DOI":"10.18653\/v1\/P17-1015","article-title":"Program induction by rationale generation: Learning to solve and explain algebraic word problems","author":"Ling Wang","year":"2017","unstructured":"Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017. Program induction by rationale generation: Learning to solve and explain algebraic word problems. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).158\u2013167.","journal-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)."},{"key":"e_1_3_3_70_2","article-title":"Deductive verification of chain-of-thought reasoning","author":"Ling Zhan","year":"2023","unstructured":"Zhan Ling, Yunhao Fang, Xuanlin Li, Zhiao Huang, Mingu Lee, Roland Memisevic, and Hao Su. 2023. Deductive verification of chain-of-thought reasoning. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 36407\u201336433.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 36407\u201336433."},{"key":"e_1_3_3_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02484"},{"key":"e_1_3_3_72_2","article-title":"Visual instruction tuning","author":"Liu Haotian","year":"2023","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 34892\u201334916.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 34892\u201334916."},{"key":"e_1_3_3_73_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Liu Xiao","year":"2024","unstructured":"Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et\u00a0al. 2024. AgentBench: Evaluating LLMs as agents. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_74_2","article-title":"Prompt injection attack against LLM-integrated applications","volume":"2306","author":"Liu Yi","year":"2023","unstructured":"Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, and Yang Liu. 2023. Prompt injection attack against LLM-integrated applications. arXiv preprint abs\/2306.05499 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_75_2","article-title":"ControlLLM: Augment language models with tools by searching on graphs","volume":"2310","author":"Liu Zhaoyang","year":"2023","unstructured":"Zhaoyang Liu, Zeqiang Lai, Zhangwei Gao, Erfei Cui, Xizhou Zhu, Lewei Lu, Qifeng Chen, Yu Qiao, Jifeng Dai, and Wenhai Wang. 2023. ControlLLM: Augment language models with tools by searching on graphs. arXiv preprint abs\/2310.17796 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_76_2","first-page":"14605","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Lu Pan","year":"2023","unstructured":"Pan Lu, Liang Qiu, Wenhao Yu, Sean Welleck, and Kai-Wei Chang. 2023. A survey of deep learning for mathematical reasoning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 14605\u201314631."},{"key":"e_1_3_3_77_2","first-page":"4765","article-title":"A unified approach to interpreting model predictions","author":"Lundberg Scott M.","year":"2017","unstructured":"Scott M. Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS\u201917).4765\u20134774.","journal-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS\u201917)."},{"key":"e_1_3_3_78_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.lindif.2013.01.010"},{"key":"e_1_3_3_79_2","doi-asserted-by":"crossref","first-page":"525","DOI":"10.1038\/s42256-024-00832-8","article-title":"Augmenting large language models with chemistry tools","author":"Bran Andres M.","year":"2024","unstructured":"Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D. White, and Philippe Schwaller. 2024. Augmenting large language models with chemistry tools. Nature Machine Intelligence 6 (2024), 525\u2013535.","journal-title":"Nature Machine Intelligence"},{"key":"e_1_3_3_80_2","first-page":"1105","volume-title":"AMIA Annual Symposium Proceedings","author":"Ma Zilin","year":"2023","unstructured":"Zilin Ma, Yiyang Mei, and Zhaoyuan Su. 2023. Understanding the benefits and challenges of using large language model-based conversational agents for mental well-being support. AMIA Annual Symposium Proceedings 2023 (2023), 1105\u20131114."},{"key":"e_1_3_3_81_2","first-page":"811","volume-title":"Readings in Human\u2013Computer Interaction: Toward the Year 2000","author":"Maes Pattie","year":"1995","unstructured":"Pattie Maes. 1995. Agents that reduce work and information overload. In Readings in Human\u2013Computer Interaction: Toward the Year 2000 (2nd ed.), Ronald M. Baecker, Jonathan Grudin, William A. S. Buxton, and Saul Greenberg (Eds.). Elsevier, 811\u2013821."},{"key":"e_1_3_3_82_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-short.151"},{"key":"e_1_3_3_83_2","first-page":"286","volume-title":"Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA\u201924)","author":"Mandi Zhao","year":"2024","unstructured":"Zhao Mandi, Shreeya Jain, and Shuran Song. 2024. RoCo: Dialectic multi-robot collaboration with large language models. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA\u201924). 286\u2013299."},{"key":"e_1_3_3_84_2","volume-title":"Proceedings of the 13th National CCF Conference on Natural Language Processing and Chinese Computing (NLPCC\u201924)","author":"Mao Shengyu","year":"2024","unstructured":"Shengyu Mao, Ningyu Zhang, Xiaohan Wang, Mengru Wang, Yunzhi Yao, Yong Jiang, Pengjun Xie, Fei Huang, and Huajun Chen. 2024. Editing personality for LLMs. In Proceedings of the 13th National CCF Conference on Natural Language Processing and Chinese Computing (NLPCC\u201924). 241\u2013254."},{"key":"e_1_3_3_85_2","first-page":"1306","volume-title":"Findings of the Association for Computational Linguistics: EACL 2024","author":"Mehta Nikhil","year":"2024","unstructured":"Nikhil Mehta, Milagro Teruel, Xin Deng, Sergio Figueroa Sanz, Ahmed Awadallah, and Julia Kiseleva. 2024. Improving grounded language understanding in a collaborative environment by interacting with agents through help feedback. In Findings of the Association for Computational Linguistics: EACL 2024. 1306\u20131321."},{"key":"e_1_3_3_86_2","article-title":"Locating and editing factual associations in GPT","author":"Meng Kevin","year":"2022","unstructured":"Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS\u201922). 17359\u201317372.","journal-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS\u201922). 17359\u201317372."},{"key":"e_1_3_3_87_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)","author":"Meng Kevin","year":"2023","unstructured":"Kevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov, and David Bau. 2023. Mass-editing memory in a transformer. In Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)."},{"key":"e_1_3_3_88_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_3_89_2","unstructured":"Seungwhan Moon Andrea Madotto Zhaojiang Lin Tushar Nagarajan Matt Smith Shashank Jain Chun-Fu Yeh Prakash Murugesan Peyman Heidari Yue Liu et\u00a0al. 2023. AnyMAL: An efficient and scalable any-modality augmented language model. arxiv:2309.16058[cs.LG] (2023)."},{"key":"e_1_3_3_90_2","unstructured":"Louis-Philippe Morency Amir Zadeh and Paul Liang. 2022. Advanced Multimodal Machine Learning. Retrieved February 26 2025 from https:\/\/cmu-multicomp-lab.github.io\/adv-mmml-course\/spring2022\/schedule\/11877_week5.pdf"},{"key":"e_1_3_3_91_2","unstructured":"Yohei Nakajima. 2023. BabyAGI. Retrieved February 26 2025 from https:\/\/github.com\/yoheinakajima\/babyagi"},{"key":"e_1_3_3_92_2","article-title":"WebGPT: Browser-assisted question-answering with human feedback","volume":"2112","author":"Nakano Reiichiro","year":"2021","unstructured":"Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et\u00a0al. 2021. WebGPT: Browser-assisted question-answering with human feedback. arXiv preprint abs\/2112.09332 (2021).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_93_2","volume-title":"Proceedings of the 40th International Conference on Machine Learning (ICML\u201923)","author":"Ni Ansong","year":"2023","unstructured":"Ansong Ni, Srini Iyer, Dragomir Radev, Ves Stoyanov, Wen-Tau Yih, Sida I. Wang, and Xi Victoria Lin. 2023. LEVER: Learning to verify language-to-code generation with execution. In Proceedings of the 40th International Conference on Machine Learning (ICML\u201923)."},{"key":"e_1_3_3_94_2","article-title":"Skeleton-of-thought: Large language models can do parallel decoding","author":"Ning Xuefei","year":"2023","unstructured":"Xuefei Ning, Zinan Lin, Zixuan Zhou, Zifu Wang, Huazhong Yang, and Yu Wang. 2023. Skeleton-of-thought: Large language models can do parallel decoding. In Proceedings of the Efficient Natural Language and Speech Processing Workshop (ENLSP-III).","journal-title":"Proceedings of the Efficient Natural Language and Speech Processing Workshop (ENLSP-III)."},{"key":"e_1_3_3_95_2","volume-title":"Proceedings of the Deep Learning for Code Workshop","author":"Nye Maxwell","year":"2022","unstructured":"Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et\u00a0al. 2022. Show your work: Scratchpads for intermediate computation with language models. In Proceedings of the Deep Learning for Code Workshop."},{"key":"e_1_3_3_96_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Olausson Theo X.","year":"2024","unstructured":"Theo X. Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao, and Armando Solar-Lezama. 2024. Is self-repair a silver bullet for code generation? In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_97_2","doi-asserted-by":"publisher","DOI":"10.5555\/3692070.3693649"},{"key":"e_1_3_3_98_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00660"},{"key":"e_1_3_3_99_2","doi-asserted-by":"publisher","DOI":"10.1145\/3586183.3606763"},{"key":"e_1_3_3_100_2","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2080\u20132094.","author":"Patel Arkil","year":"2021","unstructured":"Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are NLP models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2080\u20132094."},{"key":"e_1_3_3_101_2","article-title":"Gorilla: Large language model connected with massive APIs","volume":"2305","author":"Patil Shishir G.","year":"2023","unstructured":"Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez. 2023. Gorilla: Large language model connected with massive APIs. arXiv preprint abs\/2305.15334 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_102_2","article-title":"Why think step by step? Reasoning emerges from the locality of experience","author":"Prystawski Ben","year":"2023","unstructured":"Ben Prystawski, Michael Li, and Noah Goodman. 2023. Why think step by step? Reasoning emerges from the locality of experience. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 70926\u201370947.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 70926\u201370947."},{"key":"e_1_3_3_103_2","article-title":"Investigate-Consolidate-Exploit: A general strategy for inter-task agent self-evolution","volume":"2401","author":"Qian Cheng","year":"2024","unstructured":"Cheng Qian, Shihao Liang, Yujia Qin, Yining Ye, Xin Cong, Yankai Lin, Yesai Wu, Zhiyuan Liu, and Maosong Sun. 2024. Investigate-Consolidate-Exploit: A general strategy for inter-task agent self-evolution. CoRR abs\/2401.13996 (2024).","journal-title":"CoRR"},{"key":"e_1_3_3_104_2","doi-asserted-by":"crossref","first-page":"15174","DOI":"10.18653\/v1\/2024.acl-long.810","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Qian Chen","year":"2024","unstructured":"Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, et\u00a0al. 2024. ChatDev: Communicative agents for software development. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 15174\u201315186."},{"key":"e_1_3_3_105_2","doi-asserted-by":"crossref","first-page":"5368","DOI":"10.18653\/v1\/2023.acl-long.294","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Qiao Shuofei","year":"2023","unstructured":"Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen. 2023. Reasoning with language model prompting: A survey. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 5368\u20135393."},{"key":"e_1_3_3_106_2","volume-title":"Proceedings of the ICLR 2024 Workshop on Large Language Model (LLM) Agents","author":"Qiao Shuofei","year":"2024","unstructured":"Shuofei Qiao, Ningyu Zhang, Runnan Fang, Yujie Luo, Wangchunshu Zhou, Yuchen Eleanor Jiang, Huajun Chen, et\u00a0al. 2024. AutoAct: Automatic agent learning from scratch for QA via self-planning. In Proceedings of the ICLR 2024 Workshop on Large Language Model (LLM) Agents."},{"key":"e_1_3_3_107_2","first-page":"1339","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201923)","author":"Qin Chengwei","year":"2023","unstructured":"Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang. 2023. Is ChatGPT a general-purpose natural language processing task solver? In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201923). 1339\u20131384."},{"key":"e_1_3_3_108_2","first-page":"2695","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201923)","author":"Qin Libo","year":"2023","unstructured":"Libo Qin, Qiguang Chen, Fuxuan Wei, Shijue Huang, and Wanxiang Che. 2023. Cross-lingual prompting: Improving zero-shot chain-of-thought reasoning across languages. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201923). 2695\u20132709."},{"key":"e_1_3_3_109_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Qin Yujia","year":"2024","unstructured":"Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et\u00a0al. 2024. ToolLLM: Facilitating large language models to master 16000+ real-world APIs. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_110_2","first-page":"8748","volume-title":"Proceedings of the 38th International Conference on Machine Learning (ICML\u201921)","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et\u00a0al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML\u201921). 8748\u20138763."},{"issue":"7371","key":"e_1_3_3_111_2","doi-asserted-by":"crossref","first-page":"113","DOI":"10.1038\/nature10514","article-title":"Verbal and non-verbal intelligence changes in the teenage brain","volume":"479","author":"Ramsden Sue","year":"2011","unstructured":"Sue Ramsden, Fiona M. Richardson, Goulven Josse, Michael S. C. Thomas, Caroline Ellis, Clare Shakeshaft, Mohamed L. Seghier, and Cathy J. Price. 2011. Verbal and non-verbal intelligence changes in the teenage brain. Nature 479, 7371 (2011), 113\u2013116.","journal-title":"Nature"},{"key":"e_1_3_3_112_2","article-title":"Android in the Wild: A large-scale dataset for Android device control","author":"Rawles Christopher","year":"2023","unstructured":"Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy Lillicrap. 2023. Android in the Wild: A large-scale dataset for Android device control. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 59708\u201359728.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 59708\u201359728."},{"key":"e_1_3_3_113_2","unstructured":"Reworkd. 2023. AgentGPT. Retrieved February 26 2025 from https:\/\/github.com\/reworkd\/AgentGPT"},{"key":"e_1_3_3_114_2","unstructured":"Toran Bruce Richards. 2023. Auto-GPT: An autonomous GPT-4 experiment. Retrieved February 26 2025 from https:\/\/github.com\/Significant-Gravitas\/Auto-GPT"},{"key":"e_1_3_3_115_2","article-title":"Visual chain of thought: Bridging logical gaps with multimodal infillings","volume":"2305","author":"Rose Daniel","year":"2023","unstructured":"Daniel Rose, Vaishnavi Himakunthala, Andy Ouyang, Ryan He, Alex Mei, Yujie Lu, Michael Saxon, Chinmay Sonar, Diba Mirza, and William Yang Wang. 2023. Visual chain of thought: Bridging logical gaps with multimodal infillings. arXiv preprint abs\/2305.02317 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_116_2","volume-title":"Proceedings of the NeurIPS 2023 Foundation Models for Decision Making Workshop","author":"Ruan Jingqing","year":"2023","unstructured":"Jingqing Ruan, Yihong Chen, Bin Zhang, Zhiwei Xu, Tianpeng Bao, Guoqing Du, Shiwei Shi, Hangyu Mao, Ziyue Li, Xingyu Zeng, et\u00a0al. 2023. TPTU: Task planning and tool usage of large language model-based AI agents. In Proceedings of the NeurIPS 2023 Foundation Models for Decision Making Workshop."},{"key":"e_1_3_3_117_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Ruan Yangjun","year":"2024","unstructured":"Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J. Maddison, and Tatsunori Hashimoto. 2024. Identifying the risks of LM agents with an LM-emulated sandbox. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_118_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)","author":"Rust Phillip","year":"2023","unstructured":"Phillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, and Desmond Elliott. 2023. Language modelling with pixels. In Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)."},{"key":"e_1_3_3_119_2","doi-asserted-by":"crossref","first-page":"19","DOI":"10.1007\/978-1-4615-1185-4_2","volume-title":"Understanding Psychological Assessment","author":"Ryan Joseph J.","year":"2001","unstructured":"Joseph J. Ryan and Shane J. Lopez. 2001. Wechsler Adult Intelligence Scale-III. In Understanding Psychological Assessment, William I. Dorfman and Michel Hersen (Eds.). Perspectives in Individual Differences. Springer, 19\u201342."},{"key":"e_1_3_3_120_2","article-title":"Toolformer: Language models can teach themselves to use tools","author":"Schick Timo","year":"2023","unstructured":"Timo Schick, Jane Dwivedi-Yu, Roberto Dess\u00ec, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 68539\u201368551.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 68539\u201368551."},{"key":"e_1_3_3_121_2","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9781139173438"},{"key":"e_1_3_3_122_2","first-page":"2683","volume-title":"Proceedings of the Conference on Robot Learning","author":"Shah Dhruv","year":"2023","unstructured":"Dhruv Shah, Michael Robert Equi, B\u0142a\u017cej Osi\u0144ski, Fei Xia, Brian Ichter, and Sergey Levine. 2023. Navigation with large language models: Semantic guesswork as a heuristic for planning. In Proceedings of the Conference on Robot Learning. 2683\u20132699."},{"key":"e_1_3_3_123_2","doi-asserted-by":"crossref","first-page":"4454","DOI":"10.18653\/v1\/2023.acl-long.244","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Shaikh Omar","year":"2023","unstructured":"Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang. 2023. On second thought, let\u2019s not think step by step! Bias and toxicity in zero-shot reasoning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 4454\u20134470."},{"key":"e_1_3_3_124_2","doi-asserted-by":"crossref","first-page":"13153","DOI":"10.18653\/v1\/2023.emnlp-main.814","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201923)","author":"Shao Yunfan","year":"2023","unstructured":"Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023. Character-LLM: A trainable agent for role-playing. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201923). 13153\u201313187."},{"key":"e_1_3_3_125_2","article-title":"HuggingGPT: Solving AI tasks with ChatGPT and its friends in Hugging Face","author":"Shen Yongliang","year":"2023","unstructured":"Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023. HuggingGPT: Solving AI tasks with ChatGPT and its friends in Hugging Face. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 38154\u201338180.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 38154\u201338180."},{"key":"e_1_3_3_126_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)","author":"Shi Freda","year":"2023","unstructured":"Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, et\u00a0al. 2023. Language models are multilingual chain-of-thought reasoners. In Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)."},{"key":"e_1_3_3_127_2","article-title":"Reflexion: Language agents with verbal reinforcement learning","author":"Shinn Noah","year":"2023","unstructured":"Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 8634\u20138652.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 8634\u20138652."},{"key":"e_1_3_3_128_2","doi-asserted-by":"crossref","first-page":"12113","DOI":"10.18653\/v1\/2023.findings-emnlp.811","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Shum Kashun","year":"2023","unstructured":"Kashun Shum, Shizhe Diao, and Tong Zhang. 2023. Automatic prompt augmentation and selection with chain-of-thought from labeled data. In Findings of the Association for Computational Linguistics: EMNLP 2023. 12113\u201312139."},{"key":"e_1_3_3_129_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-023-06291-2"},{"key":"e_1_3_3_130_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2408.03314"},{"key":"e_1_3_3_131_2","article-title":"Beyond the imitation game: Quantifying and extrapolating the capabilities of language models","author":"Srivastava Aarohi","year":"2023","unstructured":"Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md. Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adri\u00e0 Garriga-Alonso, et\u00a0al. 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research. Published Online, May 11, 2023.","journal-title":"Transactions on Machine Learning Research."},{"key":"e_1_3_3_132_2","volume-title":"Proceedings of the NeurIPS 2023 Foundation Models for Decision Making Workshop","author":"Stechly Kaya","year":"2023","unstructured":"Kaya Stechly, Matthew Marquez, and Subbarao Kambhampati. 2023. GPT-4 doesn\u2019t know it\u2019s wrong: An analysis of iterative prompting for reasoning problems. In Proceedings of the NeurIPS 2023 Foundation Models for Decision Making Workshop."},{"issue":"2","key":"e_1_3_3_133_2","doi-asserted-by":"crossref","first-page":"155","DOI":"10.1016\/0304-422X(82)90031-6","article-title":"The nature of verbal comprehension","volume":"11","author":"Sternberg Robert J.","year":"1982","unstructured":"Robert J. Sternberg, Janet S. Powell, and Daniel B. Kaye. 1982. The nature of verbal comprehension. Poetics 11, 2 (1982), 155\u2013187.","journal-title":"Poetics"},{"key":"e_1_3_3_134_2","article-title":"Cognitive architectures for language agents","author":"Sumers Theodore","year":"2024","unstructured":"Theodore Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas Griffiths. 2024. Cognitive architectures for language agents. Transactions on Machine Learning Research. Published Online, February 22, 2024.","journal-title":"Transactions on Machine Learning Research."},{"key":"e_1_3_3_135_2","first-page":"5636","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics (COLING\u201922)","author":"Sunkara Srinivas","year":"2022","unstructured":"Srinivas Sunkara, Maria Wang, Lijuan Liu, Gilles Baechler, Yu-Chung Hsiao, Jindong Chen, Abhanshu Sharma, and James W. W. Stout. 2022. Towards better semantic understanding of mobile interfaces. In Proceedings of the 29th International Conference on Computational Linguistics (COLING\u201922). 5636\u20135650."},{"key":"e_1_3_3_136_2","first-page":"11888","volume-title":"Proceedings of the 2023 IEEE\/CVF International Conference on Computer Vision (ICCV\u201923)","author":"Sur\u00eds D\u00eddac","year":"2023","unstructured":"D\u00eddac Sur\u00eds, Sachit Menon, and Carl Vondrick. 2023. ViperGPT: Visual inference via Python execution for reasoning. In Proceedings of the 2023 IEEE\/CVF International Conference on Computer Vision (ICCV\u201923). 11888\u201311898."},{"key":"e_1_3_3_137_2","first-page":"4149","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long and Short Papers.","author":"Talmor Alon","year":"2019","unstructured":"Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long and Short Papers.4149\u20134158."},{"key":"e_1_3_3_138_2","volume-title":"Proceedings of the 10th International Conference on Learning Representations (ICLR\u201922)","author":"Tay Yi","year":"2022","unstructured":"Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Steven Zheng, et\u00a0al. 2022. UL2: Unifying language learning paradigms. In Proceedings of the 10th International Conference on Learning Representations (ICLR\u201922)."},{"key":"e_1_3_3_139_2","article-title":"XAgent: An autonomous agent for complex task solving","author":"Team XAgent","year":"2023","unstructured":"XAgent Team. 2023. XAgent: An autonomous agent for complex task solving. XAgent Blog.","journal-title":"XAgent Blog."},{"key":"e_1_3_3_140_2","article-title":"LaMDA: Language models for dialog applications","volume":"2201","author":"Thoppilan Romal","year":"2022","unstructured":"Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et\u00a0al. 2022. LaMDA: Language models for dialog applications. arXiv preprint abs\/2201.08239 (2022).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_141_2","doi-asserted-by":"crossref","first-page":"10014","DOI":"10.18653\/v1\/2023.acl-long.557","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Trivedi Harsh","year":"2023","unstructured":"Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2023. Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 10014\u201310037."},{"key":"e_1_3_3_142_2","volume-title":"Proceedings of the NeurIPS 2023 Foundation Models for Decision Making Workshop","author":"Valmeekam Karthik","year":"2023","unstructured":"Karthik Valmeekam, Matthew Marquez, and Subbarao Kambhampati. 2023. Investigating the effectiveness of self-critiquing in LLMs solving planning tasks. In Proceedings of the NeurIPS 2023 Foundation Models for Decision Making Workshop."},{"key":"e_1_3_3_143_2","doi-asserted-by":"crossref","first-page":"2717","DOI":"10.18653\/v1\/2023.acl-long.153","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wang Boshi","year":"2023","unstructured":"Boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, and Huan Sun. 2023. Towards understanding chain-of-thought prompting: An empirical study of what matters. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2717\u20132739."},{"key":"e_1_3_3_144_2","article-title":"Voyager: An open-ended embodied agent with large language models","author":"Wang Guanzhi","year":"2024","unstructured":"Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2024. Voyager: An open-ended embodied agent with large language models. Transactions on Machine Learning Research. Published Online, March 19, 2024.","journal-title":"Transactions on Machine Learning Research."},{"key":"e_1_3_3_145_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2410.09671"},{"issue":"6","key":"e_1_3_3_146_2","doi-asserted-by":"crossref","first-page":"186345","DOI":"10.1007\/s11704-024-40231-1","article-title":"A survey on large language model based autonomous agents","volume":"18","author":"Wang Lei","year":"2024","unstructured":"Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et\u00a0al. 2024. A survey on large language model based autonomous agents. Frontiers of Computer Science 18, 6 (2024), 186345.","journal-title":"Frontiers of Computer Science"},{"key":"e_1_3_3_147_2","doi-asserted-by":"crossref","first-page":"2609","DOI":"10.18653\/v1\/2023.acl-long.147","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wang Lei","year":"2023","unstructured":"Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023. Plan-and-Solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2609\u20132634."},{"key":"e_1_3_3_148_2","article-title":"User behavior simulation with large language model based agents","author":"Wang Lei","year":"2023","unstructured":"Lei Wang, Jingsen Zhang, Hao Yang, Zhiyuan Chen, Jiakai Tang, Zeyu Zhang, Xu Chen, Yankai Lin, Ruihua Song, Wayne Xin Zhao, et\u00a0al. 2023. User behavior simulation with large language model based agents. arXiv preprint arXiv:2306.02552 (2023).","journal-title":"arXiv preprint arXiv:2306.02552"},{"key":"e_1_3_3_149_2","volume-title":"Proceedings of the Workshop on Efficient Systems for Foundation Models (ICML\u201923)","author":"Wang Xinyi","year":"2023","unstructured":"Xinyi Wang and William Yang Wang. 2023. Reasoning ability emerges in large language models as aggregation of reasoning paths: A case study with knowledge graphs. In Proceedings of the Workshop on Efficient Systems for Foundation Models (ICML\u201923)."},{"key":"e_1_3_3_150_2","article-title":"Rationale-augmented ensembles in language models","volume":"2207","author":"Wang Xuezhi","year":"2022","unstructured":"Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022. Rationale-augmented ensembles in language models. arXiv preprint abs\/2207.00747 (2022).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_151_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)","author":"Wang Xuezhi","year":"2023","unstructured":"Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)."},{"key":"e_1_3_3_152_2","doi-asserted-by":"crossref","first-page":"8640","DOI":"10.18653\/v1\/2023.acl-long.482","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wang Yiming","year":"2023","unstructured":"Yiming Wang, Zhuosheng Zhang, and Rui Wang. 2023. Element-aware summarization with large language models: Expert-aligned evaluation and chain-of-thought method. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 8640\u20138665."},{"key":"e_1_3_3_153_2","article-title":"Jailbroken: How does LLM safety training fail?","author":"Wei Alexander","year":"2023","unstructured":"Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How does LLM safety training fail? In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 80079\u201380110.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 80079\u201380110."},{"key":"e_1_3_3_154_2","article-title":"Emergent abilities of large language models","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et\u00a0al. 2022. Emergent abilities of large language models. Transactions on Machine Learning Research. Published Online, August 31, 2022.","journal-title":"Transactions on Machine Learning Research."},{"key":"e_1_3_3_155_2","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS\u201922).24824\u201324837.","journal-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS\u201922)."},{"key":"e_1_3_3_156_2","unstructured":"Lilian Weng. 2023. LLM Powered Autonomous Agents. Retrieved February 26 2025 from https:\/\/lilianweng.github.io"},{"key":"e_1_3_3_157_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.167"},{"key":"e_1_3_3_158_2","volume-title":"Practical Planning: Extending the Classical AI Planning Paradigm","author":"Wilkins David E.","year":"2014","unstructured":"David E. Wilkins. 2014. Practical Planning: Extending the Classical AI Planning Paradigm. Elsevier."},{"key":"e_1_3_3_159_2","doi-asserted-by":"publisher","DOI":"10.1017\/S0269888900008122"},{"key":"e_1_3_3_160_2","article-title":"Scaling virtual world with delta-engine","author":"Wu Hongqiu","year":"2024","unstructured":"Hongqiu Wu, Zekai Xu, Tianyang Xu, Jiale Hong, Weiqi Wu, Hai Zhao, Min Zhang, and Zhezhi He. 2024. Scaling virtual world with delta-engine. arXiv preprint arXiv:2408.05842 (2024).","journal-title":"arXiv preprint arXiv:2408.05842"},{"key":"e_1_3_3_161_2","volume-title":"Proceedings of the 41st International Conference on Machine Learning (ICML\u201924)","author":"Wu Shengqiong","year":"2024","unstructured":"Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. 2024. NExT-GPT: Any-to-any multimodal LLM. In Proceedings of the 41st International Conference on Machine Learning (ICML\u201924)."},{"key":"e_1_3_3_162_2","first-page":"3271","volume-title":"Findings of the Association for Computational Linguistics: ACL 2024","author":"Wu Weiqi","year":"2024","unstructured":"Weiqi Wu, Hongqiu Wu, Lai Jiang, Xingyuan Liu, Hai Zhao, and Min Zhang. 2024. From role-play to drama-interaction: An LLM solution. In Findings of the Association for Computational Linguistics: ACL 2024. 3271\u20133290."},{"key":"e_1_3_3_163_2","article-title":"The rise and potential of large language model based agents: A survey","volume":"2309","author":"Xi Zhiheng","year":"2023","unstructured":"Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et\u00a0al. 2023. The rise and potential of large language model based agents: A survey. arXiv preprint abs\/2309.07864 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_164_2","volume-title":"Proceedings of the 10th International Conference on Learning Representations (ICLR\u201922)","author":"Xie Sang Michael","year":"2022","unstructured":"Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2022. An explanation of in-context learning as implicit Bayesian inference. In Proceedings of the 10th International Conference on Learning Representations (ICLR\u201922)."},{"key":"e_1_3_3_165_2","doi-asserted-by":"crossref","first-page":"7572","DOI":"10.18653\/v1\/2023.findings-emnlp.508","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Xiong Kai","year":"2023","unstructured":"Kai Xiong, Xiao Ding, Yixin Cao, Ting Liu, and Bing Qin. 2023. Examining inter-consistency of large language models collaboration: An in-depth analysis via debate. In Findings of the Association for Computational Linguistics: EMNLP 2023. 7572\u20137590."},{"key":"e_1_3_3_166_2","doi-asserted-by":"crossref","first-page":"4643","DOI":"10.18653\/v1\/2024.naacl-long.260","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)","author":"Xiong Wenhan","year":"2024","unstructured":"Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, et\u00a0al. 2024. Effective long-context scaling of foundation models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 4643\u20134663."},{"key":"e_1_3_3_167_2","article-title":"ReWOO: Decoupling reasoning from observations for efficient augmented language models","volume":"2305","author":"Xu Binfeng","year":"2023","unstructured":"Binfeng Xu, Zhiyuan Peng, Bowen Lei, Subhabrata Mukherjee, Yuchen Liu, and Dongkuan Xu. 2023. ReWOO: Decoupling reasoning from observations for efficient augmented language models. arXiv preprint abs\/2305.18323 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_168_2","article-title":"ExpertPrompting: Instructing large language models to be distinguished experts","volume":"2305","author":"Xu Benfeng","year":"2023","unstructured":"Benfeng Xu, An Yang, Junyang Lin, Quan Wang, Chang Zhou, Yongdong Zhang, and Zhendong Mao. 2023. ExpertPrompting: Instructing large language models to be distinguished experts. arXiv preprint abs\/2305.14688 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_169_2","article-title":"SC-Safety: A multi-round open-ended question adversarial safety benchmark for large language models in Chinese","volume":"2310","author":"Xu Liang","year":"2023","unstructured":"Liang Xu, Kangkang Zhao, Lei Zhu, and Hang Xue. 2023. SC-Safety: A multi-round open-ended question adversarial safety benchmark for large language models in Chinese. arXiv preprint abs\/2310.05818 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_170_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Xu Yiheng","year":"2024","unstructured":"Yiheng Xu, Hongjin Su, Chen Xing, Boyu Mi, Qian Liu, Weijia Shi, Binyuan Hui, Fan Zhou, Yitao Liu, Tianbao Xie, et\u00a0al. 2024. Lemur: Harmonizing natural language and code for language agents. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_171_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Yang Chengrun","year":"2024","unstructured":"Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou, and Xinyun Chen. 2024. Large language models as optimizers. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_172_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Yang Sherry","year":"2024","unstructured":"Sherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson, Leslie Pack Kaelbling, Dale Schuurmans, and Pieter Abbeel. 2024. Learning interactive real-world simulators. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_173_2","article-title":"Learn to interpret Atari agents","volume":"1812","author":"Yang Zhao","year":"2018","unstructured":"Zhao Yang, Song Bai, Li Zhang, and Philip H. S. Torr. 2018. Learn to interpret Atari agents. arXiv preprint abs\/1812.11276 (2018).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_174_2","article-title":"Tree of thoughts: Deliberate problem solving with large language models","author":"Yao Shunyu","year":"2023","unstructured":"Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 11809\u201311822.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 11809\u201311822."},{"key":"e_1_3_3_175_2","volume-title":"Proceedings of the 10th International Conference on Learning Representations (ICLR\u201922)","author":"Yao Shunyu","year":"2022","unstructured":"Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2022. ReAct: Synergizing reasoning and acting in language models. In Proceedings of the 10th International Conference on Learning Representations (ICLR\u201922)."},{"key":"e_1_3_3_176_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Yao Weiran","year":"2024","unstructured":"Weiran Yao, Shelby Heinecke, Juan Carlos Niebles, Zhiwei Liu, Yihao Feng, Le Xue, Rithesh Murthy, Zeyuan Chen, Jianguo Zhang, Devansh Arpit, et\u00a0al. 2024. Retroformer: Retrospective large language agents with policy gradient optimization. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_177_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.hcc.2024.100211"},{"key":"e_1_3_3_178_2","doi-asserted-by":"crossref","first-page":"2901","DOI":"10.18653\/v1\/2024.findings-naacl.183","volume-title":"Findings of the Association for Computational Linguistics: NAACL 2024","author":"Yao Yao","year":"2024","unstructured":"Yao Yao, Zuchao Li, and Hai Zhao. 2024. GoT: Effective graph-of-thought reasoning in language models. In Findings of the Association for Computational Linguistics: NAACL 2024. 2901\u20132921."},{"key":"e_1_3_3_179_2","volume-title":"Proceedings of the ICLR 2024 Workshop on Large Language Model (LLM) Agents","author":"Yin Da","year":"2024","unstructured":"Da Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Chandu, Kai-Wei Chang, Yejin Choi, and Bill Yuchen Lin. 2024. Lumos: Learning agents with unified data, modular design, and open-source LLMs. In Proceedings of the ICLR 2024 Workshop on Large Language Model (LLM) Agents."},{"key":"e_1_3_3_180_2","article-title":"Natural language reasoning, a survey","author":"Yu Fei","year":"2024","unstructured":"Fei Yu, Hongbo Zhang, Prayag Tiwari, and Benyou Wang. 2024. Natural language reasoning, a survey. ACM Computing Surveys 56, 12 (2024), Article 304, 39 pages.","journal-title":"ACM Computing Surveys"},{"key":"e_1_3_3_181_2","article-title":"Towards better chain-of-thought prompting strategies: A survey","volume":"2310","author":"Yu Zihan","year":"2023","unstructured":"Zihan Yu, Liang He, Zhen Wu, Xinyu Dai, and Jiajun Chen. 2023. Towards better chain-of-thought prompting strategies: A survey. arXiv preprint abs\/2310.04959 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_182_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Yuan Lifan","year":"2024","unstructured":"Lifan Yuan, Yangyi Chen, Xingyao Wang, Yi Fung, Hao Peng, and Heng Ji. 2024. CRAFT: Customizing LLMs by creating and retrieving from specialized toolsets. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_183_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Yue Xiang","year":"2024","unstructured":"Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2024. MAmmoTH: Building math generalist models through hybrid instruction tuning. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_184_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2410.04343"},{"key":"e_1_3_3_185_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Zhang Renrui","year":"2024","unstructured":"Renrui Zhang, Jiaming Han, Chris Liu, Aojun Zhou, Pan Lu, Yu Qiao, Hongsheng Li, and Peng Gao. 2024. LLaMA-Adapter: Efficient fine-tuning of large language models with zero-initialized attention. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_186_2","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445186"},{"key":"e_1_3_3_187_2","doi-asserted-by":"crossref","first-page":"15537","DOI":"10.18653\/v1\/2024.acl-long.830","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhang Zhexin","year":"2024","unstructured":"Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2024. SafetyBench: Evaluating the safety of large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 15537\u201315553."},{"key":"e_1_3_3_188_2","volume-title":"Findings of the Association for Computational Linguistics: ACL 2024","author":"Zhang Zhuosheng","year":"2024","unstructured":"Zhuosheng Zhang and Aston Zhang. 2024. You only look at screens: Multimodal chain-of-action agents. In Findings of the Association for Computational Linguistics: ACL 2024."},{"key":"e_1_3_3_189_2","article-title":"Multimodal chain-of-thought reasoning in language models","author":"Zhang Zhuosheng","year":"2024","unstructured":"Zhuosheng Zhang, Aston Zhang, Mu Li, hai zhao, George Karypis, and Alex Smola. 2024. Multimodal chain-of-thought reasoning in language models. Transactions on Machine Learning Research. Published Online, June 3, 2024.","journal-title":"Transactions on Machine Learning Research."},{"key":"e_1_3_3_190_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)","author":"Zhang Zhuosheng","year":"2023","unstructured":"Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2023. Automatic chain of thought prompting in large language models. In Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)."},{"key":"e_1_3_3_191_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Zhao Haozhe","year":"2024","unstructured":"Haozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma, Kaikai An, Liang Chen, Zixuan Liu, Sheng Wang, Wenjuan Han, and Baobao Chang. 2024. MMICL: Empowering vision-language model with multi-modal in-context learning. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_192_2","doi-asserted-by":"crossref","first-page":"5823","DOI":"10.18653\/v1\/2023.acl-long.320","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhao Ruochen","year":"2023","unstructured":"Ruochen Zhao, Xingxuan Li, Shafiq Joty, Chengwei Qin, and Lidong Bing. 2023. Verify-and-Edit: A knowledge-enhanced chain-of-thought framework. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 5823\u20135840."},{"key":"e_1_3_3_193_2","article-title":"Large language models as commonsense knowledge for large-scale task planning","author":"Zhao Zirui","year":"2023","unstructured":"Zirui Zhao, Wee Sun Lee, and David Hsu. 2023. Large language models as commonsense knowledge for large-scale task planning. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 31967\u201331987.","journal-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS\u201923). 31967\u201331987."},{"key":"e_1_3_3_194_2","doi-asserted-by":"crossref","first-page":"2299","DOI":"10.18653\/v1\/2024.findings-naacl.149","volume-title":"Findings of the Association for Computational Linguistics: NAACL 2024","author":"Zhong Wanjun","year":"2024","unstructured":"Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2024. AGIEval: A human-centric benchmark for evaluating foundation models. In Findings of the Association for Computational Linguistics: NAACL 2024. 2299\u20132314."},{"key":"e_1_3_3_195_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Zhou Aojun","year":"2024","unstructured":"Aojun Zhou, Ke Wang, Zimu Lu, Weikang Shi, Sichun Luo, Zipeng Qin, Shaoqing Lu, Anya Jia, Linqi Song, Mingjie Zhan, et\u00a0al. 2024. Solving challenging math word problems using GPT-4 code interpreter with code-based self-verification. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_196_2","doi-asserted-by":"publisher","DOI":"10.5555\/3692070.3694642"},{"key":"e_1_3_3_197_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)","author":"Zhou Shuyan","year":"2024","unstructured":"Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et\u00a0al. 2024. WebArena: A realistic web environment for building autonomous agents. In Proceedings of the 12th International Conference on Learning Representations (ICLR\u201924)."},{"key":"e_1_3_3_198_2","volume-title":"Proceedings of the ICLR 2024 Workshop on Large Language Model (LLM) Agents","author":"Zhou Wangchunshu","year":"2024","unstructured":"Wangchunshu Zhou, Yuchen Eleanor Jiang, Long Li, Jialong Wu, Tiannan Wang, Shuai Wang, Jiamin Chen, Jintian Zhang, Jing Chen, Xiangru Tang, et\u00a0al. 2024. Agents: An open-source framework for autonomous language agents. In Proceedings of the ICLR 2024 Workshop on Large Language Model (LLM) Agents."},{"key":"e_1_3_3_199_2","article-title":"LLM as DBA","volume":"2308","author":"Zhou Xuanhe","year":"2023","unstructured":"Xuanhe Zhou, Guoliang Li, and Zhiyuan Liu. 2023. LLM as DBA. arXiv preprint abs\/2308.05481 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_200_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)","author":"Zhou Yongchao","year":"2023","unstructured":"Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2023. Large language models are human-level prompt engineers. In Proceedings of the 11th International Conference on Learning Representations (ICLR\u201923)."},{"key":"e_1_3_3_201_2","article-title":"Ghost in the Minecraft: Generally capable agents for open-world environments via large language models with text-based knowledge and memory","volume":"2305","author":"Zhu Xizhou","year":"2023","unstructured":"Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, et\u00a0al. 2023. Ghost in the Minecraft: Generally capable agents for open-world environments via large language models with text-based knowledge and memory. arXiv preprint abs\/2305.17144 (2023).","journal-title":"arXiv preprint"},{"key":"e_1_3_3_202_2","doi-asserted-by":"crossref","first-page":"10259","DOI":"10.18653\/v1\/2023.findings-acl.651","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Ziqi Jin","year":"2023","unstructured":"Jin Ziqi and Wei Lu. 2023. Tab-CoT: Zero-shot tabular chain of thought. In Findings of the Association for Computational Linguistics: ACL 2023. 10259\u201310277."},{"key":"e_1_3_3_203_2","first-page":"1801","volume-title":"Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources, and Evaluation (LREC-COLING\u201924)","author":"Zou Anni","year":"2024","unstructured":"Anni Zou, Zhuosheng Zhang, and Hai Zhao. 2024. AuRoRA: A one-for-all platform for augmented reasoning and refining with task-adaptive chain-of-thought prompting. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources, and Evaluation (LREC-COLING\u201924). 1801\u20131807."}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3719341","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3719341","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T18:43:21Z","timestamp":1750272201000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3719341"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,22]]},"references-count":202,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2025,8,31]]}},"alternative-id":["10.1145\/3719341"],"URL":"https:\/\/doi.org\/10.1145\/3719341","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,22]]},"assertion":[{"value":"2023-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-14","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-22","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}