{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,3]],"date-time":"2026-08-03T07:28:07Z","timestamp":1785742087163,"version":"3.56.0"},"reference-count":146,"publisher":"American Association for the Advancement of Science (AAAS)","content-domain":{"domain":["spj.science.org"],"crossmark-restriction":true},"short-container-title":["Intell Comput"],"published-print":{"date-parts":[[2025,1]]},"abstract":"<jats:p>\n            This paper presents a comprehensive survey of the current status and opportunities for large language models (LLMs) in task planning, a sophisticated process of reasoning and decision-making that organizes a sequence of actions to accomplish predefined goals. Task planning centers on identifying suitable solutions for a task, with a specific emphasis on ensuring its successful completion. Although task planning plays a crucial role in enabling systems to function effectively in dynamic and complex environments, there is a lack of systematic reviews on this topic. We explore the theories, methodologies, and applications related to task planning with LLMs, highlighting the burgeoning development in this field and the interdisciplinary approaches that enhance their ability to complete tasks. This survey aims to systematize and clarify the fragmented literature, provide a systematic review that underscores the importance of task planning as a critical capability, and offer insights into future research directions and potential improvements. We hope to help researchers gain a clear understanding of the field and spark greater interest in this highly impactful research direction. A continuously updated resource is available in our GitHub repository at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/ZhaiWenShuo\/Survey-of-Task-Planning\">https:\/\/github.com\/ZhaiWenShuo\/Survey-of-Task-Planning<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.34133\/icomputing.0124","type":"journal-article","created":{"date-parts":[[2025,5,1]],"date-time":"2025-05-01T01:21:44Z","timestamp":1746062504000},"update-policy":"https:\/\/doi.org\/10.34133\/aaas_crossmark_01","source":"Crossref","is-referenced-by-count":9,"title":["A Survey of Task Planning with Large Language Models"],"prefix":"10.34133","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-4735-6985","authenticated-orcid":true,"given":"Wenshuo","family":"Zhai","sequence":"first","affiliation":[{"name":"Laboratory for Big Data and Decision, \rNational University of Defense Technology, Changsha, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinzhi","family":"Liao","sequence":"additional","affiliation":[{"name":"Laboratory for Big Data and Decision, \rNational University of Defense Technology, Changsha, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1714-0304","authenticated-orcid":false,"given":"Ziyang","family":"Chen","sequence":"additional","affiliation":[{"name":"Laboratory for Big Data and Decision, \rNational University of Defense Technology, Changsha, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-8317-1766","authenticated-orcid":false,"given":"Bolun","family":"Su","sequence":"additional","affiliation":[{"name":"Laboratory for Big Data and Decision, \rNational University of Defense Technology, Changsha, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6339-0219","authenticated-orcid":true,"given":"Xiang","family":"Zhao","sequence":"additional","affiliation":[{"name":"Laboratory for Big Data and Decision, \rNational University of Defense Technology, Changsha, China."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"221","published-online":{"date-parts":[[2025,5,23]]},"reference":[{"issue":"1","key":"e_1_3_2_2_2","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1016\/0004-3702(94)90081-7","article-title":"The computational complexity of propositional STRIPS planning","volume":"69","author":"Bylander T","year":"1994","unstructured":"Bylander T. The computational complexity of propositional STRIPS planning. Artif Intell. 1994;69(1-2):165\u2013204.","journal-title":"Artif Intell"},{"key":"e_1_3_2_3_2","doi-asserted-by":"crossref","unstructured":"Ghallab M Nau DS Traverso P. Automated planning\u2014Theory and practice. San Francisco (CA): Morgan Kaufmann Publishers; 2004.","DOI":"10.1016\/B978-155860856-6\/50021-1"},{"key":"e_1_3_2_4_2","doi-asserted-by":"crossref","unstructured":"LaValle SM. Planning algorithms. Cambridge (UK): Cambridge University Press; 2006.","DOI":"10.1017\/CBO9780511546877"},{"key":"e_1_3_2_5_2","doi-asserted-by":"crossref","unstructured":"Ghallab M Nau DS Traverso P. Automated planning and acting. Cambridge (UK): Cambridge University Press; 2016.","DOI":"10.1017\/CBO9781139583923"},{"key":"e_1_3_2_6_2","unstructured":"Wang X Li C Wang Z Bai F Luo H Zhang J Jojic N Xing E Hu Z. PromptAgent: Strategic planning with language models enables expert-level prompt optimization. Paper presented at: The Twelfth International Conference on Learning Representations ICLR 2024; 2024 May 7\u201311; Vienna Austria."},{"key":"e_1_3_2_7_2","doi-asserted-by":"crossref","unstructured":"Zhao Z Cheng S Ding Y Zhou Z Zhang S Xu D Zhao Y. A survey of optimization-based task and motion planning: From classical to learning approaches. IEEE\/ASME Trans. Mechatron. 2024.","DOI":"10.1109\/TMECH.2024.3452509"},{"key":"e_1_3_2_8_2","unstructured":"Zhao Z Lee WS Hsu D. Large language models as commonsense knowledge for largescale task planning. Paper presented at: Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023 NeurIPS 2023; 2023 Dec 10\u201316; New Orleans LA USA."},{"issue":"1","key":"e_1_3_2_9_2","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1146\/annurev-control-091420-084139","article-title":"Integrated task and motion planning","volume":"4","author":"Garrett CR","year":"2021","unstructured":"Garrett CR, Chitnis R, Holladay RM, Kim B, Silver T, Kaelbling LP, Lozano-P\u00e9rez T. Integrated task and motion planning. Annu Rev Control Robot Auton Syst. 2021;4(1):265\u2013293.","journal-title":"Annu Rev Control Robot Auton Syst"},{"key":"e_1_3_2_10_2","unstructured":"Rana K Haviland J Garg S Abou-Chakra J Reid ID S\u00fcnderhauf N. SayPlan: Grounding large language models using 3D scene graphs for scalable robot task planning. Paper presented at: Conference on Robot Learning CoRL 2023; 2023 Nov 6\u20139; Atlanta GA USA."},{"issue":"213","key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3564696","article-title":"Multiple mobile robot task and motion planning: A survey","volume":"55","author":"Antonyshyn L","year":"2023","unstructured":"Antonyshyn L, Silveira J, Givigi S, Marshall J. Multiple mobile robot task and motion planning: A survey. ACM Comput Surv. 2023;55(213):1\u201335.","journal-title":"ACM Comput Surv"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1146\/annurev-control-082619-100135"},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","DOI":"10.1007\/s11704-024-40231-1","article-title":"A survey on large language model based autonomous agents","volume":"18","author":"Wang L","year":"2024","unstructured":"Wang L, Ma C, Feng X, Zhang Z, Yang H, Zhang J, Chen Z, Tang J, Chen X, Lin Y, et al. A survey on large language model based autonomous agents. Front Comput Sci. 2024;18: Article 186345.","journal-title":"Front Comput Sci"},{"key":"e_1_3_2_14_2","unstructured":"Cheng Y Zhang C Zhang Z Meng X Hong S Li W Wang Z Wang Z Yin F Zhao J et\u00a0al. Exploring large language model based intelligent agents: Definitions methods and prospects. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2401.03428"},{"key":"e_1_3_2_15_2","unstructured":"Xu X Wang Y Xu C Ding Z Jiang J Ding Z Karlsson BF. A survey on game playing agents and large models: Methods applications and challenges. CoRR. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv. 2403.10249"},{"key":"e_1_3_2_16_2","unstructured":"Huang X Liu W Chen X Wang X Hang H Lian D Wang Y Tang R Chen E et\u00a0al. Understanding the planning of LLM agents: A survey. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2402.02716"},{"key":"e_1_3_2_17_2","unstructured":"Zhang Y Mao S Ge T Wang X de Wynter A Xia Y Wu W Song T Lan M Wei F et\u00a0al. LLM as a mastermind: A survey of strategic reasoning with large language models. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2404.01230"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","unstructured":"Pallagani V Muppasani BC Roy K Fabiano F Loreggia A Murugesan K Srivastava B Rossi F Horesh L Sheth AP. On the prospects of incorporating large language models (LLMs) in automated planning and scheduling (APS). Paper presented at: Proceedings of the Thirty-Fourth International Conference on Automated Planning and Scheduling ICAPS 2024; 2024 Jun 1\u20136; Banff Alberta Canada.","DOI":"10.1609\/icaps.v34i1.31503"},{"key":"e_1_3_2_19_2","unstructured":"Wei J Wang X Schuurmans D Bosma M Ichter B Xia F Chi E Le Q Zhou D. Chain-of-thought prompting elicits reasoning in large language models. Paper presented at: Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022 NeurIPS 2022; 2022 Nov 28\u2013Dec 9; New Orleans LA USA."},{"key":"e_1_3_2_20_2","unstructured":"Fu Y Peng H Sabharwal A Clark P Khot T. Complexity-based prompting for multistep reasoning. Paper presented at: The Eleventh International Conference on Learning Representations ICLR 2023; 2023 May 1\u20135; Kigali Rwanda."},{"key":"e_1_3_2_21_2","unstructured":"Kojima T Gu SS Reid M Matsuo Y Iwasawa Y. Large language models are zero-shot reasoners. Paper presented at: Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022 NeurIPS 2022; 2022 Nov 28\u2013Dec 9; New Orleans LA USA."},{"key":"e_1_3_2_22_2","unstructured":"Zhang Z Zhang A Li M Smola A. Automatic chain of thought prompting in large language models. Paper presented at: The Eleventh International Conference on Learning Representations ICLR 2023; 2023 May 1\u20135; Kigali Rwanda."},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","unstructured":"Wang L Xu W Lan Y Hu Z Lan Y Lee RKW Lim EP. Plan-and-Solve Prompting: Improving zero-shot chain-of-thought reasoning by large language models. Paper presented at: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) ACL 2023; 2023 Jul 9\u201314; Toronto Canada.","DOI":"10.18653\/v1\/2023.acl-long.147"},{"key":"e_1_3_2_24_2","unstructured":"Yao S Yu D Zhao J Shafran I Griffiths T Cao Y Narasimhan K. Tree of thoughts: Deliberate problem solving with large language models. Paper presented at: Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023 NeurIPS 2023; 2023 Dec 10\u201316; New Orleans LA USA."},{"issue":"16","key":"e_1_3_2_25_2","first-page":"17682","article-title":"Graph of thoughts: Solving elaborate problems with large language models","volume":"38","author":"Besta M","year":"2024","unstructured":"Besta M, Blach N, Kubicek A, Gerstenberger R, Podstawski M, Gianinazzi L, Gajda J, Lehmann T, Niewiadomski H, Nyczyk P, et al. Graph of thoughts: Solving elaborate problems with large language models. AAAI Conf Artif Intell. 2024;38(16):17682\u201317690.","journal-title":"AAAI Conf Artif Intell"},{"key":"e_1_3_2_26_2","unstructured":"Yao S Zhao J Yu D Du N Shafran I Narasimhan KR Cao Y. ReAct: Synergizing reasoning and acting in language models. Paper presented at: The Eleventh International Conference on Learning Representations ICLR 2023; 2023 May 1\u20135; Kigali Rwanda."},{"key":"e_1_3_2_27_2","unstructured":"Qin Y Liang S Ye Y Zhu K Yan L Lu Y Lin Y Cong X Tang X Qian B et\u00a0al. ToolLLM: Facilitating large language models to master 16000+ real-world APIs. Paper presented at: The Twelfth International Conference on Learning Representations ICLR 2024; 2024 May 7\u201311; Vienna Austria. OpenReview.net 2024. url: https:\/\/openreview.net\/forum?id=dHng2O0Jjr"},{"key":"e_1_3_2_28_2","unstructured":"Ye Y Cong X Tian S Qin Y Liu C Lin Y Liu Z Sun M. Rational decision-making agent with internalized utility judgment. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2308.12519"},{"key":"e_1_3_2_29_2","unstructured":"Hou X Yang M Jiao W Wang X Tu Z Zhao WX. CoAct: A global-local hierarchy for autonomous agent collaboration. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2406.13381"},{"key":"e_1_3_2_30_2","unstructured":"Zhou D Sch\u00e4rli N Hou L Wei J Scales N Wang X Schuurmans D Cui C Bousquet O Le QV et\u00a0al. Least-to-most prompting enables complex reasoning in large language models. Paper presented at: The Eleventh International Conference on Learning Representations ICLR 2023; 2023 May 1\u20135; Kigali Rwanda."},{"key":"e_1_3_2_31_2","unstructured":"Team K Du A Gao B Xing B Jiang C Chen C Li C Xiao C Du C Liao C et\u00a0al. Kimi k1.5: Scaling reinforcement learning with LLMs. arXiv. 2025. https:\/\/doi.org\/10.48550\/arXiv. 2501.12599"},{"key":"e_1_3_2_32_2","unstructured":"DeepSeek-AI Guo D Yang D Zhang H Song J Zhang R Xu R Zhu Q Ma S Wang P et\u00a0al. DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. arXiv. 2025. https:\/\/doi.org\/10.48550\/arXiv.2501.12948"},{"key":"e_1_3_2_33_2","unstructured":"El-Kishky A Wei A Saraiva A Minaiev B Selsam D Dohan D Song F Lightman H Ignasi C Pachocki J et\u00a0al. Competitive programming with large reasoning models. arXiv. 2025. https:\/\/doi.org\/10.48550\/arXiv.2502.06807"},{"key":"e_1_3_2_34_2","unstructured":"Wang X Wei J Schuurmans D Le QV Chi EH Narang S Chowdhery A Zhou D. Self-consistency improves chain of thought reasoning in language models. Paper presented at: The Eleventh International Conference on Learning Representations ICLR 2023; 2023 May 1\u20135; Kigali Rwanda. OpenReview.net 2023."},{"issue":"17","key":"e_1_3_2_35_2","first-page":"19525","article-title":"PREFER: Prompt ensemble learning via feedback-reflectrefine","volume":"38","author":"Zhang C","year":"2024","unstructured":"Zhang C, Liu L, Wang C, Sun X, Wang H, Wang J, Cai M. PREFER: Prompt ensemble learning via feedback-reflectrefine. Proc AAAI Conf Artif Intell. 2024;38(17):19525\u201319532.","journal-title":"Proc AAAI Conf Artif Intell"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","unstructured":"Huang J Gu S Hou L Wu Y Wang X Yu H Han J. Large language models can self-improve. Paper presented at: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing EMNLP 2023; 2023 Dec 6\u201310; Singapore.","DOI":"10.18653\/v1\/2023.emnlp-main.67"},{"key":"e_1_3_2_37_2","unstructured":"Madaan A Tandon N Gupta P Gupta P Hallinan S Gao L Wiegreffe S Alon U Dziri N Prabhumoye S Yang S et\u00a0al. Self-refine: Iterative refinement with self-feedback. Paper presented at: Advances in Neural Information Processing Systems 36: Annual Conference on Neual Information Processing Systems 2023 NeurIPS 2023; 2023 Dec 10\u201316; New Orleans LA USA."},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","unstructured":"Zhang W Shen Y Wu L Peng Q Wang J Zhuang Y Lu W. Self-contrast: Better reflection through inconsistent solving perspectives. Paper presented at: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) ACL 2024; 2024 Aug 11\u201316; Bangkok Thailand.","DOI":"10.18653\/v1\/2024.acl-long.197"},{"key":"e_1_3_2_39_2","unstructured":"Gou Z Shao Z Gong Y Shen Y Yang Y Duan N Chen W. CRITIC: Large language models can self-correct with toolinteractive critiquing. Paper presented at: The Twelfth International Conference on Learning Representations ICLR 2024; 2024 May 7\u201311; Vienna Austria."},{"key":"e_1_3_2_40_2","unstructured":"Wang G Xie Y Jiang Y Mandlekar A Xiao C Zhu Y Fan L Anandkumar A. Voyager: An open-ended embodied agent with large language models. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2305.16291"},{"key":"e_1_3_2_41_2","unstructured":"Tan W Ding Z Zhang W Li B ZhouB Yue J Xia H Jiang J Zheng L Xu X et\u00a0al. Towards general computer control: A multimodal agent for red dead redemption II as a case study. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2403.03186"},{"key":"e_1_3_2_42_2","unstructured":"Shinn N Cassano F Gopinath A Narasimhan K Yao S. Reflexion: Language agents with verbal reinforcement learning. Paper presented at: Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023 NeurIPS 2023; 2023 Dec 10\u201316; New Orleans LA USA."},{"key":"e_1_3_2_43_2","unstructured":"Lightman H Kosaraju V Burda Y Edwards H Baker B Lee T Leike J Schulman J Sutskever I Cobbe K. Let\u2019s verify step by step. Paper presented at: The Twelfth International Conference on Learning Representations ICLR 2024; 2024 May 7\u201311; Vienna Austria."},{"key":"e_1_3_2_44_2","doi-asserted-by":"crossref","unstructured":"Feng Y Wang Y Liu J Zheng S Lu Z. LLaMA-Rider: Spurring large language models to explore the open world. Paper presented at: Findings of the Association for Computational Linguistics: NAACL 2024; 2024 Jun 16\u201321; Mexico City Mexico.","DOI":"10.18653\/v1\/2024.findings-naacl.292"},{"key":"e_1_3_2_45_2","unstructured":"Liu H Sferrazza C Abbeel P. Chain of hindsight aligns language models with feedback. Paper presented at: The Twelfth International Conference on Learning Representations ICLR 2024; 2024 May 7\u201311; Vienna Austria."},{"key":"e_1_3_2_46_2","unstructured":"Gao Y Xiong Y Gao X Jia K Pan J Bi Y Dai Y Sun J Wang M Wang H. Retrieval-augmented generation for large language models: A survey. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv. 2312.10997"},{"key":"e_1_3_2_47_2","unstructured":"Zhao P Zhang H Yu Q Wang Z Geng Y Fu F Yan L Zhang W Jiang J Cui B. Retrieval-augmented generation for AI-generated content: A survey. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2402.19473"},{"key":"e_1_3_2_48_2","doi-asserted-by":"crossref","unstructured":"Park JS O\u2019Brien JC Cai CJ Morris MR Liang P Bernstein MS. Generative agents: Interactive simulacra of human behavior. Paper presented at: Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology UIST 2023; 2023 Oct 29\u2013Nov 1; San Francisco CA USA.","DOI":"10.1145\/3586183.3606763"},{"key":"e_1_3_2_49_2","unstructured":"Xu Y Wang S Li Ps. Exploring large language models for communication games: An empirical study on werewolf. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2309.04658"},{"key":"e_1_3_2_50_2","unstructured":"Zhou W Jiang YE Li L Wu J Wang T Qiu S Zhang J Chen J Wu R Wang S et al. Agents: An open-source framework for autonomous language agents. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2309.07870"},{"key":"e_1_3_2_51_2","unstructured":"Zhu X Chen Y Tian H Tao C Su W Yang C Huang G Li B Lu L Wang X et al. Ghost in the Minecraft: Generally capable agents for OpenWorld environments via large language models with text-based knowledge and memory. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2305.17144"},{"key":"e_1_3_2_52_2","unstructured":"Piterbarg U Pinto L Fergus R. diff history for neural language agents. Paper presented at: Forty-first International Conference on Machine Learning ICML 2024; 2024 Jul 21\u201327; Vienna Austria."},{"key":"e_1_3_2_53_2","unstructured":"Aeronautiques C Howe A Knoblock C McDermott ISID Ram A Veloso M Weld D Sri DW Barrett A Christianson D. Pddl| the planning domain definition language. Technical Report; 1998."},{"key":"e_1_3_2_54_2","unstructured":"Silver T Hariprasad V Shuttleworth RS Kumar N Lozano-P\u00e9rez T Kaelbling LP. PDDL planning with pretrained large language models. Paper presented at: NeurIPS 2022 Foundation Models for Decision Making Workshop; 2022; New Orleans LA USA."},{"issue":"18","key":"e_1_3_2_55_2","first-page":"20256","article-title":"Generalized planning in PDDL domains with pretrained large language models","volume":"38","author":"Silver T","year":"2024","unstructured":"Silver T, Dan S, Srinivas K, Tenenbaum JB, Kaelbling LP, Katz M. Generalized planning in PDDL domains with pretrained large language models. Proc AAAI Conf Artif Intell. 2024;38(18):20256\u201320264.","journal-title":"Proc AAAI Conf Artif Intell"},{"key":"e_1_3_2_56_2","unstructured":"Valmeekam K Marquez M Sreedharan S Kambhampati S. On the planning abilities of large language models\u2014A critical investigation. Paper presented at: Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023 NeurIPS 2023; 2023 Dec 10\u201316; New Orleans LA USA."},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","unstructured":"Chen G Yang L Jia R Hu Z Chen Y Zhang W Wang W Pan J. Language-augmented symbolic planner for open-world task planning. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2407.09792","DOI":"10.15607\/RSS.2024.XX.037"},{"key":"e_1_3_2_58_2","doi-asserted-by":"crossref","unstructured":"Han M Zhu Y Zhu S Wu YN Zhu Y. InterPreT: Interactive predicate learning from language feedback for generalizable task planning. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2405.19758","DOI":"10.15607\/RSS.2024.XX.034"},{"key":"e_1_3_2_59_2","unstructured":"Guan L Valmeekam K Sreedharan S Kambhampati S. Leveraging Pre-trained large language models to construct and utilize world models for model-based task planning. Paper presented at: Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023 NeurIPS 2023; 2023 Dec 10\u201316; New Orleans LA USA."},{"key":"e_1_3_2_60_2","unstructured":"Kambhampati S Valmeekam K Guan L Verma M Stechly K Bhambri S Saldyt LP Murthy AB. Position: LLMs can\u2019t plan but can help planning in LLM-modulo frameworks. Paper presented at: Forty-first International Conference on Machine Learning ICML 2024; 2024 Jul 21\u201327; Vienna Austria."},{"key":"e_1_3_2_61_2","unstructured":"Singh I Traum D Thomason J. TwoStep: Multi-agent task planning using classical planners and large language models. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2403.17246."},{"key":"e_1_3_2_62_2","unstructured":"Liu B Jiang Y Zhang X Liu Q Zhang S Biswas J Stone P LLM+P: Empowering large language models with optimal planning proficiency. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2304.11477"},{"key":"e_1_3_2_63_2","unstructured":"Xie Y Yu C Zhu T Bai J Gong Z Soh H. Translating natural language to planning goals with large-language models. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2302.05128"},{"key":"e_1_3_2_64_2","unstructured":"Oswald JT Srinivas K Kokel H Lee J Katz M Sohrabi S. Large language models as planning domain generators (student abstract). Paper presented at: Thirty-Eighth AAAI Conference on Artificial Intelligence AAAI 2024 Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence IAAI 2024 Fourteenth Symposium on Educational Advances in Artificial Intelligence EAAI 2014; 2024 Feb 20\u201327; Vancouver Canada."},{"key":"e_1_3_2_65_2","doi-asserted-by":"crossref","unstructured":"Yang Z Ishay A Lee J. Coupling large language models with logic programming for robust and general reasoning from text. Paper presented at: Findings of the Association for Computational Linguistics: ACL 2023; 2023 Jul 9\u201314; Toronto Canada.","DOI":"10.18653\/v1\/2023.findings-acl.321"},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","unstructured":"Pan L Albalak A Wang X Wang WY. Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning. Paper presented at: Findings of the Association for Computational Linguistics: EMNLP 2023; 2023 Dec 6\u201310; Singapore.","DOI":"10.18653\/v1\/2023.findings-emnlp.248"},{"key":"e_1_3_2_67_2","doi-asserted-by":"crossref","unstructured":"Zhou Z Song J Yao K Shu Z Ma L. ISR-LLM: Iterative self-refined large language model for long-horizon sequential task planning. Paper presented at: IEEE International Conference on Robotics and Automation ICRA 2024; 2024 May 13\u201317; Yokohama Japan.","DOI":"10.1109\/ICRA57147.2024.10610065"},{"key":"e_1_3_2_68_2","doi-asserted-by":"crossref","unstructured":"Yang Y Tomar A. On the planning search and memorization capabilities of large language models. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2309.01868","DOI":"10.1007\/978-3-031-71391-0_3"},{"key":"e_1_3_2_69_2","doi-asserted-by":"crossref","unstructured":"Birr T Pohl C Younes A Asfour T. AutoGPT+P: Affordance-based task planning with large language models. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2402.10778","DOI":"10.15607\/RSS.2024.XX.112"},{"key":"e_1_3_2_70_2","doi-asserted-by":"crossref","unstructured":"Cao Y Zhao H Cheng Y Shu T Chen Y Liu G Liang G Zhao J Yan J Li Y. Survey on large language model-enhanced reinforcement learning: Concept taxonomy and methods. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2404.00282","DOI":"10.1109\/TNNLS.2024.3497992"},{"key":"e_1_3_2_71_2","doi-asserted-by":"crossref","unstructured":"Hao S Gu Y Ma H Hong J Wang Z Wang D Hu Z. Reasoning with language model is planning with world model. Paper presented at: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing EMNLP 2023; 2023 Dec 6\u201310; Singapore.","DOI":"10.18653\/v1\/2023.emnlp-main.507"},{"key":"e_1_3_2_72_2","unstructured":"Wan Z Feng X Wen M Mc Aleer SM Wen Y Zhang W Wang J. AlphaZero-like tree-search can guide large language model decoding and training. Paper presented at: Forty-first International Conference on Machine Learning ICML 2024; 2024 Jul 21\u201327; Vienna Austria."},{"key":"e_1_3_2_73_2","doi-asserted-by":"crossref","unstructured":"Deng Y Zhang W Lam W Ng SK Chua TS. Plug-and-play policy planner for large language model powered dialogue agents. Paper presented at: The Twelfth International Conference on Learning Representations ICLR 2024; 2024 May 7\u201311; Vienna Austria.","DOI":"10.1145\/3589335.3641240"},{"key":"e_1_3_2_74_2","unstructured":"Hazra R Martires PZD Raedt LD. SayCanPay: Heuristic planning with large language models using learnable domain knowledge. Paper presented at: Thirty-Eighth AAAI Conference on Artificial Intelligence AAAI 2024 Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence IAAI 2024 Fourteenth Symposium on Educational Advances in Artificial Intelligence EAAI 2014; 2024 Feb 20\u201327; Vancouver Canada."},{"key":"e_1_3_2_75_2","unstructured":"Liu Z Hu H Zhang S Guo H Ke S Liu B Wang Z. Reason for future act for now: A principled architecture for autonomous LLM agents. Paper presented at: Forty-first International Conference on Machine Learning ICML 2024; 2024 Jul 21\u201327; Vienna Austria."},{"key":"e_1_3_2_76_2","unstructured":"Murthy R Heinecke S Niebles JC Liu Z Xue L Yao W Feng Y Chen Z Gokul A Arpit D et\u00a0al. REX: Rapid exploration and eXploitation for AI agents. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2307.08962"},{"key":"e_1_3_2_77_2","unstructured":"Silver D Hubert T Schrittwieser J Antonoglou I Lai M Guez A Lanctot M Sifire L Kumaran D Graepel T et\u00a0al. Mastering chess and shogi by self-play with a general reinforcement learning algorithm. arXiv. 2017. https:\/\/doi.org\/10.48550\/arXiv.1712.01815"},{"key":"e_1_3_2_78_2","unstructured":"Ouyang L Wu J Jiang X Almeida D Wainwright CL Mishkin P Zhang C Agarwal S Slama K Ray A et\u00a0al. Training language models to follow instructions with human feedback. Paper presented at: Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022 NeurIPS 2022; 2022 Nov 28\u2013Dec 9; New Orleans LA USA."},{"key":"e_1_3_2_79_2","unstructured":"Lee H Phatale S Mansoor H Lu KR Mesnard T Ferret J Bishop C Hall E Carbune V Rastogi A. RLAIF vs. RLHF: Scaling reinforcement learning from human feedback with AI feedback. Paper presented at: Forty-First International Conference on Machine Learning ICML 2024; 2024 Jul 21\u201327; Vienna Austria."},{"key":"e_1_3_2_80_2","unstructured":"Yao W Heinecke S Niebles JC Liu Z Feng Y Xue L Rithesh RN Chen Z Zhang J Arpit D et\u00a0al. Retroformer: Retrospective large language agents with policy gradient optimization. Paper presented at: The Twelfth International Conference on Learning Representations ICLR 2024; 2024 May 7\u201311; Vienna Austria."},{"key":"e_1_3_2_81_2","doi-asserted-by":"crossref","unstructured":"Yang J Dong Y Liu S Li B Wang Z Tan H Jiang C Kang J Zhang Y Zhou K et\u00a0al. Octopus: Embodied vision-language programmer from environmental feedback. Paper presented at: Computer Vision\u2014ECCV 2024\u201418th European Conference; 2024 Sep 29\u2013Oct 4; Milan Italy.","DOI":"10.1007\/978-3-031-73232-4_2"},{"key":"e_1_3_2_82_2","unstructured":"Du Y Watkins O Wang Z Colas C Darrell T Abbeel P Gupta A Andreas J. Guiding pretraining in reinforcement learning with large language models. Paper presented at: International Conference on Machine Learning ICML 2023; 2023 July 23\u201329; Honolulu Hawaii USA."},{"key":"e_1_3_2_83_2","unstructured":"Carta T Romac C Wolf T Lamprier S Sigaud O Oudeyer P. Grounding large language models in interactive environments with online reinforcement learning. Paper presented at: International Conference on Machine Learning ICML 2023; 2023 July 23\u201329; Honolulu Hawaii USA."},{"key":"e_1_3_2_84_2","unstructured":"Wang K Lu Y Santacroce M Gong Y Zhang C Shen Y. Adapting LLM agents through communication. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2310.01444"},{"key":"e_1_3_2_85_2","unstructured":"Xu Z Yu C Fang F Wang Y Wu Y. Language agents with reinforcement learning for strategic play in the werewolf game. Paper presented at: Forty-first International Conference on Machine Learning ICML 2024; 2024 Jul 21\u201327; Vienna Austria."},{"key":"e_1_3_2_86_2","unstructured":"Zhang D Chen L Zhang S Xu H Zhao Z Yu K. Large language models are semiparametric reinforcement learning agents. Paper presented at: Advances in Neural Information ProcessingSystems 36: Annual Conference on Neural Information Processing Systems 2023 NeurIPS 2023; 2023 Dec 10\u201316; New Orleans LA USA."},{"key":"e_1_3_2_87_2","unstructured":"Schulman J Wolski F Dhariwal P Radford A Klimov O. Proximal policy optimization algorithms. arXiv. 2017. https:\/\/doi.org\/10.48550\/arXiv.1707.06347"},{"key":"e_1_3_2_88_2","unstructured":"Zhang Y Mao S Ge T Wang X Xia Y Lan M Wei F. K-level reasoning with large language models. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2402.01521"},{"key":"e_1_3_2_89_2","unstructured":"Guo J Yang B Yoo P Lin BY Iwasawa Y Matsuo Y. Suspicion-agent: Playing imperfect information games with theory of mind aware GPT-4. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2309.17277"},{"key":"e_1_3_2_90_2","unstructured":"Gemp I Patel R Bachrach Y Lanctot M Dasagi V Marris L Piliouras G Liu S Tuyls K. States as strings as strategies: Steering language models with game-theoretic solvers. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2402.01704"},{"key":"e_1_3_2_91_2","unstructured":"Mao S Cai Y Xia Y Wu W Wang X Wang F Ge T Wei F. ALYMPICS: Language agents meet game theory. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2311.03220"},{"key":"e_1_3_2_92_2","unstructured":"Duan J Zhang R Diffenderfer J Kaikhura B Sun L Stengel-Eskin E Bansal M Chen T Xu K. GTBench: Uncovering the strategic reasoning limitations of LLMs via game-theoretic evaluations. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv. 2402.12348"},{"key":"e_1_3_2_93_2","unstructured":"Fan C Chen J Jin Y He H. Can large language models serve as rational players in game theory? A systematic analysis. Paper presented at: Thirty-Eighth AAAI Conference on Artificial Intelligence AAAI 2024 Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence IAAI 2024 Fourteenth Symposium on Educational Advances in Artificial Intelligence EAAI 2014; 2024 Feb 20\u201327; Vancouver Canada."},{"key":"e_1_3_2_94_2","unstructured":"Huang J Li EJ Lam MH Liang T Wang W Yuan Y Jiao W Wang X Tu Z Lyu MR. How far are we on the decision-making of LLMs? Evaluating LLMs\u2019 gaming ability in multi-agent environments. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv. 2403.11807"},{"key":"e_1_3_2_95_2","unstructured":"Webb TW Mondal SS Wang C Krabach B Momennejad I. A prefrontal cortex-inspired architecture for planning in large language models. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2308.09658"},{"key":"e_1_3_2_96_2","unstructured":"Hu P Qi J Li X Li H Wang X Quan B Wang R Zhou Y. Tree-of-mixed-thought: Combining fast and slow thinking for multihop visual reasoning. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2308.09658"},{"key":"e_1_3_2_97_2","unstructured":"Lin BY Fu Y Yang K Brahman F Huang S Bhagavatula C Ammanabrolu P Choi Y Ren X. SwiftSage: A generative agent with fast and slow thinking for complex interactive tasks. Paper presented at: Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023 NeurIPS 2023; 2023 Dec 10\u201316; New Orleans LA USA."},{"key":"e_1_3_2_98_2","unstructured":"Wu X Shen Y Shan C Song K Wang S Zhang B Feng J Cheng H Chen W Xiong Y et\u00a0al. Can graph learning improve task planning? arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2405.19119"},{"issue":"11","key":"e_1_3_2_99_2","doi-asserted-by":"crossref","first-page":"959","DOI":"10.1016\/j.tics.2022.08.003","article-title":"Planning with theory of mind","volume":"26","author":"Ho MK","year":"2022","unstructured":"Ho MK, Saxe R, Cushman F. Planning with theory of mind. Trends Cogn Sci. 2022;26(11):959\u2013971.","journal-title":"Trends Cogn Sci"},{"key":"e_1_3_2_100_2","unstructured":"Akata E Schulz L Coda-Forno J Oh SJ Bethge M Schulz E. Playing repeated games with large language models. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2305.16867"},{"key":"e_1_3_2_101_2","unstructured":"Liu Y Chen W Bai Y Li G Gao W Lin L. Aligning cyber space with physical world: A comprehensive survey on embodied AI. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2407.06886"},{"key":"e_1_3_2_102_2","unstructured":"NVIDIA. Nvidia isaac sim: Robotics simulation and synthetic data. 2023. https:\/\/developer.nvidia.com\/isaac-sim."},{"key":"e_1_3_2_103_2","unstructured":"Koenig N and Howard A. Design and use paradigms for gazebo an open-source multi-robot simulator. Paper presented at: IEEE\/RSJ International Conference on Intelligent Robots and Systems; 2004; New Orleans LA USA."},{"key":"e_1_3_2_104_2","unstructured":"Coumans E Bai Y. Pybullet a python module for physics simulation for games robotics and machine learning. 2016\u20132019. http:\/\/pybullet.org."},{"key":"e_1_3_2_105_2","unstructured":"Juliani A Berges VP Teng E Cohen A Harper J Elion C goy C Gao Y Henry H Mattar M et\u00a0al. Unity: A general platform for intelligent agents. arXiv. 2020. https:\/\/doi.org\/10.48550\/arXiv.1809.02627"},{"key":"e_1_3_2_106_2","doi-asserted-by":"crossref","unstructured":"Shah S Dey D Lovett C Kapoor A. AirSim: High-fidelity visual and physical simulation for autonomous vehicles. In: Field and service robotics. Cham: Springer; 2018. p. 621\u2013635.","DOI":"10.1007\/978-3-319-67361-5_40"},{"key":"e_1_3_2_107_2","unstructured":"Makoviychuk V Wawrzyniak L Guo Y Lu M Storey K Macklin M Hoeller D Rudin N Allshire A Handa A et\u00a0al. Isaac gym: High performance gpu-based physics simulation for robot learning. arXiv. 2021. https:\/\/doi.org\/10.48550\/arXiv.2108.10470"},{"key":"e_1_3_2_108_2","unstructured":"Cyberbotics. Webots: Open-source robot simulator. https:\/\/github.com\/cyberbotics\/webots"},{"key":"e_1_3_2_109_2","doi-asserted-by":"crossref","unstructured":"Todorov E Erez T Tassa Y. MuJoCo: A physics engine for model-based control. Paper presented at: IEEE\/RSJ International Conference on Intelligent Robots and Systems; 2012; Vilamoura-Algarve Portugal.","DOI":"10.1109\/IROS.2012.6386109"},{"key":"e_1_3_2_110_2","unstructured":"ISAE-SUPAERO. MORSE: The modular open robots simulator engine."},{"key":"e_1_3_2_111_2","unstructured":"Kolve E Mottaghi R Gordon D Zhu Y Gupta A Farhadi A. AI2-THOR: An interactive 3D environment for visual AI. arXiv. 2017. https:\/\/doi.org\/10.48550\/arXiv.1712.05474"},{"key":"e_1_3_2_112_2","doi-asserted-by":"crossref","unstructured":"Chang AX Dai A Funkhouser TA Halber M Nieber M Savva M Song S Zeng A Zhang Y. Matterport3D: Learning from RGB-D data in indoor environments. Paper presented at: 2017 International Conference on 3D Vision 3DV 2017; 2017 Oct 10\u201312; Qingdao China.","DOI":"10.1109\/3DV.2017.00081"},{"key":"e_1_3_2_113_2","doi-asserted-by":"crossref","unstructured":"Puig X Ra K Boben M Li J Wang T Fidler S Torralba A. VirtualHome: Simulating household activities via programs. Paper presented at: 2018 IEEE Conference on Computer Vision and Pattern Recognition CVPR 2018; 2018 Jun 18\u201322; Salt Lake City UT USA.","DOI":"10.1109\/CVPR.2018.00886"},{"key":"e_1_3_2_114_2","doi-asserted-by":"crossref","unstructured":"Xiang F Qin Y Mo K Xia Y Zhu H Liu F Liu M Jiang H Yuan Y Wang H et\u00a0al. SAPIEN: A simulated part-based interactive environment. Paper presented at: 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition CVPR 2020; 2020 Jun 13\u201319; Seattle WA USA.","DOI":"10.1109\/CVPR42600.2020.01111"},{"key":"e_1_3_2_115_2","unstructured":"Li C Xia F Mart\u00edn-Mart\u00edn R Lingelbach M SrivastavaS Shen B Vainio KE Gokmen C Dharan G Jain T et\u00a0al. iGibson 2.0: Object-centric simulation for robot learning of everyday household tasks. Paper presented at: Conference on Robot Learning; 2021 Nov 8\u201311; London UK."},{"key":"e_1_3_2_116_2","doi-asserted-by":"crossref","unstructured":"Savva M Kadian A Maksymets O Zhao Y Wijmans E Jain B Straub J Liu J Koltun V Malik J et\u00a0al. Habitat: A platform for embodied AI research. Paper presented at: 2019 IEEE\/CVF International Conference on Computer Vision ICCV 2019; Oct 27\u2013Nov 2 2019; Seoul South Korea.","DOI":"10.1109\/ICCV.2019.00943"},{"key":"e_1_3_2_117_2","unstructured":"Gan C Schwartz J Alter S Mrowca D Schrimpf M Traer J De Freitas J Kubilius J Bhandwaldar A Haber N et\u00a0al. ThreeDWorld: A platform for interactive multi-modal physical simulation. Paper presented at: Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 NeurIPS Datasets and Benchmarks 2021; 2021 December; Virtual."},{"key":"e_1_3_2_118_2","unstructured":"Costarelli A Allen M Hauksson R Sodunker G Hariharan S Cheng C Li W Clymer J Yadav A. GameBench: Evaluating strategic reasoning abilities of LLM agents. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2406.06613"},{"key":"e_1_3_2_119_2","doi-asserted-by":"crossref","unstructured":"Sweetser P. Large language models and video games: A preliminary scoping review. Paper presented at: ACM Conversational User Interfaces 2024 CUI 2024; 2024 Jul 8\u201310; Luxembourg.","DOI":"10.1145\/3640794.3665582"},{"key":"e_1_3_2_120_2","unstructured":"Hu S Huang T Liu G Kompella RR Ilhan F Tekin SF Xu Y Yahn Z Liu L. A survey on large language model-based game agents. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2404.02039"},{"key":"e_1_3_2_121_2","doi-asserted-by":"crossref","unstructured":"Yang D Kleinman E Harteveld C. GPT for games: An updated scoping review (2020-2024). arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2411.00308","DOI":"10.1109\/TG.2025.3563780"},{"key":"e_1_3_2_122_2","unstructured":"Ma W Mi Q Zeng Y Yan X Wu Y Lin R Zhang H Wang J. Large language models play StarCraft II: Benchmarks and a chain of summarization approach. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2312.11865"},{"key":"e_1_3_2_123_2","unstructured":"Shao X Jiang W Zuo F Liu M. SwarmBrain: Embodied agent for real-time strategy game StarCraft II via large language models. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2401.17749"},{"key":"e_1_3_2_124_2","doi-asserted-by":"crossref","unstructured":"Goecks VG Waytowich NR. COA-GPT: Generative pre-trained transformers for accelerated course of action development in military operations. Paper presented at: International Conference on Military Communication and Information Systems ICMCIS 2024; 2024 Apr 23\u201324; Koblenz Germany.","DOI":"10.1109\/ICMCIS61231.2024.10540749"},{"key":"e_1_3_2_125_2","unstructured":"Wang Z Cai S Chen G Liu A Ma X Liang Y. Describe explain plan and select: Interactive planning with LLMs enables open-world multi-task agents. Paper presented at: Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023 NeurIPS 2023; 2023 Dec 10\u201316; New Orleans LA USA."},{"key":"e_1_3_2_126_2","doi-asserted-by":"crossref","unstructured":"Cai S Wang Z Ma X Liu A Liang Y. Open-world multi-task control through goal aware representation learning and adaptive horizon prediction. Paper presented at: IEEE\/CVF Conference on Computer Vision and Pattern Recognition CVPR 2023; 2023 Jun 17\u201324; Vancouver BC Canada.","DOI":"10.1109\/CVPR52729.2023.01320"},{"key":"e_1_3_2_127_2","unstructured":"Yuan H Zhang C Wang H Xie F Cai P Dong H Lu Z. Plan4MC: Skill reinforcement learning and planning for open-world Minecraft tasks. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2303.16563"},{"key":"e_1_3_2_128_2","unstructured":"Wu Y Tang X Mitchell TM Li Y. SmartPlay: A benchmark for LLMs as intelligent agents. Paper presented at: The Twelfth International Conference on Learning Representations ICLR 2024; 2024 May 7\u201311; Vienna Austria."},{"key":"e_1_3_2_129_2","unstructured":"Liu J Yu C Gao J Xie Y Liao Q Wu Y Wang Y. LLM-powered hierarchical language agent for real-time human AI coordination. Paper presented at: Proceedings of the 23rd International Conference on Autonomous Agents and Multiagent Systems AAMAS 2024; 2024 May 6\u201310; Auckland New Zealand."},{"key":"e_1_3_2_130_2","doi-asserted-by":"crossref","unstructured":"Lan Y Hu Z Wang L Ye D Zhao P Lim E-P Xiong H Wang H. LLM-based agent society investigation: Collaboration and confrontation in Avalon gameplay. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2310.14985","DOI":"10.18653\/v1\/2024.emnlp-main.7"},{"key":"e_1_3_2_131_2","unstructured":"Wang S Liu C Zheng Z Qi S Chen S Yang Q Zhao A Wang C Song S Huang G. Avalon\u2019s game of thoughts: Battle against deception through recursive contemplation. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2310.01320"},{"key":"e_1_3_2_132_2","unstructured":"Light J Cai M Shen S Hu Z. Avalonbench: Evaluating llms playing the game of avalon. Paper presented at: NeurIPS 2023 Foundation Models for Decision Making Workshop; 2023; New Orleans LA USA."},{"key":"e_1_3_2_133_2","doi-asserted-by":"crossref","unstructured":"Wu D Shi H Sun Z Liu B. Deciphering digital detectives: Understanding LLM behaviors and capabilities in multi-agent mystery games. Paper presented at: Findings of the Association for Computational Linguistics ACL 2024; 2024 Aug 11\u201316; Bangkok Thailand.","DOI":"10.18653\/v1\/2024.findings-acl.490"},{"key":"e_1_3_2_134_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.ade9097"},{"key":"e_1_3_2_135_2","doi-asserted-by":"crossref","unstructured":"Zhu A Aggarwal K Feng AH Martin LJ Callison-Burch C. FIREBALL: A dataset of dungeons and dragons actual-play with structured game state information. Paper presented at: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) ACL 2023; 2023 Jul 9\u201314; Toronto Canada.","DOI":"10.18653\/v1\/2023.acl-long.229"},{"key":"e_1_3_2_136_2","unstructured":"Wu S Zhu L Yang T Xu S Fu Q Wei Y Fu H. Enhance reasoning for large language models in the game werewolf. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2402.02330"},{"key":"e_1_3_2_137_2","doi-asserted-by":"crossref","unstructured":"Chen J Hu X Liu S Huang S Tu W-W He Z Wen L. LLMArena: Assessing capabilities of large language models in dynamic multi-agent environments. Paper presented at: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) ACL 2024; 2024 Aug 11\u201316; Bangkok Thailand.","DOI":"10.18653\/v1\/2024.acl-long.705"},{"key":"e_1_3_2_138_2","unstructured":"Huang C Cao Y Wen Y Zhou T Zhang Y. PokerGPT: An end-to-end lightweight solver for multi-player Texas Hold\u2019em via large language model. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2401.06781"},{"key":"e_1_3_2_139_2","unstructured":"Guo H Liu Z Zhang Y Wang Z. Can large language models play games? A case study of a self-play approach. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2403.05632"},{"key":"e_1_3_2_140_2","doi-asserted-by":"crossref","unstructured":"Qian C Liu W Liu H Chen N Dang Y Li J Yang C Su Y Cong X et\u00a0al. ChatDev: Communicative agents for software development. Paper presented at: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) ACL 2024; 2024 Aug 11\u201316; Bangkok Thailand.","DOI":"10.18653\/v1\/2024.acl-long.810"},{"key":"e_1_3_2_141_2","unstructured":"Chen J Yuan S Ye R Majumder BP Richardson K. Put your money where your mouth is: Evaluating strategic planning and execution of LLM agents in an auction arena. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2310.05746"},{"key":"e_1_3_2_142_2","unstructured":"Yuan Q Kazemi M Xu X Noble I Imbrasaite V Ramachandran D. TaskLAMA: Probing the complex task understanding of language models. Paper presented at: Thirty-Eighth AAAI Conference on Artificial Intelligence AAAI 2024 Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence IAAI 2024 Fourteenth Symposium on Educational Advances in Artificial Intelligence EAAI 2014; 2024 Feb 20\u201327; Vancouver Canada."},{"key":"e_1_3_2_143_2","unstructured":"Lamparth M Corso A Ganz J Mastro OS Schneider J Trinkunas H. Human vs. machine: Language models and wargames. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2403.03407"},{"key":"e_1_3_2_144_2","unstructured":"Hua W Fan L Li L Mei K Ji J Ge Y Hemphill L Zhang Y. War and peace (WarAgent): Large language model-based multiagent simulation of world wars. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2311.17227"},{"key":"e_1_3_2_145_2","doi-asserted-by":"crossref","DOI":"10.1007\/s11704-024-40678-2","article-title":"Tool learning with large language models: A survey","volume":"19","author":"Qu C","year":"2025","unstructured":"Qu C, Dai S, Wei X, Cai H, Wang S, Yin D, Xu J, Wen JR. Tool learning with large language models: A survey. Front Comput Sci. 2025;19: Article 198343.","journal-title":"Front Comput Sci"},{"key":"e_1_3_2_146_2","unstructured":"Zhang C He S Qian J Li B Li L Qin S Kang Y Ma M Liu G Lin Q et\u00a0al. Large language model-brained GUI agents: A survey. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2411.18279"},{"key":"e_1_3_2_147_2","unstructured":"Wang S Liu W Chen J Zhou Y Gan W Zeng X Che Y Yu S Hao X Shao K et\u00a0al. GUI agents with foundation models: A comprehensive survey. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2411.04890"}],"container-title":["Intelligent Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/spj.science.org\/doi\/pdf\/10.34133\/icomputing.0124","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,23]],"date-time":"2025-05-23T08:01:49Z","timestamp":1747987309000},"score":1,"resource":{"primary":{"URL":"https:\/\/spj.science.org\/doi\/10.34133\/icomputing.0124"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1]]},"references-count":146,"alternative-id":["10.34133\/icomputing.0124"],"URL":"https:\/\/doi.org\/10.34133\/icomputing.0124","relation":{},"ISSN":["2771-5892"],"issn-type":[{"value":"2771-5892","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1]]},"assertion":[{"value":"2024-12-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-17","order":1,"name":"revised","label":"Revised","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-23","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"0124"}}