{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T18:35:12Z","timestamp":1784572512988,"version":"3.55.0"},"reference-count":69,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2024,9,27]],"date-time":"2024-09-27T00:00:00Z","timestamp":1727395200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key R & D Program","award":["2023YFB4503801"],"award-info":[{"award-number":["2023YFB4503801"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62072007, 62192733, 61832009, and 62192730"],"award-info":[{"award-number":["62072007, 62192733, 61832009, and 62192730"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key Program of Hubei","award":["JD2023008"],"award-info":[{"award-number":["JD2023008"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2024,9,30]]},"abstract":"<jats:p>Although large language models (LLMs) have demonstrated impressive ability in code generation, they are still struggling to address the complicated intent provided by humans. It is widely acknowledged that humans typically employ planning to decompose complex problems and schedule solution steps prior to implementation. To this end, we introduce planning into code generation to help the model understand complex intent and reduce the difficulty of problem-solving. This paper proposes a self-planning code generation approach with large language models, which consists of two phases, namely planning phase and implementation phase. Specifically, in the planning phase, LLM plans out concise solution steps from the intent combined with few-shot prompting. Subsequently, in the implementation phase, the model generates code step by step, guided by the preceding solution steps. We conduct extensive experiments on various code-generation benchmarks across multiple programming languages. Experimental results show that self-planning code generation achieves a relative improvement of up to 25.4% in Pass@1 compared to direct code generation, and up to 11.9% compared to Chain-of-Thought of code generation. Moreover, our self-planning approach also enhances the quality of the generated code with respect to correctness, readability, and robustness, as assessed by humans.<\/jats:p>","DOI":"10.1145\/3672456","type":"journal-article","created":{"date-parts":[[2024,6,13]],"date-time":"2024-06-13T17:56:59Z","timestamp":1718301419000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":129,"title":["Self-Planning Code Generation with Large Language Models"],"prefix":"10.1145","volume":"33","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5317-865X","authenticated-orcid":false,"given":"Xue","family":"Jiang","sequence":"first","affiliation":[{"name":"Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education, Beijing, China and School of Computer Science, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6228-4019","authenticated-orcid":false,"given":"Yihong","family":"Dong","sequence":"additional","affiliation":[{"name":"Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education, Beijing, China and School of Computer Science, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-3151-6683","authenticated-orcid":false,"given":"Lecheng","family":"Wang","sequence":"additional","affiliation":[{"name":"Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education, Beijing, China and School of Computer Science, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7000-6909","authenticated-orcid":false,"given":"Zheng","family":"Fang","sequence":"additional","affiliation":[{"name":"Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education, Beijing, China and School of Computer Science, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-6501-0847","authenticated-orcid":false,"given":"Qiwei","family":"Shang","sequence":"additional","affiliation":[{"name":"Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education, Beijing, China and School of Computer Science, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5828-0186","authenticated-orcid":false,"given":"Ge","family":"Li","sequence":"additional","affiliation":[{"name":"Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education, Beijing, China and School of Computer Science, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1087-226X","authenticated-orcid":false,"given":"Zhi","family":"Jin","sequence":"additional","affiliation":[{"name":"Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education, Beijing, China and School of Computer Science, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9374-3900","authenticated-orcid":false,"given":"Wenpin","family":"Jiao","sequence":"additional","affiliation":[{"name":"Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education, Beijing, China and School of Computer Science, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,9,27]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","unstructured":"Pekka Abrahamsson Outi Salo Jussi Ronkainen and Juhani Warsta. 2017. Agile software development methods: Review and analysis. arXiv:1709.08439. Retrieved from 10.48550\/arXiv.1709.08439","DOI":"10.48550\/arXiv.1709.08439"},{"key":"e_1_3_3_3_2","unstructured":"Jacob Austin Augustus Odena Maxwell I. Nye Maarten Bosma Henryk Michalewski David Dohan Ellen Jiang Carrie J. Cai Michael Terry Quoc V. Le and Charles Sutton. 2021. Program synthesis with large language models. arXiv:2108.07732. Retrieved from https:\/\/arxiv.org\/abs\/2108.07732"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.5555\/3495724.3495883"},{"key":"e_1_3_3_5_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201923)","author":"Chen Bei","year":"2023","unstructured":"Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen. 2023. CodeT: Code generation with generated tests. In Proceedings of the International Conference on Learning Representations (ICLR \u201923)."},{"key":"e_1_3_3_6_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Pond\u00e9 de Oliveira Pinto Jared Kaplan Harrison Edwards Yuri Burda Nicholas Joseph Greg Brockman Alex Ray Raul Puri Gretchen Krueger Michael Petrov Heidy Khlaaf Girish Sastry Pamela Mishkin Brooke Chan Scott Gray Nick Ryder Mikhail Pavlov Alethea Power Lukasz Kaiser Mohammad Bavarian Clemens Winter Philippe Tillet Felipe Petroski Such Dave Cummings Matthias Plappert Fotios Chantzis Elizabeth Barnes Ariel Herbert-Voss William Hebgen Guss Alex Nichol Alex Paino Nikolas Tezak Jie Tang Igor Babuschkin Suchir Balaji Shantanu Jain William Saunders Christopher Hesse Andrew N. Carr Jan Leike Joshua Achiam Vedant Misra Evan Morikawa Alec Radford Matthew Knight Miles Brundage Mira Murati Katie Mayer Peter Welinder Bob McGrew Dario Amodei Sam McCandlish Ilya Sutskever and Wojciech Zaremba. 2021. Evaluating large language models trained on code. arXiv:2107.03374. Retrieved from https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_3_3_7_2","unstructured":"Wenhu Chen Xueguang Ma Xinyi Wang and William W. Cohen. 2022. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv:2211.12588. Retrieved from https:\/\/arxiv.org\/abs\/2211.12588"},{"key":"e_1_3_3_8_2","unstructured":"Xinyun Chen Maxwell Lin Nathanael Sch\u00e4rli and Denny Zhou. 2023. Teaching large language models to self-debug. arXiv:2304.05128. Retrieved from https:\/\/arxiv.org\/abs\/2304.05128"},{"key":"e_1_3_3_9_2","unstructured":"Hyung Won Chung Le Hou Shayne Longpre Barret Zoph Yi Tay William Fedus Eric Li Xuezhi Wang Mostafa Dehghani Siddhartha Brahma Albert Webson Shixiang Shane Gu Zhuyun Dai Mirac Suzgun Xinyun Chen Aakanksha Chowdhery Sharan Narang Gaurav Mishra Adams Yu Vincent Y. Zhao Yanping Huang Andrew M. Dai Hongkun Yu Slav Petrov Ed H. Chi Jeff Dean Jacob Devlin Adam Roberts Denny Zhou Quoc V. Le and Jason Wei. 2022. Scaling instruction-finetuned language models. arXiv:2210.11416. Retrieved from https:\/\/arxiv.org\/abs\/2210.11416"},{"key":"e_1_3_3_10_2","unstructured":"Karl Cobbe Vineet Kosaraju Mohammad Bavarian Jacob Hilton Reiichiro Nakano Christopher Hesse and John Schulman. 2021. Training verifiers to solve math word problems. arXiv:2110.14168. Retrieved from https:\/\/arxiv.org\/abs\/2110.14168"},{"key":"e_1_3_3_11_2","doi-asserted-by":"crossref","unstructured":"Yihong Dong Jiazheng Ding Xue Jiang Zhuo Li Ge Li and Zhi Jin. 2023. CodeScore: Evaluating code generation by learning code execution. arXiv:2301.09043. Retrieved from https:\/\/arxiv.org\/abs\/2301.09043","DOI":"10.1145\/3695991"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598048"},{"key":"e_1_3_3_13_2","unstructured":"Chrisantha Fernando Dylan Banarse Henryk Michalewski Simon Osindero and Tim Rockt\u00e4schel. 2023. Promptbreeder: Self-referential self-improvement via prompt evolution. arXiv:2309.16797. Retrieved from https:\/\/arxiv.org\/abs\/2309.16797"},{"key":"e_1_3_3_14_2","unstructured":"Daniel Fried Armen Aghajanyan Jessy Lin Sida Wang Eric Wallace Freda Shi Ruiqi Zhong Wen-tau Yih Luke Zettlemoyer and Mike Lewis. 2022. InCoder: A generative model for code infilling and synthesis. arXiv:2204.05999. Retrieved from https:\/\/arxiv.org\/abs\/2204.05999"},{"key":"e_1_3_3_15_2","unstructured":"Luyu Gao Aman Madaan Shuyan Zhou Uri Alon Pengfei Liu Yiming Yang Jamie Callan and Graham Neubig. 2022. PAL: Program-aided language models. arXiv:2211.10435. Retrieved from https:\/\/arxiv.org\/abs\/2211.10435"},{"key":"e_1_3_3_16_2","unstructured":"Zelalem Gero Chandan Singh Hao Cheng Tristan Naumann Michel Galley Jianfeng Gao and Hoifung Poon. 2023. Self-verification improves few-shot clinical information extraction. arXiv:2306.00024. Retrieved from https:\/\/arxiv.org\/abs\/2306.00024"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/2950290.2950334"},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2017\/514"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.499"},{"key":"e_1_3_3_20_2","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS Datasets and Benchmarks \u201921)","author":"Hendrycks Dan","year":"2021","unstructured":"Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. 2021. Measuring coding challenge competence with APPS. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS Datasets and Benchmarks \u201921)."},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.67"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-1002"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.5555\/3600270.3601883"},{"key":"e_1_3_3_24_2","first-page":"3843","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS \u201922)","volume":"35","author":"Lewkowycz Aitor","year":"2022","unstructured":"Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay V. Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022. Solving quantitative reasoning problems with language models. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS \u201922), Vol. 35. 3843\u20133857."},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","unstructured":"Jia Li Ge Li Yongmin Li and Zhi Jin. 2023. Structured chain-of-thought prompting for code generation. arXiv:2305.06599. Retrieved from 10.48550\/arXiv.2305.06599","DOI":"10.48550\/arXiv.2305.06599"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00179"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.abq1158"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-1057"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3560815"},{"key":"e_1_3_3_30_2","volume-title":"Requirements Engineering","author":"Macaulay Linda A","year":"2012","unstructured":"Linda A Macaulay. 2012. Requirements Engineering. Springer Science & Business Media."},{"key":"e_1_3_3_31_2","unstructured":"Aman Madaan Niket Tandon Prakhar Gupta Skyler Hallinan Luyu Gao Sarah Wiegreffe Uri Alon Nouha Dziri Shrimai Prabhumoye Yiming Yang Sean Welleck Bodhisattwa P. Majumder Shashank Gupta Amir Yazdanbakhsh and Peter Clark. 2023. Self-Refine: Iterative refinement with self-feedback. arXiv:2303.17651. Retrieved from https:\/\/arxiv.org\/abs\/2303.17651"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.90"},{"key":"e_1_3_3_33_2","first-page":"18984","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS \u201921)","volume":"34","author":"Mukherjee Rohan","year":"2021","unstructured":"Rohan Mukherjee, Yeming Wen, Dipak Chaudhari, Thomas W. Reps, Swarat Chaudhuri, and Christopher M. Jermaine. 2021. Neural program generation modulo static analysis. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS \u201921), Vol. 34. 18984\u201318996."},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","unstructured":"Erik Nijkamp Bo Pang Hiroaki Hayashi Lifu Tu Huan Wang Yingbo Zhou Silvio Savarese and Caiming Xiong. 2022. Codegen: An open large language model for code with multi-turn program synthesis. arXiv:2203.13474. Retrieved from 10.48550\/arXiv.2203.13474","DOI":"10.48550\/arXiv.2203.13474"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.138"},{"key":"e_1_3_3_36_2","unstructured":"OpenAI. 2023. GPT-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-02152-7_29"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613372.3614197"},{"key":"e_1_3_3_39_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201922)","author":"Poesia Gabriel","year":"2022","unstructured":"Gabriel Poesia, Alex Polozov, Vu Le, Ashish Tiwari, Gustavo Soares, Christopher Meek, and Sumit Gulwani. 2022. Synchromesh: Reliable code generation from pre-trained language models. In Proceedings of the International Conference on Learning Representations (ICLR \u201922)."},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1105"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/2594291.2594321"},{"key":"e_1_3_3_42_2","unstructured":"Shuo Ren Daya Guo Shuai Lu Long Zhou Shujie Liu Duyu Tang Neel Sundaresan Ming Zhou Ambrosio Blanco and Shuai Ma. 2020. CodeBLEU: A method for automatic evaluation of code synthesis. arXiv:2009.10297. Retrieved from https:\/\/arxiv.org\/abs\/2009.10297"},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.191"},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/1764810.1764814"},{"key":"e_1_3_3_45_2","volume-title":"Proceedings of the 10th International Conference on Learning Representations (ICLR \u201922)","author":"Sanh Victor","year":"2022","unstructured":"Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault F\u00e9vry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M. Rush. 2022. Multitask prompted training enables zero-shot task generalization. In Proceedings of the 10th International Conference on Learning Representations (ICLR \u201922). OpenReview.net."},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3558965"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","unstructured":"Yongliang Shen Kaitao Song Xu Tan Dongsheng Li Weiming Lu and Yueting Zhuang. 2023. Hugginggpt: Solving AI tasks with chatgpt and its friends in huggingface. arXiv:2303.17580. Retrieved from 10.48550\/arXiv.2303.17580","DOI":"10.48550\/arXiv.2303.17580"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48891.2023.10161317"},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00280"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33017055"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6430"},{"key":"e_1_3_3_52_2","first-page":"20227","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems (NeurIPS \u201920)","author":"Talmor Alon","year":"2020","unstructured":"Alon Talmor, Oyvind Tafjord, Peter Clark, Yoav Goldberg, and Jonathan Berant. 2020. Leap-of-thought: Teaching pre-trained models to systematically reason over implicit knowledge. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NeurIPS \u201920). 20227\u201320237."},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-industry.62"},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.754"},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.685"},{"key":"e_1_3_3_56_2","unstructured":"Zihao Wang Shaofei Cai Anji Liu Xiaojian Ma and Yitao Liang. 2023. Describe explain plan and select: Interactive planning with large language models enables open-world multi-task agents. arXiv:2302.01560. Retrieved from https:\/\/arxiv.org\/abs\/2302.01560"},{"key":"e_1_3_3_57_2","first-page":"6559","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS \u201919)","volume":"32","author":"Wei Bolin","year":"2019","unstructured":"Bolin Wei, Ge Li, Xin Xia, Zhiyi Fu, and Zhi Jin. 2019. Code generation as a dual task of code summarization. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS \u201919), Vol. 32. 6559\u20136569."},{"key":"e_1_3_3_58_2","doi-asserted-by":"publisher","unstructured":"Jason Wei Xuezhi Wang Dale Schuurmans Maarten Bosma Ed Chi Quoc Le and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv:2201.11903. Retrieved from 10.48550\/arXiv.2201.11903","DOI":"10.48550\/arXiv.2201.11903"},{"key":"e_1_3_3_59_2","first-page":"32353","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS \u201922)","volume":"35","author":"Wu Yuhuai","year":"2022","unstructured":"Yuhuai Wu, Albert Qiaochu Jiang, Wenda Li, Markus N. Rabe, Charles Staats, Mateja Jamnik, and Christian Szegedy. 2022. Autoformalization with large language models. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS \u201922), Vol. 35. 32353\u201332368."},{"key":"e_1_3_3_60_2","volume-title":"Proceedings of the 37th Conference on Neural Information Processing Systems","author":"Xie Yuxi","year":"2023","unstructured":"Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, Xu Zhao, Min-Yen Kan, Junxian He, and Qizhe Xie. 2023. Self-evaluation guided beam search for reasoning. In Proceedings of the 37th Conference on Neural Information Processing Systems."},{"key":"e_1_3_3_61_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201923)","author":"Yao Shunyu","year":"2023","unstructured":"Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing reasoning and acting in language models. In Proceedings of the International Conference on Learning Representations (ICLR \u201923). OpenReview.net."},{"key":"e_1_3_3_62_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1041"},{"key":"e_1_3_3_63_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-2002"},{"key":"e_1_3_3_64_2","unstructured":"Shun Zhang Zhenfang Chen Yikang Shen Mingyu Ding Joshua B. Tenenbaum and Chuang Gan. 2023. Planning with large language models for code generation. arXiv:2303.05510. Retrieved from https:\/\/arxiv.org\/abs\/2303.05510"},{"key":"e_1_3_3_65_2","unstructured":"Tianyi Zhang Tao Yu Tatsunori B. Hashimoto Mike Lewis Wen-tau Yih Daniel Fried and Sida I. Wang. 2022. Coder reviewer reranking for code generation. arXiv:2211.16490. Retrieved from https:\/\/arxiv.org\/abs\/2211.16490"},{"key":"e_1_3_3_66_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201923)","author":"Zhang Zhuosheng","year":"2023","unstructured":"Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2023. Automatic chain of thought prompting in large language models. In Proceedings of the International Conference on Learning Representations (ICLR \u201923). OpenReview.net."},{"key":"e_1_3_3_67_2","doi-asserted-by":"crossref","unstructured":"Qinkai Zheng Xiao Xia Xu Zou Yuxiao Dong Shan Wang Yufei Xue Zihan Wang Lei Shen Andi Wang Yang Li Teng Su Zhilin Yang and Jie Tang. 2023. CodeGeeX: A pre-trained model for code generation with multilingual evaluations on HumanEval-X. arXiv:2303.17568. Retrieved from https:\/\/arxiv.org\/abs\/2303.17568","DOI":"10.1145\/3580305.3599790"},{"key":"e_1_3_3_68_2","unstructured":"Denny Zhou Nathanael Sch\u00e4rli Le Hou Jason Wei Nathan Scales Xuezhi Wang Dale Schuurmans Olivier Bousquet Quoc Le and Ed H. Chi. 2022. Least-to-most prompting enables complex reasoning in large language models. arXiv:2205.10625. Retrieved from https:\/\/arxiv.org\/abs\/2205.10625"},{"key":"e_1_3_3_69_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201923)","author":"Zhou Yongchao","year":"2023","unstructured":"Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2023. Large language models are human-level prompt engineers. In Proceedings of the International Conference on Learning Representations (ICLR \u201923). OpenReview.net."},{"key":"e_1_3_3_70_2","doi-asserted-by":"crossref","unstructured":"Kaijie Zhu Jindong Wang Jiaheng Zhou Zichen Wang Hao Chen Yidong Wang Linyi Yang Wei Ye Neil Zhenqiang Gong Yue Zhang and Xing Xie. 2023. PromptBench: Towards evaluating the robustness of large language models on adversarial prompts. arXiv:2306.04528. Retrieved from https:\/\/arxiv.org\/abs\/2306.04528","DOI":"10.1145\/3689217.3690621"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3672456","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3672456","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:58:01Z","timestamp":1750294681000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3672456"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,27]]},"references-count":69,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2024,9,30]]}},"alternative-id":["10.1145\/3672456"],"URL":"https:\/\/doi.org\/10.1145\/3672456","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,9,27]]},"assertion":[{"value":"2023-10-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-05-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-09-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}