{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T19:24:12Z","timestamp":1784575452857,"version":"3.55.0"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2024,9,27]],"date-time":"2024-09-27T00:00:00Z","timestamp":1727395200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62192733, 61832009, 62192731, 62192730, and 62072007"],"award-info":[{"award-number":["62192733, 61832009, 62192731, 62192730, and 62072007"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key Program of Hubei","award":["JD2023008"],"award-info":[{"award-number":["JD2023008"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2024,9,30]]},"abstract":"<jats:p>Although large language models (LLMs) have demonstrated remarkable code-generation ability, they still struggle with complex tasks. In real-world software development, humans usually tackle complex tasks through collaborative teamwork, a strategy that significantly controls development complexity and enhances software quality. Inspired by this, we present a self-collaboration framework for code generation employing LLMs, exemplified by ChatGPT. Specifically, through role instructions, (1) Multiple LLM agents act as distinct \u201cexperts,\u201d each responsible for a specific subtask within a complex task; (2) Specify the way to collaborate and interact, so that different roles form a virtual team to facilitate each other\u2019s work, ultimately the virtual team addresses code generation tasks collaboratively without the need for human intervention. To effectively organize and manage this virtual team, we incorporate software-development methodology into the framework. Thus, we assemble an elementary team consisting of three LLM roles (i.e., analyst, coder, and tester) responsible for software development\u2019s analysis, coding, and testing stages. We conduct comprehensive experiments on various code-generation benchmarks. Experimental results indicate that self-collaboration code generation relatively improves 29.9\u201347.1% Pass@1 compared to the base LLM agent. Moreover, we showcase that self-collaboration could potentially enable LLMs to efficiently handle complex repository-level tasks that are not readily solved by the single LLM agent.<\/jats:p>","DOI":"10.1145\/3672459","type":"journal-article","created":{"date-parts":[[2024,6,12]],"date-time":"2024-06-12T21:28:32Z","timestamp":1718227712000},"page":"1-38","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":193,"title":["Self-Collaboration Code Generation via ChatGPT"],"prefix":"10.1145","volume":"33","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6228-4019","authenticated-orcid":false,"given":"Yihong","family":"Dong","sequence":"first","affiliation":[{"name":"Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education, Beijing, China and School of Computer Science, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5317-865X","authenticated-orcid":false,"given":"Xue","family":"Jiang","sequence":"additional","affiliation":[{"name":"Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education, Beijing, China and School of Computer Science, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1087-226X","authenticated-orcid":false,"given":"Zhi","family":"Jin","sequence":"additional","affiliation":[{"name":"Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education, Beijing, China and School of Computer Science, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5828-0186","authenticated-orcid":false,"given":"Ge","family":"Li","sequence":"additional","affiliation":[{"name":"Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education, Beijing, China and School of Computer Science, Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,9,27]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"Pekka Abrahamsson Outi Salo Jussi Ronkainen and Juhani Warsta. 2017. Agile software development methods: Review and analysis. arXiv:1709.08439. Retrieved from https:\/\/arxiv.org\/abs\/1709.08439"},{"key":"e_1_3_3_3_2","doi-asserted-by":"crossref","unstructured":"Chetan Arora John Grundy and Mohamed Abdelrazek. 2023. Advancing requirements engineering through generative AI: Assessing the role of LLMs. arXiv:2310.13976. Retrieved from https:\/\/arxiv.org\/abs\/2310.13976","DOI":"10.1007\/978-3-031-55642-5_6"},{"key":"e_1_3_3_4_2","unstructured":"Jacob Austin Augustus Odena Maxwell I. Nye Maarten Bosma Henryk Michalewski David Dohan Ellen Jiang Carrie J. Cai Michael Terry Quoc V. Le and Charles Sutton. 2021. Program synthesis with large language models. arXiv:2108.07732. Retrieved from https:\/\/arxiv.org\/abs\/2108.07732"},{"key":"e_1_3_3_5_2","unstructured":"Kent Beck Mike Beedle Arie Van Bennekum Alistair Cockburn Ward Cunningham Martin Fowler James Grenning Jim Highsmith Andrew Hunt Ron Jeffries Jon Kern Brian Marick Robert C. Martin Steve Mellor Ken Schwaber Jeff Sutherland and Dave Thomas. 2001. Manifesto for agile software development. (2001)."},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.4324\/9780080963242"},{"key":"e_1_3_3_7_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS 20)","volume":"33","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS 20), Vol. 33."},{"key":"e_1_3_3_8_2","unstructured":"Bei Chen Fengji Zhang Anh Nguyen Daoguang Zan Zeqi Lin Jian-Guang Lou and Weizhu Chen. 2022. CodeT: Code generation with generated tests. arXiv:2207.10397. Retrieved from https:\/\/arxiv.org\/abs\/2207.10397"},{"key":"e_1_3_3_9_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Pond\u00e9 de Oliveira Pinto Jared Kaplan Harrison Edwards Yuri Burda Nicholas Joseph Greg Brockman Alex Ray Raul Puri Gretchen Krueger Michael Petrov Heidy Khlaaf Girish Sastry Pamela Mishkin Brooke Chan Scott Gray Nick Ryder Mikhail Pavlov Alethea Power Lukasz Kaiser Mohammad Bavarian Clemens Winter Philippe Tillet Felipe Petroski Such Dave Cummings Matthias Plappert Fotios Chantzis Elizabeth Barnes Ariel Herbert-Voss William Hebgen Guss Alex Nichol Alex Paino Nikolas Tezak Jie Tang Igor Babuschkin Suchir Balaji Shantanu Jain William Saunders Christopher Hesse Andrew N. Carr Jan Leike Joshua Achiam Vedant Misra Evan Morikawa Alec Radford Matthew Knight Miles Brundage Mira Murati Katie Mayer Peter Welinder Bob McGrew Dario Amodei Sam McCandlish Ilya Sutskever and Wojciech Zaremba. 2021. Evaluating large language models trained on code. arXiv:2107.03374. Retrieved from https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_3_3_10_2","unstructured":"Xinyun Chen Maxwell Lin Nathanael Sch\u00e4rli and Denny Zhou. 2023. Teaching large language models to self-debug. arXiv:2304.05128. Retrieved from https:\/\/arxiv.org\/abs\/2304.05128"},{"key":"e_1_3_3_11_2","unstructured":"Aakanksha Chowdhery Sharan Narang Jacob Devlin Maarten Bosma Gaurav Mishra Adam Roberts Paul Barham Hyung Won Chung Charles Sutton Sebastian Gehrmann Parker Schuh Kensen Shi Sasha Tsvyashchenko Joshua Maynez Abhishek Rao Parker Barnes Yi Tay Noam Shazeer Vinodkumar Prabhakaran Emily Reif Nan Du Ben Hutchinson Reiner Pope James Bradbury Jacob Austin Michael Isard Guy Gur-Ari Pengcheng Yin Toju Duke Anselm Levskaya Sanjay Ghemawat Sunipa Dev Henryk Michalewski Xavier Garcia Vedant Misra Kevin Robinson Liam Fedus Denny Zhou Daphne Ippolito David Luan Hyeontaek Lim Barret Zoph Alexander Spiridonov Ryan Sepassi David Dohan Shivani Agrawal Mark Omernick Andrew M. Dai Thanumalayan Sankaranarayana Pillai Marie Pellat Aitor Lewkowycz Erica Moreira Rewon Child Oleksandr Polozov Katherine Lee Zongwei Zhou Xuezhi Wang Brennan Saeta Mark Diaz Orhan Firat Michele Catasta Jason Wei Kathy Meier-Hellstern Douglas Eck Jeff Dean Slav Petrov and Noah Fiedel. 2022. PaLM: Scaling language modeling with pathways. arXiv:2204.02311. Retrieved from https:\/\/arxiv.org\/abs\/2204.02311"},{"key":"e_1_3_3_12_2","unstructured":"Hyung W. Chung Le Hou Shayne Longpre Barret Zoph Yi Tay William Fedus Eric Li Xuezhi Wang Mostafa Dehghani Siddhartha Brahma Albert Webson Shixiang Shane Gu Zhuyun Dai Mirac Suzgun Xinyun Chen Aakanksha Chowdhery Sharan Narang Gaurav Mishra Adams Yu Vincent Y. Zhao Yanping Huang Andrew M. Dai Hongkun Yu Slav Petrov Ed H. Chi Jeff Dean Jacob Devlin Adam Roberts Denny Zhou Quoc V. Le and Jason Wei. 2022. Scaling instruction-finetuned language models. arXiv:2210.11416. Retrieved from https:\/\/arxiv.org\/abs\/2210.11416"},{"key":"e_1_3_3_13_2","first-page":"746","volume-title":"Proceedings of the 15th National\/10th Conference on Artificial Intelligence\/Innovative Applications of Artificial Intelligence (AAAI\/IAAI \u201998)","author":"Claus Caroline","year":"1998","unstructured":"Caroline Claus and Craig Boutilier. 1998. The dynamics of reinforcement learning in cooperative multiagent systems. In Proceedings of the 15th National\/10th Conference on Artificial Intelligence\/Innovative Applications of Artificial Intelligence (AAAI\/IAAI \u201998). AAAI Press \/ The MIT Press, 746\u2013752."},{"key":"e_1_3_3_14_2","volume-title":"Peopleware: Productive Projects and Teams","author":"DeMarco Tom","year":"2013","unstructured":"Tom DeMarco and Tim Lister. 2013. Peopleware: Productive Projects and Teams. Addison-Wesley."},{"key":"e_1_3_3_15_2","doi-asserted-by":"crossref","unstructured":"Yihong Dong Jiazheng Ding Xue Jiang Zhuo Li Ge Li and Zhi Jin. 2023. CodeScore: Evaluating code generation by learning code execution. arXiv:2301.09043. Retrieved from https:\/\/arxiv.org\/abs\/2301.09043","DOI":"10.1145\/3695991"},{"key":"e_1_3_3_16_2","doi-asserted-by":"crossref","unstructured":"Yihong Dong Xue Jiang Huanyu Liu Zhi Jin and Ge Li. 2024. Generalization or memorization: Data contamination and trustworthy evaluation for large language models. arXiv:2402.15938. Retrieved from https:\/\/arxiv.org\/abs\/2402.15938","DOI":"10.18653\/v1\/2024.findings-acl.716"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598048"},{"key":"e_1_3_3_18_2","unstructured":"Yihong Dong Kangcheng Luo Xue Jiang Zhi Jin and Ge Li. 2023. PACE: Improving prompt with actor-critic editing for large language model. arXiv:2308.10088. Retrieved from https:\/\/arxiv.org\/abs\/2308.10088"},{"key":"e_1_3_3_19_2","unstructured":"Daniel Fried Armen Aghajanyan Jessy Lin Sida Wang Eric Wallace Freda Shi Ruiqi Zhong Wen-Tau Yih Luke Zettlemoyer and Mike Lewis. 2022. InCoder: A Generative Model for Code Infilling and Synthesis. arXiv:2204.05999. Retrieved from https:\/\/arxiv.org\/abs\/2204.05999"},{"key":"e_1_3_3_20_2","unstructured":"GitHub. 2022. Copilot. Retrieved from https:\/\/github.com\/features\/copilot"},{"key":"e_1_3_3_21_2","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS Datasets and Benchmarks \u201921)","volume":"1","author":"Hendrycks Dan","year":"2021","unstructured":"Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. 2021. Measuring coding challenge competence with APPS. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS Datasets and Benchmarks \u201921), Vol. 1."},{"key":"e_1_3_3_22_2","unstructured":"Huggingface. 2023. HuggingChat. Retrieved from https:\/\/huggingface.co\/chat\/"},{"key":"e_1_3_3_23_2","unstructured":"IBM. 2023. Dromedary. Retrieved from https:\/\/huggingface.co\/ibm\/dromedary-65b-lora-delta-v0"},{"key":"e_1_3_3_24_2","doi-asserted-by":"crossref","unstructured":"Xue Jiang Yihong Dong Lecheng Wang Qiwei Shang and Ge Li. 2023. Self-planning code generation with large language model. arXiv:2303.06689. Retrieved from https:\/\/arxiv.org\/abs\/2303.06689","DOI":"10.1145\/3672456"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00324"},{"key":"e_1_3_3_26_2","doi-asserted-by":"crossref","unstructured":"Sungmin Kang Juyeon Yoon and Shin Yoo. 2022. Large language models are few-shot testers: Exploring LLM-based general bug reproduction. arXiv:2209.11515. Retrieved from https:\/\/arxiv.org\/abs\/2209.11515","DOI":"10.1109\/ICSE48619.2023.00194"},{"key":"e_1_3_3_27_2","volume-title":"The Wisdom of Teams: Creating the High-Performance Organization","author":"Katzenbach Jon R.","year":"2015","unstructured":"Jon R. Katzenbach and Douglas K. Smith. 2015. The Wisdom of Teams: Creating the High-Performance Organization. Harvard Business Review Press."},{"key":"e_1_3_3_28_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS \u201922)","volume":"35","author":"Kojima Takeshi","year":"2022","unstructured":"Takeshi Kojima, Shixiang S. Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS \u201922), Vol. 35."},{"key":"e_1_3_3_29_2","unstructured":"Guohao Li Hasan A. A. K. Hammoud Hani Itani Dmitrii Khizbullin and Bernard Ghanem. 2023. CAMEL: communicative agents for \u201cmind\u201d exploration of large scale language model society. arXiv:2303.17760. Retrieved from https:\/\/arxiv.org\/abs\/2303.17760"},{"key":"e_1_3_3_30_2","unstructured":"Raymond Li Loubna Ben Allal Yangtian Zi Niklas Muennighoff Denis Kocetkov Chenghao Mou Marc Marone Christopher Akiki Jia Li Jenny Chim Qian Liu Evgenii Zheltonozhskii Terry Yue Zhuo Thomas Wang Olivier Dehaene Mishig Davaadorj Joel Lamy-Poirier Jo\u00e3o Monteiro Oleh Shliazhko Nicolas Gontier Nicholas Meade Armel Zebaze Ming-Ho Yee Logesh Kumar Umapathi Jian Zhu Benjamin Lipkin Muhtasham Oblokulov Zhiruo Wang Rudra Murthy Jason Stillerman Siva Sankalp Patel Dmitry Abulkhanov Marco Zocca Manan Dey Zhihan Zhang Nour Moustafa-Fahmy Urvashi Bhattacharyya Wenhao Yu Swayam Singh Sasha Luccioni Paulo Villegas Maxim Kunakov Fedor Zhdanov Manuel Romero Tony Lee Nadav Timor Jennifer Ding Claire Schlesinger Hailey Schoelkopf Jan Ebert Tri Dao Mayank Mishra Alex Gu Jennifer Robinson Carolyn Jane Anderson Brendan Dolan-Gavitt Danish Contractor Siva Reddy Daniel Fried Dzmitry Bahdanau Yacine Jernite Carlos Mu\u00f1noz Ferrandis Sean Hughes Thomas Wolf Arjun Guha Leandro von Werra and Harm de Vries. 2023. StarCoder: May the source be with you! arXiv:2305.06161. Retrieved from https:\/\/arxiv.org\/abs\/2305.06161"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.abq1158"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3560815"},{"key":"e_1_3_3_33_2","unstructured":"LMSYS. 2023. Fastchat. Retrieved from https:\/\/huggingface.co\/lmsys\/fastchat-t5-3b-v1.0"},{"key":"e_1_3_3_34_2","unstructured":"LMSYS. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. Retrieved from https:\/\/lmsys.org\/blog\/2023-03-30-vicuna\/"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2003.10.001"},{"key":"e_1_3_3_36_2","volume-title":"The Emotion Machine: Commonsense Thinking, Artificial Intelligence, and the Future of the Human Mind","author":"Minsky Marvin","year":"2007","unstructured":"Marvin Minsky. 2007. The Emotion Machine: Commonsense Thinking, Artificial Intelligence, and the Future of the Human Mind. Simon and Schuster."},{"key":"e_1_3_3_37_2","unstructured":"MosaicML. 2023. Introducing MPT-7B: A New Standard for Open-Source ly Usable LLMs. Retrieved from www.mosaicml.com\/blog\/mpt-7b"},{"key":"e_1_3_3_38_2","unstructured":"H. Penny Nii. 1986. Blackboard Systems. AI Mag. 7 (1986)."},{"key":"e_1_3_3_39_2","unstructured":"Erik Nijkamp Bo Pang Hiroaki Hayashi Lifu Tu Huan Wang Yingbo Zhou Silvio Savarese and Caiming Xiong. 2022. Codegen: An open large language model for code with multi-turn program synthesis. arXiv:2203.13474. Retrieved from https:\/\/arxiv.org\/abs\/2203.13474"},{"key":"e_1_3_3_40_2","unstructured":"OpenAI. 2022. ChatGPT. Retrieved from https:\/\/openai.com\/blog\/chatgpt\/"},{"key":"e_1_3_3_41_2","unstructured":"OpenAI. 2023. GPT-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_3_42_2","unstructured":"Long Ouyang Jeff Wu Xu Jiang Diogo Almeida Carroll L. Wainwright Pamela Mishkin Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray John Schulman Jacob Hilton Fraser Kelton Luke Miller Maddie Simens Amanda Askell Peter Welinder Paul F. Christiano Jan Leike and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. arXiv:2203.02155. Retrieved from https:\/\/arxiv.org\/abs\/2203.02155"},{"key":"e_1_3_3_43_2","unstructured":"Long Ouyang Jeff Wu Xu Jiang Diogo Almeida Carroll L. Wainwright Pamela Mishkin Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray John Schulman Jacob Hilton Fraser Kelton Luke Miller Maddie Simens Amanda Askell Peter Welinder Paul F. Christiano Jan Leike and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. arXiv:2203.02155. Retrieved from https:\/\/arxiv.org\/abs\/2203.02155"},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-02152-7_29"},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3411763.3451760"},{"key":"e_1_3_3_46_2","unstructured":"Baptiste Rozi\u00e8re Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing Ellen Tan Yossi Adi Jingyu Liu Tal Remez J\u00e9r\u00e9my Rapin Artyom Kozhevnikov Ivan Evtimov Joanna Bitton Manish Bhatt Cristian Canton-Ferrer Aaron Grattafiori Wenhan Xiong Alexandre D\u00e9fossez Jade Copet Faisal Azhar Hugo Touvron Louis Martin Nicolas Usunier Thomas Scialom and Gabriel Synnaeve. 2023. Code Llama: Open foundation models for code. arXiv:2308.12950. Retrieved from https:\/\/arxiv.org\/abs\/2308.12950"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/1764810.1764814"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2023.3334955"},{"key":"e_1_3_3_49_2","unstructured":"Timo Schick Jane Dwivedi-Yu Zhengbao Jiang Fabio Petroni Patrick S. H. Lewis Gautier Izacard Qingfei You Christoforos Nalmpantis Edouard Grave and Sebastian Riedel. 2022. PEER: A collaborative language model. arXiv:2208.11663. Retrieved from https:\/\/arxiv.org\/abs\/2208.11663"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3558965"},{"key":"e_1_3_3_51_2","unstructured":"Yongliang Shen Kaitao Song Xu Tan Dongsheng Li Weiming Lu and Yueting Zhuang. 2023. HuggingGPT: Solving AI tasks with chatGPT and its friends in huggingface. arXiv:2303.17580. Retrieved from https:\/\/arxiv.org\/abs\/2303.17580"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.346"},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(91)90035-I"},{"key":"e_1_3_3_54_2","unstructured":"Chang-You Tai Ziru Chen Tianshu Zhang Xiang Deng and Huan Sun. 2023. Exploring chain-of-thought style prompting for text-to-SQL. arXiv:2305.14215. Retrieved from https:\/\/arxiv.org\/abs\/2305.14215"},{"key":"e_1_3_3_55_2","unstructured":"THUDM. 2023. ChatGLM. Retrieved from https:\/\/huggingface.co\/THUDM\/chatglm-6b"},{"key":"e_1_3_3_56_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale Dan Bikel Lukas Blecher Cristian Canton-Ferrer Moya Chen Guillem Cucurull David Esiobu Jude Fernandes Jeremy Fu Wenyin Fu Brian Fuller Cynthia Gao Vedanuj Goswami Naman Goyal Anthony Hartshorn Saghar Hosseini Rui Hou Hakan Inan Marcin Kardas Viktor Kerkez Madian Khabsa Isabel Kloumann Artem Korenev Punit Singh Koura Marie-Anne Lachaux Thibaut Lavril Jenya Lee Diana Liskovich Yinghai Lu Yuning Mao Xavier Martinet Todor Mihaylov Pushkar Mishra Igor Molybog Yixin Nie Andrew Poulton Jeremy Reizenstein Rashi Rungta Kalyan Saladi Alan Schelten Ruan Silva Eric Michael Smith Ranjan Subramanian Xiaoqing Ellen Tan Binh Tang Ross Taylor Adina Williams Jian Xiang Kuan Puxin Xu Zheng Yan Iliyan Zarov Yuchen Zhang Angela Fan Melanie Kambadur Sharan Narang Aur\u00e9lien Rodriguez Robert Stojnic Sergey Edunov and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_3_57_2","unstructured":"Jason Wei Xuezhi Wang Dale Schuurmans Maarten Bosma Ed Chi Quoc Le and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv:\/2201.11903. Retrieved from https:\/\/arxiv.org\/abs\/2201.11903"},{"key":"e_1_3_3_58_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS \u201922)","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS \u201922), Vol. 35."},{"key":"e_1_3_3_59_2","unstructured":"Chenfei Wu Shengming Yin Weizhen Qi Xiaodong Wang Zecheng Tang and Nan Duan. 2023. Visual chatGPT: Talking drawing and editing with visual foundation models. arXiv:2303.04671. Retrieved from https:\/\/arxiv.org\/abs\/2303.04671"},{"key":"e_1_3_3_60_2","unstructured":"Chunqiu Steven Xia Yuxiang Wei and Lingming Zhang. 2022. Practical program repair in the era of large pre-trained language models. arXiv:2210.14179. Retrieved from https:\/\/arxiv.org\/abs\/2210.14179"},{"key":"e_1_3_3_61_2","unstructured":"Chunqiu S. Xia and Lingming Zhang. 2023. Conversational automated program repair. arXiv:2301.13246. Retrieved from https:\/\/arxiv.org\/abs\/2301.13246"},{"key":"e_1_3_3_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3623316"},{"key":"e_1_3_3_63_2","unstructured":"Eric Zelikman Qian Huang Gabriel Poesia Noah D. Goodman and Nick Haber. 2022. Parsel: A unified natural language framework for algorithmic reasoning. arXiv:2212.10561. Retrieved from https:\/\/arxiv.org\/abs\/2212.10561"},{"key":"e_1_3_3_64_2","unstructured":"Tianyi Zhang Tao Yu Tatsunori B. Hashimoto Mike Lewis Wen-tau Yih Daniel Fried and Sida I. Wang. 2022. Coder reviewer reranking for code generation. arXiv:2211.16490. Retrieved from https:\/\/arxiv.org\/abs\/2211.16490"},{"key":"e_1_3_3_65_2","doi-asserted-by":"crossref","unstructured":"Qinkai Zheng Xiao Xia Xu Zou Yuxiao Dong Shan Wang Yufei Xue Zi-Yuan Wang Lei Shen Andi Wang Yang Li Teng Su Zhilin Yang and Jie Tang. 2023. CodeGeeX: A pre-trained model for code generation with multilingual evaluations on humaneval-X. arXiv:2303.17568. Retrieved from https:\/\/arxiv.org\/abs\/2303.17568","DOI":"10.1145\/3580305.3599790"},{"key":"e_1_3_3_66_2","unstructured":"Yongchao Zhou Andrei Ioan Muresanu Ziwen Han Keiran Paster Silviu Pitis Harris Chan and Jimmy Ba. 2022. Large language models are human-level prompt engineers. arXiv:2211.01910. Retrieved from https:\/\/arxiv.org\/abs\/2211.01910"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3672459","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3672459","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:58:01Z","timestamp":1750294681000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3672459"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,27]]},"references-count":65,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2024,9,30]]}},"alternative-id":["10.1145\/3672459"],"URL":"https:\/\/doi.org\/10.1145\/3672459","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,9,27]]},"assertion":[{"value":"2023-11-06","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-05-10","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-09-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}