{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T11:21:16Z","timestamp":1787484076057,"version":"build-2736575974"},"reference-count":96,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:00:00Z","timestamp":1750291200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,19]]},"abstract":"<jats:p>Code generation has largely improved development efficiency in the era of large language models (LLMs). With the ability to follow instructions, current LLMs can be prompted to generate code solutions given detailed descriptions in natural language. Many research efforts are being devoted to improving the correctness of LLM-generated code, and many benchmarks are proposed to evaluate the correctness comprehensively. Despite the focus on correctness, the time efficiency of LLM-generated code solutions is under-explored. Current correctness benchmarks are not suitable for time efficiency evaluation since their test cases cannot well distinguish the time efficiency of different code solutions. Besides, the current execution time measurement is not stable and comprehensive, threatening the validity of the time efficiency evaluation.<\/jats:p>\n                  <jats:p>To address the challenges in the time efficiency evaluation of code generation, we propose COFFE, a code generation benchmark for evaluating the time efficiency of LLM-generated code solutions. COFFE contains 398 and 358 problems for function-level and file-level code generation, respectively. To improve the distinguishability, we design a novel stressful test case generation approach with contracts and two new formats of test cases to improve the accuracy of generation. For the time evaluation metric, we propose efficienct@k based on CPU instruction count to ensure a stable and solid comparison between different solutions. We evaluate 14 popular LLMs on COFFE and identify four findings. Based on the findings, we draw some implications for LLM researchers and software practitioners to facilitate future research and usage of LLMs in code generation.<\/jats:p>","DOI":"10.1145\/3715727","type":"journal-article","created":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T11:16:02Z","timestamp":1750331762000},"page":"242-265","source":"Crossref","is-referenced-by-count":13,"title":["COFFE: A Code Efficiency Benchmark for Code Generation"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1936-5598","authenticated-orcid":false,"given":"Yun","family":"Peng","sequence":"first","affiliation":[{"name":"The Chinese University of Hong Kong, HongKong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-3294-688X","authenticated-orcid":false,"given":"Jun","family":"Wan","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-8370-644X","authenticated-orcid":false,"given":"Yichen","family":"Li","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5526-1617","authenticated-orcid":false,"given":"Xiaoxue","family":"Ren","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,19]]},"reference":[{"key":"e_1_3_1_1_2","doi-asserted-by":"publisher","unstructured":"Marah I Abdin Sam Ade Jacobs Ammar Ahmad Awan Jyoti Aneja Ahmed Awadallah Hany Awadalla Nguyen Bach Amit Bahree Arash Bakhtiari et al.. 2024. Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone. CoRR abs\/2404.14219 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2404.14219 10.48550\/ARXIV.2404.14219 arXiv:2404.14219","DOI":"10.48550\/ARXIV.2404.14219"},{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","unstructured":"Loubna Ben Allai Raymond Li Denis Kocetkov Chenghao Mou Christopher Akiki Carlos Munoz Ferrandis Niklas Muennighoff Mayank Mishra Alex Gu Manan Dey Logesh Kumar Umapathi et al.. 2023. SantaCoder: don\u2019t reach for the stars!. CoRR abs\/2301.03988 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2301.03988 10.48550\/ARXIV.2301.03988 arXiv:2301.03988","DOI":"10.48550\/ARXIV.2301.03988"},{"key":"e_1_3_1_3_2","unstructured":"Anthropic. 2024. API reference provided by Anthropic. https:\/\/docs.anthropic.com\/en\/api\/getting-started."},{"key":"e_1_3_1_4_2","unstructured":"Anthropic. 2024. Claude 3.5 Sonnet. https:\/\/www.anthropic.com\/news\/claude-3-5-sonnet."},{"key":"e_1_3_1_5_2","unstructured":"Jacob Austin Augustus Odena Maxwell I. Nye Maarten Bosma Henryk Michalewski David Dohan Ellen Jiang Carrie J. Cai Michael Terry Quoc V. Le and Charles Sutton. 2021. Program Synthesis with Large Language Models. CoRR abs\/2108.07732 (2021). arXiv:2108.07732 https:\/\/arxiv.org\/abs\/2108.07732"},{"key":"e_1_3_1_6_2","unstructured":"Ned Batchelder. 2024. The Coverage.py library. https:\/\/github.com\/nedbat\/coveragepy https:\/\/github.com\/nedbat\/coveragepy."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","unstructured":"Shreya Bhatia Tarushi Gandhi Dhruv Kumar and Pankaj Jalote. 2023. Unit Test Generation using Generative Al : A Comparative Performance Analysis of Autogeneration Tools. CoRR abs\/2312.10622 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2312.10622 10.48550\/ARXIV.2312.10622 arXiv:2312.10622","DOI":"10.48550\/ARXIV.2312.10622"},{"key":"e_1_3_1_8_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Pond\u00e9 de Oliveira Pinto Jared Kaplan Harrison Edwards Yuri Burda Nicholas Joseph Greg Brockman et al.. 2021. Evaluating Large Language Models Trained on Code. CoRR abs\/2107.03374 (2021). arXiv:2107.03374 https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_3_1_9_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Pond\u00e9 de Oliveira Pinto Jared Kaplan Harrison Edwards Yuri Burda Nicholas Joseph Greg Brockman Alex Ray et al.. 2021. Evaluating Large Language Models Trained on Code. CoRR abs\/2107.03374 (2021). arXiv:2107.03374 https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","unstructured":"Xinyun Chen Maxwell Lin Nathanael Sch\u00e4rli and Denny Zhou. 2023. Teaching Large Language Models to Self-Debug. CoRR abs\/2304.05128 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2304.05128 10.48550\/ARXIV.2304.05128 arXiv:2304.05128","DOI":"10.48550\/ARXIV.2304.05128"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","unstructured":"Yinghao Chen Zehao Hu Chen Zhi Junxiao Han Shuiguang Deng and Jianwei Yin. 2024. ChatUniTest: A Framework for LLM-Based Test Generation. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering FSE 2024 Porto de Galinhas Brazil July 15-19 2024. ACM 572\u2013576. https:\/\/doi.org\/10.1145\/3663529.3663801 10.1145\/3663529.3663801","DOI":"10.1145\/3663529.3663801"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","unstructured":"Hyung Won Chung Le Hou Shayne Longpre Barret Zoph Yi Tay William Fedus Eric Li Xuezhi Wang et al.. 2022. Scaling Instruction-Finetuned Language Models. CoRR abs\/2210.11416 (2022).https:\/\/doi.org\/10.48550\/ARXIV.2210.11416 10.48550\/ARXIV.2210.11416 arXiv:2210.11416","DOI":"10.48550\/ARXIV.2210.11416"},{"key":"e_1_3_1_13_2","unstructured":"The MITRE Corporation. 2024. Performance efficiency CWEs. https:\/\/cwe.mitre.org\/data\/definitions\/1132.html."},{"key":"e_1_3_1_14_2","unstructured":"Deepmind. 2024. Gemini 1.5 Pro. https:\/\/deepmind.google\/technologies\/gemini\/pro\/."},{"key":"e_1_3_1_15_2","unstructured":"DeepSeek. 2024. DeepSeek API. https:\/\/platform.deepseek.com\/ https:\/\/platform.deepseek.com\/."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","unstructured":"DeepSeek-AI et al.. 2024. DeepSeek-V2: A Strong Economical and Efficient Mixture-of-Experts Language Model. CoRR abs\/2405.04434 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2405.04434 10.48550\/ARXIV.2405.04434 arXiv:2405.04434","DOI":"10.48550\/ARXIV.2405.04434"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","unstructured":"Yinlin Deng Chunqiu Steven Xia Haoran Peng Chenyuan Yang and Lingming Zhang. 2023. Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language Models. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis ISSTA 2023 Seattle WA USA July 17-21 2023. ACM 423\u2013435. https:\/\/doi.org\/10.1145\/3597926.3598067 10.1145\/3597926.3598067","DOI":"10.1145\/3597926.3598067"},{"key":"e_1_3_1_18_2","unstructured":"Yangruibo Ding Zijian Wang Wasi Uddin Ahmad Hantian Ding Ming Tan Nihal Jain Murali Krishna Ramanathan Ramesh Nallapati Parminder Bhatia Dan Roth and Bing Xiang. 2023. CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023 NeurlPS 2023 New Orleans LA USA December 10-16 2023. http:\/\/papers. nips.cc\/paper_files\/paper\/2023\/hash\/920f2dced7d32ab2ba2fl970bc306af6-Abstract-Datasets_and_Benchmarks.html"},{"key":"e_1_3_1_19_2","unstructured":"Docker. 2024. Docker. https:\/\/www.docker.com\/ https:\/\/www.docker.com\/."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3660791"},{"key":"e_1_3_1_21_2","unstructured":"The Linux Foundation. 2024. The perf tool on linux. https:\/\/perf.wiki.kernel.org\/index.php\/Main_Page."},{"key":"e_1_3_1_22_2","unstructured":"Daniel Fried Armen Aghajanyan Jessy Lin Sida Wang Eric Wallace Freda Shi Ruiqi Zhong Scott Yih Luke Zettlemoyer and Mike Lewis. 2023. InCoder: A Generative Model for Code Infilling and Synthesis. In The Eleventh International Conference on Learning Representations ICLR 2023 Kigali Rwanda May 1-5 2023. OpenReview.net. https:\/\/openreview.net\/pdf?id=hQwb-lbM6EL"},{"key":"e_1_3_1_23_2","unstructured":"Inc. GitHub. 2022. GitHub Octoverse report on programming languages. https:\/\/octoverse.github.com\/2022\/top-programming-languages."},{"key":"e_1_3_1_24_2","unstructured":"Google. 2023. Sanitized version of MBPP benchmark released by Google. https:\/\/huggingface.co\/datasets\/google-research- datasets\/mbpp\/viewer\/sanitized\/test https:\/\/huggingface.co\/datasets\/google-research- datasets\/mbpp\/viewer\/sanitized\/test."},{"key":"e_1_3_1_25_2","unstructured":"Google. 2024. API reference provided by Google. https:\/\/ai.google.dev\/gemini-api\/docs\/models\/gemini https:\/\/ai.google.dev\/gemini-api\/docs\/models\/gemini."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","unstructured":"Daya Guo Qihao Zhu Dejian Yang Zhenda Xie Kai Dong Wentao Zhang Guanting Chen Xiao Bi Y. Wu Y. K. Li Fuli Luo Yingfei Xiong and Wenfeng Liang. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming - The Rise of Code Intelligence. CoRR abs\/2401.14196 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2401.14196 10.48550\/ARXIV.2401.14196 arXiv:2401.14196","DOI":"10.48550\/ARXIV.2401.14196"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","unstructured":"Nam Le Hai Dung Manh Nguyen and Nghi D. Q. Bui. 2024. REPOEXEC: Evaluate Code Generation with a Repository- Level Executable Benchmark. CoRR abs\/2406.11927 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2406.1192710.48550\/ARXIV.2406.11927 arXiv:2406.11927","DOI":"10.48550\/ARXIV.2406.11927"},{"key":"e_1_3_1_28_2","unstructured":"Dan Hendrycks Steven Basart Saurav Kadavath Mantas Mazeika Akul Arora Ethan Guo Collin Burns Samir Puranik Horace He Dawn Song and Jacob Steinhardt. 2021. Measuring Coding Challenge Competence With APPS. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 NeurlPS Datasets and Benchmarks 2021 December 2021 virtual. https:\/\/datasets-benchmarks-proceedings.neurips.cc\/paper\/2021\/hash\/c24cd76elce41366a4bbe8a49b02a028-Abstract-round2.html"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","unstructured":"Samuel Holt Max Ruiz Luyten and Mihaela van der Schaar. 2023. L2MAC: Large Language Model Automatic Computer for Unbounded Code Generation. CoRR abs\/2310.02003 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2310.02003 10.48550\/ARXIV.2310.02003 arXiv:2310.02003","DOI":"10.48550\/ARXIV.2310.02003"},{"key":"e_1_3_1_30_2","unstructured":"Sirui Hong Mingchen Zhuge Jonathan Chen Xiawu Zheng Yuheng Cheng Jinlin Wang Ceyao Zhang Zili Wang Steven Ka Shing Yau et al.. 2024. MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. In The Twelfth International Conference on Learning Representations ICLR 2024 Vienna Austria May 7-11 2024. OpenReview.net. https:\/\/openreview.net\/forum?id=VtmBAGCN7o"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","unstructured":"Soneya Binta Hossain and Matthew B. Dwyer. 2024. TOGLL: Correct and Strong Test Oracle Generation with LLMs. CoRR abs\/2405.03786 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2405.03786 10.48550\/ARXIV.2405.03786 arXiv:2405.03786","DOI":"10.48550\/ARXIV.2405.03786"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","unstructured":"Dong Huang Qingwen Bu Jie M. Zhang Michael Luck and Heming Cui. 2023. AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation. CoRR abs\/2312.13010 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2312.13010 10.48550\/ARXIV.2312.13010 arXiv:2312.13010","DOI":"10.48550\/ARXIV.2312.13010"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","unstructured":"Dong Huang Jie M. Zhang Yuhao Qing and Heming Cui. 2024. EffiBench: Benchmarking the Efficiency of Automatically Generated Code. CoRR abs\/2402.02037 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2402.02037 10.48550\/ARXIV.2402.02037 arXiv:2402.02037","DOI":"10.48550\/ARXIV.2402.02037"},{"key":"e_1_3_1_34_2","unstructured":"Deep Infra. 2024. Deep Infra API. https:\/\/deepinfra.com\/ https:\/\/deepinfra.com\/."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Antoine Roux Arthur Mensch Blanche Savary Chris Bamford Devendra Singh Chaplot Diego de Las Casas Emma Bou Hanna Florian Bressand et al.. 2024. Mixtral of Experts. CoRR abs\/2401.04088 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2401.04088 10.48550\/ARXIV.2401.04088 arXiv:2401.04088","DOI":"10.48550\/ARXIV.2401.04088"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","unstructured":"Carlos E. Jimenez John Yang Alexander Wettig Shunyu Yao Kexin Pei Ofir Press and Karthik Narasimhan. 2023. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. CoRR abs\/2310.06770 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2310.06770 10.48550\/ARXIV.2310.06770 arXiv:2310.06770","DOI":"10.48550\/ARXIV.2310.06770"},{"key":"e_1_3_1_37_2","unstructured":"Rabimba Karanjai Aftab Hussain Md Rafiqul Islam Rabin Lei Xu Weidong Shi and Mohammad Amin Alipour. 2024. Harnessing the Power of LLMs: Automating Unit Test Generation for High-Performance Computing. arXiv:2407.05202 [cs.SE] https:\/\/arxiv.org\/abs\/2407.05202"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","unstructured":"Mohammad Abdullah Matin Khan M. Saiful Bari Xuan Long Do Weishi Wang Md. Rizwan Parvez and Shafiq R. Joty. 2023. xCodeEval: A Large Scale Multilingual Multitask Benchmark for Code Understanding Generation Translation and Retrieval. CoRR abs\/2303.03004 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2303.03004 10.48550\/ARXIV.2303.03004 arXiv:2303.03004","DOI":"10.48550\/ARXIV.2303.03004"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409683"},{"key":"e_1_3_1_40_2","unstructured":"Hung Le Yue Wang Akhilesh Deepak Gotmare Silvio Savarese and Steven Chu-Hong Hoi. 2022. CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022 NeurlPS 2022 New Orleans LA USA November 28 - December 9 2022. http:\/\/papers.nips.ee\/paper_files\/paper\/2022\/hash\/8636419dealaa9fbd25fc4248e702da4-Abstract-Conference.html"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","unstructured":"Kefan Li and Yuan Yuan. 2024. Large Language Models as Test Case Generators: Performance Evaluation and Enhancement. CoRR abs\/2404.13340 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2404.13340 10.48550\/ARXIV.2404.13340 arXiv:2404.13340","DOI":"10.48550\/ARXIV.2404.13340"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","unstructured":"Raymond Li Loubna Ben Allai Yangtian Zi Niklas Muennighoff Denis Kocetkov Chenghao Mou Marc Marone Christopher Akiki Jia Li Jenny Chim et al.. 2023. StarCoder: may the source be with you!. CoRR abs\/2305.06161 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2305.06161 10.48550\/ARXIV.2305.06161 arXiv:2305.06161","DOI":"10.48550\/ARXIV.2305.06161"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.abq1158"},{"key":"e_1_3_1_44_2","unstructured":"Jiawei Liu Thanh Nguyen Mingyue Shang Hantian Ding Xiaopeng Li Yu Yu Varun Kumar and Zijian Wang. 2024. Learning Code Preference via Synthetic Evolution. arXiv:2410.03837 [cs.LG] https:\/\/arxiv.org\/abs\/2410.03837"},{"key":"e_1_3_1_45_2","unstructured":"Jiawei Liu Chunqiu Steven Xia Yuyao Wang and LINGMING ZHANG. 2023. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. In Thirty-seventh Conference on Neural Information Processing Systems. https:\/\/openreview.net\/forum?id=Mqvx610Cu7"},{"key":"e_1_3_1_46_2","unstructured":"Jiawei Liu Chunqiu Steven Xia Yuyao Wang and LINGMING ZHANG. 2023. The MBPP Plus benchmark. https:\/\/github.com\/evalplus\/evalplus\/releases\/tag\/v0.2.1 https:\/\/github.com\/evalplus\/evalplus\/releases\/tag\/v0.2.1."},{"key":"e_1_3_1_47_2","unstructured":"Jiawei Liu Chunqiu Steven Xia Yuyao Wang and LINGMING ZHANG. 2024. EvalPlus Leaderboard. https:\/\/evalplus.github.io\/leaderboard.html https:\/\/evalplus.github.io\/leaderboard.html."},{"key":"e_1_3_1_48_2","unstructured":"Jiawei Liu Songrun Xie Junhao Wang Yuxiang Wei Yifeng Ding and Lingming Zhang. 2024. Evaluating Language Models for Efficient Code Generation. arXiv:2408.06450 [cs.SE] https:\/\/arxiv.org\/abs\/2408.06450"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","unstructured":"Kaibo Liu Yiyang Liu Zhenpeng Chen Jie M. Zhang Yudong Han Yun Ma Ge Li and Gang Huang. 2024. LLM- Powered Test Case Generation for Detecting Tricky Bugs. CoRR abs\/2404.10304 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2404.10304 10.48550\/ARXIV.2404.10304 arXiv:2404.10304","DOI":"10.48550\/ARXIV.2404.10304"},{"key":"e_1_3_1_50_2","unstructured":"Shangqing Liu Yu Chen Xiaofei Xie Jing Kai Siow and Yang Liu. 2021. Retrieval-Augmented Generation for Code Summarization via Hybrid GNN. In 9th International Conference on Learning Representations ICLR 2021 Virtual Event Austria May 3-7 2021. OpenReview.net. https:\/\/openreview.net\/forum?id=zv-typlgPxA"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2306.03091"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","unstructured":"Anton Lozhkov Raymond Li Loubna Ben Allai Federico Cassano Joel Lamy-Poirier Nouamane Tazi Ao Tang Dmytro Pykhtar Jiawei Liu Yuxiang Wei et al.. 2024. StarCoder 2 and The Stack v2: The Next Generation. CoRR abs\/2402.19173 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2402.19173 10.48550\/ARXIV.2402.19173 arXiv:2402.19173","DOI":"10.48550\/ARXIV.2402.19173"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","unstructured":"Ziyang Luo Can Xu Pu Zhao Qingfeng Sun Xiubo Geng Wenxiang Hu Chongyang Tao Jing Ma Qingwei Lin and Daxin Jiang. 2023. WizardCoder: Empowering Code Large Language Models with Evol-Instruct. CoRR abs\/2306.08568 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2306.08568 10.48550\/ARXIV.2306.08568 arXiv:2306.08568","DOI":"10.48550\/ARXIV.2306.08568"},{"key":"e_1_3_1_54_2","unstructured":"Meta. 2024. Llama3. https:\/\/ai.meta.com\/blog\/meta-llama-3\/ https:\/\/ai.meta.com\/blog\/meta-llama-3\/."},{"key":"e_1_3_1_55_2","unstructured":"Meta. 2024. Llama3.1. https:\/\/ai.meta.com\/blog\/meta-llama-3-1\/ https:\/\/ai.meta.com\/blog\/meta-llama-3-1\/."},{"key":"e_1_3_1_56_2","unstructured":"Niklas Muennighoff Qian Liu Armel Randy Zebaze Qinkai Zheng Binyuan Hui Terry Yue Zhuo Swayam Singh Xiangru Tang Leandro von Werra and Shayne Longpre. 2024. OctoPack: Instruction Tuning Code Large Language Models. In The Twelfth International Conference on Learning Representations ICLR 2024 Vienna Austria May 7-11 2024. OpenReview.net. https:\/\/openreview.net\/forum?id=mwlPWNSWZP"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","unstructured":"Niels M\u00fcndler Mark Niklas M\u00fcller Jingxuan He and Martin T. Vechev. 2024. Code Agents are State of the Art Software Testers. CoRR abs\/2406.12952 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2406.12952 10.48550\/ARXIV.2406.12952 arXiv:2406.12952","DOI":"10.48550\/ARXIV.2406.12952"},{"key":"e_1_3_1_58_2","unstructured":"Ansong Ni Srini Iyer Dragomir Radev Veselin Stoyanov Wen-Tau Yih Sida I. Wang and Xi Victoria Lin. 2023. LEVER: Learning to Verify Language-to-Code Generation with Execution. In International Conference on Machine Learning ICML 2023 23-29 July 2023 Honolulu Hawaii USA (Proceedings of Machine Learning Research Vol. 202). PMLR 26106\u201326128. https:\/\/proceedings.mlr.press\/v202\/ni23b.html"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","unstructured":"Erik Nijkamp Hiroaki Hayashi Caiming Xiong Silvio Savarese and Yingbo Zhou. 2023. CodeGen2: Lessons for Training LLMs on Programming and Natural Languages. CoRR abs\/2305.02309 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2305.02309 10.48550\/ARXIV.2305.02309 arXiv:2305.02309","DOI":"10.48550\/ARXIV.2305.02309"},{"key":"e_1_3_1_60_2","unstructured":"Erik Nijkamp Bo Pang Hiroaki Hayashi Lifu Tu Huan Wang Yingbo Zhou Silvio Savarese and Caiming Xiong. 2023. CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. In The Eleventh International Conference on Learning Representations ICLR 2023 Kigali Rwanda May 1-5 2023. OpenReview.net. https:\/\/openreview.net\/pdf?id=MaYcJKpY2B_"},{"key":"e_1_3_1_61_2","unstructured":"OpenAI. 2022. ChatGPT. https:\/\/openai.com\/blog\/chatgpt."},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2303.08774"},{"key":"e_1_3_1_63_2","unstructured":"OpenAI. 2024. GPT-4o. https:\/\/openai.com\/index\/hello-gpt-4o\/ https:\/\/openai.com\/index\/hello-gpt-4o\/."},{"key":"e_1_3_1_64_2","unstructured":"OpenAI. 2024. OpenAI API. https:\/\/openai.com\/api\/ https:\/\/openai.com\/api\/."},{"key":"e_1_3_1_65_2","unstructured":"Long Ouyang Jeffrey Wu Xu Jiang Diogo Almeida Carroll L. Wainwright Pamela Mishkin Chong Zhang et al.. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022 NeurlPS 2022 New Orleans LA USA November 28 - December 9 2022. http:\/\/papers.nips.cc\/paper_files\/paper\/2022\/hash\/blefde53be364a73914f58805a001731-Abstract-Conference.html"},{"key":"e_1_3_1_66_2","unstructured":"Wendk\u00fbuni C. Ou\u00e9draogo Kader Kabor\u00e9 Haoye Tian Yewei Song Anil Koyuncu Jacques Klein David Lo and Tegawend\u00e9 F. Bissyand\u00e9. 2024. Large-scale Independent and Comprehensive study of the power of LLMs for test case generation. arXiv:2407.00225 [cs.SE] https:\/\/arxiv.org\/abs\/2407.00225"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.18653\/V1\/2021.FINDINGS-EMNLP.232"},{"key":"e_1_3_1_68_2","volume-title":"Computer Organization and Design - The Hardware \/ Software Interface (Revised 4th Edition)","author":"Patterson David A.","year":"2012","unstructured":"David A. Patterson and John L. Hennessy. 2012. Computer Organization and Design - The Hardware \/ Software Interface (Revised 4th Edition). Academic Press. http:\/\/www.elsevierdirect.com\/product.jsp?isbn=9780123747501"},{"key":"e_1_3_1_69_2","unstructured":"Yun Peng Akhilesh Deepak Gotmare Michael Lyu Caiming Xiong Silvio Savarese and Doyen Sahoo. 2024. Per-fCodeGen: Improving Performance of LLM Generated Code with Execution Feedback. arXiv:2412.03578 [cs.SE] https:\/\/arxiv.org\/abs\/2412.03578"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","unstructured":"Tal Ridnik Dedy Kredo and Itamar Friedman. 2024. Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering. CoRR abs\/2401.08500 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2401.08500 10.48550\/ARXIV.2401.08500 arXiv:2401.08500","DOI":"10.48550\/ARXIV.2401.08500"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2308.12950"},{"key":"e_1_3_1_72_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2406.07021"},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","unstructured":"Max Sch\u00e4fer Sarah Nadi Aryaz Eghbali and Frank Tip. 2024. An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation. IEEE Transactions on Software Engineering 50 1 (2024) 85\u2013105. https:\/\/doi.org\/10.1109\/TSE.2023.3334955 10.1109\/TSE.2023.3334955","DOI":"10.1109\/TSE.2023.3334955"},{"key":"e_1_3_1_74_2","unstructured":"Noah Shinn Federico Cassano Ashwin Gopinath Karthik Narasimhan and Shunyu Yao. 2023. Reflexion: language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023 NeurlPS 2023 New Orleans LA USA December 10 - 16 2023. http:\/\/papers.nips.cc\/paper_files\/paper\/2023\/hash\/lb44b878bb782e6954cd888628510e90-Abstract-Conference.html"},{"key":"e_1_3_1_75_2","unstructured":"Alexander G Shypula Aman Madaan Yimeng Zeng Uri Alon Jacob R. Gardner Yiming Yang Milad Hashemi Graham Neubig Parthasarathy Ranganathan Osbert Bastani and Amir Yazdanbakhsh. 2024. Learning Performance-Improving Code Edits. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=ix7rLVHXyY"},{"key":"e_1_3_1_76_2","unstructured":"Matt Stuchlik Bruno P. Kinoshita and Donald Lee. 2024. The Cirron library. https:\/\/github.com\/s7nfo\/Cirron https:\/\/github.com\/s7nfo\/Cirron."},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","unstructured":"Hongjin Su Shuyang Jiang Yuhang Lai Haoyuan Wu Boao Shi Che Liu Qian Liu and Tao Yu. 2024. ARKS: Active Retrieval in Knowledge Soup for Code Generation. CoRR abs\/2402.12317 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2402.12317 10.48550\/ARXIV.2402.12317 arXiv:2402.12317","DOI":"10.48550\/ARXIV.2402.12317"},{"key":"e_1_3_1_78_2","doi-asserted-by":"publisher","DOI":"10.1007\/sl0664-022-10247-x"},{"key":"e_1_3_1_79_2","doi-asserted-by":"publisher","unstructured":"Wenhan Wang Chenyuan Yang Zhijie Wang Yuheng Huang Zhaoyang Chu Da Song Lingming Zhang An Ran Chen and Lei Ma. 2024. TESTEVAL: Benchmarking Large Language Models for Test Case Generation. CoRR abs\/2406.04531 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2406.04531 10.48550\/ARXIV.2406.04531 arXiv:2406.04531","DOI":"10.48550\/ARXIV.2406.04531"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","unstructured":"Xingyao Wang Yangyi Chen Lifan Yuan Yizhe Zhang Yunzhu Li Hao Peng and Heng Ji. 2024. Executable Code Actions Elicit Better LLM Agents. CoRR abs\/2402.01030 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2402.01030 10.48550\/ARXIV.2402.01030 arXiv:2402.01030","DOI":"10.48550\/ARXIV.2402.01030"},{"key":"e_1_3_1_81_2","doi-asserted-by":"publisher","unstructured":"Yue Wang Hung Le Akhilesh Gotmare Nghi D. Q. Bui Junnan Li and Steven C. H. Hoi. 2023. CodeT5+: Open Code Large Language Models for Code Understanding and Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing EMNLP 2023 Singapore December 6-10 2023. Association for Computational Linguistics 1069\u20131088. https:\/\/doi.org\/10.18653\/Vl\/2023.EMNLP-MAIN.68 10.18653\/Vl\/2023.EMNLP-MAIN.68","DOI":"10.18653\/Vl\/2023.EMNLP-MAIN.68"},{"key":"e_1_3_1_82_2","doi-asserted-by":"publisher","DOI":"10.18653\/Vl\/2021.EMNLP-MAIN.685"},{"key":"e_1_3_1_83_2","unstructured":"Jason Wei Yi Tay Rishi Bommasani Colin Raffel Barret Zoph Sebastian Borgeaud Dani Yogatama Maarten Bosma Denny Zhou Donald Metzler Ed H. Chi Tatsunori Hashimoto Oriol Vinyals Percy Liang Jeff Dean and William Fedus. 2022. Emergent Abilities of Large Language Models. Trans. Mach. Learn. Res. 2022 (2022). https:\/\/openreview.net\/forum?id=yzkSU5zdwD"},{"key":"e_1_3_1_84_2","doi-asserted-by":"publisher","unstructured":"Yuxiang Wei Zhe Wang Jiawei Liu Yifeng Ding and Lingming Zhang. 2023. Magicoder: Source Code Is All You Need. CoRR abs\/2312.02120 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2312.02120 10.48550\/ARXIV.2312.02120 arXiv:2312.02120","DOI":"10.48550\/ARXIV.2312.02120"},{"key":"e_1_3_1_85_2","unstructured":"Papers with Code. 2024. The Leaderboard of APPS benchmark on Papers with Code. https:\/\/paperswithcode.com\/sota\/code-generation-on-apps."},{"key":"e_1_3_1_86_2","unstructured":"Papers with Code. 2024. The Leaderboard of the Code Contests benchmark. https:\/\/paperswithcode.com\/sota\/code-generation-on-codecontests."},{"key":"e_1_3_1_87_2","doi-asserted-by":"publisher","unstructured":"John Yang Carlos E. Jimenez Alexander Wettig Kilian Lieret Shunyu Yao Karthik Narasimhan and Ofir Press. 2024. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. CoRR abs\/2405.15793 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2405.15793 10.48550\/ARXIV.2405.15793 arXiv:2405.15793","DOI":"10.48550\/ARXIV.2405.15793"},{"key":"e_1_3_1_88_2","unstructured":"Xiao Yu Lei Liu Xing Hu Jacky Wai Keung Jin Liu and Xin Xia. 2024. Where Are Large Language Models for Code Generation on GitHub?. arXiv:2406.19544 [cs.SE] https:\/\/arxiv.org\/abs\/2406.19544"},{"key":"e_1_3_1_89_2","doi-asserted-by":"publisher","unstructured":"Dmitrijs Zaparanuks Milan Jovic and Matthias Hauswirth. 2009. Accuracy of performance counter measurements. In 2009 IEEE International Symposium on Performance Analysis of Systems and Software. 23\u201332. https:\/\/doi.org\/10.1109\/ISPASS.2009.4919635 10.1109\/ISPASS.2009.4919635","DOI":"10.1109\/ISPASS.2009.4919635"},{"key":"e_1_3_1_90_2","doi-asserted-by":"publisher","unstructured":"Fengji Zhang Bei Chen Yue Zhang Jacky Keung Jin Liu Daoguang Zan Yi Mao Jian-Guang Lou and Weizhu Chen. 2023. RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing EMNLP 2023 Singapore December 6-10 2023. Association for Computational Linguistics 2471\u20132484. https:\/\/doi.org\/10.18653\/Vl\/2023.EMNLP-MAIN.151 10.18653\/Vl\/2023.EMNLP-MAIN.151","DOI":"10.18653\/Vl\/2023.EMNLP-MAIN.151"},{"key":"e_1_3_1_91_2","doi-asserted-by":"publisher","unstructured":"Fengji Zhang Bei Chen Yue Zhang Jacky Keung Jin Liu Daoguang Zan Yi Mao Jian-Guang Lou and Weizhu Chen. 2023. RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing EMNLP 2023 Singapore December 6-10 2023. Association for Computational Linguistics 2471\u20132484. https:\/\/doi.org\/10.18653\/Vl\/2023.EMNLP-MAIN.151 10.18653\/Vl\/2023.EMNLP-MAIN.151","DOI":"10.18653\/Vl\/2023.EMNLP-MAIN.151"},{"key":"e_1_3_1_92_2","doi-asserted-by":"publisher","unstructured":"Yuntong Zhang Haifeng Ruan Zhiyu Fan and Abhik Roychoudhury. 2024. AutoCodeRover: Autonomous Program Improvement. CoRR abs\/2404.05427 (2024). https:\/\/doi.org\/10.48550\/ARXIV.2404.05427 10.48550\/ARXIV.2404.05427 arXiv:2404.05427","DOI":"10.48550\/ARXIV.2404.05427"},{"key":"e_1_3_1_93_2","doi-asserted-by":"publisher","unstructured":"Qinkai Zheng Xiao Xia Xu Zou Yufei Dong Shan Wang Yufei Xue Zihan Wang Lei Shen Andi Wang Yang Li Teng Su Zhilin Yang and Jie Tang. 2023. CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Evaluations on HumanEval-X. CoRR abs\/2303.17568 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2303.17568 10.48550\/ARXIV.2303.17568 arXiv:2303.17568","DOI":"10.48550\/ARXIV.2303.17568"},{"key":"e_1_3_1_94_2","doi-asserted-by":"publisher","unstructured":"Andy Zhou Kai Yan Michal Shlapentokh-Rothman Haohan Wang and Yu-Xiong Wang. 2023. Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models. CoRR abs\/2310.04406 (2023). https:\/\/doi.org\/10.48550\/ARXIV.2310.04406 10.48550\/ARXIV.2310.04406 arXiv:2310.04406","DOI":"10.48550\/ARXIV.2310.04406"},{"key":"e_1_3_1_95_2","unstructured":"Shuyan Zhou Uri Alon Frank F. Xu Zhengbao Jiang and Graham Neubig. 2023. DocPrompting: Generating Code by Retrieving the Docs. In The Eleventh International Conference on Learning Representations ICLR 2023 Kigali Rwanda May 1-5 2023. OpenReview.net. https:\/\/openreview.net\/forum?id=ZTCxT2t2Ru"},{"key":"e_1_3_1_96_2","unstructured":"Qihao Zhu Daya Guo Zhihong Shao Dejian Yang Peiyi Wang Runxin Xu Y Wu Yukun Li Huazuo Gao Shirong Ma et al.. 2024. DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence. arXiv preprint arXiv:2406.11931 (2024)."}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3715727","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T11:08:11Z","timestamp":1787483291000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715727"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,19]]},"references-count":96,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2025,6,19]]}},"alternative-id":["10.1145\/3715727"],"URL":"https:\/\/doi.org\/10.1145\/3715727","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,19]]}}}