{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,17]],"date-time":"2026-03-17T00:54:27Z","timestamp":1773708867372,"version":"3.50.1"},"reference-count":73,"publisher":"Association for Computing Machinery (ACM)","issue":"ISSTA","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,22]]},"abstract":"<jats:p>The capabilities of Large Language Models (LLMs) in code generation have been extensively studied, particularly for implementing target functionalities from natural-language descriptions. As an alternative to natural language, input-output (I\/O) examples provide an accessible, unambiguous, and flexible way to describe functionalities. However, their inherent diversity, opaqueness, and incompleteness impose greater challenges for understanding and implementing the target requirements. Therefore, generating code from I\/O examples (i.e., example-based code generation) provides a new perspective, allowing us to additionally evaluate LLMs\u2019 capability to infer target functionalities from limited information and to process new-form requirements. However, related research about LLMs in example-based code generation remains largely unexplored. To fill this gap, this paper presents the first comprehensive study on example-based code generation using LLMs. To address the incorrectness caused by the incompleteness of I\/O examples, we adopt an iterative evaluation framework and formalize the objective of example-based code generation as two sequential sub-objectives: generating code conforming to the given examples and generating code that successfully implements the target functionalities from (iteratively) given examples. We assess six state-of-the-art LLMs using a new benchmark of 172 diverse target functionalities (derived from HumanEval and CodeHunt). The results demonstrate that when requirements are described using iterative I\/O examples rather than natural language, the LLMs\u2019 score decreases by over 60%, indicating that example-based code generation remains challenging for the evaluated LLMs. Notably, the vast majority (even over 95%) of successfully implemented functionalities are achieved in the first round of the iterations, suggesting that the LLMs struggle to effectively utilize the iteratively supplemented requirements. Furthermore, we find that combining I\/O examples with even imprecise and fragmental natural language descriptions greatly improves LLM performance, and the selection of initial I\/O examples can also influence the score, suggesting opportunities for prompt optimization. These findings highlight the importance of early prompts during interactions and offer critical insights and implications for enhancing LLM-based code generation.<\/jats:p>","DOI":"10.1145\/3728947","type":"journal-article","created":{"date-parts":[[2025,6,22]],"date-time":"2025-06-22T10:52:56Z","timestamp":1750589576000},"page":"1583-1606","source":"Crossref","is-referenced-by-count":3,"title":["The First Prompt Counts the Most! An Evaluation of Large Language Models on Iterative Example-Based Code Generation"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2574-9774","authenticated-orcid":false,"given":"Yingjie","family":"Fu","sequence":"first","affiliation":[{"name":"School of Computer Science, Peking University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-7519-5733","authenticated-orcid":false,"given":"Bozhou","family":"Li","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5403-3217","authenticated-orcid":false,"given":"Linyi","family":"Li","sequence":"additional","affiliation":[{"name":"Simon Fraser University, Burnaby, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7532-5550","authenticated-orcid":false,"given":"Wentao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6731-216X","authenticated-orcid":false,"given":"Tao","family":"Xie","sequence":"additional","affiliation":[{"name":"Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education; School of Computer Science, Peking University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,6,22]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2108.11590"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2304.06815"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","unstructured":"Jacob Austin Augustus Odena Maxwell Nye Maarten Bosma Henryk Michalewski David Dohan Ellen Jiang Carrie Cai Michael Terry and Quoc Le. 2021. Program Synthesis with Large Language Models. https:\/\/doi.org\/10.48550\/arXiv.2108.07732 10.48550\/arXiv.2108.07732","DOI":"10.48550\/arXiv.2108.07732"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1611.01989"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2983990.2984020"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2015.172"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2005.14165"},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the 2008 USENIX Symposium on Operating Systems Design and Implementation. 8, 209\u2013224","author":"Cadar Cristian","year":"2008","unstructured":"Cristian Cadar, Daniel Dunbar, and Dawson R Engler. 2008. KLEE: Unassisted and Automatic Generation of High-Coverage Tests for Complex Systems Programs. In Proceedings of the 2008 USENIX Symposium on Operating Systems Design and Implementation. 8, 209\u2013224."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-NIER58687.2023.00025"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2208.08227"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2307.03109"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2107.03374"},{"key":"e_1_2_1_13_1","volume-title":"Allen Cypher","year":"2032","unstructured":"1993. Watch What I Do: Programming by Demonstration, Allen Cypher, Daniel C. Halbert, David Kurlander, Henry Lieberman, David Maulsby, Brad A. Myers, and Alan Turransky (Eds.). MIT Press. isbn:0262032139"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.14722\/bar.2020.23009"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","unstructured":"DeepSeek-AI Aixin Liu Bei Feng and Bin Wang. 2024. DeepSeek-V2: A Strong Economical and Efficient Mixture-of-Experts Language Model. https:\/\/doi.org\/10.48550\/arXiv.2405.04434 10.48550\/arXiv.2405.04434","DOI":"10.48550\/arXiv.2405.04434"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","unstructured":"Xueying Du Mingwei Liu Kaixin Wang Hanlin Wang Junwei Liu Yixuan Chen Jiayi Feng Chaofeng Sha Xin Peng and Yiling Lou. 2023. ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation. https:\/\/doi.org\/10.48550\/arXiv.2308.01861 10.48550\/arXiv.2308.01861","DOI":"10.48550\/arXiv.2308.01861"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","unstructured":"Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey and Abhishek Kadian. 2024. The Llama 3 Herd of Models. https:\/\/doi.org\/10.48550\/arXiv.2407.21783 10.48550\/arXiv.2407.21783","DOI":"10.48550\/arXiv.2407.21783"},{"key":"e_1_2_1_18_1","unstructured":"Yingjie Fu Bozhou Li Linyi Li Wentao Zhang and Tao Xie. 2025. The InterCode Project. https:\/\/sites.google.com\/view\/intercodeproj"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2016.2616877"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1064978.1065036"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2308.12950"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1926385.1926423"},{"key":"e_1_2_1_23_1","unstructured":"Sumit Gulwani. 2016. Programming by Examples (and its Applications in Data Wrangling). https:\/\/www.microsoft.com\/en-us\/research\/publication\/programming-examples-applications-data-wrangling\/"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2736282"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","unstructured":"Daya Guo Qihao Zhu Dejian Yang Zhenda Xie Kai Dong Wentao Zhang Guanting Chen Xiao Bi Yu Wu Y. K. Li Fuli Luo Yingfei Xiong and Wenfeng Liang. 2024. DeepSeek-Coder: When the Large Language Model Meets Programming \u2013 The Rise of Code Intelligence. https:\/\/doi.org\/10.48550\/arXiv.2401.14196 10.48550\/arXiv.2401.14196","DOI":"10.48550\/arXiv.2401.14196"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2006.10720"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","unstructured":"Yiyang Hao Ge Li Yongqiang Liu Xiaowei Miao He Zong Siyuan Jiang Yang Liu and He Wei. 2022. AixBench: A Code Generation Benchmark Dataset. https:\/\/doi.org\/10.48550\/arXiv.2206.13179 10.48550\/arXiv.2206.13179","DOI":"10.48550\/arXiv.2206.13179"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","unstructured":"Dan Hendrycks Steven Basart Saurav Kadavath Mantas Mazeika Akul Arora Ethan Guo Collin Burns Samir Puranik Horace He Dawn Song and Jacob Steinhardt. 2021. Measuring Coding Challenge Competence with APPS. https:\/\/doi.org\/10.48550\/arXiv.2105.09938 10.48550\/arXiv.2105.09938","DOI":"10.48550\/arXiv.2105.09938"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2005.314"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","unstructured":"Juyong Jiang Fan Wang Jiasi Shen Sungju Kim and Sunghun Kim. 2024. A Survey on Large Language Models for Code Generation. https:\/\/doi.org\/10.48550\/arXiv.2406.00515 10.48550\/arXiv.2406.00515","DOI":"10.48550\/arXiv.2406.00515"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2302.05020"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2310.06770"},{"key":"e_1_2_1_33_1","unstructured":"Valentin Knappich. 2023. Tests4J benchmark: execution-based evaluation of context-aware language models for test case generation. Master\u2019s thesis."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2211.11501"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2594291.2594333"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2305.03111"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2312.14852"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","unstructured":"Wen-Ding Li and Kevin Ellis. 2024. Is Programming by Example Solved by LLMs? https:\/\/doi.org\/10.48550\/arXiv.2406.08316 10.48550\/arXiv.2406.08316","DOI":"10.48550\/arXiv.2406.08316"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1126\/science.abq1158"},{"key":"e_1_2_1_40_1","volume-title":"Yuyao Wang, and Lingming Zhang.","author":"Liu Jiawei","year":"2024","unstructured":"Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2024. EvalPlus Leaderboard. https:\/\/evalplus.github.io\/leaderboard.html"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2305.01210"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","unstructured":"Tianyang Liu Canwen Xu and Julian McAuley. 2023. RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems. https:\/\/doi.org\/10.48550\/arXiv.2306.03091 10.48550\/arXiv.2306.03091","DOI":"10.48550\/arXiv.2306.03091"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2402.19173"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3360569"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2308.07124"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3385412.3386012"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2302.08468"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","unstructured":"Erik Nijkamp Bo Pang Hiroaki Hayashi Lifu Tu Huan Wang Yingbo Zhou Silvio Savarese and Caiming Xiong. 2023. CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. https:\/\/doi.org\/10.48550\/arXiv.2203.13474 10.48550\/arXiv.2203.13474","DOI":"10.48550\/arXiv.2203.13474"},{"key":"e_1_2_1_49_1","unstructured":"OpenAI. 2024. GPT-4o-mini. https:\/\/openai.com\/index\/gpt-4o-mini-advancing-cost-efficient-intelligence\/ Large Language Model by OpenAI"},{"key":"e_1_2_1_50_1","unstructured":"OpenAI Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad and Ilge Akkaya. 2024. GPT-4 Technical Report. https:\/\/doi.org\/10.48550\/arXiv.2303.08774 10.48550\/arXiv.2303.08774"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2103.02004"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICACS60934.2024.10473291"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.10668"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.4230\/OASIcs.PLATEAU.2018.3"},{"key":"e_1_2_1_55_1","volume-title":"Jiayi Weng, Juan Felipe Ceron Uribe, Liam Fedus, Luke Metz, and Michael Pokorny.","author":"Schulman John","year":"2022","unstructured":"John Schulman, Barret Zoph, Christina Kim, Jacob Menick Jacob Hilton, Jiayi Weng, Juan Felipe Ceron Uribe, Liam Fedus, Luke Metz, and Michael Pokorny. 2022. ChatGPT: Optimizing Language Models for Dialogue.. https:\/\/chatgpt.r4wand.eu.org\/"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/1095430.1081750"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","unstructured":"CodeGemma Team Heri Zhao Jeffrey Hui Joshua Howland Nam Nguyen Siqi Zuo Andrea Hu Christopher A Choquette-Choo Jingyue Shen and Joe Kelley. 2024. CodeGemma: Open Code Models Based on Gemma. https:\/\/doi.org\/10.48550\/arXiv.2406.11409 10.48550\/arXiv.2406.11409","DOI":"10.48550\/arXiv.2406.11409"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2403.08295"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","unstructured":"Gemma Team Morgane Riviere Shreya Pathak and Pier Giuseppe Sessa. 2024. Gemma 2: Improving Open Language Models at a Practical Size. https:\/\/doi.org\/10.48550\/arXiv.2408.00118 10.48550\/arXiv.2408.00118","DOI":"10.48550\/arXiv.2408.00118"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/2593833.2593838"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-79124-9_10"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/2642937.2642941"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava and Shruti Bhosale. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. https:\/\/doi.org\/10.48550\/arXiv.2307.09288 10.48550\/arXiv.2307.09288","DOI":"10.48550\/arXiv.2307.09288"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/3524842.3528009"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1911.09668"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3637364"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","unstructured":"Yeming Wen Pengcheng Yin Kensen Shi Henryk Michalewski Swarat Chaudhuri and Alex Polozov. 2024. Grounding Data Science Code Generation with Input-Output Specifications. https:\/\/doi.org\/10.48550\/arXiv.2402.08073 10.48550\/arXiv.2402.08073","DOI":"10.48550\/arXiv.2402.08073"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2009.5270315"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3591288"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3623322"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2303.12570"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","unstructured":"Zhuosheng Zhang Aston Zhang Mu Li and Alex Smola. 2022. Automatic Chain of Thought Prompting in Large Language Models. https:\/\/doi.org\/10.48550\/arXiv.2210.03493 10.48550\/arXiv.2210.03493","DOI":"10.48550\/arXiv.2210.03493"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2206.08474"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3728947","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,16]],"date-time":"2025-07-16T16:50:02Z","timestamp":1752684602000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728947"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,22]]},"references-count":73,"journal-issue":{"issue":"ISSTA","published-print":{"date-parts":[[2025,6,22]]}},"alternative-id":["10.1145\/3728947"],"URL":"https:\/\/doi.org\/10.1145\/3728947","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,22]]}}}