{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T02:37:37Z","timestamp":1784342257041,"version":"3.55.0"},"reference-count":43,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T00:00:00Z","timestamp":1720742400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2024,7,12]]},"abstract":"<jats:p>Testing plays a pivotal role in ensuring software quality, yet conventional Search Based Software Testing (SBST) methods often struggle with complex software units, achieving suboptimal test coverage. Recent work using large language models (LLMs) for test generation have focused on improving generation quality through optimizing the test generation context and correcting errors in model outputs, but use fixed prompting strategies that prompt the model to generate tests without additional guidance. As a result LLM-generated testsuites still suffer from low coverage.<\/jats:p>\n                  <jats:p>In this paper, we present SymPrompt, a code-aware prompting strategy for LLMs in test generation. SymPrompt\u2019s approach is based on recent work that demonstrates LLMs can solve more complex logical problems when prompted to reason about the problem in a multi-step fashion. We apply this methodology to test generation by deconstructing the testsuite generation process into a multi-stage sequence, each of which is driven by a specific prompt aligned with the execution paths of the method under test, and exposing relevant type and dependency focal context to the model. Our approach enables pretrained LLMs to generate more complete test cases without any additional training. We implement SymPrompt using the TreeSitter parsing framework and evaluate on a benchmark challenging methods from open source Python projects. SymPrompt enhances correct test generations by a factor of 5 and bolsters relative coverage by 26% for CodeGen2. Notably, when applied to GPT-4, SymPrompt improves coverage by over 2\u00d7 compared to baseline prompting strategies.<\/jats:p>","DOI":"10.1145\/3643769","type":"journal-article","created":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T10:22:09Z","timestamp":1720779729000},"page":"951-971","source":"Crossref","is-referenced-by-count":89,"title":["Code-Aware Prompting: A Study of Coverage-Guided Test Generation in Regression Setting using LLM"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-9464-587X","authenticated-orcid":false,"given":"Gabriel","family":"Ryan","sequence":"first","affiliation":[{"name":"Columbia University, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-3845-6478","authenticated-orcid":false,"given":"Siddhartha","family":"Jain","sequence":"additional","affiliation":[{"name":"AWS AI Labs, Los Angeles, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-6523-3516","authenticated-orcid":false,"given":"Mingyue","family":"Shang","sequence":"additional","affiliation":[{"name":"AWS AI Labs, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6338-1432","authenticated-orcid":false,"given":"Shiqi","family":"Wang","sequence":"additional","affiliation":[{"name":"AWS AI Labs, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-3163-0310","authenticated-orcid":false,"given":"Xiaofei","family":"Ma","sequence":"additional","affiliation":[{"name":"AWS AI Labs, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6568-017X","authenticated-orcid":false,"given":"Murali Krishna","family":"Ramanathan","sequence":"additional","affiliation":[{"name":"AWS AI Labs, San Jose, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3406-5235","authenticated-orcid":false,"given":"Baishakhi","family":"Ray","sequence":"additional","affiliation":[{"name":"AWS AI Labs, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,7,12]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"[n. d.]. Am I in the stack? https:\/\/huggingface.co\/spaces\/bigcode\/in-the-stack."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.449"},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","unstructured":"Saranya Alagarsamy Chakkrit Tantithamthavorn and Aldeida Aleti. 2023. A3Test: Assertion-Augmented Automated Test Case Generation. arXiv preprint arXiv:2302.10352 (2023).","DOI":"10.2139\/ssrn.4724885"},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","unstructured":"Miltiadis Allamanis Earl T Barr Christian Bird and Charles Sutton. 2015. Suggesting accurate method and class names. In Proceedings of the 2015 10th joint meeting on foundations of software engineering. 38-49.","DOI":"10.1145\/2786805.2786849"},{"key":"e_1_3_1_6_2","unstructured":"Jacob Austin Augustus Odena Maxwell Nye Maarten Bosma Henryk Michalewski David Dohan Ellen Jiang Carrie Cai Michael Terry Quoc Le et al. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732 (2021)."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3182657"},{"key":"e_1_3_1_8_2","unstructured":"Patrick Barei\u00df Beatriz Souza Marcelo d'Amorim and Michael Pradel. 2022. Code generation tools (almost) for free? a study of few-shot pre-trained language models on code. arXiv preprint arXiv:2206.01335 (2022)."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/566171.566191"},{"key":"e_1_3_1_10_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE.2014.11"},{"key":"e_1_3_1_12_2","unstructured":"Arghavan Moradi Dakhel Amin Nikanjam Vahid Majdinasab Foutse Khomh and Michel C Desmarais. 2023. Effective Test Generation Using Pre-trained Large Language Models and Mutation Testing. arXiv preprint arXiv:2308.16557 (2023)."},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Elizabeth Dinella Gabriel Ryan Todd Mytkowicz and Shuvendu K Lahiri. 2022. Toga: A neural method for test oracle generation. In Proceedings of the 44th International Conference on Software Engineering. 2130-2141.","DOI":"10.1145\/3510003.3510141"},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","unstructured":"Gordon Fraser and Andrea Arcuri. 2011. Evosuite: automatic test suite generation for object-oriented software. In Proceedings of the 19th ACM SIGSOFT symposium and the 13th European conference on Foundations of software engineering. 416-419.","DOI":"10.1145\/2025113.2025179"},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","first-page":"362","DOI":"10.1109\/ICST.2013.51","volume-title":"2013 IEEE sixth international conference on software testing, verification and validation","author":"Fraser Gordon","year":"2013","unstructured":"Gordon Fraser and Andrea Arcuri. 2013. Evosuite: On the challenges of test case generation in the real world. In 2013 IEEE sixth international conference on software testing, verification and validation. IEEE, 362-369."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/2685612"},{"key":"e_1_3_1_17_2","doi-asserted-by":"crossref","unstructured":"Gordon Fraser and Andrea Arcuri. 2015. 1600 faults in 100 projects: automatically finding faults while achieving high coverage with evosuite. Empirical software engineering 20 (2015) 611-639.","DOI":"10.1007\/s10664-013-9288-2"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","unstructured":"Patrice Godefroid. 2007. Compositional dynamic test generation. In Proceedings of the 34th annual ACM SIGPLAN- SIGACT symposium on Principles ofprogramming languages. 47-54.","DOI":"10.1145\/1190216.1190226"},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","unstructured":"Patrice Godefroid Nils Klarlund and Koushik Sen. 2005. DART: Directed automated random testing. In Proceedings of the 2005 ACM SIGPLAN conference on Programming language design and implementation. 213-223.","DOI":"10.1145\/1065010.1065036"},{"key":"e_1_3_1_20_2","unstructured":"Sepehr Hashtroudi Jiho Shin Hadi Hemmati and Song Wang. 2023. Automated Test Case Generation Using Code Models and Domain Adaptation. arXiv preprint arXiv:2308.08033 (2023)."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/360248.360252"},{"key":"e_1_3_1_22_2","unstructured":"Denis Kocetkov Raymond Li Loubna Ben Allal Jia Li Chenghao Mou Carlos Mu\u00f1oz Ferrandis Yacine Jernite Margaret Mitchell Sean Hughes Thomas Wolf et al. 2022. The stack: 3 tb of permissively licensed source code. arXiv preprint arXiv:2211.15533 (2022)."},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","unstructured":"Caroline Lemieux Jeevana Priya Inala Shuvendu K Lahiri and Siddhartha Sen. 2023. CODAMOSA: Escaping coverage plateaus in test generation with pre-trained large language models. In International conference on software engineering (ICSE).","DOI":"10.1109\/ICSE48619.2023.00085"},{"key":"e_1_3_1_24_2","unstructured":"Tsz-On Li Wenxi Zong Yibo Wang Haoye Tian Ying Wang and Shing-Chi Cheung. 2023. Finding Failure-Inducing Test Cases with ChatGPT. arXiv preprint arXiv:2304.11686 (2023)."},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","unstructured":"Stephan Lukasczyk and Gordon Fraser. 2022. Pynguin: Automated unit test generation for python. In Proceedings of the ACM\/IEEE 44th International Conference on Software Engineering: Companion Proceedings. 168-172.","DOI":"10.1145\/3510454.3516829"},{"key":"e_1_3_1_26_2","doi-asserted-by":"crossref","first-page":"9","DOI":"10.1007\/978-3-030-59762-7_2","volume-title":"Search-Based Software Engineering: 12th International Symposium, SSBSE 2020, Bari, Italy, October 7-8, 2020, Proceedings 12.","author":"Lukasczyk Stephan","year":"2020","unstructured":"Stephan Lukasczyk, Florian Kroi\u00df, and Gordon Fraser. 2020. Automated unit test generation for python. In Search-Based Software Engineering: 12th International Symposium, SSBSE 2020, Bari, Italy, October 7-8, 2020, Proceedings 12. Springer, 9-24."},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","unstructured":"Thomas J. McCabe. 1976. A Complexity Measure. IEEE Transactions on Software Engineering SE-2 (1976) 308-320. https:\/\/api.semanticscholar.org\/CorpusID:9116234","DOI":"10.1109\/TSE.1976.233837"},{"key":"e_1_3_1_28_2","unstructured":"Erik Nijkamp Hiroaki Hayashi Caiming Xiong Silvio Savarese and Yingbo Zhou. 2023. Codegen2: Lessons for training llms on programming and natural languages. arXiv preprint arXiv:2305.02309 (2023)."},{"key":"e_1_3_1_29_2","unstructured":"OpenAI. 2023. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL]"},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","unstructured":"Carlos Pacheco and Michael D Ernst. 2007. Randoop: feedback-directed random testing for Java. In Companion to the 22ndACM SIGPLAN conference on Object-orientedprogramming systems andapplications companion. 815-816.","DOI":"10.1145\/1297846.1297902"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2007.37"},{"key":"e_1_3_1_32_2","unstructured":"Strategic Planning. 2002. The economic impacts of inadequate infrastructure for software testing. National Institute of Standards and Technology 1 (2002)."},{"key":"e_1_3_1_33_2","unstructured":"Max Sch\u00e4fer Sarah Nadi Aryaz Eghbali and Frank Tip. 2023. Adaptive test generation using a large language model. arXiv preprint arXiv:2302.06527 (2023)."},{"key":"e_1_3_1_34_2","unstructured":"Max Sch\u00e4fer Sarah Nadi Aryaz Eghbali and Frank Tip. 2023. An empirical evaluation of using large language models for automated unit test generation. IEEE Transactions on Software Engineering (2023)."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/1095430.1081750"},{"key":"e_1_3_1_36_2","first-page":"134","volume-title":"International conference on tests andproofs","author":"Halleux Nikolai Tillmann and Jonathan De","year":"2008","unstructured":"Nikolai Tillmann and Jonathan De Halleux. 2008. Pex-white box test generation for. net. In International conference on tests andproofs. Springer, 134-153."},{"key":"e_1_3_1_37_2","unstructured":"Michele Tufano Dawn Drain Alexey Svyatkovskiy Shao Kun Deng and Neel Sundaresan. 2020. Unit test case generation with transformers and focal context. arXiv preprint arXiv:2009.05617 (2020)."},{"key":"e_1_3_1_38_2","unstructured":"Jason Wei Yi Tay Rishi Bommasani Colin Raffel Barret Zoph Sebastian Borgeaud Dani Yogatama Maarten Bosma Denny Zhou Donald Metzler et al. 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682 (2022)."},{"key":"e_1_3_1_39_2","unstructured":"Jason Wei Xuezhi Wang Dale Schuurmans Maarten Bosma Fei Xia Ed Chi Quoc V Le Denny Zhou et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (2022) 24824-24837."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3417943"},{"key":"e_1_3_1_41_2","unstructured":"Zhuokui Xie Yinghao Chen Chen Zhi Shuiguang Deng and Jianwei Yin. 2023. ChatUniTest: a ChatGPT-based automated unit test generation tool. arXiv preprint arXiv:2305.04764 (2023)."},{"key":"e_1_3_1_42_2","unstructured":"Shunyu Yao Dian Yu Jeffrey Zhao Izhak Shafran Thomas L Griffiths Yuan Cao and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601 (2023)."},{"issue":"2","key":"e_1_3_1_43_2","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1002\/stv.430","article-title":"Regression testing minimization, selection and prioritization: a survey","volume":"22","author":"Harman Shin Yoo and Mark","year":"2012","unstructured":"Shin Yoo and Mark Harman. 2012. Regression testing minimization, selection and prioritization: a survey. Software testing, verification and reliability 22, 2 (2012), 67-120.","journal-title":"Software testing, verification and reliability"},{"key":"e_1_3_1_44_2","unstructured":"Zhiqiang Yuan Yiling Lou Mingwei Liu Shiji Ding Kaixin Wang Yixuan Chen and Xin Peng. 2023. No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation. arXiv preprint arXiv:2305.04207 (2023)."}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643769","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3643769","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T07:56:58Z","timestamp":1770191818000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643769"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,12]]},"references-count":43,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2024,7,12]]}},"alternative-id":["10.1145\/3643769"],"URL":"https:\/\/doi.org\/10.1145\/3643769","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,12]]}}}