{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,20]],"date-time":"2025-06-20T04:08:53Z","timestamp":1750392533004,"version":"3.41.0"},"reference-count":33,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","funder":[{"name":"National Science Foundation","award":["CNS-2120386"],"award-info":[{"award-number":["CNS-2120386"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,19]]},"abstract":"<jats:p>While LLMs excel in understanding source code and descriptive texts for tasks like code generation, code completion, etc., they exhibit weaknesses in predicting dynamic program behavior, such as code coverage and runtime error detection, which typically require program execution. Aiming to advance the capability of LLMs in reasoning and predicting the program behavior at runtime, we present CRISPE (short for Coverage Rationalization and Intelligent Selection ProcedurE), a novel approach for code coverage prediction. CRISPE guides an LLM in simulating program execution via an execution plan based on two key factors: (1) program semantics of each statement type, and (2) the observation of the set of covered statements at the current \u201cexecution\u201d step relative to all feasible code coverage options. We formulate code coverage prediction as a process of semantic-guided execution-based planning, where feasible coverage options are utilized to assess whether the LLM is heading in the correct reasoning. We enhance the traditional generative task with the retrieval-based framework on feasible options of code coverage. Our experimental results show that CRISPE achieves high accuracy in coverage prediction in terms of both exact-match and statement-match coverage metrics, improving over the baselines. We also show that with semantic-guiding and dynamic reasoning from CRISPE, the LLM generates more correct planning steps. To demonstrate CRISPE\u2019s usefulness, we used it in the downstream task of statically detecting runtime error(s) in incomplete code snippets with the given inputs.<\/jats:p>","DOI":"10.1145\/3729401","type":"journal-article","created":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T15:15:34Z","timestamp":1750346134000},"page":"2965-2986","source":"Crossref","is-referenced-by-count":0,"title":["CRISPE: Semantic-Guided Execution Planning and Dynamic Reasoning for Enhancing Code Coverage Prediction"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-4474-2984","authenticated-orcid":false,"given":"Hridya","family":"Dhulipala","sequence":"first","affiliation":[{"name":"University of Texas at Dallas, Dallas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8785-6319","authenticated-orcid":false,"given":"Aashish","family":"Yadavally","sequence":"additional","affiliation":[{"name":"University of Texas at Dallas, Dallas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-3404-3430","authenticated-orcid":false,"given":"Smit Soneshbhai","family":"Patel","sequence":"additional","affiliation":[{"name":"University of Texas at Dallas, Dallas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-7962-6090","authenticated-orcid":false,"given":"Tien N.","family":"Nguyen","sequence":"additional","affiliation":[{"name":"University of Texas at Dallas, Dallas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,6,19]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"[n. d.]. Python Hunter. https:\/\/github.com\/ionelmc\/python-hunter Accessed: 07\/21\/2023"},{"key":"e_1_2_1_2_1","unstructured":"2024. Claude.ai. https:\/\/claude.ai"},{"key":"e_1_2_1_3_1","unstructured":"2024. CRISPE. https:\/\/github.com\/crispe-prompt-engineering\/crispe"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/MS.2020.3016773"},{"key":"e_1_2_1_5_1","unstructured":"[n. d.]. OpenAI. https:\/\/openai.com\/."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3238147.3238214"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1002\/stvr.347"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3650105.3652292"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3608140"},{"key":"e_1_2_1_10_1","volume-title":"14th USENIX Workshop on Offensive Technologies (WOOT 20)","author":"Fioraldi Andrea","year":"2020","unstructured":"Andrea Fioraldi, Dominik Maier, Heiko Ei\u00df feldt, and Marc Heuse. 2020. AFL++ : Combining Incremental Steps of Fuzzing Research. In 14th USENIX Workshop on Offensive Technologies (WOOT 20). USENIX Association. https:\/\/www.usenix.org\/conference\/woot20\/presentation\/fioraldi"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the 11th International Workshop on Search-Based Software Testing (SBST). 65\u201382","author":"Gay Gregory","year":"2017","unstructured":"Gregory Gay. 2017. Generating effective test suites by combining coverage criteria. In Proceedings of the 11th International Workshop on Search-Based Software Testing (SBST). 65\u201382."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2568225.2568278"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1"},{"key":"e_1_2_1_14_1","volume-title":"Ismini Lourentzou, and Chris Brown.","author":"Anjum Haque Md Mahim","year":"2023","unstructured":"Md Mahim Anjum Haque, Wasi Uddin Ahmad, Ismini Lourentzou, and Chris Brown. 2023. FixEval: Execution-based Evaluation of Program Fixes for Programming Problems. arxiv:2206.07796. arxiv:2206.07796"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3338906.3340459"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3468264.3468580"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2019.00069"},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the 41st International Conference on Machine Learning (ICML\u201924)","author":"Ni Ansong","year":"2024","unstructured":"Ansong Ni, Miltiadis Allamanis, Arman Cohan, Yinlin Deng, Kensen Shi, Charles Sutton, and Pengcheng Yin. 2024. NeXT: teaching large language models to reason about code execution. In Proceedings of the 41st International Conference on Machine Learning (ICML\u201924). JMLR.org, Article 1540, 28 pages."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/302405.302637"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598128"},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks","author":"Puri Ruchir","year":"2021","unstructured":"Ruchir Puri, David Kung, Geert Janssen, Wei Zhang, Giacomo Domeniconi, Vladimir Zolotov, Julian T Dolby, Jie Chen, Mihir Choudhury, Lindsey Decker, Veronika Thost, Veronika Thost, Luca Buratti, Saurabh Pujar, Shyam Ramji, Ulrich Finkler, Susan Malaika, and Frederick Reiss. 2021. CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, J. Vanschoren and S. Yeung (Eds.). 1, https:\/\/datasets-benchmarks-proceedings.neurips.cc\/paper_files\/paper\/2021\/file\/a5bfc9e07964f8dddeb95fc584cd965d-Paper-round2.pdf"},{"key":"e_1_2_1_23_1","unstructured":"Baptiste Rozi\u00e8re Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing Ellen Tan Yossi Adi Jingyu Liu Tal Remez J\u00e9r\u00e9my Rapin Artyom Kozhevnikov Ivan Evtimov Joanna Bitton Manish Bhatt Cristian Canton Ferrer Aaron Grattafiori Wenhan Xiong Alexandre D\u00e9fossez Jade Copet Faisal Azhar Hugo Touvron Louis Martin Nicolas Usunier Thomas Scialom and Gabriel Synnaeve. 2023. Code Llama: Open Foundation Models for Code. arxiv:2308.12950."},{"key":"e_1_2_1_24_1","unstructured":"Clang Team. 2023. Source-based Code Coverage. https:\/\/clang.llvm.org\/docs\/SourceBasedCodeCoverage.html Clang 15 Documentation"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2004.12.021"},{"key":"e_1_2_1_26_1","unstructured":"Michele Tufano Shubham Chandel Anisha Agarwal Neel Sundaresan and Colin Clement. 2023. Predicting Code Coverage without Execution. arxiv:2307.13383."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3491101.3519665"},{"key":"e_1_2_1_28_1","first-page":"9781713871088","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS \u201922)","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought prompting elicits reasoning in Large Language Models. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS \u201922). Curran Associates Inc., Red Hook, NY, USA. Article 1800, 14 pages. isbn:9781713871088"},{"key":"e_1_2_1_29_1","volume-title":"Is branch coverage a good measure of testing effectiveness? Springer-Verlag","author":"Wei Yi","unstructured":"Yi Wei, Bertrand Meyer, and Manuel Oriol. 2012. Is branch coverage a good measure of testing effectiveness? Springer-Verlag, Berlin, Heidelberg. 194\u2013212. isbn:9783642252303"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639121"},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the 11th International Conference on Learning Representations (ICLR).","author":"Yao Shunyu","year":"2023","unstructured":"Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. In Proceedings of the 11th International Conference on Learning Representations (ICLR)."},{"key":"e_1_2_1_32_1","unstructured":"Zhiqiang Yuan Junwei Liu Qiancheng Zi Mingwei Liu Xin Peng and Yiling Lou. 2023. Evaluating Instruction-Tuned Large Language Models on Code Comprehension and Generation. arxiv:2308.01240. arxiv:2308.01240"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2019.00046"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729401","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T15:21:55Z","timestamp":1750346515000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3729401"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,19]]},"references-count":33,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2025,6,19]]}},"alternative-id":["10.1145\/3729401"],"URL":"https:\/\/doi.org\/10.1145\/3729401","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,19]]}}}