{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,20]],"date-time":"2025-06-20T04:08:53Z","timestamp":1750392533418,"version":"3.41.0"},"reference-count":27,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","funder":[{"name":"National Science Foundation","award":["CNS-2120386"],"award-info":[{"award-number":["CNS-2120386"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,19]]},"abstract":"<jats:p>Although Large Language Models (LLMs) are highly proficient in understanding source code and descriptive texts, they have limitations in reasoning on dynamic program behaviors, such as execution trace and code coverage prediction, and runtime error prediction, which usually require actual program execution. To advance the ability of LLMs in predicting dynamic behaviors, we leverage the strengths of both approaches, Program Analysis (PA) and LLM, in building PredEx, a predictive executor for Python. Our principle is a blended analysis between PA and LLM to use PA to guide the LLM in predicting execution traces. We break down the task of predictive execution into smaller sub-tasks and leverage the deterministic nature when an execution order can be deterministically decided. When it is not certain, we use predictive backward slicing per variable, i.e., slicing the prior trace to only the parts that affect each variable separately breaks up the valuation prediction into significantly simpler problems. Our empirical evaluation on real-world datasets shows that PredEx achieves 31.5\u201347.1% relatively higher accuracy in predicting full execution traces than the state-of-the-art models. It also produces 8.6\u201353.7% more correct execution trace prefixes than those baselines. In predicting next executed statements, its relative improvement over the baselines is 15.7\u2013102.3%. Finally, we show PredEx\u2019s usefulness in two tasks: static code coverage analysis and static prediction of run-time errors for (in)complete code.<\/jats:p>","DOI":"10.1145\/3729402","type":"journal-article","created":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T15:15:34Z","timestamp":1750346134000},"page":"2987-3008","source":"Crossref","is-referenced-by-count":0,"title":["Blended Analysis for Predictive Execution"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-0143-0677","authenticated-orcid":false,"given":"Yi","family":"Li","sequence":"first","affiliation":[{"name":"University of Texas at Dallas, Dallas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-4474-2984","authenticated-orcid":false,"given":"Hridya","family":"Dhulipala","sequence":"additional","affiliation":[{"name":"University of Texas at Dallas, Dallas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8785-6319","authenticated-orcid":false,"given":"Aashish","family":"Yadavally","sequence":"additional","affiliation":[{"name":"University of Texas at Dallas, Dallas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-8457-8528","authenticated-orcid":false,"given":"Xiaokai","family":"Rong","sequence":"additional","affiliation":[{"name":"University of Texas at Dallas, Dallas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5777-7759","authenticated-orcid":false,"given":"Shaohua","family":"Wang","sequence":"additional","affiliation":[{"name":"Central University of Finance and Economics, Bejing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-7962-6090","authenticated-orcid":false,"given":"Tien N.","family":"Nguyen","sequence":"additional","affiliation":[{"name":"University of Texas at Dallas, Dallas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,6,19]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"[n. d.]. python-graphs. https:\/\/github.com\/google-research\/python-graphs"},{"key":"e_1_2_1_2_1","unstructured":"[n. d.]. Python Hunter howpublished =. https:\/\/github.com\/ionelmc\/python-hunter"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/93548.93576"},{"key":"e_1_2_1_4_1","unstructured":"Islem Bouzenia Yangruibo Ding Kexin Pei Baishakhi Ray and Michael Pradel. 2023. TraceFixer: Execution Trace-Driven Program Repair. arxiv:2304.12743."},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the 8th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201908)","author":"Cadar Cristian","year":"2008","unstructured":"Cristian Cadar, Daniel Dunbar, and Dawson Engler. 2008. KLEE: unassisted and automatic generation of high-coverage tests for complex systems programs. In Proceedings of the 8th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201908). USENIX Association, USA. 209\u2013224."},{"key":"e_1_2_1_6_1","unstructured":"[n. d.]. OpenAI. https:\/\/openai.com\/."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3650105.3652292"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3608140"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2017.31"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1065010.1065036"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1"},{"key":"e_1_2_1_12_1","volume-title":"Ismini Lourentzou, and Chris Brown.","author":"Anjum Haque Md Mahim","year":"2023","unstructured":"Md Mahim Anjum Haque, Wasi Uddin Ahmad, Ismini Lourentzou, and Chris Brown. 2023. FixEval: Execution-based Evaluation of Program Fixes for Programming Problems. arxiv:2206.07796. arxiv:2206.07796"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3485832.3488026"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1"},{"key":"e_1_2_1_15_1","unstructured":"[n. d.]. Llama. https:\/\/llama.meta.com\/."},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the 41st International Conference on Machine Learning (ICML\u201924)","author":"Ni Ansong","year":"2024","unstructured":"Ansong Ni, Miltiadis Allamanis, Arman Cohan, Yinlin Deng, Kensen Shi, Charles Sutton, and Pengcheng Yin. 2024. NExT: teaching large language models to reason about code execution. In Proceedings of the 41st International Conference on Machine Learning (ICML\u201924). JMLR.org, Article 1540, 28 pages."},{"key":"e_1_2_1_17_1","unstructured":"[n. d.]. Predictive Execution. https:\/\/github.com\/predictiveexecution\/predictive-execution."},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks","author":"Puri Ruchir","year":"2021","unstructured":"Ruchir Puri, David Kung, Geert Janssen, Wei Zhang, Giacomo Domeniconi, Vladimir Zolotov, Julian T Dolby, Jie Chen, Mihir Choudhury, Lindsey Decker, Veronika Thost, Veronika Thost, Luca Buratti, Saurabh Pujar, Shyam Ramji, Ulrich Finkler, Susan Malaika, and Frederick Reiss. 2021. CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, J. Vanschoren and S. Yeung (Eds.). 1, https:\/\/datasets-benchmarks-proceedings.neurips.cc\/paper_files\/paper\/2021\/file\/a5bfc9e07964f8dddeb95fc584cd965d-Paper-round2.pdf"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2019.2900307"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1081706.1081750"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3611643.3616254"},{"key":"e_1_2_1_22_1","unstructured":"Michele Tufano Shubham Chandel Anisha Agarwal Neel Sundaresan and Colin Clement. 2023. Predicting Code Coverage without Execution. arxiv:2307.13383."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1"},{"key":"e_1_2_1_24_1","first-page":"9781713871088","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS \u201922)","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought prompting elicits reasoning in Large Language Models. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS \u201922). Curran Associates Inc., Red Hook, NY, USA. Article 1800, 14 pages. isbn:9781713871088"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3417943"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00209"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2019.00046"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729402","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T15:22:09Z","timestamp":1750346529000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3729402"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,19]]},"references-count":27,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2025,6,19]]}},"alternative-id":["10.1145\/3729402"],"URL":"https:\/\/doi.org\/10.1145\/3729402","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,19]]}}}