{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T19:31:55Z","timestamp":1783711915568,"version":"3.55.0"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"5","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62322208 and 62232001"],"award-info":[{"award-number":["62322208 and 62232001"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Huawei Fund"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>Automatic test generation plays a critical role in software quality assurance. While the recent advances in Search-Based Software Testing (SBST) and Large Language Models (LLMs) have shown promise in generating useful tests, these techniques still struggle to cover certain branches. Reaching these hard-to-cover branches usually requires constructing complex objects and resolving intricate inter-procedural dependencies in branch conditions, which poses significant challenges for existing techniques. In this work, we propose TELPA, a novel technique aimed at addressing these challenges. Its key insight lies in extracting real usage scenarios of the target method under test to learn how to construct complex objects and extracting methods entailing inter-procedural dependencies with hard-to-cover branches to learn the semantics of branch constraints. To enhance efficiency and effectiveness, TELPA identifies a set of ineffective tests as counter-examples for LLMs and employs a feedback-based process to iteratively refine these counter-examples. Then, TELPA integrates program analysis results and counter-examples into the prompt, guiding LLMs to gain deeper understandings of the semantics of the target method and generate diverse tests that can reach the hard-to-cover branches. Our experimental results on 27 open-source Python projects demonstrate that TELPA significantly outperforms the state-of-the-art SBST and LLM-enhanced techniques, achieving an average improvement of 34.10% and 25.93% in terms of branch coverage.<\/jats:p>","DOI":"10.1145\/3748505","type":"journal-article","created":{"date-parts":[[2025,7,14]],"date-time":"2025-07-14T19:54:11Z","timestamp":1752522851000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Advancing Code Coverage: Incorporating Program Analysis with Large Language Models"],"prefix":"10.1145","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0759-940X","authenticated-orcid":false,"given":"Chen","family":"Yang","sequence":"first","affiliation":[{"name":"College of Intelligence and Computing, Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3056-9962","authenticated-orcid":false,"given":"Junjie","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Intelligence and Computing, Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6307-8460","authenticated-orcid":false,"given":"Bin","family":"Lin","sequence":"additional","affiliation":[{"name":"Hangzhou Dianzi University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-8630-8724","authenticated-orcid":false,"given":"Ziqi","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Intelligence and Computing, Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4867-5416","authenticated-orcid":false,"given":"Jianyi","family":"Zhou","sequence":"additional","affiliation":[{"name":"Huawei Cloud Computing Technologies Co Ltd, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,4,24]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"Codex. 2024. Codex Shutdown. Retrieved from https:\/\/platform.openai.com\/docs\/deprecations"},{"key":"e_1_3_3_3_2","unstructured":"Coverage.py. 2024. Coverage.py. Retrieved from https:\/\/coverage.readthedocs.io\/en\/7.6.0"},{"key":"e_1_3_3_4_2","unstructured":"Deepseek. 2024. Deepseek. Retrieved from https:\/\/github.com\/deepseek-ai\/DeepSeek-Coder"},{"key":"e_1_3_3_5_2","unstructured":"FastChat. 2024. FastChat. Retrieved from https:\/\/github.com\/lm-sys\/FastChat"},{"key":"e_1_3_3_6_2","unstructured":"Hugging Face. 2024. Hugging Face. Retrieved from https:\/\/huggingface.co"},{"key":"e_1_3_3_7_2","unstructured":"Hugging Face. 2024. Hugging Face Big Code Model Leaderboard. Retrieved from https:\/\/huggingface.co\/spaces\/bigcode\/bigcode-models-leaderboard"},{"key":"e_1_3_3_8_2","unstructured":"Hugging Face. 2024. Huggingface LLM Learderboard. Retrieved from https:\/\/huggingface.co\/spaces\/bigcode\/bigcode-models-leaderboard"},{"key":"e_1_3_3_9_2","unstructured":"Jacoco. 2024. Jacoco. Retrieved from https:\/\/www.eclemma.org\/jacoco"},{"key":"e_1_3_3_10_2","unstructured":"JavaParser. 2024. JavaParser. Retrieved from https:\/\/javaparser.org"},{"key":"e_1_3_3_11_2","unstructured":"Junit. 2024. Junit. Retrieved from https:\/\/junit.org\/junit5"},{"key":"e_1_3_3_12_2","unstructured":"Phind-CodeLlama. 2024. Phind-CodeLlama-34B-v2. Retrieved from https:\/\/huggingface.co\/Phind\/Phind-CodeLlama-34B-v2"},{"key":"e_1_3_3_13_2","unstructured":"PyTorch. 2024. PyTorch. Retrieved from http:\/\/pytorch.org"},{"key":"e_1_3_3_14_2","unstructured":"Transformers. 2024. Transformers. Retrieved from https:\/\/github.com\/huggingface\/transformers"},{"key":"e_1_3_3_15_2","unstructured":"Project Homepage. 2025. Project Homepage. Retrieved from https:\/\/zenodo.org\/records\/15410112"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.211"},{"key":"e_1_3_3_17_2","doi-asserted-by":"crossref","unstructured":"Saranya Alagarsamy Chakkrit Tantithamthavorn and Aldeida Aleti. 2023. A3test: Assertion-augmented automated test case generation. arXiv:2302.10352. Retrieved from https:\/\/arxiv.org\/abs\/2302.10352","DOI":"10.2139\/ssrn.4724885"},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3663529.3663839"},{"key":"e_1_3_3_19_2","doi-asserted-by":"crossref","unstructured":"Nadia Alshahwan Mark Harman Inna Harper Alexandru Marginean Shubho Sengupta and Eddy Wang. 2024. Assured LLM-based software engineering. arXiv:2402.04380. Retrieved from https:\/\/arxiv.org\/abs\/2402.04380","DOI":"10.1145\/3643661.3643953"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-66299-2_1"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2018.05.003"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/1985793.1985795"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2015.2490067"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3092703.3092715"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/2970276.2970366"},{"key":"e_1_3_3_26_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman et al. 2021. Evaluating large language models trained on code. arXiv:2107.03374. Retrieved from https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3663529.3663801"},{"key":"e_1_3_3_28_2","unstructured":"Yew Ken Chia Guizhen Chen Luu Anh Tuan Soujanya Poria and Lidong Bing. 2023. Contrastive chain-of-thought prompting. arXiv:2311.09277. Retrieved from https:\/\/arxiv.org\/abs\/2311.09277"},{"key":"e_1_3_3_29_2","first-page":"44","volume-title":"In Proceedings of the 2017 32nd IEEE\/ACM International Conference on Automated Software Engineering (ASE)","author":"Della Toffola Luca","year":"2017","unstructured":"Luca Della Toffola, Cristian-Alexandru Staicu, and Michael Pradel. 2017. Saying \u2018hi!\u2019 is not enough: Mining inputs for effective test generation. In Proceedings of the 2017 32nd IEEE\/ACM International Conference on Automated Software Engineering (ASE). IEEE, 44\u201349."},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510141"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/2025113.2025179"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE.2013.6698889"},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2010.62"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/2610384.2628055"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1952.10483441"},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00085"},{"key":"e_1_3_3_37_2","unstructured":"Jia Li Ge Li Yongmin Li and Zhi Jin. 2023. Structured chain-of-thought prompting for code generation. arXiv:2305.06599. Retrieved from https:\/\/arxiv.org\/abs\/2305.06599"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/2970276.2970364"},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3468264.3468619"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510454.3516829"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/1297846.1297902"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2017.2663435"},{"key":"e_1_3_3_43_2","article-title":"Direct preference optimization: Your language model is secretly a reward model","volume":"36","author":"Rafailov Rafael","year":"2024","unstructured":"Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASE56229.2023.00193"},{"key":"e_1_3_3_45_2","doi-asserted-by":"crossref","unstructured":"Gabriel Ryan Siddhartha Jain Mingyue Shang Shiqi Wang Xiaofei Ma Murali Krishna Ramanathan and Baishakhi Ray. 2024. Code-aware prompting: A study of coverage guided test generation in regression setting using LLM. arXiv:2402.00097. Retrieved from https:\/\/arxiv.org\/abs\/2402.00097","DOI":"10.1145\/3643769"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00146"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3639476.3639764"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2023.3334955"},{"key":"e_1_3_3_49_2","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv:1707.06347. Retrieved from https:\/\/arxiv.org\/abs\/1707.06347"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2024.3382365"},{"key":"e_1_3_3_51_2","first-page":"645","volume-title":"Proceedings of the 2025 IEEE\/ACM 47th International Conference on Software Engineering (ICSE)","author":"Tian Zhao","year":"2025","unstructured":"Zhao Tian, Junjie Chen, and Xiangyu Zhang. 2025. Fixing large language models\u2019 specification misunderstanding for better code generation. In Proceedings of the 2025 IEEE\/ACM 47th International Conference on Software Engineering (ICSE). IEEE Computer Society, 645\u2013645."},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-79124-9_10"},{"key":"e_1_3_3_53_2","unstructured":"Michele Tufano Dawn Drain Alexey Svyatkovskiy Shao Kun Deng and Neel Sundaresan. 2020. Unit test case generation with transformers. arXiv:2009.05617. Retrieved from https:\/\/arxiv.org\/abs\/2009.05617"},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3524481.3527220"},{"key":"e_1_3_3_55_2","doi-asserted-by":"crossref","unstructured":"Junjie Wang Yuchao Huang Chunyang Chen Zhe Liu Song Wang and Qing Wang. 2024. Software testing with large language models: Survey landscape and vision. IEEE Transactions on Software Engineering 50 4 (2024) 911\u2013936.","DOI":"10.1109\/TSE.2024.3368208"},{"key":"e_1_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3691620.3695501"},{"key":"e_1_3_3_57_2","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma,. Brian Ichter, Brian Ichter, Fei Xia, Ed Chi, Quoc V. Le, Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, Vol. 35, 24824\u201324837.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_58_2","first-page":"1","volume-title":"Wiley Encyclopedia of Clinical Trials.","author":"Woolson Robert F.","year":"2007","unstructured":"Robert F. Woolson. 2007. Wilcoxon signed-rank test. In Wiley Encyclopedia of Clinical Trials. R. B. D\u2019Agostino, L. Sullivan, and J. Massaro (Eds.), John Wiley & Sons, Ltd, 1\u20133."},{"key":"e_1_3_3_59_2","unstructured":"Chunqiu Steven Xia Matteo Paltenghi Jia Le Tian Michael Pradel and Lingming Zhang. 2023. Universal fuzzing via large language models. arXiv:2308.04748. Retrieved from https:\/\/arxiv.org\/abs\/2308.04748"},{"key":"e_1_3_3_60_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2024.107645"},{"key":"e_1_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3689794"},{"key":"e_1_3_3_62_2","unstructured":"Chenyuan Yang Yinlin Deng Runyu Lu Jiayi Yao Jiawei Liu Reyhaneh Jabbarvand and Lingming Zhang. 2023. White-box compiler fuzzing empowered by large language models. arXiv:2310.15991. Retrieved from https:\/\/arxiv.org\/abs\/2310.15991"},{"key":"e_1_3_3_63_2","unstructured":"Chen Yang Ziqi Wang Yanjie Jiang Lin Yang Yuteng Zheng Jianyi Zhou and Junjie Chen. 2025. Precisely detecting python type errors via LLM-based unit test generation. arXiv:2507.02318. Retrieved from https:\/\/arxiv.org\/abs\/2507.02318"},{"key":"e_1_3_3_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3691620.3695529"},{"key":"e_1_3_3_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3660783"},{"key":"e_1_3_3_66_2","unstructured":"Zhuosheng Zhang Aston Zhang Mu Li and Alex Smola. 2022. Automatic chain of thought prompting in large language models. arXiv:2210.03493. Retrieved from https:\/\/arxiv.org\/abs\/2210.03493"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3748505","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T06:33:40Z","timestamp":1777098820000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3748505"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,24]]},"references-count":65,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3748505"],"URL":"https:\/\/doi.org\/10.1145\/3748505","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,24]]},"assertion":[{"value":"2024-08-13","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-07","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}