{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T19:20:35Z","timestamp":1778613635834,"version":"3.51.4"},"reference-count":63,"publisher":"Association for Computing Machinery (ACM)","issue":"ISSTA","funder":[{"name":"Research on Risk Assessment and Security Governance of Heterogeneous and Complex Software Supply Chains under the Pioneer Science and Technology Program of Zhejiang Province","award":["2025C01083"],"award-info":[{"award-number":["2025C01083"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2025,6,22]]},"abstract":"<jats:p>Synchronizing production and test code, known as PT co-evolution, is critical for software quality. Given the significant manual effort involved, researchers have tried automating PT co-evolution using predefined heuristics and machine learning models. However, existing solutions are still incomplete. Most approaches only detect and flag obsolete test cases, leaving developers to manually update them. Meanwhile, existing solutions may suffer from low accuracy, especially when applied to real-world software projects.  \nIn this paper, we propose ReAccept, a novel approach leveraging large language models (LLMs), retrievalaugmented generation (RAG), and dynamic validation to fully automate PT co-evolution with high accuracy. ReAccept employs an experience-guided approach to generate prompt templates for the identification and subsequent update processes. After updating a test case, ReAccept performs dynamic validation by checking syntax, verifying semantics, and assessing test coverage. If the validation fails, ReAccept leverages the error messages to iteratively refine the patch. To evaluate ReAccept's effectiveness, we conducted extensive experiments with a dataset of 537 Java projects and compared ReAccept's performance with several stateof-the-art methods. The evaluation results show that ReAccept achieved an update accuracy of 60.16% on the correctly identified obsolete test code, surpassing the state-of-the-art technique CEPROT by 90%. These findings demonstrate that ReAccept can effectively maintain test code, improve overall software quality, and significantly reduce maintenance effort.<\/jats:p>","DOI":"10.1145\/3728930","type":"journal-article","created":{"date-parts":[[2025,6,22]],"date-time":"2025-06-22T10:52:56Z","timestamp":1750589576000},"page":"1234-1256","source":"Crossref","is-referenced-by-count":2,"title":["REACCEPT: Automated Co-evolution of Production and Test Code Based on Dynamic Validation and Large Language Models"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7298-0955","authenticated-orcid":false,"given":"Jianlei","family":"Chi","sequence":"first","affiliation":[{"name":"Xidian University, X'an, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-9938-2028","authenticated-orcid":false,"given":"Xiaotian","family":"Wang","sequence":"additional","affiliation":[{"name":"Harbin Engineering University, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-1310-9103","authenticated-orcid":false,"given":"Yuhan","family":"Huang","sequence":"additional","affiliation":[{"name":"Xidian University, Xi'an, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8306-3028","authenticated-orcid":false,"given":"Lechen","family":"Yu","sequence":"additional","affiliation":[{"name":"Microsoft, Seattle, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2859-9003","authenticated-orcid":false,"given":"Di","family":"Cui","sequence":"additional","affiliation":[{"name":"Xidian University, Xi'an, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-8216-2925","authenticated-orcid":false,"given":"Jianguo","family":"Sun","sequence":"additional","affiliation":[{"name":"Xidian University, Xi'an, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3545-1392","authenticated-orcid":false,"given":"Jun","family":"Sun","sequence":"additional","affiliation":[{"name":"Singapore Management University, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,6,22]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"REACCEPT: REasoning-Action mechanism and Code dynamic validation assisted Co-Evolution of Production and Test code. https:\/\/github.com\/Timiyang-ai\/REACCEPT.git.","author":"Augest","year":"2024","unstructured":"Augest 2, 2024. REACCEPT: REasoning-Action mechanism and Code dynamic validation assisted Co-Evolution of Production and Test code. https:\/\/github.com\/Timiyang-ai\/REACCEPT.git."},{"key":"e_1_2_1_2_1","unstructured":"July 30 2024. Chroma-the open-source embedding database. https:\/\/github.com\/chroma-core\/chroma."},{"key":"e_1_2_1_3_1","unstructured":"July 30 2024. JaCoCo Java Code Coverage Library. https:\/\/www.jacoco.org\/jacoco\/."},{"key":"e_1_2_1_4_1","unstructured":"July 30 2024. javac-Java programming language compiler. https:\/\/docs.oracle.com\/javase\/7\/docs\/technotes\/tools\/ solaris\/javac.html."},{"key":"e_1_2_1_5_1","unstructured":"July 30 2024. JUnit 4. https:\/\/junit.org\/junit4\/."},{"key":"e_1_2_1_6_1","unstructured":"July 30 2024. Langchain Framework. https:\/\/www.langchain.com\/."},{"key":"e_1_2_1_7_1","unstructured":"July 30 2024. Machine Learning Platform-Text Analysis Service | Datumbox. https:\/\/www.datumbox.com\/."},{"key":"e_1_2_1_8_1","unstructured":"July 30 2024. OpenAI Embeddings. https:\/\/platform.openai.com\/docs\/guides\/embeddings."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSTW"},{"key":"e_1_2_1_10_1","unstructured":"Yufan Cai Yun Lin Chenyan Liu Jinglian Wu Yifan Zhang Yiming Liu Yeyun Gong and Jin Song Dong. 2024. On-the-fly adapting code summarization on trainable cost-efective language models. Advances in Neural Information Processing Systems 36 ( 2024 )."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","unstructured":"Saikat Chakraborty Shuvendu K Lahiri Sarah Fakhoury Madanlal Musuvathi Akash Lal Aseem Rastogi Aditya Senthilnathan Rahul Sharma and Nikhil Swamy. 2023. Ranking llm-generated loop invariants for program verification. arXiv preprint arXiv:2310.09342 ( 2023 ). doi:10.48550\/arXiv.2310.09342","DOI":"10.48550\/arXiv.2310.09342"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","unstructured":"Dong Chen Shaoxin Lin Muhan Zeng Daoguang Zan Jian-Gang Wang Anton Cheshkov Jun Sun Hao Yu Guoliang Dong Artem Aliev et al. 2024. CodeR: Issue Resolving with Multi-Agent and Task Graphs. arXiv preprint arXiv:2406.01304 ( 2024 ). doi:10.48550\/arXiv.2406.01304","DOI":"10.48550\/arXiv.2406.01304"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3639219"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-13792-1_3"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2642937.2642982"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3549098"},{"key":"e_1_2_1_18_1","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1109\/TAIC-PART.2006.20","volume-title":"Testing: Academic & Industrial Conference-Practice And Research Techniques (TAIC PART'06)","author":"Grindal Mats","year":"2006","unstructured":"Mats Grindal, Jef Ofutt, and Jonas Mellin. 2006. On the testing maturity of software producing organizations. In Testing: Academic & Industrial Conference-Practice And Research Techniques (TAIC PART'06). IEEE, 171-180."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3611643.3617850"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","unstructured":"Qi Guo Xiaohong Li Xiaofei Xie Shangqing Liu Ze Tang Ruitao Feng Junjie Wang Jidong Ge and Lei Bu. 2024. FT2Ra: A Fine-Tuning-Inspired Approach to Retrieval-Augmented Code Completion. arXiv preprint arXiv:2404.01554 ( 2024 ). doi: 10.1145\/3650212.3652130","DOI":"10.1145\/3650212.3652130"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3624735"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE56229"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","unstructured":"Yuan Huang Zhicao Tang Xiangping Chen and Xiaocong Zhou. 2024. Towards automatically identifying the co-change of production and test code. Software Testing Verification and Reliability 34 3 ( 2024 ) e1870. doi: 10.1002\/stvr.1870","DOI":"10.1002\/stvr.1870"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CSMR"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3639477.3639733"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00107"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3643991.3645074"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/SEAA56994"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","unstructured":"Bonan Kou Shengmai Chen Zhijie Wang Lei Ma and Tianyi Zhang. 2023. Is model attention aligned with human attention? an empirical study on large language models for code generation. arXiv preprint arXiv:2306.01220 ( 2023 ). doi:10.48550\/arXiv.2306.01220","DOI":"10.48550\/arXiv.2306.01220"},{"key":"e_1_2_1_30_1","unstructured":"Patrick Lewis Ethan Perez Aleksandra Piktus Fabio Petroni Vladimir Karpukhin Naman Goyal Heinrich K\u00fcttler Mike Lewis Wen-tau Yih Tim Rockt\u00e4schel et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems 33 ( 2020 ) 9459-9474."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICST"},{"key":"e_1_2_1_32_1","volume-title":"Proceedings of the 2024 IEEE\/ACM 46th International Conference on Software Engineering: Companion Proceedings. 422-423","author":"Lian Xiaoli","year":"2024","unstructured":"Xiaoli Lian, Shuaisong Wang, Jieping Ma, Xin Tan, Fang Liu, Lin Shi, Cuiyun Gao, and Li Zhang. 2024. Imperfect Code Generation: Uncovering Weaknesses in Automatic Code Generation by Large Language Models. In Proceedings of the 2024 IEEE\/ACM 46th International Conference on Software Engineering: Companion Proceedings. 422-423. doi: 10.1145\/ 3639478.3643081"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/SANER60148"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3609437.3609449"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSR"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/SCAM"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3661167.3661273"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3660810"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","unstructured":"Wendk\u00fbuni C Ou\u00e9draogo Kader Kabor\u00e9 Haoye Tian Yewei Song Anil Koyuncu Jacques Klein David Lo and Tegawend\u00e9 F Bissyand\u00e9. 2024. Large-scale Independent and Comprehensive study of the power of LLMs for test case generation. arXiv preprint arXiv:2407.00225 ( 2024 ). doi:10.48550\/arXiv.2407.00225","DOI":"10.48550\/arXiv.2407.00225"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2931037.2931057"},{"key":"e_1_2_1_41_1","volume-title":"Software testing. Dependable Embedded Systems 5","author":"Pan Jiantao","year":"2006","unstructured":"Jiantao Pan. 1999. Software testing. Dependable Embedded Systems 5, 2006 ( 1999 ), 1."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3608132"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/2393596.2393634"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","unstructured":"Sanyogita Piya and Allison Sullivan. 2023. LLM4TDD: Best Practices for Test Driven Development Using Large Language Models. arXiv preprint arXiv:2312.04687 ( 2023 ). doi:10.48550\/arXiv.2312.04687","DOI":"10.48550\/arXiv.2312.04687"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/MS"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3524481.3527222"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","unstructured":"Jiho Shin Sepehr Hashtroudi Hadi Hemmati and Song Wang. 2024. Domain Adaptation for Code Model-based Unit Test Case Generation. doi:10.48550\/arXiv.2308.08033","DOI":"10.48550\/arXiv.2308.08033"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSM"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","unstructured":"Jeongju Sohn and Mike Papadakis. 2022. Using Evolutionary Coupling to Establish Relevance Links Between Tests and Code Units. A case study on fault localization. arXiv preprint arXiv:2203.11343 ( 2022 ). doi:10.48550\/arXiv.2203.11343","DOI":"10.48550\/arXiv.2203.11343"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3607183"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/CSMR"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/SANER50967"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2406.04531"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377811.3380921"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2210.03629"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","unstructured":"Ahmadreza Saboor Yaraghi Darren Holden Nafiseh Kahani and Lionel Briand. 2024. Automated Test Case Repair Using Language Models. arXiv preprint arXiv:2401.06765 ( 2024 ). doi:10.48550\/arXiv.2401.06765","DOI":"10.48550\/arXiv.2401.06765"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3660783"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICST"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10664-010-9143-7"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i1"},{"key":"e_1_2_1_63_1","volume-title":"2023 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW). IEEE, 62-65","author":"Zimmermann Daniel","year":"2023","unstructured":"Daniel Zimmermann and Anne Koziolek. 2023. Automating gui-based software testing with gpt-3. In 2023 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW). IEEE, 62-65. doi: 10.1109\/ ICSTW58534. 2023.00022"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3728930","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,16]],"date-time":"2025-07-16T16:50:55Z","timestamp":1752684655000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3728930"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,22]]},"references-count":63,"journal-issue":{"issue":"ISSTA","published-print":{"date-parts":[[2025,6,22]]}},"alternative-id":["10.1145\/3728930"],"URL":"https:\/\/doi.org\/10.1145\/3728930","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,22]]}}}