{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,24]],"date-time":"2026-08-24T17:24:32Z","timestamp":1787592272973,"version":"build-2736575974"},"reference-count":43,"publisher":"Association for Computing Machinery (ACM)","issue":"OOPSLA1","license":[{"start":{"date-parts":[[2025,4,9]],"date-time":"2025-04-09T00:00:00Z","timestamp":1744156800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Program. Lang."],"published-print":{"date-parts":[[2025,4,9]]},"abstract":"<jats:p>Software updates, including bug repair and feature additions, are frequent in modern applications but they often leave test suites outdated, resulting in undetected bugs and increased chances of system failures. A recent study by Meta revealed that 14%-22% of software failures stem from outdated tests that fail to reflect changes in the codebase. This highlights the need to keep tests in sync with code changes to ensure software reliability.<\/jats:p>\n                  <jats:p>\n                    In this paper, we present\n                    <jats:sc>UTFix<\/jats:sc>\n                    , a novel approach for repairing unit tests when their corresponding focal methods undergo changes.\n                    <jats:sc>UTFix<\/jats:sc>\n                    addresses two critical issues: assertion failure and reduced code coverage caused by changes in the focal method. Our approach leverages language models to repair unit tests by providing contextual information such as static code slices, dynamic code slices, and failure messages. We evaluate\n                    <jats:sc>UTFix<\/jats:sc>\n                    on our generated synthetic benchmark (Syn-Bench), and real-world benchmark. In our experiment,\n                    <jats:sc>UTFix<\/jats:sc>\n                    successfully repaired 89.2% of assertion failures and achieved 100% code coverage for 96 tests out of 369 unit tests. On the real-world benchmarks,\n                    <jats:sc>UTFix<\/jats:sc>\n                    repaired 60% of assertion failures while achieving 100% code coverage for 19 out of 30 unit tests. To the best of our knowledge, this is the first comprehensive study focused on unit test in evolving Python projects. Our contributions include the development of\n                    <jats:sc>UTFix<\/jats:sc>\n                    , the creation of Syn-Bench and real-world benchmarks, and the demonstration of the effectiveness of LLM-based methods in addressing unit test failures due to software evolution.\n                  <\/jats:p>","DOI":"10.1145\/3720419","type":"journal-article","created":{"date-parts":[[2025,4,9]],"date-time":"2025-04-09T13:48:26Z","timestamp":1744206506000},"page":"143-168","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["UTFix: Change Aware Unit Test Repairing using LLM"],"prefix":"10.1145","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0599-8215","authenticated-orcid":false,"given":"Shanto","family":"Rahman","sequence":"first","affiliation":[{"name":"University of Texas at Austin, Austin, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5739-013X","authenticated-orcid":false,"given":"Sachit","family":"Kuhar","sequence":"additional","affiliation":[{"name":"Amazon Web Services, Santa Clara, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4261-090X","authenticated-orcid":false,"given":"Berk","family":"Cirisci","sequence":"additional","affiliation":[{"name":"Amazon Web Services, Berlin, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0575-6320","authenticated-orcid":false,"given":"Pranav","family":"Garg","sequence":"additional","affiliation":[{"name":"Amazon Web Services, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6338-1432","authenticated-orcid":false,"given":"Shiqi","family":"Wang","sequence":"additional","affiliation":[{"name":"Meta, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-3163-0310","authenticated-orcid":false,"given":"Xiaofei","family":"Ma","sequence":"additional","affiliation":[{"name":"Amazon Web Services, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-4566-8767","authenticated-orcid":false,"given":"Anoop","family":"Deoras","sequence":"additional","affiliation":[{"name":"Amazon Web Services, Santa Clara, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3406-5235","authenticated-orcid":false,"given":"Baishakhi","family":"Ray","sequence":"additional","affiliation":[{"name":"Amazon Web Services, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,4,9]]},"reference":[{"key":"e_1_3_1_1_2","unstructured":"2024. Claude-3.5-sonnet. https:\/\/www.anthropic.com\/news\/claude-3-5-sonnet."},{"key":"e_1_3_1_2_2","unstructured":"2024. huggingface. https:\/\/huggingface.co\/."},{"key":"e_1_3_1_3_2","unstructured":"2024. langchain. https:\/\/api.python.langchain.com\/en\/latest\/chat_message_histories\/langchain_community.chat_message_histories.in_memory.ChatMessageHistory.html."},{"key":"e_1_3_1_4_2","unstructured":"2024. Tox. https:\/\/tox.wiki\/en\/latest\/config.html."},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","unstructured":"Mithun Acharya and Brian Robinson. 2011. Practical change impact analysis based on static program slicing for industrial software systems. In International Conference on Software Engineering. 746\u2013755.","DOI":"10.1145\/1985793.1985898"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/93548.93576"},{"key":"e_1_3_1_7_2","doi-asserted-by":"crossref","unstructured":"Nadia Alshahwan Jubin Chheda Anastasia Finogenova Beliz Gokkaya Mark Harman Inna Harper Alexandru Marginean Shubho Sengupta and Eddy Wang. 2024. Automated unit test improvement using large language models at meta. In International Symposium on Foundations of Software Engineering. 185\u2013196.","DOI":"10.1145\/3663529.3663839"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.3390\/app11104673"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/24.963124"},{"key":"e_1_3_1_10_2","doi-asserted-by":"crossref","unstructured":"Brett Daniel Danny Dig Tihomir Gvero Vilas Jagannath Johnston Jiaa Damion Mitchell Jurand Nogiec Shin Hwei Tan and Darko Marinov. 2011. Reassert: a tool for repairing broken unit tests. In International Conference on Software Engineering (Tool Demonstrations Track). 1010\u20131012.","DOI":"10.1145\/1985793.1985978"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","unstructured":"Brett Daniel Vilas Jagannath Danny Dig and Darko Marinov. 2009. ReAssert: Suggesting repairs for broken unit tests. In International Conference on Automated Software Engineering. 433\u2013444.","DOI":"10.1109\/ASE.2009.17"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","unstructured":"Elizabeth Dinella Gabriel Ryan Todd Mytkowicz and Shuvendu K Lahiri. 2022. Toga: A neural method for test oracle generation. In International Conference on Software Engineering. 2130\u20132141.","DOI":"10.1145\/3510003.3510141"},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Gordon Fraser and Andrea Arcuri. 2011. Evosuite: automatic test suite generation for object-oriented software. In International Symposium on Foundations of Software Engineering. 416\u2013419.","DOI":"10.1145\/2025113.2025179"},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","unstructured":"Gordon Fraser and Andrea Arcuri. 2013. Evosuite: On the challenges of test case generation in the real world. In International Conference on Software Testing Verification and Validation. 362\u2013369.","DOI":"10.1109\/ICST.2013.51"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/2685612"},{"key":"e_1_3_1_16_2","doi-asserted-by":"crossref","unstructured":"Malcom Gethers Bogdan Dit Huzefa Kagdi and Denys Poshyvanyk. 2012. Integrated impact analysis for managing software changes. In International Conference on Software Engineering. 430\u2013440.","DOI":"10.1109\/ICSE.2012.6227172"},{"key":"e_1_3_1_17_2","doi-asserted-by":"crossref","unstructured":"Rahul Gopinath Carlos Jensen and Alex Groce. 2014. Code coverage for suite evaluation by developers. In International Conference on Software Engineering. 72\u201382.","DOI":"10.1145\/2568225.2568278"},{"key":"e_1_3_1_18_2","unstructured":"Siqi Gu Chunrong Fang Quanjun Zhang Fangyuan Tian Jianyi Zhou and Zhenyu Chen. 2024. Improving LLM-based Unit test generation via Template-based Repair. arXiv preprint arXiv:2408.03095. (2024)."},{"key":"e_1_3_1_19_2","unstructured":"Sepehr Hashtroudi Jiho Shin Hadi Hemmati and Song Wang. 2023. Automated test case generation using code models and domain adaptation. arXiv preprint arXiv:2308.08033. (2023)."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3690928"},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","unstructured":"Xiangyu Li Marcelo d\u2019Amorim and Alessandro Orso. 2019. Intent-preserving test repair. In International Conference on Software Testing Verification and Validation. 217\u2013227.","DOI":"10.1109\/ICST.2019.00030"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3643674"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","unstructured":"Stephan Lukasczyk and Gordon Fraser. 2022. Pynguin: Automated unit test generation for python. In International Conference on Software Engineering Companion. 168\u2013172.","DOI":"10.1145\/3510454.3516829"},{"key":"e_1_3_1_24_2","doi-asserted-by":"crossref","unstructured":"Mehdi Mirzaaghaei. 2011. Automatic test suite evolution. In International Symposium on Foundations of Software Engineering. 396\u2013399.","DOI":"10.1145\/2025113.2025172"},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","unstructured":"Mehdi Mirzaaghaei Fabrizio Pastore and Mauro Pezz\u00e8. 2012. Supporting test suite evolution through test case adaptation. In International Conference on Software Testing Verification and Validation. 231\u2013240.","DOI":"10.1109\/ICST.2012.103"},{"key":"e_1_3_1_26_2","doi-asserted-by":"crossref","unstructured":"Carlos Pacheco Shuvendu K Lahiri Michael D Ernst and Thomas Ball. 2007. Feedback-directed random test generation. In International Conference on Software Engineering. 75\u201384.","DOI":"10.1109\/ICSE.2007.37"},{"key":"e_1_3_1_27_2","unstructured":"Strategic Planning. 2002. The economic impacts of inadequate infrastructure for software testing. National Institute of Standards and Technology. 1 (2002)."},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","unstructured":"Shanto Rahman Abdelrahman Baz Sasa Misailovic and August Shi. 2024. Quantizing large-language models for predicting flaky tests. In International Conference on Software Testing Verification and Validation. 93\u2013104.","DOI":"10.1109\/ICST60714.2024.00018"},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","unstructured":"Shanto Rahman Bala Naren Chanumolu Suzzana Rafi August Shi and Wing Lam. 2025. Ranking Relevant Tests for Order-Dependent Flaky Tests. In International Conference on Software Engineering.","DOI":"10.1109\/ICSE55347.2025.00178"},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","unstructured":"Shanto Rahman and August Shi. 2024. FlakeSync: Automatically Repairing Async Flaky Tests. In International Conference on Software Engineering. 1\u201312.","DOI":"10.1145\/3597503.3639115"},{"key":"e_1_3_1_31_2","doi-asserted-by":"crossref","unstructured":"Gabriel Ryan Siddhartha Jain Mingyue Shang Shiqi Wang Xiaofei Ma Murali Krishna Ramanathan and Baishakhi Ray. 2024. Code-Aware Prompting: A Study of Coverage-Guided Test Generation in Regression Setting using LLM. In International Symposium on Foundations of Software Engineering. 951\u2013971.","DOI":"10.1145\/3643769"},{"key":"e_1_3_1_32_2","unstructured":"Max Sch\u00e4fer Sarah Nadi Aryaz Eghbali and Frank Tip. 2023. An empirical evaluation of using large language models for automated unit test generation. IEEE Transactions on Software Engineering. (2023)."},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","unstructured":"D\u00e1vid Tengeri \u00c1rp\u00e1d Besz\u00e9des Tam\u00e1s Gergely L\u00e1szl\u00f3 Vid\u00e1cs D\u00e1vid Havas and Tibor Gyim\u00f3thy. 2015. Beyond code coverage\u2014An approach for test suite assessment and improvement. In International Conference on Software Testing Verification and Validation Workshops. 1\u20137.","DOI":"10.1109\/ICSTW.2015.7107476"},{"key":"e_1_3_1_34_2","unstructured":"UTFix. 2025. UTFix. https:\/\/sites.google.com\/view\/utfix."},{"key":"e_1_3_1_35_2","doi-asserted-by":"crossref","unstructured":"Arash Vahabzadeh Amin Milani Fard and Ali Mesbah. 2015. An empirical study of bugs in test code. In International Conference on Software Maintenance and Evolution. 101\u2013110.","DOI":"10.1109\/ICSM.2015.7332456"},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","unstructured":"Zejun Wang Kaibo Liu Ge Li and Zhi Jin. 2024. HITS: High-coverage LLM-based Unit Test Generation via Method Slicing. In International Conference on Automated Software Engineering. 1258\u20131268.","DOI":"10.1145\/3691620.3695501"},{"key":"e_1_3_1_37_2","doi-asserted-by":"crossref","unstructured":"Jason Wei Xuezhi Wang Dale Schuurmans Maarten Bosma Fei Xia Ed Chi Quoc V Le Denny Zhou et al.. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022) 24824\u201324837.","DOI":"10.52202\/068431-1800"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.1984.5010248"},{"key":"e_1_3_1_39_2","unstructured":"Zhuokui Xie Yinghao Chen Chen Zhi Shuiguang Deng and Jianwei Yin. 2023. ChatUniTest: a ChatGPT-based automated unit test generation tool. arXiv preprint arXiv:2305.04764 (2023)."},{"key":"e_1_3_1_40_2","doi-asserted-by":"crossref","unstructured":"Yong Xu Bo Huang Guoqing Wu and Mengting Yuan. 2014. Using genetic algorithms to repair JUnit test cases. In Asia-Pacific Software Engineering Conference. 287\u2013294.","DOI":"10.1109\/APSEC.2014.51"},{"key":"e_1_3_1_41_2","unstructured":"Chen Yang Junjie Chen Bin Lin Jianyi Zhou and Ziqi Wang. 2024. Enhancing LLM-based Test Generation for Hard-to-Cover Branches via Program Analysis. arXiv preprint arXiv:2404.04966 (2024)."},{"key":"e_1_3_1_42_2","article-title":"Automated Test Case Repair Using Language Models","author":"Yaraghi Ahmadreza Saboor","year":"2024","unstructured":"Ahmadreza Saboor Yaraghi, Darren Holden, Nafiseh Kahani, and Lionel Briand. 2024. Automated Test Case Repair Using Language Models. IEEE Transactions on Software Engineering (2024).","journal-title":"IEEE Transactions on Software Engineering"},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","unstructured":"Jialu Zhang Todd Mytkowicz Mike Kaufman Ru\u017eica Piskac and Shuvendu K Lahiri. 2022. Using pre-trained language models to resolve textual and semantic merge conflicts (experience paper). In International Symposium on Software Testing and Analysis. 77\u201388.","DOI":"10.1145\/3533767.3534396"}],"container-title":["Proceedings of the ACM on Programming Languages"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3720419","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3720419","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,24]],"date-time":"2026-08-24T16:28:56Z","timestamp":1787588936000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3720419"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,9]]},"references-count":43,"journal-issue":{"issue":"OOPSLA1","published-print":{"date-parts":[[2025,4,9]]}},"alternative-id":["10.1145\/3720419"],"URL":"https:\/\/doi.org\/10.1145\/3720419","relation":{},"ISSN":["2475-1421"],"issn-type":[{"value":"2475-1421","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,9]]},"assertion":[{"value":"2024-10-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-18","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-09","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}