{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T02:41:08Z","timestamp":1784342468722,"version":"3.55.0"},"reference-count":89,"publisher":"Association for Computing Machinery (ACM)","issue":"FSE","license":[{"start":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T00:00:00Z","timestamp":1720742400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Softw. Eng."],"published-print":{"date-parts":[[2024,7,12]]},"abstract":"<jats:p>\n                    Unit testing plays an essential role in detecting bugs in functionally-discrete program units (\n                    <jats:italic toggle=\"yes\">e.g.<\/jats:italic>\n                    , methods). Manually writing high-quality unit tests is time-consuming and laborious. Although the traditional techniques are able to generate tests with reasonable coverage, they are shown to exhibit low readability and still cannot be directly adopted by developers in practice. Recent work has shown the large potential of large language models (LLMs) in unit test generation. By being pre-trained on a massive developer-written code corpus, the models are capable of generating more human-like and meaningful test code.\n                  <\/jats:p>\n                  <jats:p>\n                    In this work, we perform the first empirical study to evaluate the capability of ChatGPT (\n                    <jats:italic toggle=\"yes\">i.e<\/jats:italic>\n                    ., one of the most representative LLMs with outstanding performance in code generation and comprehension) in unit test generation. In particular, we conduct both a quantitative analysis and a user study to systematically investigate the quality of its generated tests in terms of correctness, sufficiency, readability, and usability. We find that the tests generated by ChatGPT still suffer from correctness issues, including diverse compilation errors and execution failures (mostly caused by incorrect assertions); but the passing tests generated by ChatGPT almost resemble manually-written tests by achieving comparable coverage, readability, and even sometimes developers\u2019 preference. Our findings indicate that generating unit tests with ChatGPT could be very promising if the correctness of its generated tests could be further improved.\n                  <\/jats:p>\n                  <jats:p>\n                    Inspired by our findings above, we further propose\n                    <jats:sc>ChatTester<\/jats:sc>\n                    , a novel ChatGPT-based unit test generation approach, which leverages ChatGPT itself to improve the quality of its generated tests. Chat Tester incorporates an initial test generator and an iterative test refiner. Our evaluation demonstrates the effectiveness of\n                    <jats:sc>ChatTester<\/jats:sc>\n                    by generating\n                    <jats:inline-formula>\n                      <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" display=\"inline\">\n                        <mml:mn>34.3<\/mml:mn>\n                        <mml:mo>%<\/mml:mo>\n                      <\/mml:math>\n                    <\/jats:inline-formula>\n                    more compilable tests and\n                    <jats:inline-formula>\n                      <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" display=\"inline\">\n                        <mml:mn>18.7<\/mml:mn>\n                        <mml:mo>%<\/mml:mo>\n                      <\/mml:math>\n                    <\/jats:inline-formula>\n                    more tests with correct assertions than the default ChatGPT. In addition to ChatGPT, we further investigate the generalization capabilities of\n                    <jats:sc>ChatTester<\/jats:sc>\n                    by applying it to two recent open-source LLMs (\n                    <jats:italic toggle=\"yes\">i.e.<\/jats:italic>\n                    , CodeLlama-Instruct and CodeFuse) and our results show that\n                    <jats:sc>ChatTester<\/jats:sc>\n                    can also improve the quality of tests generated by these LLMs.\n                  <\/jats:p>","DOI":"10.1145\/3660783","type":"journal-article","created":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T10:22:09Z","timestamp":1720779729000},"page":"1703-1726","source":"Crossref","is-referenced-by-count":123,"title":["Evaluating and Improving ChatGPT for Unit Test Generation"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6497-9380","authenticated-orcid":false,"given":"Zhiqiang","family":"Yuan","sequence":"first","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3462-997X","authenticated-orcid":false,"given":"Mingwei","family":"Liu","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7781-4095","authenticated-orcid":false,"given":"Shiji","family":"Ding","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-6414-2174","authenticated-orcid":false,"given":"Kaixin","family":"Wang","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-8804-1828","authenticated-orcid":false,"given":"Yixuan","family":"Chen","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3376-2581","authenticated-orcid":false,"given":"Xin","family":"Peng","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4066-3365","authenticated-orcid":false,"given":"Yiling","family":"Lou","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,7,12]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2019. http:\/\/javaparser.org\/. (2019)."},{"key":"e_1_3_1_3_2","unstructured":"2022. https:\/\/the-decoder.com\/chatgpt-guide-prompt-strategies\/. (2022)."},{"key":"e_1_3_1_4_2","unstructured":"2022. https:\/\/www.jacoco.org\/jacoco\/. (2022)."},{"key":"e_1_3_1_5_2","unstructured":"CodeLlama 34b Instruct. 2023. (2023). https:\/\/huggingface.co\/codellama\/CodeLlama-34b-Instruct-hf"},{"key":"e_1_3_1_6_2","first-page":"2655","article-title":"Unified Pre-training for Program Understanding and Generation","author":"Ahmad Wasi Uddin","year":"2021","unstructured":"Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021. Unified Pre-training for Program Understanding and Generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, Fune 6-11, 2021. Association for Computational Linguistics, 2655\u20132668.","journal-title":"In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, Fune 6-11, 2021"},{"key":"e_1_3_1_7_2","doi-asserted-by":"crossref","unstructured":"Mohammad Moein Almasi Hadi Hemmati Gordon Fraser Andrea Arcuri and Janis Benefelds. 2017. An Industrial Evaluation of Unit Test Generation: Finding Real Faults in a Financial Application. In 39th IEEE\/ACM International Conference on Software Engineering: Software Engineering in Practice Track ICSE-SEIP 2017 Buenos Aires Argentina May 20-28 2017. IEEE Computer Society 263\u2013272.","DOI":"10.1109\/ICSE-SEIP.2017.27"},{"key":"e_1_3_1_8_2","unstructured":"Yuntao Bai Andy Jones Kamal Ndousse Amanda Askell Anna Chen Nova DasSarma Dawn Drain Stanislav Fort Deep Ganguli Tom Henighan et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862 (2022)."},{"issue":"5","key":"e_1_3_1_9_2","doi-asserted-by":"crossref","first-page":"507","DOI":"10.1109\/TSE.2014.2372785","article-title":"The Oracle Problem in Software Testing","volume":"41","author":"Barr Earl T.","year":"2015","unstructured":"Earl T. Barr, Mark Harman, Phil McMinn, Muzammil Shahbaz, and Shin Yoo. 2015. The Oracle Problem in Software Testing: A Survey. IEEE Trans. Software Eng. 41, 5 (2015), 507\u2013525.","journal-title":"A Survey. IEEE Trans. Software Eng"},{"key":"e_1_3_1_10_2","doi-asserted-by":"crossref","unstructured":"Arianna Blasi Alessandra Gorla Michael D. Ernst and Mauro Pezz\u00e8. 2022. Call Me Maybe: Using NLP to Automatically Generate Unit Test Cases Respecting Temporal Constraints. In 37th IEEE\/ACM International Conference on Automated Software Engineering ASE 2022 Rochester MI USA October 10-14 2022. ACM 19:1-19:11.","DOI":"10.1145\/3551349.3556961"},{"key":"e_1_3_1_11_2","first-page":"19:1","article-title":"Call Me Maybe: Using NLP to Automatically Generate Unit Test Cases Respecting Temporal Constraints","author":"Blasi Arianna","year":"2022","unstructured":"Arianna Blasi, Alessandra Gorla, Michael D. Ernst, and Mauro Pezz\u00e8. 2022. Call Me Maybe: Using NLP to Automatically Generate Unit Test Cases Respecting Temporal Constraints. In 37th IEEE\/ACM International Conference on Automated Software Engineering, ASE 2022, Rochester, MI, USA, October 10-14, 2022. ACM, 19:1\u201319:11.","journal-title":"In 37th IEEE\/ACM International Conference on Automated Software Engineering, ASE 2022, Rochester, MI, USA, October 10-14, 2022. ACM"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1145\/3540250.3549162","article-title":"NatGen: generative pre-training by \u201cnaturalizing\u201d source code","author":"Chakraborty Saikat","year":"2022","unstructured":"Saikat Chakraborty, Toufique Ahmed, Yangruibo Ding, Premkumar T. Devanbu, and Baishakhi Ray. 2022. NatGen: generative pre-training by \u201cnaturalizing\u201d source code. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, ESEC\/FSE 2022, Singapore, Singapore, November 14-18, 2022. ACM, 18\u201330.","journal-title":"In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, ESEC\/FSE 2022"},{"key":"e_1_3_1_13_2","unstructured":"ChatTester. 2023. (2023). https:\/\/github.com\/FudanSELab\/ChatTester\/tree\/main"},{"key":"e_1_3_1_14_2","first-page":"321","volume-title":"GPTutor: A ChatGPT-Powered Programming Tool for Code Explanation","author":"Chen Eason","year":"2023","unstructured":"Eason Chen, Ray Huang, Han-Shin Chen, Yuen-Hsien Tseng, and Liang-Yi Li. 2023. GPTutor: A ChatGPT-Powered Programming Tool for Code Explanation (Communications in Computer and Information Science, Vol. 1831). Springer, 321\u2013327."},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","unstructured":"Hugh A Chipman Edward I George and Robert E McCulloch. 2010. BART: Bayesian additive regression trees. (2010).","DOI":"10.1214\/09-AOAS285"},{"key":"e_1_3_1_16_2","unstructured":"CodeFuse-CodeLlama-34B. 2023. (2023). https:\/\/huggingface.co\/codefuse-ai\/CodeFuse-CodeLlama-34B"},{"key":"e_1_3_1_17_2","doi-asserted-by":"crossref","unstructured":"Christoph Csallner Nikolai Tillmann and Yannis Smaragdakis. 2008. DySy: dynamic symbolic execution for invariant inference. In 30th International Conference on Software Engineering (ICSE 2008) Leipzig Germany May 10-18 2008. ACM 281\u2013290.","DOI":"10.1145\/1368088.1368127"},{"key":"e_1_3_1_18_2","first-page":"3079","volume-title":"Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015","author":"Dai Andrew M.","year":"2015","unstructured":"Andrew M. Dai and Quoc V. Le. 2015. Semi-supervised Sequence Learning. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada. 3079\u20133087."},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1145\/2786805.2786838","article-title":"Modeling readability to improve unit tests","author":"Daka Ermira","year":"2015","unstructured":"Ermira Daka, Jos\u00e9 Campos, Gordon Fraser, Jonathan Dorn, and Westley Weimer. 2015. Modeling readability to improve unit tests. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering, ESEC\/FSE 2015, Bergamo, Italy, August 30 - September 4, 2015. ACM, 107\u2013118.","journal-title":"In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering, ESEC\/FSE 2015"},{"key":"e_1_3_1_20_2","first-page":"201","article-title":"A Survey on Unit Testing Practices and Problems","author":"Daka Ermira","year":"2014","unstructured":"Ermira Daka and Gordon Fraser. 2014. A Survey on Unit Testing Practices and Problems. In 25th IEEE International Symposium on Software Reliability Engineering, ISSRE 2014, Naples, Italy, November 3-6, 2014. IEEE Computer Society, 201\u2013211.","journal-title":"In 25th IEEE International Symposium on Software Reliability Engineering, ISSRE 2014, Naples, Italy, November 3-6, 2014"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2022.3227418"},{"key":"e_1_3_1_22_2","first-page":"423","article-title":"Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language Models","author":"Deng Yinlin","year":"2023","unstructured":"Yinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang, and Lingming Zhang. 2023. Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language Models. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis, ISSTA 2023, Seattle, WA, USA, fuly 17-21, 2023. ACM, 423\u2013435.","journal-title":"In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis, ISSTA 2023"},{"key":"e_1_3_1_23_2","first-page":"70:1","article-title":"Large Language Models are Edge-Case Generators: Crafting Unusual Programs for Fuzzing Deep Learning Libraries","author":"Deng Yinlin","year":"2024","unstructured":"Yinlin Deng, Chunqiu Steven Xia, Chenyuan Yang, Shizhuo Dylan Zhang, Shujing Yang, and Lingming Zhang. 2024. Large Language Models are Edge-Case Generators: Crafting Unusual Programs for Fuzzing Deep Learning Libraries. In Proceedings of the 46th IEEE\/ACM International Conference on Software Engineering, ICSE 2024, Lisbon, Portugal, April 14-20, 2024. ACM, 70:1\u201370:13.","journal-title":"In Proceedings of the 46th IEEE\/ACM International Conference on Software Engineering, ICSE 2024, Lisbon, Portugal"},{"key":"e_1_3_1_24_2","first-page":"4171","article-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","volume":"1","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers). Association for Computational Linguistics, 4171\u20134186.","journal-title":"In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019"},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","unstructured":"Elizabeth Dinella Gabriel Ryan Todd Mytkowicz and Shuvendu K. Lahiri. 2022. TOGA: A Neural Method for Test Oracle Generation. In 44th IEEE\/ACM 44th International Conference on Software Engineering ICSE 2022 Pittsburgh PA USA May 25-27 2022. ACM 2130\u20132141.","DOI":"10.1145\/3510003.3510141"},{"key":"e_1_3_1_26_2","unstructured":"Yihong Dong Xue Jiang Zhi Jin and Ge Li. 2023. Self-collaboration Code Generation via ChatGPT. CoRR abs\/2304.07590 (2023). arXiv:2304.07590"},{"key":"e_1_3_1_27_2","unstructured":"Xueying Du Mingwei Liu Juntao Li Hanlin Wang Xin Peng and Yiling Lou. 2023. Resolving Crash Bugs via Large Language Models: An Empirical Study. CoRR abs\/2312.10448 (2023). arXiv:2312.10448"},{"key":"e_1_3_1_28_2","unstructured":"Xueying Du Mingwei Liu Kaixin Wang Hanlin Wang Junwei Liu Yixuan Chen Jiayi Feng Chaofeng Sha Xin Peng and Yiling Lou. 2023. ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation. CoRR abs\/2308.01861 (2023). arXiv:2308.01861"},{"issue":"1","key":"e_1_3_1_29_2","first-page":"35","article-title":"The Daikon system for dynamic detection of likely invariants. Sci. Comput","volume":"69","author":"Ernst Michael D.","year":"2007","unstructured":"Michael D. Ernst, Jeff H. Perkins, Philip J. Guo, Stephen McCamant, Carlos Pacheco, Matthew S. Tschantz, and Chen Xiao. 2007. The Daikon system for dynamic detection of likely invariants. Sci. Comput. Program. 69, 1-3 (2007), 35\u201345.","journal-title":"Program."},{"key":"e_1_3_1_30_2","first-page":"1536","article-title":"CodeBERT: A Pre-Trained Model for Programming and Natural Languages","author":"Feng Zhangyin","year":"2020","unstructured":"Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020 (Findings of ACL, Vol. EMNLP 2020). Association for Computational Linguistics, 1536\u20131547.","journal-title":"In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020"},{"key":"e_1_3_1_31_2","article-title":"A Review of ChatGPT Applications in Education, Marketing, Software Engineering, and Healthcare","author":"Fraiwan Mohammad","year":"2023","unstructured":"Mohammad Fraiwan and Natheer Khasawneh. 2023. A Review of ChatGPT Applications in Education, Marketing, Software Engineering, and Healthcare: Benefits, Drawbacks, and Research Directions. CoRR abs\/2305.00237 (2023). arXiv:2305.00237","journal-title":"Benefits, Drawbacks, and Research Directions. CoRR abs\/2305.00237 (2023)"},{"key":"e_1_3_1_32_2","first-page":"416","article-title":"EvoSuite: automatic test suite generation for object-oriented software","author":"Fraser Gordon","year":"2011","unstructured":"Gordon Fraser and Andrea Arcuri. 2011. EvoSuite: automatic test suite generation for object-oriented software. In SIGSOFT\/FSE\u201911 19th ACM SIGSOFT Symposium on the Foundations of Software Engineering (FSE-19) and ESEC\u201911: 13th European Software Engineering Conference (ESEC-13), Szeged, Hungary, September 5-9, 2011. ACM, 416\u2013419.","journal-title":"In SIGSOFT\/FSE\u201911 19th ACM SIGSOFT Symposium on the Foundations of Software Engineering (FSE-19) and ESEC\u201911: 13th European Software Engineering Conference (ESEC-13), Szeged"},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","unstructured":"Shuzheng Gao Xin-Cheng Wen Cuiyun Gao Wenxuan Wang Hongyu Zhang and Michael R. Lyu. 2023. What Makes Good In-Context Demonstrations for Code Intelligence Tasks with LLMs?. In 38th IEEE ACM International Conference on Automated Software Engineering ASE 2023 Luxembourg September 11-15 2023. IEEE 761\u2013773.","DOI":"10.1109\/ASE56229.2023.00109"},{"key":"e_1_3_1_34_2","unstructured":"Shuzheng Gao Hongyu Zhang Cuiyun Gao and Chaozheng Wang. 2023. Keeping Pace with Ever-Increasing Data: Towards Continual Learning of Code Intelligence Models. CoRR abs\/2302.03482 (2023). arXiv:2302.03482"},{"key":"e_1_3_1_35_2","unstructured":"getEnvironment(). 2016. (2016). https:\/\/github.com\/trautonen\/coveralls-maven-plugin\/blob\/master\/src\/main\/java\/ org\/eluder\/coveralls\/maven\/plugin\/service\/Travis.java#L75"},{"key":"e_1_3_1_36_2","first-page":"312","article-title":"Scented since the beginning: On the diffuseness of test smells in automatically generated test code. J. Syst","volume":"156","author":"Grano Giovanni","year":"2019","unstructured":"Giovanni Grano, Fabio Palomba, Dario Di Nucci, Andrea De Lucia, and Harald C. Gall. 2019. Scented since the beginning: On the diffuseness of test smells in automatically generated test code. J. Syst. Softw. 156 (2019), 312\u2013327.","journal-title":"Softw."},{"key":"e_1_3_1_37_2","first-page":"348","volume-title":"26th Conference on Program Comprehension, ICPC 2018, Gothenburg, Sweden, May 27-28, 2018","author":"Grano Giovanni","year":"2018","unstructured":"Giovanni Grano, Simone Scalabrino, Harald C. Gall, and Rocco Oliveto. 2018. An empirical investigation on the readability of manual and generated test cases. In Proceedings of the 26th Conference on Program Comprehension, ICPC 2018, Gothenburg, Sweden, May 27-28, 2018. ACM, 348\u2013351."},{"key":"e_1_3_1_38_2","first-page":"226","article-title":"A Theoretical and Empirical Study of Search-Based Testing","author":"Harman Mark","year":"2010","unstructured":"Mark Harman and Phil McMinn. 2010. A Theoretical and Empirical Study of Search-Based Testing: Local, Global, and Hybrid Search. IEEE Trans. Software Eng. 36, 2 (2010), 226\u2013247.","journal-title":"Local, Global, and Hybrid Search. IEEE Trans. Software Eng. 36, 2 (2010)"},{"key":"e_1_3_1_39_2","unstructured":"HumanEval. 2021. (2021). https:\/\/github.com\/openai\/human-eval"},{"key":"e_1_3_1_40_2","article-title":"CodeSearchNet challenge: Evaluating the state of semantic code search. arXiv","author":"Husain Hamel","year":"2019","unstructured":"Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019. CodeSearchNet challenge: Evaluating the state of semantic code search. arXiv preprint arXiv:1909.09436 (2019).","journal-title":"preprint arXiv:1909.09436 (2019)"},{"key":"e_1_3_1_41_2","doi-asserted-by":"crossref","unstructured":"Sajed Jalil Suzzana Rafi Thomas D. LaToza Kevin Moran and Wing Lam. 2023. ChatGPT and Software Testing Education: Promises & Perils. CoRR abs\/2302.03287 (2023). arXiv:2302.03287","DOI":"10.1109\/ICSTW58534.2023.00078"},{"key":"e_1_3_1_42_2","unstructured":"jInstagram. 2015. (2015). https:\/\/github.com\/sachin-handiekar\/jInstagram"},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","first-page":"5131","DOI":"10.1609\/aaai.v37i4.25642","article-title":"Repair Is Nearly Generation: Multilingual Program Repair with LLMs","author":"Joshi Harshit","year":"2023","unstructured":"Harshit Joshi, Jos\u00e9 Pablo Cambronero S\u00e1nchez, Sumit Gulwani, Vu Le, Gust Verbruggen, and Ivan Radicek. 2023. Repair Is Nearly Generation: Multilingual Program Repair with LLMs. AAAI Press, 5131\u20135140.","journal-title":"AAAI Press"},{"key":"e_1_3_1_44_2","doi-asserted-by":"crossref","unstructured":"Sungmin Kang Juyeon Yoon and Shin Yoo. 2023. Large Language Models are Few-shot Testers: Exploring LLM-based General Bug Reproduction. In 45th IEEE\/ACM International Conference on Software Engineering ICSE 2023 Melbourne Australia May 14-20 2023. IEEE 2312\u20132323.","DOI":"10.1109\/ICSE48619.2023.00194"},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","unstructured":"Claus Klammer and Albin Kern. 2015. Writing unit tests: It\u2019s now or never!. In Eighth IEEE International Conference on Software Testing Verification and Validation ICST 2015 Workshops Graz Austria April 13-17 2015. IEEE Computer Society 1\u20134.","DOI":"10.1109\/ICSTW.2015.7107469"},{"key":"e_1_3_1_46_2","doi-asserted-by":"crossref","first-page":"111629","DOI":"10.1016\/j.jss.2023.111629","article-title":"Automatically generating test cases for safety-critical software via symbolic execution","volume":"199","author":"Kurian Elson","year":"2023","unstructured":"Elson Kurian, Daniela Briola, Pietro Braione, and Giovanni Denaro. 2023. Automatically generating test cases for safety-critical software via symbolic execution. J. Syst. Softw. 199 (2023), 111629.","journal-title":"J. Syst. Softw."},{"key":"e_1_3_1_47_2","doi-asserted-by":"crossref","unstructured":"Caroline Lemieux Jeevana Priya Inala Shuvendu K. Lahiri and Siddhartha Sen. 2023. CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language Models. In 45th IEEE\/ACM International Conference on Software Engineering ICSE 2023 Melbourne Australia May 14-20 2023. IEEE 919\u2013931.","DOI":"10.1109\/ICSE48619.2023.00085"},{"key":"e_1_3_1_48_2","unstructured":"Bo Li Gexiang Fang Yang Yang Quansen Wang Wei Ye Wen Zhao and Shikun Zhang. 2023. Evaluating ChatGPT\u2019s Information Extraction Capabilities: An Assessment of Performance Explainability Calibration and Faithfulness. CoRR abs\/2304.11633 (2023). arXiv:2304.11633"},{"key":"e_1_3_1_49_2","article-title":"Nuances are the Key: Unlocking ChatGPT to Find Failure-Inducing Tests with Differential Prompting","author":"Li Tsz On","year":"2023","unstructured":"Tsz On Li, Wenxi Zong, Yibo Wang, Haoye Tian, Ying Wang, Shing-Chi Cheung, and Jeff Kramer. 2023. Nuances are the Key: Unlocking ChatGPT to Find Failure-Inducing Tests with Differential Prompting. 38th IEEE ACM International Conference on Automated Software Engineering (ASE 2023), 11-15 September 2023, Kirchberg, Luxembourg (2023).","journal-title":"38th IEEE ACM International Conference on Automated Software Engineering"},{"key":"e_1_3_1_50_2","doi-asserted-by":"crossref","unstructured":"Stephan Lukasczyk and Gordon Fraser. 2022. Pynguin: Automated Unit Test Generation for Python. In 44th IEEE\/ACM International Conference on Software Engineering: Companion Proceedings ICSE Companion 2022 Pittsburgh PA USA May 22-24 2022. ACM\/IEEE 168\u2013172.","DOI":"10.1109\/ICSE-Companion55297.2022.9793730"},{"key":"e_1_3_1_51_2","doi-asserted-by":"crossref","unstructured":"Stephan Lukasczyk Florian Kroi\u00df and Gordon Fraser. 2023. An empirical study of automated unit test generation for Python. Empir. Softw. Eng. 28 2 (2023) 36.","DOI":"10.1007\/s10664-022-10248-w"},{"key":"e_1_3_1_52_2","doi-asserted-by":"crossref","unstructured":"Antonio Mastropaolo Simone Scalabrino Nathan Cooper David Nader-Palacio Denys Poshyvanyk Rocco Oliveto and Gabriele Bavota. 2021. Studying the Usage of Text-To-Text Transfer Transformer to Support Code-Related Tasks. In 43rd IEEE\/ACM International Conference on Software Engineering ICSE 2021 Madrid Spain 22-30 May 2021. IEEE 336\u2013347.","DOI":"10.1109\/ICSE43902.2021.00041"},{"key":"e_1_3_1_53_2","doi-asserted-by":"crossref","first-page":"223","DOI":"10.1109\/TSE.1976.233818","article-title":"Automatic Generation of Floating-Point Test Data","author":"Miller Webb","year":"1976","unstructured":"Webb Miller and David L. Spooner. 1976. Automatic Generation of Floating-Point Test Data. IEEE Trans. Software Eng. 2, 3 (1976), 223\u2013226.","journal-title":"IEEE Trans. Software Eng. 2, 3 (1976)"},{"key":"e_1_3_1_54_2","doi-asserted-by":"crossref","unstructured":"Noor Nashid Mifta Sintaha and Ali Mesbah. 2023. Retrieval-Based Prompt Selection for Code-Related Few-Shot Learning. In 45th IEEE\/ACM International Conference on Software Engineering ICSE 2023 Melbourne Australia May 14-20 2023. IEEE 2450\u20132462.","DOI":"10.1109\/ICSE48619.2023.00205"},{"key":"e_1_3_1_55_2","doi-asserted-by":"crossref","unstructured":"Pengyu Nie Rahul Banerjee Junyi Jessy Li Raymond J. Mooney and Milos Gligoric. 2023. Learning Deep Semantics for Test Completion. In 45th IEEE\/ACM International Conference on Software Engineering ICSE 2023 Melbourne Australia May 14-20 2023. IEEE 2111\u20132123.","DOI":"10.1109\/ICSE48619.2023.00178"},{"key":"e_1_3_1_56_2","unstructured":"Erik Nijkamp Bo Pang Hiroaki Hayashi Lifu Tu Huan Wang Yingbo Zhou Silvio Savarese and Caiming Xiong. 2023. CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. In The Eleventh International Conference on Learning Representations ICLR 2023 Kigali Rwanda May 1-5 2023. OpenReview.net."},{"key":"e_1_3_1_57_2","first-page":"319","article-title":"Unit testing: test early, test often","author":"Olan Michael","year":"2003","unstructured":"Michael Olan . 2003. Unit testing: test early, test often. Journal of Computing Sciences in Colleges 19, 2 (2003), 319\u2013328.","journal-title":"Journal of Computing Sciences in Colleges 19, 2 (2003)"},{"key":"e_1_3_1_58_2","unstructured":"OpenAI. 2023. ChatGPT: Optimizing Language Models for Dialogue. https:\/\/openai.com\/blog\/chatgpt\/ (2023)."},{"key":"e_1_3_1_59_2","volume-title":"In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022","author":"Ouyang Long","year":"2022","unstructured":"Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, and etc. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022."},{"key":"e_1_3_1_60_2","doi-asserted-by":"crossref","unstructured":"Carlos Pacheco Shuvendu K. Lahiri Michael D. Ernst and Thomas Ball. 2007. Feedback-Directed Random Test Generation. In 29th International Conference on Software Engineering (ICSE 2007) Minneapolis MN USA May 20-26 2007. IEEE Computer Society 75\u201384.","DOI":"10.1109\/ICSE.2007.37"},{"key":"e_1_3_1_61_2","first-page":"5","article-title":"On the diffusion of test smells in automatically generated test code: an empirical study","author":"Palomba Fabio","year":"2016","unstructured":"Fabio Palomba, Dario Di Nucci, Annibale Panichella, Rocco Oliveto, and Andrea De Lucia. 2016. On the diffusion of test smells in automatically generated test code: an empirical study. In Proceedings of the 9th International Workshop on Search-Based Software Testing, SBST@ICSE 2016, Austin, Texas, USA, May 14-22, 2016. ACM, 5\u201314.","journal-title":"In Proceedings of the 9th International Workshop on Search-Based Software Testing, SBST@ICSE 2016"},{"key":"e_1_3_1_62_2","first-page":"130","article-title":"Automatic test case generation: what if test code quality matters?","author":"Palomba Fabio","year":"2016","unstructured":"Fabio Palomba, Annibale Panichella, Andy Zaidman, Rocco Oliveto, and Andrea De Lucia. 2016. Automatic test case generation: what if test code quality matters?. In Proceedings of the 25th International Symposium on Software Testing and Analysis, ISSTA 2016, Saarbr\u00fccken, Germany, July 18-20, 2016. ACM, 130\u2013141.","journal-title":"In Proceedings of the 25th International Symposium on Software Testing and Analysis, ISSTA 2016"},{"key":"e_1_3_1_63_2","unstructured":"Yihao Qin Shangwen Wang Yiling Lou Jinhao Dong Kaixin Wang Xiaoling Li and Xiaoguang Mao. 2024. AgentFL: Scaling LLM-based Fault Localization to Project-Level Context. CoRR abs\/2403.16362 (2024). arXiv:2403.16362"},{"key":"e_1_3_1_64_2","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans Ilya Sutskever et al. 2018. Improving language understanding by generative pre-training. (2018)."},{"key":"e_1_3_1_65_2","unstructured":"Colin Raffel Noam Shazeer Adam Roberts Katherine Lee Sharan Narang Michael Matena Yanqi Zhou Wei Li and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. J. Mach. Learn. Res. 21 (2020) 140:1-140:67."},{"key":"e_1_3_1_66_2","first-page":"383","article-title":"Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017","author":"Ramachandran Prajit","year":"2017","unstructured":"Prajit Ramachandran, Peter J. Liu, and Quoc V. Le. 2017. Unsupervised Pretraining for Sequence to Sequence Learning. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017. Association for Computational Linguistics, 383\u2013391.","journal-title":"Association for Computational Linguistics"},{"key":"e_1_3_1_67_2","doi-asserted-by":"crossref","unstructured":"Xiaoxue Ren Xinyuan Ye Dehai Zhao Zhenchang Xing and Xiaohu Yang. 2023. From Misuse to Mastery: Enhancing Code Generation with Knowledge-Driven AI Chaining. In 38th IEEE\/ACM International Conference on Automated Software Engineering ASE 2023 Luxembourg September 11-15 2023. IEEE 976\u2013987.","DOI":"10.1109\/ASE56229.2023.00143"},{"key":"e_1_3_1_68_2","doi-asserted-by":"crossref","unstructured":"Per Runeson . 2006. A Survey of Unit Testing Practices. IEEE Softw. 23 4 (2006) 22-29.","DOI":"10.1109\/MS.2006.91"},{"key":"e_1_3_1_69_2","doi-asserted-by":"crossref","unstructured":"Simone Scalabrino Giovanni Grano Dario Di Nucci Michele Guerra Andrea De Lucia Harald C. Gall and Rocco Oliveto. 2018. OCELOT: a search-based test-data generation tool for C. In Proceedings of the 33rd ACM\/IEEE International Conference on Automated Software Engineering ASE 2018 Montpellier France September 3-7 2018. ACM 868\u2013871.","DOI":"10.1145\/3238147.3240477"},{"key":"e_1_3_1_70_2","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1109\/TSE.2023.3334955","article-title":"An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation","author":"Sch\u00e4fer Max","year":"2024","unstructured":"Max Sch\u00e4fer, Sarah Nadi, Aryaz Eghbali, and Frank Tip. 2024. An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation. IEEE Trans. Software Eng. 50, 1 (2024), 85\u2013105.","journal-title":"IEEE Trans. Software Eng. 50, 1 (2024)"},{"key":"e_1_3_1_71_2","volume-title":"Elements of survey sampling","author":"Ravindra Singh and Naurang Singh Mangat","year":"2013","unstructured":"Ravindra Singh and Naurang Singh Mangat. 2013. Elements of survey sampling. Vol. 15. Springer Science & Business Media."},{"key":"e_1_3_1_72_2","doi-asserted-by":"crossref","unstructured":"Dominik Sobania Martin Briesch Carol Hanna and Justyna Petke. 2023. An Analysis of the Automatic Bug Fixing Performance of ChatGPT. CoRR abs\/2301.08653 (2023). arXiv:2301.08653","DOI":"10.1109\/APR59189.2023.00012"},{"key":"e_1_3_1_73_2","unstructured":"tabula java. 2017. (2017). https:\/\/github.com\/tabulapdf\/tabula-java"},{"key":"e_1_3_1_74_2","unstructured":"Michele Tufano Dawn Drain Alexey Svyatkovskiy Shao Kun Deng and Neel Sundaresan. 2020. Unit test case generation with transformers and focal context. arXiv preprint arXiv:2009.05617 (2020)."},{"key":"e_1_3_1_75_2","first-page":"5998","volume-title":"Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA. 5998\u20136008."},{"key":"e_1_3_1_76_2","volume-title":"Mathematical statistics with applications","author":"Wackerly Dennis","year":"2014","unstructured":"Dennis Wackerly, William Mendenhall, and Richard L Scheaffer. 2014. Mathematical statistics with applications. Cengage Learning."},{"key":"e_1_3_1_77_2","unstructured":"Chong Wang Jianan Liu Xin Peng Yang Liu and Yiling Lou. 2023. Boosting Static Resource Leak Detection via LLM-based Resource-Oriented Intention Inference. CoRR abs\/2311.04448 (2023). arXiv:2311.04448"},{"key":"e_1_3_1_78_2","unstructured":"Junjie Wang Yuchao Huang Chunyang Chen Zhe Liu Song Wang and Qing Wang. 2023. Software Testing with Large Language Model: Survey Landscape and Vision. CoRR abs\/2307.07221 (2023). arXiv:2307.07221"},{"key":"e_1_3_1_79_2","first-page":"8696","article-title":"CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation","author":"Wang Yue","year":"2021","unstructured":"Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021. CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event \/ Punta Cana, Dominican Republic, 7-11 November, 2021. Association for Computational Linguistics, 8696\u20138708.","journal-title":"In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event \/ Punta Cana, Dominican Republic, 7-11 November, 2021"},{"key":"e_1_3_1_80_2","doi-asserted-by":"crossref","unstructured":"Cody Watson Michele Tufano Kevin Moran Gabriele Bavota and Denys Poshyvanyk. 2020. On learning meaningful assert statements for unit test cases. In ICSE\u201920: 42nd International Conference on Software Engineering Seoul South Korea 27 June - 19 July 2020. ACM 1398\u20131409.","DOI":"10.1145\/3377811.3380429"},{"key":"e_1_3_1_81_2","article-title":"Chain-of-Thought Prompting Elicits Reasoning in Large Language Models","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022.","journal-title":"In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022"},{"key":"e_1_3_1_82_2","unstructured":"Chunqiu Steven Xia Matteo Paltenghi Jia Le Tian Michael Pradel and Lingming Zhang. 2023. Universal Fuzzing via Large Language Models. CoRR abs\/2308.04748 (2023). arXiv:2308.04748"},{"key":"e_1_3_1_83_2","unstructured":"Chunqiu Steven Xia and Lingming Zhang. 2023. Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT. CoRR abs\/2304.00385 (2023). arXiv:2304.00385"},{"key":"e_1_3_1_84_2","doi-asserted-by":"crossref","unstructured":"Xusheng Xiao Sihan Li Tao Xie and Nikolai Tillmann. 2013. Characteristic studies of loop problems for structural test generation via symbolic execution. In 2013 28th IEEE\/ACM International Conference on Automated Software Engineering ASE 2013 Silicon Valley CA USA November 11-15 2013. IEEE 246\u2013256.","DOI":"10.1109\/ASE.2013.6693084"},{"key":"e_1_3_1_85_2","first-page":"5754","volume-title":"In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019","author":"Yang Zhilin","year":"2019","unstructured":"Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019. XLNet: Generalized Autoregressive Pretraining for Language Understanding. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada. 5754\u20135764."},{"key":"e_1_3_1_86_2","article-title":"Tree of Thoughts: Deliberate Problem Solving with Large Language Models","author":"Yao Shunyu","year":"2023","unstructured":"Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10-16, 2023.","journal-title":"In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023"},{"key":"e_1_3_1_87_2","unstructured":"zappos json. 2016. (2016). https:\/\/github.com\/Zappos\/zappos-json"},{"key":"e_1_3_1_88_2","unstructured":"Andreas Zeller Rahul Gopinath Marcel B\u00f6hme Gordon Fraser and Christian Holler. 2019. The fuzzing book."},{"key":"e_1_3_1_89_2","unstructured":"Wayne Xin Zhao Kun Zhou Junyi Li Tianyi Tang Xiaolei Wang and etc. 2023. A Survey of Large Language Models. CoRR abs\/2303.18223 (2023). arXiv:2303.18223"},{"key":"e_1_3_1_90_2","doi-asserted-by":"crossref","unstructured":"Hong Zhu Patrick A. V. Hall and John H. R. May. 1997. Software Unit Test Coverage and Adequacy. ACM Comput. Surv. 29 4 (1997) 366\u2013427.","DOI":"10.1145\/267580.267590"}],"container-title":["Proceedings of the ACM on Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3660783","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3660783","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,4]],"date-time":"2026-02-04T07:56:03Z","timestamp":1770191763000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3660783"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,12]]},"references-count":89,"journal-issue":{"issue":"FSE","published-print":{"date-parts":[[2024,7,12]]}},"alternative-id":["10.1145\/3660783"],"URL":"https:\/\/doi.org\/10.1145\/3660783","relation":{},"ISSN":["2994-970X"],"issn-type":[{"value":"2994-970X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,12]]}}}