{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T04:47:29Z","timestamp":1784695649829,"version":"3.55.0"},"reference-count":40,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,1,21]],"date-time":"2025-01-21T00:00:00Z","timestamp":1737417600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,1,21]],"date-time":"2025-01-21T00:00:00Z","timestamp":1737417600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100000781","name":"European Research Council","doi-asserted-by":"publisher","award":["741278."],"award-info":[{"award-number":["741278."]}],"id":[{"id":"10.13039\/501100000781","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Autom Softw Eng"],"published-print":{"date-parts":[[2025,5]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Ever since the first large language models (LLMs) have become available, both academics and practitioners have used them to aid software engineering tasks. However, little research as yet has been done in combining search-based software engineering (SBSE) and LLMs. In this paper, we evaluate the use of LLMs as mutation operators for genetic improvement (GI), an SBSE approach, to improve the GI search process. In a preliminary work, we explored the feasibility of combining the\n                    <jats:italic>Gin<\/jats:italic>\n                    Java GI toolkit with OpenAI LLMs in order to generate an edit for the  tool. Here we extend this investigation involving three LLMs and three types of prompt, and five real-world software projects. We sample the edits at random, as well as using local search. We also conducted a qualitative analysis to understand why LLM-generated code edits break as part of our evaluation. Our results show that, compared with conventional statement GI edits, LLMs produce fewer unique edits, but these compile and pass tests more often, with the  model finding test-passing edits 77% of the time. The  and  LLMs are roughly equal in finding the best run-time improvements. Simpler prompts are more successful than those providing more context and examples. The qualitative analysis reveals a wide variety of areas where LLMs typically fail to produce valid edits commonly including inconsistent formatting, generating non-Java syntax, or refusing to provide a solution.\n                  <\/jats:p>","DOI":"10.1007\/s10515-024-00473-6","type":"journal-article","created":{"date-parts":[[2025,1,21]],"date-time":"2025-01-21T13:59:02Z","timestamp":1737467942000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Large language model based mutations in genetic improvement"],"prefix":"10.1007","volume":"32","author":[{"given":"Alexander E. I.","family":"Brownlee","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"James","family":"Callan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Karine","family":"Even-Mendoza","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alina","family":"Geiger","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Carol","family":"Hanna","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Justyna","family":"Petke","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Federica","family":"Sarro","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dominik","family":"Sobania","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,1,21]]},"reference":[{"key":"473_CR1","doi-asserted-by":"crossref","unstructured":"Albuquerque, L., Gheyi, R., Ribeiro, M.: Evaluating the capability of llms in identifying compilation errors in configurable systems. arXiv preprint arXiv:2407.19087 (2024)","DOI":"10.5753\/sbes.2024.3560"},{"issue":"1145\/3638530","key":"473_CR2","first-page":"3654408","volume":"10","author":"G Birna","year":"2024","unstructured":"Birna, G., Saemundur, S., Haraldsson, O.: Large language models as all-in-one operators for genetic improvement 10(1145\/3638530), 3654408 (2024)","journal-title":"Large language models as all-in-one operators for genetic improvement"},{"key":"473_CR3","doi-asserted-by":"publisher","unstructured":"Blot, A., Petke, J.: A Comprehensive Survey of Benchmarks for Automated Improvement of Software\u2019s Non-Functional Properties (2022). https:\/\/doi.org\/10.48550\/ARXIV.2212.08540","DOI":"10.48550\/ARXIV.2212.08540"},{"key":"473_CR4","doi-asserted-by":"publisher","unstructured":"Blyth, S., Treude, C., Wagner, M.: Creative and correct: Requesting diverse code solutions from ai foundation models. In: Proceedings of the 2024 IEEE\/ACM First International Conference on AI Foundation Models and Software Engineering. FORGE \u201924, pp. 119\u2013123. Association for Computing Machinery, New York, NY, USA (2024). https:\/\/doi.org\/10.1145\/3650105.3652302","DOI":"10.1145\/3650105.3652302"},{"key":"473_CR5","doi-asserted-by":"publisher","unstructured":"B\u00f6hme, M., Soremekun, E.O., Chattopadhyay, S., Ugherughe, E., Zeller, A.: Where is the bug and how is it fixed? An experiment with practitioners. In: Proc. ACM Symposium on the Foundations of Software Engineering, pp. 117\u2013128 (2017). https:\/\/doi.org\/10.1145\/3106237.3106255","DOI":"10.1145\/3106237.3106255"},{"key":"473_CR6","unstructured":"Bouzenia, I., Devanbu, P., Pradel, M.: Repairagent: An autonomous, llm-based agent for program repair (2024)"},{"key":"473_CR7","doi-asserted-by":"publisher","unstructured":"Brownlee, A.E.I., Callan, J., Even-Mendoza, K., Geiger, A., Hanna, C., Petke, J., Sarro, F., Sobania, D.: Artifact of Large Language Model Based Mutations in Genetic Improvement. Zenodo, Switzerland (2024). https:\/\/doi.org\/10.5281\/zenodo.13381774","DOI":"10.5281\/zenodo.13381774"},{"key":"473_CR8","doi-asserted-by":"publisher","unstructured":"Brownlee, A.E.I., Callan, J., Even-Mendoza, K., Geiger, A., Hanna, C., Petke, J., Sarro, F., Sobania, D.: Enhancing genetic improvement mutations using large language models. In: Search-Based Software Engineering, pp. 153\u2013159. Springer, Switzerland (2024a). https:\/\/doi.org\/10.1007\/978-3-031-48796-5_13","DOI":"10.1007\/978-3-031-48796-5_13"},{"key":"473_CR9","doi-asserted-by":"publisher","unstructured":"Brownlee, A.E., Petke, J., Alexander, B., Barr, E.T., Wagner, M., White, D.R.: Gin: genetic improvement research made easy. In: GECCO, pp. 985\u2013993 (2019). https:\/\/doi.org\/10.1145\/3321707.3321841","DOI":"10.1145\/3321707.3321841"},{"key":"473_CR10","doi-asserted-by":"publisher","unstructured":"Brownlee, A.E.I., Petke, J., Rasburn, A.F.: Injecting shortcuts for faster running Java code. In: IEEE CEC 2020, pp. 1\u20138. https:\/\/doi.org\/10.1109\/CEC48606.2020.9185708","DOI":"10.1109\/CEC48606.2020.9185708"},{"key":"473_CR11","unstructured":"Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H.P.d.O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al.: Evaluating large language models trained on code. ArXiv abs\/2107.03374 (2021)"},{"key":"473_CR12","doi-asserted-by":"crossref","unstructured":"Fan, A., Gokkaya, B., Harman, M., Lyubarskiy, M., Sengupta, S., Yoo, S., Zhang, J.M.: Large Language Models for Software Engineering: Survey and Open Problems (2023)","DOI":"10.1109\/ICSE-FoSE59343.2023.00008"},{"key":"473_CR13","doi-asserted-by":"crossref","unstructured":"Fernando, C., Banarse, D.S., Michalewski, H., Osindero, S., Rockt\u00e4schel, T.: Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution (2024). https:\/\/openreview.net\/forum?id=HKkiX32Zw1","DOI":"10.1177\/10597123241262534"},{"key":"473_CR14","doi-asserted-by":"publisher","unstructured":"Guo, Q., Wang, R., Guo, J., Li, B., Song, K., Tan, X., Liu, G., Bian, J., Yang, Y.: Connecting large language models with evolutionary algorithms yields powerful prompt optimizers. CoRR abs\/2309.08532 (2023) https:\/\/doi.org\/10.48550\/ARXIV.2309.08532arXiv:2309.08532","DOI":"10.48550\/ARXIV.2309.08532"},{"issue":"2","key":"473_CR15","doi-asserted-by":"publisher","first-page":"21","DOI":"10.1007\/s10710-024-09494-2","volume":"25","author":"E Hemberg","year":"2024","unstructured":"Hemberg, E., Moskal, S., O\u2019Reilly, U.M.: Evolving code with a large language model. Genet. Program. Evol. Mach. 25(2), 21 (2024)","journal-title":"Genet. Program. Evol. Mach."},{"key":"473_CR16","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2021.3071193","author":"M Hort","year":"2021","unstructured":"Hort, M., Kechagia, M., Sarro, F., Harman, M.: A survey of performance optimization for mobile applications. IEEE Trans. Softw. Eng. (2021). https:\/\/doi.org\/10.1109\/TSE.2021.3071193","journal-title":"IEEE Trans. Softw. Eng."},{"key":"473_CR17","unstructured":"Hou, X., Liu, Y., Yang, Z., Grundy, J., Zhao, Y., Li, L., Wang, K., Luo, X., Lo, D., Wang, H.: Large language models for software engineering: A systematic literature review. arXiv:2308.10620 (2023)"},{"key":"473_CR18","doi-asserted-by":"publisher","unstructured":"Hu, J., Zhang, Q., Yin, H.: Augmenting Greybox Fuzzing with Generative AI (2023). https:\/\/doi.org\/10.48550\/arXiv.2306.06782","DOI":"10.48550\/arXiv.2306.06782"},{"key":"473_CR19","doi-asserted-by":"publisher","unstructured":"Jin, M., Shahriar, S., Tufano, M., Shi, X., Lu, S., Sundaresan, N., Svyatkovskiy, A.: Inferfix: End-to-end program repair with llms. ESEC\/FSE 2023 - Proceedings of the 31st ACM Joint Meeting European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 1646\u20131656 (2023) https:\/\/doi.org\/10.1145\/3611643.3613892","DOI":"10.1145\/3611643.3613892"},{"key":"473_CR20","doi-asserted-by":"publisher","unstructured":"Jin, M., Shahriar, S., Tufano, M., Shi, X., Lu, S., Sundaresan, N., Svyatkovskiy, A.: Inferfix: End-to-end program repair with llms. In: Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. ESEC\/FSE 2023, pp. 1646\u20131656. Association for Computing Machinery, New York, NY, USA (2023). https:\/\/doi.org\/10.1145\/3611643.3613892","DOI":"10.1145\/3611643.3613892"},{"key":"473_CR21","doi-asserted-by":"crossref","unstructured":"Kang, S., Yoo, S.: Towards objective-tailored genetic improvement through large language models. arXiv:2304.09386 (2023)","DOI":"10.1109\/GI59320.2023.00013"},{"key":"473_CR22","doi-asserted-by":"crossref","unstructured":"Kim, D., Nam, J., Song, J., Kim, S.: Automatic Patch Generation Learned from Human-Written Patches (2013). http:\/\/logging.apache.org\/log4j\/","DOI":"10.1109\/ICSE.2013.6606626"},{"issue":"4","key":"473_CR23","doi-asserted-by":"publisher","first-page":"43","DOI":"10.1109\/MS.2021.3071086","volume":"38","author":"S Kirbas","year":"2021","unstructured":"Kirbas, S., Windels, E., Mcbello, O., Kells, K., Pagano, M., Szalanski, R., Nowack, V., Winter, E., Counsell, S., Bowes, D., Hall, T., Haraldsson, S., Woodward, J.: On the introduction of automatic program repair in bloomberg. IEEE Softw. 38(4), 43\u201351 (2021). https:\/\/doi.org\/10.1109\/MS.2021.3071086","journal-title":"IEEE Softw."},{"key":"473_CR24","doi-asserted-by":"publisher","unstructured":"Lemieux, C., Inala, J.P., Lahiri, S.K., Sen, S.: Codamosa: Escaping coverage plateaus in test generation with pre-trained large language models. Proceedings - International Conference on Software Engineering, 919\u2013931 (2023) https:\/\/doi.org\/10.1109\/ICSE48619.2023.00085","DOI":"10.1109\/ICSE48619.2023.00085"},{"key":"473_CR25","doi-asserted-by":"publisher","unstructured":"Marginean, A., Bader, J., Chandra, S., Harman, M., Jia, Y., Mao, K., Mols, A., Scott, A.: Sapfix: Automated end-to-end repair at scale. In: ICSE-SEIP, pp. 269\u2013278 (2019). https:\/\/doi.org\/10.1109\/ICSE-SEIP.2019.00039","DOI":"10.1109\/ICSE-SEIP.2019.00039"},{"key":"473_CR26","doi-asserted-by":"publisher","unstructured":"Murtaza, S.B., Mccoy, A., Ren, Z., Murphy, A., Banzhaf, W.: Llm fault localisation within evolutionary computation based automated program repair. Proceedings of the Genetic and Evolutionary Computation Conference Companion, 1824\u20131829 (2024) https:\/\/doi.org\/10.1145\/3638530.3664174","DOI":"10.1145\/3638530.3664174"},{"key":"473_CR27","unstructured":"Ouyang, S., Zhang, J.M., Harman, M., Wang, M.: LLM is Like a Box of Chocolates: the Non-determinism of ChatGPT in Code Generation (2023). https:\/\/arxiv.org\/abs\/2308.02828"},{"key":"473_CR28","doi-asserted-by":"crossref","unstructured":"Pearce, H., Tan, B., Ahmad, B., Karri, R., Dolan-Gavitt, B.: Examining Zero-Shot Vulnerability Repair with Large Language Models (2022). https:\/\/arxiv.org\/abs\/2112.02125","DOI":"10.1109\/SP46215.2023.10179324"},{"key":"473_CR29","doi-asserted-by":"crossref","unstructured":"Petke, J., Alexander, B., Barr, E.T., Brownlee, A.E.I., Wagner, M., White, D.R.: A survey of genetic improvement search spaces. In: Proceedings of the Genetic and Evolutionary Computation Conference Companion, pp. 1715\u20131721 (2019)","DOI":"10.1145\/3319619.3326870"},{"key":"473_CR30","doi-asserted-by":"publisher","first-page":"415","DOI":"10.1109\/TEVC.2017.2693219","volume":"22","author":"J Petke","year":"2018","unstructured":"Petke, J., Haraldsson, S.O., Harman, M., Langdon, W.B., White, D.R., Woodward, J.R.: Genetic improvement of software: a comprehensive survey. IEEE Trans. Evolut. Comput. 22, 415\u2013432 (2018). https:\/\/doi.org\/10.1109\/TEVC.2017.2693219","journal-title":"IEEE Trans. Evolut. Comput."},{"issue":"4","key":"473_CR31","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s10664-023-10344-5","volume":"28","author":"J Petke","year":"2023","unstructured":"Petke, J., Alexander, B., Barr, E.T., Brownlee, A.E., Wagner, M., White, D.R.: Program transformation landscapes for automated program modification using Gin. Empir. Softw. Eng. 28(4), 1\u201341 (2023). https:\/\/doi.org\/10.1007\/s10664-023-10344-5","journal-title":"Empir. Softw. Eng."},{"key":"473_CR32","doi-asserted-by":"publisher","unstructured":"Sarro, F.: Search-based software engineering in the era of modern software systems. In: 2023 IEEE 31st International Requirements Engineering Conference (RE), pp. 3\u20135 (2023). https:\/\/doi.org\/10.1109\/RE57278.2023.00010","DOI":"10.1109\/RE57278.2023.00010"},{"key":"473_CR33","doi-asserted-by":"publisher","unstructured":"Sobania, D., Briesch, M., Hanna, C., Petke, J.: An analysis of the automatic bug fixing performance of chatgpt. In: 2023 IEEE\/ACM International Workshop on Automated Program Repair (APR), pp. 23\u201330. IEEE Computer Society, Los Alamitos, CA, USA (2023).https:\/\/doi.org\/10.1109\/APR59189.2023.00012","DOI":"10.1109\/APR59189.2023.00012"},{"key":"473_CR34","doi-asserted-by":"publisher","DOI":"10.1007\/s10515-024-00423-2","author":"M Watkinson","year":"2024","unstructured":"Watkinson, M., Brownlee, A.E.I.: Comparing apples and oranges? Investigating the consistency of cpu and memory profiler results across multiple java versions. Autom. Softw. Eng. (2024). https:\/\/doi.org\/10.1007\/s10515-024-00423-2","journal-title":"Autom. Softw. Eng."},{"key":"473_CR35","unstructured":"Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E.H., Le, Q.V., Zhou, D.: Chain-of-thought prompting elicits reasoning in large language models. In: Advances in Neural Information Processing Systems, vol. 35, pp. 24824\u201324837. Curran Associates, Inc., New Orleans, LA, USA (2022)"},{"key":"473_CR36","doi-asserted-by":"publisher","unstructured":"Xia, C.S., Paltenghi, M., Le\u00a0Tian, J., Pradel, M., Zhang, L.: Fuzz4all: Universal fuzzing with large language models. In: Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering. ICSE \u201924. Association for Computing Machinery, New York, NY, USA (2024). https:\/\/doi.org\/10.1145\/3597503.3639121","DOI":"10.1145\/3597503.3639121"},{"key":"473_CR37","doi-asserted-by":"publisher","unstructured":"Xia, C.S., Wei, Y., Zhang, L.: Automated program repair in the era of large pre-trained language models. Proceedings - International Conference on Software Engineering, 1482\u20131494 (2023) https:\/\/doi.org\/10.1109\/ICSE48619.2023.00129","DOI":"10.1109\/ICSE48619.2023.00129"},{"key":"473_CR38","doi-asserted-by":"crossref","unstructured":"Xia, C.S., Zhang, L.: Keep the conversation going: Fixing 162 out of 337 bugs for $0.42 each using chatgpt. arXiv preprint arXiv:2304.00385 (2023)","DOI":"10.1145\/3650212.3680323"},{"key":"473_CR39","unstructured":"Zhang, S., Chen, Z., Shen, Y., Ding, M., Tenenbaum, J.B., Gan, C.: Planning with large language models for code generation. In: The Eleventh International Conference on Learning Representations (ICLR 2023) (2023). https:\/\/openreview.net\/forum?id=Lr8cOOtYbfL"},{"key":"473_CR40","unstructured":"Zhang, Q., Fang, C., Xie, Y., Ma, Y., Sun, W., Yang, Y., Chen, Z.: A systematic literature review on large language models for automated program repair (2024)"}],"container-title":["Automated Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10515-024-00473-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10515-024-00473-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10515-024-00473-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,4,5]],"date-time":"2025-04-05T21:34:51Z","timestamp":1743888891000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10515-024-00473-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,21]]},"references-count":40,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,5]]}},"alternative-id":["473"],"URL":"https:\/\/doi.org\/10.1007\/s10515-024-00473-6","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-4437272\/v1","asserted-by":"object"}]},"ISSN":["0928-8910","1573-7535"],"issn-type":[{"value":"0928-8910","type":"print"},{"value":"1573-7535","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,21]]},"assertion":[{"value":"17 May 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 October 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 January 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"15"}}