{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T02:35:39Z","timestamp":1784342139742,"version":"3.55.0"},"reference-count":47,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2025,7,25]],"date-time":"2025-07-25T00:00:00Z","timestamp":1753401600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,7,25]],"date-time":"2025-07-25T00:00:00Z","timestamp":1753401600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100005760","name":"University of Gothenburg","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100005760","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Empir Software Eng"],"published-print":{"date-parts":[[2025,9]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Automatic Program Repair (APR) has garnered significant attention as a practical research domain focused on automatically fixing bugs in programs. While existing APR techniques primarily target imperative programming languages like C and Java, there is a growing need for effective solutions applicable to declarative software specification languages. This paper systematically investigates the capacity of Large Language Models (LLMs) to repair declarative specifications in Alloy, a declarative formal language used for software specification. We designed six different repair settings, encompassing single-agent and dual-agent paradigms, utilizing various LLMs. These configurations also incorporate different levels of feedback, including an auto-prompting mechanism for generating prompts autonomously using LLMs. Our study reveals that dual-agent with auto-prompting setup outperforms the other settings, albeit with a marginal increase in the number of iterations and token usage. This dual-agent setup demonstrated superior effectiveness compared to state-of-the-art Alloy APR techniques when evaluated on a comprehensive set of benchmarks. This work is the first to empirically evaluate LLM capabilities to repair declarative specifications, while taking into account recent trending LLM concepts such as LLM-based agents, feedback, auto-prompting, and tools, thus paving the way for future agent-based techniques in software engineering.<\/jats:p>","DOI":"10.1007\/s10664-025-10687-1","type":"journal-article","created":{"date-parts":[[2025,7,25]],"date-time":"2025-07-25T14:46:11Z","timestamp":1753454771000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["An empirical evaluation of pre-trained large language models for repairing declarative formal specifications"],"prefix":"10.1007","volume":"30","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7108-3809","authenticated-orcid":false,"given":"Mohannad","family":"Alhanahnah","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-3417-4352","authenticated-orcid":false,"given":"Md","family":"Rashedul Hasan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3465-4056","authenticated-orcid":false,"given":"Lisong","family":"Xu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6686-466X","authenticated-orcid":false,"given":"Hamid","family":"Bagheri","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,7,25]]},"reference":[{"key":"10687_CR1","doi-asserted-by":"publisher","unstructured":"Alhanahnah M, Stevens C, Bagheri H (2020) Scalable analysis of interaction threats in iot systems. In: Proceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis, Association for Computing Machinery, New York, NY, USA, ISSTA 2020, p 272\u2013285.\u00a0https:\/\/doi.org\/10.1145\/3395363.3397347\u00a0","DOI":"10.1145\/3395363.3397347"},{"key":"10687_CR2","unstructured":"AutoGPT (2024) AutoGPT. https:\/\/github.com\/Significant-Gravitas\/AutoGPT Accessed: 30-March-2024"},{"issue":"3","key":"10687_CR3","doi-asserted-by":"publisher","first-page":"54","DOI":"10.1007\/s10664-020-09932-6","volume":"26","author":"H Bagheri","year":"2021","unstructured":"Bagheri H, Wang J, Aerts J, Ghorbani N, Malek S (2021) Flair: efficient analysis of android inter-component vulnerabilities in response to incremental changes. Empir Softw Eng 26(3):54","journal-title":"Empir Softw Eng"},{"key":"10687_CR4","doi-asserted-by":"crossref","unstructured":"Bagheri H, Kang E, Mansoor N (2020) Synthesis of assurance cases for software certification. In: ICSE-NIER 42nd international conference on software engineering, New Ideas and Emerging Results, Seoul, South Korea, 27 June - 19 July, 2020, ACM, pp 61\u201364","DOI":"10.1145\/3377816.3381728"},{"key":"10687_CR5","doi-asserted-by":"publisher","unstructured":"Bagheri H, Sadeghi A, Jabbarvand R, Malek S (2016) Practical, formal synthesis and automatic enforcement of security policies for android. In: 2016 46th Annual IEEE\/IFIP International Conference on Dependable Systems and Networks (DSN), pp 514\u2013525. https:\/\/doi.org\/10.1109\/DSN.2016.53","DOI":"10.1109\/DSN.2016.53"},{"key":"10687_CR6","doi-asserted-by":"publisher","unstructured":"Bouzenia I, Devanbu P, Pradel M (2025) RepairAgent: An Autonomous, LLM-Based Agent for Program Repair. In: 2025 IEEE\/ACM 47th International Conference on Software Engineering (ICSE), IEEE Computer Society, Los Alamitos, CA, USA, pp 2188\u20132200. https:\/\/doi.org\/10.1109\/ICSE55347.2025.00157, URL https:\/\/doi.ieeecomputersociety.org\/10.1109\/ICSE55347.2025.00157","DOI":"10.1109\/ICSE55347.2025.00157"},{"key":"10687_CR7","doi-asserted-by":"publisher","unstructured":"Brida SG, Regis G, Zheng G, Bagheri H, Nguyen T, Aguirre N, Frias M (2021) Bounded exhaustive search of alloy specification repairs. In: Proceedings of the 43rd international conference on software engineering, IEEE Press, ICSE \u201921, p 1135\u20131147, https:\/\/doi.org\/10.1109\/ICSE43902.2021.00105","DOI":"10.1109\/ICSE43902.2021.00105"},{"key":"10687_CR8","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A et al (2020) Language models are few-shot learners. Adv Neural Inf Process Syst 33:1877\u20131901","journal-title":"Adv Neural Inf Process Syst"},{"key":"10687_CR9","unstructured":"Bubeck S, Chandrasekaran V, Eldan R, Gehrke J, Horvitz E, Kamar E, Lee P, Lee YT, Li Y, Lundberg S, Nori H, Palangi H, Ribeiro MT, Zhang Y (2023) Sparks of artificial general intelligence: Early experiments with gpt-4. https:\/\/arxiv.org\/abs\/2303.12712,2303.12712"},{"key":"10687_CR10","unstructured":"Chase H (2022) LangChain. https:\/\/github.com\/langchain-ai\/langchain Accessed: 27-Feb-2024"},{"key":"10687_CR11","unstructured":"Chen M, Tworek J, Jun H, Yuan Q, Pinto HP, Kaplan J, Edwards H, Burda Y, Joseph N, Brockman G, Ray A (2021) Evaluating large language models trained on code.\u00a0ArXiv abs\/2107.03374. https:\/\/api.semanticscholar.org\/CorpusID:235755472"},{"key":"10687_CR12","unstructured":"Chowdhery A, Narang S, Devlin J, Bosma M, Mishra G, Roberts A, Barham P, Chung HW, Sutton C, Gehrmann S, et al. (2023) Palm: Scaling language modeling with pathways.\u00a0J Mach Learn Res\u00a024(240):1\u2013113"},{"key":"10687_CR13","doi-asserted-by":"publisher","unstructured":"Fan Z, Gao X, Mirchev M, Roychoudhury A, Tan SH (2023) Automated repair of programs from large language models. In: Proceedings of the 45th International Conference on Software Engineering, IEEE Press, ICSE \u201923, pp 1469\u20131481. https:\/\/doi.org\/10.1109\/ICSE48619.2023.00128","DOI":"10.1109\/ICSE48619.2023.00128"},{"key":"10687_CR14","doi-asserted-by":"publisher","unstructured":"Guti\u00e9rrez\u00a0Brida S, Regis G, Zheng G, Bagheri H, Nguyen T, Aguirre N, Frias M (2023) Icebar: Feedback-driven iterative repair of alloy specifications. In: Proceedings of the 37th IEEE\/ACM international conference on automated software engineering, association for computing machinery, New York, NY, USA, ASE \u201922. https:\/\/doi.org\/10.1145\/3551349.3556944","DOI":"10.1145\/3551349.3556944"},{"key":"10687_CR15","doi-asserted-by":"publisher","unstructured":"Hasan MR, Li J, Ahmed I, Bagheri H (2025) Automated repair of declarative software specifications in the era of large language models. https:\/\/doi.org\/10.21227\/yt7g-bk72","DOI":"10.21227\/yt7g-bk72"},{"key":"10687_CR16","unstructured":"Jackson D (2006) Software abstractions: logic, language, and analysis. The MIT Press"},{"key":"10687_CR17","doi-asserted-by":"publisher","unstructured":"Jain N, Vaidyanath S, Iyer A, Natarajan N, Parthasarathy S, Rajamani S, Sharma R (2022) Jigsaw: large language models meet program synthesis. In: Proceedings of the 44th international conference on software engineering, Association for Computing Machinery, New York, NY, USA, ICSE \u201922, pp 1219\u20131231. https:\/\/doi.org\/10.1145\/3510003.3510203","DOI":"10.1145\/3510003.3510203"},{"key":"10687_CR18","doi-asserted-by":"publisher","unstructured":"Kang S, Chen B, Yoo S, Lou J (2025) Explainable automated debugging via large language model-driven scientific debugging.\u00a0Empir Softw Eng\u00a030(2). https:\/\/doi.org\/10.1007\/s10664-024-10594-x","DOI":"10.1007\/s10664-024-10594-x"},{"key":"10687_CR19","doi-asserted-by":"publisher","unstructured":"Khurshid S, Marinov D (2004) Testera: Specification-based testing of java programs using sat. Automated Software Engineering 11(4):403\u2013434. https:\/\/doi.org\/10.1023\/B:AUSE.0000038938.10589.b9","DOI":"10.1023\/B:AUSE.0000038938.10589.b9"},{"key":"10687_CR20","doi-asserted-by":"publisher","unstructured":"Kong J, Xie X, Cheng M, Liu S, Du X, Guo Q (2025) Contrastrepair: enhancing conversation-based automated program repair via contrastive test case pairs. ACM Trans Softw Eng Methodol. https:\/\/doi.org\/10.1145\/3719345","DOI":"10.1145\/3719345"},{"key":"10687_CR21","unstructured":"Langroid (2024). https:\/\/github.com\/langroid\/langroid Accessed: 27-Feb-2024"},{"key":"10687_CR22","doi-asserted-by":"crossref","unstructured":"Liu Z, Tang Y, Li M, Jin X, Long Y, Zhang LF, Luo X (2024) Llm-compdroid: repairing configuration compatibility bugs in android apps with pre-trained large language models. 2402.15078","DOI":"10.1145\/3736406"},{"key":"10687_CR23","doi-asserted-by":"publisher","unstructured":"Macedo N, Cunha A, Pereira J, Carvalho R, Silva R, Paiva AC, Sozinho\u00a0Ramalho M, Silva D (2021) Experiences on teaching alloy with an automated assessment platform. Sci Comput Program 211(C). https:\/\/doi.org\/10.1016\/j.scico.2021.102690","DOI":"10.1016\/j.scico.2021.102690"},{"key":"10687_CR24","unstructured":"Madaan A, Tandon N, Gupta P, Hallinan S, Gao L, Wiegreffe S, Alon U, Dziri N, Prabhumoye S, Yang Y, Gupta S, Majumder BP, Hermann K, Welleck S, Yazdanbakhsh A, Clark P (2023) Self-refine: iterative refinement with self-feedback. In: Proceedings of the 37th International Conference on Neural Information Processing Systems, Curran Associates Inc., Red Hook, NY, USA, NIPS \u201923"},{"key":"10687_CR25","doi-asserted-by":"publisher","unstructured":"Mirzaei N, Garcia J, Bagheri H, Sadeghi A, Malek S (2016) Reducing combinatorics in gui testing of android applications. In: Proceedings of the 38th International Conference on Software Engineering, Association for Computing Machinery, New York, NY, USA, ICSE \u201916, pp 559\u2013570. https:\/\/doi.org\/10.1145\/2884781.2884853","DOI":"10.1145\/2884781.2884853"},{"key":"10687_CR26","unstructured":"OpenAI, Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, Aleman FL, Almeida D, Altenschmidt J, Altman S, Anadkat S, Avila R, Babuschkin I, Balaji S, Balcom V, Baltescu P, Bao H, Bavarian M, Belgum J, Bello I, Berdine J, Bernadett-Shapiro G, Berner C, Bogdonoff L, Boiko O, Boyd M, Brakman AL, Brockman G, Brooks T, Brundage M, Button K, Cai T, Campbell R, Cann A, Carey B, Carlson C, Carmichael R, Chan B, Chang C, Chantzis F, Chen D, Chen S, Chen R, Chen J, Chen M, Chess B, Cho C, Chu C, Chung HW, Cummings D, Currier J, Dai Y, Decareaux C, Degry T, Deutsch N, Deville D, Dhar A, Dohan D, Dowling S, Dunning S, Ecoffet A, Eleti A, Eloundou T, Farhi D, Fedus L, Felix N, Fishman SP, Forte J, Fulford I, Gao L, Georges E, Gibson C, Goel V, Gogineni T, Goh G, Gontijo-Lopes R, Gordon J, Grafstein M, Gray S, Greene R, Gross J, Gu SS, Guo Y, Hallacy C, Han J, Harris J, He Y, Heaton M, Heidecke J, Hesse C, Hickey A, Hickey W, Hoeschele P, Houghton B, Hsu K, Hu S, Hu X, Huizinga J, Jain S, Jain S, Jang J, Jiang A, Jiang R, Jin H, Jin D, Jomoto S, Jonn B, Jun H, Kaftan T, Lukasz Kaiser, Kamali A, Kanitscheider I, Keskar NS, Khan T, Kilpatrick L, Kim JW, Kim C, Kim Y, Kirchner JH, Kiros J, Knight M, Kokotajlo D, Lukasz Kondraciuk, Kondrich A, Konstantinidis A, Kosic K, Krueger G, Kuo V, Lampe M, Lan I, Lee T, Leike J, Leung J, Levy D, Li CM, Lim R, Lin M, Lin S, Litwin M, Lopez T, Lowe R, Lue P, Makanju A, Malfacini K, Manning S, Markov T, Markovski Y, Martin B, Mayer K, Mayne A, McGrew B, McKinney SM, McLeavey C, McMillan P, McNeil J, Medina D, Mehta A, Menick J, Metz L, Mishchenko A, Mishkin P, Monaco V, Morikawa E, Mossing D, Mu T, Murati M, Murk O, M\u00b4ely D, Nair A, Nakano R, Nayak R, Neelakantan A, Ngo R, Noh H, Ouyang L, O\u2019Keefe C, Pachocki J, Paino A, Palermo J, Pantuliano A, Parascandolo G, Parish J, Parparita E, Passos A, Pavlov M, Peng A, Perelman A, de Avila Belbute Peres F, Petrov M, de Oliveira Pinto HP, Michael, Pokorny, Pokrass M, Pong VH, Powell T, Power A, Power B, Proehl E, Puri R, Radford A, Rae J, Ramesh A, Raymond C, Real F, Rimbach K, Ross C, Rotsted B, Roussez H, Ryder N, Saltarelli M, Sanders T, Santurkar S, Sastry G, Schmidt H, Schnurr D, Schulman J, Selsam D, Sheppard K, Sherbakov T, Shieh J, Shoker S, Shyam P, Sidor S, Sigler E, Simens M, Sitkin J, Slama K, Sohl I, Sokolowsky B, Song Y, Staudacher N, Such FP, Summers N, Sutskever I, Tang J, Tezak N, Thompson MB, Tillet P, Tootoonchian A, Tseng E, Tuggle P, Turley N, Tworek J, Uribe JFC, Vallone A, Vijayvergiya A, Voss C, Wainwright C, Wang JJ, Wang A, Wang B, Ward J, Wei J, Weinmann C, Welihinda A, Welinder P, Weng J, Weng L, Wiethoff M, Willner D, Winter C, Wolrich S, Wong H, Workman L, Wu S, Wu J, Wu M, Xiao K, Xu T, Yoo S, Yu K, Yuan Q, Zaremba W, Zellers R, Zhang C, Zhang M, Zhao S, Zheng T, Zhuang J, Zhuk W, Zoph B (2024) Gpt-4 technical report. URL https:\/\/arxiv.org\/abs\/2303.08774,2303.08774"},{"key":"10687_CR27","unstructured":"OpenAI (2023b) New models and developer products announced at DevDay \u2014 openai.com. https:\/\/openai.com\/blog\/new-models-and-developer-products-announced-at-devday. Accessed 29-March-2024"},{"key":"10687_CR28","unstructured":"Paul R, Hossain MM, Siddiq ML, Hasan M, Iqbal A, Santos JCS (2023) Enhancing automated program repair through fine-tuning and prompt engineering. URL https:\/\/arxiv.org\/abs\/2304.07840,2304.07840"},{"key":"10687_CR29","doi-asserted-by":"publisher","unstructured":"Salerno F, Al-Kaswan A, Izadi M (2025) How much do code language models remember? An investigation on data extraction attacks before and after fine-tuning. In: 2025 IEEE\/ACM 22nd International Conference on Mining Software Repositories (MSR), IEEE Computer Society, Los Alamitos, CA, USA, pp 465\u2013477. https:\/\/doi.org\/10.1109\/MSR66628.2025.00080, URL https:\/\/doi.ieeecomputersociety.org\/10.1109\/MSR66628.2025.00080","DOI":"10.1109\/MSR66628.2025.00080"},{"issue":"1","key":"10687_CR30","doi-asserted-by":"publisher","first-page":"85","DOI":"10.1109\/TSE.2023.3334955","volume":"50","author":"M Sch\u00e4fer","year":"2024","unstructured":"Sch\u00e4fer M, Nadi S, Eghbali A, Tip F (2024) An empirical evaluation of using large language models for automated unit test generation. IEEE Trans Softw Eng 50(1):85\u201310. https:\/\/doi.org\/10.1109\/TSE.2023.3334955","journal-title":"IEEE Trans Softw Eng"},{"key":"10687_CR31","unstructured":"Shinn N, Cassano F, Gopinath A, Narasimhan K, Yao S (2023) Reflexion: language agents with verbal reinforcement learning. In: Proceedings of the 37th International Conference on Neural Information Processing Systems, Curran Associates Inc., Red Hook, NY, USA, NIPS \u201923"},{"key":"10687_CR32","doi-asserted-by":"publisher","unstructured":"Shin T, Razeghi Y, Logan IV RL, Wallace E, Singh S (2020)\u00a0AutoPrompt: eliciting knowledge from language models with automatically generated prompts. In: Webber B, Cohn T, He Y, Liu Y (eds) Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics, Online, pp 4222\u20134235. https:\/\/doi.org\/10.18653\/v1\/2020.emnlp-main.346, URL https:\/\/aclanthology.org\/2020.emnlp-main.346\/","DOI":"10.18653\/v1\/2020.emnlp-main.346"},{"key":"10687_CR33","unstructured":"Thoppilan R, Freitas DD, Hall J, Shazeer N, Kulshreshtha A, Cheng HT, Jin A, Bos T, Baker L, Du Y, Li Y, Lee H,Zheng HS, Ghafouri A, Menegali M, Huang Y, Krikun M, Lepikhin D, Qin J, Chen D, Xu Y, Chen Z, Roberts A, Bosma M, Zhao V, Zhou Y, Chang CC, Krivokon I, Rusch W, Pickett M, Srinivasan P, Man L, Meier-Hellstern K, Morris MR, Doshi T, Santos RD, Duke T, Soraker J, Zevenbergen B, Prabhakaran V, Diaz M, Hutchinson B, Olson K, Molina A, Hoffman-John E, Lee J, Aroyo L, Rajakumar R, Butryna A, Lamm M, Kuzmina V, Fenton J, Cohen A, Bernstein R, Kurzweil R, Aguera-Arcas B, Cui C, Croak M, Chi E, Le Q (2022) Lamda: Language models for dialog applications. URL https:\/\/arxiv.org\/abs\/2201.08239,2201.08239"},{"key":"10687_CR34","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Lu, Polosukhin I (2017) Attention is all you need. In: Guyon I, Luxburg UV, Bengio S, Wallach H, Fergus R, Vishwanathan S, Garnett R (eds) Advances in Neural Information Processing Systems, Curran Associates, Inc., vol 30. URL https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2017\/file\/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf"},{"key":"10687_CR35","doi-asserted-by":"publisher","unstructured":"Wang L, Ma C, Feng X, Zhang Z, Yang H, Zhang J, Chen Z, Tang J, Chen X, Lin Y, et al. (2024) A survey on large language model based autonomous agents.\u00a0Front Comput Sci 18(6):186345. https:\/\/doi.org\/10.1007\/s11704-024-40231-1","DOI":"10.1007\/s11704-024-40231-1"},{"key":"10687_CR36","doi-asserted-by":"publisher","unstructured":"Wang K, Sullivan A, Khurshid S (2019) Arepair: a repair framework for alloy. In: 2019 IEEE\/ACM 41st international conference on software engineering: companion proceedings (ICSE-Companion), pp 103\u2013106. https:\/\/doi.org\/10.1109\/ICSE-Companion.2019.00049","DOI":"10.1109\/ICSE-Companion.2019.00049"},{"key":"10687_CR37","doi-asserted-by":"publisher","unstructured":"Wu Y, Li Z, Zhang JM, Liu Y (2024) Condefects: A complementary dataset to address the data leakage concern for llm-based fault localization and program repair. In: Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, Association for Computing Machinery, New York, NY, USA, FSE 2024, pp 642\u2013646.\u00a0https:\/\/doi.org\/10.1145\/3663529.3663815","DOI":"10.1145\/3663529.3663815"},{"key":"10687_CR38","doi-asserted-by":"crossref","unstructured":"Xia CS, Paltenghi M, Tian JL, Pradel M, Zhang L (2024) Fuzz4ALL: Universal Fuzzing with Large Language Models. In: 2024 IEEE\/ACM 46th International Conference on Software Engineering (ICSE), IEEE Computer Society, Los Alamitos, CA, USA, pp 1547\u20131559. URL https:\/\/doi.ieeecomputersociety.org\/","DOI":"10.1145\/3597503.3639121"},{"key":"10687_CR39","doi-asserted-by":"publisher","unstructured":"Xia CS, Wei Y, Zhang L (2023) Automated program repair in the era of large pre-trained language models. In: Proceedings of the 45th international conference on software engineering, IEEE Press, ICSE \u201923, p 1482\u20131494. https:\/\/doi.org\/10.1109\/ICSE48619.2023.00129","DOI":"10.1109\/ICSE48619.2023.00129"},{"key":"10687_CR40","doi-asserted-by":"publisher","unstructured":"Xia CS, Zhang L (2022) Less training, more repairing please: Revisiting automated program repair via zero-shot learning. In: Proceedings of the 30th ACM joint european software engineering conference and symposium on the foundations of software engineering, Association for Computing Machinery, New York, NY, USA, ESEC\/FSE 2022, pp 959\u2013971. https:\/\/doi.org\/10.1145\/3540250.3549101","DOI":"10.1145\/3540250.3549101"},{"key":"10687_CR41","unstructured":"Xia CS, Zhang L (2023a) Conversational automated program repair. URL https:\/\/arxiv.org\/abs\/2301.13246,2301.13246"},{"key":"10687_CR42","doi-asserted-by":"publisher","unstructured":"Xia CS, Zhang L (2024) Automated program repair via conversation: Fixing 162 out of 337 bugs for $0.42 each using chatgpt. In: Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, Association for Computing Machinery, New York, NY, USA, ISSTA 2024, pp 819\u2013831. https:\/\/doi.org\/10.1145\/3650212.3680323","DOI":"10.1145\/3650212.3680323"},{"key":"10687_CR43","doi-asserted-by":"publisher","unstructured":"Zhang Q, Fang C, Ma Y, Sun W, Chen Z (2023) A survey of learning-based automated program repair. ACM Trans Softw Eng Methodol 33(2). https:\/\/doi.org\/10.1145\/3631974","DOI":"10.1145\/3631974"},{"key":"10687_CR44","doi-asserted-by":"publisher","unstructured":"Zhang Y, Ruan H, Fan Z, Roychoudhury A (2024) Autocoderover: Autonomous program improvement. In: Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, Association for Computing Machinery, New York, NY, USA, ISSTA 2024, pp 1592\u20131604. https:\/\/doi.org\/10.1145\/3650212.3680384","DOI":"10.1145\/3650212.3680384"},{"key":"10687_CR45","doi-asserted-by":"publisher","unstructured":"Zheng G, Nguyen T, Brida SG, Regis G, Aguirre N, Frias MF, Bagheri H (2022) Atr: Template-based repair for alloy specifications. In: Proceedings of the 31st ACM SIGSOFT international symposium on software testing and analysis, Association for Computing Machinery, New York, NY, USA, ISSTA 2022, pp 666\u2013677. https:\/\/doi.org\/10.1145\/3533767.3534369","DOI":"10.1145\/3533767.3534369"},{"key":"10687_CR46","unstructured":"Zhou W, Jiang YE, Li L, Wu J, Wang T, Qiu S, Zhang J, Chen J, Wu R, Wang S, Zhu S, Chen J, Zhang W, Tang X, Zhang N, Chen H, Cui P, Sachan M (2023a) Agents: An open-source framework for autonomous language agents. https:\/\/arxiv.org\/abs\/2309.07870,2309.07870"},{"key":"10687_CR47","unstructured":"Zhou W, Jiang YE, Li L, Wu J, Wang T, Qiu S, Zhang J, Chen J, Wu R, Wang S, Zhu S, Chen J, Zhang W, Zhang N, Chen H, Cui P, Sachan M (2023b) Agents: an open-source framework for autonomous language agents"}],"container-title":["Empirical Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10664-025-10687-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10664-025-10687-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10664-025-10687-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,13]],"date-time":"2025-09-13T08:53:49Z","timestamp":1757753629000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10664-025-10687-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,25]]},"references-count":47,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2025,9]]}},"alternative-id":["10687"],"URL":"https:\/\/doi.org\/10.1007\/s10664-025-10687-1","relation":{},"ISSN":["1382-3256","1573-7616"],"issn-type":[{"value":"1382-3256","type":"print"},{"value":"1573-7616","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,25]]},"assertion":[{"value":"10 June 2025","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 July 2025","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of Interests"}}],"article-number":"149"}}