{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T15:09:25Z","timestamp":1784300965780,"version":"3.55.0"},"reference-count":117,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2024,12,21]],"date-time":"2024-12-21T00:00:00Z","timestamp":1734739200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2024,12,21]],"date-time":"2024-12-21T00:00:00Z","timestamp":1734739200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Empir Software Eng"],"published-print":{"date-parts":[[2025,3]]},"DOI":"10.1007\/s10664-024-10590-1","type":"journal-article","created":{"date-parts":[[2024,12,21]],"date-time":"2024-12-21T08:17:38Z","timestamp":1734769058000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":39,"title":["How secure is AI-generated code: a large-scale comparison of large language models"],"prefix":"10.1007","volume":"30","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9002-5935","authenticated-orcid":false,"given":"Norbert","family":"Tihanyi","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tamas","family":"Bisztray","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohamed Amine","family":"Ferrag","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ridhi","family":"Jain","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lucas C.","family":"Cordeiro","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,12,21]]},"reference":[{"key":"10590_CR1","volume-title":"Compilers: Principles, Techniques, And Tools","author":"AV Aho","year":"2006","unstructured":"Aho AV, Lam MS, Sethi R, Ullman JD (2006) Compilers: Principles, Techniques, And Tools, 2nd edn. Addison-Wesley Longman Publishing Co., Inc, Boston, MA","edition":"2"},{"key":"10590_CR2","unstructured":"Almazrouei E, Alobeidli H, Alshamsi A, Cappelli A, Cojocaru R, Debbah M, Goffinet \u00c9, Hesslow D, Launay J, Malartic Q et al (2023) The falcon series of open language models. arXiv preprint arXiv:2311.16867"},{"key":"10590_CR3","doi-asserted-by":"crossref","unstructured":"Alshmrany KM, Aldughaim M, Bhayat A, Cordeiro LC (2021) Fusebmc: An energy-efficient test generator for finding security vulnerabilities in C programs. In: Loulergue F, Wotawa F (eds) Tests and Proofs - 15th International Conference, TAP 2021, Held as Part of STAF 2021, Virtual Event, June 21-22, 2021, Proceedings. Lecture Notes in Computer Science, vol 12740, pp 85\u2013105. Springer","DOI":"10.1007\/978-3-030-79379-1_6"},{"key":"10590_CR4","unstructured":"Anwar U, Saparov A, Rando J, Paleka D, Turpin M, Hase P, Lubana ES, Jenner E, Casper S, Sourbut O et al (2024) Foundational challenges in assuring alignment and safety of large language models. arXiv preprint arXiv:2404.09932"},{"key":"10590_CR5","unstructured":"Austin J, Odena A, Nye M, Bosma M, Michalewski H, Dohan D, Jiang E, Cai C, Terry M, Le Q et al (2021) Program synthesis with large language models"},{"key":"10590_CR6","doi-asserted-by":"publisher","first-page":"495","DOI":"10.1007\/978-3-031-30820-8_29","volume-title":"Tools and Algorithms for the Construction and Analysis of Systems","author":"D Beyer","year":"2023","unstructured":"Beyer D (2023) Competition on software verification and witness validation: Sv-comp 2023. In: Sankaranarayanan S, Sharygina N (eds) Tools and Algorithms for the Construction and Analysis of Systems. Springer, Cham, pp 495\u2013522"},{"key":"10590_CR7","doi-asserted-by":"publisher","unstructured":"Black PE (2018) A Software Assurance Reference Dataset: Thousands of Programs With Known Bugs. Journal of Research of the National Institute of Standards and Technology. 123:1\u20133. https:\/\/doi.org\/10.6028\/jres.123.005. Accessed 27 Jun 2023","DOI":"10.6028\/jres.123.005"},{"key":"10590_CR8","unstructured":"Braberman VA, Bonomo-Braberman F, Charalambous Y, Colonna JG, Cordeiro LC, Freitas R (2024) Tasks People Prompt: A Taxonomy of LLM Downstream Tasks in Software Verification and Falsification Approaches"},{"key":"10590_CR9","unstructured":"Bui NDQ, Le H, Wang Y, Li J, Gotmare AD, Hoi SCH (2023) CodeTF: One-stop Transformer Library for State-of-the-art Code LLM. arXiv. arxiv:2306.00029. Accessed 22 Jun 2023"},{"key":"10590_CR10","unstructured":"Cao J, Li M, Wen M, Cheung S-c (2023) A study on prompt design, advantages and limitations of chatgpt for deep learning program repair. arXiv preprint arXiv:2304.08191"},{"issue":"9","key":"10590_CR11","doi-asserted-by":"publisher","first-page":"3280","DOI":"10.1109\/TSE.2021.3087402","volume":"48","author":"S Chakraborty","year":"2022","unstructured":"Chakraborty S, Krishna R, Ding Y, Ray B (2022) Deep Learning Based Vulnerability Detection: Are We There Yet? IEEE Trans Software Eng 48(9):3280\u20133296. https:\/\/doi.org\/10.1109\/TSE.2021.3087402","journal-title":"IEEE Trans Software Eng"},{"key":"10590_CR12","unstructured":"Chan A, Kharkar A, Moghaddam RZ, Mohylevskyy Y, Helyar A, Kamal E, Elkamhawy M, Sundaresan N (2023) Transformer-based vulnerability detection in code at edittime: Zero-shot, few-shot, or fine-tuning? arXiv preprint arXiv:2306.01754"},{"key":"10590_CR13","doi-asserted-by":"publisher","unstructured":"Charalambous Y, Tihanyi N, Jain R, Sun Y, Ferrag MA, Cordeiro LC (2023) A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification. arXiv . https:\/\/doi.org\/10.48550\/arXiv.2305.14752. Accessed 31 May 2023","DOI":"10.48550\/arXiv.2305.14752"},{"key":"10590_CR14","doi-asserted-by":"publisher","unstructured":"Chavez MR, Butler TS, Rekawek P, Heo H, Kinzler WL (2023) Chat Generative Pre-trained Transformer: why we should embrace this technology. Am J Obstet Gynecol 228(6):706\u2013711. https:\/\/doi.org\/10.1016\/j.ajog.2023.03.010. Accessed 22 Jun 2023","DOI":"10.1016\/j.ajog.2023.03.010"},{"key":"10590_CR15","doi-asserted-by":"publisher","unstructured":"Chen Y, Ding Z, Alowain L, Chen X, Wagner D (2023) DiverseVul: A New Vulnerable Source Code Dataset for Deep Learning Based Vulnerability Detection. In: Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses. RAID \u201923, pp 654\u2013668. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3607199.3607242","DOI":"10.1145\/3607199.3607242"},{"key":"10590_CR16","unstructured":"Chen M, Tworek J, Jun H, Yuan Q, Oliveira\u00a0Pinto HP, Kaplan J, Edwards H, Burda Y, Joseph N, Brockman G, Ray A, Puri R, Krueger G, Petrov M, Khlaaf H, Sastry G, Mishkin P, Chan B, Gray S, Ryder N, Pavlov M, Power A, Kaiser L, Bavarian M, Winter C, Tillet P, Such FP, Cummings D, Plappert M, Chantzis F, Barnes E, Herbert-Voss A, Guss WH, Nichol A, Paino A, Tezak N, Tang J, Babuschkin I, Balaji S, Jain S, Saunders W, Hesse C, Carr AN, Leike J, Achiam J, Misra V, Morikawa E, Radford A, Knight M, Brundage M, Murati M, Mayer K, Welinder P, McGrew B, Amodei D, McCandlish S, Sutskever I, Zaremba W (2021) Evaluating large language models trained on code. arXiv:2107.03374. [cs.LG]"},{"key":"10590_CR17","doi-asserted-by":"crossref","unstructured":"Cordeiro LC, Kroening D, Schrammel P (2019) JBMC: bounded model checking for java bytecode - (competition contribution). In: Tools and Algorithms for the Construction and Analysis of Systems (TACAS). LNCS, vol 11429, pp 219\u2013223. Springer","DOI":"10.1007\/978-3-030-17502-3_17"},{"issue":"4","key":"10590_CR18","doi-asserted-by":"publisher","first-page":"957","DOI":"10.1109\/TSE.2011.59","volume":"38","author":"L Cordeiro","year":"2012","unstructured":"Cordeiro L, Fischer B, Marques-Silva J (2012) SMT-Based Bounded Model Checking for Embedded ANSI-C Software. IEEE Trans Software Eng 38(4):957\u2013974. https:\/\/doi.org\/10.1109\/TSE.2011.59","journal-title":"IEEE Trans Software Eng"},{"issue":"1","key":"10590_CR19","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1049\/IET-CPS.2018.5006","volume":"5","author":"LC Cordeiro","year":"2020","unstructured":"Cordeiro LC, Lima Filho EB, Bessa IV (2020) Survey on automated symbolic verification and its application for synthesising cyber-physical systems. IET Cyper-Phys Syst Theory Appl 5(1):1\u201324. https:\/\/doi.org\/10.1049\/IET-CPS.2018.5006","journal-title":"IET Cyper-Phys Syst Theory Appl"},{"key":"10590_CR20","doi-asserted-by":"crossref","unstructured":"Cordy JR, Roy CK (2011) The nicad clone detector. 2011 IEEE 19th International Conference on Program Comprehension 219\u2013220","DOI":"10.1109\/ICPC.2011.26"},{"key":"10590_CR21","unstructured":"Deligiannis P, Lal A, Mehrotra N, Rastogi A (2023) Fixing rust compilation errors using llms. arXiv preprint arXiv:2308.05177"},{"issue":"7","key":"10590_CR22","doi-asserted-by":"publisher","first-page":"1165","DOI":"10.1109\/TCAD.2008.923410","volume":"27","author":"V D\u2019Silva","year":"2008","unstructured":"D\u2019Silva V, Kroening D, Weissenbacher G (2008) A Survey of Automated Techniques for Formal Software Verification. IEEE Trans Comput Aided Des Integr Circuits Syst 27(7):1165\u20131178. https:\/\/doi.org\/10.1109\/TCAD.2008.923410","journal-title":"IEEE Trans Comput Aided Des Integr Circuits Syst"},{"key":"10590_CR23","doi-asserted-by":"crossref","unstructured":"Fan Z, Gao X, Mirchev M, Roychoudhury A, Tan SH (2023) Automated repair of programs from large language models. In: 2023 IEEE\/ACM 45th International Conference on Software Engineering (ICSE), pp 1469\u20131481. IEEE","DOI":"10.1109\/ICSE48619.2023.00128"},{"key":"10590_CR24","doi-asserted-by":"publisher","unstructured":"Fan J, Li Y, Wang S, Nguyen TN (2020) A C\/C++ Code Vulnerability Dataset with Code Changes and CVE Summaries. In: Proceedings of the 17th International Conference on Mining Software Repositories. MSR \u201920, pp 508\u2013512. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3379597.3387501 . Accessed 27 Jun 2023","DOI":"10.1145\/3379597.3387501"},{"key":"10590_CR25","doi-asserted-by":"crossref","unstructured":"Gadelha MR, Monteiro FR, Morse J, Cordeiro LC, Fischer B, Nicole DA (2018) Esbmc 5.0: an industrial-strength c model checker. In: Proceedings of the 33rd ACM\/IEEE International Conference on Automated Software Engineering, pp 888\u2013891. ACM, Montpellier, France","DOI":"10.1145\/3238147.3240481"},{"key":"10590_CR26","doi-asserted-by":"publisher","unstructured":"Gadelha MR, Monteiro FR, Morse J, Cordeiro LC, Fischer B, Nicole DA (2018) Esbmc 5.0: an industrial-strength c model checker. In: Proceedings of the 33rd ACM\/IEEE International Conference on Automated Software Engineering. ASE \u201918, pp 888\u2013891. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3238147.3240481","DOI":"10.1145\/3238147.3240481"},{"key":"10590_CR27","doi-asserted-by":"crossref","unstructured":"Gadelha MYR, Monteiro FR, Cordeiro LC, Nicole DA (2019) ESBMC v6.0: Verifying C programs using k-induction and invariant inference - (competition contribution). In: Beyer D, Huisman M, Kordon F, Steffen B (eds) Tools and Algorithms for the Construction and Analysis of Systems (TACAS). LNCS, vol 11429, pp 209\u2013213. Springer","DOI":"10.1007\/978-3-030-17502-3_15"},{"key":"10590_CR28","doi-asserted-by":"publisher","unstructured":"Gadelha MYR, Steffinlongo E, Cordeiro LC, Fischer B, Nicole DA (2019) Smt-based refutation of spurious bug reports in the clang static analyzer. In: Atlee JM, Bultan T, Whittle J (eds) Proceedings of the 41st International Conference on Software Engineering, pp 11\u201314. IEEE \/ ACM, Montreal, QC, Canada. https:\/\/doi.org\/10.1109\/ICSE-Companion.2019.00026","DOI":"10.1109\/ICSE-Companion.2019.00026"},{"issue":"1","key":"10590_CR29","doi-asserted-by":"publisher","first-page":"97","DOI":"10.1007\/s10009-015-0407-9","volume":"19","author":"MYR Gadelha","year":"2017","unstructured":"Gadelha MYR, Ismail HI, Cordeiro LC (2017) Handling loops in bounded model checking of C programs via k-induction. Int J Softw Tools Technol Transf 19(1):97\u2013114. https:\/\/doi.org\/10.1007\/s10009-015-0407-9","journal-title":"Int J Softw Tools Technol Transf"},{"key":"10590_CR30","doi-asserted-by":"crossref","unstructured":"Gao S, Mao W, Gao C, Li L, Hu X, Xia X, Lyu MR (2024) Learning in the wild: Towards leveraging unlabeled data for effectively tuning pre-trained code models. In: Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering, pp 1\u201313","DOI":"10.1145\/3597503.3639216"},{"key":"10590_CR31","unstructured":"Gao Z, Wang H, Zhou Y, Zhu W, Zhang C (2023) How far have we gone in vulnerability detection using large language models. arXiv preprint arXiv:2311.12420"},{"key":"10590_CR32","doi-asserted-by":"crossref","unstructured":"Grishina A, Hort M, Moonen L (2023) The earlybird catches the bug: On exploiting early layers of encoder models for more efficient code classification. In: Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pp 895\u2013907","DOI":"10.1145\/3611643.3616304"},{"key":"10590_CR33","unstructured":"Guo D, Zhu Q, Yang D, Xie Z, Dong K, Zhang W, Chen G, Bi X, Wu Y, Li Y et al (2024) Deepseek-coder: When the large language model meets programming\u2013the rise of code intelligence. arXiv preprint arXiv:2401.14196"},{"key":"10590_CR34","unstructured":"Hao Y, Chen W, Zhou Z, Cui W (2023) E &v: Prompting large language models to perform static analysis by pseudo-code execution and verification. arXiv preprint arXiv:2312.08477"},{"key":"10590_CR35","unstructured":"Honarvar S, Wilk M, Donaldson A (2023) Turbulence: Systematically and automatically testing instruction-tuned large language models for code. arXiv preprint arXiv:2312.14856"},{"key":"10590_CR36","doi-asserted-by":"crossref","unstructured":"Hou X, Zhao Y, Liu Y, Yang Z, Wang K, Li L, Luo X, Lo D, Grundy J, Wang H (2023) Large language models for software engineering: A systematic literature review. ACM Trans Softw Eng Method","DOI":"10.1145\/3695988"},{"key":"10590_CR37","unstructured":"Huang D, Bu Q, Zhang JM, Luck M, Cui H (2023) Agentcoder: Multi-agent-based code generation with iterative testing and optimisation. arXiv preprint arXiv:2312.13010"},{"key":"10590_CR38","unstructured":"Huang Q, Zhu J, Xing Z, Jin H, Wang C, Xu X (2023) A chain of ai-based solutions for resolving fqns and fixing syntax errors in partial code. arXiv preprint arXiv:2306.11981"},{"key":"10590_CR39","doi-asserted-by":"crossref","unstructured":"Imani S, Du L, Shrivastava H (2023) Mathprompter: Mathematical reasoning using large language models. https:\/\/doi.org\/10.48550\/arXiv.2303.05398","DOI":"10.18653\/v1\/2023.acl-industry.4"},{"key":"10590_CR40","unstructured":"Islam NT, Najafirad P (2024) Code security vulnerability repair using reinforcement learning with large language models. arXiv preprint arXiv:2401.07031"},{"key":"10590_CR41","doi-asserted-by":"crossref","unstructured":"Jain R, Gervasoni N, Ndhlovu M, Rawat S (2023) A code centric evaluation of c\/c++ vulnerability datasets for deep learning based vulnerability detection techniques. In: Proceedings of the 16th Innovations in Software Engineering Conference, pp 1\u201310. ACM, Prayagraj, India","DOI":"10.1145\/3578527.3578530"},{"key":"10590_CR42","doi-asserted-by":"crossref","unstructured":"Jain N, Vaidyanath S, Iyer A, Natarajan N, Parthasarathy S, Rajamani S, Sharma R (2022) Jigsaw: Large language models meet program synthesis. In: Proceedings of the 44th International Conference on Software Engineering, pp 1219\u20131231","DOI":"10.1145\/3510003.3510203"},{"key":"10590_CR43","doi-asserted-by":"crossref","unstructured":"Jin M, Shahriar S, Tufano M, Shi X, Lu S, Sundaresan N, Svyatkovskiy A (2023) Inferfix: End-to-end program repair with llms. In: Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pp 1646\u20131656","DOI":"10.1145\/3611643.3613892"},{"key":"10590_CR44","doi-asserted-by":"crossref","unstructured":"Jr FEB, Black PE (2012) The Juliet 1.1 C\/C++ and Java Test Suite. NIST. 45(10):88\u201390. Last Modified: 2021-10-12T11:10-04:00 Publisher: Frederick E. Boland Jr., Paul E. Black. Accessed 2023-05-28","DOI":"10.1109\/MC.2012.345"},{"key":"10590_CR45","unstructured":"Khare A, Dutta S, Li Z, Solko-Breslin A, Alur R, Naik M (2023) Understanding the effectiveness of large language models in detecting security vulnerabilities. arXiv preprint arXiv:2311.16169"},{"key":"10590_CR46","doi-asserted-by":"publisher","unstructured":"Khoury R, Avila AR, Brunelle J, Camara BM (2023) How secure is code generated by chatgpt? In: 2023 IEEE International Conference on Systems, Man, and Cybernetics (SMC), pp 2445\u20132451. https:\/\/doi.org\/10.1109\/SMC53992.2023.10394237","DOI":"10.1109\/SMC53992.2023.10394237"},{"key":"10590_CR47","unstructured":"Kim L, Russell R (2018) Draper VDISC Dataset - Vulnerability Detection in Source Code. Publisher: OSF. https:\/\/osf.io\/d45bw\/ Accessed 27 Jun 2023"},{"key":"10590_CR48","doi-asserted-by":"publisher","unstructured":"Kirova VD, Ku CS, Laracy JR, Marlowe TJ (2024) Software engineering education must adapt and evolve for an llm environment. In: Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1. SIGCSE 2024, pp 666\u2013672. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3626252.3630927","DOI":"10.1145\/3626252.3630927"},{"key":"10590_CR49","doi-asserted-by":"crossref","unstructured":"Kroening D, Tautschnig M (2014) Cbmc\u2013c bounded model checker: (competition contribution). In: Tools and Algorithms for the Construction and Analysis of Systems: TACAS 2014, pp 389\u2013391. Springer, Grenoble, France","DOI":"10.1007\/978-3-642-54862-8_26"},{"key":"10590_CR50","doi-asserted-by":"crossref","unstructured":"Lajk\u00f3 M, Csuvik V, Vid\u00e1cs L (2022) Towards javascript program repair with generative pre-trained transformer (gpt-2). In: Proceedings of the Third International Workshop on Automated Program Repair, pp 61\u201368. IEEE, ???","DOI":"10.1145\/3524459.3527350"},{"key":"10590_CR51","unstructured":"Li T-O, Zong W, Wang Y, Tian H, Wang Y, Cheung S-C (2023) Finding Failure-Inducing Test Cases with ChatGPT"},{"key":"10590_CR52","unstructured":"Liang X, Song S, Zheng Z, Wang H, Yu Q, Li X, Li R-H, Xiong F, Li Z (2024) Internal consistency and self-feedback in large language models: A survey. arXiv preprint arXiv:2407.14507"},{"key":"10590_CR53","unstructured":"Lin F, Kim DJ et al (2024) When llm-based code generation meets the software development process. arXiv preprint arXiv:2403.15852"},{"key":"10590_CR54","unstructured":"Lu S, Guo D, Ren S, Huang J, Svyatkovskiy A, Blanco A, Clement C, Drain D, Jiang D, Tang D et al (2021) Codexglue: A machine learning benchmark dataset for code understanding and generation. arXiv preprint arXiv:2102.04664"},{"key":"10590_CR55","doi-asserted-by":"publisher","unstructured":"Ma W, Liu S, Wang W, Hu Q, Liu Y, Zhang C, Nie L, Liu Y (2023) The Scope of ChatGPT in Software Engineering: A Thorough Investigation. arXiv. https:\/\/doi.org\/10.48550\/arXiv.2305.12138. Accessed 10 Jun 2023","DOI":"10.48550\/arXiv.2305.12138"},{"key":"10590_CR56","unstructured":"Marjam\u00e4ki D (2024) Cppcheck: A Tool for Static Analysis of C\/C++ Code. https:\/\/cppcheck.sourceforge.io\/. [Online], Available at: https:\/\/cppcheck.sourceforge.io\/. Accessed 12 Sept 2024"},{"issue":"4","key":"10590_CR57","doi-asserted-by":"publisher","first-page":"308","DOI":"10.1109\/TSE.1976.233837","volume":"SE\u20132","author":"TJ McCabe","year":"1976","unstructured":"McCabe TJ (1976) A complexity measure. IEEE Trans Softw Eng SE\u20132(4):308\u2013320. https:\/\/doi.org\/10.1109\/TSE.1976.233837","journal-title":"IEEE Trans Softw Eng"},{"key":"10590_CR58","doi-asserted-by":"crossref","unstructured":"Menezes RS, Aldughaim M, Farias B, Li X, Manino E, Shmarov F, Song K, Brau\u00dfe F, Gadelha MR, Tihanyi N, Korovin K, Cordeiro LC (2024) ESBMC v7.4: Harnessing the power of intervals - (competition contribution). In: Tools and Algorithms for the Construction and Analysis of Systems (TACAS). LNCS, vol 14572, pp 376\u2013380. Springer","DOI":"10.1007\/978-3-031-57256-2_24"},{"key":"10590_CR59","doi-asserted-by":"crossref","unstructured":"Menezes R, Moura D, Cavalcante H, Freitas R, Cordeiro LC (2022) Esbmc-jimple: verifying kotlin programs via jimple intermediate representation. In: Ryu S, Smaragdakis Y (eds) ISSTA \u201922: 31st ACM SIGSOFT International Symposium on Software Testing and Analysis, Virtual Event, South Korea, July 18 - 22, 2022, pp 777\u2013780. ACM","DOI":"10.1145\/3533767.3543294"},{"key":"10590_CR60","unstructured":"Mikejo5000 (2024) Code metrics - Cyclomatic complexity - Visual Studio (Windows). https:\/\/learn.microsoft.com\/en-us\/visualstudio\/code-quality\/code-metrics-cyclomatic-complexity?view=vs-2022. Accessed 18 Apr 2024"},{"key":"10590_CR61","unstructured":"Mirzadeh I, Alizadeh K, Shahrokhi H, Tuzel O, Bengio S, Farajtabar M (2024) Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models. arXiv preprint arXiv:2410.05229"},{"key":"10590_CR62","unstructured":"Mohajer MM, Aleithan R, Harzevili NS, Wei M, Belle AB, Pham HV, Wang S (2023) Skipanalyzer: An embodied agent for code analysis with large language models. arXiv preprint arXiv:2310.18532"},{"key":"10590_CR63","doi-asserted-by":"crossref","unstructured":"Morse J, Cordeiro LC, Nicole DA, Fischer B (2011) Context-bounded model checking of LTL properties for ANSI-C software. In: Barthe G, Pardo A, Schneider G (eds) Software Engineering and Formal Methods - 9th International Conference, SEFM 2011, Montevideo, Uruguay, November 14-18, 2011. Proceedings. Lecture Notes in Computer Science, vol 7041, pp 302\u2013317. Springer","DOI":"10.1007\/978-3-642-24690-6_21"},{"key":"10590_CR64","unstructured":"Muennighoff N, Liu Q, Zebaze A, Zheng Q, Hui B, Zhuo TY, Singh S, Tang X, Von\u00a0Werra L, Longpre S (2023) Octopack: Instruction tuning code large language models. arXiv preprint arXiv:2308.07124"},{"key":"10590_CR65","unstructured":"Nehorai N (2024) Analyzing Common Vulnerabilities Introduced by Code-Generative AI | HackerNoon. https:\/\/hackernoon.com\/analyzing-common-vulnerabilities-introduced-by-code-generative-ai. Accessed 28 Feb 2024"},{"key":"10590_CR66","unstructured":"Nguyen V, Yuan X, Wu T, Nepal S, Grobler M, Rudolph C (2024) Deep learning-based out-of-distribution source code data identification: How far we have gone? arXiv preprint arXiv:2404.05964"},{"key":"10590_CR67","unstructured":"Noever D (2023) Can large language models find and fix vulnerable software? arXiv preprint arXiv:2308.10345"},{"key":"10590_CR68","unstructured":"OpenAI (2023) GPT-4 Technical Report. arXiv. arxiv:2303.08774. Accessed 29 May 2023"},{"key":"10590_CR69","unstructured":"Paul R, Mohib\u00a0Hossain M, Hasan M, Iqbal A (2023) Automated program repair based on code review: How do pre-trained transformer models perform? arXiv e-prints, 2304"},{"key":"10590_CR70","doi-asserted-by":"crossref","unstructured":"Pearce H, Ahmad B, Tan B, Dolan-Gavitt B, Karri R (2022) Asleep at the keyboard? assessing the security of github copilot\u2019s code contributions. In: 2022 IEEE Symposium on Security and Privacy (SP), pp 754\u2013768. IEEE, ???","DOI":"10.1109\/SP46214.2022.9833571"},{"key":"10590_CR71","doi-asserted-by":"crossref","unstructured":"Pearce H, Tan B, Ahmad B, Karri R, Dolan-Gavitt B (2023) Examining zero-shot vulnerability repair with large language models. In: 2023 IEEE Symposium on Security and Privacy (SP), pp 2339\u20132356. IEEE, ???","DOI":"10.1109\/SP46215.2023.10179324"},{"key":"10590_CR72","doi-asserted-by":"publisher","unstructured":"Peng Y, Gao S, Gao C, Huo Y, Lyu M (2024) Domain knowledge matters: Improving prompts with fix templates for repairing python type errors. In: Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering. ICSE \u201924. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3597503.3608132","DOI":"10.1145\/3597503.3608132"},{"key":"10590_CR73","doi-asserted-by":"publisher","unstructured":"Perry N, Srivastava M, Kumar D, Boneh D (2023) Do users write more insecure code with ai assistants? In: Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security. CCS \u201923, pp 2785\u20132799. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3576915.3623157","DOI":"10.1145\/3576915.3623157"},{"key":"10590_CR74","unstructured":"Quan VLA, Phat CT, Van\u00a0Nguyen K, Duy PT, Pham V-H (2023) Xgv-bert: Leveraging contextualized language model and graph neural network for efficient software vulnerability detection. arXiv preprint arXiv:2309.14677"},{"key":"10590_CR75","doi-asserted-by":"publisher","unstructured":"Ross SI, Martinez F, Houde S, Muller M, Weisz JD (2023) The Programmer\u2019s Assistant: Conversational Interaction with a Large Language Model for Software Development. In: Proceedings of the 28th International Conference on Intelligent User Interfaces. IUI \u201923, pp. 491\u2013514. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3581641.3584037 . Accessed 22 Jun 2023","DOI":"10.1145\/3581641.3584037"},{"key":"10590_CR76","unstructured":"Roziere B, Gehring J, Gloeckle F, Sootla S, Gat I, Tan XE, Adi Y, Liu J, Remez T, Rapin J et al (2023) Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950"},{"key":"10590_CR77","doi-asserted-by":"publisher","unstructured":"Russell RL, Kim LY, Hamilton LH, Lazovich T, Harer JA, Ozdemir O, Ellingwood PM, McConley MW (2018) Automated Vulnerability Detection in Source Code Using Deep Representation Learning. In: 2018 17th IEEE International Conference on Machine Learning and Applications (ICMLA), pp. 757\u2013762. IEEE, Orlando, FL, USA. https:\/\/doi.org\/10.1109\/ICMLA.2018.00120 . https:\/\/api.semanticscholar.org\/CorpusID:49670513","DOI":"10.1109\/ICMLA.2018.00120"},{"key":"10590_CR78","doi-asserted-by":"crossref","unstructured":"Sadowski C, Yi J (2014) How developers use data race detection tools. In: Proceedings of the 5th Workshop on Evaluation and Usability of Programming Languages and Tools, pp 43\u201351. ACM, Portland, USA","DOI":"10.1145\/2688204.2688205"},{"key":"10590_CR79","unstructured":"Sandoval G, Pearce H, Nys T, Karri R, Garg S, Dolan-Gavitt B (2023) Lost at c: A user study on the security implications of large language model code assistants. In: 32nd USENIX Security Symposium (USENIX Security 23), pp 2205\u20132222. USENIX Association"},{"key":"10590_CR80","doi-asserted-by":"crossref","unstructured":"Shestov A, Cheshkov A, Levichev R, Mussabayev R, Zadorozhny P, Maslov E, Vadim C, Bulychev E (2024) Finetuning large language models for vulnerability detection. arXiv preprint arXiv:2401.17010","DOI":"10.1109\/ACCESS.2025.3546700"},{"key":"10590_CR81","unstructured":"Shumailov I, Shumaylov Z, Zhao Y, Gal Y, Papernot N, Anderson R (2023) The Curse of Recursion: Training on Generated Data Makes Models Forget. arXiv. arxiv:2305.17493. Accessed 2023-06-27"},{"key":"10590_CR82","doi-asserted-by":"publisher","unstructured":"Steenhoek B, Gao H, Le W (2024) Dataflow analysis-inspired deep learning for efficient vulnerability detection. In: Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering. ICSE \u201924. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3597503.3623345","DOI":"10.1145\/3597503.3623345"},{"issue":"10","key":"10590_CR83","doi-asserted-by":"publisher","first-page":"4691","DOI":"10.1109\/TSE.2023.3310874","volume":"49","author":"T Sun","year":"2023","unstructured":"Sun T, Allix K, Kim K, Zhou X, Kim D, Lo D, Bissyand\u00e9 TF, Klein J (2023) Dexbert: Effective, task-agnostic and fine-grained representation learning of android bytecode. IEEE Trans Software Eng 49(10):4691\u20134706. https:\/\/doi.org\/10.1109\/TSE.2023.3310874","journal-title":"IEEE Trans Software Eng"},{"key":"10590_CR84","unstructured":"Sun Y, Wu D, Xue Y, Liu H, Ma W, Zhang L, Shi M, Liu Y (2024) Llm4vuln: A unified evaluation framework for decoupling and enhancing llms\u2019 vulnerability reasoning. arXiv preprint arXiv:2401.16185"},{"key":"10590_CR85","doi-asserted-by":"publisher","unstructured":"Tang W, Tang M, Ban M, Zhao Z, Feng M (2023) Csgvd: A deep learning approach combining sequence and graph embedding for source code vulnerability detection. J Syst Softw 199(C). https:\/\/doi.org\/10.1016\/j.jss.2023.111623","DOI":"10.1016\/j.jss.2023.111623"},{"key":"10590_CR86","unstructured":"Team G, Anil R, Borgeaud S, Wu Y, Alayrac J-B, Yu J, Soricut R, Schalkwyk J, Dai AM, Hauth A et al (2023) Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805"},{"key":"10590_CR87","doi-asserted-by":"publisher","unstructured":"Thapa C, Jang SI, Ahmed ME, Camtepe S, Pieprzyk J, Nepal S (2022) Transformer-based language models for software vulnerability detection. In: Proceedings of the 38th Annual Computer Security Applications Conference. ACSAC \u201922, pp 481\u2013496. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3564625.3567985","DOI":"10.1145\/3564625.3567985"},{"key":"10590_CR88","doi-asserted-by":"publisher","unstructured":"Tian H, Liu K, Kabor\u00e9 AK, Koyuncu A, Li L, Klein J, Bissyand\u00e9 TF (2021) Evaluating representation learning of code changes for predicting patch correctness in program repair. In: Proceedings of the 35th IEEE\/ACM International Conference on Automated Software Engineering. ASE \u201920, pp 981\u2013992. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3324884.3416532","DOI":"10.1145\/3324884.3416532"},{"key":"10590_CR89","doi-asserted-by":"publisher","unstructured":"Tian H, Liu K, Li Y, Kabor\u00e9 AK, Koyuncu A, Habib A, Li L, Wen J, Klein J, Bissyand\u00e9 TF (2023) The best of both worlds: Combining learned embeddings with engineered features for accurate prediction of correct patches. ACM Trans Softw Eng Methodol 32(4). https:\/\/doi.org\/10.1145\/3576039","DOI":"10.1145\/3576039"},{"key":"10590_CR90","doi-asserted-by":"crossref","unstructured":"Tihanyi N, Bisztray T, Dubniczky RA, Toth R, Borsos B, Cherif B, Ferrag MA, Muzsai L, Jain R, Marinelli R et al (2024) Dynamic intelligence assessment: Benchmarking llms on the road to agi with a focus on model confidence. arXiv preprint arXiv:2410.15490","DOI":"10.1109\/BigData62323.2024.10825051"},{"key":"10590_CR91","doi-asserted-by":"publisher","unstructured":"Tihanyi N, Bisztray T, Jain R, Amine\u00a0Ferrag M,\u00a0Cordeiro LC, Mavroeidis V (2023) FormAI Dataset: A Large Collection of AI-Generated C Programs and Their Vulnerability Classifications. IEEE Dataport. https:\/\/doi.org\/10.21227\/vp9n-wv96","DOI":"10.21227\/vp9n-wv96"},{"key":"10590_CR92","doi-asserted-by":"publisher","unstructured":"Tihanyi N, Bisztray T, Jain R, Ferrag MA, Cordeiro LC, Mavroeidis V (2023) The formai dataset: Generative ai in software security through the lens of formal verification. In: Proceedings of the 19th International Conference on Predictive Models and Data Analytics in Software Engineering. PROMISE 2023, pp 33\u201343. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3617555.3617874","DOI":"10.1145\/3617555.3617874"},{"key":"10590_CR93","doi-asserted-by":"crossref","unstructured":"T\u00f3th R, Bisztray T, Erd\u0151di L (2024) Llms in web development: Evaluating llm-generated php code unveiling vulnerabilities and limitations. Computer Safety, Reliability, and Security. SAFECOMP 2024 Workshops. Springer, Cham, pp 425\u2013437","DOI":"10.1007\/978-3-031-68738-9_34"},{"key":"10590_CR94","doi-asserted-by":"publisher","unstructured":"Wallace DR, Fujii RU (1989) Software verification and validation: an overview. IEEE Softw 6(3):10\u201317. https:\/\/doi.org\/10.1109\/52.28119. Accessed 22 Jun 2023","DOI":"10.1109\/52.28119"},{"key":"10590_CR95","doi-asserted-by":"crossref","unstructured":"Wang J, Huang Y, Chen C, Liu Z, Wang S, Wang Q (2024) Software testing with large language models: Survey, landscape, and vision. IEEE Trans Software Eng","DOI":"10.1109\/TSE.2024.3368208"},{"key":"10590_CR96","doi-asserted-by":"crossref","unstructured":"Wang H, Liu Z, Wang S, Cui G, Ding N, Liu Z, Yu G (2023) Intervenor: Prompt the coding ability of large language models with the interactive chain of repairing. arXiv preprint arXiv:2311.09868","DOI":"10.18653\/v1\/2024.findings-acl.124"},{"key":"10590_CR97","unstructured":"Wang S, Long Z, Fan Z, Wei Z, Huang X (2024) Benchmark self-evolving: A multi-agent framework for dynamic llm evaluation. arXiv preprint arXiv:2402.11443"},{"key":"10590_CR98","doi-asserted-by":"publisher","unstructured":"Wang W, Wang Y, Joty S, Hoi SCH (2023) Rap-gen: Retrieval-augmented patch generation with codet5 for automatic program repair. In: Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. ESEC\/FSE 2023, pp 146\u2013158. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3611643.3616256","DOI":"10.1145\/3611643.3616256"},{"key":"10590_CR99","first-page":"24824","volume":"35","author":"J Wei","year":"2022","unstructured":"Wei J, Wang X, Schuurmans D, Bosma M, Xia F, Chi E, Le QV, Zhou D (2022) Chain-of-thought prompting elicits reasoning in large language models. Adv Neural Inf Process Syst 35:24824\u201324837","journal-title":"Adv Neural Inf Process Syst"},{"key":"10590_CR100","doi-asserted-by":"publisher","unstructured":"Wei Y, Xia CS, Zhang L (2023) Copiloting the copilots: Fusing large language models with completion engines for automated program repair. In: Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering. ESEC\/FSE 2023, pp 172\u2013184. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3611643.3616271","DOI":"10.1145\/3611643.3616271"},{"key":"10590_CR101","doi-asserted-by":"publisher","unstructured":"White J, Fu Q, Hays S, Sandborn M, Olea C, Gilbert H, Elnashar A, Spencer-Smith J, Schmidt DC (2023) A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT. arXiv. https:\/\/doi.org\/10.48550\/arXiv.2302.11382. Accessed 24 Jun 2023","DOI":"10.48550\/arXiv.2302.11382"},{"key":"10590_CR102","doi-asserted-by":"crossref","unstructured":"White M, Tufano M, Vendome C, Poshyvanyk D (2016) Deep learning code fragments for code clone detection. In: Proceedings of the 31st IEEE\/ACM International Conference on Automated Software Engineering, pp 87\u201398. Association for Computing Machinery, New York, USA","DOI":"10.1145\/2970276.2970326"},{"key":"10590_CR103","doi-asserted-by":"crossref","unstructured":"Widjojo P, Treude C (2023) Addressing compiler errors: Stack overflow or large language models? arXiv preprint arXiv:2307.10793","DOI":"10.2139\/ssrn.4529345"},{"key":"10590_CR104","doi-asserted-by":"crossref","unstructured":"Wu Y, Jiang N, Pham HV, Lutellier T, Davis J, Tan L, Babkin P, Shah S (2023) How effective are neural networks for fixing security vulnerabilities. In: Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis, pp 1282\u20131294","DOI":"10.1145\/3597926.3598135"},{"key":"10590_CR105","unstructured":"Wu Y, Li Z, Zhang JM, Papadakis M, Harman M, Liu Y (2023) Large language models in fault localisation. arXiv preprint arXiv:2308.15276"},{"key":"10590_CR106","unstructured":"Xia CS, Wei Y, Zhang L (2022) Practical program repair in the era of large pre-trained language models. arXiv preprint arXiv:2210.14179"},{"key":"10590_CR107","doi-asserted-by":"crossref","unstructured":"Xia CS, Zhang L (2023) Keep the conversation going: Fixing 162 out of 337 bugs for \\$0.42 each using chatgpt. arXiv preprint arXiv:2304.00385","DOI":"10.1145\/3650212.3680323"},{"key":"10590_CR108","doi-asserted-by":"crossref","unstructured":"Xu FF, Alon U, Neubig G, Hellendoorn VJ (2022) A systematic evaluation of large language models of code. In: Proceedings of the 6th ACM SIGPLAN International Symposium on Machine Programming, pp 1\u201310","DOI":"10.1145\/3520312.3534862"},{"key":"10590_CR109","doi-asserted-by":"crossref","unstructured":"Yang AZ, Le\u00a0Goues C, Martins R, Hellendoorn V (2024) Large language models for test-free fault localization. In: Proceedings of the 46th IEEE\/ACM International Conference on Software Engineering, pp 1\u201312","DOI":"10.1145\/3597503.3623342"},{"key":"10590_CR110","unstructured":"Yao S, Yu D, Zhao J, Shafran I, Griffiths T, Cao Y, Narasimhan K (2024) Tree of thoughts: Deliberate problem solving with large language models. In: Advances in Neural Information Processing Systems, vol 36"},{"issue":"3","key":"10590_CR111","doi-asserted-by":"publisher","first-page":"474","DOI":"10.1109\/TSE.2024.3354969","volume":"50","author":"Q Zhang","year":"2024","unstructured":"Zhang Q, Fang C, Sun W, Liu Y, He T, Hao X, Chen Z (2024) Appt: Boosting automated patch correctness prediction via fine-tuning pre-trained models. IEEE Trans Software Eng 50(3):474\u2013494. https:\/\/doi.org\/10.1109\/TSE.2024.3354969","journal-title":"IEEE Trans Software Eng"},{"key":"10590_CR112","doi-asserted-by":"crossref","unstructured":"Zhang Q, Fang C, Zhang T, Yu B, Sun W, Chen Z (2023) Gamma: Revisiting template-based automated program repair via mask prediction. In: Proceedings of the 38th IEEE\/ACM International Conference on Automated Software Engineering (ASE), pp 535\u2013547. IEEE","DOI":"10.1109\/ASE56229.2023.00063"},{"key":"10590_CR113","unstructured":"Zhang Y, Jin Z, Xing Y, Li G (2023) Steam: simulating the interactive behavior of programmers for automatic bug fixing. arXiv preprint arXiv:2308.14460"},{"key":"10590_CR114","unstructured":"Zhang Y, Li G, Jin Z, Xing Y (2023) Neural program repair with program dependence analysis and effective filter mechanism. arXiv preprint arXiv:2305.09315"},{"key":"10590_CR115","doi-asserted-by":"crossref","unstructured":"Zhang C, Liu H, Zeng J, Yang K, Li Y, Li H (2023) Prompt-enhanced software vulnerability detection using chatgpt. arXiv preprint arXiv:2308.12697","DOI":"10.1145\/3639478.3643065"},{"key":"10590_CR116","doi-asserted-by":"crossref","unstructured":"Zhao G, Huang J (2018) Deepsim: deep learning code functional similarity. In: Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pp 141\u2013151. ACM, Lake Buena Vista, USA","DOI":"10.1145\/3236024.3236068"},{"key":"10590_CR117","unstructured":"Zhou Y, Liu S, Siow J, Du X, Liu Y (2019) Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural Networks, pp 10197\u201310207. Curran Associates Inc., Red Hook, NY, USA"}],"container-title":["Empirical Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10664-024-10590-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10664-024-10590-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10664-024-10590-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,2]],"date-time":"2025-05-02T13:46:48Z","timestamp":1746193608000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10664-024-10590-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,21]]},"references-count":117,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,3]]}},"alternative-id":["10590"],"URL":"https:\/\/doi.org\/10.1007\/s10664-024-10590-1","relation":{},"ISSN":["1382-3256","1573-7616"],"issn-type":[{"value":"1382-3256","type":"print"},{"value":"1573-7616","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,21]]},"assertion":[{"value":"6 November 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 December 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no competing interests to declare that are relevant to the content of this article.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of interest"}}],"article-number":"47"}}