{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T03:41:57Z","timestamp":1777866117888,"version":"3.51.4"},"reference-count":46,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T00:00:00Z","timestamp":1773360000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T00:00:00Z","timestamp":1773360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Empir Software Eng"],"published-print":{"date-parts":[[2026,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>While several studies have examined the security of code generated by GPT and other Large Language Models (LLMs), most have relied on controlled experiments rather than real developer interactions. This paper investigates the security of GPT-generated code extracted from the DevGPT dataset and evaluates the ability of current LLMs to detect and repair vulnerabilities in this real-world context. We analysed 2,315 C, C++, and C# code snippets using static scanners combined with manual inspection, identifying 56 vulnerabilities across 48 files. These files were then assessed using GPT-4.1, GPT-5, and Claude Opus 4.1 to determine whether these could identify the security issues and, where applicable, to specify the corresponding Common Weakness Enumeration (CWE) numbers and propose fixes. Manual review and re-scanning of the modified code showed that GPT-4.1, GPT-5, and Claude Opus 4.1 correctly detected 46, 44, and 45 vulnerabilities, and successfully repaired 42, 44, and 43 respectively. A comparison of experiments conducted in October 2024 and September 2025 indicates substantial progress, with overall detection and remediation rates improving from roughly 50% to around 75\u201380%. We also observe that LLM-generated code is about as likely to contain vulnerabilities as developer-written code, and that LLMs may confidently provide incorrect information, posing risks for less experienced developers.<\/jats:p>","DOI":"10.1007\/s10664-026-10812-8","type":"journal-article","created":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T12:02:22Z","timestamp":1773403342000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Secure coding with AI \u2013 from detection to repair"],"prefix":"10.1007","volume":"31","author":[{"given":"Vladislav","family":"Belozerov","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peter J","family":"Barclay","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0023-9543","authenticated-orcid":false,"given":"Ashkan","family":"Sami","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,3,13]]},"reference":[{"key":"10812_CR1","unstructured":"\u201cBandit\u201d tool web site. https:\/\/github.com\/PyCQA\/bandit retrieved 31-Oct-2025"},{"key":"10812_CR2","unstructured":"2024 CWE Top 25 Most Dangerous Software Weaknesses. https:\/\/cwe.mitre.org\/top25\/archive\/2024\/2024_cwe_top25.html retrieved 29-Mar-2025"},{"key":"10812_CR3","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2025.3538050","author":"S Almanasra","year":"2025","unstructured":"Almanasra S, Suwais K (2025) Analysis of GPT-Generated Codes Across Multiple Programming Languages. IEEE Access. https:\/\/doi.org\/10.1109\/ACCESS.2025.3538050","journal-title":"IEEE Access"},{"key":"10812_CR4","unstructured":"Arora C, Sayeed AI, Licorish S, Wang F, Treude C (2024) Optimizing Large Language Model Hyperparameters for Code Generation. https:\/\/arxiv.org\/abs\/2408.10577v1 retrieved 31-Oct-2025"},{"key":"10812_CR5","unstructured":"Bakhshandeh A, Keramatfar A, Norouzi A, Chekidehkhoun MM (2023) Using GPT as a static application security testing tool. arxiv:2308.14434"},{"key":"10812_CR6","doi-asserted-by":"crossref","unstructured":"Borji A (2023) A categorical archive of GPT failures. arxiv:2302.03494","DOI":"10.21203\/rs.3.rs-2895792\/v1"},{"key":"10812_CR7","unstructured":"Cheshkov A, Zadorozhny P, Levichev R (2023) Evaluation of GPT model for vulnerability detection. arxiv:2304.07232"},{"key":"10812_CR8","unstructured":"Claude Opus model web site. https:\/\/www.anthropic.com\/claude\/opus retrieved 26-Oct-2025"},{"key":"10812_CR9","unstructured":"Cloc tool. https:\/\/github.com\/AlDanial\/cloc retrieved 17-Nov-2024"},{"key":"10812_CR10","unstructured":"CodeBert web site. https:\/\/github.com\/microsoft\/CodeBERT retrieved 25-Mar-2025"},{"key":"10812_CR11","unstructured":"CodeForces web site. https:\/\/codeforces.com\/ retrieved 09-Apr-2025"},{"key":"10812_CR12","unstructured":"CppCheck scanner. https:\/\/github.com\/danmar\/cppcheck retrieved 10-Nov-2024"},{"key":"10812_CR13","unstructured":"CWE database. https:\/\/cwe.mitre.org\/ retrieved 20-Nov-2024"},{"key":"10812_CR14","unstructured":"DeepSeek web site. https:\/\/www.deepseek.com\/ retrieved 27-Mar-2025"},{"key":"10812_CR15","unstructured":"DevGPT. https:\/\/zenodo.org\/records\/16392320 retrieved 06-Nov-2024"},{"key":"10812_CR16","unstructured":"Flawfinder scanner. https:\/\/dwheeler.com\/flawfinder retrieved 10-Nov-2024"},{"issue":"8","key":"10812_CR17","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3716848","volume":"34","author":"Y Fu","year":"2023","unstructured":"Fu Y, Liang P, Tahir A, Li Z, Shahin M, Yu J, Chen J (2023) Security Weaknesses of Copilot Generated Code in GitHub: An empirical study. ACM Trans Softw Eng Methodol 34(8):1\u201334","journal-title":"ACM Trans Softw Eng Methodol"},{"key":"10812_CR18","unstructured":"Full results analysis table for October 2024. https:\/\/github.com\/vlad2010\/MSc_DataEngineering_Results\/blob\/main\/OpenAI_API_Issues64_Experiment\/SecurityIssueAnalysis_2024.ods retrieved 09-Nov-2025"},{"key":"10812_CR19","unstructured":"Full results analysis table for September 2025. https:\/\/github.com\/vlad2010\/MSc_DataEngineering_Results\/blob\/main\/DetectAndFix_Data\/2025_09_24\/SecurityIssuesAnalysis.ods retrieved 09-Nov-2025"},{"key":"10812_CR20","doi-asserted-by":"publisher","unstructured":"Fu M, Tantithamthavorn CK, Nguyen V, Le T (2023) GPT for Vulnerability Detection, Classification, and Repair: How Far Are We? In: Proceedings - Asia-Pacific Software Engineering Conference APSEC pp 632\u2013636. https:\/\/doi.org\/10.1109\/APSEC60848.2023.00085","DOI":"10.1109\/APSEC60848.2023.00085"},{"key":"10812_CR21","unstructured":"GitHub. https:\/\/github.com. Accessed 18 Nov 2024"},{"key":"10812_CR46","unstructured":"GitHub (2025) Repository with software tools used in the study. https:\/\/github.com\/vlad2010\/MS_DataEngineering_Public. Accessed\u00a025 Jan 2026"},{"key":"10812_CR22","doi-asserted-by":"publisher","first-page":"18","DOI":"10.1016\/j.infsof.2015.08.002","volume":"68","author":"K Goseva-Popstojanova","year":"2015","unstructured":"Goseva-Popstojanova K, Perhinschi A (2015) On the capability of static code analysis to detect security vulnerabilities. Inf Softw Technol 68:18\u201333. https:\/\/doi.org\/10.1016\/j.infsof.2015.08.002","journal-title":"Inf Softw Technol"},{"key":"10812_CR23","unstructured":"Hacker news. https:\/\/news.ycombinator.com\/news retrieved 06-Nov-2024"},{"key":"10812_CR24","doi-asserted-by":"crossref","unstructured":"Hamer S, d\u2019Amorim M, Williams L (2024) Just another copy and paste? Comparing the security vulnerabilities of GPT generated code and StackOverflow answers. In 2024 IEEE Security and Privacy Workshops (SPW). IEEE, pp 87\u201394","DOI":"10.1109\/SPW63631.2024.00014"},{"key":"10812_CR25","doi-asserted-by":"publisher","unstructured":"Jamdade M, Liu Y (2024) A pilot study on secure code generation with GPT for Web applications. In: Proceedings of the 2024 ACM Southeast Conference, ACMSE 2024, pp 229\u2013234. https:\/\/doi.org\/10.1145\/3603287.3651194","DOI":"10.1145\/3603287.3651194"},{"key":"10812_CR26","doi-asserted-by":"crossref","unstructured":"Kharma M, Choi S, AlKhanafseh M, Mohaisen D (2025) Security and quality in LLM-generated code: a multi-language, multi-model analysis. arxiv:2502.01853","DOI":"10.1109\/TDSC.2026.3672745"},{"key":"10812_CR27","doi-asserted-by":"publisher","unstructured":"Kholoosi MM, Babar MA, Croft R (2024). A Qualitative Study on Using GPT for Software Security: Perception vs. Practicality. In: 2024 IEEE 6th International Conference on Trust, Privacy and Security in Intelligent Systems, and Applications (TPS-ISA), pp 107\u2013117. https:\/\/doi.org\/10.1109\/TPS-ISA62245.2024.00022","DOI":"10.1109\/TPS-ISA62245.2024.00022"},{"key":"10812_CR28","doi-asserted-by":"publisher","unstructured":"Khoury R, Avila AR, Brunelle J, Camara BM (2023) How secure is code generated by GPT? In: Conference Proceedings - IEEE international conference on systems man and cybernetics, pp 2445\u20132451. https:\/\/doi.org\/10.1109\/SMC53992.2023.10394237","DOI":"10.1109\/SMC53992.2023.10394237"},{"key":"10812_CR29","doi-asserted-by":"crossref","unstructured":"Kulenovic M, Donko D (2014) A survey of static code analysis methods for security vulnerabilities detection. In: 2014 37th International Convention on Information and Communication Technology, Electronics and Microelectronics (MIPRO). IEEE, pp 1381-1386","DOI":"10.1109\/MIPRO.2014.6859783"},{"key":"10812_CR30","unstructured":"Leetcode web site. https:\/\/leetcode.com\/ retrieved 31-Mar-2023"},{"key":"10812_CR31","unstructured":"Motaleb M, Manik H (2025) GPT vs. DeepSeek: A Comparative Study on AI-Based Code Generation. arxiv:2502.18467"},{"key":"10812_CR32","unstructured":"OpenAI models web site. https:\/\/platform.openai.com\/docs\/models retrieved 26-Oct-2025"},{"key":"10812_CR33","doi-asserted-by":"crossref","unstructured":"Pearce H, Ahmad B, Tan B, Dolan-Gavitt B, Karri R (2021). Asleep at the Keyboard? Assessing the security of GitHub copilot\u2019s code contributions. arxiv:2108.09293","DOI":"10.1109\/SP46214.2022.9833571"},{"key":"10812_CR34","doi-asserted-by":"publisher","unstructured":"Perry N, Srivastava M, Kumar D, Boneh D (2023) Do Users Write More Insecure Code with AI Assistants? CCS 2023 - Proceedings of the 2023 ACM SIGSAC conference on computer and communications security, pp 2785\u20132799. https:\/\/doi.org\/10.1145\/3576915.3623157","DOI":"10.1145\/3576915.3623157"},{"key":"10812_CR35","unstructured":"Results GitHub repository. https:\/\/github.com\/vlad2010\/MSc_DataEngineering_Results\/tree\/main\/ retrieved 09-Nov-2025"},{"key":"10812_CR36","doi-asserted-by":"crossref","unstructured":"Rrv A, Tyagi N, Uddin N, Varshney N, Baral C. Chaos with Keywords: Exposing Large Language Models Sycophancy to Misleading Keywords and Evaluating Defense Strategies. https:\/\/aclanthology.org\/2024.findings-acl.755\/ retrieved 19-May-2025","DOI":"10.18653\/v1\/2024.findings-acl.755"},{"key":"10812_CR37","doi-asserted-by":"crossref","unstructured":"Sajadi A, Le B, Nguyen A, Damevski K, Chatterjee P (2025) Do LLMs Consider Security? An Empirical Study on Responses to Programming Questions. arxiv:2502.14202","DOI":"10.1007\/s10664-025-10658-6"},{"key":"10812_CR38","doi-asserted-by":"crossref","unstructured":"Sarif format. https:\/\/docs.oasis-open.org\/sarif\/sarif\/v2.1.0\/sarif-v2.1.0.html retrieved 31-Oct-2024","DOI":"10.46339\/al-mulk.v2i2.1404"},{"key":"10812_CR39","unstructured":"Semgrep scanner. https:\/\/semgrep.dev. Accessed 10 Nov 2024"},{"key":"10812_CR40","unstructured":"Snyk scanner. retrieved 10-Nov-2024 https:\/\/snyk.io"},{"key":"10812_CR41","unstructured":"SonarQube web site. https:\/\/www.sonarsource.com\/products\/sonarqube\/ retrieved 31-Oct-2025"},{"key":"10812_CR42","unstructured":"StackOverflow web site. https:\/\/stackoverflow.com\/, retrieved 18-Nov-2024"},{"issue":"3","key":"10812_CR43","doi-asserted-by":"publisher","first-page":"443","DOI":"10.1348\/014466604X17092","volume":"44","author":"LM Van Swol","year":"2005","unstructured":"Van Swol LM, Sniezek JA (2005) Factors affecting the acceptance of expert advice. Br J Soc Psychol 44(3):443\u2013461","journal-title":"Br J Soc Psychol"},{"key":"10812_CR44","unstructured":"Wu F, Zhang Q, Bajaj AP, Bao T, Zhang N, Wang R, Fish, Xiao C (2023) Exploring the limits of GPT in software security applications. arxiv:2312.05275"},{"issue":"12","key":"10812_CR45","doi-asserted-by":"publisher","first-page":"3435","DOI":"10.1109\/TSE.2024.3492204","volume":"50","author":"X Yu","year":"2024","unstructured":"Yu X, Liu L, Hu X, Keung JW, Liu J, Xia X (2024) Fight Fire With Fire: How Much Can We Trust GPT on Source Code-Related Tasks? IIEEE Trans Software Eng 50(12):3435\u20133453. https:\/\/doi.org\/10.1109\/TSE.2024.3492204","journal-title":"IIEEE Trans Software Eng"}],"container-title":["Empirical Software Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10664-026-10812-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10664-026-10812-8","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10664-026-10812-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T08:03:04Z","timestamp":1777536184000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10664-026-10812-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,13]]},"references-count":46,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,7]]}},"alternative-id":["10812"],"URL":"https:\/\/doi.org\/10.1007\/s10664-026-10812-8","relation":{},"ISSN":["1382-3256","1573-7616"],"issn-type":[{"value":"1382-3256","type":"print"},{"value":"1573-7616","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,13]]},"assertion":[{"value":"18 April 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 January 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 March 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no competing interests to declare.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of Interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical Approval"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Informed Consent"}},{"value":"Clinical trial number: not applicable.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Clinical Trial Number"}}],"article-number":"93"}}