{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T15:02:05Z","timestamp":1782313325278,"version":"3.54.5"},"reference-count":51,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2025,12,5]],"date-time":"2025-12-05T00:00:00Z","timestamp":1764892800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["JCP"],"abstract":"<jats:p>Large language models (LLMs) have shown remarkable potential for automatic code generation. Yet, these models share a weakness with their human counterparts: inadvertently generating code with security vulnerabilities that could allow unauthorized attackers to access sensitive data or systems. In this work, we propose Feedback-Driven Security Patching (FDSP), wherein LLMs automatically refine vulnerable generated code. The key to our approach is a unique framework that leverages automatic static code analysis to enable the LLM to create and implement potential solutions to code vulnerabilities. Further, we curate a novel benchmark, PythonSecurityEval, that can accelerate progress in the field of code generation by covering diverse, real-world applications, including databases, websites, and operating systems. Our proposed FDSP approach achieves the strongest improvements, reducing vulnerabilities by up to 33% when evaluated with Bandit and 12% with CodeQL and outperforming baseline refinement methods.<\/jats:p>","DOI":"10.3390\/jcp5040110","type":"journal-article","created":{"date-parts":[[2025,12,5]],"date-time":"2025-12-05T15:11:24Z","timestamp":1764947484000},"page":"110","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Leveraging Static Analysis for Feedback-Driven Security Patching in LLM-Generated Code"],"prefix":"10.3390","volume":"5","author":[{"given":"Kamel","family":"Alrashedy","sequence":"first","affiliation":[{"name":"School of Interactive Computing, Georgia Institute of Technology, Atlanta, GA 30332, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Abdullah","family":"Aljasser","sequence":"additional","affiliation":[{"name":"School of Interactive Computing, Georgia Institute of Technology, Atlanta, GA 30332, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pradyumna","family":"Tambwekar","sequence":"additional","affiliation":[{"name":"School of Interactive Computing, Georgia Institute of Technology, Atlanta, GA 30332, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5321-6038","authenticated-orcid":false,"given":"Matthew","family":"Gombolay","sequence":"additional","affiliation":[{"name":"School of Interactive Computing, Georgia Institute of Technology, Atlanta, GA 30332, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,12,5]]},"reference":[{"key":"ref_1","unstructured":"Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., and Askell, A. (2020, January 6\u201312). Language models are few-shot learners. Proceedings of the Advances in Neural Information Processing Systems 33, Online."},{"key":"ref_2","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017). Attention is All you Need. Advances in Neural Information Processing Systems, Curran Associates, Inc."},{"key":"ref_3","unstructured":"Rozi\u00e8re, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X.E., Adi, Y., Liu, J., Remez, T., and Rapin, J. (2023). Code llama: Open foundation models for code. arXiv."},{"key":"ref_4","unstructured":"Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., and Bhosale, S. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., and Roman, S. (2018). Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. arXiv.","DOI":"10.18653\/v1\/D18-1425"},{"key":"ref_6","unstructured":"Lachaux, M.A., Roziere, B., Chanussot, L., and Lample, G. (2020). Unsupervised translation of programming languages. arXiv."},{"key":"ref_7","unstructured":"Shypula, A., Madaan, A., Zeng, Y., Alon, U., Gardner, J., Hashemi, M., Neubig, G., Ranganathan, P., Bastani, O., and Yazdanbakhsh, A. (2023). Learning performance-improving code edits. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Pearce, H., Tan, B., Ahmad, B., Karri, R., and Dolan-Gavitt, B. (2023, January 22\u201325). Examining Zero-Shot Vulnerability Repair with Large Language Models. Proceedings of the IEEE Symposium on Security and Privacy (SP), San Francisco, CA, USA.","DOI":"10.1109\/SP46215.2023.10179324"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Wong, M.F., Guo, S., Hang, C.N., Ho, S.W., and Tan, C.W. (2023). Natural language generation and understanding of big code for ai-assisted programming: A review. Entropy, 25.","DOI":"10.3390\/e25060888"},{"key":"ref_10","unstructured":"Hermann, K., Peldszus, S., Stegh\u00f6fer, J.P., and Berger, T. (May, January 27). An Exploratory Study on the Engineering of Security Features. Proceedings of the International Conference on Software Engineering (ICSE), Ottawa, ON, Canada."},{"key":"ref_11","unstructured":"Spiess, C., Gros, D., Pai, K.S., Pradel, M., Rabin, M.R.I., Alipour, A., Jha, S., Devanbu, P., and Ahmed, T. (2024, January 14\u201320). Calibration and correctness of language models for code. Proceedings of the International Conference on Software Engineering (ICSE), Lisbon, Portugal."},{"key":"ref_12","unstructured":"Zhang, T., Yu, Y., Mao, X., Wang, S., Yang, K., Lu, Y., Zhang, Z., and Zhao, Y. (2024, January 14\u201320). Instruct or Interact? Exploring and Eliciting LLMs\u2019 Capability in Code Snippet Adaptation Through Prompt Engineering. Proceedings of the International Conference on Software Engineering (ICSE), Lisbon, Portugal."},{"key":"ref_13","unstructured":"Chen, X., Lin, M., Sch\u00e4rli, N., and Zhou, D. (2023). Teaching large language models to self-debug. arXiv."},{"key":"ref_14","unstructured":"Athiwaratkun, B., Gouda, S.K., and Wang, Z. (2023, January 1\u20135). Multi-lingual Evaluation of Code Generation Models. Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Siddiq, M.L., Casey, B., and Santos, J.C.S. (2023). A Lightweight Framework for High-Quality Code Generation. arXiv.","DOI":"10.1109\/SCAM63643.2024.00020"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Le, H., Sahoo, D., Zhou, Y., Xiong, C., and Savarese, S. (2024, January 9\u201315). INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness. Proceedings of the Thirty-Eighth Annual Conference on Neural Information Processing Systems, Vancouver, BC, Canada.","DOI":"10.52202\/079017-2717"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Tony, C., Mutas, M., D\u00edaz Ferreyra, N., and Scandariato, R. (2023, January 15\u201316). LLMSecEval: A Dataset of Natural Language Prompts for Security Evaluations. Proceedings of the 2023 IEEE\/ACM 20th International Conference on Mining Software Repositories (MSR), Melbourne, Australia.","DOI":"10.1109\/MSR59073.2023.00084"},{"key":"ref_18","unstructured":"Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K. (2023, January 1\u20135). Let\u2019s verify step by step. Proceedings of the Twelfth International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_19","unstructured":"Huang, J., Chen, X., Mishra, S., Zheng, H.S., Yu, A.W., Song, X., and Zhou, D. (2023, January 1\u20135). Large Language Models Cannot Self-Correct Reasoning Yet. Proceedings of the Twelfth International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Siddiq, M., and Santos, J. (2022, January 18). SecurityEval Dataset: Mining Vulnerability Examples to Evaluate Machine Learning-Based Code Generation Techniques. Proceedings of the 1st International Workshop on Mining Software Repositories Applications for Privacy and Security (MSR4P S22), Virtually.","DOI":"10.1145\/3549035.3561184"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zheng, Q., Xia, X., Zou, X., Dong, Y., Wang, S., Xue, Y., Wang, Z., Shen, L., Wang, A., and Li, Y. (2023). Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x. arXiv.","DOI":"10.1145\/3580305.3599790"},{"key":"ref_22","unstructured":"Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., and Le, Q. (2021). Program synthesis with large language models. arXiv."},{"key":"ref_23","unstructured":"Zhou, S., Alon, U., Xu, F.F., Wang, Z., Jiang, Z., and Neubig, G. (2023, January 1\u20135). DocPrompting: Generating Code by Retrieving the Docs. Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda."},{"key":"ref_24","unstructured":"Nijkamp, E., Hayashi, H., Xiong, C., Savarese, S., and Zhou, Y. (2023, January 1\u20135). CodeGen2: Lessons for Training LLMs on Programming and Natural Languages. Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda."},{"key":"ref_25","first-page":"27865","article-title":"Self-supervised bug detection and repair","volume":"34","author":"Allamanis","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_26","unstructured":"Rasooli, M.S., and Tetreault, J.R. (2015). Yara Parser: A Fast and Accurate Dependency Parser. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Nam, D., Macvean, A., Hellendoorn, V., Vasilescu, B., and Myers, B. (2024, January 14\u201320). Using an LLM to Help With Code Understanding. Proceedings of the 2024 IEEE\/ACM 46th International Conference on Software Engineering (ICSE), Lisbon, Portugal.","DOI":"10.1145\/3597503.3639187"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"911","DOI":"10.1109\/TSE.2024.3368208","article-title":"Software testing with large language models: Survey, landscape, and vision","volume":"50","author":"Wang","year":"2024","journal-title":"IEEE Trans. Softw. Eng."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Aggarwal, P., Madaan, A., Yang, Y. (2023, January 6\u201310). Let\u2019s Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMs. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore.","DOI":"10.18653\/v1\/2023.emnlp-main.761"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Alrashedy, K., Hellendoorn, V.J., and Orso, A. (2023). Learning Defect Prediction from Unrealistic Data. arXiv.","DOI":"10.1109\/SANER60148.2024.00063"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"3280","DOI":"10.1109\/TSE.2021.3087402","article-title":"Deep learning based vulnerability detection: Are we there yet?","volume":"48","author":"Chakraborty","year":"2021","journal-title":"IEEE Trans. Softw. Eng."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Andrew, G., and Gao, J. (2007, January 20\u201324). Scalable training of L1-regularized log-linear models. Proceedings of the 24th International Conference on Machine Learning, Corvallis, OR, USA.","DOI":"10.1145\/1273496.1273501"},{"key":"ref_33","first-page":"1817","article-title":"A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data","volume":"6","author":"Ando","year":"2005","journal-title":"J. Mach. Learn. Res."},{"key":"ref_34","unstructured":"Bhatt, M., Chennabasappa, S., Nikolaidis, C., Wan, S., Evtimov, I., Gabi, D., Song, D., Ahmad, F., Aschermann, C., and Fontana, L. (2023). Purple llama cyberseceval: A secure coding benchmark for language models. arXiv."},{"key":"ref_35","unstructured":"Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., and Yang, Y. (2023, January 10\u201316). Self-refine: Iterative refinement with self-feedback. Proceedings of the Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_36","unstructured":"Olausson, T.X., Inala, J.P., Wang, C., Gao, J., and Solar-Lezama, A. (2023, January 1\u20135). Is Self-Repair a Silver Bullet for Code Generation?. Proceedings of the Twelfth International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_37","unstructured":"Gou, Z., Shao, Z., Gong, Y., Shen, Y., Yang, Y., Duan, N., and Chen, W. (2023). CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Elgohary, A., Meek, C., Richardson, M., Fourney, A., Ramos, G., and Awadallah, A.H. (2021, January 6\u201311). NL-EDIT: Correcting Semantic Parse Errors through Natural Language Interaction. Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), Online.","DOI":"10.18653\/v1\/2021.naacl-main.444"},{"key":"ref_39","unstructured":"Bai, Y., Jones, A., and Ndousse, K. (2023). Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv."},{"key":"ref_40","unstructured":"Yasunaga, M., and Liang, P. (2020, January 13\u201318). Graph-based, self-supervised program repair from diagnostic feedback. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Bafatakis, N., Boecker, N., Boon, W., Salazar, M.C., Krinke, J., Oznacar, G., and White, R. (2019, January 25\u201331). Python coding style compliance on stack overflow. Proceedings of the 2019 IEEE\/ACM 16th International Conference on Mining Software Repositories (MSR), Montreal, QC, Canada.","DOI":"10.1109\/MSR.2019.00042"},{"key":"ref_42","unstructured":"Peng, B., Galley, M., He, P., Cheng, H., Xie, Y., Hu, Y., Huang, Q., Liden, L., Yu, Z., and Chen, W. (2023). Check your facts and try again: Improving large language models with external knowledge and automated feedback. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Yang, K., Tian, Y., Peng, N., and Klein, D. (2022, January 7\u201311). Re3: Generating longer stories with recursive reprompting and revision. Proceedings of the Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates.","DOI":"10.18653\/v1\/2022.emnlp-main.296"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Wang, B., Shin, R., Liu, X., Polozov, O., and Richardson, M. (2020, January 5\u201310). RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online.","DOI":"10.18653\/v1\/2020.acl-main.677"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Scholak, T., Schucher, N., and Bahdanau, D. (2021, January 7\u201311). PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from Language Models. Proceedings of the Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic.","DOI":"10.18653\/v1\/2021.emnlp-main.779"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Gusfield, D. (1997). Algorithms on Strings, Trees and Sequences, Cambridge University Press.","DOI":"10.1017\/CBO9780511574931"},{"key":"ref_47","unstructured":"Zhuo, T.Y., Vu, M.C., Chim, J., Hu, H., Yu, W., Widyasari, R., Yusuf, I.N.B., Zhan, H., He, J., and Paul, I. (2024). BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions. arXiv."},{"key":"ref_48","unstructured":"Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G. (2023, January 23\u201329). Pal: Program-aided language models. Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Aky\u00fcrek, A.F., Aky\u00fcrek, E., Kalyan, A., Clark, P., Wijaya, D., and Tandon, N. (2023, January 9\u201314). RL4F: Generating natural language feedback with reinforcement learning for repairing model outputs. Proceedings of the Annual Meeting of the Association of Computational Linguistics 2023, Toronto, ON, Canada.","DOI":"10.18653\/v1\/2023.acl-long.427"},{"key":"ref_50","unstructured":"Aho, A.V., and Ullman, J.D. (1972). The Theory of Parsing, Translation and Compiling, Prentice-Hall."},{"key":"ref_51","unstructured":"(2025, March 04). CodeQL. Available online: https:\/\/codeql.github.com."}],"container-title":["Journal of Cybersecurity and Privacy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2624-800X\/5\/4\/110\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T05:24:26Z","timestamp":1765257866000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2624-800X\/5\/4\/110"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,5]]},"references-count":51,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["jcp5040110"],"URL":"https:\/\/doi.org\/10.3390\/jcp5040110","relation":{},"ISSN":["2624-800X"],"issn-type":[{"value":"2624-800X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,5]]}}}