{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T16:40:40Z","timestamp":1784738440759,"version":"3.55.0"},"reference-count":66,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2025,9,7]],"date-time":"2025-09-07T00:00:00Z","timestamp":1757203200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["JCP"],"abstract":"<jats:p>Large language models (LLMs) have become essential in various use cases, such as code generation, reasoning, or translation. Applications vary from language understanding to decision making. Despite this rapid evolution, significant concerns appear regarding the security of these models and the vulnerabilities they present. In this research, we present an overview of the common LLM models, and their design components and architectures. Moreover, we present their domains of applications. Following that, we present the main security concerns associated with LLMs as defined in different security referentials and standards such as OWASP, MITRE, and NIST. Moreover, we present prior research that focuses on the security concerns in LLMs. Finally, we conduct a comparative study of the performance and robustness of several models against various attack scenarios. We highlight the behavior differences of these models, which prove the importance of giving more attention for the security aspect when using or designing LLMs.<\/jats:p>","DOI":"10.3390\/jcp5030071","type":"journal-article","created":{"date-parts":[[2025,9,10]],"date-time":"2025-09-10T09:32:01Z","timestamp":1757496721000},"page":"71","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Vulnerability Detection in Large Language Models: Addressing Security Concerns"],"prefix":"10.3390","volume":"5","author":[{"given":"Sahar","family":"Ben Yaala","sequence":"first","affiliation":[{"name":"InnoV\u2019COM Laboratory-Sup\u2019Com, University of Carthage, Ariana 2083, Tunisia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ridha","family":"Bouallegue","sequence":"additional","affiliation":[{"name":"InnoV\u2019COM Laboratory-Sup\u2019Com, University of Carthage, Ariana 2083, Tunisia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,9,7]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1955","DOI":"10.1109\/COMST.2024.3465447","article-title":"Large language model (llm) for telecommunications: A comprehensive survey on principles, key techniques, and opportunities","volume":"27","author":"Zhou","year":"2024","journal-title":"IEEE Commun. Surv. Tutor."},{"key":"ref_2","unstructured":"Shao, J. (2025, March 05). First Token Probabilities are Unreliable Indicators for LLM Knowledge. Available online: https:\/\/www2.eecs.berkeley.edu\/Pubs\/TechRpts\/2024\/EECS-2024-114.pdf."},{"key":"ref_3","unstructured":"Yang, K., Liu, J., Wu, J., Yang, C., Fung, Y.R., Li, S., Huang, Z., Cao, X., Wang, X., and Wang, Y. (2024). If llm is the wizard, then code is the wand: A survey on how code empowers large language models to serve as intelligent agents. arXiv."},{"key":"ref_4","unstructured":"Zhao, H., Liu, Z., Wu, Z., Li, Y., Yang, T., Shu, P., Xu, S., Dai, H., Zhao, L., and Jiang, H. (2024). Revolutionizing finance with llms: An overview of applications and insights. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"55","DOI":"10.3991\/ijet.v11i06.5644","article-title":"Choosing the right learning management system (LMS) for the higher education institution context: A systematic review","volume":"11","author":"Kasim","year":"2016","journal-title":"Int. J. Emerg. Technol. Learn."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Goyal, S., Rastogi, E., Rajagopal, S.P., Yuan, D., Zhao, F., Chintagunta, J., Naik, G., and Ward, J. (2024, January 4\u20138). Healai: A healthcare llm for effective medical documentation. Proceedings of the 17th ACM International Conference on Web Search and Data Mining, Merida, Mexico.","DOI":"10.1145\/3616855.3635739"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Ye, Q., Axmed, M., Pryzant, R., and Khani, F. (2023). Prompt engineering a prompt engineer. arXiv.","DOI":"10.18653\/v1\/2024.findings-acl.21"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Wu, X.-K., Chen, M., Li, W., Wang, R., Lu, L., Liu, J., Hwang, K., Hao, Y., Pan, Y., and Meng, Q. (2025). LLM Fine-Tuning: Concepts, Opportunities, and Challenges. Big Data Cogn. Comput., 9.","DOI":"10.3390\/bdcc9040087"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Islam, R., and Moushi, O.M. (2025, April 13). Gpt-4o: The cutting-edge advancement in multimodal llm. Authorea Preprints. Available online: https:\/\/www.techrxiv.org\/users\/771522\/articles\/1121145-gpt-4o-the-cutting-edge-advancement-in-multimodal-llm.","DOI":"10.36227\/techrxiv.171986596.65533294\/v1"},{"key":"ref_10","first-page":"19","article-title":"Large Language Model (LLM) Comparison Between Gpt-3 and Palm-2 to Produce Indonesian Cultural Content","volume":"130","author":"Erlansyah","year":"2024","journal-title":"East.-Eur. J. Enterp. Technol."},{"key":"ref_11","unstructured":"Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., and Bhosale, S. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"100184","DOI":"10.1016\/j.mcpdig.2024.11.005","article-title":"Fine-tuning llms for specialized use cases","volume":"3","author":"Anisuzzaman","year":"2024","journal-title":"Mayo Clin. Proc. Digit. Health"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"100211","DOI":"10.1016\/j.hcc.2024.100211","article-title":"A survey on large language model (llm) security and privacy: The good, the bad, and the ugly","volume":"4","author":"Yao","year":"2024","journal-title":"High-Confid. Comput."},{"key":"ref_14","unstructured":"Xu, H., Wang, S., Li, N., Wang, K., Zhao, Y., Chen, K., Yu, T., Liu, Y., and Wang, H. (2024). Large language models for cyber security: A systematic literature review. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"126176","DOI":"10.1109\/ACCESS.2024.3450388","article-title":"A security risk taxonomy for prompt-based interaction with large language models","volume":"12","author":"Derner","year":"2024","journal-title":"IEEE Access"},{"key":"ref_16","first-page":"9","article-title":"Extracting Training Data: Risks and solutions in the context of LLM security","volume":"12","author":"Gerasimenko","year":"2024","journal-title":"Int. J. Open Inf. Technol."},{"key":"ref_17","unstructured":"Pedro, R., Castro, D., Carreira, P., and Santos, N. (2023). From prompt injections to sql injection attacks: How protected is your llm-integrated web application?. arXiv."},{"key":"ref_18","unstructured":"Jiang, M., Liu, K.Z., Zhong, M., Schaeffer, R., Ouyang, S., Han, J., and Koyejo, S. (2024). Investigating data contamination for pre-training language models. arXiv."},{"key":"ref_19","unstructured":"Yu, M., Fang, J., Zhou, Y., Fan, X., Wang, K., Pan, S., and Wen, Q. (2024). LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models. arXiv."},{"key":"ref_20","unstructured":"Wu, S., Fei, H., Li, X., Ji, J., Zhang, H., Chua, T.-S., and Yan, S. (2024). Towards semantic equivalence of tokenization in multimodal llm. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C.D. (2014, January 25\u201329). Glove: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"155","DOI":"10.1017\/S1351324916000334","article-title":"Word2Vec","volume":"23","author":"Church","year":"2017","journal-title":"Nat. Lang. Eng."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Luo, Q., Zeng, W., Chen, M., Peng, G., Yuan, X., and Yin, Q. (2023, January 21\u201324). Self-Attention and Transformers: Driving the Evolution of Large Language Models. Proceedings of the 2023 IEEE 6th International Conference on Electronic Information and Communication Technology (ICEICT), Qingdao, China.","DOI":"10.1109\/ICEICT57916.2023.10245906"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Geva, M., Caciularu, A., Wang, K.R., and Goldberg, Y. (2022). Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. arXiv.","DOI":"10.18653\/v1\/2022.emnlp-main.3"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1111\/emip.12590","article-title":"Using OpenAI GPT to generate reading comprehension items","volume":"43","author":"Sayin","year":"2024","journal-title":"Educ. Meas. Issues Pract."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Islam, R., and Ahmed, I. (2024, January 10\u201312). Gemini-the most powerful LLM: Myth or Truth. Proceedings of the 2024 5th Information Communication Technologies Conference (ICTC), Nanjing, China.","DOI":"10.1109\/ICTC61510.2024.10602253"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"100063","DOI":"10.1016\/j.ajoint.2024.100063","article-title":"Testing the power of Google DeepMind: Gemini versus ChatGPT 4 facing a European ophthalmology examination","volume":"1","author":"Giannuzzi","year":"2024","journal-title":"AJO Int."},{"key":"ref_28","unstructured":"LearnLM Team, Modi, A., Veerubhotla, A.S., Rysbek, A., Huber, A., Wiltshire, B., Veprek, B., Gillick, D., Kasenberg, D., and Ahmed, D. (2024). LearnLM: Improving gemini for learning. arXiv."},{"key":"ref_29","unstructured":"Gemini Team, Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A.M., Hauth, A., and Millican, K. (2023). Gemini: A family of highly capable multimodal models. arXiv."},{"key":"ref_30","unstructured":"Priyanshu, A., Maurya, Y., and Hong, Z. (2024). AI Governance and Accountability: An Analysis of Anthropic\u2019s Claude. arXiv."},{"key":"ref_31","first-page":"184","article-title":"Information Security, Ethics, and Integrity in LLM Agent Interaction","volume":"16","author":"Chen","year":"2024","journal-title":"J. Inf. Secur."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"1132","DOI":"10.1038\/s41433-024-03545-9","article-title":"Benchmarking the performance of large language models in uveitis: A comparative analysis of ChatGPT-3.5, ChatGPT-4.0, Google Gemini, and Anthropic Claude3","volume":"39","author":"Zhao","year":"2024","journal-title":"Eye"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Kumar, V., Srivastava, P., Dwivedi, A., Budhiraja, I., Ghosh, D., Goyal, V., and Arora, R. (2023, January 7\u20138). Large-language-models (llm)-based ai chatbots: Architecture, in-depth analysis and their performance evaluation. Proceedings of the International Conference on Recent Trends in Image Processing and Pattern Recognition, Derby, UK.","DOI":"10.1007\/978-3-031-53085-2_20"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Hassan, E., Bhatnagar, R., and Shams, M.Y. (2023, January 23\u201324). Advancing scientific research in computer science by chatgpt and llama\u2014A review. Proceedings of the International Conference on Intelligent Manufacturing and Energy Sustainability, Hyderabad, India.","DOI":"10.1007\/978-981-99-6774-2_3"},{"key":"ref_35","unstructured":"Roque, L. (2025, March 17). The Evolution of Llama: From Llama 1 to Llama 3.1. Available online: https:\/\/medium.com\/data-science\/the-evolution-of-llama-from-llama-1-to-llama-3-1-13c4ebe96258."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Sasaki, M., Watanabe, N., and Komanaka, T. (2025, March 11). Enhancing Contextual Understanding of Mistral LLM with External Knowledge Bases. Available online: https:\/\/www.researchsquare.com\/article\/rs-4215447\/v1.","DOI":"10.21203\/rs.3.rs-4215447\/v1"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Aydin, O., Karaarslan, E., Erenay, F.S., and Bacanin, N. (2025). Generative AI in Academic Writing: A Comparison of DeepSeek, Qwen, ChatGPT, Gemini, Llama, Mistral, and Gemma. arXiv.","DOI":"10.2139\/ssrn.5133368"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lieberum, T., Rajamanoharan, S., Conmy, A., Smith, L., Sonnerat, N., Varma, V., Kram\u00e1r, J., Dragan, A., Shah, R., and Nanda, N. (2024). Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2. arXiv.","DOI":"10.18653\/v1\/2024.blackboxnlp-1.19"},{"key":"ref_39","unstructured":"(2025, February 07). 2025 Top 10 Risk & Mitigations for LLMs and Gen AI Apps. Available online: https:\/\/genai.owasp.org\/llm-top-10\/."},{"key":"ref_40","unstructured":"(2025, February 15). Navigate Threats to AI Systems Through reAl-World Insights. Available online: https:\/\/atlas.mitre.org\/."},{"key":"ref_41","unstructured":"(2025, February 03). AI Risk Management Framework, Available online: https:\/\/www.nist.gov\/itl\/ai-risk-management-framework."},{"key":"ref_42","unstructured":"Mireshghallah, N., Antoniak, M., More, Y., Choi, Y., and Farnadi, G. (2024). Trust no bot: Discovering personal disclosures in human-llm conversations in the wild. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Shah, S.P., and Deshpande, A.V. (2024, January 20\u201321). Addressing Data Poisoning and Model Manipulation Risks using LLM Models in Web Security. Proceedings of the 2024 International Conference on Distributed Systems, Computer Networks and Cybersecurity (ICDSCNC), Bengaluru, India.","DOI":"10.1109\/ICDSCNC62492.2024.10941696"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3714464","article-title":"Research directions in software supply chain security","volume":"34","author":"Williams","year":"2025","journal-title":"ACM Trans. Softw. Eng. Methodol."},{"key":"ref_45","unstructured":"Zhang, Y., Rando, J., Evtimov, I., Chi, J., Smith, E.M., Carlini, N., Tram\u00e8r, F., and Ippolito, D. (2024). Persistent Pre-Training Poisoning of LLMs. arXiv."},{"key":"ref_46","unstructured":"Wan, A., Wallace, E., Shen, S., and Klein, D. (2023, January 23\u201329). Poisoning language models during instruction tuning. Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA."},{"key":"ref_47","unstructured":"Liu, Y., Deng, G., Li, Y., Wang, K., Wang, Z., Wang, X., Zhang, T., Liu, Y., Wang, H., and Zheng, Y. (2023). Prompt Injection attack against LLM-integrated Applications. arXiv."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Shi, J., Yuan, Z., Liu, Y., Huang, Y., Zhou, P., Sun, L., and Gong, N.Z. (2024, January 14\u201318). Optimization-based prompt injection attack to llm-as-a-judge. Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, Salt Lake City, UT, USA.","DOI":"10.1145\/3658644.3690291"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"2773","DOI":"10.3390\/ai5040134","article-title":"Trustworthy AI: Securing sensitive data in large language models","volume":"5","author":"Feretzakis","year":"2024","journal-title":"AI"},{"key":"ref_50","unstructured":"Bezabih, A., Nourriz, S., and Smith, C. (2024). Toward LLM-Powered Social Robots for Supporting Sensitive Disclosures of Stigmatized Health Conditions. arXiv."},{"key":"ref_51","first-page":"1","article-title":"Large language model supply chain: A research agenda","volume":"34","author":"Wang","year":"2024","journal-title":"ACM Trans. Softw. Eng. Methodol."},{"key":"ref_52","unstructured":"Li, B., Mellou, K., Zhang, B., Pathuri, J., and Menache, I. (2023). Large language models for supply chain optimization. arXiv."},{"key":"ref_53","unstructured":"He, P., Xing, Y., Xu, H., Xiang, Z., and Tang, J. (2025). Multi-Faceted Studies on Data Poisoning can Advance LLM Development. arXiv."},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"618","DOI":"10.1038\/s41591-024-03445-1","article-title":"Medical large language models are vulnerable to data-poisoning attacks","volume":"31","author":"Alber","year":"2025","journal-title":"Nat. Med."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Pathmanathan, P., Chakraborty, S., Liu, X., Liang, Y., and Huang, F. (2024). Is poisoning a real threat to LLM alignment? Maybe more so than you think. arXiv.","DOI":"10.1609\/aaai.v39i26.34968"},{"key":"ref_56","unstructured":"Stoica, I., Zaharia, M., Gonzalez, J., Goldberg, K., Sen, K., Zhang, H., Angelopoulos, A., Patil, S.G., Chen, L., and Chiang, W.-L. (2024). Specifications: The missing link to making the development of llm systems an engineering discipline. arXiv."},{"key":"ref_57","unstructured":"John, S., Del, R.R.F., Evgeniy, K., Helen, O., Idan, H., Kayla, U., Ken, H., Peter, S., Rakshith, A., and Ron, B. (2025, March 10). OWASP Top 10 for LLM Apps & Gen AI Agentic Security Initiative. OWASP. Available online: https:\/\/hal.science\/hal-04985337v1."},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Hui, B., Yuan, H., Gong, N., Burlina, P., and Cao, Y. (2024, January 14\u201318). Pleak: Prompt leaking attacks against large language model applications. Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, Salt Lake City, UT, USA.","DOI":"10.1145\/3658644.3670370"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Agarwal, D., Fabbri, A.R., Risher, B., Laban, P., Joty, S., and Wu, C.-S. (2024, January 12\u201316). Prompt Leakage effect and mitigation strategies for multi-turn LLM Applications. Proceedings of the 2024 Conference on Empirical Methods in Natural Language, Miami, FL, USA. Procdunne2024weaknessesessing: Industry Track.","DOI":"10.18653\/v1\/2024.emnlp-industry.94"},{"key":"ref_60","unstructured":"Jiang, Z., Jin, Z., and He, G. (2024). Safeguarding System Prompts for LLMs. arXiv."},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Dunne, M., Schram, K., and Fischmeister, S. (2024, January 1\u20135). Weaknesses in LLM-Generated Code for Embedded Systems Networking. Proceedings of the 2024 IEEE 24th International Conference on Software Quality, Reliability and Security (QRS), Cambridge, UK.","DOI":"10.1109\/QRS62785.2024.00033"},{"key":"ref_62","unstructured":"Jeyaraman, J. (2025). Vector Databases Unleashed: Isolating Data in Multi-Tenant LLM Systems, Libertatem Media Private Limited."},{"key":"ref_63","unstructured":"Chen, C., and Shu, K. (2023). Can llm-generated misinformation be detected?. arXiv."},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"Huang, T., Yi, J., Yu, P., and Xu, X. (2025). Unmasking Digital Falsehoods: A Comparative Analysis of LLM-Based Misinformation Detection Strategies. arXiv.","DOI":"10.1109\/ICAACE65325.2025.11020217"},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Wan, H., Feng, S., Tan, Z., Wang, H., Tsvetkov, Y., and Luo, M. (2024). Dell: Generating reactions and explanations for llm-based misinformation detection. arXiv.","DOI":"10.18653\/v1\/2024.findings-acl.155"},{"key":"ref_66","doi-asserted-by":"crossref","first-page":"1248","DOI":"10.1109\/TSE.2025.3548168","article-title":"SecureFalcon: Are we there yet in automated software vulnerability detection with LLMs?","volume":"51","author":"Ferrag","year":"2025","journal-title":"IEEE Trans. Softw. Eng."}],"container-title":["Journal of Cybersecurity and Privacy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2624-800X\/5\/3\/71\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:41:28Z","timestamp":1760035288000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2624-800X\/5\/3\/71"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,7]]},"references-count":66,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2025,9]]}},"alternative-id":["jcp5030071"],"URL":"https:\/\/doi.org\/10.3390\/jcp5030071","relation":{},"ISSN":["2624-800X"],"issn-type":[{"value":"2624-800X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,7]]}}}