{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,27]],"date-time":"2026-08-27T01:16:53Z","timestamp":1787793413908,"version":"build-2784847793"},"reference-count":58,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2025,8,10]],"date-time":"2025-08-10T00:00:00Z","timestamp":1754784000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["JCP"],"abstract":"<jats:p>The increasing complexity and volume of cybersecurity logs demand advanced analytical techniques capable of accurate threat detection and explainability. This paper investigates the application of Large Language Models (LLMs), specifically qwen2.5:7b, gemma3:4b, llama3.2:3b, qwen3:8b and qwen2.5:32b to cybersecurity log classification, demonstrating their superior performance compared to traditional machine learning models such as XGBoost, Random Forest, and LightGBM. We present a comprehensive evaluation pipeline that integrates domain-specific prompt engineering, robust parsing of free-text LLM outputs, and uncertainty quantification to enable scalable, automated benchmarking. Our experiments on a vulnerability detection task show that the LLM achieves an F1-score of 0.928 ([0.913, 0.942] 95% CI), significantly outperforming XGBoost (0.555 [0.520, 0.590]) and LightGBM (0.432 [0.380, 0.484]). In addition to superior predictive performance, the LLM generates structured, domain-relevant explanations aligned with classical interpretability methods. These findings highlight the potential of LLMs as interpretable, adaptive tools for operational cybersecurity, making advanced threat detection feasible for SMEs and paving the way for their deployment in dynamic threat environments.<\/jats:p>","DOI":"10.3390\/jcp5030055","type":"journal-article","created":{"date-parts":[[2025,8,11]],"date-time":"2025-08-11T09:59:13Z","timestamp":1754906353000},"page":"55","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Leveraging Large Language Models for Scalable and Explainable Cybersecurity Log Analysis"],"prefix":"10.3390","volume":"5","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1330-4735","authenticated-orcid":false,"given":"Giulia","family":"Palma","sequence":"first","affiliation":[{"name":"Dipartimento di Scienze Sociali Politiche e Cognitive, Universit\u00e0 degli Studi di Siena, 53100 Siena, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-5892-0156","authenticated-orcid":false,"given":"Gaia","family":"Cecchi","sequence":"additional","affiliation":[{"name":"Dipartimento di Scienze Sociali Politiche e Cognitive, Universit\u00e0 degli Studi di Siena, 53100 Siena, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mario","family":"Caronna","sequence":"additional","affiliation":[{"name":"Dipartimento di Scienze Sociali Politiche e Cognitive, Universit\u00e0 degli Studi di Siena, 53100 Siena, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4096-8150","authenticated-orcid":false,"given":"Antonio","family":"Rizzo","sequence":"additional","affiliation":[{"name":"Dipartimento di Scienze Sociali Politiche e Cognitive, Universit\u00e0 degli Studi di Siena, 53100 Siena, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,8,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Khraisat, A., Gondal, I., Vamplew, P., and Kamruzzaman, J. (2019). A Novel Ensemble of Hybrid Intrusion Detection System for Detecting Internet of Things Attacks. Electronics, 8.","DOI":"10.3390\/electronics8111210"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Ban, T., Takahashi, T., Ndichu, S., and Inoue, D. (2023). Breaking Alert Fatigue: AI-Assisted SIEM Framework for Effective Incident Response. Appl. Sci., 13.","DOI":"10.3390\/app13116610"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Krzyszto\u0144, E., Rojek, I., and Miko\u0142ajewski, D. (2024). A Comparative Analysis of Anomaly Detection Methods in IoT Networks: An Experimental Study. Appl. Sci., 14.","DOI":"10.3390\/app142411545"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Henriques, J., Caldeira, F., Cruz, T., and Sim\u00f5es, P. (2020). Combining K-Means and XGBoost Models for Anomaly Detection Using Log Datasets. Electronics, 9.","DOI":"10.3390\/electronics9071164"},{"key":"ref_5","first-page":"1","article-title":"Log-based anomaly detection: A survey","volume":"55","author":"Landauer","year":"2022","journal-title":"ACM Comput. Surv."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"21954","DOI":"10.1109\/ACCESS.2017.2762418","article-title":"A deep learning approach for intrusion detection using recurrent neural networks","volume":"5","author":"Yin","year":"2017","journal-title":"IEEE Access"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zhou, Y., Chen, Y., Rao, X., Zhou, Y., Li, Y., and Hu, C. (2024). Leveraging Large Language Models and BERT for Log Parsing and Anomaly Detection. Mathematics, 12.","DOI":"10.3390\/math12172758"},{"key":"ref_8","first-page":"1877","article-title":"Language Models are Few-Shot Learners","volume":"33","author":"Brown","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Guti\u00e9rrez-Galeano, L., Dom\u00ednguez-Jim\u00e9nez, J.-J., Sch\u00e4fer, J., and Medina-Bulo, I. (2025). LLM-Based Cyberattack Detection Using Network Flow Statistics. Appl. Sci., 15.","DOI":"10.3390\/app15126529"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Wen, X., Zhang, H., Zheng, S., Xu, W., and Bian, J. (2024, January 25\u201329). From Supervised to Generative: A Novel Paradigm for Tabular Deep Learning with Large Language Models. Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD \u201924), Barcelona, Spain.","DOI":"10.1145\/3637528.3671975"},{"key":"ref_11","first-page":"1","article-title":"Explainability for Large Language Models: A Survey","volume":"15","author":"Zhao","year":"2024","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"248","DOI":"10.1145\/3571730","article-title":"Survey of Hallucination in Natural Language Generation","volume":"55","author":"Ji","year":"2023","journal-title":"ACM Comput. Surv."},{"key":"ref_13","unstructured":"Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H.P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., and Brockman, G. (2021). Evaluating Large Language Models Trained on Code. arXiv."},{"key":"ref_14","unstructured":"Tsai, C.-P., Teng, G., Wallis, P., and Ding, W. (2024, January 7\u201311). AnoLLM: Large Language Models for Tabular Anomaly Detection. Proceedings of the 12th International Conference on Learning Representations (ICLR), Vienna, Austria."},{"key":"ref_15","unstructured":"Li, A., Zhao, Y., Qiu, C., Kloft, M., Smyth, P., Rudolph, M., and Mandt, S. (2024). Anomaly Detection of Tabular Data Using LLMs. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Ficili, I., Giacobbe, M., Tricomi, G., and Puliafito, A. (2025). From Sensors to Data Intelligence: Leveraging IoT, Cloud, and Edge Computing with AI. Sensors, 25.","DOI":"10.3390\/s25061763"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Al-Dhamari, N., and Clarke, N. (2024). GPT-Enabled Cybersecurity Training: A Tailored Approach for Effective Awareness. Information Security Education-Challenges in the Digital Age, WISE 2024, Springer. IFIP Advances in Information and Communication Technology.","DOI":"10.1007\/978-3-031-62918-1_1"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Chan, K.C., Gururajan, R., and Carmignani, F. (2025). A Human\u2013AI Collaborative Framework for Cybersecurity Consulting in Capstone Projects for Small Businesses. J. Cybersecur. Priv., 5.","DOI":"10.3390\/jcp5020021"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Pedersen, K.T., Pepke, L., St\u00e6rmose, T., Papaioannou, M., Choudhary, G., and Dragoni, N. (2025). Deepfake-Driven Social Engineering: Threats, Detection Techniques, and Defensive Strategies in Corporate Environments. J. Cybersecur. Priv., 5.","DOI":"10.3390\/jcp5020018"},{"key":"ref_20","first-page":"20","article-title":"The AI-Based Cyber Threat Landscape: A Survey","volume":"53","author":"Kaloudi","year":"2020","journal-title":"ACM Comput. Surv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Kwon, H., and Pak, W. (2024). Text-Based Prompt Injection Attack Using Mathematical Functions in Modern Large Language Models. Electronics, 13.","DOI":"10.3390\/electronics13245008"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Jabbar, H., and Al-Janabi, S. (2025). AI-Driven Phishing Detection: Enhancing Cybersecurity with Reinforcement Learning. J. Cybersecur. Priv., 5.","DOI":"10.3390\/jcp5020026"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Alali, A., and Theodorakopoulos, G. (2025). Partial Fake Speech Attacks in the Real World Using Deepfake Audio. J. Cybersecur. Priv., 5.","DOI":"10.3390\/jcp5010006"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3487890","article-title":"Artificial Intelligence Security: Threats and Countermeasures","volume":"55","author":"Hu","year":"2021","journal-title":"ACM Comput. Surv."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1893","DOI":"10.1109\/JBHI.2014.2344095","article-title":"Systematic Poisoning Attacks on and Defenses for Machine Learning in Healthcare","volume":"19","author":"Raghunathan","year":"2015","journal-title":"IEEE J. Biomed. Health Inform."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1109\/TMSCS.2015.2494021","article-title":"Energy-Efficient Long-term Continuous Personal Health Monitoring","volume":"1","author":"Raghunathan","year":"2015","journal-title":"IEEE Trans. Multi-Scale Comput. Syst."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Gonz\u00e1lez-Granadillo, G., Gonz\u00e1lez-Zarzosa, S., and Diaz, R. (2021). Security Information and Event Management (SIEM): Analysis, Trends, and Usage in Critical Infrastructures. Sensors, 21.","DOI":"10.3390\/s21144759"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Wu, Y., Zou, B., and Cao, Y. (2024). Current Status and Challenges and Future Trends of Deep Learning-Based Intrusion Detection Models. J. Imaging, 10.","DOI":"10.3390\/jimaging10100254"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Balogh, \u0160., Mlyn\u010dek, M., Vra\u0148\u00e1k, O., and Zajac, P. (2024). Using Generative AI Models to Support Cybersecurity Analysts. Electronics, 13.","DOI":"10.3390\/electronics13234718"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Bhusal, D., Alam, M.T., Nguyen, L., Mahara, A., Lightcap, Z., Frazier, R., Fieblinger, R., To-rales, G.L., Blakely, B.A., and Rastogi, N. (2024, January 9\u201313). SECURE: Benchmarking Large Language Models for Cybersecurity. Proceedings of the 2024 Annual Computer Security Applications Conference (ACSAC), Honolulu, HI, USA.","DOI":"10.1109\/ACSAC63791.2024.00019"},{"key":"ref_31","unstructured":"Liu, Z., Shi, J., and Buford, J.F. (2024, January 20\u201327). CyberBench: A Multi-Task Benchmark for Evaluating Large Language Models in Cybersecurity. Proceedings of the AAAI-24 Workshop on Artificial Intelligence for Cyber Security (AICS), Vancouver, BC, Canada."},{"key":"ref_32","unstructured":"Chang, Y., Wang, X., Wang, J., Wu, Y., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., and Ye, W. (2023). A Survey on Evaluation of Large Language Models. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Kasri, W., Himeur, Y., Alkhazaleh, H.A., Tarapiah, S., Atalla, S., Mansoor, W., and Al-Ahmad, H. (2025). From Vulnerability to Defense: The Role of Large Language Models in Enhancing Cybersecurity. Computation, 13.","DOI":"10.3390\/computation13020030"},{"key":"ref_34","unstructured":"Akhtar, S., Khan, S., and Parkinson, S. (2025). LLM-based event log analysis techniques: A survey. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Sayegh, H.R., Dong, W., and Al-madani, A.M. (2024). Enhanced Intrusion Detection with LSTM-Based Model, Feature Selection, and SMOTE for Imbalanced Data. Appl. Sci., 14.","DOI":"10.3390\/app14020479"},{"key":"ref_36","unstructured":"Joshi, B., Ren, X., Swayamdipta, S., Koncel-Kedziorski, R., and Paek, T. (2025). Improving LLM Personas via Rationalization with Psychological Scaffolds. arXiv."},{"key":"ref_37","first-page":"100948","article-title":"Foundation models and intelligent decision-making: Progress, challenges, and perspectives, The Innovation","volume":"6","author":"Huang","year":"2025","journal-title":"Sci. Direct"},{"key":"ref_38","unstructured":"Ollama (2025, June 17). Ollama Run Documentation. Retrieved from the Official Ollama GitHub Repository. Available online: https:\/\/github.com\/ollama\/ollama\/tree\/main\/docs."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Chen, T., and Guestrin, C. (2016, January 13\u201317). XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA.","DOI":"10.1145\/2939672.2939785"},{"key":"ref_40","unstructured":"Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T.-Y. (2017, January 4\u20139). LightGBM: A highly efficient gradient boosting decision tree. Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS\u201917), Long Beach, CA, USA."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/A:1010933404324","article-title":"Random Forests","volume":"45","author":"Breiman","year":"2001","journal-title":"Mach. Learn."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Liu, F.T., Ting, K.M., and Zhou, Z.H. (2008, January 15\u201319). Isolation Forest. Proceedings of the 2008 Eighth IEEE International Conference on Data Mining, Pisa, Italy.","DOI":"10.1109\/ICDM.2008.17"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Joloudari, J.H., Marefat, A., Nematollahi, M.A., Oyelere, S.S., and Hussain, S. (2023). Effective Class-Imbalance Learning Based on SMOTE and Convolutional Neural Networks. Appl. Sci., 13.","DOI":"10.3390\/app13064006"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"193473","DOI":"10.1109\/ACCESS.2024.3520159","article-title":"Influence-Balanced XGBoost: Improving XGBoost for Imbalanced Data Using Influence Functions","volume":"12","author":"Sutou","year":"2024","journal-title":"IEEE Access"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Patwardhan, N., Marrone, S., and Sansone, C. (2023). Transformers in the Real World: A Survey on NLP Applications. Information, 14.","DOI":"10.3390\/info14040242"},{"key":"ref_46","unstructured":"Qwen Team (2025, June 16). Qwen2.5 Model Collection. Hugging Face., Available online: https:\/\/huggingface.co\/collections\/Qwen\/qwen25-66e81a666513e518adb90d9e."},{"key":"ref_47","unstructured":"Google (2025, June 16). Gemma 3 Release Model Collection. Hugging Face., Available online: https:\/\/huggingface.co\/collections\/google\/gemma-3-release-67c6c6f89c4f76621268bb6d."},{"key":"ref_48","unstructured":"Meta Llama (2025, June 16). Llama 3.2 Model Collection. Hugging Face., Available online: https:\/\/huggingface.co\/collections\/meta-llama\/llama-32-66f448ffc8c32f949b04c8cf."},{"key":"ref_49","unstructured":"Qwen Team (2025, June 16). Qwen\/Qwen3-8B Model Card. Hugging Face., Available online: https:\/\/huggingface.co\/Qwen\/Qwen3-8B."},{"key":"ref_50","unstructured":"Microsoft (2025, June 16). Microsoft\/Phi-3.5-Mini-Instruct Model Card. Hugging Face. Available online: https:\/\/huggingface.co\/microsoft\/Phi-3.5-mini-instruct."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Zhang, W., and Zhang, J. (2025). Hallucination Mitigation for Retrieval-Augmented Large Language Models: A Review. Mathematics, 13.","DOI":"10.3390\/math13050856"},{"key":"ref_52","unstructured":"Kritz, J., Robinson, V., Vacareanu, R., Varjavand, B., Choi, M., Gogov, B., Team, S.R., Yue, S., Primack, W.E., and Wang, Z. (2025). Jailbreaking to Jailbreak. arXiv."},{"key":"ref_53","first-page":"9459","article-title":"Retrieval-augmented generation for knowledge-intensive nlp tasks","volume":"33","author":"Lewis","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Xu, K., Zhang, K., Li, J., Huang, W., and Wang, Y. (2025). CRP-RAG: A Retrieval-Augmented Generation Framework for Supporting Complex Logical Reasoning and Knowledge Planning. Electronics, 14.","DOI":"10.20944\/preprints202411.1648.v1"},{"key":"ref_55","unstructured":"Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L.A., Welbl, J., and Clark, A. (December, January 28). Training compute-optimal large language models. Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS \u201922), New Orleans, LA, USA."},{"key":"ref_56","unstructured":"Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., and Kumar, A. (2023, January 10\u201316). Holistic Evaluation of Language Models. Proceedings of the Thirty-seventh Conference on Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_57","unstructured":"Shen, Y., Chen, Z., Zhao, W., Zhang, K., and Yang, M. (2024). Security and Privacy Challenges of Large Language Models: A Survey. arXiv."},{"key":"ref_58","unstructured":"IBM (2025, June 17). Cost of a Data Breach Report 2024. Ponemon Institute., Available online: https:\/\/www.ibm.com\/it-it\/reports\/data-breach."}],"container-title":["Journal of Cybersecurity and Privacy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2624-800X\/5\/3\/55\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:27:28Z","timestamp":1760034448000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2624-800X\/5\/3\/55"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,10]]},"references-count":58,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2025,9]]}},"alternative-id":["jcp5030055"],"URL":"https:\/\/doi.org\/10.3390\/jcp5030055","relation":{},"ISSN":["2624-800X"],"issn-type":[{"value":"2624-800X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,10]]}}}