{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T17:25:49Z","timestamp":1783790749423,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":44,"publisher":"ACM","funder":[{"name":"Lorraine Universit\uc3a9 d?Excellence (LUE)","award":["ANR-15-IDEX-04-LUE"],"award-info":[{"award-number":["ANR-15-IDEX-04-LUE"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,10,13]]},"DOI":"10.1145\/3733799.3762978","type":"proceedings-article","created":{"date-parts":[[2025,12,30]],"date-time":"2025-12-30T11:38:49Z","timestamp":1767094729000},"page":"194-205","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["LLM-CVX: A Benchmarking Framework for Assessing the Offensive Potential of LLMs in Exploiting CVEs"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-7575-3392","authenticated-orcid":false,"given":"Mohamed Amine","family":"El yagouby","sequence":"first","affiliation":[{"name":"CNRS, Inria, LORIA, Universit\u00e9 de Lorraine, Nancy, France and TICLab, International University of Rabat, Rabat, Morocco"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3882-1560","authenticated-orcid":false,"given":"Abdelkader","family":"Lahmadi","sequence":"additional","affiliation":[{"name":"CNRS, Inria, LORIA, Universit\u00e9 de Lorraine, Nancy, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2891-8588","authenticated-orcid":false,"given":"Mehdi","family":"Zakroum","sequence":"additional","affiliation":[{"name":"TICLab, International University of Rabat, Rabat, Morocco"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3181-7967","authenticated-orcid":false,"given":"Olivier","family":"Festor","sequence":"additional","affiliation":[{"name":"CNRS, Inria, LORIA,, Universit\u00e9 de Lorraine, Nancy, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0055-7867","authenticated-orcid":false,"given":"Mounir","family":"Ghogho","sequence":"additional","affiliation":[{"name":"University Mohammed VI Polytechnic, Rabat, Morocco"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,12,30]]},"reference":[{"key":"e_1_3_3_2_2_2","unstructured":"2024. Decoding the Threat Landscape: ChatGPT FraudGPT and WormGPT in Social Engineering Attacks. https:\/\/arxiv.org\/abs\/2310.05595."},{"key":"e_1_3_3_2_3_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia\u00a0Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et\u00a0al. 2023. Gpt-4 technical report. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2303.08774 (2023)."},{"key":"e_1_3_3_2_4_2","unstructured":"Adaptive Security. 2023. ChatGPT & the Surge in AI Phishing: Here\u2019s Why It\u2019s Happening. https:\/\/www.adaptivesecurity.com\/resources\/4151-increase-in-phishing-since-the-launch-of-chatgpt."},{"key":"e_1_3_3_2_5_2","unstructured":"Mohammad Almukaynizi Kevin Borgolte and Vyas Sekar. 2021. Cybersecurity red teaming and the role of controlled adversarial simulations. IEEE Security & Privacy 19 1 (2021) 58\u201366."},{"key":"e_1_3_3_2_6_2","doi-asserted-by":"publisher","unstructured":"Nuno Antunes Rui Neves and Marco Vieira. 2024. Enhancing Cybersecurity Resilience Through Advanced Red-Teaming. Computers & Security 138 (2024) 103123. 10.1016\/j.cose.2024.103123","DOI":"10.1016\/j.cose.2024.103123"},{"key":"e_1_3_3_2_7_2","unstructured":"Manish Bhatt Sahana Chennabasappa Yue Li Cyrus Nikolaidis Daniel Song Shengye Wan Faizan Ahmad Cornelius Aschermann Yaohui Chen Dhaval Kapil et\u00a0al. 2024. Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2404.13161 (2024)."},{"key":"e_1_3_3_2_8_2","unstructured":"Rishi Bommasani Drew\u00a0A Hudson Ehsan Adeli Russ Altman Simran Arora Michael von Arx et\u00a0al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2108.07258 (2021)."},{"key":"e_1_3_3_2_9_2","first-page":"1877","volume-title":"Advances in neural information processing systems","author":"Brown Tom\u00a0B","year":"2020","unstructured":"Tom\u00a0B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et\u00a0al. 2020. Language models are few-shot learners. In Advances in neural information processing systems , Vol.\u00a033. 1877\u20131901."},{"key":"e_1_3_3_2_10_2","volume-title":"Proceedings of the USENIX Summit on Gaming, Games, and Gamification in Security Education (3GSE)","author":"Chapman Peter","year":"2014","unstructured":"Peter Chapman et\u00a0al. 2014. Playing attacker and defender in the cyber war game. In Proceedings of the USENIX Summit on Gaming, Games, and Gamification in Security Education (3GSE)."},{"key":"e_1_3_3_2_11_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de\u00a0Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman et\u00a0al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2107.03374 (2021)."},{"key":"e_1_3_3_2_12_2","unstructured":"Cyware Alerts. 2023. Introducing FraudGPT: The Latest AI Cybercrime Tool in the Dark Web. https:\/\/social.cyware.com\/news\/introducing-fraudgpt-the-latest-ai-cybercrime-tool-in-the-dark-web-1ba40d2f. Accessed 2025-08-25."},{"key":"e_1_3_3_2_13_2","unstructured":"Dark Reading. 2024. \u2019FraudGPT\u2019 Malicious Chatbot Now for Sale on Dark Web. https:\/\/www.darkreading.com\/threat-intelligence\/fraudgpt-malicious-chatbot-for-sale-dark-web."},{"key":"e_1_3_3_2_14_2","unstructured":"Edward Durant Richard Ford and Eugene Spafford. 2011. Using red teams to measure effectiveness of security training. IEEE Security & Privacy 9 6 (2011) 73\u201375."},{"key":"e_1_3_3_2_15_2","unstructured":"Arthur Erzberger. 2023. WormGPT and FraudGPT \u2013 The Rise of Malicious LLMs. https:\/\/www.trustwave.com\/en-us\/resources\/blogs\/spiderlabs-blog\/wormgpt-and-fraudgpt-the-rise-of-malicious-llms\/. Accessed 2025-08-25."},{"key":"e_1_3_3_2_16_2","unstructured":"Richard Fang Rohan Bindu Akul Gupta and Daniel Kang. 2024. Llm agents can autonomously exploit one-day vulnerabilities. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2404.08144 13 (2024) 14."},{"key":"e_1_3_3_2_17_2","unstructured":"Richard Fang Rohan Bindu Akul Gupta Qiusi Zhan and Daniel Kang. 2024. Llm agents can autonomously hack websites. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2402.06664 (2024)."},{"key":"e_1_3_3_2_18_2","unstructured":"Luca Gioacchini Marco Mellia Idilio Drago Alexander Delsanto Giuseppe Siracusano and Roberto Bifulco. 2024. AutoPenBench: Benchmarking Generative Agents for Penetration Testing. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.03225 (2024)."},{"key":"e_1_3_3_2_19_2","unstructured":"Jia He Mukund Rungta David Koleczek Arshdeep Sekhon Franklin\u00a0X Wang and Sadid Hasan. 2024. Does Prompt Formatting Have Any Impact on LLM Performance? arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2411.10541 (2024)."},{"key":"e_1_3_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.5555\/3485754.3485760"},{"key":"e_1_3_3_2_21_2","unstructured":"Imperva. 2024. 2024 Bad Bot Report. https:\/\/www.imperva.com\/resources\/resource-library\/reports\/2024-bad-bot-report\/."},{"key":"e_1_3_3_2_22_2","unstructured":"Cyentia Institute and FIRST. 2024. A Visual Exploration of Exploits in the Wild. https:\/\/www.cyentia.com\/wp-content\/uploads\/2024\/07\/EPSS-Exploration-Of-Exploits.pdf. Accessed: 2025-06-27."},{"key":"e_1_3_3_2_23_2","unstructured":"Cheonsu Jeong. 2024. Fine-tuning and utilization methods of domain-specific llms. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2401.02981 (2024)."},{"key":"e_1_3_3_2_24_2","doi-asserted-by":"crossref","unstructured":"K\u00fcbra\u00a0Nilg\u00fcn Karaca and Ayd\u0131n \u00c7etin. 2025. Systematic Review of Current Approaches and Innovative Solutions for Combating Zero-day Vulnerabilities and Zero-day Attacks. IEEE Access (2025).","DOI":"10.1109\/ACCESS.2025.3577941"},{"key":"e_1_3_3_2_25_2","unstructured":"Takeshi Kojima Shixiang\u00a0Shane Gu Machel Reid Yutaka Matsuo and Yusuke Iwasawa. 2022. Large Language Models are Zero-Shot Reasoners. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2205.11916 (2022)."},{"key":"e_1_3_3_2_26_2","first-page":"9","volume-title":"2nd Workshop on Research with Security Vulnerability Databases, Purdue University, West Lafayette, Indiana","author":"Mann David\u00a0E","year":"1999","unstructured":"David\u00a0E Mann and Steven\u00a0M Christey. 1999. Towards a common enumeration of vulnerabilities. In 2nd Workshop on Research with Security Vulnerability Databases, Purdue University, West Lafayette, Indiana. 9."},{"key":"e_1_3_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/HICSS.2016.677"},{"key":"e_1_3_3_2_28_2","doi-asserted-by":"crossref","unstructured":"Peter Mell Karen Scarfone and Sasha Romanosky. 2006. Common vulnerability scoring system. IEEE Security & Privacy 4 6 (2006) 85\u201389.","DOI":"10.1109\/MSP.2006.145"},{"key":"e_1_3_3_2_29_2","unstructured":"Lajos Muzsai David Imolai and Andr\u00e1s Luk\u00e1cs. 2024. HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2412.01778 (2024)."},{"key":"e_1_3_3_2_30_2","unstructured":"Chris Rohlf. 2024. No LLM Agents Cannot Autonomously Exploit One-day Vulnerabilities. https:\/\/struct.github.io\/auto_agents_1_day.html. Accessed: 2025-06-06."},{"key":"e_1_3_3_2_31_2","unstructured":"Chris Rohlf. 2024. No LLM Agents Cannot Autonomously \u2019Hack\u2019 Websites. https:\/\/struct.github.io\/llm_auto_hax.html. Accessed: 2025-06-06."},{"key":"e_1_3_3_2_32_2","unstructured":"Timo Schick Shibani\u00a0Santurkar Dwivedi-Yu Hinrich Sch\u00fctze Yujia Ma Nathan Scales Yi Tay Rohan Puri and Neil Houlsby. 2023. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2302.04761 (2023)."},{"key":"e_1_3_3_2_33_2","unstructured":"Minghao Shao Boyuan Chen Sofija Jancheska Brendan Dolan-Gavitt Siddharth Garg Ramesh Karri and Muhammad Shafique. 2024. An empirical evaluation of llms for solving offensive security challenges. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2402.11814 (2024)."},{"key":"e_1_3_3_2_34_2","volume-title":"MITRE ATT&CK\u00ae: Design and Philosophy","author":"Strom Blake\u00a0E.","year":"2020","unstructured":"Blake\u00a0E. Strom, Andy Applebaum, Doug\u00a0P. Miller, Kathryn\u00a0C. Nickels, Adam\u00a0G. Pennington, and Cody\u00a0B. Thomas. 2020. MITRE ATT&CK\u00ae: Design and Philosophy. Technical Report. The MITRE Corporation. https:\/\/attack.mitre.org\/docs\/ATTACK_Design_and_Philosophy_March_2020.pdf"},{"key":"e_1_3_3_2_35_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et\u00a0al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2302.13971 (2023)."},{"key":"e_1_3_3_2_36_2","doi-asserted-by":"crossref","unstructured":"Yan Wang Wei Wu Chao Zhang Xinyu Xing Xiaorui Gong and Wei Zou. 2019. From proof-of-concept to exploitable: One step towards automatic exploitability assessment. Cybersecurity 2 (2019) 1\u201325.","DOI":"10.1186\/s42400-019-0028-9"},{"key":"e_1_3_3_2_37_2","unstructured":"Jason Wei Xuezhi Wang Dale Schuurmans Maarten Bosma Fei Xia Ed Chi Quoc\u00a0V Le Denny Zhou et\u00a0al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022) 24824\u201324837."},{"key":"e_1_3_3_2_38_2","unstructured":"Laura Weidinger Joe Mellor Mathias Rauh Christopher Griffin Jonathan Uesato Po-Sen Huang Aslan Glaese Borja Balle Atoosa Kasirzadeh Nicholas Carlini et\u00a0al. 2021. Ethical and social risks of harm from language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2112.04359 (2021)."},{"key":"e_1_3_3_2_39_2","unstructured":"Jiacen Xu Jack\u00a0W Stokes Geoff McDonald Xuesong Bai David Marshall Siyue Wang Adith Swaminathan and Zhou Li. 2024. Autoattacker: A large language model guided system to implement automatic cyber-attacks. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2403.01038 (2024)."},{"key":"e_1_3_3_2_40_2","unstructured":"Zihao Xu Yi Liu Gelei Deng Yuekang Li and Stjepan Picek. 2024. LLM Jailbreak Attack versus Defense Techniques\u2013A Comprehensive Study. arXiv e-prints (2024) arXiv\u20132402."},{"key":"e_1_3_3_2_41_2","unstructured":"Soufian\u00a0El Yadmani Robin The and Olga Gadyatskaya. 2022. Beyond the Surface: Investigating Malicious CVE Proof of Concept Exploits on GitHub. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2210.08374 (2022)."},{"key":"e_1_3_3_2_42_2","volume-title":"Proceedings of the EvalLM 2025 Workshop on Evaluating Large Language Models","author":"Yagouby Mohamed Amine\u00a0El","year":"2025","unstructured":"Mohamed Amine\u00a0El Yagouby, Abdelkader Lahmadi, Mehdi Zakroum, Olivier Festor, and Mounir Ghogho. 2025. Evaluating LLM Efficiency Using Successive Attempts on Binary-Outcome Tasks. In Proceedings of the EvalLM 2025 Workshop on Evaluating Large Language Models. Accepted to EvalLM 2025. https:\/\/evalllm2025.sciencesconf.org."},{"key":"e_1_3_3_2_43_2","volume-title":"International Conference on Learning Representations (ICLR)","author":"Yao Shunyu","year":"2023","unstructured":"Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_3_2_44_2","unstructured":"Andy\u00a0K Zhang Neil Perry Riya Dulepet Joey Ji Celeste Menders Justin\u00a0W Lin Eliot Jones Gashon Hussein Samantha Liu Donovan Jasper et\u00a0al. 2024. Cybench: A framework for evaluating cybersecurity capabilities and risks of language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2408.08926 (2024)."},{"key":"e_1_3_3_2_45_2","unstructured":"Yuxuan Zhu Antony Kellermann Dylan Bowman Philip Li Akul Gupta Adarsh Danda Richard Fang Conner Jensen Eric Ihli Jason Benn et\u00a0al. 2025. CVE-Bench: A Benchmark for AI Agents\u2019 Ability to Exploit Real-World Web Application Vulnerabilities. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2503.17332 (2025)."}],"event":{"name":"AISec '25: Proceedings of the 2025 Workshop on Artificial Intelligence and Security","location":"Taipei , Taiwan","acronym":"AISec '25","sponsor":["SIGSAC ACM Special Interest Group on Security, Audit, and Control"]},"container-title":["Proceedings of the 18th ACM Workshop on Artificial Intelligence and Security"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3733799.3762978","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,30]],"date-time":"2025-12-30T11:51:53Z","timestamp":1767095513000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3733799.3762978"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,13]]},"references-count":44,"alternative-id":["10.1145\/3733799.3762978","10.1145\/3733799"],"URL":"https:\/\/doi.org\/10.1145\/3733799.3762978","relation":{},"subject":[],"published":{"date-parts":[[2025,10,13]]},"assertion":[{"value":"2025-12-30","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}