{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T10:53:05Z","timestamp":1781952785522,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":33,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,6,22]],"date-time":"2026-06-22T00:00:00Z","timestamp":1782086400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"Helmholtz Association (HGF)","award":["46.23.02"],"award-info":[{"award-number":["46.23.02"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,6,22]]},"DOI":"10.1145\/3765611.3815141","type":"proceedings-article","created":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T09:39:33Z","timestamp":1781948373000},"page":"55-64","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2323-0533","authenticated-orcid":false,"given":"Gustav","family":"Keppler","sequence":"first","affiliation":[{"name":"Karlsruhe Institute of Technology, Karlsruhe, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-4036-9032","authenticated-orcid":false,"given":"Moritz","family":"Gst\u00fcr","sequence":"additional","affiliation":[{"name":"Karlsruhe Institute of Technology, Karlsruhe, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3572-9083","authenticated-orcid":false,"given":"Veit","family":"Hagenmeyer","sequence":"additional","affiliation":[{"name":"Karlsruhe Institute of Technology, Karlsruhe, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,22]]},"reference":[{"key":"e_1_3_3_1_2_2","doi-asserted-by":"publisher","unstructured":"Chuadhry\u00a0Mujeeb Ahmed. 2025. AttackLLM: LLM-based Attack Pattern Generation for an Industrial Control System. arxiv:https:\/\/arXiv.org\/abs\/2504.04187\u00a0[cs] 10.48550\/arXiv.2504.04187","DOI":"10.48550\/arXiv.2504.04187"},{"key":"e_1_3_3_1_3_2","doi-asserted-by":"publisher","unstructured":"Andrey Anurin Jonathan Ng Kibo Schaffer Jason Schreiber and Esben Kran. 2024. Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities. (2024). 10.48550\/ARXIV.2410.09114","DOI":"10.48550\/ARXIV.2410.09114"},{"key":"e_1_3_3_1_4_2","first-page":"847","volume-title":"33rd USENIX Security Symposium (USENIX Security 24)","author":"Deng Gelei","year":"2024","unstructured":"Gelei Deng, Yi Liu, V\u00edctor Mayoral-Vilches, Peng Liu, Yuekang Li, Yuan Xu, Tianwei Zhang, Yang Liu, Martin Pinzger, and Stefan Rass. 2024. {PentestGPT}: Evaluating and Harnessing Large Language Models for Automated Penetration Testing. In 33rd USENIX Security Symposium (USENIX Security 24). 847\u2013864."},{"key":"e_1_3_3_1_5_2","doi-asserted-by":"publisher","unstructured":"Richard Fang Rohan Bindu Akul Gupta and Daniel Kang. 2024. LLM Agents Can Autonomously Exploit One-day Vulnerabilities. arxiv:https:\/\/arXiv.org\/abs\/2404.08144\u00a0[cs] 10.48550\/arXiv.2404.08144","DOI":"10.48550\/arXiv.2404.08144"},{"key":"e_1_3_3_1_6_2","doi-asserted-by":"publisher","unstructured":"Richard Fang Rohan Bindu Akul Gupta Qiusi Zhan and Daniel Kang. 2024. LLM Agents Can Autonomously Hack Websites. arxiv:https:\/\/arXiv.org\/abs\/2402.06664\u00a0[cs] 10.48550\/arXiv.2402.06664","DOI":"10.48550\/arXiv.2402.06664"},{"key":"e_1_3_3_1_7_2","unstructured":"Richard Fang Rohan Bindu Akul Gupta Qiusi Zhan and Daniel Kang. 2024. Teams of LLM Agents Can Exploit Zero-Day Vulnerabilities. arxiv:https:\/\/arXiv.org\/abs\/2406.01637\u00a0[cs]"},{"key":"e_1_3_3_1_8_2","doi-asserted-by":"publisher","unstructured":"Luca Gioacchini Marco Mellia Idilio Drago Alexander Delsanto Giuseppe Siracusano and Roberto Bifulco. 2024. AutoPenBench: Benchmarking Generative Agents for Penetration Testing. arxiv:https:\/\/arXiv.org\/abs\/2410.03225\u00a0[cs] 10.48550\/arXiv.2410.03225","DOI":"10.48550\/arXiv.2410.03225"},{"key":"e_1_3_3_1_9_2","unstructured":"Andreas Happe and J\u00fcrgen Cito. 2024. Got Root? A Linux Priv-Esc Benchmark. arxiv:https:\/\/arXiv.org\/abs\/2405.02106\u00a0[cs]"},{"key":"e_1_3_3_1_10_2","doi-asserted-by":"publisher","unstructured":"Andreas Happe and J\u00fcrgen Cito. 2025. Benchmarking Practices in LLM-driven Offensive Security: Testbeds Metrics and Experiment Design. arxiv:https:\/\/arXiv.org\/abs\/2504.10112\u00a0[cs] 10.48550\/arXiv.2504.10112","DOI":"10.48550\/arXiv.2504.10112"},{"key":"e_1_3_3_1_11_2","doi-asserted-by":"publisher","unstructured":"Andreas Happe and J\u00fcrgen Cito. 2025. On the Surprising Efficacy of LLMs for Penetration-Testing. arxiv:https:\/\/arXiv.org\/abs\/2507.00829\u00a0[cs] 10.48550\/arXiv.2507.00829","DOI":"10.48550\/arXiv.2507.00829"},{"key":"e_1_3_3_1_12_2","unstructured":"Andreas Happe Aaron Kaplan and J\u00fcrgen Cito. 2024. LLMs as Hackers: Autonomous Linux Privilege Escalation Attacks. arxiv:https:\/\/arXiv.org\/abs\/2310.11409\u00a0[cs]"},{"key":"e_1_3_3_1_13_2","doi-asserted-by":"publisher","unstructured":"Nourhan Ibrahim and Rasha Kashef. 2025. Exploring the Emerging Role of Large Language Models in Smart Grid Cybersecurity: A Survey of Attacks Detection Mechanisms and Mitigation Strategies. Frontiers in Energy Research 13 (March 2025) 1531655. 10.3389\/fenrg.2025.1531655","DOI":"10.3389\/fenrg.2025.1531655"},{"key":"e_1_3_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPS65515.2025.11087903"},{"key":"e_1_3_3_1_15_2","doi-asserted-by":"publisher","unstructured":"Michael Kouremetis Marissa Dotter Alex Byrne Dan Martin Ethan Michalak Gianpaolo Russo Michael Threet and Guido Zarrella. 2025. OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities. arxiv:https:\/\/arXiv.org\/abs\/2502.15797\u00a0[cs] 10.48550\/arXiv.2502.15797","DOI":"10.48550\/arXiv.2502.15797"},{"key":"e_1_3_3_1_16_2","doi-asserted-by":"publisher","unstructured":"Stela Kucek and Maria Leitner. 2020. An Empirical Survey of Functions and Configurations of Open-Source Capture the Flag (CTF) Environments. Journal of Network and Computer Applications 151 (Feb. 2020) 102470. 10.1016\/j.jnca.2019.102470","DOI":"10.1016\/j.jnca.2019.102470"},{"key":"e_1_3_3_1_17_2","doi-asserted-by":"publisher","unstructured":"V\u00edctor Mayoral-Vilches Andreas Makris and Kevin Finisterre. 2025. Cybersecurity AI: Humanoid Robots as Attack Vectors. arxiv:https:\/\/arXiv.org\/abs\/2509.14139\u00a0[cs] 10.48550\/arXiv.2509.14139","DOI":"10.48550\/arXiv.2509.14139"},{"key":"e_1_3_3_1_18_2","volume-title":"MITRE ATT&CK for Industrial Control Systems","author":"ATT&CK MITRE","year":"2020","unstructured":"MITRE ATT&CK. 2020. MITRE ATT&CK for Industrial Control Systems. The MITRE Corporation. https:\/\/attack.mitre.org\/matrices\/ics\/"},{"key":"e_1_3_3_1_19_2","doi-asserted-by":"publisher","unstructured":"Lajos Muzsai David Imolai and Andr\u00e1s Luk\u00e1cs. 2024. HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing. arxiv:https:\/\/arXiv.org\/abs\/2412.01778\u00a0[cs] 10.48550\/arXiv.2412.01778","DOI":"10.48550\/arXiv.2412.01778"},{"key":"e_1_3_3_1_20_2","unstructured":"Sho Nakatani. 2025. RapidPen: Fully Automated IP-to-Shell Penetration Testing with LLM-based Agents. https:\/\/arxiv.org\/abs\/2502.16730v1."},{"key":"e_1_3_3_1_21_2","doi-asserted-by":"publisher","unstructured":"Mary Phuong Matthew Aitchison Elliot Catt Sarah Cogan Alexandre Kaskasoli Victoria Krakovna David Lindner Matthew Rahtz Yannis Assael Sarah Hodkinson Heidi Howard Tom Lieberum Ramana Kumar Maria\u00a0Abi Raad Albert Webson Lewis Ho Sharon Lin Sebastian Farquhar Marcus Hutter Gregoire Deletang Anian Ruoss Seliem El-Sayed Sasha Brown Anca Dragan Rohin Shah Allan Dafoe and Toby Shevlane. 2024. Evaluating Frontier Models for Dangerous Capabilities. arxiv:https:\/\/arXiv.org\/abs\/2403.13793\u00a0[cs] 10.48550\/arXiv.2403.13793","DOI":"10.48550\/arXiv.2403.13793"},{"key":"e_1_3_3_1_22_2","doi-asserted-by":"publisher","unstructured":"Mikel Rodriguez Raluca\u00a0Ada Popa Four Flynn Lihao Liang Allan Dafoe and Anna Wang. 2025. A Framework for Evaluating Emerging Cyberattack Capabilities of AI. arxiv:https:\/\/arXiv.org\/abs\/2503.11917\u00a0[cs] 10.48550\/arXiv.2503.11917","DOI":"10.48550\/arXiv.2503.11917"},{"key":"e_1_3_3_1_23_2","doi-asserted-by":"publisher","unstructured":"Mar\u00eda Sanz-G\u00f3mez V\u00edctor Mayoral-Vilches Francesco Balassone Luis\u00a0Javier Navarrete-Lozano Crist\u00f3bal R. J.\u00a0Veas Chavez and Maite del\u00a0Mundo de Torres. 2025. Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents. arxiv:https:\/\/arXiv.org\/abs\/2510.24317\u00a0[cs] 10.48550\/arXiv.2510.24317","DOI":"10.48550\/arXiv.2510.24317"},{"key":"e_1_3_3_1_24_2","doi-asserted-by":"publisher","unstructured":"Minghao Shao Boyuan Chen Sofija Jancheska Brendan Dolan-Gavitt Siddharth Garg Ramesh Karri and Muhammad Shafique. 2024. An Empirical Evaluation of LLMs for Solving Offensive Security Challenges. (2024). 10.48550\/ARXIV.2402.11814","DOI":"10.48550\/ARXIV.2402.11814"},{"key":"e_1_3_3_1_25_2","doi-asserted-by":"publisher","unstructured":"Minghao Shao Sofija Jancheska Meet Udeshi Brendan Dolan-Gavitt Haoran Xi Kimberly Milner Boyuan Chen Max Yin Siddharth Garg Prashanth Krishnamurthy Farshad Khorrami Ramesh Karri and Muhammad Shafique. 2024. NYU CTF Dataset: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security. (2024). 10.48550\/ARXIV.2406.05590","DOI":"10.48550\/ARXIV.2406.05590"},{"key":"e_1_3_3_1_26_2","doi-asserted-by":"publisher","unstructured":"Brian Singer Keane Lucas Lakshmi Adiga Meghna Jain Lujo Bauer and Vyas Sekar. 2025. On the Feasibility of Using LLMs to Execute Multistage Network Attacks. arxiv:https:\/\/arXiv.org\/abs\/2501.16466\u00a0[cs] 10.48550\/arXiv.2501.16466","DOI":"10.48550\/arXiv.2501.16466"},{"key":"e_1_3_3_1_27_2","doi-asserted-by":"publisher","unstructured":"Meet Udeshi Minghao Shao Haoran Xi Nanda Rani Kimberly Milner Venkata Sai\u00a0Charan Putrevu Brendan Dolan-Gavitt Sandeep\u00a0Kumar Shukla Prashanth Krishnamurthy Farshad Khorrami Ramesh Karri and Muhammad Shafique. 2025. D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security. arxiv:https:\/\/arXiv.org\/abs\/2502.10931\u00a0[cs] 10.48550\/arXiv.2502.10931","DOI":"10.48550\/arXiv.2502.10931"},{"key":"e_1_3_3_1_28_2","doi-asserted-by":"publisher","unstructured":"Christoforos Vasilatos Dunia\u00a0J. Mahboobeh Hithem Lamri Manaar Alam and Michail Maniatakos. 2025. LLMPot: Dynamically Configured LLM-based Honeypot for Industrial Protocol and Physical Process Emulation. arxiv:https:\/\/arXiv.org\/abs\/2405.05999\u00a0[cs] 10.48550\/arXiv.2405.05999","DOI":"10.48550\/arXiv.2405.05999"},{"key":"e_1_3_3_1_29_2","doi-asserted-by":"publisher","unstructured":"Shengye Wan Cyrus Nikolaidis Daniel Song David Molnar James Crnkovich Jayson Grace Manish Bhatt Sahana Chennabasappa Spencer Whitman Stephanie Ding Vlad Ionescu Yue Li and Joshua Saxe. 2024. CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models. arxiv:https:\/\/arXiv.org\/abs\/2408.01605\u00a0[cs] 10.48550\/arXiv.2408.01605","DOI":"10.48550\/arXiv.2408.01605"},{"key":"e_1_3_3_1_30_2","doi-asserted-by":"publisher","unstructured":"Zhun Wang Tianneng Shi Jingxuan He Matthew Cai Jialin Zhang and Dawn Song. 2025. CyberGym: Evaluating AI Agents\u2019 Cybersecurity Capabilities with Real-World Vulnerabilities at Scale. arxiv:https:\/\/arXiv.org\/abs\/2506.02548\u00a0[cs] 10.48550\/arXiv.2506.02548","DOI":"10.48550\/arXiv.2506.02548"},{"key":"e_1_3_3_1_31_2","doi-asserted-by":"publisher","unstructured":"Jiacen Xu Jack\u00a0W. Stokes Geoff McDonald Xuesong Bai David Marshall Siyue Wang Adith Swaminathan and Zhou Li. 2024. AutoAttacker: A Large Language Model Guided System to Implement Automatic Cyber-attacks. (2024). 10.48550\/ARXIV.2403.01038","DOI":"10.48550\/ARXIV.2403.01038"},{"key":"e_1_3_3_1_32_2","doi-asserted-by":"publisher","unstructured":"Aydin Zaboli Seong\u00a0Lok Choi Tai-Jin Song and Junho Hong. 2024. ChatGPT and Other Large Language Models for Cybersecurity of Smart Grid Applications. arxiv:https:\/\/arXiv.org\/abs\/2311.05462\u00a0[cs] 10.48550\/arXiv.2311.05462","DOI":"10.48550\/arXiv.2311.05462"},{"key":"e_1_3_3_1_33_2","doi-asserted-by":"publisher","unstructured":"Andy\u00a0K. Zhang Neil Perry Riya Dulepet Joey Ji Celeste Menders Justin\u00a0W. Lin Eliot Jones Gashon Hussein Samantha Liu Donovan Jasper Pura Peetathawatchai Ari Glenn Vikram Sivashankar Daniel Zamoshchin Leo Glikbarg Derek Askaryar Mike Yang Teddy Zhang Rishi Alluri Nathan Tran Rinnara Sangpisit Polycarpos Yiorkadjis Kenny Osele Gautham Raghupathi Dan Boneh Daniel\u00a0E. Ho and Percy Liang. 2025. Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models. arxiv:https:\/\/arXiv.org\/abs\/2408.08926\u00a0[cs] 10.48550\/arXiv.2408.08926","DOI":"10.48550\/arXiv.2408.08926"},{"key":"e_1_3_3_1_34_2","doi-asserted-by":"publisher","unstructured":"Zhenyong Zhang Mengxiang Liu Mingyang Sun Ruilong Deng Peng Cheng Dusit Niyato Mo-Yuen Chow and Jiming Chen. 2024. Vulnerability of Machine Learning Approaches Applied in IoT-Based Smart Grid: A Review. IEEE Internet of Things Journal 11 11 (June 2024) 18951\u201318975. 10.1109\/JIOT.2024.3349381","DOI":"10.1109\/JIOT.2024.3349381"}],"event":{"name":"ACM Sustainability Week '26: ACM Sustainability Week 2026","location":"Banff , Alberta , Canada","acronym":"ACM Sustainability Week '26","sponsor":["SIGENERGY ACM Special Interest Group on Energy Systems and Informatics"]},"container-title":["Proceedings of the 2026 ACM Sustainability Week"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3765611.3815141","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T10:46:27Z","timestamp":1781952387000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3765611.3815141"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,22]]},"references-count":33,"alternative-id":["10.1145\/3765611.3815141","10.1145\/3765611"],"URL":"https:\/\/doi.org\/10.1145\/3765611.3815141","relation":{},"subject":[],"published":{"date-parts":[[2026,6,22]]},"assertion":[{"value":"2026-06-22","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}