{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T14:28:18Z","timestamp":1784644098810,"version":"3.55.0"},"reference-count":144,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,1,31]]},"abstract":"<jats:p>In the ongoing network cybersecurity arms race, the defenders face a significant disadvantage as they must detect and counteract every attack. Conversely, the attacker only needs to succeed once to achieve their goal. To balance the odds, Autonomous Cyber Network Defence (ACND) employs autonomous agents for proactive and intelligent cyber-attack response. This article surveys the state-of-the-art of Autonomous Blue and Red Teaming agents, as well as cyber operations environments. We begin by presenting a detailed set of criteria for ACND algorithms and systems that evaluate the preparedness of integrating autonomous agents into real-world networked environments. Our analysis identifies critical research gaps and challenges within the ACND landscape, including issues of autonomous agent explainability, continuous learning capability under evolving threats, and the development of realistic agent training environments. Based on these insights, we discuss promising research directions and open challenges that need to be addressed for the deployment of ACND agents in real-world networks.<\/jats:p>","DOI":"10.1145\/3729213","type":"journal-article","created":{"date-parts":[[2025,5,24]],"date-time":"2025-05-24T07:29:52Z","timestamp":1748071792000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Towards the Deployment of Realistic Autonomous Cyber Network Defence: A Systematic Review"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5603-6868","authenticated-orcid":false,"given":"Sanyam","family":"Vyas","sequence":"first","affiliation":[{"name":"School of Computer Science and Informatics, Cardiff University","place":["Cardiff, United Kingdom of Great Britain and Northern Ireland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2667-5906","authenticated-orcid":false,"given":"Vasilios","family":"Mavroudis","sequence":"additional","affiliation":[{"name":"AICD Research Centre, The Alan Turing Institute","place":["London, United Kingdom of Great Britain and Northern Ireland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0396-633X","authenticated-orcid":false,"given":"Pete","family":"Burnap","sequence":"additional","affiliation":[{"name":"School of Computer Science & Informatics, Cardiff University School of Computer Science and Informatics","place":["Cardiff, United Kingdom of Great Britain and Northern Ireland"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,8,30]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"2022. Cyber Operations Research Gym. Retrieved from https:\/\/github.com\/cage-challenge\/CybORG. Created by Maxwell Standen David Bowman Son Hoang Toby Richer Martin Lucas Richard Van Tassel Phillip Vu Mitchell Kiely KC C. Natalie Konschnik Joshua Collyer."},{"key":"e_1_3_3_3_2","unstructured":"2023. Cyber Agents for Security Testing and Learning Environments. Retrieved from https:\/\/sam.gov\/opp\/9c4593776a9b44e98b9bc734a3e16976\/view#description. Created by Defense Advanced Research Projects Agency."},{"key":"e_1_3_3_4_2","volume-title":"Proceedings of the NeurIPS 2023 Workshop on Backdoors in Deep Learning-The Good, the Bad, and the Ugly","author":"Acharya Manoj","year":"2023","unstructured":"Manoj Acharya, Weichao Zhou, Anirban Roy, Xiao Lin, Wenchao Li, and Susmit Jha. 2023. Universal Trojan signatures in reinforcement learning. In Proceedings of the NeurIPS 2023 Workshop on Backdoors in Deep Learning-The Good, the Bad, and the Ugly."},{"key":"e_1_3_3_5_2","unstructured":"Queensland Defence Science Alliance. 2022. Artificial Intelligence for Decision Making Initiative (2022). Retrieved from https:\/\/queenslanddefencesciencealliance.com.au\/federal-and-state-defence-funding-opportunities-2\/artificial-intelligence-for-decision-making-initiative-round-2022\/"},{"key":"e_1_3_3_6_2","doi-asserted-by":"crossref","unstructured":"Prithviraj Ammanabrolu and Mark O. Riedl. 2018. Playing text-adventure games with graph-based deep reinforcement learning. arXiv preprint arXiv:1812.01628 (2018).","DOI":"10.18653\/v1\/N19-1358"},{"key":"e_1_3_3_7_2","unstructured":"Alex Andrew Sam Spillard Joshua Collyer and Neil Dhir. 2022. Developing optimal causal cyber-defence agents via cyber security simulation. arXiv preprint arXiv:2207.12355."},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3560830.3563732"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNSM.2020.3031843"},{"key":"e_1_3_3_10_2","unstructured":"Chace Ashcraft and Kiran Karra. 2021. Poisoning deep reinforcement learning agents with in-distribution triggers. arXiv preprint arXiv:2106.07798 (2021)."},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3605764.3623916"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/1553374.1553380"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/BLISS.2008.17"},{"key":"e_1_3_3_14_2","unstructured":"Christopher Berner Greg Brockman Brooke Chan Vicki Cheung Przemys\u0142aw D\u0119biak Christy Dennison David Farhi Quirin Fischer Shariq Hashme Chris Hesse Rafal J\u00f3zefowicz Scott Gray Catherine Olsson Jakub Pachocki Michael Petrov Henrique P. d.O. Pinto Jonathan Raiman Tim Salimans Jeremy Schlatter Jonas Schneider Szymon Sidor Ilya Sutskever Jie Tang Filip Wolski and Susan Zhang. 2019. Dota 2 with large scale deep reinforcement learning. arXiv preprint arXiv:1912.06680 (2019)."},{"key":"e_1_3_3_15_2","first-page":"14704","article-title":"Provable defense against backdoor policies in reinforcement learning","volume":"35","author":"Bharti Shubham","year":"2022","unstructured":"Shubham Bharti, Xuezhou Zhang, Adish Singla, and Jerry Zhu. 2022. Provable defense against backdoor policies in reinforcement learning. Advances in Neural Information Processing Systems 35 (2022), 14704\u201314714.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_16_2","article-title":"A model-based, decision-theoretic perspective on automated cyber response","author":"Booker Lashon","year":"2022","unstructured":"Lashon Booker and Scott Musman. 2022. A model-based, decision-theoretic perspective on automated cyber response. In Proceedings of the International Conference on Autonomous Intelligent Cyber-defence agents.","journal-title":"International Conference on Autonomous Intelligent Cyber-defence agents"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/MILCOM.2016.7795375"},{"key":"e_1_3_3_18_2","unstructured":"Miles Brundage Shahar Avin Jasmine Wang Haydn Belfield Gretchen Krueger Gillian Hadfield Heidy Khlaaf Jingying Yang Helen Toner Ruth Fong Tegan Maharaj Pang Wei Koh Sara Hooker Jade Leung Andrew Trask Emma Bluemke Jonathan Lebensold Cullen O\u2019Keefe Mark Koren Th\u00e9o Ryffel JB Rubinovitz Tamay Besiroglu Federica Carugati Jack Clark Peter Eckersley Sarah de Haas Maritza Johnson Ben Laurie Alex Ingerman Igor Krawczuk Amanda Askell Rosario Cammarota Andrew Lohn David Krueger Charlotte Stix Peter Henderson Logan Graham Carina Prunkl Bianca Martin Elizabeth Seger Noa Zilberman Se\u00e1n \u00d3 h\u00c9igeartaigh Frens Kroeger Girish Sastry Rebecca Kagan Adrian Weller Brian Tse Elizabeth Barnes Allan Dafoe Paul Scharre Ariel Herbert-Voss Martijn Rasser Shagun Sodhani Carrick Flynn Thomas Krendl Gilbert Lisa Dyer Saif Khan Yoshua Bengio and Markus Anderljung. 2020. Toward trustworthy AI development: Mechanisms for supporting verifiable claims. arXiv preprint arXiv:2004.07213."},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/BigData52589.2021.9671918"},{"key":"e_1_3_3_20_2","article-title":"Robust Artificial Intelligence for Active Cyber Defence","author":"Burke A.","year":"2017","unstructured":"A. Burke. 2017 [Online]. Robust Artificial Intelligence for Active Cyber Defence. Alan Turing Insitute. Retrieved from https:\/\/www.turing.ac.uk\/sites\/default\/files\/2020-08\/public_ai_acd_techreport_final.pdf","journal-title":"Alan Turing Insitute"},{"key":"e_1_3_3_21_2","unstructured":"CAGE. 2021. CAGE Challenge 1. In Proceedings of the IJCAI-21 1st International Workshop on Adaptive Cyber Defense. arXiv. Available at https:\/\/arxiv.org\/abs\/placeholder"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1117\/12.2559319"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/MILCOM.2016.7795481"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","unstructured":"Xinzhong Chai Yasen Wang Chuanxu Yan Yuan Zhao Wenlong Chen and Xiaolei Wang. 2020. DQ-MOTAG: Deep reinforcement learning-based moving target defense against DDoS attacks. In Proceedings of the 2020 IEEE 5th International Conference on Data Science in Cyberspace (DSC). 375\u2013379. DOI:10.1109\/DSC50466.2020.00065","DOI":"10.1109\/DSC50466.2020.00065"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3068768"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/1276958.1277345"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSN50589.2020.00086"},{"key":"e_1_3_3_28_2","unstructured":"Ankur Chowdhary Dijiang Huang Abdulhakim Sabur Neha Vadnere Myong Kang and Bruce Montrose. 2021. SDN-based moving target defense using multi-agent reinforcement learning. In Proceedings of the first International Conference on Autonomous Intelligent Cyber Defense Agents (AICA 2021). Paris France. 15\u201316."},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1177\/1548512918795061"},{"key":"e_1_3_3_30_2","unstructured":"Josh Collyer Alex Andrew and Duncan Hodges. 2022. ACD-G: Enhancing autonomous cyber defense agent generalization through graph embedded network representation. International Conference on Machine Learning."},{"key":"e_1_3_3_31_2","volume-title":"Cybersecurity Workforce Gap","author":"Crumpler William","year":"2022","unstructured":"William Crumpler and James A. Lewis. 2022. Cybersecurity Workforce Gap. JSTOR."},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i10.29052"},{"key":"e_1_3_3_33_2","doi-asserted-by":"crossref","unstructured":"Richard Dazeley Peter Vamplew and Francisco Cruz. 2023. Explainable reinforcement learning for broad-xai: a conceptual framework and survey. Neural Computing and Applications 35 23 (2023) 16893\u201316916.","DOI":"10.1007\/s00521-023-08423-1"},{"key":"e_1_3_3_34_2","unstructured":"National Defence. 2021. Government of Canada. Retrieved from https:\/\/www.canada.ca\/en\/department-national-defence\/programs\/defence-ideas\/element\/innovation-networks\/challenge\/autonomous-systems-defence-security-trust-barriers-adoption.html"},{"key":"e_1_3_3_35_2","doi-asserted-by":"crossref","unstructured":"S. Kate Devitt and Damian Copeland. 2023. Australia\u2019s approach to AI governance in security and defence. In The AI Wave in Defence Innovation. Routledge 217\u2013250.","DOI":"10.4324\/9781003218326-11"},{"key":"e_1_3_3_36_2","unstructured":"Neil Dhir Henrique Hoeltgebaum Niall Adams Mark Briers Anthony Burke and Paul Jones. 2021. Prospective artificial intelligence approaches for active cyber defence. arXiv preprint arXiv:2104.09981."},{"key":"e_1_3_3_37_2","unstructured":"Maxwell Dondo and Natalia Nakhla. 2021. Towards a framework for autonomous defensive cyber operations in a Network Operations Centre. https:\/\/cradpdf.drdc-rddc.gc.ca\/PDFS\/unc382\/p814083_A1b.pdf"},{"key":"e_1_3_3_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-64793-3_4"},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","unstructured":"Thomas C. Eskridge Marco M. Carvalho Evan Stoner Troy Toggweiler and Adrian Granados. 2015. VINE: A cyber emulation environment for MTD experimentation(MTD\u201915). Association for Computing Machinery New York NY USA 43\u201347. DOI:10.1145\/2808475.2808486","DOI":"10.1145\/2808475.2808486"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2908033"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3488932.3527286"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","unstructured":"Making AI Work for Cyber Defense: The Accuracy-Robustness Tradeoff. 2021. DOI:10.51593\/2021CA007","DOI":"10.51593\/2021CA007"},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9781139046855"},{"key":"e_1_3_3_44_2","unstructured":"Angelo Furfaro Antonio Piccolo and Domenico Sacca. 2016. Smallworld: A test and training system for the cybersecurity. European Scientific Journal (2016)."},{"key":"e_1_3_3_45_2","doi-asserted-by":"crossref","unstructured":"Ariel Futoransky Fernando Miranda Jos\u00e9 Orlicki and Carlos Sarraute. 2010. Simulating cyber-attacks for fun and profit. arXiv preprint arXiv:1006.1919 (2010).","DOI":"10.4108\/ICST.SIMUTOOLS2009.5773"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/SSCI50451.2021.9659947"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/1812\/1\/012039"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1155\/2021\/6378218"},{"key":"e_1_3_3_49_2","unstructured":"Maxime Gasse Damien Grasset Guillaume Gaudron and Pierre-Yves Oudeyer. 2021. Causal reinforcement learning using observational and interventional data. arXiv preprint arXiv:2106.14421 (2021)."},{"key":"e_1_3_3_50_2","doi-asserted-by":"crossref","unstructured":"Claire Glanois Paul Weng Matthieu Zimmer Dong Li Tianpei Yang Jianye Hao and Wulong Liu. 2021. A survey on interpretable reinforcement learning. Machine Learning 113 8 (2024) 5847\u20135890.","DOI":"10.1007\/s10994-024-06543-w"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1504\/IJICS.2018.095298"},{"key":"e_1_3_3_52_2","article-title":"TTCP CAGE Challenge 2","author":"Grouo TTCP Cage Working","year":"2022","unstructured":"TTCP Cage Working Grouo. 2022. TTCP CAGE Challenge 2. Retrieved from https:\/\/github.com\/cage-challenge\/cage-challenge-2","journal-title":"R"},{"key":"e_1_3_3_53_2","article-title":"TTCP CAGE Challenge 3","author":"Group TTCP CAGE Working","year":"2022","unstructured":"TTCP CAGE Working Group. 2022. TTCP CAGE Challenge 3. Retrieved from https:\/\/github.com\/cage-challenge\/cage-challenge-3","journal-title":"R"},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00433"},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.23919\/CNSM50824.2020.9269092"},{"key":"e_1_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.23919\/CNSM52442.2021.9615542"},{"key":"e_1_3_3_57_2","unstructured":"Ronan Hamon Henrik Junklewitz Ignacio Sanchez et\u00a0al. 2020. Robustness and explainability of artificial intelligence. Publications Office of the European Union 207 (2020) 2020."},{"key":"e_1_3_3_58_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-30164-8_363"},{"key":"e_1_3_3_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3605764.3623986"},{"key":"e_1_3_3_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/HST47167.2019.9032927"},{"key":"e_1_3_3_61_2","unstructured":"Robert R. Hoffman Shane T. Mueller Gary Klein and Jordan Litman. 2018. Metrics for Explainable AI: Challenges and Prospects. arXiv preprint arXiv:1812.04608 (2018)."},{"key":"e_1_3_3_62_2","unstructured":"Xing Hu Rui Zhang Ke Tang Jiaming Guo Qi Yi Ruizhi Chen Zidong Du Ling Li Qi Guo Yunji Chen et\u00a0al. 2022. Causality-driven hierarchical structure discovery for reinforcement learning. Advances in Neural Information Processing Systems 35 (2022) 20064\u201320076."},{"key":"e_1_3_3_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3305218.3305239"},{"key":"e_1_3_3_64_2","doi-asserted-by":"crossref","unstructured":"Yunhan Huang Linan Huang and Quanyan Zhu. 2021. Reinforcement learning for feedback-enabled cyber resilience. Annual Reviews in Control 53 (2022) 273\u2013295.","DOI":"10.1016\/j.arcontrol.2022.01.001"},{"key":"e_1_3_3_65_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-32430-8_14"},{"key":"e_1_3_3_66_2","doi-asserted-by":"crossref","unstructured":"Jarom\u00edr Janisch Tom\u00e1\u0161 Pevn\u1ef3 and Viliam Lis\u1ef3. 2023. NASimEmu: Network attack simulator & emulator for training agents generalizing to novel scenarios. In European Symposium on Research in Computer Security. Springer 589\u2013608.","DOI":"10.1007\/978-3-031-54129-2_35"},{"key":"e_1_3_3_67_2","doi-asserted-by":"publisher","DOI":"10.32604\/iasc.2021.016240"},{"key":"e_1_3_3_68_2","unstructured":"Jean Kaddour Aengus Lynch Qi Liu Matt J. Kusner and Ricardo Silva. 2022. Causal Machine Learning: A Survey and Open Problems. arXiv preprint arXiv:2206.15475."},{"key":"e_1_3_3_69_2","volume-title":"Guidelines for Performing Systematic Literature Reviews in Software Engineering","author":"Keele Staffs","year":"2007","unstructured":"Staffs Keele2007. Guidelines for Performing Systematic Literature Reviews in Software Engineering. Technical Report. Technical report, ver. 2.3 ebse technical report. ebse."},{"key":"e_1_3_3_70_2","doi-asserted-by":"publisher","DOI":"10.3390\/make3040045"},{"key":"e_1_3_3_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18072.2020.9218663"},{"key":"e_1_3_3_72_2","article-title":"Procedures for performing systematic reviews","volume":"33","author":"Kitchenham Barbara","year":"2004","unstructured":"Barbara Kitchenham. 2004. Procedures for performing systematic reviews. Keele, UK, Keele Univ. 33, 2004 (082004), 1\u201326.","journal-title":"Keele, UK, Keele Univ."},{"key":"e_1_3_3_73_2","doi-asserted-by":"publisher","DOI":"10.4324\/9780367808846-14"},{"key":"e_1_3_3_74_2","unstructured":"Alexander Kott. 2023. Autonomous intelligent cyber defense agent (aica). Springer."},{"key":"e_1_3_3_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474369.3486877"},{"key":"e_1_3_3_76_2","article-title":"Maintainable log datasets for evaluation of intrusion detection systems","author":"Landauer Max","year":"2022","unstructured":"Max Landauer, Florian Skopik, Maximilian Frank, Wolfgang Hotwagner, Markus Wurzenberger, and Andreas Rauber. 2022. Maintainable log datasets for evaluation of intrusion detection systems. IEEE Transactions on Dependable and Secure Computing 20, 4 (2022), 3466\u20133482.","journal-title":"IEEE Transactions on Dependable and Secure Computing"},{"key":"e_1_3_3_77_2","doi-asserted-by":"crossref","unstructured":"Max Landauer Florian Skopik and Markus Wurzenberger. 2024. Introducing a new alert data set for multi-step attack analysis. In Proceedings of the 17th Cyber Security Experimentation and Test Workshop. 41\u201353.","DOI":"10.1145\/3675741.3675748"},{"key":"e_1_3_3_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/TR.2020.3031317"},{"key":"e_1_3_3_79_2","doi-asserted-by":"publisher","DOI":"10.1049\/ise2.12050"},{"key":"e_1_3_3_80_2","unstructured":"Li Li Raed Fayad and Adrian Taylor. 2021. Cygil: A cyber gym for training autonomous agents over emulated network systems. arXiv preprint arXiv:2109.03331 (2021)."},{"key":"e_1_3_3_81_2","unstructured":"Michael Littman. 2009. Algorithms for Sequential Decision Making."},{"key":"e_1_3_3_82_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33012970"},{"key":"e_1_3_3_83_2","doi-asserted-by":"crossref","unstructured":"Prashan Madumal Tim Miller Liz Sonenberg and Frank Vetere. 2020. Explainable reinforcement learning through a causal lens. In Proceedings of the AAAI Conference on Artificial Intelligence 34 (2020) 2493\u20132500.","DOI":"10.1609\/aaai.v34i03.5631"},{"key":"e_1_3_3_84_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cose.2020.102108"},{"key":"e_1_3_3_85_2","doi-asserted-by":"publisher","DOI":"10.1145\/3339252.3339282"},{"key":"e_1_3_3_86_2","unstructured":"Kleanthis Malialis and Daniel Kudenko. 2013. Large-Scale DDoS response using cooperative reinforcement learning."},{"key":"e_1_3_3_87_2","doi-asserted-by":"publisher","DOI":"10.1145\/2808475.2808482"},{"key":"e_1_3_3_88_2","unstructured":"Stephanie Milani Nicholay Topin Manuela Veloso and Fei Fang. 2022. A survey of explainable reinforcement learning. arXiv preprint arXiv:2202.08434 (2022)."},{"key":"e_1_3_3_89_2","doi-asserted-by":"publisher","DOI":"10.1109\/THS.2010.5655108"},{"key":"e_1_3_3_90_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-022-06142-7"},{"key":"e_1_3_3_91_2","unstructured":"Andres Molina-Markham Cory Miniter Becky Powell and Ahmad Ridley. 2021. Network environment design for autonomous cyberdefense. arXiv preprint arXiv:2103.07583."},{"key":"e_1_3_3_92_2","unstructured":"NATO. 2021. Artificial Intelligence and Autonomy in the Military. Retrieved from https:\/\/ccdcoe.org\/uploads\/2021\/12\/Strategies_and_Deployment_A4.pdf"},{"key":"e_1_3_3_93_2","unstructured":"NATO. 2022. Cooperative Cyber Defence Centre of Excellence. Retrieved from https:\/\/ccdcoe.org\/library\/publications\/"},{"key":"e_1_3_3_94_2","doi-asserted-by":"publisher","DOI":"10.1145\/3440749.3442660"},{"key":"e_1_3_3_95_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2021.3121870"},{"key":"e_1_3_3_96_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2018.2885530"},{"key":"e_1_3_3_97_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMM48946.2020.9141989"},{"key":"e_1_3_3_98_2","article-title":"The Cyber Security Learning Environment","author":"Technology KTH Royal Institute of","year":"2023","unstructured":"KTH Royal Institute of Technology and DARPA. 2023. The Cyber Security Learning Environment. Retrieved from https:\/\/github.com\/Limmen\/csle.","journal-title":"R"},{"key":"e_1_3_3_99_2","unstructured":"Cabinet Office. 2022. Government Cyber Security Strategyy. Retrieved from https:\/\/www.gov.uk\/government\/publications\/government-cyber-security-strategy-2022-to-2030"},{"key":"e_1_3_3_100_2","unstructured":"Hamed Okhravi Thomas R. Hobson William W. Streilein George K. Baah Shannon C. Roberts and Sophia Yuditskaya. 2015. Title of the report. Technical Report AD1034028. Defense Technical Information Center. https:\/\/apps.dtic.mil\/sti\/html\/tr\/AD1034028\/index.html. Accessed: 2025-06-03."},{"key":"e_1_3_3_101_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2021.103455"},{"key":"e_1_3_3_102_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459991"},{"key":"e_1_3_3_103_2","doi-asserted-by":"crossref","unstructured":"Deepak Pathak Pulkit Agrawal Alexei A. Efros and Trevor Darrell. 2017. Curiosity-driven Exploration by Self-supervised Prediction. In International Conference on Machine Learning. PMLR 2778\u20132787.","DOI":"10.1109\/CVPRW.2017.70"},{"key":"e_1_3_3_104_2","doi-asserted-by":"crossref","unstructured":"Jeffrey Pawlick Edward Colbert and Quanyan Zhu. 2019. A game-theoretic taxonomy and survey of defensive deception for cybersecurity and privacy. ACM Computing Surveys (CSUR) 52 4 (2019) 1\u201328.","DOI":"10.1145\/3337772"},{"key":"e_1_3_3_105_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Peng Xiangyu","year":"2022","unstructured":"Xiangyu Peng, Mark Riedl, and Prithviraj Ammanabrolu. 2022. Inherently explainable reinforcement learning in natural language. In Proceedings of the Advances in Neural Information Processing Systems, Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (Eds.). Retrieved from https:\/\/openreview.net\/forum?id=DSEP9rCvZln"},{"key":"e_1_3_3_106_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-57321-8_5"},{"key":"e_1_3_3_107_2","first-page":"83","article-title":"Machine learning for cyber defense and attack","volume":"2018","author":"Rege Manjeet","year":"2018","unstructured":"Manjeet Rege and Raymond Blanch K Mbah. 2018. Machine learning for cyber defense and attack. Data Analytics 2018 (2018), 83.","journal-title":"Data Analytics"},{"key":"e_1_3_3_108_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-022-19366-3"},{"key":"e_1_3_3_109_2","unstructured":"Danilo J. Rezende Ivo Danihelka George Papamakarios Nan Rosemary Ke Ray Jiang Theophane Weber Karol Gregor Hamza Merzic Fabio Viola Jane Wang Jovana Mitrovic Frederic Besse Ioannis Antonoglou and Lars Buesing. 2020. Causally Correct Partial Models for Reinforcement Learning. arXiv preprint arXiv:2002.02836 (2020)."},{"key":"e_1_3_3_110_2","doi-asserted-by":"publisher","DOI":"10.1109\/SmartGridComm47815.2020.9302997"},{"key":"e_1_3_3_111_2","doi-asserted-by":"publisher","unstructured":"George Rush Daniel R. Tauritz and Alexander D. Kent. 2015. Coevolutionary agent-based network defense lightweight event system (CANDLES)(GECCO Companion\u201915). Association for Computing Machinery New York NY USA 859\u2013866. DOI:10.1145\/2739482.2768429","DOI":"10.1145\/2739482.2768429"},{"key":"e_1_3_3_112_2","unstructured":"Kevin Schoonover Eric Michalak Sean Harris Adam Gausmann Hannah Reinbolt Daniel Tauritz Chris Rawlings and Aaron Pope. 2018. Galaxy: A network emulation framework for cybersecurity."},{"key":"e_1_3_3_113_2","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal Policy Optimization Algorithms. arXiv preprint arXiv:1707.06347 (2017)."},{"key":"e_1_3_3_114_2","unstructured":"Jonathon Schwartz and Hanna Kurniawati. 2019. Autonomous penetration testing using reinforcement learning. ArXiv abs\/1905.05965 (2019)."},{"key":"e_1_3_3_115_2","unstructured":"Jonathon Schwartz and Hanna Kurniawatti. 2019. NASim: Network Attack Simulator."},{"key":"e_1_3_3_116_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-97532-6_4"},{"key":"e_1_3_3_117_2","volume-title":"Advances in Neural Information Processing Systems","author":"Shafahi Ali","year":"2018","unstructured":"Ali Shafahi, W. Ronny Huang, Mahyar Najibi, Octavian Suciu, Christoph Studer, Tudor Dumitras, and Tom Goldstein. 2018. Poison Frogs! Targeted clean-label poisoning attacks on neural networks. In Advances in Neural Information Processing Systems, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.). Vol. 31. Curran Associates, Inc.Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2018\/file\/22722a343513ed45f14905eb07621686-Paper.pdf"},{"key":"e_1_3_3_118_2","unstructured":"Tianmin Shu Caiming Xiong and Richard Socher. 2017. Hierarchical and interpretable skill acquisition in multi-task reinforcement learning. arXiv preprint arXiv:1712.07294 (2017)."},{"key":"e_1_3_3_119_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCPS54341.2022.00036"},{"key":"e_1_3_3_120_2","volume-title":"Game Theory 101: The Complete Textbook","author":"Spaniel William","year":"2014","unstructured":"William Spaniel. 2014. Game Theory 101: The Complete Textbook. CreateSpace."},{"key":"e_1_3_3_121_2","unstructured":"Ben Spencer and Steve Cooper. 2021. $10 million to build defence\u2019s AI capability and Support Critical Tech for Australia. https:\/\/www.minister.defence.gov.au\/media-releases\/2021-11-18\/10-million-build-defences-ai-capabilityand-support-critical-tech-australia"},{"key":"e_1_3_3_122_2","unstructured":"Maxwell Standen Martin Lucas David Bowman Toby J. Richer Junae Kim and Damian Marriott. 2021. CybORG: A Gym for the Development of Autonomous Cyber Agents. arXiv preprint arXiv:2108.09118 (2021)."},{"key":"e_1_3_3_123_2","doi-asserted-by":"publisher","DOI":"10.1117\/12.2585173"},{"key":"e_1_3_3_124_2","doi-asserted-by":"publisher","DOI":"10.5555\/551283"},{"key":"e_1_3_3_125_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2016.7759424"},{"key":"e_1_3_3_126_2","unstructured":"Microsoft Defender Research Team. 2021. CyberBattleSim. https:\/\/github.com\/microsoft\/cyberbattlesim. Created by Christian Seifert Michael Betser William Blum James Bono Kate Farris Emily Goren Justin Grana Kristian Holsheimer Brandon Marken Joshua Neil Nicole Nichols Jugal Parikh Haoran Wei."},{"key":"e_1_3_3_127_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2021.103627"},{"key":"e_1_3_3_128_2","unstructured":"Khuong Tran Ashlesha Akella Maxwell Standen Junae Kim David Bowman Toby Richer and Chin-Teng Lin. 2021. Deep Hierarchical Reinforcement Agents for Automated Penetration Testing. arXiv preprint arXiv:2109.06449 (2021)."},{"key":"e_1_3_3_129_2","doi-asserted-by":"publisher","DOI":"10.3389\/fpsyg.2020.01049"},{"key":"e_1_3_3_130_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-019-1724-z"},{"key":"e_1_3_3_131_2","unstructured":"Ben Wallace. 2022. Defence Artificial Intelligence Strategy. Retrieved from https:\/\/www.gov.uk\/government\/publications\/defence-artificial-intelligence-strategy\/defence-artificial-intelligence-strategy"},{"key":"e_1_3_3_132_2","unstructured":"Erich Walter Kimberly Ferguson-Walter and Ahmad Ridley. 2021. Incorporating deception into cyberbattlesim for autonomous defense. arXiv preprint arXiv:2108.13980 (2021)."},{"key":"e_1_3_3_133_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2974786"},{"key":"e_1_3_3_134_2","doi-asserted-by":"publisher","DOI":"10.3390\/a15040134"},{"key":"e_1_3_3_135_2","doi-asserted-by":"publisher","DOI":"10.1145\/2601248.2601268"},{"key":"e_1_3_3_136_2","unstructured":"Melody Wolk Andy Applebaum Camron Denver Patrick Dwyer Marina Moskowitz Harold Nguyen Nicole Nichols Nicole Park Paul Rachwalski Frank Rau and Adrian Webster. 2022. Beyond CAGE: Investigating generalization of learned autonomous network defense policies. arXiv preprint arXiv:2211.15557 (2022)."},{"key":"e_1_3_3_137_2","doi-asserted-by":"crossref","unstructured":"Annie Wong Thomas B\u00e4ck Anna V. Kononova and Aske Plaat. 2023. Deep multiagent reinforcement learning: Challenges and directions. Artificial Intelligence Review 56 6 (2023) 5023\u20135056.","DOI":"10.1007\/s10462-022-10299-x"},{"key":"e_1_3_3_138_2","unstructured":"Paul L. Yu. 2023. Multidisciplinary university research initiative: Adversarial and uncertain reasoning for adaptive cyber defense (summary technical report 2013\u20132021). (2023)."},{"key":"e_1_3_3_139_2","doi-asserted-by":"publisher","DOI":"10.1109\/GLOBECOM48099.2022.10000751"},{"key":"e_1_3_3_140_2","unstructured":"Mattia Zago V\u00edctor S\u00e1nchez Manuel P\u00e9rez and Gregorio Martinez Perez. 2017. Tackling cyber threats with automatic decisions and reactions based on machine-learning techniques."},{"key":"e_1_3_3_141_2","unstructured":"Matej Ze\u010devi\u0107 Devendra Singh Dhami Petar Veli\u010dkovi\u0107 and Kristian Kersting. 2021. Relating Graph Neural Networks to Structural Causal Models. arXiv preprint arXiv:2109.04173 (2021)."},{"key":"e_1_3_3_142_2","doi-asserted-by":"publisher","DOI":"10.1109\/SSCI47803.2020.9308468"},{"key":"e_1_3_3_143_2","doi-asserted-by":"publisher","DOI":"10.1631\/FITEE.1800532"},{"key":"e_1_3_3_144_2","series-title":"Proceedings of Machine Learning Research","first-page":"7614","volume-title":"Proceedings of the 36th International Conference on Machine Learning","volume":"97","author":"Zhu Chen","year":"2019","unstructured":"Chen Zhu, W. Ronny Huang, Hengduo Li, Gavin Taylor, Christoph Studer, and Tom Goldstein. 2019. Transferable clean-label poisoning attacks on deep neural nets. In Proceedings of the 36th International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol. 97), Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, 7614\u20137623. Retrieved from https:\/\/proceedings.mlr.press\/v97\/zhu19a.html"},{"key":"e_1_3_3_145_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2009.5270307"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729213","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,30]],"date-time":"2025-08-30T13:37:43Z","timestamp":1756561063000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3729213"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,30]]},"references-count":144,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1,31]]}},"alternative-id":["10.1145\/3729213"],"URL":"https:\/\/doi.org\/10.1145\/3729213","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,30]]},"assertion":[{"value":"2023-02-24","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-30","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}