{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,9,15]],"date-time":"2026-09-15T02:33:15Z","timestamp":1789439595389,"version":"build-2803163510"},"reference-count":23,"publisher":"Wiley","issue":"3","license":[{"start":{"date-parts":[[2025,6,27]],"date-time":"2025-06-27T00:00:00Z","timestamp":1750982400000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Applied AI Letters"],"published-print":{"date-parts":[[2025,10]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>Artificial Intelligence (AI) is set to become an essential tool for defending against machine\u2010speed attacks on increasingly connected cyber networks and systems. It will allow self\u2010defending and self\u2010recovering cyber\u2010defence agents to be developed, which can respond to attacks in a timely manner. But how can these agents be trusted to perform as expected, and how can they be evaluated responsibly and thoroughly? To answer these questions, a Test and Evaluation (T&amp;E) process has been developed to assess cyber\u2010defence agents. The process evaluates the performance, effectiveness, resilience, and generalizability of agents in both low\u2010 and high\u2010fidelity cyber environments. This paper demonstrates the low\u2010fidelity part of the process by performing an example evaluation in the Cyber Operations Research Gym (CybORG) environment on Reinforcement Learning (RL) agents trained as part of Cyber Autonomy Gym for Experimentation (CAGE) Challenge 2. The process makes use of novel Measures of Effectiveness (MoE) metrics, which can be used in combination with performance metrics such as the RL reward. MoE are tailored for cyber defence, allowing a greater understanding of agents' defensive abilities within a cyber environment. Agents are evaluated against multiple conditions that perturb the environment to investigate their robustness to scenarios not seen during training. The results from this evaluation process will help inform decisions around the benefits and risks of integrating autonomous agents into existing or future cyber systems.<\/jats:p>","DOI":"10.1002\/ail2.125","type":"journal-article","created":{"date-parts":[[2025,6,27]],"date-time":"2025-06-27T03:35:06Z","timestamp":1750995306000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Evaluating Reinforcement Learning Agents for Autonomous Cyber Defence"],"prefix":"10.1002","volume":"6","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-4389-6622","authenticated-orcid":false,"given":"Abby","family":"Morris","sequence":"first","affiliation":[{"name":"QinetiQ  Great Malvern UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rachael","family":"Procter","sequence":"additional","affiliation":[{"name":"QinetiQ  Great Malvern UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Caroline","family":"Wallbank","sequence":"additional","affiliation":[{"name":"QinetiQ  Farnborough UK"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2025,6,27]]},"reference":[{"key":"e_1_2_1_8_2_1","first-page":"12","volume-title":"Global Cybersecurity Outlook 2023","year":"2023"},{"key":"e_1_2_1_8_3_1","doi-asserted-by":"crossref","unstructured":"M.Kalash M.Rochan N.Mohammed N. D. B.Bruce Y.Wang andF.Iqbal \u201cMalware Classification With Deep Convolutional Neural Networks \u201d9th IFIP International Conference on New Technologies Mobility and Security (NTMS) Paris France 2018 1\u20135 https:\/\/doi.org\/10.1109\/NTMS.2018.8328749.","DOI":"10.1109\/NTMS.2018.8328749"},{"key":"e_1_2_1_8_4_1","doi-asserted-by":"crossref","unstructured":"N.Elmrabit F.Zhou F.Li andH.Zhou \u201cEvaluation of Machine Learning Algorithms for Anomaly Detection \u201d2020 International Conference on Cyber Security and Protection of Digital Services (Cyber Security) Dublin Ireland 1\u20138 https:\/\/doi.org\/10.1109\/CyberSecurity49315.2020.9138871.","DOI":"10.1109\/CyberSecurity49315.2020.9138871"},{"key":"e_1_2_1_8_5_1","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton R.","year":"2018"},{"key":"e_1_2_1_8_6_1","volume-title":"The Path to Autonomous Cyber Defense","author":"Oesch S.","year":"2024"},{"key":"e_1_2_1_8_7_1","doi-asserted-by":"crossref","unstructured":"S.Iannucci O. D.Barba V.Cardellini et\u00a0al. \u201cA Performance Evaluation of Deep Reinforcement Learning for Model\u2010Based Intrusion Response \u201d2019 IEEE 4th International Workshops on Foundations and Applications of Self* Systems (FAS*W) Umea Sweden 158\u2013163 https:\/\/doi.org\/10.1109\/FAS\u2010W.2019.00047.","DOI":"10.1109\/FAS-W.2019.00047"},{"key":"e_1_2_1_8_8_1","doi-asserted-by":"crossref","unstructured":"S.Oesch A.Chaulagain B.Weber et\u00a0al. \u201cTowards a High Fidelity Training Environment for Autonomous Cyber Defense Agents \u201d2024 In Proceedings of the 17th Cyber Security Experimentation and Test Workshop 91\u201399.","DOI":"10.1145\/3675741.3675752"},{"key":"e_1_2_1_8_9_1","volume-title":"Evaluation of Reinforcement Learning for Autonomous Penetration Testing Using A3C, Q\u2010Learning and DQN","author":"Becker N.","year":"2024"},{"key":"e_1_2_1_8_10_1","unstructured":"S.Cheung J.Claypoole P.Sharma et\u00a0al. \u201cEIReLaND: Evaluating and Interpreting Reinforcement\u2010Learning\u2010Based Network Defenses \u201d2024."},{"key":"e_1_2_1_8_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-41188-6_3"},{"key":"e_1_2_1_8_12_1","doi-asserted-by":"publisher","DOI":"10.22541\/au.173271087.78484560\/v1"},{"key":"e_1_2_1_8_13_1","doi-asserted-by":"crossref","unstructured":"E.Bates V.Mavroudis andC.Hicks \u201cReward Shaping for Happier Autonomous Cyber Security Agents \u201d2023 Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security 221\u2013232.","DOI":"10.1145\/3605764.3623916"},{"key":"e_1_2_1_8_14_1","doi-asserted-by":"publisher","DOI":"10.3390\/a15040134"},{"key":"e_1_2_1_8_15_1","doi-asserted-by":"crossref","unstructured":"K.HammarandR.Stadler \u201cFinding Effective Security Strategies Through Reinforcement Learning and Self\u2010Play \u201d2020 16th International Conference on Network and Service Management (CNSM) 1\u20139 IEEE.","DOI":"10.23919\/CNSM50824.2020.9269092"},{"key":"e_1_2_1_8_16_1","unstructured":"Ministry of Defence \u201cJSP 936 V1.1: Dependable Artificial Intelligence (AI) in Defence Part 1: Directive \u201d2024."},{"key":"e_1_2_1_8_17_1","volume-title":"On Autonomous Agents in a Cyber Defence Environment","author":"Kiely M.","year":"2023"},{"key":"e_1_2_1_8_18_1","volume-title":"Cyborg: A Gym for the Development of Autonomous Cyber Agents","author":"Standen M.","year":"2021"},{"key":"e_1_2_1_8_19_1","unstructured":"M.Kiely D.Bowman M.Standen andC.Moir \u201cTTCP CAGE Challenge 2 \u201daccessed September 16 2024 https:\/\/github.com\/cage\u2010challenge\/cage\u2010challenge\u20102."},{"key":"e_1_2_1_8_20_1","unstructured":"NATO Research and Technology Organisation Research Group (SAS\u2010026) \u201cNATO Code of Best Practice for C2 Assessment \u201d2002 http:\/\/www.dodccrp.org\/files\/NATO_COBP.pdf."},{"key":"e_1_2_1_8_21_1","unstructured":"J.Hannay \u201cCyborg\u2010Cage\u20102 \u201daccessed September 16 2024 https:\/\/github.com\/john\u2010cardiff\/\u2010cyborg\u2010cage\u20102."},{"key":"e_1_2_1_8_22_1","doi-asserted-by":"crossref","unstructured":"M.Foley C.Hicks K.Highnam andV.Mavroudis \u201cAutonomous Network Defence Using Reinforcement Learning \u201d2022 In Proceedings of the 2022 ACM on Asia Conference on Computer and Communications Security 1252\u20131254.","DOI":"10.1145\/3488932.3527286"},{"key":"e_1_2_1_8_23_1","volume-title":"Proximal Policy Optimization Algorithms","author":"Schulman J.","year":"2017"},{"key":"e_1_2_1_8_24_1","volume-title":"Beyond Cage: Investigating Generalization of Learned Autonomous Network Defense Policies","author":"Wolk M.","year":"2022"}],"container-title":["Applied AI Letters"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/ail2.125","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T08:54:04Z","timestamp":1760691244000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/ail2.125"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,27]]},"references-count":23,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,10]]}},"alternative-id":["10.1002\/ail2.125"],"URL":"https:\/\/doi.org\/10.1002\/ail2.125","archive":["Portico"],"relation":{},"ISSN":["2689-5595","2689-5595"],"issn-type":[{"value":"2689-5595","type":"print"},{"value":"2689-5595","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,27]]},"assertion":[{"value":"2025-02-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-14","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e125"}}