{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T22:58:53Z","timestamp":1783119533756,"version":"3.54.6"},"reference-count":35,"publisher":"Wiley","issue":"3","license":[{"start":{"date-parts":[[2025,9,5]],"date-time":"2025-09-05T00:00:00Z","timestamp":1757030400000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100010418","name":"Defence Science and Technology Laboratory","doi-asserted-by":"publisher","award":["U75AT030"],"award-info":[{"award-number":["U75AT030"]}],"id":[{"id":"10.13039\/100010418","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Applied AI Letters"],"published-print":{"date-parts":[[2025,10]]},"abstract":"<jats:title>ABSTRACT<\/jats:title><jats:p>Cyber\u2010attacks pose a security threat to military command and control networks, Intelligence, Surveillance, and Reconnaissance (ISR) systems, and civilian critical national infrastructure. The use of artificial intelligence and autonomous agents in these attacks increases the scale, range, and complexity of this threat and the subsequent disruption they cause. Autonomous Cyber Defence (ACD) agents aimto mitigate this threat by responding at machine speed and at the scale required to address the problem. Additionally, they reduce the burden on the limited number of human cyber experts available to respond to an attack. Sequential decision\u2010making algorithms such as Deep Reinforcement Learning (RL) provide a promising route to create ACD agents. These algorithms focus on a single objectivesuch as minimising the intrusion of red agents on the network, by using a handcrafted weighted sum of rewards. This approach removes the ability to adapt the model during inference, and fails to address the many competing objectivespresent when operating and protecting these networks. Conflicting objectives, such as restoring a machine from a back\u2010up image, must be carefully balanced with the cost of associated down\u2010time or the disruption to network traffic or services that might result. Instead of pursuing a Single\u2010Objective RL (SORL) approach, here we present a simple example of a multi\u2010objective network defense game that requires consideration of both defending the network against red\u2010agents and maintaining the critical functionality of green\u2010agents. Two Multi\u2010Objective Reinforcement Learning (MORL) algorithms, namely Multi\u2010Objective Proximal Policy Optimization (MOPPO) and Pareto\u2010Conditioned Networks (PCN), are used to create two trained ACD agents whose performance is compared on our Multi\u2010Objective Cyber Defense game. The benefits and limitations of MORL ACD agents in comparison to SORL ACD agents are discussed based on the investigations of this game.<\/jats:p>","DOI":"10.1002\/ail2.70007","type":"journal-article","created":{"date-parts":[[2025,9,5]],"date-time":"2025-09-05T09:21:27Z","timestamp":1757064087000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Multi\u2010Objective Reinforcement Learning for Automated Resilient Cyber Defence"],"prefix":"10.1002","volume":"6","author":[{"given":"Ross","family":"O'Driscoll","sequence":"first","affiliation":[{"name":"Roke Manor Research Ltd.  Woking UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Claudia","family":"Hagen","sequence":"additional","affiliation":[{"name":"Roke Manor Research Ltd.  Woking UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Joe","family":"Bater","sequence":"additional","affiliation":[{"name":"Roke Manor Research Ltd.  Woking UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5299-6378","authenticated-orcid":false,"given":"James","family":"Adams","sequence":"additional","affiliation":[{"name":"School of Mathematics and Physics University of Surrey  Guildford UK"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2025,9,5]]},"reference":[{"key":"e_1_2_1_9_2_1","unstructured":"S.Vyas J.Hannay A.Bolton andP. P.Burnap \u201cAutomated Cyber Defence: A Review \u201d(2023) https:\/\/doi.org\/10.48550\/arXiv.2303.04926."},{"key":"e_1_2_1_9_3_1","volume-title":"Autonomous Intelligent Cyber\u2010Defense Agent (AICA) Reference Architecture, Release 2.0","author":"Kott A.","year":"2019"},{"key":"e_1_2_1_9_4_1","unstructured":"M. J.Weisman A.Kott J. E.Ellis et\u00a0al. \u201cQuantitative Measurement of Cyber Resilience: Modeling and Experimentation \u201d(2023) https:\/\/doi.org\/10.48550\/arXiv.2303.16307."},{"key":"e_1_2_1_9_5_1","unstructured":"M.Standen M.Lucas D.Bowman T. J.Richer J.Kim andD.Marriott \u201cCybORG: A Gym for the Development of Autonomous Cyber Agents \u201d(2021)."},{"key":"e_1_2_1_9_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3605764.3623916"},{"key":"e_1_2_1_9_7_1","unstructured":"M.Wolk \u201cBeyond CAGE: Investigating Generalization of Learned Autonomous Network Defense Policies \u201d(2022)."},{"key":"e_1_2_1_9_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10458\u2010022\u201009552\u2010y"},{"key":"e_1_2_1_9_9_1","volume-title":"Proceedings of the 36th International Conference on Machine Learning","author":"Abels A.","year":"2019"},{"key":"e_1_2_1_9_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10994\u2010010\u20105232\u20105"},{"key":"e_1_2_1_9_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ADPRL.2013.6615007"},{"key":"e_1_2_1_9_12_1","first-page":"10607","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Xu J.","year":"2020"},{"key":"e_1_2_1_9_13_1","first-page":"3663","article-title":"Multi\u2010Objective Reinforcement Learning Using Sets of Pareto Dominating Policies","volume":"15","author":"Van Moffaert K.","year":"2014","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_1_9_14_1","unstructured":"M.Reymond E.Bargiacchi andA.Now\u00e9 \u201cPareto Conditioned Networks \u201d(2022) https:\/\/doi.org\/10.48550\/arXiv.2204.05036."},{"key":"e_1_2_1_9_15_1","unstructured":"R.Yang X.Sun andK.Narasimhan \u201cA Generalized Algorithm for Multi\u2010Objective Reinforcement Learning and Policy Adaptation \u201d(2019) https:\/\/doi.org\/10.48550\/arXiv.1908.08342."},{"key":"e_1_2_1_9_16_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2022\/476"},{"key":"e_1_2_1_9_17_1","unstructured":"A.Andrew S.Spillard J.Collyer andN.Dhir \u201cDeveloping Optimal Causal Cyber\u2010Defence Agents via Cyber Security Simulation \u201d(2022)."},{"key":"e_1_2_1_9_18_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586\u2010020\u201003051\u20104"},{"key":"e_1_2_1_9_19_1","unstructured":"J.Schulman F.Wolski P.Dhariwal A.Radford andO.Klimov \u201cProximal Policy Optimization Algorithms \u201d(2017) https:\/\/doi.org\/10.48550\/arXiv.1707.06347."},{"key":"e_1_2_1_9_20_1","unstructured":"R.Agarwal D.Schuurmans andM.Norouzi \u201cAn Optimistic Perspective on Offline Reinforcement Learning \u201d(2019) https:\/\/doi.org\/10.48550\/arX1907.04543iv."},{"key":"e_1_2_1_9_21_1","unstructured":"L.Chen K.Lu A.Rajeswaran et\u00a0al. \u201cDecision Transformer: Reinforcement Learning via Sequence Modeling \u201d(2021) https:\/\/doi.org\/10.48550\/arXiv.2106.01345."},{"key":"e_1_2_1_9_22_1","doi-asserted-by":"crossref","unstructured":"M.Pternea P.Singh A.Chakraborty et\u00a0al. \u201cThe RL\/LLM Taxonomy Tree: Reviewing Synergies Between Reinforcement Learning and Large Language Models \u201d(2024) https:\/\/doi.org\/10.48550\/arXiv.2402.01874.","DOI":"10.1613\/jair.1.15960"},{"key":"e_1_2_1_9_23_1","doi-asserted-by":"crossref","unstructured":"Z.Shi X.Chen X.Qiu andX.Huang \u201cToward Diverse Text Generation With Inverse Reinforcement Learning \u201d(2018) https:\/\/doi.org\/10.48550\/arXiv.1804.11258.","DOI":"10.24963\/ijcai.2018\/606"},{"key":"e_1_2_1_9_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3488932.3527286"},{"key":"e_1_2_1_9_25_1","doi-asserted-by":"crossref","unstructured":"M.Foley M.Wang Z.M C.Hicks andV.Mavroudis \u201cInroads Into Autonomous Network Defence Using Explained Reinforcement Learning \u201d(2023).","DOI":"10.1145\/3488932.3527286"},{"key":"e_1_2_1_9_26_1","volume-title":"CAMLIS 2023: Conference on Applied Machine Learning in Information Security","author":"Acuto A.","year":"2023"},{"key":"e_1_2_1_9_27_1","unstructured":"Y.YangandX.Liu \u201cBehaviour\u2010Diverse Automatic Penetration Testing: A Curiosity\u2010Driven Multi\u2010Objective Deep Reinforcement Learning Approach \u201d(2022)."},{"key":"e_1_2_1_9_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICMLA.2015.144"},{"key":"e_1_2_1_9_29_1","doi-asserted-by":"publisher","DOI":"10.3390\/app8010136"},{"key":"e_1_2_1_9_30_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2020.104021"},{"key":"e_1_2_1_9_31_1","unstructured":"L. N.Alegre F.Felten E.\u2010G.Talbi G.Danoy A.Now\u00e9 andA. L. C.Bazzan \u201cMO\u2010Gym: A Library of Multi\u2010Objective Reinforcement Learning Environments \u201d(2022)."},{"key":"e_1_2_1_9_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11047\u2010018\u20109685\u2010y"},{"key":"e_1_2_1_9_33_1","doi-asserted-by":"crossref","unstructured":"S.WeiandM.Niethammer \u201cThe Fairness\u2010Accuracy Pareto Front \u201d(2021).","DOI":"10.1002\/sam.11560"},{"key":"e_1_2_1_9_34_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2014.08.071"},{"key":"e_1_2_1_9_35_1","unstructured":"K.Paster S.McIlraith andJ.Ba \u201cYou Can't Count on Luck: Why Decision Transformers and RvS Fail in Stochastic Environments \u201d(2022)."},{"key":"e_1_2_1_9_36_1","volume-title":"Adaptive and Learning Agents Workshop (ALA)","author":"Delgrange F.","year":"2023"}],"container-title":["Applied AI Letters"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/ail2.70007","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T08:53:56Z","timestamp":1760691236000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/ail2.70007"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,5]]},"references-count":35,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,10]]}},"alternative-id":["10.1002\/ail2.70007"],"URL":"https:\/\/doi.org\/10.1002\/ail2.70007","archive":["Portico"],"relation":{},"ISSN":["2689-5595","2689-5595"],"issn-type":[{"value":"2689-5595","type":"print"},{"value":"2689-5595","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,5]]},"assertion":[{"value":"2025-01-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-05","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70007"}}