{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,22]],"date-time":"2026-03-22T10:00:44Z","timestamp":1774173644522,"version":"3.50.1"},"reference-count":52,"publisher":"Maximum Academic Press","license":[{"start":{"date-parts":[[2023,3,6]],"date-time":"2023-03-06T00:00:00Z","timestamp":1678060800000},"content-version":"unspecified","delay-in-days":64,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":["cambridge.org"],"crossmark-restriction":true},"short-container-title":["The Knowledge Engineering Review"],"published-print":{"date-parts":[[2023]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>We investigate artificial intelligence and machine learning methods for optimizing the adversarial behavior of agents in cybersecurity simulations. Our cybersecurity simulations integrate the modeling of agents launching Advanced Persistent Threats (APTs) with the modeling of agents using detection and mitigation mechanisms against APTs. This simulates the phenomenon of how attacks and defenses coevolve. The simulations and machine learning are used to search for optimal agent behaviors. The central question is: under what circumstances, is one training method more advantageous than another? We adapt and compare a variety of deep reinforcement learning (DRL), evolutionary strategies (ES) and Monte Carlo Tree Search methods within Connect 4, a baseline game environment, and on both a simulation supporting a simple APT threat model, SNAPT, as well as CyberBattleSim, an open-source cybersecurity simulation. Our results show that when attackers are trained by DRL and ES algorithms, as well as when they are trained with both algorithms being used in alternation, they are able to effectively choose complex exploits that thwart a defense. The algorithm that combines DRL and ES achieves the best comparative performance when attackers and defenders are simultaneously trained, rather than when each is trained against its non-learning counterpart.<\/jats:p>","DOI":"10.1017\/s0269888923000012","type":"journal-article","created":{"date-parts":[[2023,3,6]],"date-time":"2023-03-06T03:41:59Z","timestamp":1678074119000},"update-policy":"https:\/\/doi.org\/10.1017\/policypage","source":"Crossref","is-referenced-by-count":11,"title":["Adversarial agent-learning for cybersecurity: a comparison of algorithms"],"prefix":"10.48130","volume":"38","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6651-7705","authenticated-orcid":false,"given":"Alexander","family":"Shashkov","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2153-3506","authenticated-orcid":false,"given":"Erik","family":"Hemberg","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Miguel","family":"Tulla","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Una-May","family":"O\u2019Reilly","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"27968","published-online":{"date-parts":[[2023,3,6]]},"reference":[{"key":"S0269888923000012_ref1","first-page":"165","article-title":"A knowledge-based approach of connect-four","volume":"11","author":"Allis","year":"1988","journal-title":"Journal of the International Computer Games Association"},{"key":"S0269888923000012_ref43","doi-asserted-by":"publisher","DOI":"10.1109\/CIG.2008.5035662"},{"key":"S0269888923000012_ref9","unstructured":"Corporation, T. M. (n.d.b). Mitre engage. https:\/\/engage.mitre.org."},{"key":"S0269888923000012_ref27","unstructured":"Metz, L. , Ibarz, J. , Jaitly, N. & Davidson, J. 2017. Discrete sequential prediction of continuous actions for deep RL. ArXiv abs\/1705.05035."},{"key":"S0269888923000012_ref42","unstructured":"Salimans, T. , Ho, J. , Chen, X. & Sutskever, I. 2017. Evolution strategies as a scalable alternative to reinforcement learning. ArXiv abs\/1703.03864."},{"key":"S0269888923000012_ref11","doi-asserted-by":"publisher","DOI":"10.1016\/j.cose.2022.102681"},{"key":"S0269888923000012_ref28","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"S0269888923000012_ref10","doi-asserted-by":"publisher","DOI":"10.1016\/S0304-3975(01)00182-7"},{"key":"S0269888923000012_ref8","unstructured":"Corporation, T. M. (n.d.a). Mitre att&ck. https:\/\/attack.mitre.org."},{"key":"S0269888923000012_ref5","unstructured":"Baillie, C. , Standen, M. , Schwartz, J. , Docking, M. , Bowman, D. & Kim, J. 2020. Cyborg: An autonomous cyber operations research gym. ArXiv abs\/2002.10667."},{"key":"S0269888923000012_ref17","first-page":"1898","volume-title":"Competitive Coevolution for Defense and Security: Elo-Based Similar-Strength Opponent Sampling","author":"Harris","year":"2021"},{"key":"S0269888923000012_ref36","doi-asserted-by":"publisher","DOI":"10.1162\/106365600568086"},{"key":"S0269888923000012_ref22","unstructured":"Lillicrap, T. P. , Hunt, J. J. , Pritzel, A. , Heess, N. M. O. , Erez, T. , Tassa, Y. , Silver, D. & Wierstra, D. 2016. Continuous control with deep reinforcement learning. ArXiv abs\/1509.02971."},{"key":"S0269888923000012_ref50","unstructured":"Walter, E. C. , Ferguson-Walter, K. J. & Ridley, A. D. 2021. Incorporating deception into cyberbattlesim for autonomous defense. ArXiv abs\/2108.13980."},{"key":"S0269888923000012_ref26","article-title":"Fully distributed actor-critic architecture for multitask deep reinforcement learning","author":"Macua","year":"2021","journal-title":"The Knowledge Engineering Review"},{"key":"S0269888923000012_ref6","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-46133-1_9"},{"key":"S0269888923000012_ref40","unstructured":"Reinstadler, B. 2021. Ai Attack Planning for Emulated Networks. Master\u2019s thesis, Massachusetts Institute of Technology."},{"key":"S0269888923000012_ref34","unstructured":"Partalas, I. , Vrakas, D. & Vlahavas, I. 2012. Reinforcement learning and automated planning: a survey. In Artificial Intelligence for Advanced Problem Solving Techniques."},{"key":"S0269888923000012_ref47","unstructured":"Team, M. D. R. 2021. Cyberbattlesim. https:\/\/github.com\/microsoft\/cyberbattlesim. Created by Christian Seifert, Michael Betser, William Blum, James Bono, Kate Farris, Emily Goren, Justin Grana, Kristian Holsheimer, Brandon Marken, Joshua Neil, Nicole Nichols, Jugal Parikh, Haoran Wei."},{"key":"S0269888923000012_ref51","first-page":"1713","article-title":"Effective repair strategy against advanced persistent threat: a differential game approach","author":"Yang","year":"2018","journal-title":"IEEE Transactions on Information Forensics and Security"},{"key":"S0269888923000012_ref32","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2016.2520371"},{"key":"S0269888923000012_ref3","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2017.2743240"},{"key":"S0269888923000012_ref46","first-page":"1","article-title":"Achieving long-term progress in competitive co-evolution","author":"Simione","year":"2017","journal-title":"2017 IEEE Symposium Series on Computational Intelligence (SSCI)"},{"key":"S0269888923000012_ref37","unstructured":"Pourchot, A. & Sigaud, O. 2018. CEM-RL: combining evolutionary and gradient-based methods for policy search. ArXiv abs\/1810.01222."},{"key":"S0269888923000012_ref16","unstructured":"Hansen, N. (2016). The CMA evolution strategy: a tutorial. ArXiv abs\/1604.00772."},{"key":"S0269888923000012_ref12","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2867556"},{"key":"S0269888923000012_ref18","doi-asserted-by":"publisher","DOI":"10.1016\/j.cose.2019.101660"},{"key":"S0269888923000012_ref19","doi-asserted-by":"publisher","DOI":"10.1017\/S026988891200001X"},{"key":"S0269888923000012_ref23","first-page":"174","volume-title":"2016 8th Computer Science and Electronic Engineering","author":"Liu","year":"2016"},{"key":"S0269888923000012_ref7","unstructured":"Brockman, G. , Cheung, V. , Pettersson, L. , Schneider, J. , Schulman, J. , Tang, J. & Zaremba, W. 2016. Openai gym. ArXiv abs\/1606.01540."},{"key":"S0269888923000012_ref24","article-title":"Improving software security awareness using a serious game","author":"Liu","year":"2018","journal-title":"IET Software"},{"key":"S0269888923000012_ref29","unstructured":"Molina-Markham, A. , Winder, R. K. & Ridley, A. 2021. Network defense is not a game. ArXiv abs\/2104.10262."},{"key":"S0269888923000012_ref21","first-page":"10124","volume-title":"Advances in Neural Information Processing Systems","volume":"33","author":"Lee","year":"2020"},{"key":"S0269888923000012_ref13","volume-title":"Genetic Algorithms in Search, Optimization and Machine Learning","author":"Goldberg","year":"1989"},{"key":"S0269888923000012_ref38","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-72699-7_18"},{"key":"S0269888923000012_ref39","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-83814-9_6"},{"key":"S0269888923000012_ref14","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCC.2012.2218595"},{"key":"S0269888923000012_ref33","unstructured":"Panait, L. & Luke, S. 2002. A comparison of two competitive fitness functions. In Proceedings of the 4th Annual Conference on Genetic and Evolutionary Computation, GECCO\u201902, Morgan Kaufmann Publishers Inc., 503\u2013511."},{"key":"S0269888923000012_ref41","doi-asserted-by":"publisher","DOI":"10.1145\/2739482.2768429"},{"key":"S0269888923000012_ref48","unstructured":"The MITRE Corporation 2020. Caldera. https:\/\/github.com\/mitre\/caldera."},{"key":"S0269888923000012_ref4","unstructured":"Backes, M. , Hoffmann, J. , K\u00fcnnemann, R. , Speicher, P. & Steinmetz, M. 2017. Simulated penetration testing and mitigation analysis. ArXiv abs\/1705.05088."},{"key":"S0269888923000012_ref25","doi-asserted-by":"publisher","DOI":"10.1007\/s11416-019-00342-x"},{"key":"S0269888923000012_ref30","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2021.3121870"},{"key":"S0269888923000012_ref15","unstructured":"Group, A. 2022. Adversarial agent-learning for cybersecurity. https:\/\/github.com\/ALFA-group\/adversarial_agent_learning_for_cybersecurity."},{"key":"S0269888923000012_ref35","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-92910-9_31"},{"key":"S0269888923000012_ref20","first-page":"283","volume-title":"A Coevolutionary Approach to Deep Multi-Agent Reinforcement Learning","author":"Klijn","year":"2021"},{"key":"S0269888923000012_ref52","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2814481"},{"key":"S0269888923000012_ref31","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-72699-7_33"},{"key":"S0269888923000012_ref2","doi-asserted-by":"crossref","unstructured":"Applebaum, A. , Miller, D. , Strom, B. , Korban, C. & Wolf, R. 2016. Intelligent, automated red team emulation. In Proceedings of the 32nd Annual Conference on Computer Security Applications, 363\u2013373.","DOI":"10.1145\/2991079.2991111"},{"key":"S0269888923000012_ref49","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-019-1724-z"},{"key":"S0269888923000012_ref44","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2019.01.011"},{"key":"S0269888923000012_ref45","doi-asserted-by":"publisher","DOI":"10.1126\/science.aar6404"}],"container-title":["The Knowledge Engineering Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S0269888923000012","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,5]],"date-time":"2026-01-05T14:42:23Z","timestamp":1767624143000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S0269888923000012\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023]]},"references-count":52,"alternative-id":["S0269888923000012"],"URL":"https:\/\/doi.org\/10.1017\/s0269888923000012","relation":{},"ISSN":["0269-8889","1469-8005"],"issn-type":[{"value":"0269-8889","type":"print"},{"value":"1469-8005","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023]]},"assertion":[{"value":"\u00a9 The Author(s), 2023. Published by Cambridge University Press","name":"copyright","label":"Copyright","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}}],"article-number":"e3"}}