{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T06:10:53Z","timestamp":1784873453392,"version":"3.55.0"},"reference-count":28,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2022,9,12]],"date-time":"2022-09-12T00:00:00Z","timestamp":1662940800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,9,12]],"date-time":"2022-09-12T00:00:00Z","timestamp":1662940800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Intell Inf Syst"],"published-print":{"date-parts":[[2023,4]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Penetration testing (PT) is a method for assessing and evaluating the security of digital assets by planning, generating, and executing possible attacks that aim to discover and exploit vulnerabilities. In large networks, penetration testing becomes repetitive, complex and resource consuming despite the use of automated tools. This paper investigates reinforcement learning (RL) to make penetration testing more intelligent, targeted, and efficient. The proposed approach called Intelligent Automated Penetration Testing Framework (IAPTF) utilizes model-based RL to automate sequential decision making. Penetration testing tasks are treated as a partially observed Markov decision process (POMDP) which is solved with an external POMDP-solver using different algorithms to identify the most efficient options. A major difficulty encountered was solving large POMDPs resulting from large networks. This was overcome by representing networks hierarchically as a group of clusters and treating each cluster separately. This approach is tested through simulations of networks of various sizes. The results show that IAPTF with hierarchical network modeling outperforms previous approaches as well as human performance in terms of time, number of tested vectors and accuracy, and the advantage increases with the network size. Another advantage of IAPTF is the ease of repetition for retesting similar networks, which is often encountered in real PT. The results suggest that IAPTF is a promising approach to offload work from and ultimately replace human pen testing.<\/jats:p>","DOI":"10.1007\/s10844-022-00738-0","type":"journal-article","created":{"date-parts":[[2022,9,12]],"date-time":"2022-09-12T13:11:27Z","timestamp":1662988287000},"page":"281-303","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":63,"title":["Hierarchical reinforcement learning for efficient and effective automated penetration testing of large networks"],"prefix":"10.1007","volume":"60","author":[{"given":"Mohamed C.","family":"Ghanem","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thomas M.","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Erivelton G.","family":"Nepomuceno","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,9,12]]},"reference":[{"key":"738_CR1","doi-asserted-by":"crossref","unstructured":"Abu-Dabaseh, F., & Alshammari, E. (2018). Automated penetration testing : an overview computer science and information technology.","DOI":"10.5121\/csit.2018.80610"},{"key":"738_CR2","doi-asserted-by":"publisher","first-page":"2210","DOI":"10.12785\/ijcds\/040207","volume":"4","author":"M Al-Emran","year":"2015","unstructured":"Al-Emran, M. (2015). Hierarchical reinforcement learning: a survey. International Journal of Computing and Digital Systems, 4, 2210\u2013142. https:\/\/doi.org\/10.12785\/ijcds\/040207.","journal-title":"International Journal of Computing and Digital Systems"},{"key":"738_CR3","doi-asserted-by":"publisher","unstructured":"Babenko, L., & Kirillov, A. (2022). Development of automated malware detection system. izvestiya SFedu. Engineering Sciences:153\u2013167. https:\/\/doi.org\/10.18522\/2311-3103-2021-7-153-167.","DOI":"10.18522\/2311-3103-2021-7-153-167"},{"key":"738_CR4","unstructured":"Backes, M., Hoffmann, J., K\u00fcnnemann, R, Speicher, P., & Steinmetz, M. (2017). Simulated penetration testing and mitigation analysis. arXiv:1705.05088."},{"issue":"1-2","key":"738_CR5","doi-asserted-by":"publisher","first-page":"19","DOI":"10.5121\/ijnsa.2011.3602","volume":"3","author":"A Bacudio","year":"2011","unstructured":"Bacudio, A., Yuan, X., Chu, B., & Jones, M. (2011). An overview of penetration testing. International Journal of Network Security & Its Applications, 3 (1-2), 19\u201338. https:\/\/doi.org\/10.5121\/ijnsa.2011.3602.","journal-title":"International Journal of Network Security & Its Applications"},{"issue":"1","key":"738_CR6","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s13173-016-0050-7","volume":"23","author":"DD Bertoglio","year":"2017","unstructured":"Bertoglio, D.D., & Zorzo, A.F. (2017). Overview and open issues on penetration test. Journal of the Brazilian Computer Society, 23(1), 1\u201316.","journal-title":"Journal of the Brazilian Computer Society"},{"key":"738_CR7","unstructured":"Boddy, M., Gohde, J., Haigh, T., & Harp, S. (2005). Course of action generation for cyber security using classical planning. Proceedings of the 15th international conference on automated planning and scheduling. ICAPS\u201905:12\u201321."},{"key":"738_CR8","unstructured":"Cassandra, A.R., Littman, M.L., & Zhang, N.L. (2013). Incremental pruning: A simple, fast, exact method for partially observable markov decision processes. arXiv preprint arXiv:1302.1525"},{"key":"738_CR9","doi-asserted-by":"publisher","first-page":"6","DOI":"10.3390\/info11010006","volume":"11","author":"M Ghanem","year":"2019","unstructured":"Ghanem, M., & Chen, T. (2019). Reinforcement learning for efficient network penetration testing. Information, 11, 6. https:\/\/doi.org\/10.3390\/info11010006.","journal-title":"Information"},{"key":"738_CR10","doi-asserted-by":"crossref","unstructured":"He, L., & Bode, N. (2006). Network penetration testing. In A Blyth (Ed.) EC2ND 2005. Springer, pp 3\u201312.","DOI":"10.1007\/1-84628-352-3_1"},{"key":"738_CR11","unstructured":"Jain, A., & Niekum, S. (2018). Efficient hierarchical robot motion planning under uncertainty and hybrid dynamics."},{"key":"738_CR12","doi-asserted-by":"publisher","unstructured":"Joglekar, N. (2008). Hierarchical planning under uncertainty: Real options and heuristics, pp 291\u2013313. https:\/\/doi.org\/10.1016\/B978-0-7506-8552-8.50014-1.","DOI":"10.1016\/B978-0-7506-8552-8.50014-1"},{"key":"738_CR13","doi-asserted-by":"publisher","first-page":"100","DOI":"10.1016\/j.cose.2020.102108","volume":"102108","author":"R Maeda","year":"2021","unstructured":"Maeda, R., & Mimura, M. (2021). Automating post-exploitation with deep reinforcement learning. Computers & Security, 102108, 100. https:\/\/doi.org\/10.1016\/j.cose.2020.102108.","journal-title":"Computers & Security"},{"key":"738_CR14","doi-asserted-by":"publisher","unstructured":"Moerland, T., Broekens, J., & Jonker, C. (2020). Model-based reinforcement learning: a survey. https:\/\/doi.org\/10.48550\/arXiv.2006.16712.","DOI":"10.48550\/arXiv.2006.16712"},{"key":"738_CR15","unstructured":"Pineau, J., & Gordon, G. (2003). Point-based value iteration: an anytime algorithm for pomdps. In Proceedings international joint conference of artificial intelligence, pp. 1025\u20131032."},{"key":"738_CR16","doi-asserted-by":"publisher","first-page":"50","DOI":"10.4018\/ijdcf.2014100104","volume":"6","author":"C Phong","year":"2014","unstructured":"Phong, C., & Yan, W. (2014). An overview of penetration testing. International Journal of Digital Crime and Forensics (IJDCF), 6, 50\u201374. https:\/\/doi.org\/10.4018\/ijdcf.2014100104.","journal-title":"International Journal of Digital Crime and Forensics (IJDCF)"},{"key":"738_CR17","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1613\/jair.3987","volume":"48","author":"DM Roijers","year":"2013","unstructured":"Roijers, D.M., Vamplew, P., Whiteson, S., & Dazeley, R. (2013). A survey of multi-objective sequential decision-making. Journal of Artificial Intelligence Research, 48, 67\u2013113.","journal-title":"Journal of Artificial Intelligence Research"},{"key":"738_CR18","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1613\/jair.1659","volume":"24","author":"M Spaan","year":"2005","unstructured":"Spaan, M., & Vlassis, N. (2005). Perseus: Randomized point-based value iteration for pomdps. Journal Artificial Intelligent Research (JAIR), 24, 195\u2013220.","journal-title":"Journal Artificial Intelligent Research (JAIR)"},{"key":"738_CR19","doi-asserted-by":"crossref","unstructured":"Spaan, M.T. (2012). Partially observable markov decision processes. In Reinforcement learning, pp 387\u2013414.","DOI":"10.1007\/978-3-642-27645-3_12"},{"key":"738_CR20","doi-asserted-by":"crossref","unstructured":"Sarraute, C., Buffet, O., & Hoffmann, J. (2012). Pomdps make better hackers: accounting for uncertainty in penetration testing. Proceedings of the twenty-sixth aaai conference on artificial intelligence, pp 1816\u20131824.","DOI":"10.1609\/aaai.v26i1.8363"},{"key":"738_CR21","doi-asserted-by":"publisher","unstructured":"Sarraute, C., Richarte, G., & Luc\u00e1ngeli Obes, J. (2013). An algorithm to find optimal attack paths in nondeterministic scenarios. Proceedings of the acm conference on computer and communications security. https:\/\/doi.org\/10.1145\/2046684.2046695.","DOI":"10.1145\/2046684.2046695"},{"issue":"4","key":"738_CR22","doi-asserted-by":"publisher","first-page":"373","DOI":"10.1007\/s13218-017-0507-7","volume":"31","author":"S Stock","year":"2017","unstructured":"Stock, S. (2017). Hierarchical hybrid planning for mobile robots. KI-K\u00fcnstliche Intelligenz, 31(4), 373\u2013376.","journal-title":"KI-K\u00fcnstliche Intelligenz"},{"key":"738_CR23","doi-asserted-by":"crossref","unstructured":"Walraven, E., & Spaan, M.T.J. (2017). Accelerated vector pruning for optimal pomdp solvers. In AAAI.","DOI":"10.1609\/aaai.v31i1.11032"},{"key":"738_CR24","doi-asserted-by":"crossref","unstructured":"Walraven, E., & Spaan, M. (2017). Accelerated vector pruning for optimal pomdp solvers. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 31(1).","DOI":"10.1609\/aaai.v31i1.11032"},{"key":"738_CR25","first-page":"12","volume":"7","author":"I Yaqoob","year":"2017","unstructured":"Yaqoob, I., Hussain, S., Mamoon, S., Naseer, N., & Akram, J. (2017). Penetration testing and vulnerability assessment. Journal of Network Communications and Emerging Technologies, 7, 12\u201321.","journal-title":"Journal of Network Communications and Emerging Technologies"},{"key":"738_CR26","unstructured":"Zennaro, F.M., & Erdodi, L. (2020). Modeling penetration testing with reinforcement learning using capture-the-flag challenges and tabular q-learning. arXiv:2005.12632."},{"key":"738_CR27","doi-asserted-by":"publisher","unstructured":"Zhang, W., Nevin, D., Zhang, N., Supervisor, T., & Golin, G. (2003). Algorithms for partially observable markov decision processes. https:\/\/doi.org\/10.14288\/1.0098252.","DOI":"10.14288\/1.0098252"},{"key":"738_CR28","doi-asserted-by":"publisher","unstructured":"Zhou, R., Pan, J., Tan, X., & Xi, H. (2008). Application of clips expert system to malware detection system, vol. 1, pp. 309\u2013314. https:\/\/doi.org\/10.1109\/CIS.2008.100.","DOI":"10.1109\/CIS.2008.100"}],"container-title":["Journal of Intelligent Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10844-022-00738-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10844-022-00738-0\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10844-022-00738-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,4,24]],"date-time":"2023-04-24T21:04:06Z","timestamp":1682370246000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10844-022-00738-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,12]]},"references-count":28,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,4]]}},"alternative-id":["738"],"URL":"https:\/\/doi.org\/10.1007\/s10844-022-00738-0","relation":{},"ISSN":["0925-9902","1573-7675"],"issn-type":[{"value":"0925-9902","type":"print"},{"value":"1573-7675","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,12]]},"assertion":[{"value":"23 May 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 August 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 August 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 September 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Authors give full consent for publication.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"<!--Emphasis Type='Bold' removed-->Consent for publication"}},{"value":"The authors declare that they have no known competing interests or personal relationships that could have appeared to influence the work reported in this paper.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"<!--Emphasis Type='Bold' removed-->Competing interests"}}]}}