{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T15:28:52Z","timestamp":1785511732399,"version":"3.56.0"},"reference-count":32,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2019,12,20]],"date-time":"2019-12-20T00:00:00Z","timestamp":1576800000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Penetration testing (also known as pentesting or PT) is a common practice for actively assessing the defenses of a computer network by planning and executing all possible attacks to discover and exploit existing vulnerabilities. Current penetration testing methods are increasingly becoming non-standard, composite and resource-consuming despite the use of evolving tools. In this paper, we propose and evaluate an AI-based pentesting system which makes use of machine learning techniques, namely reinforcement learning (RL) to learn and reproduce average and complex pentesting activities. The proposed system is named Intelligent Automated Penetration Testing System (IAPTS) consisting of a module that integrates with industrial PT frameworks to enable them to capture information, learn from experience, and reproduce tests in future similar testing cases. IAPTS aims to save human resources while producing much-enhanced results in terms of time consumption, reliability and frequency of testing. IAPTS takes the approach of modeling PT environments and tasks as a partially observed Markov decision process (POMDP) problem which is solved by POMDP-solver. Although the scope of this paper is limited to network infrastructures PT planning and not the entire practice, the obtained results support the hypothesis that RL can enhance PT beyond the capabilities of any human PT expert in terms of time consumed, covered attacking vectors, accuracy and reliability of the outputs. In addition, this work tackles the complex problem of expertise capturing and re-use by allowing the IAPTS learning module to store and re-use PT policies in the same way that a human PT expert would learn but in a more efficient way.<\/jats:p>","DOI":"10.3390\/info11010006","type":"journal-article","created":{"date-parts":[[2019,12,20]],"date-time":"2019-12-20T09:50:33Z","timestamp":1576835433000},"page":"6","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":112,"title":["Reinforcement Learning for Efficient Network Penetration Testing"],"prefix":"10.3390","volume":"11","author":[{"given":"Mohamed C.","family":"Ghanem","sequence":"first","affiliation":[{"name":"School of Mathematics Computer Science and Engineering, University of London, London EC1V 0HB, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thomas M.","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Mathematics Computer Science and Engineering, University of London, London EC1V 0HB, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2019,12,20]]},"reference":[{"key":"ref_1","unstructured":"Creasey, J., and Glover, I. (2017). A Guide for Running an Effective Penetration Testing Program, CREST Publication. Available online: http:\/\/www.crest-approved.org."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Almubairik, N., and Wills, G. (2016, January 5\u20137). Automated penetration testing based on a threat model. Proceedings of the 11th International Conference for Internet Technologies and Secured Transactions, ICITST, Barcelona, Spain.","DOI":"10.1109\/ICITST.2016.7856742"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Applebaum, A., Miller, D., Strom, B., Korban, C., and Wol, R. (2016, January 5\u20138). Intelligent, automated red team emulation. Proceedings of the 32nd Annual Conference on Computer Security Applications (ACSAC \u201916), Los Angeles, CA, USA.","DOI":"10.1145\/2991079.2991111"},{"key":"ref_4","unstructured":"Obes, J., Richarte, G., and Sarraute, C. (2013). Attack planning in the real world. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Spaan, M. (2012). Partially Observable Markov Decision Processes, Reinforcement Learning: State of the Art, Springer.","DOI":"10.1007\/978-3-642-27645-3_12"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Hoffmann, J. (2015, January 7\u201311). Simulated penetration testing: From Dijkstra to aaTuring Test++. Proceedings of the 25th International Conference on Automated Planning and Scheduling, Israel.","DOI":"10.1609\/icaps.v25i1.13684"},{"key":"ref_7","unstructured":"Sarraute, C. (2019, December 20). Automated Attack Planning. Available online: https:\/\/arxiv.org\/abs\/1307.7808."},{"key":"ref_8","unstructured":"Qiu, X., Jia, Q., Wang, S., Xia, C., and Shuang, L. (2014, January 8\u201310). Automatic generation algorithm of penetration graph in penetration testing. Proceedings of the Ninth International Conference on P2P, Parallel, Grid, Cloud and Internet Computing, Guangdong, China."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Heinl, C. (2014, January 3\u20136). Artificial (intelligent) agents and active cyber defence: Policy implications. Proceedings of the 6th International Conference On Cyber Conflict (CyCon 2014), Tallinn, Estonia.","DOI":"10.1109\/CYCON.2014.6916395"},{"key":"ref_10","unstructured":"Sarraute, C., Buffet, O., and Hoffmann, J. (2019, December 20). POMDPs Make Better Hackers: Accounting for Uncertainty in Penetration Testing. Available online: https:\/\/arxiv.org\/abs\/1307.8182."},{"key":"ref_11","unstructured":"Backes, M., Hoffmann, J., Kunnemann, R., Speicher, P., and Steinmetz, M. (2017). Simulated Penetration Testing and Mitigation Analysis. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Ghanem, M., and Chen, T. (2018, January 30\u201331). Reinforcement Learning for Intelligent Penetration Testing. Proceedings of the WS4 the World Conference on Smart Trends in Systems, Security and Sustainability, London, UK.","DOI":"10.1109\/WorldS4.2018.8611595"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Durkota, K., Lisy, V., Bosansk, B., and Kiekintveld, C. (2015, January 25\u201331). Optimal network security hardening using attack graph games. Proceedings of the 24th International Joint Conference on on Artificial Intelligence (IJCAI-2015), Buenos Aires, Argentina.","DOI":"10.1109\/MIS.2016.74"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Veeramachaneni, K., Arnaldo, I., Cuesta-Infante, A., Korrapati, V., Bassias, C., and Li, K. (2016). AI2: Training a Big Data Machine to Defend, CSAIL, MIT Cambridge.","DOI":"10.1109\/BigDataSecurity-HPSC-IDS.2016.79"},{"key":"ref_15","unstructured":"Sutton, R.S., and Barto, A.G. (2018). Reinforcement Learning: An Introduction, MIT Press."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"433","DOI":"10.1017\/S026988891200001X","article-title":"A review of machine learning for automated planning","volume":"27","author":"Jimenez","year":"2009","journal-title":"Knowl. Eng. Rev."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"307","DOI":"10.1613\/jair.1.11324","article-title":"Point-Based Value Iteration for Finite-Horizon POMDPs","volume":"65","author":"Walraven","year":"2019","journal-title":"J. Artif. Intell. Res."},{"key":"ref_18","unstructured":"Walraven, E., and Spaan, M. (2019). Planning under Uncertainty in Constrained and Partially Observable Environments. [Ph.D. Thesis, Delft University of Technology]."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1613\/jair.1659","article-title":"PERSEUS: Randomized point-based value iteration for POMDPs","volume":"24","author":"Spaan","year":"2005","journal-title":"J. Artif. Intell. Res."},{"key":"ref_20","unstructured":"Andrew, Y., and Jordan, M. (July, January 30). PEGASUS: A policy search method for large MDPs and POMDPs. Proceedings of the 16th Conference on Uncertainty in Artificial Intelligence, Stanford, CA, USA."},{"key":"ref_21","unstructured":"Meuleau, N., Kim, K., Kaelbling, L., and Cassandra, A. (2013, January 11\u201315). Solving POMDPs by searching the space of finite policies. Proceedings of the 15th Conference on Uncertainty in Artificial Intelligence, Bellevue, WA, USA."},{"key":"ref_22","unstructured":"Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2015). Prioritized experience replay, Google DeepMind. arXiv."},{"key":"ref_23","unstructured":"Dimitrakakis, C., and Ortner, R. (2019, December 20). Decision Making Under Uncertainty and Reinforcement Learning. Available online: http:\/\/www.cse.chalmers.se\/~chrdimi\/downloads\/book.pdf."},{"key":"ref_24","unstructured":"Osband, I., Russo, D., and van Roy, B. (2019, December 20). Efficient Reinforcement Learning via Posterior Sampling. Available online: https:\/\/papers.nips.cc\/paper\/5185-more-efficient-reinforcement-learning-via-posterior-sampling.pdf."},{"key":"ref_25","unstructured":"Grande, R., Walsh, T., and How, J. (2014, January 21\u201326). Sample efficient reinforcement learning with gaussian processes. Proceedings of the International Conference on Machine Learning, Beijing, China."},{"key":"ref_26","unstructured":"Agrawal, S., and Jia, R. (2017, January 4\u20139). Optimistic posterior sampling for reinforcement learning: Worst-case regret bounds. Proceedings of the Annual Conference on Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_27","unstructured":"Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A.Y. (July, January 28). Multimodal deep learning. Proceedings of the 28th International Conference on Machine Learning (ICML-11), Washington, DC, USA."},{"key":"ref_28","unstructured":"Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Hsu, D., Lee, W., Lim, Z., and Bai, A. (2015, January 7\u201311). PLEASE: Palm Leaf Search for POMDPs with Large Observation Spaces. Proceedings of the International Conference on Automated Planning and Scheduling, Israel.","DOI":"10.1609\/icaps.v25i1.13706"},{"key":"ref_30","unstructured":"NIST (2019, December 18). Computer Security Resource Center\u2014NATIONAL VULNERABILITY DATABASE (CVSS), Available online: https:\/\/nvd.nist.gov."},{"key":"ref_31","unstructured":"MITRE (2019, December 18). The MITRE Corporation\u2014Common Vulnerabilities and Exposures (CVE) Database. Available online: https:\/\/cve.mitre.org."},{"key":"ref_32","unstructured":"Lyu, D. (2019, January 3\u20137). Knowledge-Based Sequential Decision-Making Under Uncertainty. Proceedings of the 15th International Conference on Logic Programming and Non-monotonic Reasoning, LPNMR, Philadelphia, PA, USA."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/11\/1\/6\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:44:14Z","timestamp":1760190254000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/11\/1\/6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,12,20]]},"references-count":32,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2020,1]]}},"alternative-id":["info11010006"],"URL":"https:\/\/doi.org\/10.3390\/info11010006","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,12,20]]}}}