{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,28]],"date-time":"2026-07-28T21:53:57Z","timestamp":1785275637330,"version":"3.55.0"},"reference-count":267,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2022,2,17]],"date-time":"2022-02-17T00:00:00Z","timestamp":1645056000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100007660","name":"University of Antwerp","doi-asserted-by":"publisher","award":["x"],"award-info":[{"award-number":["x"]}],"id":[{"id":"10.13039\/501100007660","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["MAKE"],"abstract":"<jats:p>Reinforcement learning (RL) allows an agent to solve sequential decision-making problems by interacting with an environment in a trial-and-error fashion. When these environments are very complex, pure random exploration of possible solutions often fails, or is very sample inefficient, requiring an unreasonable amount of interaction with the environment. Hierarchical reinforcement learning (HRL) utilizes forms of temporal- and state-abstractions in order to tackle these challenges, while simultaneously paving the road for behavior reuse and increased interpretability of RL systems. In this survey paper we first introduce a selection of problem-specific approaches, which provided insight in how to utilize often handcrafted abstractions in specific task settings. We then introduce the Options framework, which provides a more generic approach, allowing abstractions to be discovered and learned semi-automatically. Afterwards we introduce the goal-conditional approach, which allows sub-behaviors to be embedded in a continuous space. In order to further advance the development of HRL agents, capable of simultaneously learning abstractions and how to use them, solely from interaction with complex high dimensional environments, we also identify a set of promising research directions.<\/jats:p>","DOI":"10.3390\/make4010009","type":"journal-article","created":{"date-parts":[[2022,2,17]],"date-time":"2022-02-17T20:25:45Z","timestamp":1645129545000},"page":"172-221","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":110,"title":["Hierarchical Reinforcement Learning: A Survey and Open Research Challenges"],"prefix":"10.3390","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6091-294X","authenticated-orcid":false,"given":"Matthias","family":"Hutsebaut-Buysse","sequence":"first","affiliation":[{"name":"Department of Computer Science, University of Antwerp\u2014imec, Sint-Pietersvliet 7, 2000 Antwerp, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4812-4841","authenticated-orcid":false,"given":"Kevin","family":"Mets","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Antwerp\u2014imec, Sint-Pietersvliet 7, 2000 Antwerp, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Steven","family":"Latr\u00e9","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Antwerp\u2014imec, Sint-Pietersvliet 7, 2000 Antwerp, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,2,17]]},"reference":[{"key":"ref_1","unstructured":"Sutton, R.S., and Barto, A.G. (2018). Reinforcement Learning: An Introduction, MIT Press."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"237","DOI":"10.1613\/jair.301","article-title":"Reinforcement Learning: A Survey","volume":"4","author":"Kaelbling","year":"1996","journal-title":"J. Artif. Intell. Res."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1016\/j.neunet.2014.09.003","article-title":"Deep Learning in Neural Networks: An Overview","volume":"61","author":"Schmidhuber","year":"2015","journal-title":"Neural Netw."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"219","DOI":"10.1561\/2200000071","article-title":"An introduction to deep reinforcement learning","volume":"11","author":"Henderson","year":"2018","journal-title":"Found. Trends\u00ae Mach. Learn."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"26","DOI":"10.1109\/MSP.2017.2743240","article-title":"Deep Reinforcement Learning: A Brief Survey","volume":"34","author":"Arulkumaran","year":"2017","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_7","unstructured":"Li, Y. (2017). Deep Reinforcement Learning: An Overview. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TG.2019.2896986","article-title":"Deep learning for video game playing","volume":"12","author":"Justesen","year":"2019","journal-title":"IEEE Trans. Games"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Luketina, J., Nardelli, N., Farquhar, G., Foerster, J., Andreas, J., Grefenstette, E., Whiteson, S., and Rockt\u00e4schel, T. (2019). A Survey of Reinforcement Learning Informed by Natural Language. arXiv.","DOI":"10.24963\/ijcai.2019\/880"},{"key":"ref_11","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal policy optimization algorithms. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D. (2018, January 2\u20137). Rainbow: Combining improvements in deep reinforcement learning. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11796"},{"key":"ref_13","unstructured":"Badia, A.P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, D., and Blundell, C. (2020). Agent57: Outperforming the Atari Human Benchmark. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1038\/nature16961","article-title":"Mastering the game of Go with deep neural networks and tree search","volume":"529","author":"Silver","year":"2016","journal-title":"Nature"},{"key":"ref_15","unstructured":"OpenAI (2019). Dota 2 with Large Scale Deep Reinforcement Learning. arXiv."},{"key":"ref_16","unstructured":"OpenAI (2021, December 09). OpenAI Five. Available online: https:\/\/blog.openai.com\/openai-five\/."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"350","DOI":"10.1038\/s41586-019-1724-z","article-title":"Grandmaster Level in StarCraft II Using Multi-Agent Reinforcement Learning","volume":"575","author":"Vinyals","year":"2019","journal-title":"Nature"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"859","DOI":"10.1126\/science.aau6249","article-title":"Human-level performance in 3D multiplayer games with population-based reinforcement learning","volume":"364","author":"Jaderberg","year":"2019","journal-title":"Science"},{"key":"ref_19","unstructured":"OpenAI (2019). Learning Dexterous In-Hand Manipulation. arXiv."},{"key":"ref_20","unstructured":"Mahmood, A.R., Korenkevych, D., Vasan, G., Ma, W., and Bergstra, J. (2018, January 29\u201331). Benchmarking Reinforcement Learning Algorithms on Real-World Robots. Proceedings of the Conference on Robot Learning, CoRL18, Z\u00fcrich, Switzerland."},{"key":"ref_21","unstructured":"Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., and Vanhoucke, V. (2018). QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation. arXiv."},{"key":"ref_22","first-page":"1334","article-title":"End-to-End Training of Deep Visuomotor Policies","volume":"17","author":"Levine","year":"2016","journal-title":"J. Mach. Learn. Res."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"421","DOI":"10.1177\/0278364917710318","article-title":"Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection","volume":"37","author":"Levine","year":"2016","journal-title":"Int. J. Robot. Res."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1238","DOI":"10.1177\/0278364913495721","article-title":"Reinforcement learning in robotics: A survey","volume":"32","author":"Kober","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"395","DOI":"10.1016\/j.csl.2009.07.001","article-title":"Evaluation of a hierarchical reinforcement learning spoken dialogue system","volume":"24","author":"Renals","year":"2010","journal-title":"Comput. Speech Lang."},{"key":"ref_26","unstructured":"Mandel, T., Liu, Y.E., Levine, S., Brunskill, E., and Popovic, Z. (2014, January 5\u20139). Offline Policy Evaluation Across Representations with Applications to Educational Games. Proceedings of the 2014 International Conference on Autonomous Agents and Multi-Agent Systems, Paris, France."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Pan, X., You, Y., Wang, Z., and Lu, C. (2017). Virtual to real reinforcement learning for autonomous driving. arXiv.","DOI":"10.5244\/C.31.11"},{"key":"ref_28","unstructured":"Ng, A.Y., Kim, H.J., Jordan, M.I., and Sastry, S. (2021, December 09). Autonomous Helicopter Flight via Reinforcement Learning. Advances in Neural Information Processing Systems 16 (NIPS 2003). Available online: https:\/\/proceedings.neurips.cc\/paper\/2003\/hash\/b427426b8acd2c2e53827970f2c2f526-Abstract.html."},{"key":"ref_29","unstructured":"Fran\u00e7ois-Lavet, V., Taralla, D., Ernst, D., and Fonteneau, R. (2016, January 3\u20134). Deep reinforcement learning solutions for energy microgrids management. Proceedings of the European Workshop on Reinforcement Learning (EWRL 2016), Barcelona, Spain."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Lin, K., Zhao, R., Xu, Z., and Zhou, J. (2018, January 19\u201323). Efficient large-scale fleet management via multi-agent deep reinforcement learning. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, London, UK.","DOI":"10.1145\/3219819.3219993"},{"key":"ref_31","unstructured":"Mirhoseini, A., Goldie, A., Pham, H., Steiner, B., Le, Q.V., and Dean, J. (May, January 30). A Hierarchical Model for Device Placement. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Mao, H., Alizadeh, M., Menache, I., and Kandula, S. (2016, January 9\u201310). Resource management with deep reinforcement learning. Proceedings of the 15th ACM Workshop on Hot Topics in Networks, Atlanta, GA, USA.","DOI":"10.1145\/3005745.3005750"},{"key":"ref_33","unstructured":"Zhang, J., Hao, B., Chen, B., Li, C., Chen, H., and Sun, J. (February, January 27). Hierarchical reinforcement learning for course recommendation in MOOCs. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_34","unstructured":"Zhu, H., Yu, J., Gupta, A., Shah, D., Hartikainen, K., Singh, A., Kumar, V., and Levine, S. (May, January 26). The Ingredients of Real World Robotic Reinforcement Learning. Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia."},{"key":"ref_35","unstructured":"Lee, S.Y., Choi, S., and Chung, S.Y. (2021, December 09). Sample-Efficient Deep Reinforcement Learning via Episodic Backward Update. Advances in Neural Information Processing Systems 32 (NeurIPS 2019). Available online: https:\/\/proceedings.neurips.cc\/paper\/2019\/hash\/e6d8545daa42d5ced125a4bf747b3688-Abstract.html."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Yu, Y. (2018, January 13\u201319). Towards Sample Efficient Reinforcement Learning. Proceedings of the IJCAI, Stockholm, Sweden.","DOI":"10.24963\/ijcai.2018\/820"},{"key":"ref_37","unstructured":"Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N. (2017, January 24\u201326). Sample Efficient Actor-Critic with Experience Replay. Proceedings of the ICLR17, Toulon, France."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Strehl, A.L., Li, L., Wiewiora, E., Langford, J., and Littman, M.L. (2006, January 7\u201310). PAC Model-Free Reinforcement Learning. Proceedings of the 23rd International Conference on Machine Learning\u2014ICML\u201906, Pittsburgh, PA, USA.","DOI":"10.1145\/1143844.1143955"},{"key":"ref_39","unstructured":"Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Man\u00e9, D. (2016). Concrete problems in AI safety. arXiv."},{"key":"ref_40","first-page":"1437","article-title":"A Comprehensive Survey on Safe Reinforcement Learning","volume":"16","year":"2015","journal-title":"J. Mach. Learn. Res."},{"key":"ref_41","first-page":"1","article-title":"Reinforcement Learning in Continuous Time and Space: A Stochastic Control Approach","volume":"21","author":"Wang","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"ref_42","unstructured":"Ramstedt, S., and Pal, C. (2019, January 8\u201314). Real-time reinforcement learning. Proceedings of the 33rd International Conference on Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"219","DOI":"10.1162\/089976600300015961","article-title":"Reinforcement learning in continuous time and space","volume":"12","author":"Doya","year":"2000","journal-title":"Neural Comput."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"124","DOI":"10.1016\/j.artint.2015.02.002","article-title":"HTN planning: Overview, comparison, and beyond","volume":"222","author":"Georgievski","year":"2015","journal-title":"Artif. Intell."},{"key":"ref_45","unstructured":"Dean, T., and Lin, S.H. (1995, January 20\u201325). Decomposition techniques for planning in stochastic domains. Proceedings of the Intemational Joint Conference on Artijicial Intelligence, Montreal, QC, Canada."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"115","DOI":"10.1016\/0004-3702(74)90026-5","article-title":"Planning in a Hierarchy of Abstraction Spaces","volume":"5","author":"Sacerdoti","year":"1973","journal-title":"Artif. Intell."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"e253","DOI":"10.1017\/S0140525X16001837","article-title":"Building machines that learn and think like people","volume":"40","author":"Lake","year":"2017","journal-title":"Behav. Brain Sci."},{"key":"ref_48","unstructured":"Ng, A.Y., Harada, D., and Russell, S.J. (1999, January 27\u201330). Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping. Proceedings of the International Conference on Machine Learning, Bled, Slovenia."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"230","DOI":"10.1109\/TAMD.2010.2056368","article-title":"Formal Theory of Creativity, Fun, and Intrinsic Motivation (1990\u20132010)","volume":"2","author":"Schmidhuber","year":"2010","journal-title":"IEEE Trans. Auton. Ment. Dev."},{"key":"ref_50","unstructured":"Burda, Y., Edwards, H., Storkey, A., and Klimov, O. (2018). Exploration by Random Network Distillation. arXiv."},{"key":"ref_51","unstructured":"Burda, Y., Edwards, H., Pathak, D., Storkey, A., Darrell, T., and Efros, A.A. (May, January 30). Large-Scale Study of Curiosity-Driven Learning. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Pathak, D., Agrawal, P., Efros, A.A., and Darrell, T. (2017, January 6\u201311). Curiosity-Driven Exploration by Self-Supervised Prediction. Proceedings of the International Conference on Machine Learning, Sydney, Australia.","DOI":"10.1109\/CVPRW.2017.70"},{"key":"ref_53","unstructured":"Bellemare, M.G., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. (2016). Unifying Count-Based Exploration and Intrinsic Motivation. arXiv."},{"key":"ref_54","unstructured":"Nachum, O., Tang, H., Lu, X., Gu, S., Lee, H., and Levine, S. (2019). Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?. arXiv."},{"key":"ref_55","unstructured":"Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018, January 10\u201315). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. Proceedings of the International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_56","unstructured":"Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R.E., and Levine, S. (2017, January 24\u201326). Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic. Proceedings of the International Conference on Learning Representations, Toulon, France."},{"key":"ref_57","first-page":"15","article-title":"An Introduction to Intertask Transfer for Reinforcement Learning","volume":"32","author":"Taylor","year":"2011","journal-title":"AI Mag."},{"key":"ref_58","unstructured":"Yosinski, J., Clune, J., Bengio, Y., and Lipson, H. (2014, January 8\u201313). How Transferable Are Features in Deep Neural Networks?. Proceedings of the 27th International Conference on Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_59","unstructured":"Bengio, Y. (2012, January 2). Deep learning of representations for unsupervised and transfer learning. Proceedings of the ICML Workshop on Unsupervised and Transfer Learning, Bellevue, DC, USA."},{"key":"ref_60","unstructured":"Hausman, K., Springenberg, J.T., Wang, Z., Heess, N., and Riedmiller, M. (May, January 30). Learning an Embedding Space for Transferable Robot Skills. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Tessler, C., Givony, S., Zahavy, T., Mankowitz, D.J., and Mannor, S. (2017, January 4\u20139). A Deep Hierarchical Approach to Lifelong Learning in Minecraft. Proceedings of the AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.10744"},{"key":"ref_62","unstructured":"Silver, D.L., Yang, Q., and Li, L. (2013, January 25\u201327). Lifelong Machine Learning Systems: Beyond Learning Algorithms. Proceedings of the AAAI Spring Symposium: Lifelong Machine Learning, Stanford, CA, USA."},{"key":"ref_63","unstructured":"Mott, A., Zoran, D., Chrzanowski, M., Wierstra, D., and Rezende, D.J. (2019). Towards Interpretable Reinforcement Learning Using Attention Augmented Agents. arXiv."},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"Gilpin, L.H., Bau, D., Yuan, B.Z., Bajwa, A., Specter, M.A., and Kagal, L. (2018, January 1\u20134). Explaining Explanations: An Overview of Interpretability of Machine Learning. Proceedings of the 2018 IEEE 5th International Conference on Data Science and Advanced Analytics (DSAA), Turin, Italy.","DOI":"10.1109\/DSAA.2018.00018"},{"key":"ref_65","unstructured":"Garnelo, M., Arulkumaran, K., and Shanahan, M. (2016). Towards Deep Symbolic Reinforcement Learning. arXiv."},{"key":"ref_66","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1023\/A:1022140919877","article-title":"Recent advances in hierarchical reinforcement learning","volume":"13","author":"Barto","year":"2003","journal-title":"Discret. Event Dyn. Syst."},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Puterman, M. (1994). Markov Decision Processes: Discrete Stochastic Dynamic Programming, John Wiley & Sons, Inc.","DOI":"10.1002\/9780470316887"},{"key":"ref_68","unstructured":"Bradtke, S.J., and Duff, M.O. (1994, January 1). Reinforcement learning methods for continuous-time Markov decision problems. Proceedings of the 7th International Conference on Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_69","unstructured":"Mahadevan, S., Das, T.K., and Gosavi, A. (1997, January 8\u201312). Self-Improving Factory Simulation using Continuous-time Average-Reward Reinforcement Learning. Proceedings of the International Conference on Machine Learning, Nashville, TN, USA."},{"key":"ref_70","unstructured":"Parr, R.E., and Russell, S. (1998). Hierarchical Control and Learning for Markov Decision Processes, University of California."},{"key":"ref_71","doi-asserted-by":"crossref","first-page":"716","DOI":"10.1073\/pnas.38.8.716","article-title":"On the theory of dynamic programming","volume":"38","author":"Bellman","year":"1952","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_72","unstructured":"Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R.Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M. (May, January 30). Parameter Space Noise for Exploration. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_73","unstructured":"Fortunato, M., Azar, M.G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., and Pietquin, O. (May, January 30). Noisy Networks for Exploration. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_74","unstructured":"Houthooft, R., Chen, X., Chen, X., Duan, Y., Schulman, J., Turck, F.D., and Abbeel, P. (2016). VIME: Variational Information Maximizing Exploration. arXiv."},{"key":"ref_75","doi-asserted-by":"crossref","unstructured":"Dabney, W., Ostrovski, G., Silver, D., and Munos, R. (2018, January 10\u201315). Implicit Quantile Networks for Distributional Reinforcement Learning. Proceedings of the International Conference on Machine Learning, Stockholm, Sweden.","DOI":"10.1609\/aaai.v32i1.11791"},{"key":"ref_76","unstructured":"Bellemare, M.G., Dabney, W., and Munos, R. (2017, January 6\u201311). A Distributional Perspective on Reinforcement Learning. Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_77","unstructured":"Harutyunyan, A., Dabney, W., Mesnard, T., Heess, N., Azar, M.G., Piot, B., van Hasselt, H., Singh, S., Wayne, G., and Precup, D. (2019). Hindsight Credit Assignment. arXiv."},{"key":"ref_78","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1007\/BF00992698","article-title":"Q-learning","volume":"8","author":"Watkins","year":"1992","journal-title":"Mach. Learn."},{"key":"ref_79","unstructured":"Rummery, G.A., and Niranjan, M. (1994). On-Line Q-Learning Using Connectionist Systems, University of Cambridge, Department of Engineering."},{"key":"ref_80","doi-asserted-by":"crossref","first-page":"229","DOI":"10.1007\/BF00992696","article-title":"Simple statistical gradient-following algorithms for connectionist reinforcement learning","volume":"8","author":"Williams","year":"1992","journal-title":"Mach. Learn."},{"key":"ref_81","unstructured":"Sutton, R.S., McAllester, D.A., Singh, S.P., and Mansour, Y. (2000, January 27\u201330). Policy Gradient Methods for Reinforcement Learning with Function Approximation. Proceedings of the Advances in Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_82","unstructured":"Konda, V.R., and Tsitsiklis, J.N. (2000, January 27\u201330). Actor-critic algorithms. Proceedings of the Advances in Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_83","doi-asserted-by":"crossref","unstructured":"Diuk, C., Cohen, A., and Littman, M.L. (2008, January 5\u20139). An object-oriented representation for efficient reinforcement learning. Proceedings of the International Conference on Machine Learning, Helsinki, Finland.","DOI":"10.1145\/1390156.1390187"},{"key":"ref_84","first-page":"503","article-title":"Tree-Based Batch Mode Reinforcement Learning","volume":"6","author":"Damien","year":"2005","journal-title":"J. Mach. Learn. Res."},{"key":"ref_85","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A. (2016). Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning. arXiv.","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"ref_86","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_87","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going Deeper with Convolutions. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_88","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20138). ImageNet Classification with Deep Convolutional Neural Networks. Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, CA, USA."},{"key":"ref_89","doi-asserted-by":"crossref","unstructured":"Karpathy, A., and Fei-Fei, L. (2015, January 8\u201310). Deep visual-semantic alignments for generating image descriptions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"ref_90","doi-asserted-by":"crossref","unstructured":"Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., and Parikh, D. (2015, January 7\u201313). VQA: Visual Question Answering. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.279"},{"key":"ref_91","unstructured":"Donahue, J., and Simonyan, K. (2019). Large Scale Adversarial Representation Learning. arXiv."},{"key":"ref_92","unstructured":"Radford, A., Metz, L., and Chintala, S. (2016). Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks. arXiv."},{"key":"ref_93","doi-asserted-by":"crossref","unstructured":"Gatys, L.A., Ecker, A.S., and Bethge, M. (2015). A Neural Algorithm of Artistic Style. arXiv.","DOI":"10.1167\/16.12.326"},{"key":"ref_94","unstructured":"van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. (2016). WaveNet: A Generative Model for Raw Audio. arXiv."},{"key":"ref_95","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 14\u201317). You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_96","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_97","doi-asserted-by":"crossref","unstructured":"Schneider, S., Baevski, A., Collobert, R., and Auli, M. (2019). wav2vec: Unsupervised Pre-training for Speech Recognition. arXiv.","DOI":"10.21437\/Interspeech.2019-1873"},{"key":"ref_98","doi-asserted-by":"crossref","unstructured":"Bahdanau, D., Chorowski, J., Serdyuk, D., Brakel, P., and Bengio, Y. (2016). End-to-End Attention-Based Large Vocabulary Speech Recognition. arXiv.","DOI":"10.1109\/ICASSP.2016.7472618"},{"key":"ref_99","doi-asserted-by":"crossref","unstructured":"Sainath, T.N., Mohamed, A.r., Kingsbury, B., and Ramabhadran, B. (2013, January 26\u201331). Deep Convolutional Neural Networks for LVCSR. Proceedings of the 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6639347"},{"key":"ref_100","doi-asserted-by":"crossref","unstructured":"Hinton, G., Deng, L., Yu, D., Dahl, G., Mohamed, A.R., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., and Sainath, T. (2012). Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups. IEEE Signal Process. Mag., 29.","DOI":"10.1109\/MSP.2012.2205597"},{"key":"ref_101","unstructured":"Sutskever, I., Vinyals, O., and Le, Q.V. (2014). Sequence to Sequence Learning with Neural Networks. arXiv."},{"key":"ref_102","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019, January 2\u20137). BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA."},{"key":"ref_103","unstructured":"Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2021, December 09). Language Models Are Unsupervised Multitask Learners. Available online: https:\/\/www.persagen.com\/files\/misc\/radford2019language.pdf."},{"key":"ref_104","doi-asserted-by":"crossref","first-page":"70","DOI":"10.1016\/j.compag.2018.02.016","article-title":"Deep learning in agriculture: A survey","volume":"147","author":"Kamilaris","year":"2018","journal-title":"Comput. Electron. Agric."},{"key":"ref_105","doi-asserted-by":"crossref","first-page":"60","DOI":"10.1016\/j.media.2017.07.005","article-title":"A survey on deep learning in medical image analysis","volume":"42","author":"Litjens","year":"2017","journal-title":"Med. Image Anal."},{"key":"ref_106","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_107","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long Short-Term Memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_108","doi-asserted-by":"crossref","unstructured":"Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014). Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation. arXiv.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_109","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural Machine Translation by Jointly Learning to Align and Translate. arXiv."},{"key":"ref_110","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention Is All You Need. arXiv."},{"key":"ref_111","doi-asserted-by":"crossref","first-page":"307","DOI":"10.1561\/2200000056","article-title":"An Introduction to Variational Autoencoders","volume":"12","author":"Kingma","year":"2019","journal-title":"Found. Trends Mach. Learn."},{"key":"ref_112","doi-asserted-by":"crossref","unstructured":"Rumelhart, D.E., Hinton, G.E., and Williams, R.J. (1985). Learning Internal Representations by Error Propagation, California Univ San Diego La Jolla Inst for Cognitive Science. Technical Report.","DOI":"10.21236\/ADA164453"},{"key":"ref_113","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014, January 8\u201313). Generative Adversarial Nets. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_114","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1109\/TNN.2008.2005605","article-title":"The graph neural network model","volume":"20","author":"Scarselli","year":"2008","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_115","unstructured":"Gori, M., Monfardini, G., and Scarselli, F. (August, January 30). A new model for learning in graph domains. Proceedings of the 2005 IEEE International Joint Conference on Neural Networks, Edinburgh, Scotland."},{"key":"ref_116","unstructured":"Van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J. (2018). Deep Reinforcement Learning and the Deadly Triad. arXiv."},{"key":"ref_117","unstructured":"Fu, J., Kumar, A., Soh, M., and Levine, S. (2019, January 9\u201315). Diagnosing Bottlenecks in Deep Q-learning Algorithms. Proceedings of the 36rd International Conference on Machine Learning, Long Beach, CA, USA."},{"key":"ref_118","doi-asserted-by":"crossref","first-page":"253","DOI":"10.1613\/jair.3912","article-title":"The arcade learning environment: An evaluation platform for general agents","volume":"47","author":"Bellemare","year":"2013","journal-title":"J. Artif. Intell. Res."},{"key":"ref_119","unstructured":"Tsitsiklis, J.N., and Van Roy, B. (1997, January 1\u20136). Analysis of temporal-diffference learning with function approximation. Proceedings of the Advances in Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_120","unstructured":"Zhang, S., and Sutton, R.S. (2017). A Deeper Look at Experience Replay. arXiv."},{"key":"ref_121","doi-asserted-by":"crossref","first-page":"293","DOI":"10.1007\/BF00992699","article-title":"Self-improving reactive agents based on reinforcement learning, planning and teaching","volume":"8","author":"Lin","year":"1992","journal-title":"Mach. Learn."},{"key":"ref_122","doi-asserted-by":"crossref","unstructured":"Van Hasselt, H., Guez, A., and Silver, D. (2016, January 12\u201317). Deep Reinforcement Learning with Double Q-Learning. Proceedings of the AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA.","DOI":"10.1609\/aaai.v30i1.10295"},{"key":"ref_123","unstructured":"Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2016). Prioritized Experience Replay. arXiv."},{"key":"ref_124","unstructured":"Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N. (2016, January 19\u201324). Dueling Network Architectures for Deep Reinforcement Learning. Proceedings of the 33rd International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_125","doi-asserted-by":"crossref","first-page":"9","DOI":"10.1007\/BF00115009","article-title":"Learning to predict by the methods of temporal differences","volume":"3","author":"Sutton","year":"1988","journal-title":"Mach. Learn."},{"key":"ref_126","doi-asserted-by":"crossref","unstructured":"Dabney, W., Rowland, M., Bellemare, M., and Munos, R. (2018, January 2\u20137). Distributional Reinforcement Learning with Quantile Regression. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11791"},{"key":"ref_127","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2016, January 2\u20134). Continuous control with deep reinforcement learning. Proceedings of the International Conference on Learning Representations, San Juan, Puerto Rico."},{"key":"ref_128","unstructured":"Schulman, J., Levine, S., Moritz, P., Jordan, M.I., and Abbeel, P. (2015, January 6\u201311). Trust Region Policy Optimization. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_129","unstructured":"Barreto, A., Borsa, D., Hou, S., Comanici, G., Ayg\u00fcn, E., Hamel, P., Toyama, D., Hunt, J., Mourad, S., and Silver, D. (2019, January 8\u201314). The Option Keyboard: Combining Skills in Reinforcement Learning. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_130","unstructured":"Li, L., Walsh, T.J., and Littman, M.L. (2006, January 4\u20136). Towards a Unified Theory of State Abstraction for MDPs. Proceedings of the Ninth International Symposium on Artificial Intelligence and Mathematics, Fort Lauderdale, FL, USA."},{"key":"ref_131","unstructured":"Konidaris, G., and Barto, A.G. (2009, January 11\u201317). Efficient Skill Learning using Abstraction Selection. Proceedings of the International Joint Conferences on Artificial Intelligence, Pasadena, CA, USA."},{"key":"ref_132","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/0890-5401(91)90030-6","article-title":"Bisimulation through probabilistic testing","volume":"94","author":"Larsen","year":"1991","journal-title":"Inf. Comput."},{"key":"ref_133","unstructured":"Ferns, N., Panangaden, P., and Precup, D. (2004, January 9\u201311). Metrics for Finite Markov Decision Processes. Proceedings of the Twentieth Conference on Uncertainty in Artificial Intelligence, Banff, AB, Canada."},{"key":"ref_134","unstructured":"Castro, P.S., and Precup, D. (2010, January 11\u201315). Using bisimulation for policy transfer in MDPs. Proceedings of the Twenty-Fourth AAAI Conference on Artificial Intelligence, Atlanta, GA, USA."},{"key":"ref_135","doi-asserted-by":"crossref","first-page":"140","DOI":"10.1007\/978-3-642-29946-9_16","article-title":"Automatic Construction of Temporally Extended Actions for MDPs Using Bisimulation Metrics","volume":"Volume 7188","author":"Hutchison","year":"2012","journal-title":"Recent Advances in Reinforcement Learning"},{"key":"ref_136","unstructured":"Anders, J., and Andrew, G.B. (2000, January 27\u201330). Automated State Abstraction for Options using the U-Tree Algorithm. Proceedings of the Advances in Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_137","doi-asserted-by":"crossref","first-page":"181","DOI":"10.1016\/S0004-3702(99)00052-1","article-title":"Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning","volume":"112","author":"Sutton","year":"1999","journal-title":"Artif. Intell."},{"key":"ref_138","unstructured":"McGovern, A., Sutton, R.S., and Fagg, A.H. (1997, January 19\u201321). Roles of macro-actions in accelerating reinforcement learning. Proceedings of the Grace Hopper Celebration of Women in Computing, San Jose, CA, USA."},{"key":"ref_139","unstructured":"Jong, N.K., Hester, T., and Stone, P. (2008, January 12\u201316). The utility of temporal abstraction in reinforcement learning. Proceedings of the International Conference on Autonomous Agents and Multiagent Systems, Estoril, Portugal."},{"key":"ref_140","doi-asserted-by":"crossref","unstructured":"Bacon, P.L., Harb, J., and Precup, D. (2017, January 4\u20139). The option-critic architecture. Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.10916"},{"key":"ref_141","unstructured":"Heess, N., Wayne, G., Tassa, Y., Lillicrap, T., Riedmiller, M., and Silver, D. (2016). Learning and transfer of modulated locomotor controllers. arXiv."},{"key":"ref_142","unstructured":"Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015, January 6\u201311). Universal value function approximators. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_143","unstructured":"Johnson, M., Hofmann, K., Hutton, T., and Bignell, D. (2016, January 9\u201315). The Malmo Platform for Artificial Intelligence Experimentation. Proceedings of the International Joint Conference on Artificial Intelligence, New York, NY, USA."},{"key":"ref_144","doi-asserted-by":"crossref","unstructured":"Bengio, Y., Louradour, J., Collobert, R., and Weston, J. (2009, January 14\u201318). Curriculum learning. Proceedings of the 26th Annual International Conference on Machine Learning, Montreal, QC, Canada.","DOI":"10.1145\/1553374.1553380"},{"key":"ref_145","doi-asserted-by":"crossref","first-page":"323","DOI":"10.1007\/BF00992700","article-title":"Transfer of learning by composing solutions of elemental sequential tasks","volume":"8","author":"Singh","year":"1992","journal-title":"Machine Learning"},{"key":"ref_146","doi-asserted-by":"crossref","first-page":"311","DOI":"10.1016\/0004-3702(92)90058-6","article-title":"Automatic programming of behavior-based robots using reinforcement learning","volume":"55","author":"Mahadevan","year":"1992","journal-title":"Artif. Intell."},{"key":"ref_147","unstructured":"Maes, P., and Brooks, R.A. (August, January 29). Learning to Coordinate Behaviors. Proceedings of the National Conference on Artificial Intelligence, Boston, MA, USA."},{"key":"ref_148","unstructured":"Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S. (2019). Diversity Is All You Need: Learning Skills without a Reward Function. arXiv."},{"key":"ref_149","unstructured":"Nachum, O., Gu, S.S., Lee, H., and Levine, S. (2018, January 3\u20138). Data-efficient hierarchical reinforcement learning. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_150","unstructured":"Vezhnevets, A.S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K. (2017, January 6\u201311). Feudal networks for hierarchical reinforcement learning. Proceedings of the International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_151","unstructured":"Jinnai, Y., Abel, D., Hershkowitz, D.E., Littman, M.L., and Konidaris, G. (2019, January 10\u201315). Finding Options That Minimize Planning Time. Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA."},{"key":"ref_152","unstructured":"Abel, D., Umbanhowar, N., Khetarpal, K., Arumugam, D., Precup, D., and Littman, M.L. (2020, January 26\u201328). Value Preserving State-Action Abstractions. Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, Palermo, Italy."},{"key":"ref_153","unstructured":"Machado, M.C., Bellemare, M.G., and Bowling, M.H. (2017, January 6\u201311). A Laplacian Framework for Option Discovery in Reinforcement Learning. Proceedings of the International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_154","unstructured":"McGovern, A., and Barto, A.G. (July, January 28). Automatic discovery of subgoals in reinforcement learning using diverse density. Proceedings of the International Conference on Machine Learning, Williamstown, MA, USA."},{"key":"ref_155","unstructured":"Haarnoja, T., Hartikainen, K., Abbeel, P., and Levine, S. (2018, January 10\u201315). Latent Space Policies for Hierarchical Reinforcement Learning. Proceedings of the International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_156","unstructured":"Kulkarni, T.D., Narasimhan, K., Saeedi, A., and Tenenbaum, J.B. (2016, January 5\u201310). Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation. Proceedings of the Advances in Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_157","doi-asserted-by":"crossref","unstructured":"Peng, X.B., Abbeel, P., Levine, S., and van de Panne, M. (2018). DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills. ACM Trans. Graph., 37.","DOI":"10.1145\/3197517.3201311"},{"key":"ref_158","unstructured":"Ross, S., Gordon, G.J., and Bagnell, J.A. (2011, January 11\u201313). A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, Lauderdale, FL, USA."},{"key":"ref_159","doi-asserted-by":"crossref","unstructured":"Argall, B.D., Chernova, S., Veloso, M., and Browning, B. (2009). A Survey of Robot Learning from Demonstration. Robot. Auton. Syst., 57.","DOI":"10.1016\/j.robot.2008.10.024"},{"key":"ref_160","first-page":"103","article-title":"A framework for behavioural cloning","volume":"15","author":"Bain","year":"1999","journal-title":"Mach. Intell."},{"key":"ref_161","unstructured":"Le, H.M., Jiang, N., Agarwal, A., Dud\u00edk, M., Yue, Y., and Daum\u00e9, H. (2018, January 10\u201315). Hierarchical Imitation and Reinforcement Learning. Proceedings of the International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_162","unstructured":"Fox, R., Krishnan, S., Stoica, I., and Goldberg, K. (2017). Multi-level discovery of deep options. arXiv."},{"key":"ref_163","unstructured":"Krishnan, S., Fox, R., Stoica, I., and Goldberg, K.Y. (2017, January 13\u201315). DDCO: Discovery of Deep Continuous Options for Robot Learning from Demonstrations. Proceedings of the Conference on Robot Learning, Mountain View, CA, USA."},{"key":"ref_164","unstructured":"Ho, J., and Ermon, S. (2016, January 5\u201310). Generative Adversarial Imitation Learning. Proceedings of the Advances in Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_165","unstructured":"Ng, A.Y., and Russell, S.J. (July, January 29). Algorithms for Inverse Reinforcement Learning. Proceedings of the International Conference on Machine Learning, Stanford, CA, USA."},{"key":"ref_166","unstructured":"Pan, X., Ohn-Bar, E., Rhinehart, N., Xu, Y., Shen, Y., and Kitani, K.M. (2018). Human-Interactive Subgoal Supervision for Efficient Inverse Reinforcement Learning. arXiv."},{"key":"ref_167","unstructured":"Krishnan, S., Garg, A., Liaw, R., Miller, L., Pokorny, F.T., and Goldberg, K. (2016). Hirl: Hierarchical inverse reinforcement learning for long-horizon tasks with delayed rewards. arXiv."},{"key":"ref_168","unstructured":"Dayan, P., and Hinton, G.E. (December, January 30). Feudal reinforcement learning. Proceedings of the Advances in Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_169","unstructured":"Kumar, A., Swersky, K., and Hinton, G. (2017, January 9). Feudal Learning for Large Discrete Action Spaces with Recursive Substructure. Proceedings of the NIPS Workshop Hierarchical Reinforcement Learning, Long Beach, CA, USA."},{"key":"ref_170","unstructured":"Parr, R., and Russell, S.J. (1997, January 1\u20136). Reinforcement learning with hierarchies of machines. Proceedings of the Advances in Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_171","unstructured":"Ghavamzadeh, M., and Mahadevan, S. (July, January 28). Continuous-time hierarchical reinforcement learning. Proceedings of the International Conference on Machine Learning, Williamstown, MA, USA."},{"key":"ref_172","doi-asserted-by":"crossref","first-page":"227","DOI":"10.1613\/jair.639","article-title":"Hierarchical reinforcement learning with the MAXQ value function decomposition","volume":"13","author":"Dietterich","year":"2000","journal-title":"J. Artif. Intell. Res."},{"key":"ref_173","unstructured":"Hengst, B. (2002, January 8\u201312). Discovering Hierarchy in Reinforcement Learning with HEXQ. Proceedings of the International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_174","unstructured":"Jonsson, A., and Barto, A.G. (2006). Causal Graph Based Decomposition of Factored MDPs. J. Mach. Learn. Res., 7."},{"key":"ref_175","doi-asserted-by":"crossref","unstructured":"Mehta, N., Ray, S., Tadepalli, P., and Dietterich, T.G. (2008, January 5\u20139). Automatic discovery and transfer of MAXQ hierarchies. Proceedings of the International Conference on Machine Learning, Helsinki, Finland.","DOI":"10.1145\/1390156.1390238"},{"key":"ref_176","unstructured":"Mann, T.A., and Mannor, S. (2014, January 21\u201326). Scaling Up Approximate Value Iteration with Options: Better Policies with Fewer Iterations. Proceedings of the International Conference on Machine Learning, Beijing, China."},{"key":"ref_177","unstructured":"Silver, D., and Ciosek, K. (July, January 26). Compositional Planning Using Optimal Option Models. Proceedings of the International Conference on Machine Learning, Edinburgh, Scotland."},{"key":"ref_178","unstructured":"Precup, D., and Sutton, R.S. (1997, January 1\u20136). Multi-time Models for Temporally Abstract Planning. Proceedings of the Advances in Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_179","unstructured":"Brunskill, E., and Li, L. (2014, January 21\u201326). Pac-inspired option discovery in lifelong reinforcement learning. Proceedings of the International Conference on Machine Learning, Beijing, China."},{"key":"ref_180","unstructured":"Guo, Z., Thomas, P.S., and Brunskill, E. (2017, January 4\u20139). Using options and covariance testing for long horizon off-policy policy evaluation. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_181","unstructured":"Konidaris, G., and Barto, A.G. (2009, January 7\u20139). Skill Discovery in Continuous Reinforcement Learning Domains using Skill Chaining. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_182","doi-asserted-by":"crossref","unstructured":"Neumann, G., Maass, W., and Peters, J. (2009, January 14\u201318). Learning complex motions by sequencing simpler motion templates. Proceedings of the International Conference on Machine Learning, Montreal, QC, Canada.","DOI":"10.1145\/1553374.1553471"},{"key":"ref_183","doi-asserted-by":"crossref","unstructured":"\u015eim\u015fek, \u00d6., and Barto, A.G. (2004, January 5\u20138). Using relative novelty to identify useful temporal abstractions in reinforcement learning. Proceedings of the Twenty-First International Conference on Machine Learning, Banff, AB, Canada.","DOI":"10.1145\/1015330.1015353"},{"key":"ref_184","unstructured":"Comanici, G., and Precup, D. (2010, January 10\u201314). Optimal Policy Switching Algorithms for Reinforcement Learning. Proceedings of the 9th International Conference on Autonomous Agents and Multiagent Systems, Toronto, ON, Canada."},{"key":"ref_185","unstructured":"Sutton, R.S., Precup, D., and Singh, S.P. (1998, January 24\u201327). Intra-Option Learning about Temporally Abstract Actions. Proceedings of the International Conference on Machine Learning, Madison, WI, USA."},{"key":"ref_186","unstructured":"Mankowitz, D.J., Mann, T.A., and Mannor, S. (2014, January 21\u201326). Time regularized interrupting options. Proceedings of the International Conference on Machine Learning, Beijing, China."},{"key":"ref_187","doi-asserted-by":"crossref","unstructured":"Harb, J., Bacon, P.L., Klissarov, M., and Precup, D. (2018, January 2\u20137). When waiting is not an option: Learning options with a deliberation cost. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11831"},{"key":"ref_188","doi-asserted-by":"crossref","unstructured":"Kaelbling, L.P. (1993, January 27\u201329). Hierarchical Learning in Stochastic Domains: Preliminary Results. Proceedings of the International Conference on Machine Learning, Amherst, MA, USA.","DOI":"10.1016\/B978-1-55860-307-3.50028-9"},{"key":"ref_189","doi-asserted-by":"crossref","unstructured":"Digney, B.L. (1998, January 17\u201321). Learning hierarchical control structures for multiple tasks and changing environments. Proceedings of the Fifth International Conference on Simulation of Adaptive Behavior on from Animals to Animats, Zurich, Switzerland.","DOI":"10.7551\/mitpress\/3119.003.0050"},{"key":"ref_190","doi-asserted-by":"crossref","unstructured":"Stolle, M., and Precup, D. (2002, January 2\u20134). Learning options in reinforcement learning. Proceedings of the International Symposium on Abstraction, Reformulation, and Approximation, Kananaskis, AB, Canada.","DOI":"10.1007\/3-540-45622-8_16"},{"key":"ref_191","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1016\/S0004-3702(96)00034-3","article-title":"Solving the multiple instance problem with axis-parallel rectangles","volume":"89","author":"Dietterich","year":"1997","journal-title":"Artif. Intell."},{"key":"ref_192","unstructured":"Maron, O., and Lozano-P\u00e9rez, T. (1997, January 1\u20136). A framework for multiple-instance learning. Proceedings of the Advances in Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_193","unstructured":"Kulkarni, T.D., Saeedi, A., Gautam, S., and Gershman, S.J. (2016). Deep successor reinforcement learning. arXiv."},{"key":"ref_194","doi-asserted-by":"crossref","first-page":"613","DOI":"10.1162\/neco.1993.5.4.613","article-title":"Improving Generalization for Temporal Difference Learning: The Successor Representation","volume":"5","author":"Dayan","year":"1993","journal-title":"Neural Comput."},{"key":"ref_195","unstructured":"Goel, S., and Huber, M. (2003, January 11\u201315). Subgoal Discovery for Hierarchical Reinforcement Learning Using Learned Policies. Proceedings of the FLAIRS Conference, St. Augustine, FL, USA."},{"key":"ref_196","doi-asserted-by":"crossref","unstructured":"Menache, I., Mannor, S., and Shimkin, N. (2002, January 19\u201323). Q-Cut\u2014Dynamic Discovery of Sub-goals in Reinforcement Learning. Proceedings of the European Conference on Machine Learning, Helsinki, Finland.","DOI":"10.1007\/3-540-36755-1_25"},{"key":"ref_197","unstructured":"Waissi, G.R. (1994). Network Flows: Theory, Algorithms, and Applications, Prentice-Hall."},{"key":"ref_198","doi-asserted-by":"crossref","unstructured":"\u015eim\u015fek, \u00d6., Wolfe, A.P., and Barto, A.G. (2005, January 7\u201311). Identifying useful subgoals in reinforcement learning by local graph partitioning. Proceedings of the 22nd International Conference on Machine Learning, Bonn, Germany.","DOI":"10.1145\/1102351.1102454"},{"key":"ref_199","doi-asserted-by":"crossref","unstructured":"Mahadevan, S. (2005, January 7\u201311). Proto-value functions: Developmental reinforcement learning. Proceedings of the International Conference on Machine Learning, Bonn, Germany.","DOI":"10.1145\/1102351.1102421"},{"key":"ref_200","unstructured":"Machado, M.C., Rosenbaum, C., Guo, X., Liu, M., Tesauro, G., and Campbell, M. (May, January 30). Eigenoption Discovery through the Deep Successor Representation. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_201","unstructured":"Lakshminarayanan, A.S., Krishnamurthy, R., Kumar, P., and Ravindran, B. (2016). Option discovery in hierarchical reinforcement learning using spatio-temporal clustering. arXiv."},{"key":"ref_202","unstructured":"Da Silva, B.C., Konidaris, G., and Barto, A.G. (July, January 26). Learning Parameterized Skills. Proceedings of the International Conference on Machine Learning, Edinburgh, Scotland."},{"key":"ref_203","unstructured":"Vezhnevets, A., Mnih, V., Agapiou, J., Osindero, S., Graves, A., Vinyals, O., and Kavukcuoglu, K. (2016, January 5\u201310). Strategic Attentive Writer for Learning Macro-Actions. Proceedings of the Advances in Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_204","unstructured":"Gregor, K., Danihelka, I., Graves, A., Rezende, D.J., and Wierstra, D. (2015, January 6\u201311). DRAW: A Recurrent Neural Network For Image Generation. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_205","doi-asserted-by":"crossref","unstructured":"Sharma, S., Lakshminarayanan, A.S., and Ravindran, B. (2017, January 24\u201326). Learning to Repeat: Fine Grained Action Repetition for Deep Reinforcement Learning. Proceedings of the International Conference on Learning Representations, Toulon, France.","DOI":"10.1609\/aaai.v31i1.10918"},{"key":"ref_206","unstructured":"Rusu, A.A., Colmenarejo, S.G., Gulcehre, C., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., and Hadsell, R. (2015). Policy distillation. arXiv."},{"key":"ref_207","unstructured":"Daniel, C., Neumann, G., and Peters, J. (2012, January 21\u201323). Hierarchical relative entropy policy search. Proceedings of the Artificial Intelligence and Statistics, La Palma, Canary Islands."},{"key":"ref_208","unstructured":"Peters, J., M\u00fclling, K., and Altun, Y. (2010, January 11\u201315). Relative Entropy Policy Search. Proceedings of the Twenty-Fourth AAAI Conference on Artificial Intelligence, Atlanta, GA, USA."},{"key":"ref_209","doi-asserted-by":"crossref","first-page":"337","DOI":"10.1007\/s10994-016-5580-x","article-title":"Probabilistic inference for determining options in reinforcement learning","volume":"104","author":"Daniel","year":"2016","journal-title":"Mach. Learn."},{"key":"ref_210","unstructured":"Sutton, R.S. (1984). Temporal Credit Assignment in Reinforcement Learning. [Ph.D. Thesis, University of Massachusetts Amherst]."},{"key":"ref_211","unstructured":"Baird, L.C. (2021, December 09). Advantage Updating, Wright Lab. Technical Report WL-TR-93-1l46. Available online: https:\/\/citeseerx.ist.psu.edu\/viewdoc\/download?doi=10.1.1.49.4950&rep=rep1&type=pdf."},{"key":"ref_212","unstructured":"Klissarov, M., Bacon, P.L., Harb, J., and Precup, D. (2017, January 4\u20139). Learnings Options End-to-End for Continuous Action Tasks. Proceedings of the Hierarchical Reinforcement Learning Workshop, Long Beach, CA, USA."},{"key":"ref_213","unstructured":"Harutyunyan, A., Dabney, W., Borsa, D., Heess, N., Munos, R., and Precup, D. (2019, January 16\u201318). The Termination Critic. Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics, Naha, Japan."},{"key":"ref_214","unstructured":"Wang, J.X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J.Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M. (2017). Learning to Reinforcement Learn. arXiv."},{"key":"ref_215","unstructured":"Finn, C., Abbeel, P., and Levine, S. (2017, January 6\u201311). Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. Proceedings of the International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_216","unstructured":"Duan, Y., Schulman, J., Chen, X., Bartlett, P.L., Sutskever, I., and Abbeel, P. (2016). RL$2^$: Fast Reinforcement Learning via Slow Reinforcement Learning. arXiv."},{"key":"ref_217","unstructured":"Frans, K., Ho, J., Chen, X., Abbeel, P., and Schulman, J. (May, January 30). Meta Learning Shared Hierarchies. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_218","unstructured":"Li, A.C., Florensa, C., Clavera, I., and Abbeel, P. (2020). Sub-Policy Adaptation for Hierarchical Reinforcement Learning. arXiv."},{"key":"ref_219","unstructured":"Sohn, S., Woo, H., Choi, J., and Lee, H. (2020). Meta Reinforcement Learning with Autonomous Inference of Subtask Dependencies. arXiv."},{"key":"ref_220","unstructured":"Sutton, R.S., Modayil, J., Delp, M., Degris, T., Pilarski, P.M., White, A., and Precup, D. (2011, January 2\u20136). Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction. Proceedings of the International Conference on Autonomous Agents and Multiagent Systems, Taipei, Taiwan."},{"key":"ref_221","unstructured":"Bengio, E., Thomas, V., Pineau, J., Precup, D., and Bengio, Y. (2017). Independently Controllable Features. arXiv."},{"key":"ref_222","doi-asserted-by":"crossref","first-page":"504","DOI":"10.1126\/science.1127647","article-title":"Reducing the Dimensionality of Data with Neural Networks","volume":"313","author":"Hinton","year":"2006","journal-title":"Science"},{"key":"ref_223","unstructured":"Levy, A., Platt, R., and Saenko, K. (2019, January 6\u20139). Hierarchical Reinforcement Learning with Hindsight. Proceedings of the International Conference on Learning Representations, New Orleans, LA, USA."},{"key":"ref_224","unstructured":"Van Seijen, H., Fatemi, M., Romoff, J., Laroche, R., Barnes, T., and Tsang, J. (2017). Hybrid reward architecture for reinforcement learning. arXiv."},{"key":"ref_225","unstructured":"Jaderberg, M., Mnih, V., Czarnecki, W.M., Schaul, T., Leibo, J.Z., Silver, D., and Kavukcuoglu, K. (2017, January 24\u201326). Reinforcement learning with unsupervised auxiliary tasks. Proceedings of the International Conference on Learning Representations, Toulon, France."},{"key":"ref_226","unstructured":"Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T.P., Harley, T., Silver, D., and Kavukcuoglu, K. (2016, January 19\u201324). Asynchronous Methods for Deep Reinforcement Learning. Proceedings of the International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_227","unstructured":"Riedmiller, M.A., Hafner, R., Lampe, T., Neunert, M., Degrave, J., de Wiele, T.V., Mnih, V., Heess, N., and Springenberg, J.T. (2018, January 10\u201315). Learning by Playing\u2014Solving Sparse Reward Tasks from Scratch. Proceedings of the International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_228","unstructured":"Ziebart, B.D., Maas, A.L., Bagnell, J.A., and Dey, A.K. (2008, January 13\u201317). Maximum Entropy Inverse Reinforcement Learning. Proceedings of the AAAI, Chicago, IL, USA."},{"key":"ref_229","unstructured":"Todorov, E. (2006, January 4\u20135). Linearly-solvable Markov decision problems. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_230","unstructured":"Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. (2017, January 6\u201311). Reinforcement Learning with Deep Energy-Based Policies. Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_231","unstructured":"Nachum, O., Gu, S., Lee, H., and Levine, S. (2018, January 3\u20138). Near-Optimal Representation Learning for Hierarchical Reinforcement Learning. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_232","unstructured":"Sukhbaatar, S., Denton, E., Szlam, A., and Fergus, R. (2018). Learning Goal Embeddings via Self-Play for Hierarchical Reinforcement Learning. arXiv."},{"key":"ref_233","unstructured":"Florensa, C., Duan, Y., and Abbeel, P. (2017, January 24\u201326). Stochastic neural networks for hierarchical reinforcement learning. Proceedings of the International Conference on Learning Representations, Toulon, France."},{"key":"ref_234","unstructured":"Tang, Y., and Salakhutdinov, R.R. (2021, December 09). Learning Stochastic Feedforward Neural Networks. Advances in Neural Information Processing Systems 26. Available online: https:\/\/proceedings.neurips.cc\/paper\/2013\/hash\/d81f9c1be2e08964bf9f24b15f0e4900-Abstract.html."},{"key":"ref_235","unstructured":"Radford, N. (1990). Learning Stochastic Feedforward Networks, Department of Computer Science University of Toronto."},{"key":"ref_236","unstructured":"Gregor, K., Rezende, D.J., and Wierstra, D. (2016). Variational Intrinsic Control. arXiv."},{"key":"ref_237","doi-asserted-by":"crossref","unstructured":"Salge, C., Glackin, C., and Polani, D. (2014). Empowerment\u2014An introduction. Guided Self-Organization: Inception, Springer.","DOI":"10.1007\/978-3-642-53734-9_4"},{"key":"ref_238","unstructured":"Achiam, J., Edwards, H., Amodei, D., and Abbeel, P. (2018). Variational option discovery algorithms. arXiv."},{"key":"ref_239","unstructured":"Kingma, D.P., and Welling, M. (2014, January 14\u201316). Auto-Encoding Variational Bayes. Proceedings of the International Conference on Learning Representations, Banff, AB, Canada."},{"key":"ref_240","unstructured":"Dibya, G., Abhishek, G., and Sergey, L. (2019, January 6\u20139). Learning Actionable Representations with Goal-Conditioned Policies. Proceedings of the International Conference on Learning Representations, New Orleans, LA, USA."},{"key":"ref_241","unstructured":"Menashe, J., and Stone, P. (2019, January 13\u201317). Escape Room: A Configurable Testbed for Hierarchical Reinforcement Learning. Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems, Montreal, QC, Canada."},{"key":"ref_242","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TSMC.1983.6313077","article-title":"Neuronlike adaptive elements that can solve difficult learning control problems","volume":"5","author":"Barto","year":"1983","journal-title":"IEEE Trans. Syst. Man, Cybern."},{"key":"ref_243","unstructured":"Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J. (2019, January 10\u201315). Quantifying Generalization in Reinforcement Learning. Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA."},{"key":"ref_244","doi-asserted-by":"crossref","unstructured":"Kempka, M., Wydmuch, M., Runc, G., Toczek, J., and Ja\u015bkowski, W. Vizdoom: A doom-based ai research platform for visual reinforcement learning. Proceedings of the 2016 IEEE Conference on Computational Intelligence and Games (CIG).","DOI":"10.1109\/CIG.2016.7860433"},{"key":"ref_245","unstructured":"Beattie, C., Leibo, J.Z., Teplyashin, D., Ward, T., Wainwright, M., K\u00fcttler, H., Lefrancq, A., Green, S., Vald\u00e9s, V., and Sadik, A. (2016). Deepmind lab. arXiv."},{"key":"ref_246","unstructured":"Vinyals, O., Ewalds, T., Bartunov, S., Georgiev, P., Vezhnevets, A.S., Yeo, M., Makhzani, A., K\u00fcttler, H., Agapiou, J., and Schrittwieser, J. (2017). Starcraft ii: A new challenge for reinforcement learning. arXiv."},{"key":"ref_247","doi-asserted-by":"crossref","unstructured":"Santiago, O., Gabriel, S., Alberto, U., Florian, R., David, C., and Mike, P. (2013). A survey of real-time strategy game AI research and competition in StarCraft. IEEE Trans. Comput. Intell. AI Games, 5.","DOI":"10.1109\/TCIAIG.2013.2286295"},{"key":"ref_248","doi-asserted-by":"crossref","unstructured":"Todorov, E., Erez, T., and Tassa, Y. (2012, January 7\u201312). MuJoCo: A physics engine for model-based control. Proceedings of the International Conference on Intelligent Robots and Systems, Vilamoura, Portugal.","DOI":"10.1109\/IROS.2012.6386109"},{"key":"ref_249","unstructured":"Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., and Lefrancq, A. (2018). DeepMind Control Suite. arXiv."},{"key":"ref_250","unstructured":"Coumans, E., and Bai, Y. (2021, December 09). PyBullet, a Python Module for Physics Simulation for Games, Robotics and Machine Learning. Available online: http:\/\/pybullet.org."},{"key":"ref_251","unstructured":"Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P. (2016, January 19\u201324). Benchmarking deep reinforcement learning for continuous control. Proceedings of the International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_252","unstructured":"Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A. (May, January 26). Implementation Matters in Deep RL: A Case Study on PPO and TRPO. Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia."},{"key":"ref_253","doi-asserted-by":"crossref","unstructured":"Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D. (2017). Deep Reinforcement Learning That Matters. arXiv.","DOI":"10.1609\/aaai.v32i1.11694"},{"key":"ref_254","doi-asserted-by":"crossref","unstructured":"Zeiler, M.D., and Fergus, R. (2014). Visualizing and understanding convolutional networks. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10590-1_53"},{"key":"ref_255","doi-asserted-by":"crossref","first-page":"3521","DOI":"10.1073\/pnas.1611835114","article-title":"Overcoming catastrophic forgetting in neural networks","volume":"114","author":"Kirkpatrick","year":"2017","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_256","unstructured":"Savinov, N., Raichuk, A., Vincent, D., Marinier, R., Pollefeys, M., Lillicrap, T., and Gelly, S. (2019, January 6\u20139). Episodic Curiosity through Reachability. Proceedings of the International Conference on Learning Representations, New Orleans, LA, USA."},{"key":"ref_257","unstructured":"Jinnai, Y., Park, J.W., Machado, M.C., and Konidaris, G. (May, January 26). Exploration in Reinforcement Learning with Deep Covering Options. Proceedings of the International Conference on Learning Representations, ICLR2020, Addis Ababa, Ethiopia."},{"key":"ref_258","unstructured":"Greydanus, S., Koul, A., Dodge, J., and Fern, A. (2018, January 10\u201315). Visualizing and Understanding Atari Agents. Proceedings of the International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_259","unstructured":"Atrey, A., Clary, K., and Jensen, D. (May, January 26). Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement Learning. Proceedings of the ICLR2020, Addis Ababa, Ethiopia."},{"key":"ref_260","unstructured":"Rupprecht, C., Ibrahim, C., and Pal, C.J. (May, January 26). Finding and Visualizing Weaknesses of Deep Reinforcement Learning Agents. Proceedings of the ICLR20, Addis Ababa, Ethiopia."},{"key":"ref_261","unstructured":"Shu, T., Xiong, C., and Socher, R. (May, January 30). Hierarchical and Interpretable Skill Acquisition in Multi-Task Reinforcement Learning. Proceedings of the International Conference on Learning Representations 2018, Vancouver, BC, Canada."},{"key":"ref_262","unstructured":"Osband, I., Doron, Y., Hessel, M., Aslanides, J., Sezener, E., Saraiva, A., McKinney, K., Lattimore, T., Szepezvari, C., and Singh, S. (2019). Behaviour Suite for Reinforcement Learning. arXiv."},{"key":"ref_263","unstructured":"Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., Wu, Y., and Zhokhov, P. (2021, December 09). OpenAI Baselines. Available online: https:\/\/github.com\/openai\/baselines."},{"key":"ref_264","unstructured":"Jiang, Y., Gu, S., Murphy, K., and Finn, C. (2019, January 8\u201314). Language as an Abstraction for Hierarchical Deep Reinforcement Learning. Proceedings of the NeurIPS19, Vancouver, BC, Canada."},{"key":"ref_265","unstructured":"Shang, W., Trott, A., Zheng, S., Xiong, C., and Socher, R. (2019). Learning World Graphs to Accelerate Hierarchical Reinforcement Learning. arXiv."},{"key":"ref_266","unstructured":"Zambaldi, V., Raposo, D., Santoro, A., Bapst, V., Li, Y., Babuschkin, I., Tuyls, K., Reichert, D., Lillicrap, T., and Lockhart, E. (2019, January 6\u20139). Deep Reinforcement Learning with Relational Inductive Biases. Proceedings of the ICLR19, New Orleans, LA, USA."},{"key":"ref_267","unstructured":"Santoro, A., Raposo, D., Barrett, D.G., Malinowski, M., Pascanu, R., Battaglia, P., and Lillicrap, T. (2017, January 4\u20139). A Simple Neural Network Module for Relational Reasoning. Proceedings of the NIPS17, Long Beach, CA, USA."}],"container-title":["Machine Learning and Knowledge Extraction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2504-4990\/4\/1\/9\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:21:52Z","timestamp":1760134912000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2504-4990\/4\/1\/9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,17]]},"references-count":267,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2022,3]]}},"alternative-id":["make4010009"],"URL":"https:\/\/doi.org\/10.3390\/make4010009","relation":{},"ISSN":["2504-4990"],"issn-type":[{"value":"2504-4990","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,17]]}}}