{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T17:04:58Z","timestamp":1781370298499,"version":"3.54.1"},"reference-count":51,"publisher":"Maximum Academic Press","license":[{"start":{"date-parts":[[2021,4,7]],"date-time":"2021-04-07T00:00:00Z","timestamp":1617753600000},"content-version":"unspecified","delay-in-days":96,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":["cambridge.org"],"crossmark-restriction":true},"short-container-title":["The Knowledge Engineering Review"],"published-print":{"date-parts":[[2021]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Designing hierarchical reinforcement learning algorithms that exhibit safe behaviour is not only vital for practical applications but also facilitates a better understanding of an agent\u2019s decisions. We tackle this problem in the options framework (Sutton, Precup &amp; Singh, 1999), a particular way to specify temporally abstract actions which allow an agent to use sub-policies with start and end conditions. We consider a behaviour as\n                    <jats:italic>safe<\/jats:italic>\n                    that avoids regions of state space with high uncertainty in the outcomes of actions. We propose an optimization objective that learns\n                    <jats:italic>safe<\/jats:italic>\n                    options by encouraging the agent to visit states with higher behavioural consistency. The proposed objective results in a trade-off between maximizing the standard expected return and minimizing the effect of model uncertainty in the return. We propose a policy gradient algorithm to optimize the constrained objective function. We examine the quantitative and qualitative behaviours of the proposed approach in a tabular grid world, continuous-state puddle world, and three games from the Arcade Learning Environment: Ms. Pacman, Amidar, and Q*Bert. Our approach achieves a reduction in the variance of return, boosts performance in environments with intrinsic variability in the reward structure, and compares favourably both with primitive actions and with risk-neutral options.\n                  <\/jats:p>","DOI":"10.1017\/s0269888921000035","type":"journal-article","created":{"date-parts":[[2021,4,7]],"date-time":"2021-04-07T09:32:16Z","timestamp":1617787936000},"update-policy":"https:\/\/doi.org\/10.1017\/policypage","source":"Crossref","is-referenced-by-count":11,"title":["Safe option-critic: learning safety in the option-critic architecture"],"prefix":"10.48130","volume":"36","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3081-0421","authenticated-orcid":false,"given":"Arushi","family":"Jain","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9975-6438","authenticated-orcid":false,"given":"Khimya","family":"Khetarpal","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Doina","family":"Precup","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"27968","published-online":{"date-parts":[[2021,4,7]]},"reference":[{"key":"S0269888921000035_ref8","doi-asserted-by":"publisher","DOI":"10.1613\/jair.639"},{"key":"S0269888921000035_ref13","first-page":"1037","article-title":"Smart exploration in reinforcement learning using absolute temporal difference errors","volume":"2013","author":"Gehring","year":"2013","journal-title":"International Conference on Autonomous Agents and Multi-agent Systems"},{"key":"S0269888921000035_ref6","doi-asserted-by":"publisher","DOI":"10.1287\/moor.27.1.192.334"},{"key":"S0269888921000035_ref28","first-page":"701","article-title":"Reinforcement learning in robust Markov decision processes","volume":"26","author":"Lim","year":"2013","journal-title":"Advances in Neural Information Processing Systems"},{"key":"S0269888921000035_ref15","doi-asserted-by":"crossref","unstructured":"Harb, J. , Bacon, P.-L. , Klissarov, M. & Precup, D. 2018. When waiting is not an option: learning options with a deliberation cost. In AAAI.","DOI":"10.1609\/aaai.v32i1.11831"},{"key":"S0269888921000035_ref27","unstructured":"Law, E. L. , Coggan, M. , Precup, D. & Ratitch, B. 2005. Risk-directed exploration in reinforcement learning. In Planning and Learning in A Priori Unknown or Dynamic Domains, 97."},{"key":"S0269888921000035_ref11","year":"2017"},{"key":"S0269888921000035_ref46","first-page":"1","article-title":"Learning the variance of the reward-to-go","volume":"17","author":"Tamar","year":"2016","journal-title":"Journal of Machine Learning Research"},{"key":"S0269888921000035_ref2","unstructured":"Bacon, P.-L. , Harb, J. & Precup, D. 2017. The option-critic architecture. In AAAI, 1726\u20131734."},{"key":"S0269888921000035_ref45","unstructured":"Tamar, A. , Di Castro, D. & Mannor, S. 2012. Policy gradients with variance related risk criteria. In Proceedings of the Twenty-Ninth International Conference on Machine Learning, 387\u2013396."},{"key":"S0269888921000035_ref5","doi-asserted-by":"publisher","DOI":"10.1613\/jair.3912"},{"key":"S0269888921000035_ref49","first-page":"3486","article-title":"Strategic attentive writer for learning macro-actions","author":"Vezhnevets","year":"2016","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"S0269888921000035_ref50","author":"Wang","year":"2015"},{"key":"S0269888921000035_ref7","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-016-5580-x"},{"key":"S0269888921000035_ref29","doi-asserted-by":"publisher","DOI":"10.1613\/jair.5699"},{"key":"S0269888921000035_ref9","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(72)90051-3"},{"key":"S0269888921000035_ref18","doi-asserted-by":"publisher","DOI":"10.1007\/BF00116836"},{"key":"S0269888921000035_ref14","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1666"},{"key":"S0269888921000035_ref24","doi-asserted-by":"crossref","unstructured":"Konidaris, G. , Kuindersma, S. , Grupen, R. A. & Barto, A. G. 2011. Autonomous skill acquisition on a mobile manipulator. In AAAI.","DOI":"10.1609\/aaai.v25i1.7982"},{"key":"S0269888921000035_ref30","first-page":"1588","article-title":"Adaptive skills adaptive partitions (ASAP)","author":"Mankowitz","year":"2016","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"S0269888921000035_ref3","first-page":"13052","article-title":"The option keyboard: combining skills in reinforcement learning","author":"Barreto","year":"2019","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"S0269888921000035_ref12","first-page":"1437","article-title":"A comprehensive survey on safe reinforcement learning","volume":"16","author":"Garca","year":"2015","journal-title":"Journal of Machine Learning Research"},{"key":"S0269888921000035_ref4","doi-asserted-by":"publisher","DOI":"10.1023\/A:1025696116075"},{"key":"S0269888921000035_ref33","unstructured":"Mnih, V. , Badia, A. P. , Mirza, M. , Graves, A. , Lillicrap, T. , Harley, T. , Silver, D. & Kavukcuoglu, K. 2016. Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning, 1928\u20131937."},{"key":"S0269888921000035_ref37","volume-title":"Temporal abstraction in reinforcement learning","author":"Precup","year":"2000"},{"key":"S0269888921000035_ref21","unstructured":"Jain, A. & Precup, D. 2018. Eligibility traces for options. In Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems, 1008\u20131016."},{"key":"S0269888921000035_ref40","first-page":"212","author":"Stolle","year":"2002"},{"key":"S0269888921000035_ref38","first-page":"10424","article-title":"Learning abstract options","author":"Riemer","year":"2018","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"S0269888921000035_ref41","doi-asserted-by":"publisher","DOI":"10.1007\/BF00115009"},{"key":"S0269888921000035_ref42","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.1998.712192"},{"key":"S0269888921000035_ref26","first-page":"3675","article-title":"Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation","author":"Kulkarni","year":"2016","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"S0269888921000035_ref23","first-page":"895","article-title":"Building portable options: Skill transfer in reinforcement learning","volume":"7","author":"Konidaris","year":"2007","journal-title":"In IJCAI"},{"key":"S0269888921000035_ref35","doi-asserted-by":"publisher","DOI":"10.1287\/opre.1050.0216"},{"key":"S0269888921000035_ref39","unstructured":"Sherstan, C. , Ashley, D. R. , Bennett, B. , Young, K. , White, A. , White, M. & Sutton, R. S. 2018. Comparing direct and indirect temporal-difference methods for estimating the variance of the return. In Proceedings of Uncertainty in Artificial Intelligence, 63\u201372."},{"key":"S0269888921000035_ref25","unstructured":"Korf, R. E. 1983. Learning to Solve Problems by Searching for Macro-operators. PhD thesis, Pittsburgh, PA, USA. AAI8425820."},{"key":"S0269888921000035_ref44","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(99)00052-1"},{"key":"S0269888921000035_ref10","volume-title":"Readings in Artificial Intelligence","author":"Fikes","year":"1981"},{"key":"S0269888921000035_ref43","first-page":"1057","article-title":"Policy gradient methods for reinforcement learning with function approximation","author":"Sutton","year":"2000","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"S0269888921000035_ref36","first-page":"1043","article-title":"Reinforcement learning with hierarchies of machines","author":"Parr","year":"1998","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"S0269888921000035_ref47","unstructured":"Tamar, A. , Xu, H. & Mannor, S. 2013. Scaling up robust MDPs by reinforcement learning. arXiv preprint arXiv:1306.6189."},{"key":"S0269888921000035_ref48","first-page":"2094","article-title":"Deep reinforcement learning with double Q-learning","volume":"16","author":"Van Hasselt","year":"2016","journal-title":"AAAI"},{"key":"S0269888921000035_ref51","doi-asserted-by":"publisher","DOI":"10.1007\/BF01719453"},{"key":"S0269888921000035_ref17","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.18.7.356"},{"key":"S0269888921000035_ref1","author":"Amodei","year":"2016"},{"key":"S0269888921000035_ref16","first-page":"105","author":"Heger","year":"1994"},{"key":"S0269888921000035_ref32","first-page":"295","author":"Menache","year":"2002"},{"key":"S0269888921000035_ref19","doi-asserted-by":"publisher","DOI":"10.1287\/moor.1040.0129"},{"key":"S0269888921000035_ref22","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5871"},{"key":"S0269888921000035_ref20","doi-asserted-by":"crossref","unstructured":"Jain, A. , Patil, G. , Jain, A. , Khetarpal, K. & Precup, D. 2021. Variance penalized on-policy and off-policy actor-critic. arXiv preprint arXiv:2102.01985.","DOI":"10.1609\/aaai.v35i9.16964"},{"key":"S0269888921000035_ref31","first-page":"361","article-title":"Automatic discovery of subgoals in reinforcement learning using diverse density","volume":"1","author":"McGovern","year":"2001","journal-title":"ICML"},{"key":"S0269888921000035_ref34","author":"Nair","year":"2015"}],"container-title":["The Knowledge Engineering Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S0269888921000035","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,5]],"date-time":"2026-01-05T14:42:21Z","timestamp":1767624141000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S0269888921000035\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021]]},"references-count":51,"alternative-id":["S0269888921000035"],"URL":"https:\/\/doi.org\/10.1017\/s0269888921000035","relation":{},"ISSN":["0269-8889","1469-8005"],"issn-type":[{"value":"0269-8889","type":"print"},{"value":"1469-8005","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021]]},"assertion":[{"value":"\u00a9 The Author(s), 2021. Published by Cambridge University Press","name":"copyright","label":"Copyright","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}}],"article-number":"e4"}}