{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,23]],"date-time":"2026-03-23T22:45:25Z","timestamp":1774305925361,"version":"3.50.1"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2022,3,23]],"date-time":"2022-03-23T00:00:00Z","timestamp":1647993600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,3,23]],"date-time":"2022-03-23T00:00:00Z","timestamp":1647993600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"ING Bank N.V."}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2022,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Deep reinforcement learning (DRL) has shown remarkable success in artificial domains and in some real-world applications. However, substantial challenges remain such as learning efficiently under safety constraints. Adherence to safety constraints is a hard requirement in many high-impact application domains such as healthcare and finance. These constraints are preferably represented symbolically to ensure clear semantics at a suitable level of abstraction. Existing approaches to safe DRL assume that being unsafe leads to low rewards. We show that this is a special case of symbolically constrained RL and analyze a generic setting in which total reward and being safe may or may not be correlated. We analyze the impact of symbolic constraints and identify a connection between expected future reward and distance towards a goal in an automaton representation of the constraints. We use this connection in an algorithm for learning complex behaviors safely and efficiently. This algorithm relies on symbolic reasoning over safety constraints to improve the efficiency of a subsymbolic learner with a symbolically obtained measure of progress. We measure sample efficiency on a grid world and a conversational product recommender with real-world constraints. The so-called Planning for Potential algorithm converges quickly and significantly outperforms all baselines. Specifically, we find that symbolic reasoning is necessary for safety during and after learning and can be effectively used to guide a neural learner towards promising areas of the solution space. We conclude that RL can be applied both safely and efficiently when combined with symbolic reasoning.<\/jats:p>","DOI":"10.1007\/s10994-022-06143-6","type":"journal-article","created":{"date-parts":[[2022,3,23]],"date-time":"2022-03-23T21:25:31Z","timestamp":1648070731000},"page":"2255-2274","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Planning for potential: efficient safe reinforcement learning"],"prefix":"10.1007","volume":"111","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2092-9904","authenticated-orcid":false,"given":"Floris","family":"den Hengst","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vincent","family":"Fran\u00e7ois-Lavet","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mark","family":"Hoogendoorn","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Frank","family":"van Harmelen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,3,23]]},"reference":[{"key":"6143_CR1","doi-asserted-by":"crossref","unstructured":"Alshiekh, M., Bloem, R., Ehlers, R., K\u00f6nighofer, B., Niekum, S., & Topcu, U. (2018). Safe reinforcement learning via shielding. In Proceedings of the AAAI conference on artificial intelligence (Vol. 32).","DOI":"10.1609\/aaai.v32i1.11797"},{"key":"6143_CR2","unstructured":"Andreas, J., Klein, D., & Levine, S. (2017). Modular multitask reinforcement learning with policy sketches. In Proceedings of the 36th international conference on machine learning vonference (pp. 166\u2013175). PMLR."},{"key":"6143_CR3","volume-title":"Principles of model checking","author":"Christel Baier","year":"2008","unstructured":"Baier, C., & Katoen, J.-P. (2008). Principles of model checking. MIT."},{"key":"6143_CR4","first-page":"1471","volume":"29","author":"Marc Bellemare","year":"2016","unstructured":"Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., & Munos, R. (2016). Unifying count-based exploration and intrinsic motivation. Advances in Neural Information Processing Systems, 29, 1471\u20131479.","journal-title":"Advances in neural information processing systems"},{"key":"6143_CR5","doi-asserted-by":"crossref","unstructured":"Bloem, R., K\u00f6nighofer, B., K\u00f6nighofer, R., & Wang, C. (2015). Shield synthesis. In International conference on tools and algorithms for the construction and analysis of systems (pp. 533\u2013548). Springer.","DOI":"10.1007\/978-3-662-46681-0_51"},{"key":"6143_CR6","doi-asserted-by":"crossref","unstructured":"Brafman, R. I., De Giacomo, G., & Patrizi, F. (2018). LTLf\/LDLf non-Markovian rewards. In Proceedings of the AAAI conference on artificial intelligence (Vol. 32).","DOI":"10.1609\/aaai.v32i1.11572"},{"key":"6143_CR7","unstructured":"Burda, Y., Edwards, H., Pathak, D., Storkey, A., Darrell, T., & Efros, A. A. (2019). Large-scale study of curiosity-driven learning. In International conference on learning representations."},{"key":"6143_CR8","unstructured":"Camacho, A., Chen, O., Sanner, S., & McIlraith, S. A. (2017). Non-markovian rewards expressed in LTL: Guiding search via reward shaping. In Tenth annual symposium on combinatorial search."},{"key":"6143_CR9","doi-asserted-by":"crossref","unstructured":"Camacho, A., Icarte, R. T., Klassen, T. Q., Valenzano, R. A., & McIlraith, S. A. (2019). Ltl and beyond: Formal languages for reward function specification in reinforcement learning. In Proceedings of the 28th joint conference on artificial intelligence (Vol. 19, pp. 6065\u20136073).","DOI":"10.24963\/ijcai.2019\/840"},{"key":"6143_CR10","doi-asserted-by":"crossref","first-page":"128","DOI":"10.1609\/icaps.v29i1.3549","volume":"29","author":"Giuseppe De Giacomo","year":"2019","unstructured":"De Giacomo, G., Iocchi, L., Favorito, M., & Patrizi, F. (2019). Foundations for restraining bolts: Reinforcement learning with LTLf\/LDLf restraining specifications. In Proceedings of the international conference on automated planning and scheduling (Vol. 29, pp. 128\u2013136).","journal-title":"In Proceedings of the International Conference on Automated Planning and Scheduling"},{"key":"6143_CR11","doi-asserted-by":"crossref","first-page":"517","DOI":"10.1609\/icaps.v30i1.6747","volume":"30","author":"Giuseppe De Giacomo","year":"2020","unstructured":"De Giacomo, G., Favorito, M., Iocchi, L., & Patrizi, F. (2020). Imitation learning over heterogeneous agents with restraining bolts. In Proceedings of the international conference on automated planning and scheduling (Vol. 30, pp. 517\u2013521).","journal-title":"In Proceedings of the International Conference on Automated Planning and Scheduling"},{"key":"6143_CR12","doi-asserted-by":"crossref","unstructured":"den Hengst, F., Hoogendoorn, M., Van Harmelen, F., & Bosman, J. (2019). Reinforcement learning for personalized dialogue management. In International conference on web intelligence (pp. 59\u201367). IEEE\/WIC\/ACM.","DOI":"10.1145\/3350546.3352501"},{"issue":"1","key":"6143_CR13","doi-asserted-by":"publisher","first-page":"107","DOI":"10.3233\/DS-200028","volume":"3","author":"Floris den Hengst","year":"2020","unstructured":"den Hengst, F., Grua, E. M., el Hassouni, A., & Hoogendoorn, M. (2020). Reinforcement learning for personalization: A systematic literature review. Data Science, 3(1), 107\u2013147.","journal-title":"Data Science"},{"key":"6143_CR14","unstructured":"Dulac-Arnold, G., Mankowitz, D., & Hester, T. (2019). Challenges of real-world reinforcement learning. In ICML workshop on real-life reinforcement learning."},{"key":"6143_CR15","doi-asserted-by":"crossref","unstructured":"Fu, J., & Topcu, U. (2014). Probably approximately correct mdp learning and control with temporal logic constraints. In Proceedings of robotics: Science and systems (Vol. 10).","DOI":"10.15607\/RSS.2014.X.039"},{"key":"6143_CR16","doi-asserted-by":"publisher","first-page":"3980","DOI":"10.1609\/aaai.v34i04.5814","volume":"34","author":"Maor Gaon","year":"2020","unstructured":"Gaon, M., & Brafman, R. (2020). Reinforcement learning with non-markovian rewards. In Proceedings of the AAAI conference on artificial intelligence, (Vol. 34, pp. 3980\u20133987).","journal-title":"In Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"6143_CR17","doi-asserted-by":"crossref","unstructured":"Grzes, M., & Kudenko, D. (2008). Plan-based reward shaping for reinforcement learning. In International IEEE conference intelligent systems (Vol. 2, pp. 10\u201322). IEEE.","DOI":"10.1109\/IS.2008.4670492"},{"key":"6143_CR18","doi-asserted-by":"crossref","unstructured":"Gu, S., Holly, E., Lillicrap, T., & Levine, S. (2017). Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In 2017 IEEE international conference on robotics and automation (ICRA) (pp. 3389\u20133396). IEEE.","DOI":"10.1109\/ICRA.2017.7989385"},{"key":"6143_CR19","unstructured":"Hasanbeig, M., Abate, A., & Kroening, D. (2020). Cautious reinforcement learning with logical constraints. In Proceedings of the 19th international conference on autonomous agents and multiagent systems (pp. 483\u2013491)."},{"key":"6143_CR20","doi-asserted-by":"crossref","unstructured":"Hasanbeig, M., Jeppu, N. Y., Abate, A., Melham, T., & Kroening, D. (2021).Deepsynth: Automata synthesis for automatic task segmentation in deep reinforcement learning. In The 35th AAAI conference on artificial intelligence, AAAI (Vol. 2, p. 36).","DOI":"10.1609\/aaai.v35i9.16935"},{"key":"6143_CR21","doi-asserted-by":"crossref","unstructured":"Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., & Silver, D. (2018). Rainbow: Combining improvements in deep reinforcement learning. In Proceedings of the AAAI conference on artificial intelligence (Vol. 32).","DOI":"10.1609\/aaai.v32i1.11796"},{"key":"6143_CR22","unstructured":"Icarte, R. T., Klassen, T., Valenzano, R., & McIlraith, S. (2018). Using reward machines for high-level task specification and decomposition in reinforcement learning. In Proceedings of the 37th international conference on machine learning conference (pp. 2107\u20132116)."},{"key":"6143_CR23","doi-asserted-by":"crossref","unstructured":"Illanes, L., Yan, X., Icarte, R. T., & McIlraith, S. A. (2020). Symbolic plans as high-level instructions for reinforcement learning. In Proceedings of the international conference on automated planning and scheduling (Vol. 30, pp. 540\u2013550).","DOI":"10.1609\/icaps.v30i1.6750"},{"key":"6143_CR24","doi-asserted-by":"crossref","unstructured":"Junges, S., Jansen, N., Dehnert, C., Topcu, U., & Katoen, J.-P.. (2016). Safety-constrained reinforcement learning for mdps. In International conference on tools and algorithms for the construction and analysis of systems (pp. 130\u2013146). Springer.","DOI":"10.1007\/978-3-662-49674-9_8"},{"key":"6143_CR25","doi-asserted-by":"crossref","unstructured":"K\u00f6nighofer, B., Lorber, F., Jansen, N., & Bloem, R. (2020). Shield synthesis for reinforcement learning. In International symposium on leveraging applications of formal methods (pp. 290\u2013306). Springer.","DOI":"10.1007\/978-3-030-61362-4_16"},{"key":"6143_CR26","doi-asserted-by":"crossref","unstructured":"Mazala, R. (2002). Infinite games (pp. 23\u201338). Springer. ISBN 978-3-540-36387-3.","DOI":"10.1007\/3-540-36387-4_2"},{"key":"6143_CR27","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., & Riedmiller, M. (2013). Playing atari with deep reinforcement learning. In NIPS deep learning workshop."},{"issue":"7540","key":"6143_CR28","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"Volodymyr Mnih","year":"2015","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529\u2013533.","journal-title":"Nature"},{"key":"6143_CR29","unstructured":"Ng, A. Y., Harada, D., & Russell, S. (1999). Policy invariance under reward transformations: Theory and application to reward shaping. In Proceedings of the 16th international conference on machine learning (pp. 278\u2013287)."},{"key":"6143_CR30","doi-asserted-by":"crossref","unstructured":"Pnueli, A. (1977). The temporal logic of programs. In 18th Annual symposium on foundations of computer science (pp. 46\u201357). IEEE.","DOI":"10.1109\/SFCS.1977.32"},{"key":"6143_CR31","doi-asserted-by":"crossref","unstructured":"Pnueli, A., & Rosner, R. (1989). On the synthesis of a reactive module. In ACM SIGPLAN-SIGACT (pp. 179\u2013190).","DOI":"10.1145\/75277.75293"},{"issue":"6419","key":"6143_CR32","doi-asserted-by":"publisher","first-page":"1140","DOI":"10.1126\/science.aar6404","volume":"362","author":"David Silver","year":"2018","unstructured":"Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., & Hassabis, D. (2018). A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science, 362(6419), 1140\u20131144.","journal-title":"Science"},{"key":"6143_CR33","volume-title":"Reinforcement learning: An introduction","author":"Richard S Sutton","year":"2018","unstructured":"Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction. MIT."},{"key":"6143_CR34","unstructured":"Tomic, S., Pecora, F., & Saffiotti, A. (2020). Learning normative behaviors through abstraction. In Proceedings of the 24th European conference on artificial intelligence."},{"issue":"3\u20134","key":"6143_CR35","first-page":"279","volume":"8","author":"Christopher JCH Watkins","year":"1992","unstructured":"Watkins, C. J. C. H., & Dayan, P. (1992). Q-learning. Machine Learning, 8(3\u20134), 279\u2013292.","journal-title":"Machine Learning"},{"key":"6143_CR36","doi-asserted-by":"crossref","unstructured":"Wen, M., Ehlers, R., & Topcu, U. (2015). Correct-by-synthesis reinforcement learning with temporal logic constraints. In 2015 IEEE\/RSJ international conference on intelligent robots and systems (IROS) (pp. 4983\u20134990). RSJ\/IEEE.","DOI":"10.1109\/IROS.2015.7354078"},{"key":"6143_CR37","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1007\/978-3-642-27645-3_1","volume":"12","author":"Marco Wiering","year":"2012","unstructured":"Wiering, M., & Van Otterlo, M. (2012). Reinforcement learning. Adaptation, Learning, and Optimization, 12, 3.","journal-title":"Adaptation, learning, and optimization"},{"key":"6143_CR38","unstructured":"Zhang, H., Gao, Z., Zhou, Y., Zhang, H., Wu, K., & Lin, F. (2019). Faster and safer training by embedding high-level knowledge into deep reinforcement learning. arXiv preprint. arXiv:1910.09986"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06143-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-022-06143-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06143-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,29]],"date-time":"2023-01-29T22:29:23Z","timestamp":1675031363000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-022-06143-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,23]]},"references-count":38,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2022,6]]}},"alternative-id":["6143"],"URL":"https:\/\/doi.org\/10.1007\/s10994-022-06143-6","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,3,23]]},"assertion":[{"value":"4 March 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 October 2021","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 February 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 March 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant financial or non-financial interests to disclose.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"No ethical approval was sought for this study on the grounds of no involvement of any human or animal subjects.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Research involving human or animal subjects"}}]}}