{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,30]],"date-time":"2026-05-30T04:38:11Z","timestamp":1780115891097,"version":"3.54.0"},"reference-count":32,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2018,9,30]],"date-time":"2018-09-30T00:00:00Z","timestamp":1538265600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100003329","name":"Spanish Ministerio de Econom\u00eda y Competitividad","doi-asserted-by":"crossref","award":["TIN2015-65686-C5-1-R"],"award-info":[{"award-number":["TIN2015-65686-C5-1-R"]}],"id":[{"id":"10.13039\/501100003329","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100012818","name":"Comunidad de Madrid","doi-asserted-by":"crossref","award":["2016-T2\/TIC-1712"],"award-info":[{"award-number":["2016-T2\/TIC-1712"]}],"id":[{"id":"10.13039\/100012818","id-type":"DOI","asserted-by":"crossref"}]},{"name":"European Union's Horizon 2020 Research and Innovation programme","award":["730086 (ERGO)"],"award-info":[{"award-number":["730086 (ERGO)"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Auton. Adapt. Syst."],"published-print":{"date-parts":[[2018,9,30]]},"abstract":"<jats:p>\n            This work introduces\n            <jats:italic>Policy Reuse for Safe Reinforcement Learning<\/jats:italic>\n            , an algorithm that combines Probabilistic Policy Reuse and teacher advice for safe exploration in dangerous and continuous state and action reinforcement learning problems in which the dynamic behavior is reasonably smooth and the space is Euclidean. The algorithm uses a continuously increasing monotonic risk function that allows for the identification of the probability to end up in failure from a given state. Such a risk function is defined in terms of how far such a state is from the state space known by the learning agent. Probabilistic Policy Reuse is used to safely balance the exploitation of actual learned knowledge, the exploration of new actions, and the request of teacher advice in parts of the state space considered dangerous. Specifically, the \u03c0-reuse exploration strategy is used. Using experiments in the helicopter hover task and a business management problem, we show that the \u03c0-reuse exploration strategy can be used to completely avoid the visit to undesirable situations while maintaining the performance (in terms of the classical long-term accumulated reward) of the final policy achieved.\n          <\/jats:p>","DOI":"10.1145\/3310090","type":"journal-article","created":{"date-parts":[[2019,3,15]],"date-time":"2019-03-15T12:13:05Z","timestamp":1552651985000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Probabilistic Policy Reuse for Safe Reinforcement Learning"],"prefix":"10.1145","volume":"13","author":[{"given":"Javier","family":"Garc\u00eda","sequence":"first","affiliation":[{"name":"Universidad Carlos III de Madrid, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fernando","family":"Fern\u00e1ndez","sequence":"additional","affiliation":[{"name":"Universidad Carlos III de Madrid, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,3,15]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/196108.196115"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2008.10.024"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.dss.2009.06.009"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2017.06.066"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1561\/2300000021"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1160633.1160762"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1007\/s13748-012-0026-6"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1142\/S0219622012500277"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/2444851.2444864"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/1622519.1622522"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-012-5322-7"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems 29","author":"Ho Jonathan","year":"2016"},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of 2005 International Conference on Machine Learning and Cybernetics","volume":"1","author":"Huang Bing-Qiang","year":"2005"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/1622737.1622748"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-010-5223-6"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12065-011-0066-z"},{"key":"e_1_2_1_18_1","first-page":"1","article-title":"End-to-end training of deep visuomotor policies","volume":"17","author":"Levine Sergey","year":"2016","journal-title":"Journal Machine Learning Research"},{"key":"e_1_2_1_19_1","volume-title":"IProceedings of the 35th Annual Conference of IEEE Industrial Electronics (IECON\u201909)","author":"de Lope J. A.","year":"2009"},{"key":"e_1_2_1_20_1","volume-title":"Computer Aided Systems Theory\u2014EUROCAST","author":"Javier de Lope Jos\u00e9 Antonio","year":"2009"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3067695.3067716"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3067695.3082052"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/2981345.2981445"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2008.02.003"},{"key":"e_1_2_1_25_1","volume-title":"In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics. 627--635","author":"Ross Stephane","year":"2011"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1177\/105971239700600201"},{"key":"e_1_2_1_27_1","unstructured":"Bradly C. Stadie Sergey Levine and Pieter Abbeel. 2015. Incentivizing exploration in reinforcement learning with deep predictive models. In Advances in Neural Information Processing Systems. 2750--2759.  Bradly C. Stadie Sergey Levine and Pieter Abbeel. 2015. Incentivizing exploration in reinforcement learning with deep predictive models. In Advances in Neural Information Processing Systems. 2750--2759."},{"key":"e_1_2_1_28_1","volume-title":"Barto","author":"Sutton Richard S.","year":"1998"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2017.8202134"},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the International Conference on Autonomous Agents and Multiagent Systems (AAMAS\u201911)","author":"Taylor Matthew E.","year":"2011"},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the 10th International Conference on Autonomous Agents and Multiagent Systems (AAMAS\u201911)","volume":"2","author":"Taylor Matthew E.","year":"2011"},{"key":"e_1_2_1_32_1","volume-title":"Proceedings of the Adaptive and Learning Agents Workshop (AAMAS\u201912)","author":"Torrey Lisa"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ADPRL.2007.368199"}],"container-title":["ACM Transactions on Autonomous and Adaptive Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3310090","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3310090","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:53:36Z","timestamp":1750204416000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3310090"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,9,30]]},"references-count":32,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2018,9,30]]}},"alternative-id":["10.1145\/3310090"],"URL":"https:\/\/doi.org\/10.1145\/3310090","relation":{},"ISSN":["1556-4665","1556-4703"],"issn-type":[{"value":"1556-4665","type":"print"},{"value":"1556-4703","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,9,30]]},"assertion":[{"value":"2017-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-01-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-03-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}