{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T18:52:32Z","timestamp":1783795952169,"version":"3.55.0"},"reference-count":92,"publisher":"Springer Science and Business Media LLC","issue":"11","license":[{"start":{"date-parts":[[2022,10,20]],"date-time":"2022-10-20T00:00:00Z","timestamp":1666224000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,10,20]],"date-time":"2022-10-20T00:00:00Z","timestamp":1666224000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100018967","name":"Universitat Pompeu Fabra","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100018967","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2022,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Policy gradient (PG) algorithms are among the best candidates for the much-anticipated applications of reinforcement learning to real-world control tasks, such as robotics. However, the trial-and-error nature of these methods poses safety issues whenever the learning process itself must be performed on a physical system or involves any form of human-computer interaction. In this paper, we address a specific safety formulation, where both goals and dangers are encoded in a scalar reward signal and the learning agent is constrained to never worsen its performance, measured as the expected sum of rewards. By studying actor-only PG from a stochastic optimization perspective, we establish improvement guarantees for a wide class of parametric policies, generalizing existing results on Gaussian policies. This, together with novel upper bounds on the variance of PG estimators, allows us to identify meta-parameter schedules that guarantee monotonic improvement with high probability. The two key meta-parameters are the step size of the parameter updates and the batch size of the gradient estimates. Through a joint, adaptive selection of these meta-parameters, we obtain a PG algorithm with monotonic improvement guarantees.<\/jats:p>","DOI":"10.1007\/s10994-022-06232-6","type":"journal-article","created":{"date-parts":[[2022,10,20]],"date-time":"2022-10-20T20:10:35Z","timestamp":1666296635000},"page":"4081-4137","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Smoothing policies and safe policy gradients"],"prefix":"10.1007","volume":"111","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3807-3171","authenticated-orcid":false,"given":"Matteo","family":"Papini","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Matteo","family":"Pirotta","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marcello","family":"Restelli","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,10,20]]},"reference":[{"issue":"13","key":"6232_CR1","doi-asserted-by":"publisher","first-page":"1608","DOI":"10.1177\/0278364910371999","volume":"29","author":"P Abbeel","year":"2010","unstructured":"Abbeel, P., Coates, A., & Ng, A. Y. (2010). Autonomous helicopter aerobatics through apprenticeship learning. The International Journal of Robotics Research, 29(13), 1608\u20131639.","journal-title":"The International Journal of Robotics Research"},{"key":"6232_CR2","first-page":"22","volume":"70","author":"J Achiam","year":"2017","unstructured":"Achiam, J., Held, D., Tamar, A., & Abbeel, P. (2017). Constrained policy optimization. ICML, 70, 22\u201331. PMLR.","journal-title":"ICML"},{"key":"6232_CR3","first-page":"64","volume":"125","author":"A Agarwal","year":"2020","unstructured":"Agarwal, A., Kakade, S. M., Lee, J. D., & Mahajan, G. (2020). Optimality and approximation with policy gradient methods in Markov decision processes. COLT, 125, 64\u201366. PMLR.","journal-title":"COLT"},{"key":"6232_CR4","unstructured":"Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., Mane, D. (2016). Concrete problems in ai safety. arXiv preprint arXiv:1606.06565 ."},{"issue":"5","key":"6232_CR5","doi-asserted-by":"publisher","first-page":"834","DOI":"10.1109\/TSMC.1983.6313077","volume":"13","author":"AG Barto","year":"1983","unstructured":"Barto, A. G., Sutton, R. S., & Anderson, C. W. (1983). Neuronlike adaptive elements that can solve dicult learning control problems. IEEE Transactions on Systems, Man, and Cybernetics, 13(5), 834\u2013846.","journal-title":"IEEE Trans. Syst. Man Cybern."},{"key":"6232_CR6","doi-asserted-by":"crossref","unstructured":"Baxter, J., & Bartlett, P.L. (2001). Infinite-horizon policy-gradient estimation. Journal of Artificial Intelligence Research, 15 .","DOI":"10.1613\/jair.806"},{"key":"6232_CR7","unstructured":"Bedi, A.S., Parayil, A., Zhang, J.,Wang, M., & Koppel, A. (2021). On the sample complexity and metastability of heavy-tailed policy search in continuous control. CoRR. https:\/\/arxiv.org\/abs\/2106.08414 ."},{"key":"6232_CR8","unstructured":"Berkenkamp, F. (2019). Safe exploration in reinforcement learning: Theory and applications in robotics (Unpublished doctoral dissertation). ETH Zurich."},{"key":"6232_CR9","unstructured":"Berkenkamp, F., Turchetta, M., Schoellig, A.P., & Krause, A. (2017). Safe modelbased reinforcement learning with stability guarantees. NIPS (pp. 908\u2013 919)."},{"issue":"3","key":"6232_CR10","doi-asserted-by":"publisher","first-page":"310","DOI":"10.1007\/s11768-011-1005-3","volume":"9","author":"DP Bertsekas","year":"2011","unstructured":"Bertsekas, D. P. (2011). Approximate policy iteration: A survey and some new methods. Journal of Control Theory and Applications, 9(3), 310\u2013335.","journal-title":"Journal of Control Theory and Applications"},{"key":"6232_CR11","unstructured":"Bertsekas, D.P., & Shreve, S. (2004). Stochastic optimal control: the discrete- time case."},{"key":"6232_CR12","unstructured":"Bhandari, J., & Russo, D. (2019). Global optimality guarantees for policy gradient methods. CoRR.  https:\/\/arxiv.org\/abs\/1906.01786."},{"key":"6232_CR13","doi-asserted-by":"crossref","unstructured":"Bisi, L., Sabbioni, L., Vittori, E., Papini, M., & Restelli, M. (2020). Risk-averse trust region optimization for reward-volatility reduction. IJCAI (pp. 4583\u20134589). ijcai.org.","DOI":"10.24963\/ijcai.2020\/632"},{"key":"6232_CR14","unstructured":"Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., & Zaremba, W. (2016). Openai gym."},{"key":"6232_CR15","unstructured":"Castro, D.D., Tamar, A., & Mannor, S. (2012). Policy gradients with variance related risk criteria. ICML. icml.cc \/ Omnipress."},{"key":"6232_CR16","first-page":"834","volume":"70","author":"P Chou","year":"2017","unstructured":"Chou, P., Maturana, D., & Scherer, S. A. (2017). Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution. ICML, 70, 834\u2013843. PMLR.","journal-title":"ICML"},{"key":"6232_CR17","unstructured":"Chow, Y., Nachum, O., Due\u00a0nez-Guzman, E.A., & Ghavamzadeh, M. (2018). A lyapunov-based approach to safe reinforcement learning. Neurips (pp. 8103\u20138112)."},{"key":"6232_CR18","first-page":"52:1","volume":"21","author":"K Ciosek","year":"2020","unstructured":"Ciosek, K., & Whiteson, S. (2020). Expected policy gradients for reinforcement learning. Journal of Machine Learning Research 21, 52:1-52:51.","journal-title":"J. Mach. Learn. Res."},{"key":"6232_CR19","first-page":"92","volume":"1992","author":"JA Clouse","year":"1992","unstructured":"Clouse, J. A., & Utgo, P. E. (1992). A teaching method for reinforcement learning. Machine Learning Proceedings, 1992, 92\u2013101. Elsevier.","journal-title":"Machine learning proceedings"},{"key":"6232_CR20","doi-asserted-by":"crossref","unstructured":"Cohen, A., Yu, L., & Wright, R. (2018). Diverse exploration for fast and safe policy improvement. arXiv preprint arXiv:1802.08331 .","DOI":"10.1609\/aaai.v32i1.11758"},{"key":"6232_CR21","unstructured":"Dalal, G., Dvijotham, K., Vecerk, M., Hester, T., Paduraru, C., & Tassa, Y. (2018). Safe exploration in continuous action spaces. CoRR. https:\/\/arxiv.org\/abs\/1801.08757."},{"key":"6232_CR22","doi-asserted-by":"crossref","unstructured":"Deisenroth, M.P., Neumann, G., & Peters, J., et al. (2013). A survey on policy search for robotics. Foundations and Trends\u00aein Robotics, 2 (1-2), 1-142.","DOI":"10.1561\/2300000021"},{"key":"6232_CR23","unstructured":"Dorato, P., Cerone, V., & Abdallah, C. (1994). Linear-quadratic control: an introduction. Simon & Schuster, Inc."},{"key":"6232_CR24","first-page":"1329","volume":"48","author":"Y Duan","year":"2016","unstructured":"Duan, Y., Chen, X., Houthooft, R., Schulman, J., & Abbeel, P. (2016). Benchmarking deep reinforcement learning for continuous control. ICML, 48, 1329\u20131338. JMLR.org.","journal-title":"ICML"},{"key":"6232_CR25","unstructured":"Fruit, R., Lazaric, A., & Pirotta, M. (2019). Regret minimization in infinite-horizon finite markov decision processes. Tutorial at ALT\u201919. Retrieved from https:\/\/rlgammazero.github.io\/"},{"key":"6232_CR26","unstructured":"Furmston, T., & Barber, D. (2012). A unifying perspective of parametric policy search methods for markov decision processes. Advances in neural information processing systems (pp. 2717-2725)."},{"key":"6232_CR27","first-page":"1431","volume":"108","author":"E Garcelon","year":"2020","unstructured":"Garcelon, E., Ghavamzadeh, M., Lazaric, A., & Pirotta, M. (2020). Conservative exploration in reinforcement learning. AISTATS, 108, 1431\u20131441. PMLR.","journal-title":"AISTATS"},{"key":"6232_CR28","doi-asserted-by":"crossref","unstructured":"Garcelon, E., Ghavamzadeh, M., Lazaric, A., & Pirotta, M. (2020b). Improved algorithms for conservative exploration in bandits. AAAI (pp. 3962\u2013 3969). AAAI Press.","DOI":"10.1609\/aaai.v34i04.5812"},{"key":"6232_CR29","first-page":"1437","volume":"16","author":"J Garca","year":"2015","unstructured":"Garca, J., & Fernandez, F. (2015). A comprehensive survey on safe reinforcement learning. Journal of Machine Learning Research 16, 1437\u20131480.","journal-title":"J. Mach. Learn. Res."},{"key":"6232_CR30","unstructured":"Gehring, C., & Precup, D. (2013). Smart exploration in reinforcement learning using absolute temporal difference errors. In Proceedings of the 2013 international conference on autonomous agents and multi-agent systems (pp. 1037\u20131044)."},{"key":"6232_CR31","doi-asserted-by":"publisher","first-page":"81","DOI":"10.1613\/jair.1666","volume":"24","author":"P Geibel","year":"2005","unstructured":"Geibel, P., & Wysotzki, F. (2005). Risk-sensitive reinforcement learning applied to control under constraints. Journal of Artificial Intelligence Research, 24, 81\u2013108.","journal-title":"Journal of Artificial Intelligence Research"},{"key":"6232_CR32","doi-asserted-by":"crossref","unstructured":"Glynn, P.W. (1986). Stochastic approximation for monte carlo optimization. WSC (pp. 356\u2013365).","DOI":"10.1145\/318242.318459"},{"key":"6232_CR33","volume-title":"Probability and random processes","author":"G Grimmett","year":"2020","unstructured":"Grimmett, G., & Stirzaker, D. (2020). Probability and random processes. Oxford University Press."},{"key":"6232_CR34","first-page":"1856","volume":"80","author":"T Haarnoja","year":"2018","unstructured":"Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. ICML, 80, 1856\u20131865. JMLR.org.","journal-title":"ICML"},{"key":"6232_CR35","unstructured":"Hans, A., Schneega, D., Sch\u00e4fer, A.M., & Udluft, S. (2008). Safe exploration for reinforcement learning. Esann (pp. 143\u2013148)."},{"issue":"2","key":"6232_CR36","doi-asserted-by":"publisher","first-page":"279","DOI":"10.1016\/j.camwa.2005.11.013","volume":"51","author":"Y Kadota","year":"2006","unstructured":"Kadota, Y., Kurano, M., & Yasuda, M. (2006). Discounted markov decision processes with utility constraints. Computers & Mathematics with Applications, 51(2), 279\u2013284.","journal-title":"Computers & Mathematics with Applications"},{"key":"6232_CR37","doi-asserted-by":"crossref","unstructured":"Kakade, S. (2001). Optimizing average reward using discounted rewards. Inter- national conference on computational learning theory (pp. 605\u2013615).","DOI":"10.1007\/3-540-44581-1_40"},{"key":"6232_CR38","unstructured":"Kakade, S. (2002). A natural policy gradient. Advances in neural information processing systems (pp. 1531\u20131538)."},{"key":"6232_CR39","unstructured":"Kakade, S., & Langford, J. (2002). Approximately optimal approximate reinforcement learning.."},{"key":"6232_CR40","volume-title":"On the sample complexity of reinforcement learning (Unpublished doctoral dissertation)","author":"SM Kakade","year":"2003","unstructured":"Kakade, S. M., et al. (2003). On the sample complexity of reinforcement learning (Unpublished doctoral dissertation). England: University of London London."},{"key":"6232_CR41","unstructured":"Kazerouni, A., Ghavamzadeh, M., Abbasi, Y., & Roy, B.V. (2017). Conservative contextual linear bandits. NIPS (pp. 3913\u20133922)."},{"key":"6232_CR42","volume-title":"Probability theory: A comprehensive course","author":"A Klenke","year":"2013","unstructured":"Klenke, A. (2013). Probability theory: A comprehensive course. Springer Science & Business Media."},{"issue":"11","key":"6232_CR43","doi-asserted-by":"publisher","first-page":"1238","DOI":"10.1177\/0278364913495721","volume":"32","author":"J Kober","year":"2013","unstructured":"Kober, J., Bagnell, J. A., & Peters, J. (2013). Reinforcement learning in robotics: A survey. The International Journal of Robotics Research, 32(11), 1238\u20131274.","journal-title":"The International Journal of Robotics Research"},{"key":"6232_CR44","unstructured":"Konda, V.R., & Tsitsiklis, J.N. (1999). Actor-critic algorithms. NeurIPS (pp. 1008\u20131014)."},{"key":"6232_CR45","unstructured":"Laroche, R., Trichelair, P., & Des Combes, R.T. (2019). Safe policy improvement with baseline bootstrapping. In International conference on machine learning (pp. 3652\u20133661)."},{"issue":"3","key":"6232_CR46","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1145\/2512962","volume":"46","author":"B Li","year":"2014","unstructured":"Li, B., & Hoi, S. C. (2014). Online portfolio selection: A survey. ACM Computing Surveys (CSUR), 46(3), 35.","journal-title":"ACM Computing Surveys (CSUR)"},{"key":"6232_CR47","unstructured":"Maurer, A., & Pontil, M. (2009). Empirical Bernstein bounds and samplevariance penalization. COLT."},{"issue":"97","key":"6232_CR48","first-page":"1","volume":"22","author":"AM Metelli","year":"2021","unstructured":"Metelli, A. M., Pirotta, M., Calandriello, D., & Restelli, M. (2021). Safe policy iteration: A monotonically improving approximate policy iteration approach. Journal of Machine Learning Research, 22(97), 1\u201383.","journal-title":"Journal of Machine Learning Research"},{"issue":"7540","key":"6232_CR49","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., et al. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529.","journal-title":"Nature"},{"key":"6232_CR50","unstructured":"Moldovan, T.M., & Abbeel, P. (2012). Safe exploration in markov decision processes. In Proceedings of the 29th international conference on international conference on machine learning (pp. 1451-+1458)."},{"key":"6232_CR51","unstructured":"Nesterov, Y. (1998). Introductory lectures on convex programming volume i: Basic course. Lecture notes."},{"key":"6232_CR52","volume-title":"Introductory lectures on convex optimization: A basic course (Vol. 87)","author":"Y Nesterov","year":"2013","unstructured":"Nesterov, Y. (2013). Introductory lectures on convex optimization: A basic course (Vol. 87). Springer Science & Business Media."},{"key":"6232_CR53","unstructured":"Neu, G., Jonsson, A., & Gomez, V. (2017). A unified view of entropy-regularized markov decision processes. CoRR. https:\/\/arxiv.org\/abs\/1705.07798."},{"key":"6232_CR54","unstructured":"Nota, C., & Thomas, P.S. (2020). Is the policy gradient a gradient? AAMAS (pp. 939\u2013947). International Foundation for Autonomous Agents and Multiagent Systems."},{"key":"6232_CR55","doi-asserted-by":"crossref","unstructured":"Okuda, R., Kajiwara, Y., & Terashima, K. (2014). A survey of technical trend of adas and autonomous driving. Technical papers of 2014 international symposium on VLSI design, automation and test (pp. 1\u20134).","DOI":"10.1109\/VLSI-DAT.2014.6834940"},{"key":"6232_CR56","unstructured":"OpenAI (2018). Openai five. https:\/\/blog.openai.com\/openai-ve\/."},{"key":"6232_CR57","doi-asserted-by":"crossref","unstructured":"Pajarinen, J., Thai, H.L., Akrour, R., Peters, J., & Neumann, G. (2019). Compatible natural gradient policy search. arXiv preprint arXiv:1902.02823 .","DOI":"10.1007\/s10994-019-05807-0"},{"key":"6232_CR58","first-page":"1188","volume":"108","author":"M Papini","year":"2020","unstructured":"Papini, M., Battistello, A., & Restelli, M. (2020). Balancing learning speed and stability in policy gradient via adaptive exploration. AISTATS, 108, 1188\u20131199. PMLR.","journal-title":"AISTATS"},{"key":"6232_CR59","first-page":"4023","volume":"80","author":"M Papini","year":"2018","unstructured":"Papini, M., Binaghi, D., Canonaco, G., Pirotta, M., & Restelli, M. (2018). Stochastic variance-reduced policy gradient. ICML, 80, 4023\u20134032. JMLR.org.","journal-title":"ICML"},{"key":"6232_CR60","unstructured":"Papini, M., Pirotta, M., & Restelli, M. (2017). Adaptive batch size for safe policy gradients. In Advances in neural information processing systems (pp. 3591\u20133600)."},{"key":"6232_CR61","unstructured":"Paul, S., Kurin, V., & Whiteson, S. (2019). Fast efficient hyperparameter tuning for policy gradients. CoRR. https:\/\/arxiv.org\/abs\/1902.06583."},{"key":"6232_CR62","doi-asserted-by":"crossref","unstructured":"Pecka, M., & Svoboda, T. (2014). Safe exploration techniques for reinforcement learning-an overview. In: International workshop on modelling and simulation for autonomous systems (pp. 357\u2013375).","DOI":"10.1007\/978-3-319-13823-7_31"},{"key":"6232_CR63","unstructured":"Peters, J. (2002). Policy gradient methods for control applications (Tech. Rep.). Technical Report TR-CLMC-2007-1,. University of Southern California."},{"issue":"4","key":"6232_CR64","doi-asserted-by":"publisher","first-page":"682","DOI":"10.1016\/j.neunet.2008.02.003","volume":"21","author":"J Peters","year":"2008","unstructured":"Peters, J., & Schaal, S. (2008). Reinforcement learning of motor skills with policy gradients. Neural Networks, 21(4), 682\u2013697.","journal-title":"Neural Networks"},{"key":"6232_CR65","first-page":"1394","volume":"26","author":"M Pirotta","year":"2013","unstructured":"Pirotta, M., Restelli, M., & Bascetta, L. (2013). Adaptive step-size for policy gradient methods. Advances in Neural Information Processing Systems, 26, 1394\u20131402.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"2\u20133","key":"6232_CR66","doi-asserted-by":"publisher","first-page":"255","DOI":"10.1007\/s10994-015-5484-1","volume":"100","author":"M Pirotta","year":"2015","unstructured":"Pirotta, M., Restelli, M., & Bascetta, L. (2015). Policy gradient in Lipschitz Markov decision processes. Machine Learning, 100(2\u20133), 255\u2013283.","journal-title":"Machine Learning"},{"key":"6232_CR67","unstructured":"Pirotta, M., Restelli, M., Pecorino, A., & Calandriello, D. (2013). Safe policy iteration. In: International conference on machine learning (pp. 307-315)."},{"key":"6232_CR68","volume-title":"Markov decision processes: Discrete stochastic dynamic programming","author":"ML Puterman","year":"2014","unstructured":"Puterman, M. L. (2014). Markov decision processes: Discrete stochastic dynamic programming. Wiley."},{"key":"6232_CR69","doi-asserted-by":"crossref","unstructured":"Recht, B. (2019). A tour of reinforcement learning: The view from continuous control. Annual Review of Control, Robotics, and Autonomous Systems.","DOI":"10.1146\/annurev-control-053018-023825"},{"key":"6232_CR70","doi-asserted-by":"publisher","first-page":"400","DOI":"10.1214\/aoms\/1177729586","volume":"22","author":"H Robbins","year":"1951","unstructured":"Robbins, H., & Monro, S. (1951). A stochastic approximation method. The Annals of Mathematical Statistics, 22, 400\u2013407.","journal-title":"The Annals of Mathematical Statistics"},{"key":"6232_CR71","first-page":"1889","volume":"37","author":"J Schulman","year":"2015","unstructured":"Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., & Moritz, P. (2015). Trust region policy optimization. ICML, 37, 1889\u20131897. JMLR.org.","journal-title":"ICML"},{"key":"6232_CR72","unstructured":"Shamir, O. (2011). A variant of Azuma\u2019s inequality for martingales with subGaussian tails. CoRR. https:\/\/arxiv.org\/abs\/1110.2392."},{"key":"6232_CR73","doi-asserted-by":"crossref","unstructured":"Shani, L., Efroni, Y., & Mannor, S. (2020). Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps. AAAI (pp. 5668\u20135675). AAAI Press.","DOI":"10.1609\/aaai.v34i04.6021"},{"key":"6232_CR74","first-page":"5729","volume":"97","author":"Z Shen","year":"2019","unstructured":"Shen, Z., Ribeiro, A., Hassani, H., Qian, H., & Mi, C. (2019). Hessian aided policy gradient. ICML, 97, 5729\u20135738. PMLR.","journal-title":"ICML"},{"issue":"6419","key":"6232_CR75","doi-asserted-by":"publisher","first-page":"1140","DOI":"10.1126\/science.aar6404","volume":"362","author":"D Silver","year":"2018","unstructured":"Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., et al. (2018). A general reinforcement learning algorithm that masters chess, shogi, and go through self-play. Science, 362(6419), 1140\u20131144.","journal-title":"Science"},{"key":"6232_CR76","volume-title":"Reinforcement learning: An introduction","author":"RS Sutton","year":"2018","unstructured":"Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction. MIT Press."},{"key":"6232_CR77","unstructured":"Sutton, R.S., McAllester, D.A., Singh, S.P., & Mansour, Y. (2000). Policy gradient methods for reinforcement learning with function approximation. Advances in neural information processing systems (pp. 1057\u20131063)."},{"key":"6232_CR78","doi-asserted-by":"crossref","unstructured":"Tan, J., Zhang, T., Coumans, E., Iscen, A., Bai, Y., Hafner, D., . . . & Vanhoucke, V. (2018). Sim-to-real: Learning agile locomotion for quadruped robots. arXiv preprint arXiv:1804.10332 .","DOI":"10.15607\/RSS.2018.XIV.010"},{"issue":"6468","key":"6232_CR79","doi-asserted-by":"publisher","first-page":"999","DOI":"10.1126\/science.aag3311","volume":"366","author":"PS Thomas","year":"2019","unstructured":"Thomas, P. S., da Silva, B. C., Barto, A. G., Giguere, S., Brun, Y., & Brunskill, E. (2019). Preventing undesirable behavior of intelligent machines. Science, 366(6468), 999\u20131004.","journal-title":"Science"},{"key":"6232_CR80","first-page":"2380","volume":"37","author":"PS Thomas","year":"2015","unstructured":"Thomas, P. S., Theocharous, G., & Ghavamzadeh, M. (2015). High confidence policy improvement. ICML, 37, 2380\u20132388. JMLR.org.","journal-title":"ICML"},{"key":"6232_CR81","unstructured":"Tucker, G., Bhupatiraju, S., Gu, S., Turner, R., Ghahramani, Z., & Levine, S. (2018). The mirage of action-dependent baselines in reinforcement learning. In International conference on machine learning (pp. 5015\u20135024)."},{"key":"6232_CR82","unstructured":"Turchetta, M., Berkenkamp, F., & Krause, A. (2016). Safe exploration in finite markov decision processes with Gaussian processes. Advances in neural information processing systems (pp. 4312\u20134320)."},{"key":"6232_CR83","unstructured":"Vinyals, O., Babuschkin, I., Chung, J., Mathieu, M., Jaderberg, M., Czarnecki, W.M., . . . & Silver, D. (2019). AlphaStar: Mastering the real-time strategy game starCraft II. https:\/\/deepmind.com\/blog\/alphastar -mastering-real-time-strategy-game-starcraft-ii\/."},{"key":"6232_CR84","unstructured":"Wagner, P. (2011). A reinterpretation of the policy oscillation phenomenon in approximate policy iteration. Advances in neural information processing systems (pp. 2573\u20132581)."},{"issue":"3\u20134","key":"6232_CR85","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1007\/BF00992696","volume":"8","author":"RJ Williams","year":"1992","unstructured":"Williams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3\u20134), 229\u2013256.","journal-title":"Machine Learning"},{"key":"6232_CR86","first-page":"1254","volume":"48","author":"Y Wu","year":"2016","unstructured":"Wu, Y., Shari, R., Lattimore, T., & Szepesvari, C. (2016). Conservative bandits. ICML, 48, 1254\u20131262. JMLR.org.","journal-title":"ICML"},{"key":"6232_CR87","unstructured":"Xu, P., Gao, F., & Gu, Q. (2020). Sample efficient policy gradient methods with recursive variance reduction. ICLR: OpenReview.net."},{"key":"6232_CR88","unstructured":"Yu, J., Aberdeen, D., & Schraudolph, N.N. (2006). Fast online policy gradient learning with SMD gain vector adaptation. Advances in neural information processing systems (pp. 1185\u20131192)."},{"key":"6232_CR89","unstructured":"Yuan, H., Lian, X., Liu, J., & Zhou, Y. (2020). Stochastic recursive momentum for policy gradient methods. CoRR. https:\/\/arxiv.org\/abs\/2003.04302."},{"key":"6232_CR90","unstructured":"Yuan, R., Gower, R.M., & Lazaric, A. (2021). A general sample complexity analysis of vanilla policy gradient. CoRR. https:\/\/arxiv.org\/abs\/2107.11433."},{"key":"6232_CR91","unstructured":"Zhang, J., Kim, J., O\u2019Donoghue, B., & Boyd, S.P. (2020). Sample efficient reinforcement learning with REINFORCE. CoRR. https:\/\/arxiv.org\/abs\/2010.11364."},{"key":"6232_CR92","unstructured":"Zhao, T., Hachiya, H., Niu, G., & Sugiyama, M. (2011). Analysis and improvement of policy gradient estimation. NIPS (pp. 262\u2013270)."}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06232-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-022-06232-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06232-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,11,23]],"date-time":"2022-11-23T19:18:46Z","timestamp":1669231126000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-022-06232-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,20]]},"references-count":92,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2022,11]]}},"alternative-id":["6232"],"URL":"https:\/\/doi.org\/10.1007\/s10994-022-06232-6","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,10,20]]},"assertion":[{"value":"18 November 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 May 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 August 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 October 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"There are no known conflicts of interest associated with this publication.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval:"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate:"}},{"value":"Not applicable.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication:"}},{"value":"custom code was used for the numerical evaluations and is available at .","order":6,"name":"Ethics","group":{"name":"EthicsHeading","label":"Code availability"}}]}}