{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T07:32:47Z","timestamp":1740123167055,"version":"3.37.3"},"reference-count":72,"publisher":"Springer Science and Business Media LLC","issue":"11","license":[{"start":{"date-parts":[[2023,8,24]],"date-time":"2023-08-24T00:00:00Z","timestamp":1692835200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,8,24]],"date-time":"2023-08-24T00:00:00Z","timestamp":1692835200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2023,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In this paper, we propose cautious policy programming (CPP), a novel value-based reinforcement learning (RL) algorithm that exploits the idea of monotonic policy improvement during learning. Based on the nature of entropy-regularized RL, we derive a new entropy-regularization-aware lower bound of policy improvement that depends on the expected policy advantage function but not on state-action-space-wise maximization as in prior work. CPP leverages this lower bound as a criterion for adjusting the degree of a policy update for alleviating policy oscillation. Different from similar algorithms that are mostly theory-oriented, we also propose a novel interpolation scheme that makes CPP better scale in high dimensional control problems. We demonstrate that the proposed algorithm can trade off performance and stability in both didactic classic control problems and challenging high-dimensional Atari games.<\/jats:p>","DOI":"10.1007\/s10994-023-06368-z","type":"journal-article","created":{"date-parts":[[2023,8,24]],"date-time":"2023-08-24T21:01:26Z","timestamp":1692910886000},"page":"4527-4562","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Cautious policy programming: exploiting KL regularization for monotonic policy improvement in reinforcement learning"],"prefix":"10.1007","volume":"112","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9514-6760","authenticated-orcid":false,"given":"Lingwei","family":"Zhu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Takamitsu","family":"Matsubara","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,8,24]]},"reference":[{"key":"6368_CR1","unstructured":"Abbasi-Yadkori, Y., Bartlett, P.\u00a0L., & Wright, S.\u00a0J. (2016). A fast and reliable policy improvement algorithm. In Proceedings of the 19th International conference on artificial intelligence and statistics, Proceedings of machine learning research, (vol.\u00a051 , pp. 1338\u20131346)."},{"key":"6368_CR2","unstructured":"Agarwal, A., Kakade, S.\u00a0M., Lee, J.\u00a0D., & Mahajan, G. (2020). Optimality and approximation with policy gradient methods in markov decision processes. In Proceedings of thirty third conference on learning theory, Proceedings of machine learning research, (vol. 125, pp. 64\u201366)."},{"issue":"98","key":"6368_CR3","first-page":"1","volume":"22","author":"A Agarwal","year":"2021","unstructured":"Agarwal, A., Kakade, S. M., Lee, J. D., & Mahajan, G. (2021). On the theory of policy gradient methods: Optimality, approximation, and distribution shift. Journal of Machine Learning Research, 22(98), 1\u201376.","journal-title":"Journal of Machine Learning Research"},{"key":"6368_CR4","unstructured":"Ahmed, Z., Le\u00a0Roux, N., Norouzi, M., & Schuurmans, D. (2019). Understanding the impact of entropy on policy optimization. In Proceedings of 36th international conference on machine learning, (vol.\u00a097, pp. 151\u2013160)."},{"key":"6368_CR5","doi-asserted-by":"crossref","unstructured":"Akiba, T., Sano, S., Yanase, T., Ohta, T., & Koyama, M. (2019). Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery data mining, KDD \u201919, (pp. 2623\u20132631) Association for Computing Machinery.","DOI":"10.1145\/3292500.3330701"},{"issue":"14","key":"6368_CR6","first-page":"1","volume":"19","author":"R Akrour","year":"2018","unstructured":"Akrour, R., Abdolmaleki, A., Abdulsamad, H., Peters, J., & Neumann, G. (2018). Model-free trajectory-based policy optimization with monotonic improvement. Journal of Machine Learning Research, 19(14), 1\u201325.","journal-title":"Journal of Machine Learning Research"},{"issue":"1","key":"6368_CR7","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1177\/0278364919887447","volume":"39","author":"OM Andrychowicz","year":"2020","unstructured":"Andrychowicz, O. M., Baker, B., & Chociej, M. Others. (2020). Learning dexterous in-hand manipulation. The International Journal of Robotics Research, 39(1), 3\u201320.","journal-title":"The International Journal of Robotics Research"},{"key":"6368_CR8","unstructured":"Asadi, K. & Littman, M.\u00a0L. (2017). An alternative softmax operator for reinforcement learning. In Proceedings of the 34th International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, (vol.\u00a070, pp. 243\u2013252) International Convention Centre, 06\u201311 Aug 2017. PMLR."},{"issue":"1","key":"6368_CR9","first-page":"3207","volume":"13","author":"MG Azar","year":"2012","unstructured":"Azar, M. G., G\u00f3mez, V., & Kappen, H. J. (2012). Dynamic policy programming. Journal of Machine Learning Research, 13(1), 3207\u20133245.","journal-title":"Journal of Machine Learning Research"},{"key":"6368_CR10","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611974997","volume-title":"First-Order Methods in Optimization","author":"A Beck","year":"2017","unstructured":"Beck, A. (2017). First-Order Methods in Optimization. Society for Industrial and Applied Mathematics."},{"issue":"1","key":"6368_CR11","doi-asserted-by":"publisher","first-page":"253","DOI":"10.1613\/jair.3912","volume":"47","author":"MG Bellemare","year":"2013","unstructured":"Bellemare, M. G., Naddaf, Y., Veness, J., & Bowling, M. (2013). The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research, 47(1), 253\u2013279. ISSN 1076-9757.","journal-title":"Journal of Artificial Intelligence Research"},{"key":"6368_CR12","doi-asserted-by":"publisher","first-page":"310","DOI":"10.1007\/s11768-011-1005-3","volume":"9","author":"D Bertsekas","year":"2011","unstructured":"Bertsekas, D. (2011). Approximate policy iteration: A survey and some new methods. Journal of Control Theory and Applications, 9, 310\u2013335.","journal-title":"Journal of Control Theory and Applications"},{"key":"6368_CR13","unstructured":"Bertsekas, D.\u00a0P. (2005). Dynamic programming and optimal control. ISBN 1886529264."},{"key":"6368_CR14","volume-title":"Neuro-dynamic programming","author":"DP Bertsekas","year":"1996","unstructured":"Bertsekas, D. P., & Tsitsiklis, J. N. (1996). Neuro-dynamic programming (1st ed.). Athena Scientific.","edition":"1"},{"key":"6368_CR15","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511804441","volume-title":"Convex Optimization","author":"S Boyd","year":"2004","unstructured":"Boyd, S., & Vandenberghe, L. (2004). Convex Optimization. Cambridge University Press."},{"key":"6368_CR16","doi-asserted-by":"publisher","first-page":"342","DOI":"10.1007\/BFb0064610","volume":"12","author":"J Bretagnolle","year":"1978","unstructured":"Bretagnolle, J., & Huber, C. (1978). Estimation des densit\u00e9s\u202f: risque minimax. S\u00e9minaire de Probabilit\u00e9s de Strasbourg, 12, 342\u2013363.","journal-title":"S\u00e9minaire de Probabilit\u00e9s de Strasbourg"},{"issue":"4","key":"6368_CR17","doi-asserted-by":"publisher","first-page":"2563","DOI":"10.1287\/opre.2021.2151","volume":"70","author":"S Cen","year":"2022","unstructured":"Cen, S., Cheng, C., Chen, Y., Wei, Y., & Chi, Y. (2022). Fast global convergence of natural policy gradient methods with entropy regularization. Operations Research, 70(4), 2563\u20132578.","journal-title":"Operations Research"},{"issue":"253","key":"6368_CR18","first-page":"1","volume":"23","author":"A Chan","year":"2022","unstructured":"Chan, A., Silva, H., Lim, S., Kozuno, T., Mahmood, A. R., & White, M. (2022). Greedification operators for policy optimization: Investigating forward and reverse kl divergences. Journal of Machine Learning Research, 23(253), 1\u201379.","journal-title":"Journal of Machine Learning Research"},{"key":"6368_CR19","doi-asserted-by":"publisher","first-page":"13","DOI":"10.1016\/j.neunet.2017.06.007","volume":"94","author":"Y Cui","year":"2017","unstructured":"Cui, Y., Matsubara, T., & Sugimoto, K. (2017). Kernel dynamic policy programming: Applicable reinforcement learning to robot systems with high dimensional states. Neural Networks, 94, 13\u201323.","journal-title":"Neural Networks"},{"key":"6368_CR20","unstructured":"Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L. & Madry, A. (2019). Implementation matters in deep policy gradients: A case study on ppo and trpo. In International conference on learning representations (ICLR), (pp. 1\u201312)."},{"key":"6368_CR21","unstructured":"Fortunato, M., Azar, M.\u00a0G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., Blundell, C. & Legg, S. (2018). Noisy networks for exploration. In Proceedings of the international conference on representation learning (ICLR 2018)"},{"key":"6368_CR22","unstructured":"Fox, R., Pakman, A. & Tishby, N. (2016). Taming the noise in reinforcement learning via soft updates. In Proceedings of the thirty-second conference on uncertainty in artificial intelligence, (pp. 202\u2013211)."},{"key":"6368_CR23","unstructured":"Fu, J., Kumar, A., Soh, M. & Levine, S. (2019). Diagnosing bottlenecks in deep q-learning algorithms. vol.\u00a097 of Proceedings of 36th international conference on machine learning, (pp. 2021\u20132030)."},{"key":"6368_CR24","unstructured":"Fujimoto, S., van Hoof, H. & Meger, D. (2018). Addressing function approximation error in actor-critic methods. In Proceedings of the 35th international conference on machine learning, (vol.\u00a080, pp. 1587\u20131596)."},{"key":"6368_CR25","unstructured":"Geist, M., Scherrer, B. & Pietquin, O. (2019). A theory of regularized Markov decission processes. In 36th international conference on machine learning, (vol.\u00a097, pp. 2160\u20132169)."},{"key":"6368_CR26","unstructured":"Haarnoja, T., Tang, H., Abbeel, P., & Levine, S. (2017). Reinforcement learning with deep energy-based policies. In 34th international conference on machine learning, (pp. 1352\u20131361)."},{"key":"6368_CR27","unstructured":"Haarnoja, T., Zhou, A., Abbeel, P. & Levine, S. (2018). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th international conference on machine learning, (pp. 1861\u20131870)."},{"key":"6368_CR28","unstructured":"Kakade, S. & Langford, J. (2002). Approximately optimal approximate reinforcement learning. In 19th international conference on machine learning (ICML), (pp. 267\u2013274)."},{"key":"6368_CR29","unstructured":"Kingma, D.\u00a0P. & Ba, J. (2015). Adam: A method for stochastic optimization. In 3rd international conference on learning representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference track proceedings."},{"key":"6368_CR30","unstructured":"Kozuno, T., Uchibe, E., & Doya, K. (2019). Theoretical analysis of efficiency and robustness of softmax and gap-increasing operators in reinforcement learning. In proceedings of the twenty-second international conference on artificial intelligence and statistics, (vol.\u00a097, pp. 2995\u20133003)."},{"key":"6368_CR31","unstructured":"Kozuno, T., Yang, W., Vieillard, N., Kitamura, T., Tang, Y., Mei, J., M\u00e9nard, P., Azar, M.\u00a0G., Valko, M. Munos, R., Pietquin, O., Geist, M., & Szepesv\u00e1ri, C. (2022). Kl-entropy-regularized rl with a generative model is minimax optimal. arXiv:2205.14211."},{"issue":"44","key":"6368_CR32","first-page":"1107","volume":"4","author":"MG Lagoudakis","year":"2003","unstructured":"Lagoudakis, M. G., & Parr, R. (2003). Least-squares policy iteration. The Journal of Machine Learning Research, 4(44), 1107\u20131149.","journal-title":"The Journal of Machine Learning Research"},{"issue":"19","key":"6368_CR33","first-page":"1","volume":"17","author":"A Lazaric","year":"2016","unstructured":"Lazaric, A., Ghavamzadeh, M., & Munos, R. (2016). Analysis of classification-based policy iteration algorithms. Journal of Machine Learning Research, 17(19), 1\u201330.","journal-title":"Journal of Machine Learning Research"},{"key":"6368_CR34","doi-asserted-by":"crossref","unstructured":"Mei, J., Xiao, C., Huang, R., Schuurmans, D., & M\u00fcller, M. (2019). On principled entropy exploration in policy optimization. In Proceedings of the twenty-eighth international joint conference on artificial intelligence, IJCAI-19, (pp. 3130\u20133136).","DOI":"10.24963\/ijcai.2019\/434"},{"key":"6368_CR35","unstructured":"Mei, J., Xiao, C., Szepesvari, C., & Schuurmans, D. (2020). On the global convergence rates of softmax policy gradient methods. In Proceedings of the 37th international conference on machine learning, (vol. 119, pp. 6820\u20136829)."},{"key":"6368_CR36","unstructured":"Metelli, A.\u00a0M., Mutti, M. & Restelli, M. (2018). Configurable Markov decision processes. In Proceedings of the 35th international conference on machine learning, proceedings of machine learning research, (vol.\u00a080, pp. 3491\u20133500)."},{"issue":"7540","key":"6368_CR37","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih, V., Kavukcuoglu, K., & Silver, D. Others. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529\u2013533.","journal-title":"Nature"},{"key":"6368_CR38","unstructured":"Mnih, V., Badia, A.\u00a0P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D. & Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. In Proceedings of The 33rd international conference on machine learning, (vol.\u00a048, pp. 1928\u20131937)."},{"key":"6368_CR39","unstructured":"Munos, R. (2005). Error bounds for approximate value iteration. In Proceedings of the 20th national conference on artificial intelligence - Volume 2, AAAI\u201905, (pp. 1006\u20131011)."},{"key":"6368_CR40","first-page":"2775","volume":"30","author":"O Nachum","year":"2017","unstructured":"Nachum, O., Norouzi, M., Xu, K., & Schuurmans, D. (2017). Bridging the gap between value and policy based reinforcement learning. In Advances in Neural Information Processing Systems, 30, 2775\u20132785.","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"6368_CR41","unstructured":"Nachum, O., Norouzi, M., Xu, K. & Schuurmans, D. (2018). Trust-pcl: An off-policy trust region method for continuous control. In International conference on learning representations, (pp. 1\u201311)."},{"key":"6368_CR42","unstructured":"Neu, G., Jonsson, A. & G\u00f3mez, V. (2017). A unified view of entropy-regularized markov decision processes. arXiv:1705.07798."},{"key":"6368_CR43","unstructured":"O\u2019Donoghue, B., Munos, R., Kavukcuoglu, K., & Mnih, V. (2016). PGQ: Combining policy gradient and Q-learning. arXiv preprint arXiv:1611.01626."},{"key":"6368_CR44","first-page":"1","volume":"30","author":"M Papini","year":"2017","unstructured":"Papini, M., Pirotta, M., & Restelli, M. (2017). Adaptive batch size for safe policy gradients. In Advances in Neural Information Processing Systems, 30, 1\u201310.","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"6368_CR45","unstructured":"Papini, M., Battistello, A., & Restelli, M. (2020). Balancing learning speed and stability in policy gradient via adaptive exploration. In Proceedings of the twenty third international conference on artificial intelligence and statistics, proceedings of machine learning research, (vol. 108, pp. 1188\u20131199)."},{"key":"6368_CR46","first-page":"1","volume":"26","author":"M Pirotta","year":"2013","unstructured":"Pirotta, M., Restelli, M., & Bascetta, L. (2013a). Adaptive step-size for policy gradient methods. In Advances in Neural Information Processing Systems, 26, 1\u20139.","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"6368_CR47","unstructured":"Pirotta, M., Restelli, M., Pecorino, A. & Calandriello, D. (2013b). Safe policy iteration. In Proceedings of the 30th international conference on machine learning, (vol.\u00a028, pp. 307\u2013315)."},{"key":"6368_CR48","first-page":"80","volume-title":"Eligibility traces for off-policy policy evaluation","author":"D Precup","year":"2000","unstructured":"Precup, D. (2000). Eligibility traces for off-policy policy evaluation (p. 80). Computer Science Department Faculty Publication Series."},{"key":"6368_CR49","doi-asserted-by":"publisher","DOI":"10.1002\/9780470316887","volume-title":"Markov Decision Processes: Discrete Stochastic Dynamic Programming","author":"ML Puterman","year":"1994","unstructured":"Puterman, M. L. (1994). Markov Decision Processes: Discrete Stochastic Dynamic Programming (1st ed.). Wiley.","edition":"1"},{"issue":"268","key":"6368_CR50","first-page":"1","volume":"22","author":"A Raffin","year":"2021","unstructured":"Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., & Dormann, N. (2021). Stable-baselines3: Reliable reinforcement learning implementations. Journal of Machine Learning Research, 22(268), 1\u20138.","journal-title":"Journal of Machine Learning Research"},{"key":"6368_CR51","doi-asserted-by":"publisher","first-page":"5973","DOI":"10.1109\/TIT.2016.2603151","volume":"62","author":"I Sason","year":"2016","unstructured":"Sason, I., & Verd\u00fa, S. (2016). f-divergence inequalities. IEEE Transactions on Information Theory, 62, 5973\u20136006.","journal-title":"IEEE Transactions on Information Theory"},{"key":"6368_CR52","doi-asserted-by":"crossref","unstructured":"Scherrer, B. & Geist, M. (2014). Local policy search in a convex space and conservative policy iteration as boosted policy search. In Machine learning and knowledge discovery in databases, (pp. 35\u201350).","DOI":"10.1007\/978-3-662-44845-8_3"},{"issue":"1","key":"6368_CR53","first-page":"1629","volume":"16","author":"B Scherrer","year":"2015","unstructured":"Scherrer, B., Ghavamzadeh, M., Gabillon, V., Lesner, B., & Geist, M. (2015). Approximate modified policy iteration and its application to the game of tetris. Journal of Machine Learning Research, 16(1), 1629\u20131676.","journal-title":"Journal of Machine Learning Research"},{"key":"6368_CR54","unstructured":"Schulman, J., Levine, S., Abbeel, P., Jordan, M. & Moritz, P. (2015). Trust region policy optimization. In Proceedings of the 32nd international conference on machine learning, (vol. 37, pp. 1889\u20131897)."},{"key":"6368_CR55","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A. & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv:1707.06347."},{"key":"6368_CR56","unstructured":"Shani, L., Efroni, Y. & Mannor, S. (2019). Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps. CoRR, arXiv:abs\/1909.02769."},{"key":"6368_CR57","volume-title":"Reinforcement learning: An introduction","author":"RS Sutton","year":"2018","unstructured":"Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction. A Bradford Book."},{"key":"6368_CR58","doi-asserted-by":"crossref","unstructured":"Todorov, E. (2006). Linearly-solvable Markov decision problems. In Advances in neural information processing systems (NIPS), (pp. 1369\u20131376).","DOI":"10.7551\/mitpress\/7503.003.0176"},{"key":"6368_CR59","doi-asserted-by":"publisher","first-page":"72","DOI":"10.1016\/j.robot.2018.11.004","volume":"112","author":"Y Tsurumine","year":"2019","unstructured":"Tsurumine, Y., Cui, Y., Uchibe, E., & Matsubara, T. (2019). Deep reinforcement learning with smooth policy update: Application to robotic cloth manipulation. Robotics and Autonomous Systems, 112, 72\u201383.","journal-title":"Robotics and Autonomous Systems"},{"key":"6368_CR60","volume-title":"Introduction to nonparametric estimation","author":"AB Tsybakov","year":"2008","unstructured":"Tsybakov, A. B. (2008). Introduction to nonparametric estimation. Springer."},{"issue":"3","key":"6368_CR61","doi-asserted-by":"publisher","first-page":"891","DOI":"10.1007\/s11063-017-9702-7","volume":"47","author":"E Uchibe","year":"2018","unstructured":"Uchibe, E. (2018). Model-free deep inverse reinforcement learning by logistic regression. Neural Processing Letters, 47(3), 891\u2013905.","journal-title":"Neural Processing Letters"},{"key":"6368_CR62","doi-asserted-by":"publisher","first-page":"138","DOI":"10.1016\/j.neunet.2021.08.017","volume":"144","author":"E Uchibe","year":"2021","unstructured":"Uchibe, E., & Doya, K. (2021). Forward and inverse reinforcement learning sharing network weights and hyperparameters. Neural Networks, 144, 138\u2013153.","journal-title":"Neural Networks"},{"key":"6368_CR63","first-page":"1","volume":"33","author":"N Vieillard","year":"2020","unstructured":"Vieillard, N., Kozuno, T., Scherrer, B., Pietquin, O., Munos, R., & Geist, M. (2020a). Leverage the average: An analysis of regularization in rl. In Advances in Neural Information Processing Systems, 33, 1\u201312.","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"6368_CR64","doi-asserted-by":"crossref","unstructured":"Vieillard, N., Pietquin, O., & Geist, M. (2020b). Deep conservative policy iteration. In The thirty-fourth AAAI conference on artificial intelligence, AAAI\u201920, (pp. 6070\u20136077).","DOI":"10.1609\/aaai.v34i04.6070"},{"key":"6368_CR65","first-page":"1","volume":"33","author":"N Vieillard","year":"2020","unstructured":"Vieillard, N., Pietquin, O., & Geist, M. (2020). Munchausen reinforcement learning. In Advances in Neural Information Processing Systems, 33, 1\u201311.","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"6368_CR66","unstructured":"Vieillard, N., Scherrer, B., Pietquin, O., & Geist, M. (2020). Momentum in reinforcement learning. In Proceedings of the Twenty third international conference on artificial intelligence and statistics, (Vol. 1, pp. 2529\u20132538)."},{"key":"6368_CR67","first-page":"2573","volume":"24","author":"P Wagner","year":"2011","unstructured":"Wagner, P. (2011). A reinterpretation of the policy oscillation phenomenon in approximate policy iteration. In Advances in Neural Information Processing Systems, 24, 2573\u20132581.","journal-title":"In Advances in Neural Information Processing Systems"},{"key":"6368_CR68","unstructured":"Wen, J., Dai, B., Li, L. & Schuurmans, D. (2020). Batch stationary distribution estimation. In Proceedings of the 37th international conference on machine learning, ( vol. 119, pp. 10203\u201310213)"},{"key":"6368_CR69","doi-asserted-by":"publisher","first-page":"593","DOI":"10.1287\/moor.1110.0516","volume":"36","author":"Y Ye","year":"2011","unstructured":"Ye, Y. (2011). The simplex and policy-iteration methods are strongly polynomial for the markov decision problem with a fixed discount rate. Mathematics of Operations Research, 36, 593\u2013603.","journal-title":"Mathematics of Operations Research"},{"key":"6368_CR70","doi-asserted-by":"publisher","first-page":"104331","DOI":"10.1016\/j.conengprac.2020.104331","volume":"97","author":"L Zhu","year":"2020","unstructured":"Zhu, L., Cui, Y., Takami, G., Kanokogi, H., & Matsubara, T. (2020). Scalable reinforcement learning for plant-wide control of vinyl acetate monomer process. Control Engineering Practice, 97, 104331\u2013104340.","journal-title":"Control Engineering Practice"},{"key":"6368_CR71","doi-asserted-by":"publisher","DOI":"10.1016\/j.compchemeng.2022.107658","volume":"158","author":"L Zhu","year":"2022","unstructured":"Zhu, L., Takami, G., Kawahara, M., Kanokogi, H., & Matsubara, T. (2022). Alleviating parameter-tuning burden in reinforcement learning for large-scale process control. Computers & Chemical Engineering, 158, 107658.","journal-title":"Computers & Chemical Engineering"},{"key":"6368_CR72","unstructured":"Ziebart, B.\u00a0D. (2010). Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy. PhD thesis, Carnegie Mellon University."}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-023-06368-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-023-06368-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-023-06368-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,10,27]],"date-time":"2023-10-27T08:03:37Z","timestamp":1698393817000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-023-06368-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,24]]},"references-count":72,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2023,11]]}},"alternative-id":["6368"],"URL":"https:\/\/doi.org\/10.1007\/s10994-023-06368-z","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"type":"print","value":"0885-6125"},{"type":"electronic","value":"1573-0565"}],"subject":[],"published":{"date-parts":[[2023,8,24]]},"assertion":[{"value":"15 March 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 April 2023","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 July 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 August 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declaration"}},{"value":"Lingwei Zhu and Takamitsu Matsubara declared no conflicting interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}