{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,11]],"date-time":"2026-08-11T07:30:34Z","timestamp":1786433434482,"version":"build-2736575974"},"reference-count":55,"publisher":"Society for Industrial & Applied Mathematics (SIAM)","issue":"3","funder":[{"DOI":"10.13039\/100000181","name":"Air Force Office of Scientific Research","doi-asserted-by":"publisher","award":["FA9550-22-1-0447"],"award-info":[{"award-number":["FA9550-22-1-0447"]}],"id":[{"id":"10.13039\/100000181","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000015","name":"U.S. Department of Energy","doi-asserted-by":"publisher","award":["DE-SC0022158"],"award-info":[{"award-number":["DE-SC0022158"]}],"id":[{"id":"10.13039\/100000015","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["SIAM J. Optim."],"published-print":{"date-parts":[[2026,9,30]]},"abstract":"<jats:p>Abstract.<\/jats:p>\n                  <jats:p>Reinforcement learning (RL) problems over general\u00a0state and action spaces are notoriously challenging. In contrast to the tableau setting, one cannot enumerate all the states and then iteratively update the policies for each state. This prevents the application of many well-studied RL methods, especially those with provable convergence guarantees. In this paper, we first present a substantial generalization of the recently developed policy mirror descent method to deal with general\u00a0state and action spaces. We introduce new approaches to incorporate function approximation into this method so that we do not need to use explicit policy parameterization at all. Moreover, we present a novel policy dual averaging method for which possibly simpler function approximation techniques can be applied. We establish a linear convergence rate to global optimality or sublinear convergence to stationarity for these methods applied to solve different classes of RL problems under exact policy evaluation. We then define proper notions of approximation errors for policy evaluation and investigate their impact on the convergence of these methods applied to general state RL problems with either finite-action or continuous-action spaces. To the best of our knowledge, the development of these algorithmic frameworks as well as their convergence analysis appear to be new in the literature. Preliminary numerical results demonstrate the robustness of the aforementioned methods and show they can be competitive with state-of-the-art RL algorithms.<\/jats:p>","DOI":"10.1137\/25m1791585","type":"journal-article","created":{"date-parts":[[2026,8,11]],"date-time":"2026-08-11T07:00:36Z","timestamp":1786431636000},"page":"1731-1772","source":"Crossref","is-referenced-by-count":0,"title":["Policy Optimization over General State and Action Spaces"],"prefix":"10.1137","volume":"36","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9532-9030","authenticated-orcid":true,"given":"Caleb","family":"Ju","sequence":"first","affiliation":[{"name":"H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology, Atlanta, GA 30332 USA."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guanghui","family":"Lan","sequence":"additional","affiliation":[{"name":"H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology, Atlanta, GA 30332 USA."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"351","published-online":{"date-parts":[[2026,8,11]]},"reference":[{"key":"ref1","unstructured":"Y. Abbasi-Yadkori, P. Bartlett, K. Bhatia, N. Lazic, C. Szepesvari, and G. Weisz, POLITEX: Regret bounds for policy iteration using expert prediction, in International Conference on Machine Learning, PMLR, 2019, pp. 3692\u20133702."},{"key":"ref2","first-page":"1","volume":"22","author":"Agarwal A.","year":"2021","journal-title":"J. Mach. Learn. Res."},{"key":"ref3","unstructured":"Z. Ahmed, N. Le Roux, M. Norouzi, and D. Schuurmans, Understanding the impact of entropy on policy optimization, in International Conference on Machine Learning, PMLR, 2019, pp. 151\u2013160."},{"key":"ref4","doi-asserted-by":"publisher","DOI":"10.1109\/TSMC.1983.6313077"},{"key":"ref5","volume-title":"Stochastic Optimal Control: The Discrete-time Case","volume":"5","author":"Bertsekas D.","year":"1996"},{"key":"ref6","volume-title":"Stochastic Optimal Control: The Discrete-Time Case","author":"Bertsekas D. P.","year":"1996","edition":"2"},{"key":"ref7","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2021.0014"},{"key":"ref8","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4939-7049-0"},{"key":"ref9","volume-title":"Perturbation Analysis of Optimization Problems","author":"Bonnans J. F.","year":"2013"},{"key":"ref10","doi-asserted-by":"publisher","DOI":"10.1137\/22M1540156"},{"key":"ref11","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2021.2151"},{"key":"ref12","first-page":"809","volume":"15","author":"Dann C.","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref13","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-018-1311-3"},{"key":"ref14","doi-asserted-by":"publisher","DOI":"10.1287\/moor.1090.0396"},{"key":"ref15","doi-asserted-by":"crossref","unstructured":"A.m. Farahmand, S. Nabi, and D. N. Nikovski, Deep reinforcement learning for partial differential equation control, in 2017 American Control Conference (ACC), IEEE, 2017, pp. 3120\u20133127.","DOI":"10.23919\/ACC.2017.7963427"},{"key":"ref16","doi-asserted-by":"publisher","DOI":"10.1007\/BF02055196"},{"key":"ref17","unstructured":"M. Geist, B. Scherrer, and O. Pietquin, A theory of regularized Markov decision processes, in International Conference on Machine Learning, PMLR, 2019, pp. 2160\u20132169."},{"key":"ref18","series-title":"Appl. Math. (N.Y.) 30","volume-title":"Discrete-time Markov Control Processes: Basic Optimality Criteria","author":"Hern\u00e1ndez-Lerma O.","year":"2012"},{"key":"ref19","doi-asserted-by":"publisher","DOI":"10.1137\/23M1554771"},{"key":"ref20","unstructured":"S. Kakade and J. Langford, Approximately optimal approximate reinforcement learning, in International Conference on Machine Learning, Morgan Kaufmann, 2002, pp. 267\u2013274."},{"key":"ref21","volume":"14","author":"Kakade S. M.","year":"2001","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"ref22","doi-asserted-by":"crossref","unstructured":"S. Khodadadian, P. R. Jhunjhunwala, S. M. Varma, and S. T. Maguluri, On the linear convergence of natural policy gradient algorithm, in 2021 60th IEEE Conference on Decision and Control (CDC), IEEE, 2021, pp. 3794\u20133799.","DOI":"10.1109\/CDC45484.2021.9682908"},{"key":"ref23","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-39568-1"},{"key":"ref24","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-022-01816-5"},{"key":"ref25","doi-asserted-by":"publisher","DOI":"10.1137\/22M1480409"},{"key":"ref26","first-page":"1","author":"Li T.","year":"2025","journal-title":"Math. Program."},{"key":"ref27","doi-asserted-by":"publisher","DOI":"10.1287\/moor.2022.0241"},{"key":"ref28","doi-asserted-by":"publisher","DOI":"10.1137\/23M1560215"},{"key":"ref29","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-023-02017-4"},{"key":"ref30","unstructured":"Y. Li, T. Zhao, and G. Lan, First-Order Policy Optimization for Robust Markov Decision Process, preprint, arXiv:2209.10579, 2023."},{"key":"ref31","unstructured":"J. Mei, C. Xiao, C. Szepesvari, and D. Schuurmans, On the global convergence rates of softmax policy gradient methods, in International Conference on Machine Learning, PMLR, 2020, pp. 6820\u20136829."},{"key":"ref32","volume":"7","author":"Micchelli C. A.","year":"2006","journal-title":"J. Mach. Learn. Res."},{"key":"ref33","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"ref34","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-007-0149-x"},{"key":"ref35","series-title":"Appl. Optim. 87","volume-title":"Introductory Lectures on Convex Optimization: A Basic Course","author":"Nesterov Y.","year":"2013"},{"key":"ref36","unstructured":"G. Neu, A. Jonsson, and V. G\u00f3mez, A Unified View of Entropy-Regularized Markov Decision Processes, preprint, arXiv:1705.07798, 2017."},{"key":"ref37","doi-asserted-by":"publisher","DOI":"10.1002\/9780470316887"},{"key":"ref38","first-page":"1","volume":"22","author":"Raffin A.","year":"2021","journal-title":"J. Mach. Learn. Res."},{"key":"ref39","volume":"20","author":"Rahimi A.","year":"2007","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"ref40","volume":"30","author":"Rudi A.","year":"2017","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"ref41","unstructured":"J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, Proximal Policy Optimization Algorithms, preprint, arXiv:1707.06347, 2017."},{"key":"ref42","doi-asserted-by":"crossref","unstructured":"L. Shani, Y. Efroni, and S. Mannor, Adaptive trust region policy optimization: Global convergence and faster rates for regularized MDPs, in The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, AAAI Press, 2020, pp. 5668\u20135675, https:\/\/aaai.org\/ojs\/index.php\/AAAI\/article\/view\/6021.","DOI":"10.1609\/aaai.v34i04.6021"},{"key":"ref43","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611976595"},{"key":"ref44","doi-asserted-by":"publisher","DOI":"10.1137\/19M129406X"},{"key":"ref45","doi-asserted-by":"publisher","DOI":"10.1126\/science.aar6404"},{"key":"ref46","doi-asserted-by":"publisher","DOI":"10.1038\/nature24270"},{"key":"ref47","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton R. S.","year":"2018","edition":"2"},{"key":"ref48","unstructured":"R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, Policy gradient methods for reinforcement learning with function approximation, in Proceedings of the 13th International Conference on Neural Information Processing Systems, MIT Press, 1999, pp. 1057\u20131063."},{"key":"ref49","doi-asserted-by":"crossref","unstructured":"Y. Tassa, T. Erez, and E. Todorov, Synthesis and stabilization of complex behaviors through online trajectory optimization, in 2012 IEEE\/RSJ International Conference on Intelligent Robots and Systems, IEEE, 2012, pp. 4906\u20134913.","DOI":"10.1109\/IROS.2012.6386025"},{"key":"ref50","doi-asserted-by":"publisher","DOI":"10.1007\/BF01016429"},{"key":"ref51","unstructured":"L. Wang, Q. Cai, Z. Yang, and Z. Wang, Neural Policy Gradient Methods: Global Optimality and Rates of Convergence, preprint, arXiv:1909.01150, 2020."},{"key":"ref52","first-page":"2543","author":"Xiao L.","year":"2010","journal-title":"J. Mach. Learn. Res."},{"key":"ref53","first-page":"1","volume":"23","author":"Xiao L.","year":"2022","journal-title":"J. Mach. Learn. Res."},{"key":"ref54","first-page":"4358","volume":"33","author":"Xu T.","year":"2020","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"ref55","doi-asserted-by":"publisher","DOI":"10.1137\/21M1456789"}],"container-title":["SIAM Journal on Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/epubs.siam.org\/doi\/pdf\/10.1137\/25M1791585","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,11]],"date-time":"2026-08-11T07:00:49Z","timestamp":1786431649000},"score":1,"resource":{"primary":{"URL":"https:\/\/epubs.siam.org\/doi\/10.1137\/25M1791585"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,8,11]]},"references-count":55,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,9,30]]}},"alternative-id":["10.1137\/25M1791585"],"URL":"https:\/\/doi.org\/10.1137\/25m1791585","relation":{},"ISSN":["1052-6234","1095-7189"],"issn-type":[{"value":"1052-6234","type":"print"},{"value":"1095-7189","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,8,11]]}}}