{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T14:39:43Z","timestamp":1787323183870,"version":"build-2736575974"},"reference-count":50,"publisher":"Society for Industrial & Applied Mathematics (SIAM)","issue":"1","funder":[{"name":"Coleman Fung endowment chair fund"},{"name":"QRT"},{"DOI":"10.13039\/501100000266","name":"EPSRC","doi-asserted-by":"crossref","award":["EP\/Y028872\/1"],"award-info":[{"award-number":["EP\/Y028872\/1"]}],"id":[{"id":"10.13039\/501100000266","id-type":"DOI","asserted-by":"crossref"}]},{"name":"NSF","award":["DMS-2524465"],"award-info":[{"award-number":["DMS-2524465"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["SIAM J. Control Optim."],"published-print":{"date-parts":[[2026,2,28]]},"abstract":"<jats:p>Abstract.<\/jats:p>\n                  <jats:p>This paper proposes and analyzes two new policy learning methods,\u00a0regularized policy gradient and iterative policy optimization (IPO), for a class of discounted linear-quadratic control (LQC) problems over an infinite time horizon with entropy regularization. Assuming access to the exact policy evaluation, both proposed approaches are proved to converge linearly in finding optimal policies of the regularized LQC. Moreover, the IPO method can achieve a superlinear convergence rate once it enters a local region around the optimal policy. Finally, when the optimal policy for a reinforcement learning (RL)\u00a0problem with a known environment is appropriately transferred as the initial policy to an RL problem with an unknown environment, the IPO method is shown to converge at a superlinear rate if the two environments are sufficiently close. A model-free version of the policy-based methods is also discussed. Performances of these proposed algorithms are supported by numerical examples.<\/jats:p>","DOI":"10.1137\/23m1621071","type":"journal-article","created":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T08:25:53Z","timestamp":1767947153000},"page":"124-151","source":"Crossref","is-referenced-by-count":0,"title":["Fast Policy Learning for Linear-Quadratic Control with Entropy Regularization"],"prefix":"10.1137","volume":"64","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3350-4606","authenticated-orcid":false,"given":"Xin","family":"Guo","sequence":"first","affiliation":[{"name":"Department of Industrial Engineering and Operations Research, University of California, Berkeley, Berkeley, CA 94720 USA."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-1206-716X","authenticated-orcid":false,"given":"Xinyu","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Mathematics, University of Oxford, Oxford OX2 6GG, UK."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4293-3450","authenticated-orcid":false,"given":"Renyuan","family":"Xu","sequence":"additional","affiliation":[{"name":"Department of Management Science and Engineering, Stanford University, Stanford, CA 94305 USA."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"351","published-online":{"date-parts":[[2026,1,9]]},"reference":[{"key":"ref1","unstructured":"A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan, Optimality and approximation with policy gradient methods in Markov decision processes, in Conference on Learning Theory, PMLR, 2020, pp. 64\u201366."},{"key":"ref2","unstructured":"Z. Ahmed, N. Le Roux, M. Norouzi, and D. Schuurmans, Understanding the impact of entropy on policy optimization, in International Conference on Machine Learning, PMLR, 2019, pp. 151\u2013160."},{"key":"ref3","volume-title":"Advances in Neural Information Processing Systems","volume":"31","author":"Balasubramanian K.","year":"2018"},{"key":"ref4","first-page":"8015","volume":"23","author":"Basei M.","year":"2022","journal-title":"J. Mach. Learn. Res."},{"key":"ref5","first-page":"679","author":"Bellman R.","year":"1957","journal-title":"J. Math. Mech."},{"key":"ref6","volume-title":"Dynamic Programming and Optimal Control","volume":"4","author":"Bertsekas D.","year":"2017"},{"key":"ref7","volume-title":"Neuro-Dynamic Programming","author":"Bertsekas D.","year":"1996"},{"key":"ref8","unstructured":"J. Bhandari and D. Russo, Global Optimality Guarantees for Policy Gradient Methods, preprint, arXiv:1906.01786, 2019."},{"key":"ref9","unstructured":"J. Bu, A. Mesbahi, M. Fazel, and M. Mesbahi, LQR through the Lens of First Order Methods: Discrete-Time Case, preprint, https:\/\/arxiv.org\/abs\/1907.08921, 2019."},{"key":"ref10","unstructured":"J. Bu, A. Mesbahi, and M. Mesbahi, Policy Gradient-Based Algorithms for Continuous-Time Linear Quadratic Control, preprint, arXiv:2006.09178, 2020."},{"key":"ref11","unstructured":"H. Cao, H. Gu, and X. Guo, Feasibility of Transfer Learning: A Mathematical Framework, preprint, arXiv:2305.12985, 2023."},{"key":"ref12","doi-asserted-by":"crossref","unstructured":"H. Cao, H. Gu, X. Guo, and M. Rosenbaum, Risk of Transfer Learning and Its Applications in Finance, preprint, https:\/\/arxiv.org\/abs\/2311.03283, 2023.","DOI":"10.2139\/ssrn.4624427"},{"key":"ref13","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2021.2151"},{"key":"ref14","first-page":"14620","volume":"34","author":"Cheng S.","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref15","unstructured":"M. Fazel, R. Ge, S. Kakade, and M. Mesbahi, Global convergence of policy gradient methods for the linear quadratic regulator, in International Conference on Machine Learning, PMLR, 2018, pp. 1467\u20131476."},{"key":"ref16","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.2020.3037046"},{"key":"ref17","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2022.2395"},{"key":"ref18","first-page":"1","author":"Guo X.","year":"2025","journal-title":"IEEE Trans. Automat. Control"},{"key":"ref19","doi-asserted-by":"publisher","DOI":"10.1287\/moor.2021.1238"},{"key":"ref20","unstructured":"T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, in International Conference on Machine Learning, PMLR, 2018, pp. 1861\u20131870."},{"key":"ref21","doi-asserted-by":"publisher","DOI":"10.1137\/20M1382386"},{"key":"ref22","doi-asserted-by":"publisher","DOI":"10.1111\/mafi.12382"},{"key":"ref23","doi-asserted-by":"publisher","DOI":"10.1080\/1350486X.2023.2239850"},{"key":"ref24","doi-asserted-by":"publisher","DOI":"10.1137\/23M1560781"},{"key":"ref25","unstructured":"E. Hazan, S. Kakade, K. Singh, and A. Van Soest, Provably efficient maximum entropy exploration, in International Conference on Machine Learning, PMLR, 2019, pp. 2681\u20132691."},{"key":"ref26","first-page":"12603","volume":"23","author":"Jia Y.","year":"2022","journal-title":"J. Mach. Learn. Res."},{"key":"ref27","first-page":"161","volume":"24","author":"Jia Y.","year":"2023","journal-title":"J. Mach. Learn. Res."},{"key":"ref28","doi-asserted-by":"publisher","DOI":"10.1007\/s40305-024-00546-z"},{"key":"ref29","first-page":"1531","volume":"14","author":"Kakade S. M.","year":"2001","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref30","doi-asserted-by":"publisher","DOI":"10.1017\/S0962492919000060"},{"key":"ref31","unstructured":"D. Malik, A. Pananjady, K. Bhatia, K. Khamaru, P. Bartlett, and M. Wainwright, Derivative-free methods for policy optimization: Guarantees for linear quadratic systems, in 22nd International Conference on Artificial Intelligence and Statistics, PMLR, 2019, pp. 2916\u20132925."},{"key":"ref32","unstructured":"J. Mei, C. Xiao, C. Szepesvari, and D. Schuurmans, On the global convergence rates of softmax policy gradient methods, in International Conference on Machine Learning, PMLR, 2020, pp. 6820\u20136829."},{"key":"ref33","doi-asserted-by":"publisher","DOI":"10.1109\/TAI.2021.3054609"},{"key":"ref34","doi-asserted-by":"crossref","unstructured":"J. Peters, K. Mulling, and Y. Altun, Relative entropy policy search, in Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 24, AAAI, 2010, pp. 1607\u20131612.","DOI":"10.1609\/aaai.v24i1.7727"},{"key":"ref35","volume-title":"Derivative Free Optimization Method","author":"Scheinberg K.","year":"2000"},{"key":"ref36","doi-asserted-by":"crossref","unstructured":"L. Shani, Y. Efroni, and S. Mannor, Adaptive trust region policy optimization: Global convergence and faster rates for regularized MDPs, in Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34, AAAI, 2020, pp. 5668\u20135675.","DOI":"10.1609\/aaai.v34i04.6021"},{"key":"ref37","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton R. S.","year":"2018"},{"key":"ref38","unstructured":"L. Szpruch, T. Treetanthiploet, and Y. Zhang, Exploration-Exploitation Trade-Off for Continuous-Time Episodic Reinforcement Learning with Linear-Convex Models, preprint, arXiv:2112.10264, 2021."},{"key":"ref39","doi-asserted-by":"publisher","DOI":"10.1137\/22M1515744"},{"key":"ref40","doi-asserted-by":"publisher","DOI":"10.1137\/21M1448185"},{"key":"ref41","doi-asserted-by":"crossref","unstructured":"A. Tsiamis, D. S. Kalogerias, L. F. Chamon, A. Ribeiro, and G. J. Pappas, Risk-constrained linear-quadratic regulators, in 59th IEEE Conference on Decision and Control (CDC), IEEE, 2020, pp. 3040\u20133047.","DOI":"10.1109\/CDC42340.2020.9303967"},{"key":"ref42","doi-asserted-by":"crossref","unstructured":"H. Wang, T. Zariphopoulou, and X. Zhou, Exploration Versus Exploitation in Reinforcement Learning: A Stochastic Control Approach, preprint, arXiv:1812.01552, 2018.","DOI":"10.2139\/ssrn.3316387"},{"key":"ref43","first-page":"8145","volume":"21","author":"Wang H.","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"ref44","unstructured":"Z. Wang, Y. Gao, S. Wang, M. M. Zavlanos, A. Abate, and K. H. Johansson, Policy evaluation in distributional LQR, in Learning for Dynamics and Control Conference, PMLR, 2023, pp. 1245\u20131256."},{"key":"ref45","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-016-0043-6"},{"key":"ref46","doi-asserted-by":"publisher","DOI":"10.1080\/09540099108946587"},{"key":"ref47","doi-asserted-by":"publisher","DOI":"10.1145\/3477600"},{"key":"ref48","doi-asserted-by":"publisher","DOI":"10.1137\/20M1347942"},{"key":"ref49","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.2023.3234176"},{"key":"ref50","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3292075"}],"container-title":["SIAM Journal on Control and Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/epubs.siam.org\/doi\/pdf\/10.1137\/23M1621071","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T13:46:27Z","timestamp":1787319987000},"score":1,"resource":{"primary":{"URL":"https:\/\/epubs.siam.org\/doi\/10.1137\/23M1621071"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,9]]},"references-count":50,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,2,28]]}},"alternative-id":["10.1137\/23M1621071"],"URL":"https:\/\/doi.org\/10.1137\/23m1621071","relation":{},"ISSN":["0363-0129","1095-7138"],"issn-type":[{"value":"0363-0129","type":"print"},{"value":"1095-7138","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,9]]}}}