{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T14:39:43Z","timestamp":1787323183875,"version":"build-2736575974"},"reference-count":48,"publisher":"Society for Industrial & Applied Mathematics (SIAM)","issue":"5","funder":[{"DOI":"10.13039\/501100000266","name":"Engineering and Physical Sciences Research Council","doi-asserted-by":"publisher","award":["EP\/L015803\/1"],"award-info":[{"award-number":["EP\/L015803\/1"]}],"id":[{"id":"10.13039\/501100000266","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["SIAM J. Control Optim."],"published-print":{"date-parts":[[2021,1]]},"abstract":"<jats:p>We explore reinforcement learning methods for finding the optimal policy in the linear quadratic regulator (LQR) problem. In particular we consider the convergence of policy gradient methods in the setting of known and unknown parameters. We are able to produce a global linear convergence guarantee for this approach in the setting of finite time horizon and stochastic state dynamics under weak assumptions. The convergence of a projected policy gradient method is also established in order to handle problems with constraints. We illustrate the performance of the algorithm with two examples. The first example is the optimal liquidation of a holding in an asset. We show results for the case where we assume a model for the underlying dynamics and where we apply the method to the data directly. The empirical evidence suggests that the policy gradient method can learn the global optimal solution for a larger class of stochastic systems containing the LQR framework, and that it is more robust with respect to model misspecification when compared to a model-based approach. The second example is an LQR system in a higher dimensional setting with synthetic data.<\/jats:p>","DOI":"10.1137\/20m1382386","type":"journal-article","created":{"date-parts":[[2021,9,28]],"date-time":"2021-09-28T11:58:19Z","timestamp":1632830299000},"page":"3359-3391","source":"Crossref","is-referenced-by-count":57,"title":["Policy Gradient Methods for the Noisy Linear Quadratic Regulator over a Finite Horizon"],"prefix":"10.1137","volume":"59","author":[{"given":"Ben","family":"Hambly","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4293-3450","authenticated-orcid":true,"given":"Renyuan","family":"Xu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huining","family":"Yang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"351","published-online":{"date-parts":[[2021,9,28]]},"reference":[{"key":"atypb1","unstructured":"Y. Abbasi-Yadkori and C. Szepesv\u00e1ri,\n                      Regret bounds for the adaptive control of linear quadratic systems\n                      , in Proceedings of the 24th Annual Conference on Learning Theory, 2011, pp. 1-26."},{"key":"atypb2","unstructured":"M. Abeille and A. Lazaric,\n                      Thompson sampling for linear-quadratic control problems\n                      , in AISTATS 2017 - 20th International Conference on Artificial Intelligence and Statistics, 2017, pp. 1246-1254."},{"key":"atypb3","doi-asserted-by":"crossref","unstructured":"M. Abeille, E. S\u00e9ri\u00e9, A. Lazaric, and X. Brokmann,\n                      LQG for Portfolio Optimization\n                      , 2016; available at SSRN 2863925,https:\/\/ssrn.com\/abstract=2863925.","DOI":"10.2139\/ssrn.2863925"},{"key":"atypb4","doi-asserted-by":"publisher","DOI":"10.1214\/ECP.v20-3829"},{"key":"atypb5","doi-asserted-by":"publisher","DOI":"10.1080\/14697680802595700"},{"key":"atypb6","doi-asserted-by":"publisher","DOI":"10.1080\/135048602100056"},{"key":"atypb7","doi-asserted-by":"publisher","DOI":"10.21314\/JOR.2001.041"},{"key":"atypb8","first-page":"58","volume":"18","author":"Almgren R.","year":"2005","journal-title":"Risk"},{"key":"atypb9","unstructured":"B. D. O. Anderson and J. B. Moore,\n                      Optimal Control: Linear Quadratic Methods\n                      , Courier Corporation, 2007."},{"key":"atypb10","unstructured":"K. J. \\AAstr\u00f6m and B. Wittenmark,\n                      Adaptive Control\n                      , Courier Corporation, 2013."},{"key":"atypb11","unstructured":"W. Bao and X.y. Liu,\n                      Multi-Agent Deep Reinforcement Learning for Liquidation Strategy Analysis\n                      , preprint,https:\/\/arxiv.org\/abs\/1906.11046, 2019."},{"key":"atypb12","unstructured":"D. Bertsekas,\n                      Dynamic Programming and Optimal Control\n                      , Vol. 1, 3rd ed., Athena Scientific, 2005,http:\/\/gen.lib.rus.ec\/book\/index.php?md5=f28152a94f3313576017f55b6bb9ffe8."},{"key":"atypb13","unstructured":"J. Bhandari and D. Russo,\n                      Global Optimality Guarantees for Policy Gradient Methods\n                      , preprint,https:\/\/arxiv.org\/abs\/1906.01786, 2019."},{"key":"atypb14","unstructured":"J. Bu, A. Mesbahi, M. Fazel, and M. Mesbahi,\n                      LQR through the Lens of First Order Methods: Discrete-time Case\n                      , preprint,https:\/\/arxiv.org\/abs\/1907.08921, 2019."},{"key":"atypb15","unstructured":"J. Bu, A. Mesbahi, and M. Mesbahi,\n                      Policy Gradient-based Algorithms for Continuous-time Linear Quadratic Control\n                      , preprint,https:\/\/arxiv.org\/abs\/2006.09178, 2020."},{"key":"atypb16","unstructured":"J. Bu, L. J. Ratliff, and M. Mesbahi,\n                      Global Convergence of Policy Gradient for Sequential Zero-Sum Linear Quadratic Dynamic Games\n                      , preprint,https:\/\/arxiv.org\/abs\/1911.04672, 2019."},{"key":"atypb17","unstructured":"R. Carmona, M. Lauri\u00e8re, and Z. Tan,\n                      Linear-Quadratic Mean-Field Reinforcement Learning: Convergence of Policy Gradient Methods\n                      , preprint,https:\/\/arxiv.org\/abs\/1910.04295, 2019."},{"key":"atypb18","unstructured":"A. Charpentier, R. Elie, and C. Remlinger,\n                      Reinforcement Learning in Economics and Finance\n                      , preprint,https:\/\/arxiv.org\/abs\/2003.10014, 2020."},{"key":"atypb19","doi-asserted-by":"publisher","DOI":"10.1007\/s10208-019-09426-y"},{"key":"atypb20","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.2020.2998952"},{"key":"atypb21","doi-asserted-by":"publisher","DOI":"10.1137\/19M1291108"},{"key":"atypb22","unstructured":"M. Fazel, R. Ge, S. M. Kakade, and M. Mesbahi,\n                      Global convergence of policy gradient methods for the linear quadratic regulator\n                      , in Proceedings of the 35th International Conference on Machine Learning, 2018, pp. 1467-1476."},{"key":"atypb23","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1145\/267460.267481","author":"Fiechter C.-N.","year":"1997","journal-title":"Proceedings of the Tenth Annual Conference on Computational Learning Theory"},{"key":"atypb24","first-page":"385","author":"Flaxman A. D.","year":"2005","journal-title":"Philadelphia"},{"key":"atypb25","doi-asserted-by":"publisher","DOI":"10.1142\/S0219024911006577"},{"key":"atypb26","unstructured":"B. Gravell, P. M. Esfahani, and T. Summers,\n                      Learning Robust Controllers for Linear Quadratic Systems with Multiplicative Noise via Policy Gradient\n                      , preprint,https:\/\/arxiv.org\/abs\/1905.13547, 2019."},{"key":"atypb27","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.2011.2104999"},{"key":"atypb28","doi-asserted-by":"crossref","unstructured":"X. Guo, R. Xu, and T. Zariphopoulou,\n                      Entropy Regularization for Mean Field Games with Learning\n                      , preprint,https:\/\/arxiv.org\/abs\/2010.00145, 2020.","DOI":"10.2139\/ssrn.3702956"},{"key":"atypb29","doi-asserted-by":"crossref","unstructured":"B. Hambly, R. Xu, and H. Yang,\n                      Policy Gradient Methods for the Noisy Linear Quadratic Regulator over a Finite Horizon\n                      , preprint,https:\/\/arxiv.org\/pdf\/2011.10300.pdf, 2021.","DOI":"10.2139\/ssrn.3734179"},{"key":"atypb30","first-page":"457","author":"Hendricks D.","year":"2014","journal-title":"IEEE"},{"key":"atypb31","first-page":"2636","author":"Ibrahimi M.","year":"2012","journal-title":"Advances in Neural Information Processing Systems"},{"key":"atypb32","unstructured":"Z. Jin, J. M. Schmitt, and Z. Wen,\n                      On the Analysis of Model-free Methods for the Linear Quadratic Regulator\n                      , preprint,https:\/\/arxiv.org\/abs\/2007.03861, 2020."},{"key":"atypb33","unstructured":"L. Leal, M. Lauri\u00e8re, and C.A. Lehalle,\n                      Learning a Functional Control for High-Frequency Finance\n                      , preprint,https:\/\/arxiv.org\/abs\/2006.09611, 2020."},{"key":"atypb34","doi-asserted-by":"crossref","unstructured":"W. Li and E. Todorov,\n                      Iterative linear quadratic regulator design for nonlinear biological movement systems\n                      , in International Conference on Informatics in Control, Automation and Robotics (ICINCO), 2004, pp. 222-229.","DOI":"10.5220\/0001143902220229"},{"key":"atypb35","first-page":"21","volume":"21","author":"Malik D.","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"atypb36","doi-asserted-by":"crossref","unstructured":"Y. Nesterov,\n                      Introductory Lectures on Convex Optimization: A Basic Course\n                      , Appl. Optim. 87, Springer Science & Business Media, 2003.","DOI":"10.1007\/978-1-4419-8853-9"},{"key":"atypb37","doi-asserted-by":"crossref","unstructured":"Y. Nevmyvaka, Y. Feng, and M. Kearns,\n                      Reinforcement learning for optimized trade execution\n                      , in Proceedings of the 23rd International Conference on Machine Learning, 2006, pp. 673-680.","DOI":"10.1145\/1143844.1143929"},{"key":"atypb38","unstructured":"B. Ning, F. H. T. Ling, and S. Jaimungal,\n                      Double Deep Q-Learning for Optimal Execution\n                      , preprint,https:\/\/arxiv.org\/abs\/1812.06600, 2018."},{"key":"atypb39","first-page":"1198","author":"Ouyang Y.","year":"2017","journal-title":"IEEE"},{"key":"atypb40","first-page":"7111","author":"Patrinos P.","year":"2011","journal-title":"IEEE"},{"key":"atypb41","doi-asserted-by":"publisher","DOI":"10.3905\/jpm.1988.409150"},{"key":"atypb42","doi-asserted-by":"publisher","DOI":"10.1146\/annurev-control-053018-023825"},{"key":"atypb43","first-page":"5005","author":"Tu S.","year":"2018","journal-title":"International Conference on Machine Learning"},{"key":"atypb44","first-page":"3036","author":"Tu S.","year":"2019","journal-title":"Conference on Learning Theory"},{"key":"atypb45","first-page":"1044","author":"Wasa Y.","year":"2017","journal-title":"IEEE"},{"key":"atypb46","unstructured":"Z. Yang, Y. Chen, M. Hong, and Z. Wang,\n                      On the Global Convergence of Actor-Critic: A Case for Linear Quadratic Regulator with Ergodic Cost\n                      , preprint,https:\/\/arxiv.org\/abs\/1907.06246, 2019."},{"key":"atypb47","first-page":"11602","author":"Zhang K.","year":"2019","journal-title":"Advances in Neural Information Processing Systems"},{"key":"atypb48","doi-asserted-by":"publisher","DOI":"10.3905\/jfds.2020.1.030"}],"container-title":["SIAM Journal on Control and Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/epubs.siam.org\/doi\/pdf\/10.1137\/20M1382386","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T13:46:38Z","timestamp":1787319998000},"score":1,"resource":{"primary":{"URL":"https:\/\/epubs.siam.org\/doi\/10.1137\/20M1382386"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,1]]},"references-count":48,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2021,1]]}},"alternative-id":["10.1137\/20M1382386"],"URL":"https:\/\/doi.org\/10.1137\/20m1382386","relation":{},"ISSN":["0363-0129","1095-7138"],"issn-type":[{"value":"0363-0129","type":"print"},{"value":"1095-7138","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,1]]}}}