{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T14:36:54Z","timestamp":1787323014885,"version":"build-2736575974"},"reference-count":43,"publisher":"Society for Industrial & Applied Mathematics (SIAM)","issue":"4","funder":[{"name":"National Key Research and Development Program of China","award":["2018YFA0703900"],"award-info":[{"award-number":["2018YFA0703900"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["11801084"],"award-info":[{"award-number":["11801084"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["11871121"],"award-info":[{"award-number":["11871121"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["11701369"],"award-info":[{"award-number":["11701369"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["12071292"],"award-info":[{"award-number":["12071292"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100007129","name":"Natural Science Foundation of Shandong Province","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004731","name":"Natural Science Foundation of Zhejiang Province","doi-asserted-by":"publisher","award":["LZ22A010005"],"award-info":[{"award-number":["LZ22A010005"]}],"id":[{"id":"10.13039\/501100004731","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004731","name":"Natural Science Foundation of Zhejiang Province","doi-asserted-by":"publisher","award":["LY21A010001"],"award-info":[{"award-number":["LY21A010001"]}],"id":[{"id":"10.13039\/501100004731","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100013285","name":"Program for Professor of Special Appointment at Shanghai Institutions of Higher Learning at Shanghai Institutions of Higher Learning)","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100013285","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["SIAM J. Control Optim."],"published-print":{"date-parts":[[2022,8]]},"abstract":"<jats:p>This paper studies an infinite horizon optimal control problem for discrete-time linear systems and quadratic criteria, both with random parameters which are independent and identically distributed with respect to time. A classical approach is to solve an algebraic Riccati equation that involves mathematical expectations and requires certain statistical information of the parameters. In this paper, we propose an iterative algorithm in the spirit of Q-learning for the situation where only one random sample of parameters emerges at each time step. The first theorem proves the equivalence of three properties: the convergence of the learning sequence, the well-posedness of the control problem, and the solvability of the algebraic Riccati equation. The second theorem shows that the adaptive feedback control in terms of the learning sequence stabilizes the system as long as the control problem is well-posed. Numerical examples are presented to illustrate our results.<\/jats:p>","DOI":"10.1137\/20m1379605","type":"journal-article","created":{"date-parts":[[2022,7,7]],"date-time":"2022-07-07T15:23:02Z","timestamp":1657207382000},"page":"1991-2015","source":"Crossref","is-referenced-by-count":15,"title":["A Q-Learning Algorithm for Discrete-Time Linear-Quadratic Control with Random Parameters of Unknown Distribution: Convergence and Stabilization"],"prefix":"10.1137","volume":"60","author":[{"given":"Kai","family":"Du","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qingxin","family":"Meng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3968-0870","authenticated-orcid":true,"given":"Fu","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"351","published-online":{"date-parts":[[2022,7,7]]},"reference":[{"key":"atypb1","first-page":"1","volume-title":"Proceedings of the 24th Annual Conference on Learning Theory, JMLR Workshop Conf. Proc. 19","author":"Abbasi-Yadkori Y.","year":"2011"},{"key":"atypb2","doi-asserted-by":"publisher","DOI":"10.1109\/CDC.1973.269246"},{"key":"atypb3","volume-title":"Optimization of Stochastic Systems","author":"Aoki M.","year":"1967"},{"key":"atypb4","volume-title":"Optimal Control and System Theory in Dynamic Economic Analysis","volume":"1","author":"Aoki M.","year":"1976"},{"key":"atypb5","volume-title":"Adaptive Control","author":"Astr\u00f6m K. J.","year":"2013"},{"key":"atypb6","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.1977.1101526"},{"key":"atypb7","doi-asserted-by":"publisher","DOI":"10.1016\/S0005-1098(98)00044-2"},{"key":"atypb8","volume-title":"Reinforcement Learning and Optimal Control","author":"Bertsekas D. P.","year":"2019"},{"key":"atypb9","doi-asserted-by":"publisher","DOI":"10.1007\/978-93-86279-38-5"},{"key":"atypb10","volume-title":"Nonconvex Optim. Appl. 64","author":"Chen H.-F.","year":"2002"},{"key":"atypb11","doi-asserted-by":"publisher","DOI":"10.1080\/00207178608933508"},{"key":"atypb12","doi-asserted-by":"publisher","DOI":"10.1137\/S0363012996310478"},{"key":"atypb13","volume-title":"Analysis and Control of Dynamic Economic Systems","author":"Chow G. C.","year":"1975"},{"key":"atypb14","doi-asserted-by":"publisher","DOI":"10.1016\/0005-1098(82)90072-3"},{"key":"atypb15","first-page":"4192","volume-title":"Proceedings of the 32nd International Conference on Neural Information Processing Systems","author":"Dean S.","year":"2018"},{"key":"atypb16","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.1964.1105691"},{"key":"atypb17","doi-asserted-by":"publisher","DOI":"10.1109\/9.788532"},{"key":"atypb18","doi-asserted-by":"publisher","DOI":"10.1525\/9780520313880-007"},{"key":"atypb19","doi-asserted-by":"publisher","DOI":"10.1080\/00207177808922359"},{"key":"atypb20","doi-asserted-by":"publisher","DOI":"10.1109\/WCICA.2006.1712311"},{"key":"atypb21","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1994.6.6.1185"},{"key":"atypb22","first-page":"287","volume-title":"Proc. Sympos. Appl. Math. 13","author":"Kalman R. E.","year":"1961"},{"key":"atypb23","doi-asserted-by":"publisher","DOI":"10.1016\/0016-0032(59)90093-6"},{"key":"atypb24","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.1977.1101633"},{"key":"atypb25","volume-title":"2nd. ed., Stoch. Model. Appl. Probab. 35","author":"Kushner H. J.","year":"2003"},{"key":"atypb26","first-page":"10154","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems, MIT Press","author":"Mania H.","year":"2019"},{"key":"atypb27","doi-asserted-by":"publisher","DOI":"10.1109\/CDC.1975.270671"},{"key":"atypb28","doi-asserted-by":"publisher","DOI":"10.1080\/07362998308809005"},{"key":"atypb29","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.2015.2509958"},{"key":"atypb30","doi-asserted-by":"publisher","DOI":"10.1109\/9.506238"},{"key":"atypb31","volume-title":"Markov Decision Processes: Discrete Stochastic Dynamic Programming","author":"Puterman M. L.","year":"2014"},{"key":"atypb32","doi-asserted-by":"publisher","DOI":"10.1137\/S0363012900371083"},{"key":"atypb33","doi-asserted-by":"publisher","DOI":"10.1137\/0614055"},{"key":"atypb34","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177729586"},{"key":"atypb35","doi-asserted-by":"publisher","DOI":"10.1016\/j.automatica.2020.108982"},{"key":"atypb36","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton R. S.","year":"2018"},{"key":"atypb37","doi-asserted-by":"publisher","DOI":"10.1080\/00207178408933286"},{"key":"atypb38","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.1973.1100242"},{"key":"atypb39","doi-asserted-by":"publisher","DOI":"10.1007\/BF00993306"},{"key":"atypb40","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2015.06.053"},{"key":"atypb41","unstructured":"C. J. C. H. Watkins,\n                      Learning From Delayed Rewards\n                      , Ph.D. thesis, University of Cambridge, 1989."},{"key":"atypb42","doi-asserted-by":"publisher","DOI":"10.1080\/00207178508933344"},{"key":"atypb43","doi-asserted-by":"publisher","DOI":"10.1109\/9.192203"}],"container-title":["SIAM Journal on Control and Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/epubs.siam.org\/doi\/pdf\/10.1137\/20M1379605","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T13:37:13Z","timestamp":1787319433000},"score":1,"resource":{"primary":{"URL":"https:\/\/epubs.siam.org\/doi\/10.1137\/20M1379605"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,7]]},"references-count":43,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2022,8]]}},"alternative-id":["10.1137\/20M1379605"],"URL":"https:\/\/doi.org\/10.1137\/20m1379605","relation":{},"ISSN":["0363-0129","1095-7138"],"issn-type":[{"value":"0363-0129","type":"print"},{"value":"1095-7138","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,7,7]]}}}