{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,5,31]],"date-time":"2025-05-31T04:11:23Z","timestamp":1748664683432,"version":"3.41.0"},"reference-count":42,"publisher":"Walter de Gruyter GmbH","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,6,26]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>This article investigates the core mechanisms of indirect data-driven control for unknown systems, focusing on the application of policy iteration (PI) within the context of the linear quadratic regulator (LQR) optimal control problem. Specifically, we consider a setting where data is collected sequentially from a linear system subject to exogenous process noise, and is then used to refine estimates of the optimal control policy. We integrate recursive least squares (RLS) for online model estimation within a certainty-equivalent framework, and employ PI to iteratively update the control policy. In this work, we investigate first the convergence behavior of RLS under two different models of adversarial noise, namely point-wise and energy bounded noise, and then we provide a closed-loop analysis of the combined model identification and control design process. This iterative scheme is formulated as an algorithmic dynamical system consisting of the feedback interconnection between two algorithms expressed as discrete-time systems. This system-theoretic viewpoint on indirect data-driven control allows us to establish convergence guarantees to the optimal controller in the face of uncertainty caused by noisy data. Simulations illustrate the theoretical results.<\/jats:p>","DOI":"10.1515\/auto-2024-0164","type":"journal-article","created":{"date-parts":[[2025,5,30]],"date-time":"2025-05-30T20:52:17Z","timestamp":1748638337000},"page":"398-412","source":"Crossref","is-referenced-by-count":0,"title":["Robustness of online identification-based policy iteration to noisy data"],"prefix":"10.1515","volume":"73","author":[{"given":"Bowen","family":"Song","sequence":"first","affiliation":[{"name":"Institute for Systems Theory and Automatic Control, University of Stuttgart , Stuttgart , Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrea","family":"Iannelli","sequence":"additional","affiliation":[{"name":"Institute for Systems Theory and Automatic Control, University of Stuttgart , Stuttgart , Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"374","published-online":{"date-parts":[[2025,5,28]]},"reference":[{"key":"2025053020521305691_j_auto-2024-0164_ref_001","doi-asserted-by":"crossref","unstructured":"Z.-S. Hou and Z. Wang, \u201cFrom model-based control to data-driven control: survey, classification and perspective,\u201d Inf. Sci., vol.\u00a0235, pp.\u00a03\u201335, 2013. https:\/\/doi.org\/10.1016\/j.ins.2012.07.014.","DOI":"10.1016\/j.ins.2012.07.014"},{"key":"2025053020521305691_j_auto-2024-0164_ref_002","doi-asserted-by":"crossref","unstructured":"A. Khaki-Sedigh, An Introduction to Data-Driven Control Systems, New Jersey, U.S., Wiley, 2023.","DOI":"10.1002\/9781394196432"},{"key":"2025053020521305691_j_auto-2024-0164_ref_003","doi-asserted-by":"crossref","unstructured":"D. Soudbakhsh, et al.., \u201cData-driven control: theory and applications,\u201d in 2023 American Control Conference (ACC), 2023.","DOI":"10.23919\/ACC55779.2023.10156081"},{"key":"2025053020521305691_j_auto-2024-0164_ref_004","doi-asserted-by":"crossref","unstructured":"F. D\u00f6rfler, \u201cData-driven control: part two of two: hot take: why not go with models?\u201d IEEE Contr. Syst. Mag., vol.\u00a043, no.\u00a06, pp.\u00a027\u201331, 2023. https:\/\/doi.org\/10.1109\/mcs.2023.3310302.","DOI":"10.1109\/MCS.2023.3310302"},{"key":"2025053020521305691_j_auto-2024-0164_ref_005","doi-asserted-by":"crossref","unstructured":"T. Faulwasser, R. Ou, G. Pan, P. Schmitz, and K. Worthmann, \u201cBehavioral theory for stochastic systems? A data-driven journey from willems to wiener and back again,\u201d Annu. Rev. Control, vol.\u00a055, pp.\u00a092\u2013117, 2023. https:\/\/doi.org\/10.1016\/j.arcontrol.2023.03.005.","DOI":"10.1016\/j.arcontrol.2023.03.005"},{"key":"2025053020521305691_j_auto-2024-0164_ref_006","doi-asserted-by":"crossref","unstructured":"J. Berberich and F. Allg\u00f6wer, \u201cAn overview of systems-theoretic guarantees in data-driven model predictive control,\u201d Annu. Rev. Control Robot. Auton. Syst., vol. 8, 2024.","DOI":"10.1146\/annurev-control-030323-024328"},{"key":"2025053020521305691_j_auto-2024-0164_ref_007","doi-asserted-by":"crossref","unstructured":"S. Dean, H. Mania, N. Matni, B. Recht, and S. Tu, \u201cOn the sample complexity of the linear quadratic regulator,\u201d Found. Math. Comput., vol. 20, no. 4, pp. 10\u2013679, 2017. https:\/\/doi.org\/10.1007\/s10208-019-09426-y.","DOI":"10.1007\/s10208-019-09426-y"},{"key":"2025053020521305691_j_auto-2024-0164_ref_008","doi-asserted-by":"crossref","unstructured":"M. Ferizbegovic, J. Umenberger, H. Hjalmarsson, and T. B. Sch\u00f6n, \u201cLearning robust LQ-controllers using application oriented exploration,\u201d IEEE Control Syst. Lett., vol.\u00a04, no.\u00a01, pp.\u00a019\u201324, 2020. https:\/\/doi.org\/10.1109\/lcsys.2019.2921512.","DOI":"10.1109\/LCSYS.2019.2921512"},{"key":"2025053020521305691_j_auto-2024-0164_ref_009","unstructured":"N. Chatzikiriakos, R. Str\u00e4sser, F. Allg\u00f6wer, and A. Iannelli, \u201cEnd-to-end guarantees for indirect data-driven control of bilinear systems with finite stochastic data,\u201d arXiv preprint arXiv:2409.18010, 2024."},{"key":"2025053020521305691_j_auto-2024-0164_ref_010","unstructured":"A. Tsiamis, I. M. Ziemann, M. Morari, N. Matni, and G. J. Pappas, \u201cLearning to control linear systems can be hard,\u201d in Proceedings of Thirty Fifth Conference on Learning Theory, 2022."},{"key":"2025053020521305691_j_auto-2024-0164_ref_011","unstructured":"L. Ljung, System Identification: Theory for the User. Prentice Hall Information and System Sciences Series, New Jersey, U.S., Prentice Hall PTR, 1999."},{"key":"2025053020521305691_j_auto-2024-0164_ref_012","unstructured":"K. J. \u00c5str\u00f6m and B. Wittenmark, Adaptive Control. Dover Books on Electrical Engineering, New York, U.S., Dover Publications, 2008."},{"key":"2025053020521305691_j_auto-2024-0164_ref_013","doi-asserted-by":"crossref","unstructured":"A. M. Annaswamy, \u201cAdaptive control and intersections with reinforcement learning,\u201d in Annual Review of Control, Robotics, and Autonomous Systems, 6(Volume 6, 2023), 2023, pp.\u00a065\u201393.","DOI":"10.1146\/annurev-control-062922-090153"},{"key":"2025053020521305691_j_auto-2024-0164_ref_014","doi-asserted-by":"crossref","unstructured":"F. L. Lewis, D. Vrabie, and V. L. Syrmos, Optimal Control, New Jersey, U.S., John Wiley & Sons, 2012.","DOI":"10.1002\/9781118122631"},{"key":"2025053020521305691_j_auto-2024-0164_ref_015","unstructured":"D. Bertsekas, Abstract Dynamic Programming, 3rd ed. Nashua, U.S.A., Athena Scientific, 2022."},{"key":"2025053020521305691_j_auto-2024-0164_ref_016","unstructured":"Y. Park, R. A. Rossi, Z. Wen, G. Wu, and H. Zhao, \u201cStructured policy iteration for linear quadratic regulator,\u201d in Proceedings of the 37th International Conference on Machine Learning, 2020."},{"key":"2025053020521305691_j_auto-2024-0164_ref_017","doi-asserted-by":"crossref","unstructured":"D. Lee, \u201cConvergence of dynamic programming on the semidefinite cone for discrete-time infinite-horizon LQR,\u201d IEEE Trans. Autom. Control, vol.\u00a067, no.\u00a010, pp.\u00a05661\u20135668, 2022. https:\/\/doi.org\/10.1109\/tac.2022.3181752.","DOI":"10.1109\/TAC.2022.3181752"},{"key":"2025053020521305691_j_auto-2024-0164_ref_018","doi-asserted-by":"crossref","unstructured":"B. Pang, T. Bian, and Z.-P. Jiang, \u201cRobust policy iteration for continuous-time linear quadratic regulation,\u201d IEEE Trans. Autom. Control, vol.\u00a067, no.\u00a01, pp.\u00a0504\u2013511, 2022. https:\/\/doi.org\/10.1109\/tac.2021.3085510.","DOI":"10.1109\/TAC.2021.3085510"},{"key":"2025053020521305691_j_auto-2024-0164_ref_019","doi-asserted-by":"crossref","unstructured":"D. Bertsekas, \u201cNewton\u2019s method for reinforcement learning and model predictive control,\u201d Results Control Optim., vol.\u00a07, 2022, Art. no. 100121. https:\/\/doi.org\/10.1016\/j.rico.2022.100121.","DOI":"10.1016\/j.rico.2022.100121"},{"key":"2025053020521305691_j_auto-2024-0164_ref_020","unstructured":"B. Song, C. Wu, and A. Iannelli, \u201cConvergence and robustness of value and policy iteration for the linear quadratic regulator,\u201d arXiv preprint arXiv:2411.04548, 2024."},{"key":"2025053020521305691_j_auto-2024-0164_ref_021","doi-asserted-by":"crossref","unstructured":"T. Y. Chun, J. Y. Lee, J. B. Park, and Y.Ho Choi, \u201cStability and monotone convergence of generalised policy iteration for discrete-time linear quadratic regulations,\u201d Int. J. Control, vol.\u00a089, no.\u00a03, pp.\u00a0437\u2013450, 2016. https:\/\/doi.org\/10.1080\/00207179.2015.1079737.","DOI":"10.1080\/00207179.2015.1079737"},{"key":"2025053020521305691_j_auto-2024-0164_ref_022","doi-asserted-by":"crossref","unstructured":"F. A. Yaghmaie, F. Gustafsson, and L. Ljung, \u201cLinear quadratic control using model-free reinforcement learning,\u201d IEEE Trans. Autom. Control, vol.\u00a068, no.\u00a02, pp.\u00a0737\u2013752, 2023. https:\/\/doi.org\/10.1109\/tac.2022.3145632.","DOI":"10.1109\/TAC.2022.3145632"},{"key":"2025053020521305691_j_auto-2024-0164_ref_023","doi-asserted-by":"crossref","unstructured":"L. Sforni, G. Carnevale, I. Notarnicola, and G. Notarstefano, \u201cOn-policy data-driven linear quadratic regulator via combined policy iteration and recursive least squares,\u201d in 2023 62nd IEEE Conference on Decision and Control (CDC), 2023.","DOI":"10.1109\/CDC49753.2023.10383604"},{"key":"2025053020521305691_j_auto-2024-0164_ref_024","doi-asserted-by":"crossref","unstructured":"B. Song and A. Iannelli, \u201cThe role of identification in data-driven policy iteration: a system theoretic study,\u201d Int. J. Robust Nonlinear Control, 2024. https:\/\/doi.org\/10.1002\/rnc.7475.","DOI":"10.1002\/rnc.7475"},{"key":"2025053020521305691_j_auto-2024-0164_ref_025","unstructured":"A. Cassel, A. Cohen, and T. Koren, \u201cLogarithmic regret for learning linear quadratic regulators efficiently,\u201d in Proceedings of the 37th International Conference on Machine Learning, Volume 119 of Proceedings of Machine Learning Research, PMLR, 2020, pp.\u00a01328\u20131337."},{"key":"2025053020521305691_j_auto-2024-0164_ref_026","unstructured":"Y. Abbasi-Yadkori and C. Szepesv\u00e1ri, \u201cRegret bounds for the adaptive control of linear quadratic systems,\u201d in Proceedings of the 24th Annual Conference on Learning Theory, Volume 19 of Proceedings of Machine Learning Research, Budapest, Hungary, PMLR, 2011, pp.\u00a01\u201326."},{"key":"2025053020521305691_j_auto-2024-0164_ref_027","unstructured":"M. Simchowitz and D. Foster, \u201cNaive exploration is optimal for online LQR,\u201d in Proceedings of the 37th International Conference on Machine Learning, Volume 119 of Proceedings of Machine Learning Research, PMLR, 2020, pp.\u00a08937\u20138948."},{"key":"2025053020521305691_j_auto-2024-0164_ref_028","doi-asserted-by":"crossref","unstructured":"M. Borghesi, A. Bosso, and G. Notarstefano, \u201cOn-policy data-driven linear quadratic regulator via model reference adaptive reinforcement learning,\u201d in 2023 62nd IEEE Conference on Decision and Control (CDC), 2023, pp.\u00a032\u201337.","DOI":"10.1109\/CDC49753.2023.10383516"},{"key":"2025053020521305691_j_auto-2024-0164_ref_029","doi-asserted-by":"crossref","unstructured":"F. D\u00f6rfler, Z. He, G. Belgioioso, S. Bolognani, J. Lygeros, and M. Muehlebach, \u201cToward a systems theory of algorithms,\u201d IEEE Control Syst. Lett., vol.\u00a08, pp.\u00a01198\u20131210, 2024. https:\/\/doi.org\/10.1109\/lcsys.2024.3406943.","DOI":"10.1109\/LCSYS.2024.3406943"},{"key":"2025053020521305691_j_auto-2024-0164_ref_030","unstructured":"M. Fazel, R. Ge, S. Kakade, and M. Mesbahi, \u201cGlobal convergence of policy gradient methods for the linear quadratic regulator,\u201d in Proceedings of the 35th International Conference on Machine Learning, Volume 80 of Proceedings of Machine Learning Research, PMLR, 2018, pp.\u00a01467\u20131476."},{"key":"2025053020521305691_j_auto-2024-0164_ref_031","doi-asserted-by":"crossref","unstructured":"S. Takakura and K. Sato, \u201cStructured output feedback control for linear quadratic regulator using policy gradient method,\u201d IEEE Trans. Autom. Control, vol.\u00a069, no.\u00a01, pp.\u00a0363\u2013370, 2024. https:\/\/doi.org\/10.1109\/tac.2023.3264176.","DOI":"10.1109\/TAC.2023.3264176"},{"key":"2025053020521305691_j_auto-2024-0164_ref_032","doi-asserted-by":"crossref","unstructured":"G. Carnevale, N. Mimmo, and G. Notarstefano, \u201cData-driven LQR with finite-time experiments via extremum-seeking policy iteration,\u201d arXiv preprint arXiv:2412.02758, 2024.","DOI":"10.1109\/CDC56724.2024.10885851"},{"key":"2025053020521305691_j_auto-2024-0164_ref_033","doi-asserted-by":"crossref","unstructured":"F. Zhao, F. D\u00f6rfler, A. Chiuso, and K. You, \u201cData-enabled policy optimization for direct adaptive learning of the LQR,\u201d arXiv preprint arXiv:2401.14871, 2024.","DOI":"10.1109\/TAC.2025.3569597"},{"key":"2025053020521305691_j_auto-2024-0164_ref_034","doi-asserted-by":"crossref","unstructured":"A. L. Bruce, A. Goel, and D. S. Bernstein, \u201cConvergence and consistency of recursive least squares with variable-rate forgetting,\u201d Automatica, vol.\u00a0119, 2020, Art. no. 109052. https:\/\/doi.org\/10.1016\/j.automatica.2020.109052.","DOI":"10.1016\/j.automatica.2020.109052"},{"key":"2025053020521305691_j_auto-2024-0164_ref_035","unstructured":"S. Tu and B. Recht, \u201cThe gap between model-based and model-free methods on the linear quadratic regulator: an asymptotic viewpoint,\u201d in Proceedings of the Thirty-Second Conference on Learning Theory, 2019."},{"key":"2025053020521305691_j_auto-2024-0164_ref_036","doi-asserted-by":"crossref","unstructured":"Y. Xie, J. Berberich, and F. Allg\u00f6wer, \u201cData-driven min-max\u202fmpc for linear systems,\u201d in 2024 American Control Conference (ACC), 2024.","DOI":"10.23919\/ACC60939.2024.10644295"},{"key":"2025053020521305691_j_auto-2024-0164_ref_037","doi-asserted-by":"crossref","unstructured":"J. Venkatasubramanian, J. K\u00f6hler, M. Cannon, and F. Allg\u00f6wer, \u201cTowards targeted exploration for non-stochastic disturbances,\u201d in 20th IFAC Symposium on System Identification SYSID, 2024.","DOI":"10.1016\/j.ifacol.2024.08.588"},{"key":"2025053020521305691_j_auto-2024-0164_ref_038","doi-asserted-by":"crossref","unstructured":"G. Hewer, \u201cAn iterative technique for the computation of the steady state gains for the discrete optimal regulator,\u201d IEEE Trans. Autom. Control, vol.\u00a016, no.\u00a04, pp.\u00a0382\u2013384, 1971. https:\/\/doi.org\/10.1109\/tac.1971.1099755.","DOI":"10.1109\/TAC.1971.1099755"},{"key":"2025053020521305691_j_auto-2024-0164_ref_039","unstructured":"K. B. Petersen and M. S. Pedersen, The Matrix Cookbook, Denmark, Technical University of Denmark, 2008, Version 20081110."},{"key":"2025053020521305691_j_auto-2024-0164_ref_040","doi-asserted-by":"crossref","unstructured":"E. D. Sontag, Input to State Stability: Basic Concepts and Results, Berlin, Heidelberg, Springer, 2008, pp.\u00a0163\u2013220.","DOI":"10.1007\/978-3-540-77653-6_3"},{"key":"2025053020521305691_j_auto-2024-0164_ref_041","doi-asserted-by":"crossref","unstructured":"Z.-P. Jiang and Y. Wang, \u201cInput-to-state stability for discrete-time nonlinear systems,\u201d Automatica, vol.\u00a037, no.\u00a06, pp.\u00a0857\u2013869, 2001. https:\/\/doi.org\/10.1016\/s0005-1098(01)00028-0.","DOI":"10.1016\/S0005-1098(01)00028-0"},{"key":"2025053020521305691_j_auto-2024-0164_ref_042","doi-asserted-by":"crossref","unstructured":"A. Iannelli and R. S. Smith, \u201cA multiobjective LQR synthesis approach to dual control for uncertain plants,\u201d IEEE Control Syst. Lett., vol.\u00a04, no.\u00a04, pp.\u00a0952\u2013957, 2020. https:\/\/doi.org\/10.1109\/lcsys.2020.2997085.","DOI":"10.1109\/LCSYS.2020.2997085"}],"container-title":["at - Automatisierungstechnik"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/auto-2024-0164\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/auto-2024-0164\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,30]],"date-time":"2025-05-30T20:53:14Z","timestamp":1748638394000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/auto-2024-0164\/html"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,28]]},"references-count":42,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2025,5,29]]},"published-print":{"date-parts":[[2025,6,26]]}},"alternative-id":["10.1515\/auto-2024-0164"],"URL":"https:\/\/doi.org\/10.1515\/auto-2024-0164","relation":{},"ISSN":["0178-2312","2196-677X"],"issn-type":[{"value":"0178-2312","type":"print"},{"value":"2196-677X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,28]]}}}