{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,21]],"date-time":"2026-01-21T20:34:10Z","timestamp":1769027650037,"version":"3.49.0"},"reference-count":26,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2026,1,18]],"date-time":"2026-01-18T00:00:00Z","timestamp":1768694400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>In the process of steel rolling production, the speed reduction compensation of the rolling mill is a key link to ensure the stability of slab rolling and product quality. This paper proposes a hybrid compensation method that integrates motor dynamic modeling with reinforcement learning to minimize mass flow error between adjacent rolling mills during slab rolling. A two-stage compensation strategy is designed, consisting of a constant-gain compensation phase followed by a decaying compensation phase, which explicitly accounts for the repetitive and consistent rolling conditions in batch slab production. Based on a motor dynamics-based theoretical model, an initial estimation of compensation parameters is first obtained, providing a physically interpretable starting point for optimization. Subsequently, a Deep Deterministic Policy Gradient (DDPG) algorithm is employed to iteratively refine the compensation parameters by learning from the mass flow error of each rolled slab, enabling data-driven adaptation while preserving physical consistency. Simulation results demonstrate that the proposed hybrid approach significantly reduces the mass flow error and achieves stable convergence, outperforming strategies with randomly initialized parameters. The results verify the effectiveness and novelty of the proposed method in combining model-based insight with reinforcement learning for intelligent and adaptive rolling mill speed drop compensation.<\/jats:p>","DOI":"10.3390\/a19010084","type":"journal-article","created":{"date-parts":[[2026,1,20]],"date-time":"2026-01-20T09:24:34Z","timestamp":1768901074000},"page":"84","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["An Iterative Reinforcement Learning Algorithm for Speed Drop Compensation in Rolling Mills"],"prefix":"10.3390","volume":"19","author":[{"given":"Shengyue","family":"Zong","sequence":"first","affiliation":[{"name":"National Engineering Research Center for Advanced Rolling Technology and Intelligent Manufacturing, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiwei","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yanpeng","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jinyan","family":"Li","sequence":"additional","affiliation":[{"name":"Nanjing Steel Group Co., Ltd., Nanjing 210035, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,1,18]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Jia, X., Wang, S., Yan, X., Wang, L., and Wang, H. (2023). Research on dynamic response of cold rolling mill with dynamic stiffness compensation. Electronics, 12.","DOI":"10.3390\/electronics12030599"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"103197","DOI":"10.1016\/j.jprocont.2024.103197","article-title":"Dynamic compensation of the threading speed drop in rolling processes","volume":"137","author":"Reinhard","year":"2024","journal-title":"J. Process Control"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"103579","DOI":"10.1016\/j.jprocont.2025.103579","article-title":"Dynamic compensation of the threading speed drop in rolling processes: Bayesian optimization of the roughing and finishing mill","volume":"156","author":"Reinhard","year":"2025","journal-title":"J. Process Control"},{"key":"ref_4","first-page":"2409","article-title":"The Dynamic Compensation Controller Implementation and Optimization in Double Loop DC System","volume":"37","author":"Yang","year":"2017","journal-title":"Proc. CSEE"},{"key":"ref_5","first-page":"1378","article-title":"Study on Roll Eccentricity Compensation Method Based on Swarm Intelligence Identification","volume":"27","author":"Li","year":"2020","journal-title":"Control Eng. China"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1117","DOI":"10.1108\/EC-08-2019-0370","article-title":"A novel forecast model based on CF-PSO-SVM approach for predicting the roll gap in acceleration and deceleration process","volume":"38","author":"Hu","year":"2021","journal-title":"Eng. Comput."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1964","DOI":"10.1007\/s11015-024-01694-6","article-title":"Development of a regulatory method for reducing the impact loads of a rolling mill based on a neural network","volume":"67","author":"Gartlib","year":"2024","journal-title":"Metallurgist"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"128407","DOI":"10.1016\/j.amc.2023.128407","article-title":"A reinforcement learning integral sliding mode control scheme against lumped disturbances in hot strip rolling","volume":"465","author":"Ding","year":"2024","journal-title":"Appl. Math. Comput."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"2495","DOI":"10.1109\/TMECH.2023.3248861","article-title":"Reinforcement-learning-based composite optimal control for looper hydraulic servo systems in hot strip rolling","volume":"28","author":"Wang","year":"2023","journal-title":"IEEE-ASME Trans. Mechatron."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"108695","DOI":"10.1016\/j.engappai.2024.108695","article-title":"A novel deep ensemble reinforcement learning based control method for strip flatness in cold rolling steel industry","volume":"134","author":"Peng","year":"2024","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"980","DOI":"10.1515\/auto-2023-0232","article-title":"Towards microstructure control in forging and rolling: Combining AI with process models for closed-loop property control","volume":"72","author":"Reinisch","year":"2024","journal-title":"Automatisierungstechnik"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"69","DOI":"10.1016\/j.optlaseng.2013.08.007","article-title":"Real-time optical detection system for monitoring roller condition with automatic error compensation","volume":"54","author":"Guo","year":"2014","journal-title":"Opt. Lasers Eng."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"27","DOI":"10.4236\/eng.2014.61005","article-title":"Strip thickness control of cold rolling mill with roll eccentricity compensation by using fuzzy neural network","volume":"6","author":"Hameed","year":"2014","journal-title":"Engineering"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"200","DOI":"10.1109\/TMECH.2016.2636337","article-title":"A plug-and-play monitoring and control architecture for disturbance compensation in rolling mills","volume":"23","author":"Luo","year":"2016","journal-title":"IEEE-ASME Trans. Mechatron."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Hou, Y., Liu, L., Wei, Q., Xu, X., and Chen, C. (2017, January 5\u20138). A novel DDPG method with prioritized experience replay. Proceedings of the IEEE International Conference on Systems, Man, and Cybernetics (SMC), Banff, AB, Canada.","DOI":"10.1109\/SMC.2017.8122622"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"18797","DOI":"10.1109\/ACCESS.2020.2968595","article-title":"Deep deterministic policy gradient (DDPG)-based resource allocation scheme for NOMA vehicular communications","volume":"8","author":"Xu","year":"2020","journal-title":"IEEE Access"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"5118","DOI":"10.1109\/TAES.2022.3216579","article-title":"Autonomous Driving for Natural Paths Using an Improved Deep Reinforcement Learning Algorithm","volume":"58","author":"Tseng","year":"2022","journal-title":"IEEE Trans. Aerosp. Electron. Syst."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"6361","DOI":"10.1109\/TCOMM.2021.3089476","article-title":"Multi-objective optimization for UAV-assisted wireless powered IoT networks based on extended DDPG algorithm","volume":"69","author":"Yu","year":"2021","journal-title":"IEEE Trans. Commun."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"3033","DOI":"10.1109\/ACCESS.2020.3047845","article-title":"Communication Emitter Motion Behavior\u2019s Cognition Based on Deep Reinforcement Learning","volume":"9","author":"Ji","year":"2021","journal-title":"IEEE Access"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"9377","DOI":"10.1109\/TII.2025.3598451","article-title":"Multi-Time Scale Consensus Algorithm of Multi-Agent Systems with Binary-valued Data under Tampering Attacks","volume":"21","author":"Jia","year":"2025","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"18875","DOI":"10.1109\/TASE.2025.3591858","article-title":"Optimal Consensus Control Strategy for Multi-Agent Systems Under Cyber Attacks via a Stackelberg Game Approach","volume":"22","author":"Yu","year":"2025","journal-title":"IEEE Trans. Autom. Sci. Eng."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"3825","DOI":"10.1109\/TAC.2020.3029325","article-title":"System identification with binary-valued observations under data tampering attacks","volume":"66","author":"Guo","year":"2021","journal-title":"IEEE Trans. Autom. Control"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"112001","DOI":"10.1016\/j.automatica.2024.112001","article-title":"Identification of FIR Systems with binary-valued observations under replay attacks","volume":"172","author":"Guo","year":"2025","journal-title":"Automatica"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Liu, W., Wang, Y., and Guo, J. (2025). Consistent Identification for FIR Systems with Multi-Level Quantized Observations Subjected to Data Tampering Attacks: A Joint Estimation Approach. IEEE Trans. Ind. Electron., 1\u201312.","DOI":"10.1109\/TIE.2025.3642258"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"9895","DOI":"10.1109\/JIOT.2020.2988033","article-title":"Multiagent DDPG-based deep learning for smart ocean federated learning IoT networks","volume":"7","author":"Kwon","year":"2020","journal-title":"IEEE Internet Things J."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"10696","DOI":"10.1109\/TVT.2023.3253905","article-title":"Intelligent resource allocation for edge-cloud collaborative networks: A hybrid DDPG-D3QN approach","volume":"72","author":"Hu","year":"2023","journal-title":"IEEE Trans. Veh. Technol."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/1\/84\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,21]],"date-time":"2026-01-21T05:14:26Z","timestamp":1768972466000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/1\/84"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,18]]},"references-count":26,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1]]}},"alternative-id":["a19010084"],"URL":"https:\/\/doi.org\/10.3390\/a19010084","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,18]]}}}