{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,27]],"date-time":"2026-07-27T12:27:04Z","timestamp":1785155224372,"version":"3.55.0"},"reference-count":28,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2018,5,5]],"date-time":"2018-05-05T00:00:00Z","timestamp":1525478400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>We developed a novel control strategy of speed servo systems based on deep reinforcement learning. The control parameters of speed servo systems are difficult to regulate for practical applications, and problems of moment disturbance and inertia mutation occur during the operation process. A class of reinforcement learning agents for speed servo systems is designed based on the deep deterministic policy gradient algorithm. The agents are trained by a significant number of system data. After learning completion, they can automatically adjust the control parameters of servo systems and compensate for current online. Consequently, a servo system can always maintain good control performance. Numerous experiments are conducted to verify the proposed control strategy. Results show that the proposed method can achieve proportional\u2013integral\u2013derivative automatic tuning and effectively overcome the effects of inertia mutation and torque disturbance.<\/jats:p>","DOI":"10.3390\/a11050065","type":"journal-article","created":{"date-parts":[[2018,5,7]],"date-time":"2018-05-07T03:12:21Z","timestamp":1525662741000},"page":"65","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":55,"title":["Control Strategy of Speed Servo Systems Based on Deep Reinforcement Learning"],"prefix":"10.3390","volume":"11","author":[{"given":"Pengzhan","family":"Chen","sequence":"first","affiliation":[{"name":"School of Electrical Engineering and Automation, East China Jiaotong University, Nanchang 330013, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9813-7616","authenticated-orcid":false,"given":"Zhiqiang","family":"He","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering and Automation, East China Jiaotong University, Nanchang 330013, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chuanxi","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering and Automation, East China Jiaotong University, Nanchang 330013, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiahong","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering and Automation, East China Jiaotong University, Nanchang 330013, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2018,5,5]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"539","DOI":"10.1016\/j.isatra.2013.03.002","article-title":"Stable adaptive PI control for permanent magnet synchronous motor drive based on improved JITL technique","volume":"52","author":"Zheng","year":"2013","journal-title":"ISA Trans."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1016\/j.conengprac.2016.01.002","article-title":"Fractional order PI control applied to level control in coupled two tank MIMO system with experimental validation","volume":"48","author":"Roy","year":"2016","journal-title":"Control Eng. Pract."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"559","DOI":"10.1109\/TCST.2005.847331","article-title":"PID control system analysis, design, and technology","volume":"13","author":"Ang","year":"2005","journal-title":"IEEE Trans. Control Syst. Technol."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"530","DOI":"10.5370\/JEET.2013.8.3.530","article-title":"Sensorless Fuzzy Direct Torque Control for High Performance Electric Vehicle with Four In-Wheel Motors","volume":"8","author":"Sekour","year":"2013","journal-title":"J. Electr. Eng. Technol."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1358","DOI":"10.1109\/TPEL.2012.2206610","article-title":"Nonlinear Speed Control for PMSM System Using Sliding-Mode Control and Disturbance Compensation Techniques","volume":"28","author":"Zhang","year":"2012","journal-title":"IEEE Trans. Power Electron."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1439","DOI":"10.5370\/JEET.2013.8.6.1439","article-title":"Neuro-Fuzzy Control of Interior Permanent Magnet Synchronous Motors","volume":"8","author":"Dang","year":"2013","journal-title":"J. Electr. Eng. Technol."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"509","DOI":"10.1016\/j.asoc.2014.02.027","article-title":"Adaptive hybrid control system using a recurrent RBFN-based self-evolving fuzzy-neural-network for PMSM servo drives","volume":"21","year":"2014","journal-title":"Appl. Soft Comput."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"3549","DOI":"10.1109\/TPEL.2012.2222675","article-title":"PMSM sliding mode FPGA-based control for torque ripple reduction","volume":"28","author":"Jezernik","year":"2012","journal-title":"IEEE Trans. Power Electron."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"900","DOI":"10.1109\/TPEL.2014.2311462","article-title":"Adaptive PID Speed Control Design for Permanent Magnet Synchronous Motor Drives","volume":"30","author":"Jung","year":"2014","journal-title":"IEEE Trans. Power Electron."},{"key":"ref_10","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2018, May 03). Playing Atari with Deep Reinforcement Learning. Available online: https:\/\/arxiv.org\/pdf\/1312.5602v1.pdf."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1038\/nature16961","article-title":"Mastering the game of Go with deep neural networks and tree search","volume":"529","author":"Silver","year":"2016","journal-title":"Nature"},{"key":"ref_13","first-page":"A187","article-title":"Continuous control with deep reinforcement learning","volume":"8","author":"Lillicrap","year":"2015","journal-title":"Comput. Sci."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"176","DOI":"10.1016\/j.neucom.2016.02.029","article-title":"Optimal tracking control for completely unknown nonlinear discrete-time Markov jump systems using data-based reinforcement learning method","volume":"194","author":"Jiang","year":"2016","journal-title":"Neurocomputing"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"183","DOI":"10.1016\/j.eswa.2017.03.002","article-title":"Incremental Q-learning strategy for adaptive PID control of mobile robots","volume":"80","author":"Carlucho","year":"2017","journal-title":"Expert Syst. Appl."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Yu, R., Shi, Z., Huang, C., Li, T., and Ma, Q. (2017, January 26\u201328). Deep reinforcement learning based optimal trajectory tracking control of autonomous underwater vehicle. Proceedings of the 2017 36th Chinese Control Conference (CCC), Dalian, China.","DOI":"10.23919\/ChiCC.2017.8028138"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zhao, H., Wang, Y., Zhao, M., Sun, C., and Tan, Q. (2017). Application of Gradient Descent Continuous Actor-Critic Algorithm for Bilateral Spot Electricity Market Modeling Considering Renewable Power Penetration. Algorithms, 10.","DOI":"10.3390\/a10020053"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Hu, Y., Li, W., Xu, K., Zahid, T., Qin, F., and Li, C. (2018). Energy Management Strategy for a Hybrid Electric Vehicle Based on Deep Reinforcement Learning. Appl. Sci., 8.","DOI":"10.3390\/app8020187"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1219","DOI":"10.1007\/s11071-015-2064-7","article-title":"Composite recurrent Laguerre orthogonal polynomials neural network dynamic control for continuously variable transmission system using altered particle swarm optimization","volume":"81","author":"Lin","year":"2015","journal-title":"Nonlinear Dyn."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"497","DOI":"10.1109\/TCYB.2014.2329495","article-title":"Adaptive NN tracking control of uncertain nonlinear discrete-time systems with nonaffine dead-zone input","volume":"45","author":"Liu","year":"2017","journal-title":"IEEE Trans. Cybern."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1146","DOI":"10.1016\/j.procs.2017.05.431","article-title":"An Adaptive Implementation of \u03b5-Greedy in Reinforcement Learning","volume":"109","year":"2017","journal-title":"Procedia Comput. Sci."},{"key":"ref_22","unstructured":"Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R.Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M. (2017). Parameter Space Noise for Exploration. Comput. Sci., Available online: http:\/\/xueshu.baidu.com\/s?wd=paperuri%3A%28c8411e5bc5e651d776e9d4997604cc3e%29&filter=sc_long_sign&tn=SE_xueshusource_2kduw22v&sc_vurl=http%3A%2F%2Farxiv.org%2Fabs%2F1706.01905&ie=utf-8&sc_us=8746754823928227025."},{"key":"ref_23","unstructured":"Nigam, K. (1999, January 1). Using maximum entropy for text classification. Proceedings of the IJCAI-99 Workshop on Machine Learning for Information Filtering, Stockholm, Sweden. Available online: http:\/\/www.kamalnigam.com\/papers\/maxent-ijcaiws99.pdf."},{"key":"ref_24","unstructured":"Ng, A.Y. (2004, January 4\u20138). Feature selection, L1 vs. L2 regularization, and rotational invariance. Proceedings of the Twenty-First International Conference on Machine Learning, Banff, AB, Canada. Available online: http:\/\/www.yaroslavvb.com\/papers\/ng-feature.pdf."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1287\/mnsc.28.1.1","article-title":"A Survey of Partially Observable Markov Decision Processes: Theory, Models, and Algorithms","volume":"28","author":"Monahan","year":"1982","journal-title":"Manag. Sci."},{"key":"ref_26","unstructured":"Sutton, R.S., and Barto, A.G. (1998). Reinforcement Learning: An Introduction, MIT Press. Available online: http:\/\/www.umiacs.umd.edu\/~hal\/courses\/2016F_RL\/RL9.pdf."},{"key":"ref_27","unstructured":"Silver, D., Lever, G., Heess, N., Thomas, D., Wierstra, D., and Riedmiller, M. (2014, January 21\u201326). Deterministic policy gradient algorithms. Proceedings of the International Conference on International Conference on Machine Learning, Beijing, China. Available online: http:\/\/www0.cs.ucl.ac.uk\/staff\/d.silver\/web\/Applications_files\/deterministic-policy-gradients.pdf."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1143","DOI":"10.1137\/S0363012901385691","article-title":"Actor-critic algorithms","volume":"42","author":"Konda","year":"2002","journal-title":"SIAM J. Control Optim."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/11\/5\/65\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T15:03:25Z","timestamp":1760195005000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/11\/5\/65"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,5,5]]},"references-count":28,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2018,5]]}},"alternative-id":["a11050065"],"URL":"https:\/\/doi.org\/10.3390\/a11050065","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,5,5]]}}}