{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T01:47:11Z","timestamp":1783561631452,"version":"3.55.0"},"reference-count":24,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2025,1,21]],"date-time":"2025-01-21T00:00:00Z","timestamp":1737417600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation","award":["62203116"],"award-info":[{"award-number":["62203116"]}]},{"name":"National Natural Science Foundation","award":["62103106"],"award-info":[{"award-number":["62103106"]}]},{"name":"National Natural Science Foundation","award":["2024A1515010222"],"award-info":[{"award-number":["2024A1515010222"]}]},{"name":"National Natural Science Foundation","award":["20231800935882"],"award-info":[{"award-number":["20231800935882"]}]},{"name":"National Natural Science Foundation","award":["20234430-01KCJ-G"],"award-info":[{"award-number":["20234430-01KCJ-G"]}]},{"name":"National Natural Science Foundation","award":["20234371-01KCJ-G"],"award-info":[{"award-number":["20234371-01KCJ-G"]}]},{"name":"GuangDong Basic and Applied Basic Research Foundation","award":["62203116"],"award-info":[{"award-number":["62203116"]}]},{"name":"GuangDong Basic and Applied Basic Research Foundation","award":["62103106"],"award-info":[{"award-number":["62103106"]}]},{"name":"GuangDong Basic and Applied Basic Research Foundation","award":["2024A1515010222"],"award-info":[{"award-number":["2024A1515010222"]}]},{"name":"GuangDong Basic and Applied Basic Research Foundation","award":["20231800935882"],"award-info":[{"award-number":["20231800935882"]}]},{"name":"GuangDong Basic and Applied Basic Research Foundation","award":["20234430-01KCJ-G"],"award-info":[{"award-number":["20234430-01KCJ-G"]}]},{"name":"GuangDong Basic and Applied Basic Research Foundation","award":["20234371-01KCJ-G"],"award-info":[{"award-number":["20234371-01KCJ-G"]}]},{"name":"Dongguan Science and Technology of Social Development Program","award":["62203116"],"award-info":[{"award-number":["62203116"]}]},{"name":"Dongguan Science and Technology of Social Development Program","award":["62103106"],"award-info":[{"award-number":["62103106"]}]},{"name":"Dongguan Science and Technology of Social Development Program","award":["2024A1515010222"],"award-info":[{"award-number":["2024A1515010222"]}]},{"name":"Dongguan Science and Technology of Social Development Program","award":["20231800935882"],"award-info":[{"award-number":["20231800935882"]}]},{"name":"Dongguan Science and Technology of Social Development Program","award":["20234430-01KCJ-G"],"award-info":[{"award-number":["20234430-01KCJ-G"]}]},{"name":"Dongguan Science and Technology of Social Development Program","award":["20234371-01KCJ-G"],"award-info":[{"award-number":["20234371-01KCJ-G"]}]},{"name":"SSL Sci-tech Commissioner Program","award":["62203116"],"award-info":[{"award-number":["62203116"]}]},{"name":"SSL Sci-tech Commissioner Program","award":["62103106"],"award-info":[{"award-number":["62103106"]}]},{"name":"SSL Sci-tech Commissioner Program","award":["2024A1515010222"],"award-info":[{"award-number":["2024A1515010222"]}]},{"name":"SSL Sci-tech Commissioner Program","award":["20231800935882"],"award-info":[{"award-number":["20231800935882"]}]},{"name":"SSL Sci-tech Commissioner Program","award":["20234430-01KCJ-G"],"award-info":[{"award-number":["20234430-01KCJ-G"]}]},{"name":"SSL Sci-tech Commissioner Program","award":["20234371-01KCJ-G"],"award-info":[{"award-number":["20234371-01KCJ-G"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>This paper proposes a deep reinforcement learning (DRL) method that combines random network distillation (RND) and long short-term memory (LSTM) to address the tracking control problem, while leveraging the inherent symmetry in robotic arm movements to eliminate the need for learning or knowing the system\u2019s dynamic model. In general, the complexity and strong coupling of robotic manipulators make trajectory tracking extremely challenging. Firstly, the prediction network and fixed network are jointly trained using the RND method. The difference in output values between the two networks acts as an internal reward for the robotic manipulator environment. This internal reward mechanism encourages the robotic arm agent to actively explore unpredictable and unknown environmental states, thereby consequently boosting the performance and efficiency of the tracking control for the robotic manipulator. Then, the Soft Actor-Critic (SAC) algorithm, the LSTM network, and the attention mechanism are integrated to resolve the instability problem during training and acquire a stable policy. The LSTM model effectively captures the symmetry and temporal changes in joint angles, while the attention mechanism dynamically prioritizes important features, thereby reducing the instability of the robotic manipulator during tracking tasks and enhancing feature extraction efficiency. The simulation outcomes demonstrate that the proposed method effectively performs the robot tracking task, confirming the efficacy and efficiency of the DRL algorithm.<\/jats:p>","DOI":"10.3390\/sym17020149","type":"journal-article","created":{"date-parts":[[2025,1,21]],"date-time":"2025-01-21T04:19:51Z","timestamp":1737433191000},"page":"149","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Trajectory Tracking Control Based on Deep Reinforcement Learning for a Robotic Manipulator with an Input Deadzone"],"prefix":"10.3390","volume":"17","author":[{"given":"Fujie","family":"Wang","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Dongguan University of Technology, Dongguan 523808, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jintao","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Dongguan University of Technology, Dongguan 523808, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8517-1137","authenticated-orcid":false,"given":"Yi","family":"Qin","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Dongguan University of Technology, Dongguan 523808, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9194-2711","authenticated-orcid":false,"given":"Fang","family":"Guo","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Dongguan University of Technology, Dongguan 523808, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ming","family":"Jiang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Dongguan University of Technology, Dongguan 523808, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,1,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Karabegovi\u0107, I., Husak, E., Karabegovi\u0107, E., and Mahmi\u0107, M. (2023). Robotic Technology as the Basis of Implementation of Industry 4.0 in Production Processes in China. International Conference \u201cNew Technologies, Development and Applications\u201d, Springer Nature.","DOI":"10.1007\/978-3-031-31066-9_1"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"13103","DOI":"10.1109\/TIE.2024.3357846","article-title":"PID-Based Event-Triggered MPC for Constrained Nonlinear Cyber-Physical Systems: Theory and Application","volume":"71","author":"He","year":"2024","journal-title":"IEEE Trans. Ind. Electron."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Ulu, B., Sava\u015f, S., Ergin, \u00d6.F., Ulu, B., K\u0131rnap, A., Bing\u00f6l, M.S., and Y\u0131ld\u0131r\u0131m, \u015e. (2024). Tuning the Proportional\u2013Integral\u2013Derivative Control Parameters of Unmanned Aerial Vehicles Using Artificial Neural Networks for Point-to-Point Trajectory Approach. Sensors, 24.","DOI":"10.3390\/s24092752"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"126","DOI":"10.1016\/j.isatra.2022.11.024","article-title":"Neural network-based adaptive second-order sliding mode control for uncertain manipulator systems with input saturation","volume":"136","author":"Hu","year":"2023","journal-title":"ISA Trans."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"7814","DOI":"10.1109\/TSMC.2023.3301662","article-title":"Improved sliding mode control for a robotic manipulator with input deadzone and deferred constraint","volume":"53","author":"Zhang","year":"2023","journal-title":"IEEE Trans. Syst. Man Cybern. Syst."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"3102","DOI":"10.1109\/LRA.2023.3264816","article-title":"Online model predictive control of robot manipulator with structured deep Koopman model","volume":"8","author":"Zhang","year":"2023","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2231","DOI":"10.1109\/TCST.2023.3277224","article-title":"Non-prehensile object transportation via model predictive non-sliding manipulation control","volume":"31","author":"Selvaggio","year":"2023","journal-title":"IEEE Trans. Control Syst. Technol."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"021003","DOI":"10.1115\/1.4064201","article-title":"Neural network-based region tracking control for a flexible-joint robot manipulator","volume":"19","author":"Yu","year":"2024","journal-title":"J. Comput. Nonlinear Dyn."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"315","DOI":"10.1007\/s11370-023-00507-0","article-title":"Control of robot manipulators with uncertain closed architecture using neural networks","volume":"17","author":"Khan","year":"2024","journal-title":"Intell. Serv. Robot."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"3770","DOI":"10.1016\/j.jfranklin.2023.01.029","article-title":"Trajectory tracking double two-loop adaptive neural network control for a Quadrotor","volume":"360","year":"2023","journal-title":"J. Frankl. Inst."},{"key":"ref_11","first-page":"1","article-title":"Study of Q-learning and deep Q-network learning control for a rotary inverted pendulum system","volume":"6","year":"2024","journal-title":"Discov. Appl. Sci."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Liu, J., Zhou, Y., Gao, J., and Yan, W. (2023, January 12\u201314). Visual Servoing Gain Tuning by Sarsa: An Application with a Manipulator. Proceedings of the 2023 3rd International Conference on Robotics and Control Engineering, Nanjing, China.","DOI":"10.1145\/3598151.3598169"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"2810","DOI":"10.1109\/TMECH.2023.3249194","article-title":"Position and Attitude Tracking Control of a Biomimetic Underwater Vehicle via Deep Reinforcement Learning","volume":"28","author":"Ma","year":"2023","journal-title":"IEEE\/ASME Trans. Mechatron."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"109902","DOI":"10.1016\/j.asoc.2022.109902","article-title":"Search and tracking strategy of autonomous surface underwater vehicle in oceanic eddies based on deep reinforcement learning","volume":"132","author":"Song","year":"2023","journal-title":"Appl. Soft Comput."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Yang, C., Yang, J., Wang, X., and Liang, B. (2019, January 6\u20138). Control of space flexible manipulator using soft actor-critic and random network distillation. Proceedings of the 2019 IEEE International Conference on Robotics and Biomimetics (ROBIO), Dali, China.","DOI":"10.1109\/ROBIO49542.2019.8961852"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2312","DOI":"10.1109\/TMECH.2021.3104504","article-title":"Improving synchronization performance of multiple Euler\u2013Lagrange systems using nonsingular terminal sliding mode control with fuzzy logic","volume":"27","author":"Wan","year":"2021","journal-title":"IEEE\/ASME Trans. Mechatron."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"2472","DOI":"10.1109\/TMECH.2020.3039967","article-title":"Fractional-order control for uncertain teleoperated cyber-physical system with actuator fault","volume":"26","author":"Ma","year":"2020","journal-title":"IEEE\/ASME Trans. Mechatron."},{"key":"ref_18","unstructured":"Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018, January 10\u201315). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. Proceedings of the International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"8783","DOI":"10.1109\/TNNLS.2022.3215596","article-title":"Improving exploration in actor\u2013critic with weakly pessimistic value estimation and optimistic policy optimization","volume":"35","author":"Li","year":"2022","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Graves, A., and Graves, A. (2012). Long short-term memory. Supervised Sequence Labelling with Recurrent Neural Networks. Studies in Computational Intelligence, Springer.","DOI":"10.1007\/978-3-642-24797-2"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Bahloul, S.N., and Mahmoudi, Y. (2021, January 27\u201328). Rainbow-RND: A Value-based Algorithm Augmented with Intrinsic Curiosity. Proceedings of the 2021 International Conference on Information Systems and Advanced Technologies (ICISAT), Tebessa, Algeria.","DOI":"10.1109\/ICISAT54145.2021.9678409"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"105589","DOI":"10.1016\/j.engappai.2022.105589","article-title":"A general motion controller based on deep reinforcement learning for an autonomous underwater vehicle with unknown disturbances","volume":"117","author":"Huang","year":"2023","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1343","DOI":"10.23919\/JSEE.2023.000113","article-title":"LSTM-DPPO based deep reinforcement learning controller for path following optimization of unmanned surface vehicle","volume":"34","author":"Jiawei","year":"2023","journal-title":"J. Syst. Eng. Electron."},{"key":"ref_24","unstructured":"Nikulin, A., Kurenkov, V., Tarasov, D., and Kolesnikov, S. (2023, January 23\u201329). Anti-exploration by random network distillation. Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/2\/149\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,8]],"date-time":"2025-10-08T10:32:39Z","timestamp":1759919559000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/2\/149"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,21]]},"references-count":24,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,2]]}},"alternative-id":["sym17020149"],"URL":"https:\/\/doi.org\/10.3390\/sym17020149","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,21]]}}}