{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T01:17:02Z","timestamp":1760059022210,"version":"build-2065373602"},"reference-count":33,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2025,5,14]],"date-time":"2025-05-14T00:00:00Z","timestamp":1747180800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Guangdong Province Key Construction Discipline Research Ability Enhancement Project","award":["2024ZDJS071","2022WGALH17","2021WGALH18"],"award-info":[{"award-number":["2024ZDJS071","2022WGALH17","2021WGALH18"]}]},{"name":"Wuyi University\u2013Hong Kong\u2013Macau Joint Funding Scheme","award":["2024ZDJS071","2022WGALH17","2021WGALH18"],"award-info":[{"award-number":["2024ZDJS071","2022WGALH17","2021WGALH18"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["JSAN"],"abstract":"<jats:p>In mobile multi-agent systems (MASs), achieving effective leader\u2013follower coordination under unknown dynamics poses significant challenges. This study proposes a two-stage cooperative strategy that integrates Gaussian Processes (GPs) for modeling and a Twin Delayed Deep Deterministic Policy Gradient (TD3) for policy optimization (GPTD3), aiming to enhance adaptability and multi-objective optimization. Initially, GPs are utilized to model the uncertain dynamics of agents based on sensor data, providing a stable and noiseless training virtual environment for the first phase of TD3 strategy network training. Subsequently, a TD3-based compensation learning mechanism is introduced to reduce consensus errors among multiple agents by incorporating the position state of other agents. Additionally, the approach employs an enhanced dual-layer reward mechanism tailored to different stages of learning, ensuring robustness and improved convergence speed. Experimental results using a differential drive robot simulation demonstrate the superiority of this method over traditional controllers. The integration of the TD3 compensation network further improves the cooperative reward among agents.<\/jats:p>","DOI":"10.3390\/jsan14030051","type":"journal-article","created":{"date-parts":[[2025,5,14]],"date-time":"2025-05-14T03:59:29Z","timestamp":1747195169000},"page":"51","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["A Two-Stage Strategy Integrating Gaussian Processes and TD3 for Leader\u2013Follower Coordination in Multi-Agent Systems"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-2374-7381","authenticated-orcid":false,"given":"Xicheng","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Mechanical and Automation Engineering, Wuyi University, Jiangmen 529000, China"},{"name":"School of Mechanical and Electrical Engineering, Guangdong University of Science and Technology, Dongguan 523668, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6157-3457","authenticated-orcid":false,"given":"Bingchun","family":"Jiang","sequence":"additional","affiliation":[{"name":"School of Mechanical and Electrical Engineering, Guangdong University of Science and Technology, Dongguan 523668, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fuqin","family":"Deng","sequence":"additional","affiliation":[{"name":"School of Electronic and Information Engineering, Wuyi University, Jiangmen 529000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1696-3277","authenticated-orcid":false,"given":"Min","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Mechanical and Electrical Engineering, Guangdong University of Science and Technology, Dongguan 523668, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,5,14]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Proskurnikov, A., and Cao, M. (2016). Consensus in Multi-Agent Systems. Wiley Encyclopedia of Electrical and Electronics Engineering, Wiley & Sons.","DOI":"10.1002\/047134608X.W8332"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"4091","DOI":"10.1109\/TIE.2016.2542134","article-title":"Data-Driven Optimal Consensus Control for Discrete-Time Multi-Agent Systems with Unknown Dynamics Using Reinforcement Learning Method","volume":"64","author":"Zhang","year":"2017","journal-title":"IEEE Trans. Ind. Electron."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Xu, Y., Yuan, Y., and Liu, H. (2018, January 18\u201320). Event-driven MPC for leader-follower nonlinear multi-agent systems. Proceedings of the 2018 3rd International Conference on Advanced Robotics and Mechatronics (ICARM), Singapore.","DOI":"10.1109\/ICARM.2018.8610731"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Bai, W., Cao, L., Dong, G., and Li, H. (2019, January 24\u201327). Adaptive Reinforcement Learning Tracking Control for Second-Order Multi-Agent Systems. Proceedings of the 2019 IEEE 8th Data Driven Control and Learning Systems Conference (DDCLS), Dali, China.","DOI":"10.1109\/DDCLS.2019.8908978"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Tan, X. (2022, January 17\u201320). Distributed Adaptive Control for Second-order Leader-following Multi-agent Systems. Proceedings of the IECON 2022\u201448th Annual Conference of the IEEE Industrial Electronics Society, Brussels, Belgium.","DOI":"10.1109\/IECON49645.2022.9968889"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"4909","DOI":"10.1109\/TITS.2021.3054625","article-title":"Deep Reinforcement Learning for Autonomous Driving: A Survey","volume":"23","author":"Kiran","year":"2022","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"7879","DOI":"10.1109\/TIE.2019.2946545","article-title":"Optimized Formation Control Using Simplified Reinforcement Learning for a Class of Multiagent Systems with Unknown Dynamics","volume":"67","author":"Wen","year":"2020","journal-title":"IEEE Trans. Ind. Electron."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"4867","DOI":"10.1109\/TASE.2024.3412188","article-title":"Optimal Robust Formation of Multi-Agent Systems as Adversarial Graphical Apprentice Games with Inverse Reinforcement Learning","volume":"22","author":"Shamaghdari","year":"2025","journal-title":"IEEE Trans. Autom. Sci. Eng."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"AlMania, Z., Sheltami, T., Ahmed, G., Mahmoud, A., and Barnawi, A. (2024). Energy-Efficient Online Path Planning for Internet of Drones Using Reinforcement Learning. J. Sens. Actuator Netw., 13.","DOI":"10.3390\/jsan13050050"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"120048","DOI":"10.1016\/j.oceaneng.2024.120048","article-title":"Adaptive predefined-time specific performance control for underactuated multi-AUVs: An edge computing-based optimized RL method","volume":"318","author":"Liu","year":"2025","journal-title":"Ocean. Eng."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"131","DOI":"10.1109\/JSAC.2020.3036962","article-title":"Multi-Agent Reinforcement Learning Based Resource Management in MEC- and UAV-Assisted Vehicular Networks","volume":"39","author":"Peng","year":"2021","journal-title":"IEEE J. Sel. Areas Commun."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Xie, J., Zhou, R., Liu, Y., Luo, J., Xie, S., Peng, Y., and Pu, H. (2021). Reinforcement-Learning-Based Asynchronous Formation Control Scheme for Multiple Unmanned Surface Vehicles. Appl. Sci., 11.","DOI":"10.3390\/app11020546"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zhang, T., Li, Y., Li, S., Ye, Q., Wang, C., and Xie, G. (June, January 30). Decentralized Circle Formation Control for Fish-like Robots in the Real-world via Reinforcement Learning. Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi\u2019an, China.","DOI":"10.1109\/ICRA48506.2021.9562019"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"2935","DOI":"10.1109\/TSG.2022.3154718","article-title":"Reinforcement Learning for Selective Key Applications in Power Systems: Recent Advances and Future Challenges","volume":"13","author":"Chen","year":"2022","journal-title":"IEEE Trans. Smart Grid"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Chang, G.N., Fu, W.X., Cui, T., Song, L.Y., and Dong, P. (2024). Distributed Consensus Multi-Distribution Filter for Heavy-Tailed Noise. J. Sens. Actuator Netw., 13.","DOI":"10.3390\/jsan13040038"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"4154","DOI":"10.1109\/TAC.2019.2958840","article-title":"Feedback Linearization Based on Gaussian Processes with Event-Triggered Online Learning","volume":"65","author":"Umlauft","year":"2020","journal-title":"IEEE Trans. Autom. Control"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Berkenkamp, F., and Schoellig, A.P. (2015, January 15\u201317). Safe and robust learning control with Gaussian processes. Proceedings of the 2015 European Control Conference (ECC), Linz, Austria.","DOI":"10.1109\/ECC.2015.7330913"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Beckers, T., Hirche, S., and Colombo, L. (2021, January 13\u201317). Online Learning-based Formation Control of Multi-Agent Systems with Gaussian Processes. Proceedings of the 2021 60th IEEE Conference on Decision and Control (CDC), Austin, TX, USA.","DOI":"10.1109\/CDC45484.2021.9683423"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"3091","DOI":"10.1109\/TAC.2022.3205424","article-title":"Cooperative Control of Uncertain Multiagent Systems via Distributed Gaussian Processes","volume":"68","author":"Lederer","year":"2023","journal-title":"IEEE Trans. Autom. Control"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Dong, Z., Shao, H., and Huang, H. (2023, January 6\u20139). Variable Impedance Control for Force Tracking Based on PILCO in Uncertain Environment. Proceedings of the 2023 IEEE International Conference on Mechatronics and Automation (ICMA), Harbin, China.","DOI":"10.1109\/ICMA57826.2023.10216082"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"120500","DOI":"10.1016\/j.apenergy.2022.120500","article-title":"Multi-agent hierarchical reinforcement learning for energy management","volume":"332","author":"Jendoubi","year":"2023","journal-title":"Appl. Energy"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1505","DOI":"10.1109\/TAES.2023.3336638","article-title":"Cooperative Fault-Tolerant Formation Tracking Control for Heterogeneous Air\u2013Ground Systems Using a Learning-Based Method","volume":"60","author":"Shi","year":"2024","journal-title":"IEEE Trans. Aerosp. Electron. Syst."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Dong, X., Shi, Y., Hua, Y., Yu, J., and Ren, Z. (2025). Robust Formation Tracking Control for Multi-Agent Systems Using Reinforcement Learning Methods. Reference Module in Materials Science and Materials Engineering, Elsevier.","DOI":"10.1016\/B978-0-443-14081-5.00128-8"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Liu, S., Wen, L., Cui, J., Yang, X., Cao, J., and Liu, Y. (October, January 27). Moving Forward in Formation: A Decentralized Hierarchical Learning Approach to Multi-Agent Moving Together. Proceedings of the 2021 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Prague, Czech Republic.","DOI":"10.1109\/IROS51168.2021.9636224"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"52","DOI":"10.1109\/MSMC.2024.3401404","article-title":"Formation Tracking of Spatiotemporal Multiagent Systems: A Decentralized Reinforcement Learning Approach","volume":"10","author":"Liu","year":"2024","journal-title":"IEEE Syst. Man Cybern. Mag."},{"key":"ref_26","unstructured":"DeepSeek-AI (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv."},{"key":"ref_27","unstructured":"Andrychowicz, M., Raichuk, A., Sta\u0144czyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., and Michalski, M. (2020). What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Panteleev, A., and Karane, M. (2023). Application of a Novel Multi-Agent Optimization Algorithm Based on PID Controllers in Stochastic Control Problems. Mathematics, 11.","DOI":"10.3390\/math11132903"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"2193","DOI":"10.1109\/LCSYS.2024.3407632","article-title":"Cooperative Multi-Agent Q-Learning Using Distributed MPC","volume":"8","author":"Esfahani","year":"2024","journal-title":"IEEE Control Syst. Lett."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhou, C., Li, J., Shi, Y., and Lin, Z. (2023). Research on Multi-Robot Formation Control Based on MATD3 Algorithm. Appl. Sci., 13.","DOI":"10.3390\/app13031874"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Garcia, G., Eskandarian, A., Fabregas, E., Vargas, H., and Farias, G. (2025). Cooperative Formation Control of a Multi-Agent Khepera IV Mobile Robots System Using Deep Reinforcement Learning. Appl. Sci., 15.","DOI":"10.3390\/app15041777"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Peng, Y., Zhang, X., Jiang, Y., Xu, X., and Liu, J. (2020, January 27\u201328). Leader-Follower Formation Control For Indoor Wheeled Robots Via Dual Heuristic Programming. Proceedings of the 2020 3rd International Conference on Unmanned Systems (ICUS), Harbin, China.","DOI":"10.1109\/ICUS50048.2020.9274823"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"4577","DOI":"10.1109\/TNNLS.2020.3023711","article-title":"Robust Formation Control for Cooperative Underactuated Quadrotors via Reinforcement Learning","volume":"32","author":"Zhao","year":"2021","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."}],"container-title":["Journal of Sensor and Actuator Networks"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2224-2708\/14\/3\/51\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T17:32:12Z","timestamp":1760031132000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2224-2708\/14\/3\/51"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,14]]},"references-count":33,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2025,6]]}},"alternative-id":["jsan14030051"],"URL":"https:\/\/doi.org\/10.3390\/jsan14030051","relation":{},"ISSN":["2224-2708"],"issn-type":[{"type":"electronic","value":"2224-2708"}],"subject":[],"published":{"date-parts":[[2025,5,14]]}}}