{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T11:17:29Z","timestamp":1781954249505,"version":"3.54.5"},"reference-count":53,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2024,12,15]],"date-time":"2024-12-15T00:00:00Z","timestamp":1734220800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No. 62273356"],"award-info":[{"award-number":["No. 62273356"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>The prevalent utilization of deterministic strategy algorithms in Multi-Agent Deep Reinforcement Learning (MADRL) for collaborative tasks has posed a significant challenge in achieving stable and high-performance cooperative behavior. Addressing the need for the balanced exploration and exploitation of multi-agent ant robots within a partially observable continuous action space, this study introduces a multi-agent centralized strategy gradient algorithm grounded in a local state transition mechanism. In order to solve this challenge, the algorithm learns local state and local state-action representation from local observations and action values, thereby establishing a \u201clocal state transition\u201d mechanism autonomously. As the input of the actor network, the automatically extracted local observation representation reduces the input state dimension, enhances the local state features closely related to the local state transition, and promotes the agent to use the local state features that affect the next observation state. To mitigate non-stationarity and reliability assignment issues in multi-agent environments, a centralized critic network evaluates the current joint strategy. The proposed algorithm, NST-FACMAC, is evaluated alongside other multi-agent deterministic strategy algorithms in a continuous control simulation environment using a multi-agent ant robot. The experimental results indicate accelerated convergence and higher average reward values in cooperative multi-agent ant simulation environments. Notably, in four simulated environments named Ant-v2 (2 \u00d7 4), Ant-v2 (2 \u00d7 4d), Ant-v2 (4 \u00d7 2), and Manyant (2 \u00d7 3), the algorithm demonstrates performance improvements of approximately 1.9%, 4.8%, 11.9%, and 36.1%, respectively, compared to the best baseline algorithm. These findings underscore the algorithm\u2019s effectiveness in enhancing the stability of multi-agent ant robot control within dynamic environments.<\/jats:p>","DOI":"10.3390\/a17120579","type":"journal-article","created":{"date-parts":[[2024,12,16]],"date-time":"2024-12-16T10:08:53Z","timestamp":1734343733000},"page":"579","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["A Multi-Agent Centralized Strategy Gradient Reinforcement Learning Algorithm Based on State Transition"],"prefix":"10.3390","volume":"17","author":[{"given":"Lei","family":"Sheng","sequence":"first","affiliation":[{"name":"National Key Laboratory of Information Systems Engineering, National University of Defense Technology, Changsha 410073, China"},{"name":"School of Command and Control Engineering, Army Engineering University, Nanjing 210007, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Honghui","family":"Chen","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Information Systems Engineering, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiliang","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Command and Control Engineering, Army Engineering University, Nanjing 210007, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,12,15]]},"reference":[{"key":"ref_1","first-page":"3338","article-title":"Deep learning for real-time Atari game play using offline Monte-Carlo tree search planning","volume":"4","author":"Guo","year":"2014","journal-title":"Int. Conf. Neural Inf. Process. Syst."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1038\/nature16961","article-title":"Mastering the game of Go with deep neural networks and tree search","volume":"529","author":"Silver","year":"2016","journal-title":"Nature"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Volodymyr","year":"2015","journal-title":"Nature"},{"key":"ref_4","unstructured":"Laskin, M., Srinivas, A., and Abbeel, P. (2020, January 13\u201318). Curl: Contrastive unsupervised representations for reinforcement learning. Proceedings of the International Conference on Machine Learning, PMLR, Virtual."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1072","DOI":"10.1016\/j.apenergy.2018.11.002","article-title":"Reinforcement learning for demand response: A review of algorithms and modeling techniques","volume":"235","author":"Nagy","year":"2019","journal-title":"Appl. Energy"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"711","DOI":"10.1177\/09544100241235824","article-title":"Single-lever control method design based on power management system and deep reinforcement learning for turboprop engines","volume":"238","author":"Ji","year":"2024","journal-title":"Proc. Inst. Mech. Eng. Part G J. Aerosp. Eng."},{"key":"ref_7","first-page":"2","article-title":"Towards risk-aware real-time security constrained economic dispatch: A tailored deep reinforcement learning approach","volume":"2","author":"Hu","year":"2024","journal-title":"IEEE Trans. Power Syst. A Publ. Power Eng. Soc."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"113573","DOI":"10.1016\/j.eswa.2020.113573","article-title":"An intelligent financial portfolio trading strategy using deep Q-learning","volume":"158","author":"Park","year":"2020","journal-title":"Expert Syst. Appl."},{"key":"ref_9","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015). Continuous control with deep reinforcement learning. ArXiv."},{"key":"ref_10","unstructured":"Fujimoto, S., Hoof, H., and Meger, D. (2018, January 10\u201315). Addressing function approximation error in actor-critic methods. Proceedings of the International Conference on Machine Learning, PMLR, Stockholm, Sweden."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1108\/IR-11-2019-0240","article-title":"Deep reinforcement learning-based attitude motion control for humanoid robots with stability constraints","volume":"47","author":"Shi","year":"2020","journal-title":"Ind. Robot. Int. J. Robot. Res. Appl."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"103905","DOI":"10.1016\/j.robot.2021.103905","article-title":"Multi-robot task allocation in disaster response: Addressing dynamic tasks with deadlines and robots with range and payload constraints","volume":"147","author":"Ghassemi","year":"2022","journal-title":"Robot. Auton. Syst."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"105","DOI":"10.1007\/s43154-020-00039-w","article-title":"Cooperative multirobot systems for military applications","volume":"2","author":"Gans","year":"2021","journal-title":"Curr. Robot. Rep."},{"key":"ref_14","first-page":"1","article-title":"Monotonic value function factorisation for deep multi-agent reinforcement learning","volume":"21","author":"Rashid","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"ref_15","first-page":"73","article-title":"A survey on multi-agent reinforcement learning and its application","volume":"3","author":"Ning","year":"2024","journal-title":"J. Autom. Intell."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"5023","DOI":"10.1007\/s10462-022-10299-x","article-title":"Deep multiagent reinforcement learning: Challenges and directions","volume":"56","author":"Wong","year":"2023","journal-title":"Artif. Intell. Rev."},{"key":"ref_17","first-page":"29142","article-title":"Towards understanding cooperative multi-agent q-learning with value factorization","volume":"34","author":"Wang","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Xu, Z., Li, D., Bai, Y., and Fan, G. (2021, January 18\u201322). Mmd-mix: Value function factorisation with maximum mean discrepancy for cooperative multi-agent reinforcement learning. Proceedings of the 2021 International Joint Conference on Neural Networks (IJCNN), Shenzhen, China.","DOI":"10.1109\/IJCNN52387.2021.9533636"},{"key":"ref_19","first-page":"24018","article-title":"Vast: Value function factorization with variable agent sub-teams","volume":"34","author":"Phan","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_20","first-page":"1","article-title":"Multi-agent actor-critic for mixed cooperative-competitive environments","volume":"30","author":"Lowe","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"206","DOI":"10.1016\/j.neucom.2020.05.097","article-title":"A TD3-based multi-agent deep reinforcement learning method in mixed cooperation-competition environment","volume":"411","author":"Zhang","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_22","first-page":"12208","article-title":"Facmac: Factored multi-agent centralised policy gradients","volume":"34","author":"Peng","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_23","first-page":"15230","article-title":"Learning to ground multi-agent communication with autoencoders","volume":"34","author":"Lin","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_24","unstructured":"Phan, T., Ritz, F., Altmann, P., Zorn, M., N\u00fc\u00dflein, J., K\u00f6lle, M., and Linnhoff-Popien, C. (2023, January 23\u201329). Attention-based recurrence for multi-agent reinforcement learning under stochastic partial observability. Proceedings of the International Conference on Machine Learning, PMLR, Honolulu, HI, USA."},{"key":"ref_25","unstructured":"Zhang, T., Li, Y., Wang, C., Xie, G., and Lu, Z. (2021, January 18\u201324). Fop: Factorizing optimal joint policy of maximum-entropy multi-agent reinforcement learning. Proceedings of the International Conference on Machine Learning, PMLR, Virtual."},{"key":"ref_26","first-page":"5510","article-title":"Towards a standardised performance evaluation protocol for cooperative marl","volume":"35","author":"Gorsane","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Canese, L., Cardarilli, G.C., Di Nunzio, L., Fazzolari, R., Giardino, D., Re, M., and Span\u00f2, S. (2021). Multi-agent reinforcement learning: A review of challenges and applications. Appl. Sci., 11.","DOI":"10.3390\/app11114948"},{"key":"ref_28","unstructured":"Iqbal, S., De Witt, C.A.S., Peng, B., B\u00f6hmer, W., Whiteson, S., and Sha, F. (2021, January 18\u201324). Randomized entity-wise factorization for multi-agent reinforcement learning. Proceedings of the International Conference on Machine Learning, PMLR, Virtual."},{"key":"ref_29","first-page":"10199","article-title":"Weighted qmix: Expanding monotonic value function factorisation for deep multi-agent reinforcement learning","volume":"33","author":"Rashid","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"191","DOI":"10.1016\/j.ins.2022.10.042","article-title":"Value function factorization with dynamic weighting for deep multi-agent reinforcement learning","volume":"615","author":"Du","year":"2022","journal-title":"Inf. Sci."},{"key":"ref_31","first-page":"5471","article-title":"ResQ: A Residual Q Function-based Approach for Multi-Agent Reinforcement Learning Value Factorization","volume":"35","author":"Shen","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"6550","DOI":"10.1109\/TSMC.2024.3370186","article-title":"SQIX: QMIX Algorithm Activated by General Softmax Operator for Cooperative Multiagent Reinforcement Learning","volume":"54","author":"Zhang","year":"2024","journal-title":"IEEE Trans. Syst. Man Cybern. Syst."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1146","DOI":"10.1016\/j.procs.2017.05.431","article-title":"An adaptive implementation of \u03b5-greedy in reinforcement learning","volume":"109","year":"2017","journal-title":"Procedia Comput. Sci."},{"key":"ref_34","first-page":"5093","article-title":"Understanding deep neural function approximation in reinforcement learning via \u03f5-greedy exploration","volume":"35","author":"Liu","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_35","unstructured":"Wang, S., Pu, Y., Yang, S., Yao, X., and Li, B. (2020, January 23\u201327). Boltzmann Exploration for Deterministic Policy Optimization. Proceedings of the Neural Information Processing: 27th International Conference, ICONIP 2020, Bangkok, Thailand. Proceedings, Part II 27."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"3103","DOI":"10.1109\/TCYB.2020.2977661","article-title":"Deep reinforcement learning for multi-objective optimization","volume":"51","author":"Li","year":"2020","journal-title":"IEEE Trans. Cybern."},{"key":"ref_37","first-page":"22405","article-title":"Swad: Domain generalization by seeking flat minima","volume":"34","author":"Cha","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_38","first-page":"1","article-title":"Is Q-learning provably efficient?","volume":"12","author":"Rastogi","year":"2018","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_39","unstructured":"Machado, M.C., Bellemare, M.G., and Bowling, M. (2020, January 7\u201312). Count-based exploration with the successor representation. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA."},{"key":"ref_40","first-page":"84","article-title":"Intrinsic motivation based on feature extractor distillation","volume":"1\u20136","year":"2022","journal-title":"Kognice Umel\u00fd Zivot XX"},{"key":"ref_41","unstructured":"Liu, B., Pu, Z., Pan, Y., Yi, J., Liang, Y., and Zhang, D. (2023, January 23\u201329). Lazy agents: A new perspective on solving sparse reward problem in multi-agent reinforcement learning. Proceedings of the International Conference on Machine Learning, PMLR, Honolulu, HI, USA."},{"key":"ref_42","unstructured":"Fujimoto, S., Meger, D., Precup, D., Nachum, O., and Gu, S.S. (2022, January 23\u201329). Why Should I Trust You, Bellman? The Bellman Error is a Poor Replacement for Value Error. Proceedings of the International Conference on Machine Learning (ICML), Honolulu, HI, USA."},{"key":"ref_43","unstructured":"Stooke, A., Lee, K., Abbeel, P., and Laskin, M. (2021, January 18\u201324). Decoupling representation learning from reinforcement learning. Proceedings of the International Conference on Machine Learning, Online."},{"key":"ref_44","unstructured":"Yarats, D., Fergus, R., Lazaric, A., and Pinto, L. (2022, January 25\u201329). Mastering visual continuous control: Improved data-augmented reinforcement learning. Proceedings of the International Conference on Learning Representations (ICLR), Online."},{"key":"ref_45","unstructured":"Liu, G., Zhang, C., Zhao, L., Qin, T., Zhu, J., Li, J., Yu, N., and Liu, T.-Y. (2021, January 4\u20138). Return-based contrastive representation learning for reinforcement learning. Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria."},{"key":"ref_46","unstructured":"Cetin, E., Ball, P.J., Roberts, S., and Celiktutan, O. (2022, January 23\u201329). Stabilizing off-policy deep reinforcement learning from pixels. Proceedings of the International Conference on Machine Learning (ICML), Honolulu, HI, USA."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"299","DOI":"10.1523\/JNEUROSCI.1327-21.2021","article-title":"Predictive representations in hippocampal and prefrontal hierarchies","volume":"42","author":"Brunec","year":"2022","journal-title":"J. Neurosci."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Kang, C., Zhang, H., Liu, Z., Huang, S., and Yin, Y. (2022). LR-GNN: A graph neural network based on link representation for predicting molecular associations. Brief. Bioinform., 23.","DOI":"10.1093\/bib\/bbab513"},{"key":"ref_49","unstructured":"Hansen, N., Wang, X., and Su, H. (2022, January 23\u201329). Temporal Difference Learning for Model Predictive Control. Proceedings of the International Conference on Machine Learning (ICML), Honolulu, HI, USA."},{"key":"ref_50","unstructured":"Li, Y., Li, S., Sitzmann, V., Agrawal, P., and Torralba, A. (2022, January 14\u201318). 3d neural scene representations for visuomotor control. Proceedings of the Conference on Robot Learning, PMLR, Auckland, New Zealand."},{"key":"ref_51","unstructured":"Gelada, C., Kumar, S., Buckman, J., Nachum, O., and Bellemare, M.G. (2019, January 10\u201315). Deepmdp: Learning continuous latent space models for representation learning. Proceedings of the International Conference on Machine Learning (ICML), Long Beach, CA, USA."},{"key":"ref_52","unstructured":"Wang, J., Ren, Z., Liu, T., Yu, Y., and Zhang, C. (2021, January 4\u20138). QPLEX: Duplex Dueling Multi-Agent Q-Learning. Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"9541","DOI":"10.1007\/s10462-022-10335-w","article-title":"High-accuracy model-based reinforcement learning, a survey","volume":"56","author":"Plaat","year":"2023","journal-title":"Artif. Intell. Rev."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/17\/12\/579\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T16:52:39Z","timestamp":1760115159000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/17\/12\/579"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,15]]},"references-count":53,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["a17120579"],"URL":"https:\/\/doi.org\/10.3390\/a17120579","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,15]]}}}