{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T03:23:26Z","timestamp":1783481006813,"version":"3.55.0"},"reference-count":34,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2024,12,25]],"date-time":"2024-12-25T00:00:00Z","timestamp":1735084800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["91948303"],"award-info":[{"award-number":["91948303"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Multi-agent systems often face challenges such as elevated communication demands, intricate interactions, and difficulties in transferability. To address the issues of complex information interaction and model scalability, we propose an innovative hierarchical graph attention actor\u2013critic reinforcement learning method. This method naturally models the interactions within a multi-agent system as a graph, employing hierarchical graph attention to capture the complex cooperative and competitive relationships among agents, thereby enhancing their adaptability to dynamic environments. Specifically, graph neural networks encode agent observations as single feature-embedding vectors, maintaining a constant dimensionality irrespective of the number of agents, which improves model scalability. Through the \u201cinter-agent\u201d and \u201cinter-group\u201d attention layers, the embedding vector of each agent is updated into an information-condensed and contextualized state representation, which extracts state-dependent relationships between agents and model interactions at both individual and group levels. We conducted experiments across several multi-agent tasks to assess our proposed method\u2019s effectiveness, stability, and scalability. Furthermore, to enhance the applicability of our method in large-scale tasks, we tested and validated its performance within a curriculum learning training framework, thereby enhancing its transferability.<\/jats:p>","DOI":"10.3390\/e27010004","type":"journal-article","created":{"date-parts":[[2024,12,25]],"date-time":"2024-12-25T19:29:52Z","timestamp":1735154992000},"page":"4","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":14,"title":["Multi-Agent Hierarchical Graph Attention Actor\u2013Critic Reinforcement Learning"],"prefix":"10.3390","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3975-6852","authenticated-orcid":false,"given":"Tongyue","family":"Li","sequence":"first","affiliation":[{"name":"Academy of Military Sciences, Beijing 100097, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8112-371X","authenticated-orcid":false,"given":"Dianxi","family":"Shi","sequence":"additional","affiliation":[{"name":"Academy of Military Sciences, Beijing 100097, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5959-0768","authenticated-orcid":false,"given":"Songchang","family":"Jin","sequence":"additional","affiliation":[{"name":"Academy of Military Sciences, Beijing 100097, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3291-0768","authenticated-orcid":false,"given":"Zhen","family":"Wang","sequence":"additional","affiliation":[{"name":"Academy of Military Sciences, Beijing 100097, China"},{"name":"Tianjin Artificial Intelligence Innovation Center (TAIIC), Tianjin 300450, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9632-2543","authenticated-orcid":false,"given":"Huanhuan","family":"Yang","sequence":"additional","affiliation":[{"name":"College of Computer, National University of Defense Technology, Changsha 410073, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3414-3328","authenticated-orcid":false,"given":"Yang","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Computer Science, Peking University, Beijing 100871, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,12,25]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Franks, N.R., Worley, A., Grant, K.A., Gorman, A.R., Vizard, V., Plackett, H., Doran, C., Gamble, M.L., Stumpe, M.C., and Sendova-Franks, A.B. (2016). Social behaviour and collective motion in plant-animal worms. Proc. R. Soc. B Biol. Sci., 283.","DOI":"10.1098\/rspb.2015.2946"},{"key":"ref_2","unstructured":"Perolat, J., Leibo, J.Z., Zambaldi, V., Beattie, C., Tuyls, K., and Graepel, T. (2017, January 4\u20139). A multi-agent reinforcement learning model of common-pool resource appropriation. Proceedings of the Advances in Neural Information Processing Systems 30 (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"326","DOI":"10.1016\/j.jmsy.2020.06.018","article-title":"Reinforcement learning for facilitating human-robot-interaction in manufacturing","volume":"56","author":"Oliff","year":"2020","journal-title":"J. Manuf. Syst."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Zhang, K., Yang, Z., and Ba\u015far, T. (2021). Multi-agent reinforcement learning: A selective overview of theories and algorithms. Handbook of Reinforcement Learning and Control, Springer.","DOI":"10.1007\/978-3-030-60990-0_12"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"103728","DOI":"10.1016\/j.trc.2022.103728","article-title":"CVLight: Decentralized learning for adaptive traffic signal control with connected vehicles","volume":"141","author":"Mo","year":"2022","journal-title":"Transp. Res. Part C Emerg. Technol."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1321","DOI":"10.1007\/s10514-016-9579-8","article-title":"Distributed on-line dynamic task assignment for multi-robot patrolling","volume":"41","author":"Farinelli","year":"2017","journal-title":"Auton. Robot."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2358","DOI":"10.1109\/TNNLS.2020.3004893","article-title":"Formation control with collision avoidance through deep reinforcement learning using model-guided demonstration","volume":"32","author":"Sui","year":"2020","journal-title":"IEEE Trans. Neural Networks Learn. Syst."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Liu, L., Luo, C., and Shen, F. (2017, January 18\u201320). Multi-agent formation control with target tracking and navigation. Proceedings of the 2017 IEEE International Conference on Information and Automation (ICIA), Macau, Chin.","DOI":"10.1109\/ICInfA.2017.8078889"},{"key":"ref_9","unstructured":"Ryu, H., Shin, H., and Park, J. (2021, January 3\u20137). Cooperative and competitive biases for multi-agent reinforcement learning. Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems, Virtual. AAMAS \u201921."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Hahn, C., Ritz, F., Wikidal, P., Phan, T., Gabor, T., and Linnhoff-Popien, C. (2020, January 13\u201318). Foraging swarms using multi-agent reinforcement learning. Proceedings of the ALIFE 2020: The 2020 Conference on Artificial Life, Online.","DOI":"10.1162\/isal_a_00267"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"934","DOI":"10.1016\/j.engappai.2011.09.025","article-title":"Bio-inspired multi-agent systems for reconfigurable manufacturing systems","volume":"25","author":"Barbosa","year":"2012","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Stadler, M., Banfi, J., and Roy, N. (2023, January 8\u201313). Approximating the value of collaborative team actions for efficient multiagent navigation in uncertain graphs. Proceedings of the International Conference on Automated Planning and Scheduling, Prague, Czech Republic.","DOI":"10.1609\/icaps.v33i1.27250"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Tassel, P., Kov\u00e1cs, B., Gebser, M., Schekotihin, K., Kohlenbrein, W., and Schrott-Kostwein, P. (2022, January 13\u201324). Reinforcement learning of dispatching strategies for large-scale industrial scheduling. Proceedings of the International Conference on Automated Planning and Scheduling, Virtual.","DOI":"10.1609\/icaps.v32i1.19852"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"102318","DOI":"10.1016\/j.inffus.2024.102318","article-title":"Hierarchical relationship modeling in multi-agent reinforcement learning for mixed cooperative\u2013competitive environments","volume":"108","author":"Xie","year":"2024","journal-title":"Inf. Fusion"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"847","DOI":"10.1007\/978-981-16-9416-5_62","article-title":"UAV collaboration for autonomous target capture","volume":"Volume 1","author":"Tony","year":"2022","journal-title":"Proceedings of the Congress on Intelligent Systems: Proceedings of CIS 2021"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1660","DOI":"10.1177\/0278364915602321","article-title":"Cooperative multi-robot control for target tracking with onboard sensing","volume":"34","author":"Hausman","year":"2015","journal-title":"Int. J. Robot. Res."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"107677","DOI":"10.1016\/j.ast.2022.107677","article-title":"All-aspect attack guidance law for agile missiles based on deep reinforcement learning","volume":"127","author":"Gong","year":"2022","journal-title":"Aerosp. Sci. Technol."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"105996","DOI":"10.1016\/j.ast.2020.105996","article-title":"Cooperative online guide-launch-guide policy in a target-missile-defender engagement using deep reinforcement learning","volume":"104","author":"Shalumov","year":"2020","journal-title":"Aerosp. Sci. Technol."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Wu, J., and Huang, Z. (2023, January 21\u201325). Promoting diversity in mixed complex cooperative and competitive multi-agent environment. Proceedings of the 32nd ACM International Conference on Information and Knowledge Management, Birmingham, UK. CIKM \u201923.","DOI":"10.1145\/3583780.3615217"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"15051","DOI":"10.1109\/TNNLS.2023.3283523","article-title":"Challenges and opportunities in deep reinforcement learning with graph neural networks: A comprehensive review of algorithms and applications","volume":"35","author":"Munikoti","year":"2022","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_21","unstructured":"Agarwal, A., Kumar, S., and Sycara, K.P. (2019). Learning transferable cooperative behavior in multi-agent teams. arXiv."},{"key":"ref_22","unstructured":"Niu, Y., Paleja, R.R., and Gombolay, M.C. (2021, January 3\u20137). Multi-agent graph-attention communication and teaming. Proceedings of the AAMAS, Virtual."},{"key":"ref_23","unstructured":"Ma, X., Yang, Y., Li, C., Lu, Y., Zhao, Q., and Jun, Y. (2021). Modeling the interaction between agents in cooperative multi-agent reinforcement learning. arXiv."},{"key":"ref_24","unstructured":"Iqbal, S., and Sha, F. (2019, January 9\u201315). Actor-attention-critic for multi-agent reinforcement learning. Proceedings of the International Conference on Machine Learning, PMLR, Long Beach, CA, USA."},{"key":"ref_25","unstructured":"Su, J., Adams, S., and Beling, P.A. (2020). Counterfactual multi-agent reinforcement learning with graph convolution communication. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Liu, Y., Wang, W., Hu, Y., Hao, J., Chen, X., and Gao, Y. (2019). Multi-agent game abstraction via graph attention neural network. arXiv.","DOI":"10.1609\/aaai.v34i05.6211"},{"key":"ref_27","unstructured":"Jiang, J., Dun, C., Huang, T., and Lu, Z. (2018). Graph convolutional reinforcement learning. arXiv."},{"key":"ref_28","first-page":"20038","article-title":"Interaction modeling with multiplex attention","volume":"35","author":"Sun","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_29","first-page":"1","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_30","unstructured":"Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., and Abbeel, P. (2018). Soft actor-critic algorithms and applications. arXiv."},{"key":"ref_31","unstructured":"Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018, January 10\u201315). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. Proceedings of the International Conference on Machine Learning, PMLR, Stockholm, Sweden."},{"key":"ref_32","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_33","unstructured":"Lowe, R., Wu, Y.I., Tamar, A., Harb, J., Abbeel, O.P., and Mordatch, I. (2017, January 4\u20139). Multi-agent actor-critic for mixed cooperative-competitive environments. Proceedings of the Advances in Neural Information Processing Systems 30 (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_34","first-page":"4555","article-title":"A survey on curriculum learning","volume":"44","author":"Wang","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/1\/4\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T16:59:51Z","timestamp":1760115591000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/1\/4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,25]]},"references-count":34,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,1]]}},"alternative-id":["e27010004"],"URL":"https:\/\/doi.org\/10.3390\/e27010004","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,25]]}}}