{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,19]],"date-time":"2026-03-19T06:20:10Z","timestamp":1773901210271,"version":"3.50.1"},"reference-count":32,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2021,12,30]],"date-time":"2021-12-30T00:00:00Z","timestamp":1640822400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Environmental information plays an important role in deep reinforcement learning (DRL). However, many algorithms do not pay much attention to environmental information. In multi-agent reinforcement learning decision-making, because agents need to make decisions combined with the information of other agents in the environment, this makes the environmental information more important. To prove the importance of environmental information, we added environmental information to the algorithm. We evaluated many algorithms on a challenging set of StarCraft II micromanagement tasks. Compared with the original algorithm, the standard deviation (except for the VDN algorithm) was smaller than that of the original algorithm, which shows that our algorithm has better stability. The average score of our algorithm was higher than that of the original algorithm (except for VDN and COMA), which shows that our work significantly outperforms existing multi-agent RL methods.<\/jats:p>","DOI":"10.3390\/fi14010017","type":"journal-article","created":{"date-parts":[[2021,12,30]],"date-time":"2021-12-30T21:41:21Z","timestamp":1640900481000},"page":"17","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["The Important Role of Global State for Multi-Agent Reinforcement Learning"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6784-0563","authenticated-orcid":false,"given":"Shuailong","family":"Li","sequence":"first","affiliation":[{"name":"State Key Laboratory of Robotics, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Institutes for Robotics and Intelligent Manufacturing, Chinese Academy of Sciences, Shenyang 110169, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Zhang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Robotics, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4063-4545","authenticated-orcid":false,"given":"Yuquan","family":"Leng","sequence":"additional","affiliation":[{"name":"Shenzhen Key Laboratory of Biomimetic Robotics and Intelligent Systems, Department of Mechanical and Energy Engineering, Southern University of Science and Technology, Shenzhen 518055, China"},{"name":"Guangdong Provincial Key Laboratory of Human-Augmentation and Rehabilitation Robotics in Universities, Southern University of Science and Technology, Shenzhen 518055, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaohui","family":"Wang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Robotics, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China"},{"name":"Institutes for Robotics and Intelligent Manufacturing, Chinese Academy of Sciences, Shenyang 110169, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,12,30]]},"reference":[{"key":"ref_1","unstructured":"Sun, R. (2019). Optimization for deep learning: Theory and algorithms. arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Whitman, J., Bhirangi, R., Travers, M., and Choset, H. (2020, January 7\u201312). Modular robot design synthesis with deep reinforcement learning. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i06.6611"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1145\/3065386","article-title":"ImageNet classification with deep convolutional neural networks","volume":"60","author":"Krizhevsky","year":"2017","journal-title":"Commun. ACM"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"2727","DOI":"10.1007\/s00521-017-3225-z","article-title":"Accurate photovoltaic power forecasting models using deep LSTM-RNN","volume":"31","author":"Mahmoud","year":"2019","journal-title":"Neural Comput. Appl."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zhou, Q., Li, H., and Wang, J. (2020, January 7\u201312). Deep model-based reinforcement learning via estimated uncertainty and conservative policy optimization. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i04.6177"},{"key":"ref_6","unstructured":"Huang, W., Pham, V.H., and Haskell, W.B. (2020, January 7\u201312). Model and reinforcement learning for markov games with risk preferences. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Kosugi, S., and Yamasaki, T. (2020, January 7\u201312). Unpaired image enhancement featuring reinforcement-learning-controlled image editing software. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6790"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Jain, V., Fedus, W., Larochelle, H., Precup, D., and Bellemare, M.G. (2020, January 7\u201312). Algorithmic improvements for deep reinforcement learning applied to interactive fiction. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i04.5857"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"G\u00e4rtner, E., Pirinen, A., and Sminchisescu, C. (2020, January 7\u201312). Deep Reinforcement Learning for Active Human Pose Estimation. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6714"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Liu, Y., Liu, Q., Zhao, H., Pan, Z., and Liu, C. (2020, January 7\u201312). Adaptive quantitative trading: An imitative deep reinforcement learning approach. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i02.5587"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1038\/nature16961","article-title":"Mastering the game of Go with deep neural networks and tree search","volume":"529","author":"Silver","year":"2016","journal-title":"Nature"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1038\/nature24270","article-title":"Mastering the game of go without human knowledge","volume":"550","author":"Silver","year":"2017","journal-title":"Nature"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"750","DOI":"10.1007\/s10458-019-09421-1","article-title":"A survey and critique of multiagent deep reinforcement learning","volume":"33","author":"Kartal","year":"2019","journal-title":"Auton. Agents Multi-Agent Syst."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"345","DOI":"10.1023\/A:1008942012299","article-title":"Multiagent systems: A survey from a machine learning perspective","volume":"8","author":"Stone","year":"2000","journal-title":"Auton. Robot."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"365","DOI":"10.1016\/j.artint.2006.02.006","article-title":"If multi-agent learning is the answer, what is the question?","volume":"171","author":"Shoham","year":"2007","journal-title":"Artif. Intell."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"277","DOI":"10.1017\/S0269888901000170","article-title":"Learning in multi-agent systems","volume":"16","author":"Alonso","year":"2001","journal-title":"Knowl. Eng. Rev."},{"key":"ref_17","first-page":"41","article-title":"Multiagent learning: Basics, challenges, and prospects","volume":"33","author":"Tuyls","year":"2012","journal-title":"AI Mag."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"156","DOI":"10.1109\/TSMCC.2007.913919","article-title":"A comprehensive survey of multiagent reinforcement learning","volume":"38","author":"Busoniu","year":"2008","journal-title":"IEEE Trans. Syst. Man Cybern. Part C (Appl. Rev.)"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Now\u00e9, A., Vrancx, P., and De Hauwere, Y.M. (2012). Game theory and multi-agent reinforcement learning. Reinforcement Learning, Springer.","DOI":"10.1007\/978-3-642-27645-3_14"},{"key":"ref_20","unstructured":"Rashid, T., Samvelyan, M., Schroeder, C., Farquhar, G., Foerster, J., and Whiteson, S. (2018, January 10\u201315). Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning. Proceedings of the International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_21","unstructured":"Russell, S.J., and Zimdars, A. (2003, January 21\u201324). Q-decomposition for reinforcement learning agents. Proceedings of the 20th International Conference on Machine Learning (ICML-03), Washington, DC, USA."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Tan, M. (1993, January 27\u201329). Multi-agent reinforcement learning: Independent vs. cooperative agents. Proceedings of the Tenth International Conference on Machine Learning, Amherst, MA, USA.","DOI":"10.1016\/B978-1-55860-307-3.50049-6"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S. (2018, January 2\u20137). Counterfactual multi-agent policy gradients. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11794"},{"key":"ref_24","unstructured":"Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W.M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J.Z., and Tuyls, K. (2017). Value-decomposition networks for cooperative multi-agent learning. arXiv."},{"key":"ref_25","unstructured":"Hausknecht, M., and Stone, P. (2015, January 12\u201314). Deep recurrent q-learning for partially observable mdps. Proceedings of the 2015 AAAI Fall Symposium Series, Arlington, VA, USA."},{"key":"ref_26","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013). Playing atari with deep reinforcement learning. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_28","unstructured":"Schaul, T., Quan, J., and Antonoglou, I. (2016, January 2\u20134). Prioritized experience replay [C\/OL]. Proceedings of the 4th Inter national Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Bellemare, M.G., Ostrovski, G., Guez, A., Thomas, P., and Munos, R. (2016, January 12\u201317). Increasing the action gap: New operators for reinforcement learning. Proceedings of the AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA.","DOI":"10.1609\/aaai.v30i1.10303"},{"key":"ref_30","first-page":"4026","article-title":"Deep exploration via bootstrapped DQN","volume":"29","author":"Osband","year":"2016","journal-title":"Adv. Neural Inf. Processing Syst."},{"key":"ref_31","unstructured":"Synnaeve, G., Nardelli, N., Auvolat, A., Chintala, S., Lacroix, T., Lin, Z., Richoux, F., and Usunier, N. (2016). Torchcraft: A library for machine learning research on real-time strategy games. arXiv."},{"key":"ref_32","unstructured":"Vinyals, O., Ewalds, T., Bartunov, S., Georgiev, P., Vezhnevets, A.S., Yeo, M., Makhzani, A., K\u00fcttler, H., Agapiou, J., and Schrittwieser, J. (2017). Starcraft ii: A new challenge for reinforcement learning. arXiv."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/1\/17\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:56:15Z","timestamp":1760169375000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/1\/17"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,30]]},"references-count":32,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2022,1]]}},"alternative-id":["fi14010017"],"URL":"https:\/\/doi.org\/10.3390\/fi14010017","relation":{},"ISSN":["1999-5903"],"issn-type":[{"value":"1999-5903","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12,30]]}}}