{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,26]],"date-time":"2026-02-26T13:35:13Z","timestamp":1772112913469,"version":"3.50.1"},"reference-count":21,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2024,1,16]],"date-time":"2024-01-16T00:00:00Z","timestamp":1705363200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key R&amp;D Program of China","award":["2022YFB3104600"],"award-info":[{"award-number":["2022YFB3104600"]}]},{"name":"National Key R&amp;D Program of China","award":["2019-YF08-00285-GX"],"award-info":[{"award-number":["2019-YF08-00285-GX"]}]},{"name":"National Key R&amp;D Program of China","award":["61976156"],"award-info":[{"award-number":["61976156"]}]},{"name":"National Key R&amp;D Program of China","award":["SCITLAB-30005"],"award-info":[{"award-number":["SCITLAB-30005"]}]},{"DOI":"10.13039\/501100019014","name":"Chengdu Science and Technology Project","doi-asserted-by":"publisher","award":["2022YFB3104600"],"award-info":[{"award-number":["2022YFB3104600"]}],"id":[{"id":"10.13039\/501100019014","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100019014","name":"Chengdu Science and Technology Project","doi-asserted-by":"publisher","award":["2019-YF08-00285-GX"],"award-info":[{"award-number":["2019-YF08-00285-GX"]}],"id":[{"id":"10.13039\/501100019014","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100019014","name":"Chengdu Science and Technology Project","doi-asserted-by":"publisher","award":["61976156"],"award-info":[{"award-number":["61976156"]}],"id":[{"id":"10.13039\/501100019014","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100019014","name":"Chengdu Science and Technology Project","doi-asserted-by":"publisher","award":["SCITLAB-30005"],"award-info":[{"award-number":["SCITLAB-30005"]}],"id":[{"id":"10.13039\/501100019014","id-type":"DOI","asserted-by":"publisher"}]},{"name":"National Natural Science Foundation of China","award":["2022YFB3104600"],"award-info":[{"award-number":["2022YFB3104600"]}]},{"name":"National Natural Science Foundation of China","award":["2019-YF08-00285-GX"],"award-info":[{"award-number":["2019-YF08-00285-GX"]}]},{"name":"National Natural Science Foundation of China","award":["61976156"],"award-info":[{"award-number":["61976156"]}]},{"name":"National Natural Science Foundation of China","award":["SCITLAB-30005"],"award-info":[{"award-number":["SCITLAB-30005"]}]},{"name":"Intelligent Terminal Key Laboratory of SiChuan Province","award":["2022YFB3104600"],"award-info":[{"award-number":["2022YFB3104600"]}]},{"name":"Intelligent Terminal Key Laboratory of SiChuan Province","award":["2019-YF08-00285-GX"],"award-info":[{"award-number":["2019-YF08-00285-GX"]}]},{"name":"Intelligent Terminal Key Laboratory of SiChuan Province","award":["61976156"],"award-info":[{"award-number":["61976156"]}]},{"name":"Intelligent Terminal Key Laboratory of SiChuan Province","award":["SCITLAB-30005"],"award-info":[{"award-number":["SCITLAB-30005"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>With the development of electronic game technology, the content of electronic games presents a larger number of units, richer unit attributes, more complex game mechanisms, and more diverse team strategies. Multi-agent deep reinforcement learning shines brightly in this type of team electronic game, achieving results that surpass professional human players. Reinforcement learning algorithms based on Q-value estimation often suffer from Q-value overestimation, which may seriously affect the performance of AI in multi-agent scenarios. We propose a multi-agent mutual evaluation method and a multi-agent softmax method to reduce the estimation bias of Q values in multi-agent scenarios, and have tested them in both the particle multi-agent environment and the multi-agent tank environment we constructed. The multi-agent tank environment we have built has achieved a good balance between experimental verification efficiency and multi-agent game task simulation. It can be easily extended for different multi-agent cooperation or competition tasks. We hope that it can be promoted in the research of multi-agent deep reinforcement learning.<\/jats:p>","DOI":"10.3390\/a17010036","type":"journal-article","created":{"date-parts":[[2024,1,16]],"date-time":"2024-01-16T04:03:30Z","timestamp":1705377810000},"page":"36","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Reducing Q-Value Estimation Bias via Mutual Estimation and Softmax Operation in MADRL"],"prefix":"10.3390","volume":"17","author":[{"given":"Zheng","family":"Li","sequence":"first","affiliation":[{"name":"Center for Future Media, School of Computer Science and Engineering, and Yibin Park, University of Electronic Science and Technology of China, Chengdu 611731, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinkai","family":"Chen","sequence":"additional","affiliation":[{"name":"Center for Future Media, School of Computer Science and Engineering, and Yibin Park, University of Electronic Science and Technology of China, Chengdu 611731, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiaqing","family":"Fu","sequence":"additional","affiliation":[{"name":"Center for Future Media, School of Computer Science and Engineering, and Yibin Park, University of Electronic Science and Technology of China, Chengdu 611731, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ning","family":"Xie","sequence":"additional","affiliation":[{"name":"Center for Future Media, School of Computer Science and Engineering, and Yibin Park, University of Electronic Science and Technology of China, Chengdu 611731, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tingting","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Tianjin University of Science and Technology, Tianjin 300457, China"},{"name":"RIKEN Center for Advanced Intelligence Project (AIP), Tokyo 103-0027, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,1,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1038\/nature16961","article-title":"Mastering the game of go with deep neural networks and tree search","volume":"529","author":"Silver","year":"2016","journal-title":"Nature"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Sutton, R.S., and Barto, A.G. (1998). Reinforcement Learning: An Introduction, MIT Press.","DOI":"10.1109\/TNN.1998.712192"},{"key":"ref_3","first-page":"2613","article-title":"Double Q-learning","volume":"23","author":"Hasselt","year":"2010","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_4","unstructured":"van Hasselt, H. (2011). Insight in Reinforcement Learning. Formal Analysis and Empirical Evaluation of Temporal-Difference Algorithms. [Ph.D. Thesis, Utrecht University]."},{"key":"ref_5","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal policy optimization algorithms. arXiv."},{"key":"ref_6","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013). Playing atari with deep reinforcement learning. arXiv."},{"key":"ref_7","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015). Continuous control with deep reinforcement learning. arXiv."},{"key":"ref_8","unstructured":"Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016, January 20\u201322). Asynchronous methods for deep reinforcement learning. Proceedings of the International Conference on Machine Learning, New York, NY, USA. PMLR."},{"key":"ref_9","unstructured":"Fujimoto, S., Hoof, H., and Meger, D. (2018, January 10\u201315). Addressing function approximation error in actor-critic methods. Proceedings of the International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_10","unstructured":"Ackermann, J., Gabler, V., Osa, T., and Sugiyama, M. (2019). Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"677","DOI":"10.1037\/0003-066X.57.9.677","article-title":"If at First You don\u2019t Succeed: False Hopes of Self-change","volume":"57","author":"Polivy","year":"2002","journal-title":"Am. Psychol."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"806","DOI":"10.1037\/0022-3514.39.5.806","article-title":"Unrealistic optimism about future life events","volume":"39","author":"Weinstein","year":"1980","journal-title":"J. Personal. Soc. Psychol."},{"key":"ref_13","unstructured":"Anschel, O., Baram, N., and Shimkin, N. (2017, January 13\u201315). Averaged-DQN: Variance Reduction and Stabilization for Deep Reinforcement Learning. Proceedings of the International Conference on Machine Learning, Mountain View, CA, USA."},{"key":"ref_14","unstructured":"Song, Z., Parr, R., and Carin, L. (2019, January 10\u201315). Revisiting the softmax bellman operator: New benefits and new perspective. Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA."},{"key":"ref_15","first-page":"1365","article-title":"Regularized softmax deep multi-agent q-learning","volume":"34","author":"Pan","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_16","first-page":"11767","article-title":"Softmax deep double deterministic policy gradients","volume":"33","author":"Pan","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_17","first-page":"6379","article-title":"Multi-agent actor-critic for mixed cooperative-competitive environments","volume":"30","author":"Lowe","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Lu, Y., Li, W., and Li, W. (2023). Official International Mahjong: A New Playground for AI Research. Algorithms, 16.","DOI":"10.3390\/a16050235"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"253","DOI":"10.1613\/jair.3912","article-title":"The arcade learning environment: An evaluation platform for general agents","volume":"47","author":"Bellemare","year":"2013","journal-title":"J. Artif. Intell. Res."},{"key":"ref_20","unstructured":"Vinyals, O., Ewalds, T., Bartunov, S., Georgiev, P., Vezhnevets, A.S., Yeo, M., Makhzani, A., K\u00fcttler, H., Agapiou, J., and Schrittwieser, J. (2017). Starcraft ii: A new challenge for reinforcement learning. arXiv."},{"key":"ref_21","first-page":"11881","article-title":"Honor of kings arena: An environment for generalization in competitive reinforcement learning","volume":"35","author":"Wei","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/17\/1\/36\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T13:47:36Z","timestamp":1760104056000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/17\/1\/36"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,16]]},"references-count":21,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,1]]}},"alternative-id":["a17010036"],"URL":"https:\/\/doi.org\/10.3390\/a17010036","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,16]]}}}