{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T17:45:12Z","timestamp":1784137512913,"version":"3.55.0"},"reference-count":188,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2025,3,15]],"date-time":"2025-03-15T00:00:00Z","timestamp":1741996800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,3,15]],"date-time":"2025-03-15T00:00:00Z","timestamp":1741996800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100013058","name":"JiangSu Provincial Key Research and Development Program","doi-asserted-by":"crossref","award":["BE2020084-1"],"award-info":[{"award-number":["BE2020084-1"]}],"id":[{"id":"10.13039\/501100013058","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Intell Rev"],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Recent years have witnessed the great achievement of the AI-driven intelligent games, such as AlphaStar defeating the human experts, and numerous intelligent games have come into the public view. Essentially, deep reinforcement learning (DRL), especially multiple-agent DRL (MADRL) has empowered a variety of artificial intelligence fields, including intelligent games. However, there is lack of systematical review on their correlations. This article provides a holistic picture on smoothly connecting intelligent games with MADRL from two perspectives: theoretical game concepts for MADRL, and MADRL for intelligent games. From the first perspective, information structure and game environmental features for MADRL algorithms are summarized; and from the second viewpoint, the challenges in intelligent games are investigated, and the existing MADRL solutions are correspondingly explored. Furthermore, the state-of-the-art (SOTA) MADRL algorithms for intelligent games are systematically categorized, especially from the perspective of credit assignment. Moreover, a comprehensively review on notorious benchmarks are conducted to facilitate the design and test of MADRL based intelligent games. Besides, a general procedure of MADRL simulations is offered. Finally, the key challenges in integrating intelligent games with MADRL, and potential future research directions are highlighted. This survey hopes to provide a thoughtful insight of developing intelligent games with the assistance of MADRL solutions and algorithms.<\/jats:p>","DOI":"10.1007\/s10462-025-11166-1","type":"journal-article","created":{"date-parts":[[2025,3,15]],"date-time":"2025-03-15T05:11:04Z","timestamp":1742015464000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Intelligent games meeting with multi-agent deep reinforcement learning: a comprehensive review"],"prefix":"10.1007","volume":"58","author":[{"given":"Yiqin","family":"Wang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yufeng","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Feng","family":"Tian","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianhua","family":"Ma","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qun","family":"Jin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,3,15]]},"reference":[{"key":"11166_CR1","unstructured":"Agarwal A, Kumar S, Sycara K et al (2020) Learning transferable cooperative behavior in multi-agent teams. In: Proceedings of the 19th international conference on autonomous agents and multiagent systems. International Foundation for Autonomous Agents and Multiagent Systems, AAMAS \u201920, pp 1741\u20131743"},{"key":"11166_CR2","doi-asserted-by":"crossref","unstructured":"Amos-Binks A, Weber BS (2023) Risk management: anticipating and reacting in starcraft. In: Proceedings of the AAAI conference on artificial intelligence and interactive digital entertainment, pp 13\u201322","DOI":"10.1609\/aiide.v19i1.27497"},{"key":"11166_CR3","unstructured":"Bai Y, Jin C (2020) Provable self-play algorithms for competitive reinforcement learning. In: International conference on machine learning, PMLR, pp 551\u2013560"},{"key":"11166_CR4","first-page":"25799","volume":"34","author":"Y Bai","year":"2021","unstructured":"Bai Y, Jin C, Wang H et al (2021) Sample-efficient learning of Stackelberg equilibria in general-sum games. Adv Neural Inf Process Syst 34:25799\u201325811","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR5","unstructured":"Balduzzi D, Garnelo M, Bachrach Y et al (2019) Open-ended learning in symmetric zero-sum games. In: International conference on machine learning, PMLR, pp 434\u2013443"},{"issue":"2","key":"11166_CR6","doi-asserted-by":"publisher","first-page":"134","DOI":"10.1109\/TG.2022.3214154","volume":"15","author":"J Barambones","year":"2022","unstructured":"Barambones J, Cano-Benito J, S, \u0301anchez-Rivero I et al (2022) Multiagent systems on virtual games: a systematic mapping study. IEEE Trans Games 15(2):134\u2013147","journal-title":"IEEE Trans Games"},{"key":"11166_CR177","unstructured":"Bellemare MG, Dabney W, Munos R (2017) A distributional perspective on reinforcement learning. In: International conference on machine learning, PMLR, pp 449\u2013458"},{"key":"11166_CR7","unstructured":"Berner C, Brockman G, Chan B et al (2019) Dota 2 with large scale deep reinforcement learning. arXiv preprint arXiv:191206680"},{"issue":"217","key":"11166_CR8","first-page":"1","volume":"25","author":"M Bettini","year":"2024","unstructured":"Bettini M, Prorok A, Moens V (2024) Benchmarl: benchmarking multi-agent reinforcement learning. J Mach Learn Res 25(217):1\u201310","journal-title":"J Mach Learn Res"},{"key":"11166_CR9","unstructured":"Bian Y, Rong Y, Xu T et al (2022) Energy-based learning for cooperative games, with applications to valuation problems in machine learning. In: International conference on learning representation"},{"key":"11166_CR10","unstructured":"Bou A, Bettini M, Dittert S et al (2024) TorchRL: a data-driven decision-making library for pytorch. In: The Twelfth international conference on learning representations"},{"issue":"11","key":"11166_CR11","doi-asserted-by":"publisher","first-page":"4948","DOI":"10.3390\/app11114948","volume":"11","author":"L Canese","year":"2021","unstructured":"Canese L, Cardarilli GC, Di Nunzio L et al (2021) Multi-agent reinforcement learning: a review of challenges and applications. Appl Sci 11(11):4948","journal-title":"Appl Sci"},{"key":"11166_CR12","doi-asserted-by":"crossref","unstructured":"Carr S, Jansen N, Junges S et al (2023) Safe reinforcement learning via shielding under partial observability. In: Proceedings of the AAAI conference on artificial intelligence, pp 14748\u201314756","DOI":"10.1609\/aaai.v37i12.26723"},{"key":"11166_CR13","doi-asserted-by":"crossref","unstructured":"Chan L, Hogaboam L, Cao R (2022) Artificial intelligence in video games and esports. Applied artificial intelligence in business: concepts and cases. Springer, pp 335\u2013352","DOI":"10.1007\/978-3-031-05740-3_22"},{"issue":"6","key":"11166_CR14","doi-asserted-by":"publisher","first-page":"590","DOI":"10.1038\/s42256-023-00657-x","volume":"5","author":"H Chen","year":"2023","unstructured":"Chen H, Covert IC, Lundberg SM et al (2023) Algorithms to estimate Shapley value feature attributions. Nat Mach Intell 5(6):590\u2013601","journal-title":"Nat Mach Intell"},{"key":"11166_CR15","unstructured":"Christianos F, Papoudakis G, Albrecht SV (2023) Pareto actor-critic for equilibrium selection in multi-agent reinforcement learning. arXiv preprint arXiv:2209.14344"},{"key":"11166_CR16","unstructured":"Chu T, Chinchali S, Katti S (2020) Multi-agent reinforcement learning for networked system control. In: International conference on learning representation"},{"key":"11166_CR17","unstructured":"Cui K, Tahir A, Ekinci G et al (2022) A survey on large-population systems and scalable multi-agent reinforcement learning. arXiv preprint arXiv:220903859"},{"key":"11166_CR18","unstructured":"Das A, Gervet T, Romoff J et al (2019) Tarmac: targeted multi-agent communication. In: International conference on machine learning, PMLR, pp 1538\u20131546"},{"issue":"23","key":"11166_CR19","doi-asserted-by":"publisher","first-page":"16893","DOI":"10.1007\/s00521-023-08423-1","volume":"35","author":"R Dazeley","year":"2023","unstructured":"Dazeley R, Vamplew P, Cruz F (2023) Explainable reinforcement learning for broad-xai: a conceptual framework and survey. Neural Comput Appl 35(23):16893\u201316916","journal-title":"Neural Comput Appl"},{"key":"11166_CR20","unstructured":"De Witt CS, Gupta T, Makoviichuk D et al (2020) Is independent learning all you need in the Starcraft multi-agent challenge? ArXiv Preprint ArXiv:201109533"},{"key":"11166_CR22","first-page":"22069","volume":"33","author":"Z Ding","year":"2020","unstructured":"Ding Z, Huang T, Lu Z (2020) Learning individually inferred communication for multi-agent Cooperation. Adv Neural Inf Process Syst 33:22069\u201322079","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR21","unstructured":"Ding D, Wei CY, Zhang K et al (2022) Independent policy gradient for large-scale markov potential games: sharper rates, function approximation, and game-agnostic convergence. In: International Conference on Machine Learning, PMLR, pp 5166\u20135220"},{"key":"11166_CR24","doi-asserted-by":"crossref","unstructured":"do Nascimento Silva V, Chaimowicz L (2015) On the development of intelligent agents for moba games. In: 2015 14th Brazilian symposium on computer games and digital entertainment (SBGames), IEEE, pp 142\u2013151","DOI":"10.1109\/SBGames.2015.33"},{"issue":"9","key":"11166_CR23","first-page":"1","volume":"55","author":"R Dwivedi","year":"2023","unstructured":"Dwivedi R, Dave D, Naik H et al (2023) Explainable Ai (xai): core ideas, techniques, and solutions. ACM-CSUR 55(9):1\u201333","journal-title":"ACM-CSUR"},{"key":"11166_CR25","unstructured":"Ellis B, Cook J, Moalla S et al (2024) Smacv2: an improved benchmark for cooperative multi-agent reinforcement learning. Adv Neural Inf Process Syst 36:37567\u201337593"},{"key":"11166_CR26","unstructured":"ElSayed-Aly I, Bharadwaj S, Amato C et al (2021) Safe multi-agent reinforcement learning via shielding. In: Proceedings of the 20th international conference on autonomous agents and multiagent systems"},{"key":"11166_CR27","doi-asserted-by":"crossref","unstructured":"Ferdous R, Kifetew F, Prandi D et al (2022) Towards agent-based testing of 3d games using reinforcement learning. In: Proceedings of the 37th IEEE\/ACM international conference on automated software engineering, pp 1\u20138","DOI":"10.1145\/3551349.3560507"},{"key":"11166_CR28","unstructured":"Foerster J, Assael IA, De Freitas N et al (2016) Learning to communicate with deep multi-agent reinforcement learning. Adv Neural Inf Process Syst 29"},{"key":"11166_CR29","doi-asserted-by":"crossref","unstructured":"Foerster J, Farquhar G, Afouras T et al (2018) Counterfactual multi-agent policy gradients. In: Proceedings of the AAAI conference on artificial intelligence","DOI":"10.1609\/aaai.v32i1.11794"},{"key":"11166_CR30","unstructured":"Fu H, Liu W, Wu S et al (2021) Actor-critic policy optimization in a large-scale imperfect-information game. In: International conference on learning representations"},{"issue":"2","key":"11166_CR31","doi-asserted-by":"publisher","first-page":"3033","DOI":"10.1016\/j.ifacol.2023.10.1431","volume":"56","author":"Q Fu","year":"2023","unstructured":"Fu Q, Ai X, Yi J et al (2023) Learning heterogeneous agent cooperation via multiagent league training. IFAC-PapersOnLine 56(2):3033\u20133040","journal-title":"IFAC-PapersOnLine"},{"key":"11166_CR178","unstructured":"Fujimoto S, Hoof H, Meger D (2018) Addressing function approximation error in actor-critic methods. In: International conference on machine learning, PMLR, pp 1587\u20131596"},{"key":"11166_CR32","doi-asserted-by":"crossref","unstructured":"Gero KI, Ashktorab Z, Dugan C et al (2020) Mental models of ai agents in a cooperative game setting. In: Proceedings of the 2020 chi conference on human factors in computing systems, pp 1\u201312","DOI":"10.1145\/3313831.3376316"},{"key":"11166_CR33","unstructured":"Ghorbani A, Zou J (2019) Data shapley: equitable valuation of data for machine learning. In: International conference on machine learning, PMLR, pp 2242\u20132251"},{"key":"11166_CR34","unstructured":"Gogineni K, Wei P (2023) Scalability bottlenecks in multi-agent reinforcement learning systems. In: FastPath 2023: International workshop on performance analysis of machine learning systems"},{"key":"11166_CR35","first-page":"5510","volume":"35","author":"R Gorsane","year":"2022","unstructured":"Gorsane R, Mahjoub O, de Kock RJ et al (2022) Towards a standardised performance evaluation protocol for cooperative marl. Adv Neural Inf Process Syst 35:5510\u20135521","journal-title":"Adv Neural Inf Process Syst"},{"issue":"2","key":"11166_CR36","doi-asserted-by":"publisher","first-page":"895","DOI":"10.1007\/s10462-021-09996-w","volume":"55","author":"S Gronauer","year":"2022","unstructured":"Gronauer S, Diepold K (2022) Multi-agent deep reinforcement learning: a survey. Artif Intell Rev 55(2):895\u2013943","journal-title":"Artif Intell Rev"},{"key":"11166_CR37","doi-asserted-by":"publisher","first-page":"103905","DOI":"10.1016\/j.artint.2023.103905","volume":"319","author":"S Gu","year":"2023","unstructured":"Gu S, Kuba JG, Chen Y et al (2023) Safe multi-agent reinforcement learning for multi-robot control. Artif Intell 319:103905","journal-title":"Artif Intell"},{"issue":"12","key":"11166_CR38","doi-asserted-by":"publisher","first-page":"11216","DOI":"10.1109\/TPAMI.2024.3457538","volume":"46","author":"S Gu","year":"2024","unstructured":"Gu S, Yang L, Du Y et al (2024) A review of safe reinforcement learning: methods, theories, and applications. IEEE Trans Pattern Anal Mach Intell 46(12):11216\u201311235","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11166_CR39","doi-asserted-by":"crossref","unstructured":"Guo X, Shi D, Fan W (2023) Scalable communication for multi-agent reinforcement learning via transformer-based email mechanism. In: Proceedings of the thirty-second international joint conference on artificial intelligence, IJCAI \u201923","DOI":"10.24963\/ijcai.2023\/15"},{"key":"11166_CR40","doi-asserted-by":"crossref","unstructured":"Gupta N, Srinivasaraghavan G, Mohalik S et al (2023) Hammer: Multi-level coordination of reinforcement learning agents via learned messaging. Neural Comput Appl 1\u201316","DOI":"10.1007\/s00521-023-09096-6"},{"issue":"7","key":"11166_CR41","doi-asserted-by":"publisher","first-page":"787","DOI":"10.1038\/s42256-024-00861-3","volume":"6","author":"L Han","year":"2024","unstructured":"Han L, Zhu Q, Sheng J et al (2024) Lifelike agility and play in quadrupedal robots using reinforcement learning and generative pre-trained models. Nat Mach Intell 6(7):787\u2013798","journal-title":"Nat Mach Intell"},{"key":"11166_CR42","unstructured":"Hansen EA, Bernstein DS, Zilberstein S (2004) Dynamic programming for partially observable stochastic games. In: AAAI, pp 709\u2013715"},{"key":"11166_CR43","doi-asserted-by":"crossref","unstructured":"Hao J, Yang T, Tang H et al (2023) Exploration in deep reinforcement learning: from single-agent to multiagent domain. IEEE Trans Neural Networks Learn Syst 35(7):8762\u20138782","DOI":"10.1109\/TNNLS.2023.3236361"},{"key":"11166_CR179","unstructured":"Hausknecht M, Stone P (2015) Deep recurrent q-learning for partially observable mdps. In: 2015 AAAI fall symposium series"},{"key":"11166_CR44","unstructured":"Heinrich J, Lanctot M, Silver D (2015) Fictitious self-play in extensive-form games. In: International conference on machine learning, PMLR, pp 805\u2013813"},{"key":"11166_CR45","doi-asserted-by":"crossref","unstructured":"Hernandez D, Denamgana \u0308\u0131 K, Gao Y et al (2019) A generalized framework for self-play training. In: 2019 IEEE Conference on Games (CoG), IEEE, pp 1\u20138","DOI":"10.1109\/CIG.2019.8848006"},{"issue":"2","key":"11166_CR46","doi-asserted-by":"publisher","first-page":"221","DOI":"10.1109\/TG.2021.3058898","volume":"14","author":"D Hernandez","year":"2021","unstructured":"Hernandez D, Denamganai K, Devlin S et al (2021) A comparison of self-play algorithms under a generalized framework. IEEE Trans Games 14(2):221\u2013231","journal-title":"IEEE Trans Games"},{"key":"11166_CR47","unstructured":"Hernandez-Leal P, Kaisers M, Baarslag T et al (2017) A survey of learning in multiagent environments: dealing with non-stationarity. arXiv preprint arXiv:170709183"},{"issue":"1","key":"11166_CR48","doi-asserted-by":"publisher","first-page":"59","DOI":"10.1109\/MCI.2021.3129959","volume":"17","author":"A Heuillet","year":"2022","unstructured":"Heuillet A, Couthouis F, \u0301\u0131az-Rodr \u0301\u0131guez D N (2022) Collective explainable Ai: explaining cooperative strategies and agent contribution in multiagent reinforcement learning with Shapley values. IEEE Comput Intell Mag 17(1):59\u201371","journal-title":"IEEE Comput Intell Mag"},{"issue":"5","key":"11166_CR49","first-page":"1","volume":"56","author":"T Hickling","year":"2023","unstructured":"Hickling T, Zenati A, Aouf N et al (2023) Explainability in deep reinforcement learning: a review into current methods and applications. ACM-CSUR 56(5):1\u201335","journal-title":"ACM-CSUR"},{"key":"11166_CR50","first-page":"32438","volume":"35","author":"Y Hong","year":"2022","unstructured":"Hong Y, Jin Y, Tang Y (2022) Rethinking individual global max in cooperative multi-agent reinforcement learning. Adv Neural Inf Process Syst 35:32438\u201332449","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR52","unstructured":"Hu J, Hu S, Liao S (2021) Policy regularization via noisy advantage values for cooperative multi-agent actor-critic methods. arXiv preprint arXiv:210614334"},{"key":"11166_CR51","unstructured":"Hu B, Zhao C, Zhang P et al (2023a) Enabling intelligent interactions between an agent and an llm: a reinforcement learning approach. arXiv preprint arXiv:230603604"},{"key":"11166_CR53","unstructured":"Hu J, Wang S, Jiang S et al (2023b) Rethinking the implementation tricks and monotonicity constraint in cooperative multi-agent reinforcement learning. In: The Second Blogpost Track at ICLR 2023"},{"issue":"315","key":"11166_CR54","first-page":"1","volume":"24","author":"S Hu","year":"2023","unstructured":"Hu S, Zhong Y, Gao M et al (2023c) Marllib: a scalable and efficient multi-agent reinforcement learning library. J Mach Learn Res 24(315):1\u201323","journal-title":"J Mach Learn Res"},{"issue":"2","key":"11166_CR55","doi-asserted-by":"publisher","first-page":"6161","DOI":"10.1007\/s11042-023-15361-6","volume":"83","author":"Z Hu","year":"2024","unstructured":"Hu Z, Liu H, Xiong Y et al (2024) Promoting human-ai interaction makes a better adoption of deep reinforcement learning: a real-world application in game industry. Multimed Tools Appl 83(2):6161\u20136182","journal-title":"Multimed Tools Appl"},{"issue":"15","key":"11166_CR56","doi-asserted-by":"publisher","first-page":"10893","DOI":"10.1016\/j.jfranklin.2023.08.032","volume":"360","author":"X Huang","year":"2023","unstructured":"Huang X (2023) Starcraft adversary-agent challenge for pursuit\u2013evasion game. J Franklin Inst 360(15):10893\u201310916","journal-title":"J Franklin Inst"},{"issue":"6443","key":"11166_CR57","doi-asserted-by":"publisher","first-page":"859","DOI":"10.1126\/science.aau6249","volume":"364","author":"M Jaderberg","year":"2019","unstructured":"Jaderberg M, Czarnecki WM, Dunning I et al (2019) Human-level performance in 3d multiplayer games with population-based reinforcement learning. Science 364(6443):859\u2013865","journal-title":"Science"},{"key":"11166_CR59","unstructured":"Ji J, Zhang B, Zhou J et al (2023) Safety gymnasium: a unified safe reinforcement learning benchmark. Adv Neural Inf Process Syst 36:18964\u201318993"},{"key":"11166_CR60","unstructured":"Jiang J, Lu Z (2018) Learning attentional communication for multi-agent cooperation. Adv Neural Inf Process Syst 31"},{"key":"11166_CR58","unstructured":"Jiang J, Dun C, Huang T et al (2020) Graph convolutional reinforcement learning. In: International conference on learning representation"},{"issue":"23","key":"11166_CR61","doi-asserted-by":"publisher","first-page":"29205","DOI":"10.1007\/s10489-023-04866-0","volume":"53","author":"K Jiang","year":"2023","unstructured":"Jiang K, Liu W, Wang Y et al (2023) Credit assignment in heterogeneous multi-agent reinforcement learning for fully cooperative tasks. Appl Intell 53(23):29205\u201329222","journal-title":"Appl Intell"},{"key":"11166_CR62","unstructured":"Kim D, Moon S, Hostallero D et al (2019) Learning to schedule communication in multi-agent reinforcement learning. In: International conference on learning representation"},{"key":"11166_CR63","unstructured":"Kim W, Park J, Sung Y (2020) Communication in multi-agent reinforcement learning: Intention sharing. In: International conference on learning representations"},{"key":"11166_CR64","unstructured":"Kuba JG, Chen R, Wen M et al (2022) Trust region policy optimisation in multi-agent reinforcement learning. In: International conference on learning representations"},{"key":"11166_CR65","doi-asserted-by":"publisher","first-page":"100862","DOI":"10.1016\/j.entcom.2024.100862","volume":"52","author":"K Kumar","year":"2025","unstructured":"Kumar K, Veena N, Aravind T et al (2025) Game-changing intelligence: unveiling the societal impact of artificial intelligence in game software. Entertain Comput 52:100862","journal-title":"Entertain Comput"},{"key":"11166_CR66","doi-asserted-by":"crossref","unstructured":"Kurach K, Raichuk A, Sta \u0301nczyk P et al (2020) Google research football: A novel reinforcement learning environment. In: Proceedings of the AAAI conference on artificial intelligence, pp 4501\u20134510","DOI":"10.1609\/aaai.v34i04.5878"},{"key":"11166_CR67","unstructured":"Kwon M, Xie SM, Bullard K et al (2023) Reward design with language models. In:The eleventh international conference on learning representations"},{"key":"11166_CR68","unstructured":"Lanctot M, Lockhart E, Lespiau JB et al (2019) Openspiel: a framework for reinforcement learning in games. arXiv preprint arXiv:190809453"},{"key":"11166_CR69","unstructured":"Leonardos S, Overman W, Panageas I et al (2022) Global convergence of multi-agent policy gradient in markov potential games. In: ICLR 2022 Workshop on gamification and multiagent solutions"},{"key":"11166_CR70","unstructured":"Li J, Koyamada S, Ye Q et al (2020) Suphx: Mastering mahjong with deep reinforcement learning. arXiv preprint arXiv:200313590"},{"key":"11166_CR71","doi-asserted-by":"crossref","unstructured":"Li J, Kuang K, Wang B et al (2021) Shapley counterfactual credits for multi-agent reinforcement learning. In: Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining, pp 934\u2013942","DOI":"10.1145\/3447548.3467420"},{"key":"11166_CR73","doi-asserted-by":"publisher","first-page":"296","DOI":"10.1016\/j.arcontrol.2022.03.003","volume":"53","author":"T Li","year":"2022","unstructured":"Li T, Zhao Y, Zhu Q (2022) The role of information structures in game-theoretic multi-agent learning. Annu Rev Control 53:296\u2013314","journal-title":"Annu Rev Control"},{"issue":"1","key":"11166_CR72","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1109\/TG.2023.3259724","volume":"16","author":"S Li","year":"2023","unstructured":"Li S, Xu J, Dong H et al (2023a) The fittest wins: a multistage framework achieving new Sota in Vizdoom competition. IEEE Trans Games 16(1):225\u2013234","journal-title":"IEEE Trans Games"},{"key":"11166_CR74","unstructured":"Li Y, Xiong K, Zhang Y et al (2023) Jiangjun: mastering Xiangqi by tackling non-transitivity in two-player zero-sum games. arXiv preprint arXiv:2308.04719"},{"key":"11166_CR180","unstructured":"Lillicrap TP, Hunt JJ, Pritzel A et al (2015) Continuous control with deep reinforcement learning. arXiv preprint arXiv:150902971"},{"key":"11166_CR76","first-page":"15230","volume":"34","author":"T Lin","year":"2021","unstructured":"Lin T, Huh J, Stauffer C et al (2021) Learning to ground multi-agent communication with autoencoders. Adv Neural Inf Process Syst 34:15230\u201315242","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR75","unstructured":"Lin F, Huang S, Pearce T et al (2023) Tizero: Mastering multi-agent football with curriculum learning and self-play. In: Proceedings of the 2023 International conference on autonomous agents and multiagent systems. International Foundation for Autonomous Agents and Multiagent Systems, AAMAS \u201923, p 67\u201376"},{"key":"11166_CR79","doi-asserted-by":"crossref","unstructured":"Liu Y, Wang W, Hu Y et al (2020) Multi-agent game abstraction via graph attention neural network. In: Proceedings of the AAAI conference on artificial intelligence, pp 7211\u20137218","DOI":"10.1609\/aaai.v34i05.6211"},{"key":"11166_CR77","doi-asserted-by":"crossref","unstructured":"Liu C, Geng N, Aggarwal V et al (2021) Cmix: Deep multi-agent reinforcement learning with peak and average constraints. In: Machine learning and knowledge discovery in databases. Research Track: European Conference, ECML PKDD 2021, pp 157\u2013173","DOI":"10.1007\/978-3-030-86486-6_10"},{"key":"11166_CR78","first-page":"18296","volume":"35","author":"Q Liu","year":"2022","unstructured":"Liu Q, Szepesv \u0301ari C, Jin C (2022) Sample-efficient reinforcement learning of partially observable Markov games. Adv Neural Inf Process Syst 35:18296\u201318308","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR80","doi-asserted-by":"crossref","unstructured":"Liu Z, Wan L, Sui X et al (2023) Deep hierarchical communication graph in multi-agent reinforcement learning. In: IJCAI, pp 208\u2013216","DOI":"10.24963\/ijcai.2023\/24"},{"key":"11166_CR81","unstructured":"Long Q, Zhou Z, Gupta A et al (2020) Evolutionary population curriculum for scaling multi-agent reinforcement learning. In: International conference on learning representation"},{"key":"11166_CR82","unstructured":"Lowe R, Wu YI, Tamar A et al (2017) Multi-agent actor-critic for mixed cooperative-competitive environments. Adv Neural Inf Process Syst 30"},{"key":"11166_CR83","unstructured":"Maddila P, Eric C, Chabrier P et al (2024) APE: an anti-poaching multi-agent reinforcement learning benchmark. In: Seventeenth European workshop on reinforcement learning"},{"key":"11166_CR84","unstructured":"Mahajan A, Rashid T, Samvelyan M et al (2019) Maven: multi-agent variational exploration. Adv Neural Inf Process Syst 32"},{"key":"11166_CR85","first-page":"36243","volume":"35","author":"W Mao","year":"2022","unstructured":"Mao W, Qiu H, Wang C et al (2022) A mean-field game approach to cloud resource management with function approximation. Adv Neural Inf Process Syst 35:36243\u201336258","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR181","unstructured":"Mnih V, Badia AP, Mirza M et al (2016) Asynchronous methods for deep reinforcement learning. In: International conference on machine learning, PMLR, pp 1928\u20131937"},{"issue":"1","key":"11166_CR182","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1561\/2200000086","volume":"16","author":"TM Moerland","year":"2023","unstructured":"Moerland TM, Broekens J, Plaat A et al (2023) Model-based reinforcement learning: a survey. Found Trends Mach Learn 16(1):1\u2013118","journal-title":"Found Trends Mach Learn"},{"issue":"6337","key":"11166_CR86","doi-asserted-by":"publisher","first-page":"508","DOI":"10.1126\/science.aam6960","volume":"356","author":"M Morav\u010d\u00edk","year":"2017","unstructured":"Morav\u010d\u00edk M, Schmid M, Burch N et al (2017) Deepstack: expert-level artificial intelligence in heads-up no-limit poker. Science 356(6337):508\u2013513. https:\/\/doi.org\/10.1126\/science.aam6960","journal-title":"Science"},{"key":"11166_CR87","first-page":"41","volume":"1","author":"J Mycielski","year":"1992","unstructured":"Mycielski J (1992) Games with perfect information. Handb Game Theory Econ Appl 1:41\u201370","journal-title":"Handb Game Theory Econ Appl"},{"key":"11166_CR183","unstructured":"Nachum O, Norouzi M, Xu K et al (2017) Bridging the gap between value and policy based reinforcement learning. Adv Neural Inf Process Syst 30"},{"key":"11166_CR184","doi-asserted-by":"crossref","unstructured":"Nagabandi A, Kahn G, Fearing RS et al (2018) Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning. In: 2018 IEEE international conference on robotics and automation (ICRA), IEEE, pp 7559\u20137566","DOI":"10.1109\/ICRA.2018.8463189"},{"key":"11166_CR88","unstructured":"Nayak S, Choi K, Ding W et al (2023) Scalable multi-agent reinforcement learning through intelligent information aggregation. In: International conference on machine learning, PMLR, pp 25817\u201325833"},{"issue":"9","key":"11166_CR89","doi-asserted-by":"publisher","first-page":"3826","DOI":"10.1109\/TCYB.2020.2977374","volume":"50","author":"TT Nguyen","year":"2020","unstructured":"Nguyen TT, Nguyen ND, Nahavandi S (2020) Deep reinforcement learning for multiagent systems: a review of challenges, solutions, and applications. IEEE Trans Cybern 50(9):3826\u20133839","journal-title":"IEEE Trans Cybern"},{"key":"11166_CR90","unstructured":"Niu Y, Paleja RR, Gombolay MC (2021) Multi-agent graph-attention communication and teaming. In: AAMAS"},{"issue":"2","key":"11166_CR91","doi-asserted-by":"publisher","first-page":"212","DOI":"10.1109\/TG.2021.3049539","volume":"14","author":"I Oh","year":"2021","unstructured":"Oh I, Rho S, Moon S et al (2021) Creating pro-level Ai for a real-time fighting game using deep reinforcement learning. IEEE Trans Games 14(2):212\u2013220","journal-title":"IEEE Trans Games"},{"issue":"11","key":"11166_CR92","doi-asserted-by":"publisher","first-page":"13677","DOI":"10.1007\/s10489-022-04105-y","volume":"53","author":"A Oroojlooy","year":"2023","unstructured":"Oroojlooy A, Hajinezhad D (2023) A review of cooperative multi-agent deep reinforcement learning. Appl Intell 53(11):13677\u201313722","journal-title":"Appl Intell"},{"key":"11166_CR93","first-page":"27862","volume":"35","author":"X Pan","year":"2022","unstructured":"Pan X, Liu M, Zhong F et al (2022) Mate: benchmarking multi-agent reinforcement learning in distributed target coverage control. Adv Neural Inf Process Syst 35:27862\u201327879","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR94","unstructured":"Papoudakis G, Christianos F, Rahman A et al (2019) Dealing with non-stationarity in multi-agent deep reinforcement learning. arXiv preprint arXiv:190604737"},{"key":"11166_CR95","unstructured":"Papoudakis G, Christianos F, Schafer L et al (2021) Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks. In: Thirty-fifth conferenceon neural information processing systems datasets and benchmarks track (Round1)"},{"key":"11166_CR96","unstructured":"Peng P, Wen Y, Yang Y et al (2017) Multiagent bidirectionally-coordinated nets: emergence of human-level coordination in learning to play starcraft combat games. arXiv preprint arXiv:170310069"},{"issue":"6623","key":"11166_CR97","doi-asserted-by":"publisher","first-page":"990","DOI":"10.1126\/science.add4679","volume":"378","author":"J Perolat","year":"2022","unstructured":"Perolat J, De Vylder B, Hennes D et al (2022) Mastering the game of stratego with model-free multiagent reinforcement learning. Science 378(6623):990\u2013996","journal-title":"Science"},{"key":"11166_CR98","doi-asserted-by":"crossref","unstructured":"Qiu Y, Jin Y, Yu L Safe multi-agent reinforcement learning via dynamic shielding. In: 2024 IEEE Conference on Artificial, Intelligence (CAI), IEEE, pp 1254\u20131257","DOI":"10.1109\/CAI59869.2024.00222"},{"key":"11166_CR99","first-page":"10088","volume":"33","author":"M Rangwala","year":"2020","unstructured":"Rangwala M, Williams R (2020) Learning multi-agent communication through structured attentive reasoning. Adv Neural Inf Process Syst 33:10088\u201310098","journal-title":"Adv Neural Inf Process Syst"},{"issue":"178","key":"11166_CR101","first-page":"1","volume":"21","author":"T Rashid","year":"2020","unstructured":"Rashid T, Samvelyan M, De Witt CS et al (2020b) Monotonic value function factorisation for deep multi-agent reinforcement learning. J Mach Learn Res 21(178):1\u201351","journal-title":"J Mach Learn Res"},{"key":"11166_CR100","first-page":"10199","volume":"33","author":"T Rashid","year":"2020","unstructured":"Rashid T, Farquhar G, Peng B et al (2020a) Weighted Qmix: expanding monotonic value function factorisation for deep multi-agent reinforcement learning. Adv Neural Inf Process Syst 33:10199\u201310210","journal-title":"Adv Neural Inf Process Syst"},{"issue":"1","key":"11166_CR102","doi-asserted-by":"publisher","first-page":"164","DOI":"10.1016\/S0899-8256(05)80020-X","volume":"8","author":"AE Roth","year":"1995","unstructured":"Roth AE, Erev I (1995) Learning in extensive-form games: experimental data and simple dynamic models in the intermediate term. Games Econ Behav 8(1):164\u2013212","journal-title":"Games Econ Behav"},{"key":"11166_CR103","doi-asserted-by":"publisher","unstructured":"Rozemberczki B, Watson L, Bayer P et al (2022) The shapley value in machine learning. In: Proceedings of the 31st international joint conference on artifical intelligence, IJCAI-ECAI 2022. International joint conferences on Artificial Intelligence Organization, pp 5572\u20135579. https:\/\/doi.org\/10.24963\/ijcai.2022\/778","DOI":"10.24963\/ijcai.2022\/778"},{"issue":"7839","key":"11166_CR104","doi-asserted-by":"publisher","first-page":"604","DOI":"10.1038\/s41586-020-03051-4","volume":"588","author":"J Schrittwieser","year":"2020","unstructured":"Schrittwieser J, Antonoglou I, Hubert T et al (2020) Mastering Atari, go, chess and Shogi by planning with a learned model. Nature 588(7839):604\u2013609","journal-title":"Nature"},{"key":"11166_CR186","unstructured":"Schulman J, Levine S, Abbeel P et al (2015) Trust region policy optimization. In: International conference on machine learning, PMLR, pp 1889\u20131897"},{"key":"11166_CR185","unstructured":"Schulman J, Chen X, Abbeel P (2017) Equivalence between policy gradients and soft q-learning. arXiv preprint arXiv:170406440"},{"key":"11166_CR105","doi-asserted-by":"crossref","unstructured":"Shapley LS (1953) Stochastic games. Proc Natl Acad Sci 39(10):1095\u20131100","DOI":"10.1073\/pnas.39.10.1095"},{"key":"11166_CR106","doi-asserted-by":"crossref","unstructured":"Shapley LS et al (1953) A value for n-person games","DOI":"10.1515\/9781400881970-018"},{"issue":"7587","key":"11166_CR107","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver D, Huang A, Maddison CJ et al (2016) Mastering the game of go with deep neural networks and tree search. Nature 529(7587):484\u2013489","journal-title":"Nature"},{"issue":"7676","key":"11166_CR109","doi-asserted-by":"publisher","first-page":"354","DOI":"10.1038\/nature24270","volume":"550","author":"D Silver","year":"2017","unstructured":"Silver D, Schrittwieser J, Simonyan K et al (2017) Mastering the game of go without human knowledge. Nature 550(7676):354\u2013359","journal-title":"Nature"},{"issue":"6419","key":"11166_CR108","doi-asserted-by":"publisher","first-page":"1140","DOI":"10.1126\/science.aar6404","volume":"362","author":"D Silver","year":"2018","unstructured":"Silver D, Hubert T, Schrittwieser J et al (2018) A general reinforcement learning algorithm that masters chess, Shogi, and go through self-play. Science 362(6419):1140\u20131144","journal-title":"Science"},{"key":"11166_CR110","unstructured":"Skrynnik A, Andreychuk A, Borzilov A et al (2024) Pogema: a benchmark platform for cooperative multi-agent navigation. arXiv preprint arXiv:240714931"},{"key":"11166_CR111","unstructured":"Son K, Kim D, Kang WJ et al (2019) Qtran: learning to factorize with transformation for cooperative multi-agent reinforcement learning. In: International conference on machine learning, PMLR, pp 5887\u20135896"},{"issue":"4","key":"11166_CR112","doi-asserted-by":"publisher","first-page":"2443","DOI":"10.3390\/app13042443","volume":"13","author":"K Souchleris","year":"2023","unstructured":"Souchleris K, Sidiropoulos GK, Papakostas GA (2023) Reinforcement learning in game industry\u2014review, prospects and challenges. Appl Sci 13(4):2443","journal-title":"Appl Sci"},{"issue":"1","key":"11166_CR113","first-page":"56","volume":"13","author":"J Subramanian","year":"2023","unstructured":"Subramanian J, Sinha A, Mahajan A (2023) Robustness and sample complexity of model-based marl for general-sum Markov games. Dyn Games Appl 13(1):56\u201388","journal-title":"Dyn Games Appl"},{"key":"11166_CR114","unstructured":"Sukhbaatar S, Fergus R et al (2016) Learning multiagent communication with backpropagation. Adv Neural Inf Process Syst 29"},{"key":"11166_CR115","unstructured":"Sun C, Huang S, Pompili D (2024) Llm-based multi-agent reinforcement learning: current and future directions. arXiv preprint arXiv:240511106"},{"key":"11166_CR116","unstructured":"Sundararajan M, Taly A, Yan Q (2017) Axiomatic attribution for deep networks. In: International conference on machine learning, PMLR, pp 3319\u20133328"},{"key":"11166_CR117","unstructured":"Sunehag P, Lever G, Gruslys A et al (2018) Value-decomposition networks for cooperative multi-agent learning based on team reward. International Foundation for Autonomous Agents and Multiagent Systems, AAMAS \u201918, pp 2085\u20132087"},{"issue":"4","key":"11166_CR187","doi-asserted-by":"publisher","first-page":"160","DOI":"10.1145\/122344.122377","volume":"2","author":"RS Sutton","year":"1991","unstructured":"Sutton RS (1991) Dyna, an integrated architecture for learning, planning, and reacting. ACM Sigart Bull 2(4):160\u2013163","journal-title":"ACM Sigart Bull"},{"issue":"4","key":"11166_CR118","doi-asserted-by":"publisher","first-page":"e0172395","DOI":"10.1371\/journal.pone.0172395","volume":"12","author":"A Tampuu","year":"2017","unstructured":"Tampuu A, Matiisen T, Kodelja D et al (2017) Multiagent Cooperation and competition with deep reinforcement learning. PLoS ONE 12(4):e0172395","journal-title":"PLoS ONE"},{"issue":"4","key":"11166_CR119","doi-asserted-by":"publisher","first-page":"835","DOI":"10.1109\/TAC.2013.2289711","volume":"59","author":"H Tembine","year":"2013","unstructured":"Tembine H, Zhu Q, Ba\u0327sar T (2013) Risk-sensitive mean-field games. IEEE Trans Autom Control 59(4):835\u2013850","journal-title":"IEEE Trans Autom Control"},{"key":"11166_CR120","first-page":"15032","volume":"34","author":"J Terry","year":"2021","unstructured":"Terry J, Black B, Grammel N et al (2021) Pettingzoo: gym for multi-agent reinforcement learning. Adv Neural Inf Process Syst 34:15032\u201315043","journal-title":"Adv Neural Inf Process Syst"},{"issue":"8","key":"11166_CR121","doi-asserted-by":"publisher","first-page":"2590","DOI":"10.1109\/JSAC.2021.3087248","volume":"39","author":"TY Tung","year":"2021","unstructured":"Tung TY, Kobus S, Roig JP et al (2021) Effective communications: a joint learning and communication framework for multi-agent reinforcement learning over noisy channels. IEEE J Sel Areas Commun 39(8):2590\u20132603","journal-title":"IEEE J Sel Areas Commun"},{"key":"11166_CR188","doi-asserted-by":"crossref","unstructured":"Van Hasselt H, Guez A, Silver D (2016) Deep reinforcement learning with double q-learning. In: Proceedings of the AAAI conference on artificial intelligence","DOI":"10.1609\/aaai.v30i1.10295"},{"key":"11166_CR122","unstructured":"Vaswani A, Shazeer N, Parmar N et al (2017) Attention is all you need. Adv Neural Inf Process Syst 30"},{"issue":"7782","key":"11166_CR123","doi-asserted-by":"publisher","first-page":"350","DOI":"10.1038\/s41586-019-1724-z","volume":"575","author":"O Vinyals","year":"2019","unstructured":"Vinyals O, Babuschkin I, Czarnecki WM et al (2019) Grandmaster level in Starcraft Ii using multi-agent reinforcement learning. Nature 575(7782):350\u2013354","journal-title":"Nature"},{"key":"11166_CR124","doi-asserted-by":"crossref","unstructured":"Von Stackelberg H (2010) Market structure and equilibrium. Springer Science & Business Media","DOI":"10.1007\/978-3-642-12586-7"},{"key":"11166_CR189","unstructured":"Wang Z, Schaul T, Hessel M et al (2016) Dueling network architectures for deep reinforcement learning. In: International conference on machine learning, PMLR, pp 1995\u20132003"},{"key":"11166_CR128","doi-asserted-by":"crossref","unstructured":"Wang J, Zhang Y, Kim TK et al (2020a) Shapley q-value: a local reward approach to solve global reward games. In: Proceedings of the AAAI conference on artificial intelligence, pp 7285\u20137292","DOI":"10.1609\/aaai.v34i05.6220"},{"key":"11166_CR129","unstructured":"Wang L, Yang Z, Wang Z (2020b) Breaking the curse of many agents: provable mean embedding q-iteration for mean-field reinforcement learning. In: International conference on machine learning, PMLR, pp 10092\u201310103"},{"key":"11166_CR131","unstructured":"Wang R, He X, Yu R et al (2020c) Learning efficient multi-agent communication: an information bottleneck approach. In: International conference on machine learning, PMLR, pp 9908\u20139918"},{"key":"11166_CR132","unstructured":"Wang T, Dong H, Lesser V et al (2020d) Roma: multi-agent reinforcement learning with emergent roles. In: Proceedings of the 37th international conference on machine learning, ICML\u201920"},{"key":"11166_CR133","unstructured":"Wang T, Wang J, Zheng C et al (2020e) Learning nearly decomposable value functions via communication minimization. In: International conference on learning representation"},{"key":"11166_CR125","unstructured":"Wang J, Ren Z, Liu T et al (2021a) QPLEX: Duplex dueling multi-agent q-learning. In: International conference on learning representations"},{"issue":"9","key":"11166_CR134","first-page":"4555","volume":"44","author":"X Wang","year":"2021","unstructured":"Wang X, Chen Y, Zhu W (2021b) A survey on curriculum learning. IEEE Trans Pattern Anal Mach Intell 44(9):4555\u20134576","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11166_CR136","unstructured":"Wang Y, Han B, Wang T et al (2021c) DOP: Off-policy multi-agent decomposed policy gradients. In: International conference on learning representations"},{"key":"11166_CR126","doi-asserted-by":"crossref","unstructured":"Wang J, Xue D, Zhao J Mastering the game of 3v3 snakes with rule-enhanced multi-agent reinforcement learning. In: 2022 IEEE Conference on Games (CoG), IEEE, pp 229\u2013236","DOI":"10.1109\/CoG51982.2022.9893608"},{"key":"11166_CR127","first-page":"5941","volume":"35","author":"J Wang","year":"2022","unstructured":"Wang J, Zhang Y, Gu Y et al (2022b) Shaq: incorporating Shapley value theory into multi-agent q-learning. Adv Neural Inf Process Syst 35:5941\u20135954","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR135","unstructured":"Wang X, Zhang Z, Zhang W (2022c) Model-based multi-agent reinforcement learning: Recent progress and prospects. arXiv preprint arXiv:220310603"},{"issue":"6","key":"11166_CR130","doi-asserted-by":"publisher","first-page":"186345","DOI":"10.1007\/s11704-024-40231-1","volume":"18","author":"L Wang","year":"2024","unstructured":"Wang L, Ma C, Feng X et al (2024) A survey on large Language model based autonomous agents. Front Comput Sci 18(6):186345","journal-title":"Front Comput Sci"},{"key":"11166_CR138","doi-asserted-by":"crossref","unstructured":"Wen G, Fu J, Dai P et al (2021) Dtde: a new cooperative multi-agent reinforcement learning framework. Innovation 2(4)","DOI":"10.1016\/j.xinn.2021.100162"},{"key":"11166_CR139","first-page":"16509","volume":"35","author":"M Wen","year":"2022","unstructured":"Wen M, Kuba J, Lin R et al (2022) Multi-agent reinforcement learning is a sequence modeling problem. Adv Neural Inf Process Syst 35:16509\u201316521","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR140","unstructured":"Whiteson S, Samvelyan M, Rashid T et al (2019) The starcraft multi-agent challenge. In: Proceedings of the international joint conference on autonomous agents and multiagent systems, AAMAS, pp 2186\u20132188"},{"key":"11166_CR190","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1007\/BF00992696","volume":"8","author":"RJ Williams","year":"1992","unstructured":"Williams RJ (1992) Simple statistical gradient-following algorithms for connectionist reinforcement learning. Mach Learn 8:229\u2013256","journal-title":"Mach Learn"},{"issue":"02n03","key":"11166_CR141","doi-asserted-by":"publisher","first-page":"265","DOI":"10.1142\/S0219525901000188","volume":"4","author":"DH Wolpert","year":"2001","unstructured":"Wolpert DH, Tumer K (2001) Optimal payoff functions for members of collectives. Adv Complex Syst 4(02n03):265\u2013279","journal-title":"Adv Complex Syst"},{"issue":"7896","key":"11166_CR142","doi-asserted-by":"publisher","first-page":"223","DOI":"10.1038\/s41586-021-04357-7","volume":"602","author":"PR Wurman","year":"2022","unstructured":"Wurman PR, Barrett S, Kawamoto K et al (2022) Outracing champion Gran turismo drivers with deep reinforcement learning. Nature 602(7896):223\u2013228","journal-title":"Nature"},{"key":"11166_CR143","unstructured":"Xu Z, Yu C, Fang F et al (2024) Language agents with reinforcement learning for strategic play in the werewolf game. In: Forty-first international conference on machine learning"},{"key":"11166_CR146","unstructured":"Yang Y, Wang J (2020) An overview of multi-agent reinforcement learning from game theoretical perspective. arXiv preprint arXiv:201100583"},{"key":"11166_CR144","unstructured":"Yang Y, Hao J, Chen G et al (2020a) Q-value path decomposition for deep multiagent reinforcement learning. In: International conference on machine learning, PMLR, pp 10706\u201310715"},{"key":"11166_CR145","unstructured":"Yang Y, Hao J, Liao B et al (2020b) Qatten: a general framework for cooperative multiagent reinforcement learning. arXiv preprint arXiv:200203939"},{"key":"11166_CR147","doi-asserted-by":"crossref","unstructured":"Yao M, Feng X, Yin Q (2023) More like real world game challenge for partially observable multi-agent cooperation. arXiv preprint arXiv:230508394","DOI":"10.1007\/978-981-97-8505-6_32"},{"key":"11166_CR148","first-page":"621","volume":"33","author":"D Ye","year":"2020","unstructured":"Ye D, Chen G, Zhang W et al (2020a) Towards playing full Moba games with deep reinforcement learning. Adv Neural Inf Process Syst 33:621\u2013632","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR149","doi-asserted-by":"crossref","unstructured":"Ye D, Liu Z, Sun M et al (2020b) Mastering complex control in moba games with deep reinforcement learning. In: Proceedings of the AAAI conference on artificial intelligence, pp 6672\u20136679","DOI":"10.1609\/aaai.v34i04.6144"},{"issue":"3","key":"11166_CR150","doi-asserted-by":"publisher","first-page":"299","DOI":"10.1007\/s11633-022-1384-6","volume":"20","author":"QY Yin","year":"2023","unstructured":"Yin QY, Yang J, Huang KQ et al (2023) Ai in human-computer gaming: techniques, challenges and opportunities. Mach Intell Res 20(3):299\u2013317","journal-title":"Mach Intell Res"},{"key":"11166_CR151","first-page":"24611","volume":"35","author":"C Yu","year":"2022","unstructured":"Yu C, Velu A, Vinitsky E et al (2022) The surprising effectiveness of Ppo in cooperative multi-agent games. Adv Neural Inf Process Syst 35:24611\u201324624","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR152","doi-asserted-by":"crossref","unstructured":"Yuan L, Wang J, Zhang F et al (2022) Multi-agent incentive communication via decentralized teammate modeling. In: Proceedings of the AAAI conference on artificial intelligence, pp 9466\u20139474","DOI":"10.1609\/aaai.v36i9.21179"},{"key":"11166_CR153","doi-asserted-by":"crossref","unstructured":"Yun WJ, Lim B, Jung S et al (2021) Attention-based reinforcement learning for real-time uav semantic communication. In: 2021 17th International symposium on wireless communication systems (ISWCS), IEEE, pp 1\u20136","DOI":"10.1109\/ISWCS49558.2021.9562230"},{"key":"11166_CR154","unstructured":"Zang Y, He J, Li K et al (2023) Sequential cooperative multi-agent reinforcement learning. In: Proceedings of the 2023 international conference on autonomous agents and multiagent systems. International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC, AAMAS \u201923, p 485\u2013493"},{"key":"11166_CR155","unstructured":"Zha D, Xie J, Ma W et al (2021) Douzero: mastering doudizhu with self-play deep reinforcement learning. In: international conference on machine learning, PMLR, pp 12333\u201312344"},{"key":"11166_CR165","unstructured":"Zhang Z (2024) Advancing sample efficiency and explainability in multi-agent reinforcement learning. In: Proceedings of the 23rd International Conference on Autonomous Agents and Multiagent Systems, pp 2791\u20132793"},{"issue":"2","key":"11166_CR164","doi-asserted-by":"publisher","first-page":"968","DOI":"10.1109\/TAC.2023.3288025","volume":"69","author":"Y Zhang","year":"2023","unstructured":"Zhang Y, Zavlanos MM (2023) Cooperative multiagent reinforcement learning with partial observations. IEEE Trans Autom Control 69(2):968\u2013981","journal-title":"IEEE Trans Autom Control"},{"key":"11166_CR162","unstructured":"Zhang SQ, Zhang Q, Lin J (2019) Efficient communication in multi-agent reinforcement learning via variance based control. Adv Neural Inf Process Syst 32"},{"key":"11166_CR157","doi-asserted-by":"crossref","unstructured":"Zhang H, Chen W, Huang Z et al (2020a) Bi-level actor-critic for multi-agent coordination. In: Proceedings of the AAAI conference on artificial intelligence, pp 7325\u20137332","DOI":"10.1609\/aaai.v34i05.6226"},{"key":"11166_CR163","first-page":"17271","volume":"33","author":"SQ Zhang","year":"2020","unstructured":"Zhang SQ, Zhang Q, Lin J (2020b) Succinct and robust multi-agent communication with temporal message control. Adv Neural Inf Process Syst 33:17271\u201317282","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR158","doi-asserted-by":"crossref","unstructured":"Zhang K, Yang Z, Ba \u0327sar T (2021a) Multi-agent reinforcement learning: a selective overview of theories and algorithms. Handbook of reinforcement learning and control pp 321\u2013384","DOI":"10.1007\/978-3-030-60990-0_12"},{"issue":"12","key":"11166_CR159","doi-asserted-by":"publisher","first-page":"5925","DOI":"10.1109\/TAC.2021.3049345","volume":"66","author":"K Zhang","year":"2021","unstructured":"Zhang K, Yang Z, Liu H et al (2021b) Finite-sample analysis for decentralized batch multiagent reinforcement learning with networked agents. IEEE Trans Autom Control 66(12):5925\u20135940","journal-title":"IEEE Trans Autom Control"},{"issue":"10","key":"11166_CR161","doi-asserted-by":"publisher","first-page":"7900","DOI":"10.1109\/TNNLS.2022.3146976","volume":"34","author":"R Zhang","year":"2022","unstructured":"Zhang R, Zong Q, Zhang X et al (2022) Game of drones: Multi-uav pursuit-evasion game with online motion planning by deep reinforcement learning. IEEE Trans Neural Networks Learn Syst 34(10):7900\u20137909","journal-title":"IEEE Trans Neural Networks Learn Syst"},{"key":"11166_CR156","doi-asserted-by":"publisher","unstructured":"Zhang B, Li L, Xu Z et al (2023a) Inducing stackelberg equilibrium through spatio-temporal sequential decision-making in multi-agent reinforcement learning. In: Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, IJCAI \u201923. https:\/\/doi.org\/10.24963\/ijcai.2023\/40","DOI":"10.24963\/ijcai.2023\/40"},{"issue":"175","key":"11166_CR160","first-page":"1","volume":"24","author":"K Zhang","year":"2023","unstructured":"Zhang K, Kakade SM, Basar T et al (2023b) Model-based multi-agent Rl in zero-summarkov games with near-optimal sample complexity. J Mach Learn Res 24(175):1\u201353","journal-title":"J Mach Learn Res"},{"issue":"2","key":"11166_CR168","doi-asserted-by":"publisher","first-page":"199","DOI":"10.1109\/TG.2020.2990865","volume":"12","author":"Y Zhao","year":"2020","unstructured":"Zhao Y, Borovikov I, de Mesentier Silva F et al (2020) Winning is not everything: enhancing game development with intelligent agents. IEEE Trans Games 12(2):199\u2013212","journal-title":"IEEE Trans Games"},{"key":"11166_CR166","doi-asserted-by":"crossref","unstructured":"Zhao E, Yan R, Li J et al (2022a) Alphaholdem: high-performance artificial intelligence for heads-up no-limit poker via end-to-end reinforcement learning. In: Proceedings of the AAAI conference on artificial intelligence, pp 4689\u2013469","DOI":"10.1609\/aaai.v36i4.20394"},{"issue":"1","key":"11166_CR167","doi-asserted-by":"publisher","first-page":"140","DOI":"10.1109\/TG.2022.3232390","volume":"16","author":"J Zhao","year":"2022","unstructured":"Zhao J, Hu X, Yang M et al (2022b) Ctds: centralized teacher with decentralized student for multiagent reinforcement learning. IEEE Trans Games 16(1):140\u2013150","journal-title":"IEEE Trans Games"},{"key":"11166_CR169","doi-asserted-by":"crossref","unstructured":"Zhao Y, Zhao J, Hu X et al (2022c) Douzero+: Improving doudizhu ai by opponent modeling and coach-guided learning. In: 2022 IEEE conference on games (CoG), IEEE, pp 127\u2013134","DOI":"10.1109\/CoG51982.2022.9893710"},{"key":"11166_CR170","doi-asserted-by":"crossref","unstructured":"Zhao Y, Zhao J, Hu X et al (2023) Full douzero+: improving Doudizhu Ai by opponent modeling, coach-guided training and bidding learning. IEEE Trans Games 16(3):518\u2013529","DOI":"10.1109\/TG.2023.3299612"},{"key":"11166_CR171","doi-asserted-by":"crossref","unstructured":"Zheng L, Yang J, Cai H et al (2018) Magent: A many-agent reinforcement learning platform for artificial collective intelligence. In: Proceedings of the AAAI conference on artificial intelligence","DOI":"10.1609\/aaai.v32i1.11371"},{"key":"11166_CR173","first-page":"11853","volume":"33","author":"M Zhou","year":"2020","unstructured":"Zhou M, Liu Z, Sui P et al (2020) Learning implicit credit assignment for cooperative multi-agent reinforcement learning. Adv Neural Inf Process Syst 33:11853\u201311864","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR172","first-page":"15757","volume":"35","author":"H Zhou","year":"2022","unstructured":"Zhou H, Lan T, Aggarwal V (2022) Pac: assisted value factorization with counterfactual predictions in multi-agent reinforcement learning. Adv Neural Inf Process Syst 35:15757\u201315769","journal-title":"Adv Neural Inf Process Syst"},{"key":"11166_CR174","unstructured":"Zhou Z, Liu G, Tang Y (2023b) Multi-agent reinforcement learning: methods, applications, visionary prospects, and challenges. arXiv preprint arXiv:230510091"},{"key":"11166_CR176","unstructured":"Zhu C, Dastani M, Wang S (2022) A survey of multi-agent reinforcement learning with communication. arXiv preprint arXiv:220308975"}],"container-title":["Artificial Intelligence Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11166-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10462-025-11166-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11166-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,4,17]],"date-time":"2025-04-17T20:03:29Z","timestamp":1744920209000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10462-025-11166-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,15]]},"references-count":188,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2025,6]]}},"alternative-id":["11166"],"URL":"https:\/\/doi.org\/10.1007\/s10462-025-11166-1","relation":{},"ISSN":["1573-7462"],"issn-type":[{"value":"1573-7462","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,15]]},"assertion":[{"value":"28 February 2025","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 March 2025","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"165"}}