{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T11:22:20Z","timestamp":1780053740258,"version":"3.54.0"},"reference-count":55,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2024,4,16]],"date-time":"2024-04-16T00:00:00Z","timestamp":1713225600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Robot. AI"],"abstract":"<jats:p><jats:bold>Introduction:<\/jats:bold> Multi-agent systems are an interdisciplinary research field that describes the concept of multiple decisive individuals interacting with a usually partially observable environment. Given the recent advances in single-agent reinforcement learning, multi-agent reinforcement learning (RL) has gained tremendous interest in recent years. Most research studies apply a fully centralized learning scheme to ease the transfer from the single-agent domain to multi-agent systems.<\/jats:p><jats:p><jats:bold>Methods:<\/jats:bold> In contrast, we claim that a decentralized learning scheme is preferable for applications in real-world scenarios as this allows deploying a learning algorithm on an individual robot rather than deploying the algorithm to a complete fleet of robots. Therefore, this article outlines a novel actor\u2013critic (AC) approach tailored to cooperative MARL problems in sparsely rewarded domains. Our approach decouples the MARL problem into a set of distributed agents that model the other agents as responsive entities. In particular, we propose using two separate critics per agent to distinguish between the joint task reward and agent-based costs as commonly applied within multi-robot planning. On one hand, the agent-based critic intends to decrease agent-specific costs. On the other hand, each agent intends to optimize the joint team reward based on the joint task critic. As this critic still depends on the joint action of all agents, we outline two suitable behavior models based on Stackelberg games: a game against nature and a dyadic game against each agent. Following these behavior models, our algorithm allows fully decentralized execution and training.<\/jats:p><jats:p><jats:bold>Results and Discussion:<\/jats:bold> We evaluate our presented method using the proposed behavior models within a sparsely rewarded simulated multi-agent environment. Although our approach already outperforms the state-of-the-art learners, we conclude this article by outlining possible extensions of our algorithm that future research may build upon.<\/jats:p>","DOI":"10.3389\/frobt.2024.1229026","type":"journal-article","created":{"date-parts":[[2024,4,16]],"date-time":"2024-04-16T15:58:30Z","timestamp":1713283110000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["Decentralized multi-agent reinforcement learning based on best-response policies"],"prefix":"10.3389","volume":"11","author":[{"given":"Volker","family":"Gabler","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dirk","family":"Wollherr","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2024,4,16]]},"reference":[{"key":"B1","unstructured":"Reducing overestimation bias in multi-agent domains using double centralized critics\n            AckermannJ.\n            GablerV.\n            OsaT.\n            SugiyamaM.\n          2019"},{"key":"B2","doi-asserted-by":"publisher","first-page":"26","DOI":"10.1109\/MSP.2017.2743240","article-title":"Deep reinforcement learning: a brief survey","volume":"34","author":"Arulkumaran","year":"2017","journal-title":"IEEE Signal Process. Mag."},{"key":"B3","doi-asserted-by":"publisher","first-page":"679","DOI":"10.1512\/iumj.1957.6.56038","article-title":"A markovian decision process","volume":"6","author":"Bellman","year":"1957","journal-title":"J. Math. Mech."},{"key":"B4","doi-asserted-by":"crossref","DOI":"10.15607\/RSS.2016.XII.033","article-title":"Combined optimization and reinforcement learning for manipulation skills","volume-title":"Robotics: science and systems (RSS)","author":"Englert","year":"2016"},{"key":"B5","unstructured":"Learning to communicate with deep multi-agent reinforcement learning\n            FoersterJ. N.\n            AssaelY. M.\n            de FreitasN.\n            WhitesonS.\n          2016"},{"key":"B6","first-page":"2974","article-title":"Counterfactual multi-agent policy gradients","volume-title":"AAAI conference on artificial intelligence","author":"Foerster","year":"2018"},{"key":"B7","volume-title":"Acquiring diverse robot skills via maximum entropy deep reinforcement learning","author":"Haarnoja","year":"2018"},{"key":"B8","volume-title":"Soft actor-critic algorithms and applications","author":"Haarnoja","year":"2018"},{"key":"B9","unstructured":"Emergence of language with multi-agent games: learning to communicate with sequences of symbols\n            HavrylovS.\n            TitovI.\n          2017"},{"key":"B10","unstructured":"Multi-agent soft actor-critic based hybrid motion planner for mobile robots\n            HeZ.\n            DongL.\n            SongC.\n            SunC.\n          2021"},{"key":"B11","doi-asserted-by":"publisher","first-page":"750","DOI":"10.1007\/s10458-019-09421-1","article-title":"A survey and critique of multiagent deep reinforcement learning","volume":"33","author":"Hernandez-Leal","year":"2019","journal-title":"Auton. Agents Multi Agent Syst."},{"key":"B12","first-page":"2146","article-title":"A very condensed survey and critique of multiagent deep reinforcement learning","volume-title":"AAMAS","author":"Hernandez-Leal","year":"2020"},{"key":"B13","first-page":"2961","article-title":"Actor-attention-critic for multi-agent reinforcement learning","volume-title":"International conference on machine learning (ICML), (PMLR), vol. 97 of proceedings of machine learning research","author":"Iqbal","year":"2019"},{"key":"B14","first-page":"3040","article-title":"Social influence as intrinsic motivation for multi-agent deep reinforcement learning","volume-title":"Proceedings of the 36th international conference on machine learning, ICML 2019, 9-15 june 2019, long beach, California, USA vol. 97 of proceedings of machine learning research","author":"Jaques","year":"2019"},{"key":"B15","doi-asserted-by":"publisher","first-page":"237","DOI":"10.1613\/jair.301","article-title":"Reinforcement learning: a survey","volume":"4","author":"Kaelbling","year":"1996","journal-title":"J. Artif. Intell. Res."},{"key":"B16","article-title":"Adam: a method for stochastic optimization","volume-title":"International conference on learning representations (ICLR)","author":"Kingma","year":"2015"},{"key":"B17","doi-asserted-by":"publisher","first-page":"1238","DOI":"10.1177\/0278364913495721","article-title":"Reinforcement learning in robotics: a survey","volume":"32","author":"Kober","year":"2013","journal-title":"J. Artif. Intell. Res."},{"key":"B18","article-title":"Policy search via the signed derivative","volume-title":"Robotics: science and systems (RSS)","author":"Kolter","year":"2009"},{"key":"B19","doi-asserted-by":"publisher","first-page":"55","DOI":"10.3233\/KES-2010-0206","article-title":"The world of independent learners is not markovian","volume":"15","author":"Laurent","year":"2011","journal-title":"Int. J. Knowl. Based Intell. Eng. Syst."},{"key":"B20","article-title":"Learning multi-level hierarchies with hindsight","author":"Levy","year":"2019"},{"key":"B21","first-page":"4213","article-title":"Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient","author":"Li","year":"2019"},{"key":"B22","article-title":"Continuous control with deep reinforcement learning","author":"Lillicrap","year":"2016"},{"key":"B23","first-page":"157","article-title":"Markov games as a framework for multi-agent reinforcement learning","author":"Littman","year":"1994"},{"key":"B24","first-page":"6379","article-title":"Multi-agent actor-critic for mixed cooperative-competitive environments","author":"Lowe","year":"2017"},{"key":"B25","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1016\/S0079-7421(08)60536-8","article-title":"Catastrophic interference in connectionist networks: the sequential learning problem","volume-title":"Psychology of learning and motivation vol. 24","author":"McCloskey","year":"1989"},{"key":"B26","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"B27","unstructured":"Emergence of grounded compositional language in multi-agent populations\n            MordatchI.\n            AbbeelP.\n          2017"},{"key":"B28","doi-asserted-by":"publisher","first-page":"48","DOI":"10.1073\/pnas.36.1.48","article-title":"Equilibrium points in N-person games","volume":"36","author":"Nash","year":"1950","journal-title":"Proc. Natl. Acad. Sci."},{"key":"B29","first-page":"363","article-title":"Autonomous inverted helicopter flight via reinforcement learning","author":"Ng","year":"2004"},{"key":"B30","doi-asserted-by":"publisher","first-page":"3826","DOI":"10.1109\/TCYB.2020.2977374","article-title":"Deep reinforcement learning for multiagent systems: a review of challenges, solutions, and applications","volume":"50","author":"Nguyen","year":"2020","journal-title":"IEEE Trans. Cybern."},{"key":"B31","doi-asserted-by":"publisher","first-page":"1627","DOI":"10.1021\/ac60214a047","article-title":"Smoothing and differentiation of data by simplified least squares procedures","volume":"36","author":"Savitzky","year":"1964","journal-title":"Anal. Chem."},{"key":"B32","first-page":"1889","article-title":"Trust region policy optimization","volume-title":"International conference on machine learning (ICML) vol. 37 of JMLR workshop and conference proceedings","author":"Schulman","year":"2015"},{"key":"B33","unstructured":"Proximal policy optimization algorithms\n            SchulmanJ.\n            WolskiF.\n            DhariwalP.\n            RadfordA.\n            KlimovO.\n          2017"},{"key":"B34","volume-title":"A value for N-person games","author":"Shapley","year":"1952"},{"key":"B35","doi-asserted-by":"publisher","first-page":"1095","DOI":"10.1073\/pnas.39.10.1953","article-title":"Stochastic games","volume":"39","author":"Shapley","year":"1953","journal-title":"Proc. Natl. Acad. Sci."},{"key":"B36","first-page":"1","article-title":"Multi-agent reinforcement learning for problems with combined individual and team reward","author":"Sheikh","year":"2020"},{"key":"B37","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511811654","volume-title":"Multiagent systems: algorithmic, game-theoretic, and logical foundations","author":"Shoham","year":"2008"},{"key":"B38","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","article-title":"Mastering the game of Go with deep neural networks and tree search","volume":"529","author":"Silver","year":"2016","journal-title":"Nature"},{"key":"B39","first-page":"387","article-title":"Deterministic policy gradient algorithms","author":"Silver","year":"2014"},{"key":"B40","article-title":"V-MPO: on-policy maximum a posteriori policy optimization for discrete and continuous control","author":"Song","year":"2020"},{"key":"B41","doi-asserted-by":"publisher","first-page":"135605","DOI":"10.1109\/ACCESS.2020.3011670","article-title":"A novel multi-agent parallel-critic network architecture for cooperative-competitive reinforcement learning","volume":"8","author":"Sun","year":"2020","journal-title":"IEEE Access"},{"key":"B42","first-page":"1057","article-title":"Policy gradient methods for reinforcement learning with function approximation","volume-title":"Annual conference on neural information processing systems (NeurIPS)","author":"Sutton","year":""},{"key":"B43","doi-asserted-by":"publisher","first-page":"181","DOI":"10.1016\/S0004-3702(99)00052-1","article-title":"Between mdps and semi-mdps: a framework for temporal abstraction in reinforcement learning","volume":"112","author":"Sutton","year":"","journal-title":"Artif. Intell."},{"key":"B44","first-page":"330","article-title":"Multi-agent reinforcement learning: independent versus cooperative agents","volume-title":"International conference on machine learning (ICML)","author":"Tan","year":"1993"},{"key":"B45","doi-asserted-by":"publisher","first-page":"42568","DOI":"10.1109\/ACCESS.2021.3062457","article-title":"A novel hierarchical soft actor-critic algorithm for multi-logistics robots task allocation","volume":"9","author":"Tang","year":"2021","journal-title":"IEEE Access"},{"key":"B46","first-page":"602","article-title":"A regularized opponent model with maximum entropy objective","volume-title":"Proceedings of the twenty-eighth international joint conference on artificial intelligence, IJCAI 2019, Macao, China, august 10-16, 2019","author":"Tian","year":"2019"},{"key":"B47","volume-title":"Stochastic dynamic programming","author":"van der Wal","year":"1980"},{"key":"B48","first-page":"2613","article-title":"Double q-learning","volume-title":"Annual conference on neural information processing systems (NeurIPS)","author":"van Hasselt","year":"2010"},{"key":"B49","doi-asserted-by":"publisher","first-page":"350","DOI":"10.1038\/s41586-019-1724-z","article-title":"Grandmaster level in StarCraft II using multi-agent reinforcement learning","volume":"575","author":"Vinyals","year":"2019","journal-title":"Nature"},{"key":"B50","doi-asserted-by":"publisher","first-page":"261","DOI":"10.1038\/s41592-019-0686-2","article-title":"SciPy 1.0: fundamental algorithms for scientific computing in Python","volume":"17","author":"Virtanen","year":"2020","journal-title":"Nat. Methods"},{"key":"B51","article-title":"Multiagent soft q-learning","author":"Wei","year":"2018"},{"key":"B52","doi-asserted-by":"publisher","first-page":"5886","DOI":"10.1109\/TCOMM.2021.3086535","article-title":"Caching transient content for iot sensing: multi-agent soft actor-critic","volume":"69","author":"Wu","year":"2021","journal-title":"IEEE Trans. Commun."},{"key":"B53","unstructured":"An overview of multi-agent reinforcement learning from game theoretical perspective\n            YangY.\n            WangJ.\n          2020"},{"key":"B54","doi-asserted-by":"publisher","first-page":"321","DOI":"10.1007\/978-3-030-60990-0_12","article-title":"Multi-agent reinforcement learning: a selective overview of theories and algorithms","author":"Zhang","year":"2021","journal-title":"Handb. Reinf. Learn. Control"},{"key":"B55","first-page":"55","article-title":"Lyapunov-based reinforcement learning for decentralized multi-agent control","volume-title":"International conference on distributed artificial intelligence (DAI) vol. 12547 of lecture notes in computer science","author":"Zhang","year":"2020"}],"container-title":["Frontiers in Robotics and AI"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frobt.2024.1229026\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,4,16]],"date-time":"2024-04-16T15:58:43Z","timestamp":1713283123000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frobt.2024.1229026\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,16]]},"references-count":55,"alternative-id":["10.3389\/frobt.2024.1229026"],"URL":"https:\/\/doi.org\/10.3389\/frobt.2024.1229026","relation":{},"ISSN":["2296-9144"],"issn-type":[{"value":"2296-9144","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,16]]},"article-number":"1229026"}}