{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T13:19:58Z","timestamp":1753881598741,"version":"3.41.2"},"reference-count":29,"publisher":"Wiley","issue":"1","license":[{"start":{"date-parts":[[2021,12,8]],"date-time":"2021-12-08T00:00:00Z","timestamp":1638921600000},"content-version":"vor","delay-in-days":341,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Journal of Sensors"],"published-print":{"date-parts":[[2021,1]]},"abstract":"<jats:p>The exponential explosion of joint actions and massive data collection are two main challenges in multiagent reinforcement learning algorithms with centralized training. To overcome these problems, in this paper, we propose a model\u2010free and fully decentralized actor\u2010critic multiagent reinforcement learning algorithm based on message diffusion. To this end, the agents are assumed to be placed in a time\u2010varying communication network. Each agent makes limited observations regarding the global state and joint actions; therefore, it needs to obtain and share information with others over the network. In the proposed algorithm, agents hold local estimations of the global state and joint actions and update them with local observations and the messages received from neighbors. Under the hypothesis of the global value decomposition, the gradient of the global objective function to an individual agent is derived. The convergence of the proposed algorithm with linear function approximation is guaranteed according to the stochastic approximation theory. In the experiments, the proposed algorithm was applied to a passive location task multiagent environment and achieved superior performance compared to state\u2010of\u2010the\u2010art algorithms.<\/jats:p>","DOI":"10.1155\/2021\/8739206","type":"journal-article","created":{"date-parts":[[2021,12,8]],"date-time":"2021-12-08T23:20:45Z","timestamp":1639005645000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Decentralized Multiagent Actor\u2010Critic Algorithm Based on Message Diffusion"],"prefix":"10.1155","volume":"2021","author":[{"given":"Siyuan","family":"Ding","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5948-4634","authenticated-orcid":false,"given":"Shengxiang","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4241-508X","authenticated-orcid":false,"given":"Guangyi","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ou","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ke","family":"Ke","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yijie","family":"Bai","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weiye","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2021,12,8]]},"reference":[{"key":"e_1_2_11_1_2","doi-asserted-by":"publisher","DOI":"10.1613\/jair.301"},{"volume-title":"Reinforcement Learning: An Introduction","year":"2018","author":"Sutton R. S.","key":"e_1_2_11_2_2"},{"key":"e_1_2_11_3_2","doi-asserted-by":"crossref","unstructured":"YangY. KiumarsiB. ModaresH. andXuC. Model-free _-policy iteration for discrete-time linear quadratic regulation IEEE Transactions on Neural Networks and Learning Systems 2021 1\u201315 https:\/\/doi.org\/10.1109\/TNNLS.2021.3098985.","DOI":"10.1109\/TNNLS.2021.3098985"},{"key":"e_1_2_11_4_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature24270"},{"key":"e_1_2_11_5_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_2_11_6_2","unstructured":"LillicrapT. P. HuntJ. J. PritzelA. HeessN. ErezT. TassaY. SilverD. andWierstraD. Continuous control with deep reinforcement learning Proceedings of 4th International Conference on Learning Representations 2016."},{"key":"e_1_2_11_7_2","doi-asserted-by":"crossref","unstructured":"BusoniuL. BabuskaR. andDe SchutterB. Multi-agent reinforcement learning: a survey Proceedings of 9th International Conference on Control Automation Robotics and Vision 2006 Singapore 1\u20136 https:\/\/doi.org\/10.1109\/ICARCV.2006.345353 2-s2.0-34547192059.","DOI":"10.1109\/ICARCV.2006.345353"},{"key":"e_1_2_11_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2020.2980815"},{"key":"e_1_2_11_9_2","unstructured":"Hernandez-LealP. KartalB. andTaylorM. E. A very condensed survey and critique of multiagent deep reinforcement learning Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems 2020 2146\u20132148."},{"key":"e_1_2_11_10_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-019-1724-z"},{"key":"e_1_2_11_11_2","doi-asserted-by":"crossref","unstructured":"FoersterJ. N. FarquharG. AfourasT. NardelliN. andWhitesonS. Counterfactual multi-agent policy gradients Proceedings of AAAI Conference on Artificial Intelligence 2018.","DOI":"10.1609\/aaai.v32i1.11794"},{"key":"e_1_2_11_12_2","doi-asserted-by":"publisher","DOI":"10.1613\/jair.2447"},{"key":"e_1_2_11_13_2","unstructured":"RashidT. SamvelyanM. SchroederC. FarquharG. FoersterJ. andWhitesonS. QMIX: monotonic value function factorisation for deep multi-agent reinforcement learning Proceedings of 35th International Conference on Machine Learning 2018 6846\u20136859."},{"key":"e_1_2_11_14_2","unstructured":"SunehagP. LeverG. GruslysA. CzarneckiW. M. ZambaldiV. JaderbergM. LanctotM. SonneratN. LeiboJ. Z. TuylsK. andGraepelT. Value-decomposition networks for cooperative multi-agent learning based on team reward Proceedings of the 2018 International Joint Conference on Autonomous Agents and Multiagent Systems 2018 2085\u20132087."},{"key":"e_1_2_11_15_2","unstructured":"LoweR. WuY. TamarA. HarbJ. AbbeelP. andMordatchI. Multi-agent actor-critic for mixed cooperative-competitive environments Proceedings of Conference on Neural Information Processing Systems 2017 6379\u20136390."},{"key":"e_1_2_11_16_2","doi-asserted-by":"crossref","unstructured":"LauerM.andRiedmillerM. Distributed reinforcement learning in multi-agent networks Proceedings of the 7th International Conference on Machine Learning 2000 St. Martin France 296\u2013299 https:\/\/doi.org\/10.1109\/CAMSAP.2013.6714066 2-s2.0-84894160317.","DOI":"10.1109\/CAMSAP.2013.6714066"},{"key":"e_1_2_11_17_2","doi-asserted-by":"crossref","unstructured":"TanM. Multi-agent reinforcement learning: independent vs. cooperative agents 1993 Proceedings ofthe tenth international conference on machine learning 330\u2013337 https:\/\/doi.org\/10.1016\/B978-1-55860-307-3.50049-6.","DOI":"10.1016\/B978-1-55860-307-3.50049-6"},{"key":"e_1_2_11_18_2","doi-asserted-by":"publisher","DOI":"10.3233\/KES-2010-0206"},{"key":"e_1_2_11_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2013.2241057"},{"key":"e_1_2_11_20_2","unstructured":"ZhangK. YangZ. LiuH. ZhangT. andBasarT. Fully decentralized multi-agent reinforcement learning with networked agents 35th International Conference on Machine Learning 2018 9340\u20139371."},{"key":"e_1_2_11_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2012.2217338"},{"key":"e_1_2_11_22_2","unstructured":"SuttonR. S. DAM. A. SinghS. P. andMansourY. Policy gradient methods for reinforcement learning with function approximation 2000 Advances in Neural Information Processing Systems 1057\u20131063."},{"volume-title":"Stochastic Processes","year":"2013","author":"Ross S. M.","key":"e_1_2_11_23_2"},{"key":"e_1_2_11_24_2","unstructured":"DegrisT. WhiteM. andSuttonR. S. Off-policy actor-critic ICML\u201912: Proceedings of the 29th International Conference on International Conference on Machine Learning 2012 179\u2013186."},{"key":"e_1_2_11_25_2","unstructured":"PrasadH. PrashanthL. A. andBhatnagarS. Two-timescale algorithms for learning Nash equilibria in general-sum stochastic games Proceedings of the 2015 International Conference on Autonomous Agents and Multiagent Systems 2015 1371\u20131379."},{"key":"e_1_2_11_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-93-86279-38-5"},{"key":"e_1_2_11_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3075475"},{"key":"e_1_2_11_28_2","doi-asserted-by":"crossref","unstructured":"KushnerH. J.andClarkD. Stochastic approximation methods for constrained and unconstrained systems 1978 Springer Science & Business Media.","DOI":"10.1007\/978-1-4684-9352-8_4"},{"volume-title":"Stochastic Approximation and Recursive Algorithms and Applications","year":"2003","author":"Kushner H. J.","key":"e_1_2_11_29_2"}],"container-title":["Journal of Sensors"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/downloads.hindawi.com\/journals\/js\/2021\/8739206.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/js\/2021\/8739206.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1155\/2021\/8739206","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,6]],"date-time":"2024-08-06T00:17:34Z","timestamp":1722903454000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1155\/2021\/8739206"}},"subtitle":[],"editor":[{"given":"Giuseppe","family":"Quero","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2021,1]]},"references-count":29,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,1]]}},"alternative-id":["10.1155\/2021\/8739206"],"URL":"https:\/\/doi.org\/10.1155\/2021\/8739206","archive":["Portico"],"relation":{},"ISSN":["1687-725X","1687-7268"],"issn-type":[{"type":"print","value":"1687-725X"},{"type":"electronic","value":"1687-7268"}],"subject":[],"published":{"date-parts":[[2021,1]]},"assertion":[{"value":"2021-06-02","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-11-16","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-12-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"8739206"}}