{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T14:43:47Z","timestamp":1740149027664,"version":"3.37.3"},"reference-count":23,"publisher":"Wiley","license":[{"start":{"date-parts":[[2021,9,3]],"date-time":"2021-09-03T00:00:00Z","timestamp":1630627200000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61973002","61902104","2008085J32","2008085QF295"],"award-info":[{"award-number":["61973002","61902104","2008085J32","2008085QF295"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003995","name":"Natural Science Foundation of Anhui Province","doi-asserted-by":"publisher","award":["61973002","61902104","2008085J32","2008085QF295"],"award-info":[{"award-number":["61973002","61902104","2008085J32","2008085QF295"]}],"id":[{"id":"10.13039\/501100003995","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Security and Communication Networks"],"published-print":{"date-parts":[[2021,9,3]]},"abstract":"<jats:p>The main objective of multiagent reinforcement learning is to achieve a global optimal policy. It is difficult to evaluate the value function with high-dimensional state space. Therefore, we transfer the problem of multiagent reinforcement learning into a distributed optimization problem with constraint terms. In this problem, all agents share the space of states and actions, but each agent only obtains its own local reward. Then, we propose a distributed optimization with fractional order dynamics to solve this problem. Moreover, we prove the convergence of the proposed algorithm and illustrate its effectiveness with a numerical example.<\/jats:p>","DOI":"10.1155\/2021\/1020466","type":"journal-article","created":{"date-parts":[[2021,9,6]],"date-time":"2021-09-06T19:05:19Z","timestamp":1630955119000},"page":"1-7","source":"Crossref","is-referenced-by-count":0,"title":["Distributed Policy Evaluation with Fractional Order Dynamics in Multiagent Reinforcement Learning"],"prefix":"10.1155","volume":"2021","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3783-7301","authenticated-orcid":true,"given":"Wei","family":"Dai","sequence":"first","affiliation":[{"name":"Anhui Engineering Laboratory of Human-Robot Integration System and Intelligent Equipment, School of Electrical Engineering and Automation, Anhui University, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4126-3571","authenticated-orcid":true,"given":"Wei","family":"Wang","sequence":"additional","affiliation":[{"name":"Center for Assessment and Demonstration Research, Academy of Military Science, Beijing 100091, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8471-2544","authenticated-orcid":true,"given":"Zhongtian","family":"Mao","sequence":"additional","affiliation":[{"name":"Anhui Engineering Laboratory of Human-Robot Integration System and Intelligent Equipment, School of Electrical Engineering and Automation, Anhui University, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2681-3777","authenticated-orcid":true,"given":"Ruwen","family":"Jiang","sequence":"additional","affiliation":[{"name":"Anhui Engineering Laboratory of Human-Robot Integration System and Intelligent Equipment, School of Electrical Engineering and Automation, Anhui University, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9604-7564","authenticated-orcid":true,"given":"Fudong","family":"Nian","sequence":"additional","affiliation":[{"name":"Hefei University, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0111-0108","authenticated-orcid":true,"given":"Teng","family":"Li","sequence":"additional","affiliation":[{"name":"Anhui Engineering Laboratory of Human-Robot Integration System and Intelligent Equipment, School of Electrical Engineering and Automation, Anhui University, Hefei 230601, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","reference":[{"volume-title":"Reinforcement Learning: An Introduction","year":"1998","author":"R. S. Sutton","key":"1"},{"key":"2","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913495721"},{"key":"3","doi-asserted-by":"publisher","DOI":"10.1109\/tsmcc.2007.913919"},{"key":"4","doi-asserted-by":"publisher","DOI":"10.1109\/tac.2020.2979274"},{"key":"5","doi-asserted-by":"publisher","DOI":"10.1109\/tac.2016.2638966"},{"key":"6","doi-asserted-by":"publisher","DOI":"10.1109\/cdc.2018.8619581"},{"key":"7","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-71682-4_5"},{"key":"8","doi-asserted-by":"publisher","DOI":"10.1007\/s11768-020-00007-x"},{"key":"9","doi-asserted-by":"publisher","DOI":"10.1109\/tac.2008.2009515"},{"key":"10","doi-asserted-by":"publisher","DOI":"10.1109\/tnnls.2019.2951790"},{"key":"11","doi-asserted-by":"publisher","DOI":"10.1016\/j.automatica.2020.109289"},{"key":"12","doi-asserted-by":"publisher","DOI":"10.1016\/j.automatica.2016.07.010"},{"key":"13","doi-asserted-by":"publisher","DOI":"10.1016\/j.arcontrol.2019.05.006"},{"key":"14","doi-asserted-by":"publisher","DOI":"10.1109\/lcsys.2020.3037038"},{"key":"15","doi-asserted-by":"publisher","DOI":"10.1109\/tac.2009.2033738"},{"key":"16","doi-asserted-by":"publisher","DOI":"10.1016\/j.automatica.2018.10.028"},{"key":"17","doi-asserted-by":"publisher","DOI":"10.1016\/j.cnsns.2018.12.023"},{"key":"18","doi-asserted-by":"publisher","DOI":"10.1016\/j.sigpro.2016.11.026"},{"key":"19","doi-asserted-by":"publisher","DOI":"10.1007\/s11768-021-00044-0"},{"article-title":"Spectral graph theory and network dependability","author":"A. Torres","key":"20","doi-asserted-by":"crossref","DOI":"10.1109\/DepCoS-RELCOMEX.2009.52"},{"key":"21","doi-asserted-by":"publisher","DOI":"10.1007\/s11760-012-0332-2"},{"key":"22","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-84996-335-0","volume-title":"Fractionalorder Systems and Controls: Fundamentals and Applications","author":"C. Monje","year":"2010"},{"key":"23","doi-asserted-by":"publisher","DOI":"10.1145\/3314578"}],"container-title":["Security and Communication Networks"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/downloads.hindawi.com\/journals\/scn\/2021\/1020466.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/scn\/2021\/1020466.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/scn\/2021\/1020466.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,6]],"date-time":"2021-09-06T19:05:24Z","timestamp":1630955124000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.hindawi.com\/journals\/scn\/2021\/1020466\/"}},"subtitle":[],"editor":[{"given":"Zhenhua","family":"Tan","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2021,9,3]]},"references-count":23,"alternative-id":["1020466","1020466"],"URL":"https:\/\/doi.org\/10.1155\/2021\/1020466","relation":{},"ISSN":["1939-0122","1939-0114"],"issn-type":[{"type":"electronic","value":"1939-0122"},{"type":"print","value":"1939-0114"}],"subject":[],"published":{"date-parts":[[2021,9,3]]}}}