{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,5,24]],"date-time":"2025-05-24T04:01:37Z","timestamp":1748059297741,"version":"3.41.0"},"reference-count":42,"publisher":"World Scientific Pub Co Pte Ltd","issue":"04","funder":[{"name":"National Key Research and Development Plan of China","award":["2018AAA0101000"],"award-info":[{"award-number":["2018AAA0101000"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62076028"],"award-info":[{"award-number":["62076028"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Un. Sys."],"published-print":{"date-parts":[[2025,7]]},"abstract":"<jats:p> Aiming at the problem of multi-agent cooperative confrontation in seize-control scenarios, we design an efficient multi-agent policy self-play (EMAP-SP) learning method. First, a multi-agent centralized policy model is constructed to command the agents to perform tasks cooperatively. Considering that the policy being trained and its historical policies usually have poor exploration capability under incomplete information in self-play trainings, the intrinsic reward mechanism based on random network distillation (RND) is introduced in the self-play learning method. In addition, we propose a multi-step on-policy deep reinforcement learning (DRL) algorithm assisted by off-policy policy evaluation (MSOAO) to learn the best response policy in the self-play. Compared with DRL algorithms commonly used in complex decision problems, MSOAO has more efficient policy evaluation capability, and efficient policy evaluation further improves the policy learning capability. The effectiveness of EMAP-SP is fully verified in MiaoSuan wargame simulation system, and the evaluation results show that EMAP-SP can learn the cooperative policy of effectively defeating the Blue side\u2019s knowledge-based policy under incomplete information. Moreover, the evaluations results in DRL benchmark environments also show that the best response policy learning algorithm MSOAO can promote the agent to learn approximately optimal policies. <\/jats:p>","DOI":"10.1142\/s230138502550061x","type":"journal-article","created":{"date-parts":[[2024,8,11]],"date-time":"2024-08-11T03:43:36Z","timestamp":1723347816000},"page":"987-1004","source":"Crossref","is-referenced-by-count":0,"title":["An Efficient Multi-Agent Policy Self-Play Learning Method Aiming at Seize-Control Scenarios"],"prefix":"10.1142","volume":"13","author":[{"given":"Huaqing","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Automation, Beijing Institute of Technology, Beijing 100081, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5734-3157","authenticated-orcid":false,"given":"Hongbin","family":"Ma","sequence":"additional","affiliation":[{"name":"School of Automation, Beijing Institute of Technology, Beijing 100081, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaofei","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Vehicle and Mobility, Tsinghua University, Beijing 100084, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Li","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Mechanical Engineering, Beijing Institute of Technology, Beijing 100081, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Minglei","family":"Han","sequence":"additional","affiliation":[{"name":"School of Automation, Beijing Institute of Technology, Beijing 100081, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hui","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Automation, Beijing Institute of Technology, Beijing 100081, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ao","family":"Ding","sequence":"additional","affiliation":[{"name":"School of Automation, Beijing Institute of Technology, Beijing 100081, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2024,9,30]]},"reference":[{"key":"S230138502550061XBIB001","doi-asserted-by":"publisher","DOI":"10.1016\/j.ast.2016.03.022"},{"key":"S230138502550061XBIB002","doi-asserted-by":"publisher","DOI":"10.1016\/j.cie.2018.05.013"},{"key":"S230138502550061XBIB003","doi-asserted-by":"publisher","DOI":"10.1016\/j.comcom.2023.07.006"},{"key":"S230138502550061XBIB004","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS52030.2021.00059"},{"key":"S230138502550061XBIB005","doi-asserted-by":"publisher","DOI":"10.1109\/JAS.2022.106007"},{"key":"S230138502550061XBIB006","doi-asserted-by":"publisher","DOI":"10.1142\/S230138502450002X"},{"key":"S230138502550061XBIB007","doi-asserted-by":"publisher","DOI":"10.1142\/S2301385024500122"},{"key":"S230138502550061XBIB008","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-019-1724-z"},{"key":"S230138502550061XBIB009","volume":"33","author":"McAleer S.","year":"2020","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"S230138502550061XBIB010","first-page":"20230","volume":"36","author":"MacQueen R.","year":"2023","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"S230138502550061XBIB011","doi-asserted-by":"publisher","DOI":"10.1109\/CoG51982.2022.9893656"},{"key":"S230138502550061XBIB012","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-10-0575-6_6"},{"key":"S230138502550061XBIB013","volume":"30","author":"Lowe R.","year":"2017","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"S230138502550061XBIB015","first-page":"24611","volume":"35","author":"Yu C.","year":"2022","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"S230138502550061XBIB016","doi-asserted-by":"publisher","DOI":"10.3390\/electronics12112396"},{"key":"S230138502550061XBIB017","volume":"30","author":"Lanctot M.","year":"2017","journal-title":"Adv. Neural Inform. Process. Syst."},{"first-page":"434","volume-title":"Int. Conf. Machine Learning (PMLR, 2019)","author":"Balduzzi D.","key":"S230138502550061XBIB018"},{"key":"S230138502550061XBIB019","first-page":"3504","volume":"34","author":"Feng X.","year":"2021","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"S230138502550061XBIB020","first-page":"23128","volume":"34","author":"McAleer S.","year":"2021","journal-title":"Adv. Neural Inform. Process. Syst."},{"journal-title":"Trans. Mach. Learn. Res.","year":"2022","author":"Dinh L. C.","key":"S230138502550061XBIB021"},{"key":"S230138502550061XBIB023","doi-asserted-by":"publisher","DOI":"10.1016\/j.aei.2021.101360"},{"key":"S230138502550061XBIB024","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2021.3129160"},{"key":"S230138502550061XBIB026","first-page":"1407","volume-title":"Int. Conf. Machine Learning","author":"Espeholt L.","year":"2018"},{"volume-title":"AAAI Fall Symp. Series","year":"2015","author":"Hausknecht M.","key":"S230138502550061XBIB027"},{"key":"S230138502550061XBIB028","volume":"12","author":"Sutton R. S.","year":"1999","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"S230138502550061XBIB029","first-page":"1889","volume-title":"Int. Conf. Machine Learning","author":"Schulman J.","year":"2015"},{"key":"S230138502550061XBIB030","first-page":"1928","volume-title":"Int. Conf. Machine Learning","author":"Mnih V.","year":"2016"},{"key":"S230138502550061XBIB031","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"S230138502550061XBIB033","volume":"32","author":"Assran M.","year":"2019","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"S230138502550061XBIB034","first-page":"1587","volume-title":"Int. Conf. Machine Learning","author":"Fujimoto S.","year":"2018"},{"key":"S230138502550061XBIB035","doi-asserted-by":"publisher","DOI":"10.1016\/B978-1-55860-335-6.50035-0"},{"key":"S230138502550061XBIB038","doi-asserted-by":"publisher","DOI":"10.1016\/j.tins.2010.01.006"},{"issue":"1","key":"S230138502550061XBIB039","volume":"30","author":"Van Hasselt H.","year":"2016","journal-title":"Proc. AAAI Conf. Artificial Intelli."},{"key":"S230138502550061XBIB041","doi-asserted-by":"publisher","DOI":"10.1007\/s11704-020-9307-6"},{"key":"S230138502550061XBIB042","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2021.04.006"},{"key":"S230138502550061XBIB043","doi-asserted-by":"publisher","DOI":"10.1016\/j.geb.2005.08.005"},{"key":"S230138502550061XBIB044","first-page":"805","volume-title":"Int. Conf. Machine Learning","author":"Heinrich J.","year":"2015"},{"key":"S230138502550061XBIB045","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-27645-3_12"},{"key":"S230138502550061XBIB048","volume":"30","author":"Vaswani A.","year":"2017","journal-title":"Adv. Neural Inform. Process. Syst."},{"key":"S230138502550061XBIB049","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"S230138502550061XBIB050","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-13841-6_41"},{"key":"S230138502550061XBIB052","doi-asserted-by":"publisher","DOI":"10.1613\/jair.5699"}],"container-title":["Unmanned Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S230138502550061X","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,23]],"date-time":"2025-05-23T04:17:40Z","timestamp":1747973860000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/10.1142\/S230138502550061X"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,30]]},"references-count":42,"journal-issue":{"issue":"04","published-print":{"date-parts":[[2025,7]]}},"alternative-id":["10.1142\/S230138502550061X"],"URL":"https:\/\/doi.org\/10.1142\/s230138502550061x","relation":{},"ISSN":["2301-3850","2301-3869"],"issn-type":[{"type":"print","value":"2301-3850"},{"type":"electronic","value":"2301-3869"}],"subject":[],"published":{"date-parts":[[2024,9,30]]}}}