{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T16:04:43Z","timestamp":1753891483658,"version":"3.41.2"},"reference-count":32,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2022,12,1]],"date-time":"2022-12-01T00:00:00Z","timestamp":1669852800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Neurorobot."],"abstract":"<jats:p>Modern air defense battlefield situations are complex and varied, requiring high-speed computing capabilities and real-time situational processing for task assignment. Current methods struggle to balance the quality and speed of assignment strategies. This paper proposes a hierarchical reinforcement learning architecture for ground-to-air confrontation (HRL-GC) and an algorithm combining model predictive control with proximal policy optimization (MPC-PPO), which effectively combines the advantages of centralized and distributed approaches. To improve training efficiency while ensuring the quality of the final decision. In a large-scale area air defense scenario, this paper validates the effectiveness and superiority of the HRL-GC architecture and MPC-PPO algorithm, proving that the method can meet the needs of large-scale air defense task assignment in terms of quality and speed.<\/jats:p>","DOI":"10.3389\/fnbot.2022.1072887","type":"journal-article","created":{"date-parts":[[2022,12,1]],"date-time":"2022-12-01T05:46:45Z","timestamp":1669873605000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Intelligent air defense task assignment based on hierarchical reinforcement learning"],"prefix":"10.3389","volume":"16","author":[{"given":"Jia-yi","family":"Liu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gang","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiang-ke","family":"Guo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Si-yuan","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qiang","family":"Fu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2022,12,1]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33019876","article-title":"A theory of state abstraction for reinforcement learning","author":"Abel","year":"2019","journal-title":"Proceedings of the 33rd AAAI conference on artificial intelligence"},{"key":"B2","doi-asserted-by":"publisher","first-page":"55","DOI":"10.1007\/s11768-015-3203-x","article-title":"Discrete-time dynamic graphical games: Model-free reinforcement learning solution.","volume":"13","author":"Abouheaf","year":"2015","journal-title":"Control Theory Technol."},{"key":"B3","doi-asserted-by":"publisher","DOI":"10.1007\/S10915-022-01876-X","article-title":"Sojourn-based approach to semi-markov reinforcement learning.","volume":"92","author":"Ascione","year":"2022","journal-title":"J. Sci. Comput."},{"key":"B4","first-page":"39","article-title":"Constructing temporal abstractions autonomously in reinforcement learning.","volume":"39","author":"Bacon","year":"2018","journal-title":"AI Mag."},{"key":"B5","doi-asserted-by":"publisher","first-page":"383","DOI":"10.20965\/jaciii.2008.p0383","article-title":"Trading rules on stock markets using genetic network programming with sarsa learning.","volume":"12","author":"Chen","year":"2008","journal-title":"J. Adv. Comput. Intell. Intell. Inform."},{"key":"B6","doi-asserted-by":"publisher","first-page":"365","DOI":"10.1016\/j.ins.2022.01.047","article-title":"Actor-critic continuous state reinforcement learning for wind-turbine control robust optimization.","volume":"591","author":"Fernandez-Gauna","year":"2022","journal-title":"Inf. Sci."},{"key":"B7","doi-asserted-by":"publisher","first-page":"87504","DOI":"10.1109\/ACCESS.2020.2993459","article-title":"Alpha C2\u2013an intelligent air defense commander independent of human decision-making.","volume":"8","author":"Fu","year":"2020","journal-title":"IEEE Access"},{"key":"B8","article-title":"Continuous deep Q-learning with model-based acceleration","author":"Gu","year":"2016","journal-title":"Proceedings of the 33rd international conference on international conference on machine learning"},{"key":"B9","first-page":"988","article-title":"Distributed task assignment algorithm for SEAD mission of heterogeneous UAVs based on CBBA algorithm.","volume":"40","author":"Lee","year":"2012","journal-title":"J. Korean Soc. Aeronaut. Space Sci."},{"key":"B10","doi-asserted-by":"publisher","DOI":"10.1016\/J.IS.2021.101772","article-title":"Deep reinforcement learning based ensemble model for rumor tracking.","volume":"103","author":"Li","year":"2022","journal-title":"Inf. Syst."},{"key":"B11","doi-asserted-by":"publisher","DOI":"10.1016\/j.dt.2022.04.001","article-title":"Task assignment in ground-to-air confrontation based on multiagent deep reinforcement learning.","author":"Liu","year":"2022","journal-title":"Def. Technol."},{"key":"B12","doi-asserted-by":"publisher","DOI":"10.1016\/J.TRC.2021.103261","article-title":"A scenario-based distributed model predictive control approach for freeway networks.","volume":"136","author":"Liu","year":"2022","journal-title":"Transp. Res. C"},{"journal-title":"Theory of neural-analog reinforcement systems and its application to the brain-model problem.","year":"1954","author":"Minsky","key":"B13"},{"key":"B14","doi-asserted-by":"publisher","first-page":"276","DOI":"10.3390\/MAKE4010013","article-title":"Robust reinforcement learning: A review of foundations and recent advances.","volume":"4","author":"Moos","year":"2022","journal-title":"Mach. Learn. Knowl. Extr."},{"key":"B15","doi-asserted-by":"crossref","DOI":"10.1109\/ICCKE.2016.7802135","article-title":"A centralized reinforcement learning method for multi-agent job scheduling in grid","author":"Moradi","year":"2016","journal-title":"Proceedings of the 6th international conference on computer and knowledge engineering (ICCKE 2016)"},{"key":"B16","doi-asserted-by":"crossref","first-page":"7559","DOI":"10.1109\/ICRA.2018.8463189","article-title":"Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning","author":"Nagabandi","year":"2018","journal-title":"Proceedings of the 2018 IEEE international conference on robotics and automation (ICRA)"},{"key":"B17","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1080\/03071846709429752","article-title":"Modern air defence: A lecture given at the RUSI on 14th December 1966.","volume":"112","author":"Rosier","year":"2009","journal-title":"R. U. Serv. Inst. J."},{"key":"B18","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/2050\/1\/012012","article-title":"Deep reinforcement learning for stock recommendation.","volume":"2050","author":"Shen","year":"2021","journal-title":"J. Phys."},{"key":"B19","doi-asserted-by":"crossref","first-page":"1549","DOI":"10.1016\/j.ifacol.2020.12.2021","article-title":"A multi-agent off-policy actor-critic algorithm for distributed reinforcement learning.","volume":"53","author":"Suttle","year":"2020","journal-title":"IFAC PapersOnLine"},{"key":"B20","article-title":"Multi-controller fusion in multi-layered reinforcement learning","author":"Takahashi","year":"2001","journal-title":"Proceedings of the conference documentation international conference on multisensor fusion and integration for intelligent systems. MFI 2001 (Cat. No.01TH8590)"},{"key":"B21","doi-asserted-by":"publisher","DOI":"10.1088\/1757-899X\/493\/1\/012056","article-title":"Research on mission assignment assurance of remote rocket barrage based on stackelberg game","author":"Wang","year":"2019","journal-title":"Proceedings of the 2nd international conference on frontiers of materials synthesis and processing IOP conference series: Materials science and engineering"},{"key":"B22","doi-asserted-by":"publisher","first-page":"279","DOI":"10.1007\/BF00992698","article-title":"Q-learning.","volume":"8","author":"Watkins","year":"1992","journal-title":"Mach. Learn."},{"key":"B23","doi-asserted-by":"crossref","DOI":"10.1142\/S0218001419510108","article-title":"Explore deep neural network and reinforcement learning to large-scale tasks processing in big data.","volume":"33","author":"Wu","year":"2019","journal-title":"Int. J. Pattern Recognit. Artif. Intell."},{"key":"B24","doi-asserted-by":"publisher","first-page":"104","DOI":"10.6911\/WSRJ.202201_8(1).0017","article-title":"Research on multi-UAV task assignment method based on reinforcement learning.","volume":"8","author":"Wu","year":"2022","journal-title":"World Sci. Res. J."},{"key":"B25","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1177\/1548512918809514","article-title":"Modeling of situation assessment in regional air defense combat.","volume":"16","author":"Yang","year":"2019","journal-title":"J. Def. Model. Simul. Appl. Methodol. Technol."},{"key":"B26","doi-asserted-by":"publisher","first-page":"699","DOI":"10.1016\/J.IFACOL.2021.08.323","article-title":"Multi-step greedy reinforcement learning based on model predictive control.","volume":"54","author":"Yang","year":"2021","journal-title":"IFAC PapersOnLine"},{"key":"B27","doi-asserted-by":"publisher","DOI":"10.26991\/d.cnki.gdllu.2021.000998","author":"Yaqi","year":"2021","journal-title":"Tensegrity robot locomotion control via reinforcement learning."},{"key":"B28","doi-asserted-by":"publisher","first-page":"102335","DOI":"10.1109\/access.2020.2997304","article-title":"IADRL: Imitation augmented deep reinforcement learning enabled UGV-UAV coalition for tasking in complex environments.","volume":"8","author":"Zhang","year":"2020","journal-title":"IEEE Access"},{"key":"B29","doi-asserted-by":"publisher","DOI":"10.1016\/J.ASOC.2021.108194","article-title":"Autonomous navigation of UAV in multi-obstacle environments based on a deep reinforcement learning approach.","volume":"115","author":"Zhang","year":"2022","journal-title":"Appl. Soft Comput. J."},{"key":"B30","doi-asserted-by":"publisher","DOI":"10.3390\/APP11188419","article-title":"End-to-end deep reinforcement learning for image-based UAV autonomous control.","volume":"11","author":"Zhao","year":"2021","journal-title":"Appl. Sci."},{"key":"B31","doi-asserted-by":"publisher","first-page":"18","DOI":"10.1016\/J.PATREC.2021.08.019","article-title":"A model-based reinforcement learning method based on conditional generative adversarial networks.","volume":"152","author":"Zhao","year":"2021","journal-title":"Pattern Recognit. Lett."},{"key":"B32","doi-asserted-by":"publisher","first-page":"588","DOI":"10.1016\/j.ast.2019.06.024","article-title":"Fast task allocation for heterogeneous unmanned aerial vehicles through reinforcement learning.","volume":"92","author":"Zhao","year":"2019","journal-title":"Aerosp. Sci. Technol."}],"container-title":["Frontiers in Neurorobotics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2022.1072887\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,1]],"date-time":"2022-12-01T05:46:55Z","timestamp":1669873615000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2022.1072887\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,1]]},"references-count":32,"alternative-id":["10.3389\/fnbot.2022.1072887"],"URL":"https:\/\/doi.org\/10.3389\/fnbot.2022.1072887","relation":{},"ISSN":["1662-5218"],"issn-type":[{"type":"electronic","value":"1662-5218"}],"subject":[],"published":{"date-parts":[[2022,12,1]]},"article-number":"1072887"}}