{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T21:03:32Z","timestamp":1783803812298,"version":"3.55.0"},"reference-count":51,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2024,1,22]],"date-time":"2024-01-22T00:00:00Z","timestamp":1705881600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Neurorobot."],"abstract":"<jats:p>Target assignment and path planning are crucial for the cooperativity of multiple unmanned aerial vehicles (UAV) systems. However, it is a challenge considering the dynamics of environments and the partial observability of UAVs. In this article, the problem of multi-UAV target assignment and path planning is formulated as a partially observable Markov decision process (POMDP), and a novel deep reinforcement learning (DRL)-based algorithm is proposed to address it. Specifically, a target assignment network is introduced into the twin-delayed deep deterministic policy gradient (TD3) algorithm to solve the target assignment problem and path planning problem simultaneously. The target assignment network executes target assignment for each step of UAVs, while the TD3 guides UAVs to plan paths for this step based on the assignment result and provides training labels for the optimization of the target assignment network. Experimental results demonstrate that the proposed approach can ensure an optimal complete target allocation and achieve a collision-free path for each UAV in three-dimensional (3D) dynamic multiple-obstacle environments, and present a superior performance in target completion and a better adaptability to complex environments compared with existing methods.<\/jats:p>","DOI":"10.3389\/fnbot.2023.1302898","type":"journal-article","created":{"date-parts":[[2024,1,22]],"date-time":"2024-01-22T04:23:58Z","timestamp":1705897438000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":49,"title":["Multi-UAV simultaneous target assignment and path planning based on deep reinforcement learning in dynamic multiple obstacles environments"],"prefix":"10.3389","volume":"17","author":[{"given":"Xiaoran","family":"Kong","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yatong","family":"Zhou","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhe","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shaohai","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2024,1,22]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"270","DOI":"10.1016\/j.comcom.2019.10.014","article-title":"Path planning techniques for unmanned aerial vehicles: a review, solutions, and challenges","volume":"149","author":"Aggarwal","year":"2020","journal-title":"Comput. Commun"},{"key":"B2","doi-asserted-by":"publisher","first-page":"156","DOI":"10.1109\/TSMCC.2007.913919","article-title":"A comprehensive survey of multiagent reinforcement learning","volume":"38","author":"Busoniu","year":"2008","journal-title":"IEEE Trans. Syst. Man. Cybern. C. Appl. Rev"},{"key":"B3","doi-asserted-by":"publisher","first-page":"102324","DOI":"10.1016\/j.adhoc.2020.102324","article-title":"A comprehensive review of unmanned aerial vehicle attacks and neutralization techniques","volume":"111","author":"Chamola","year":"2021","journal-title":"Ad hoc Netw"},{"key":"B4","first-page":"1430","article-title":"\u201cGoal-conditioned reinforcement learning with imagined subgoals,\u201d","volume-title":"International Conference on Machine Learning","author":"Chane-Sane","year":"2021"},{"key":"B5","doi-asserted-by":"publisher","first-page":"38","DOI":"10.25165\/j.ijabe.20211401.5714","article-title":"Review of agricultural spraying technologies for plant protection using unmanned aerial vehicle (UAV)","volume":"14","author":"Chen","year":"2021","journal-title":"Int. J. Agric. Biol. Eng"},{"key":"B6","doi-asserted-by":"publisher","first-page":"119137","DOI":"10.1016\/j.eswa.2022.119137","article-title":"UAV trajectory planning based on bi-directional APF-RRT* algorithm with goal-biased","volume":"213","author":"Fan","year":"2023","journal-title":"Expert Syst. Appl"},{"key":"B7","doi-asserted-by":"publisher","first-page":"19346","DOI":"10.1109\/JIOT.2022.3165278","article-title":"Autonomous cooperative search model for multi-UAV with limited communication network","volume":"9","author":"Fei","year":"2022","journal-title":"IEEE Internet Things J"},{"key":"B8","doi-asserted-by":"publisher","first-page":"108108","DOI":"10.1016\/j.asoc.2021.108108","article-title":"Trajectory planning of autonomous mobile robots applying a particle swarm optimization algorithm with peaks of diversity","volume":"116","author":"Fernandes","year":"2022","journal-title":"Appl. Soft Comput"},{"key":"B9","first-page":"1587","article-title":"\u201cAddressing function approximation error in actor-critic methods,\u201d","volume-title":"International Conference on Machine Learning","author":"Fujimoto","year":"2018"},{"key":"B10","doi-asserted-by":"publisher","first-page":"939","DOI":"10.1177\/0278364904045564","article-title":"A formal analysis and taxonomy of task allocation in multi-robot systems","volume":"23","author":"Gerkey","year":"2004","journal-title":"Int. J. Robot. Res"},{"key":"B11","first-page":"181","article-title":"\u201cA multi-label a* algorithm for multi-agent pathfinding,\u201d","volume-title":"in Proceedings of the International Conference on Automated Planning and Scheduling","author":"Grenouilleau","year":"2019"},{"key":"B12","doi-asserted-by":"crossref","first-page":"448","DOI":"10.1109\/ICRA40945.2020.9197209","article-title":"\u201cCooperative multi-robot navigation in dynamic environment with deep reinforcement learning,\u201d","volume-title":"2020 IEEE International Conference on Robotics and Automation (ICRA)","author":"Han","year":"2020"},{"key":"B13","doi-asserted-by":"publisher","first-page":"107052","DOI":"10.1016\/j.ast.2021.107052","article-title":"Explainable deep reinforcement learning for UAV autonomous path planning","volume":"118","author":"He","year":"2021","journal-title":"Aerosp. Sci. Technol"},{"key":"B14","doi-asserted-by":"publisher","first-page":"7350","DOI":"10.1007\/s10489-020-02082-8","article-title":"A novel hybrid particle swarm optimization for multi-UAV cooperate path planning","volume":"51","author":"He","year":"2021","journal-title":"Appl. Intell"},{"key":"B15","doi-asserted-by":"publisher","first-page":"9725","DOI":"10.1109\/TVT.2021.3102589","article-title":"Energy-efficient online path planning of multiple drones using reinforcement learning","volume":"70","author":"Hong","year":"2021","journal-title":"IEEE Trans. Veh. Technol"},{"key":"B16","doi-asserted-by":"publisher","first-page":"4909","DOI":"10.1109\/TITS.2021.3054625","article-title":"Deep reinforcement learning for autonomous driving: a survey","volume":"23","author":"Kiran","year":"2021","journal-title":"IEEE Trans. Intell. Transp. Syst"},{"key":"B17","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/IROS.2018.8594204","article-title":"\u201cLearning to fly by myself: a self-supervised cnn-based approach for autonomous navigation,\u201d","volume-title":"2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","author":"Kouris","year":"2018"},{"key":"B18","doi-asserted-by":"publisher","first-page":"83","DOI":"10.1002\/nav.3800020109","article-title":"The hungarian method for the assignment problem","volume":"2","author":"Kuhn","year":"1955","journal-title":"Nav. Res. Logist. Q"},{"key":"B19","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1109\/TSMCB.2003.808174","article-title":"Efficiently solving general weapon-target assignment problem by genetic algorithms with greedy eugenics","volume":"33","author":"Lee","year":"2003","journal-title":"IEEE Trans. Syst. Man Cybernet. B"},{"key":"B20","doi-asserted-by":"publisher","first-page":"826","DOI":"10.3390\/jmse10060826","article-title":"Improved rrt algorithm for auv target search in unknown 3d environment","volume":"10","author":"Li","year":"2022","journal-title":"J. Mar. Sci. Eng"},{"key":"B21","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1509.02971","article-title":"Continuous control with deep reinforcement learning","author":"Lillicrap","year":"2015","journal-title":"arXiv"},{"key":"B22","doi-asserted-by":"publisher","first-page":"10676","DOI":"10.1109\/JIOT.2021.3125784","article-title":"Cooperative path optimization for multiple uavs surveillance in uncertain environment","volume":"9","author":"Liu","year":"2021","journal-title":"IEEE Internet Things J"},{"key":"B23","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s00500-023-07981-9","article-title":"Location and tracking of environmental pollution sources under multi-UAV vision based on target motion model","volume":"27","author":"Liu","year":"2023","journal-title":"Soft Comput"},{"key":"B24","first-page":"6379","article-title":"\u201cMulti-agent actor-critic for mixed cooperative? competitive environments,\u201d","volume-title":"31st International Conference on Neural Information Processing Systems","author":"Lowe","year":"2017"},{"key":"B25","doi-asserted-by":"publisher","first-page":"4426","DOI":"10.1109\/TSMC.2021.3096997","article-title":"Learning-based policy optimization for adversarial missile-target assignment","volume":"52","author":"Luo","year":"2021","journal-title":"IEEE Trans. Syst. Man Cybernet. Syst"},{"key":"B26","doi-asserted-by":"publisher","first-page":"3266","DOI":"10.3390\/rs15133266","article-title":"Unmanned aerial vehicles for search and rescue: a survey","volume":"15","author":"Lyu","year":"2023","journal-title":"Remote Sens"},{"key":"B27","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2020.103472","article-title":"Deploying mavs for autonomous navigation in dark underground mine environments","author":"Mansouri","year":"2020","journal-title":"Robot. Auton. Syst"},{"key":"B28","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"B29","doi-asserted-by":"publisher","first-page":"7994","DOI":"10.1109\/ACCESS.2021.3049892","article-title":"A deep learning trained by genetic algorithm to improve the efficiency of path planning for data collection with multi-UAV","volume":"9","author":"Pan","year":"2021","journal-title":"IEEE Access"},{"key":"B30","doi-asserted-by":"publisher","first-page":"146264","DOI":"10.1109\/ACCESS.2019.2943253","article-title":"Joint optimization of multi-UAV target assignment and path planning based on multi-agent reinforcement learning","volume":"7","author":"Qie","year":"2019","journal-title":"IEEE Access"},{"key":"B31","doi-asserted-by":"publisher","first-page":"17290","DOI":"10.1109\/JIOT.2021.3078746","article-title":"Task selection and scheduling in UAV-enabled mec for reconnaissance with time-varying priorities","volume":"8","author":"Qin","year":"2021","journal-title":"IEEE Internet of Things Journal"},{"key":"B32","doi-asserted-by":"crossref","first-page":"148","DOI":"10.1109\/NAECON46414.2019.9057847","article-title":"\u201cCluster-based hungarian approach to task allocation for unmanned aerial vehicles\u201d","volume-title":"2019 IEEE National Aerospace and Electronics Conference (NAECON)","author":"Samiei","year":"2019"},{"key":"B33","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1707.06347","article-title":"Proximal policy optimization algorithms","author":"Schulman","year":"2017","journal-title":"arXiv"},{"key":"B34","doi-asserted-by":"publisher","first-page":"208","DOI":"10.3390\/aerospace10030208","article-title":"Survey on mission planning of multiple unmanned aerial vehicles","volume":"10","author":"Song","year":"2023","journal-title":"Aerospace"},{"key":"B35","doi-asserted-by":"crossref","first-page":"387","DOI":"10.1007\/978-3-642-27645-3_12","article-title":"Partially observable markov decision processes","volume-title":"Reinforcement learning: State-of-the-art","author":"Spaan","year":"2012"},{"key":"B36","doi-asserted-by":"publisher","first-page":"5490","DOI":"10.1080\/01431161.2018.1441570","article-title":"Using an unmanned aerial vehicle (UAV) to study wild yak in the highest desert in the world","volume":"39","author":"Su","year":"2018","journal-title":"Int. J. Remote Sens"},{"key":"B37","doi-asserted-by":"publisher","first-page":"103","DOI":"10.1609\/aimag.v21i1.1501","article-title":"Reinforcement learning: an introduction","volume":"21","author":"Thrun","year":"2000","journal-title":"AI. Mag"},{"key":"B38","doi-asserted-by":"crossref","first-page":"304","DOI":"10.1109\/CCSSE.2018.8724841","article-title":"\u201cResearch on target assignment of multiple-uavs based on improved hybrid genetic algorithm,\u201d","volume-title":"2018 IEEE 4th International Conference on Control Science and Systems Engineering (ICCSSE)","author":"Tian","year":"2018"},{"key":"B39","doi-asserted-by":"publisher","first-page":"6180","DOI":"10.1109\/JIOT.2020.2973193","article-title":"Deep-reinforcement-learning-based autonomous UAV navigation with sparse rewards","volume":"7","author":"Wang","year":"2020","journal-title":"IEEE Internet. Things. J"},{"key":"B40","doi-asserted-by":"crossref","first-page":"1647","DOI":"10.1109\/ITOEC49072.2020.9141873","article-title":"\u201cCooperative coverage reconnaissance of multi- UAV,\u201d","volume-title":"2020 IEEE 5th Information Technology and Mechatronics Engineering Conference (ITOEC)","author":"Wang","year":"2020"},{"key":"B41","doi-asserted-by":"publisher","first-page":"3362","DOI":"10.3934\/jimo.2022089","article-title":"A mini review on UAV mission planning","volume":"19","author":"Wang","year":"2023","journal-title":"J. Ind. Manag. Optim"},{"key":"B42","doi-asserted-by":"publisher","first-page":"3680","DOI":"10.1109\/TNNLS.2021.3116063","article-title":"Deep reinforcement learning on autonomous driving policy with auxiliary critic network","volume":"34","author":"Wu","year":"2021","journal-title":"IEEE Trans. Neural. Netw. Learn. Syst"},{"key":"B43","doi-asserted-by":"publisher","first-page":"102972","DOI":"10.1016\/j.ijdrr.2022.102972","article-title":"Multi-UAV cooperative system for search and rescue based on YOLOv5","volume":"76","author":"Xing","year":"2022","journal-title":"Int. J. Disaster Risk Sci"},{"key":"B44","doi-asserted-by":"publisher","first-page":"104938","DOI":"10.1016\/j.compag.2019.104938","article-title":"Online spraying quality assessment system of plant protection unmanned aerial vehicle based on android client","volume":"166","author":"Xu","year":"2019","journal-title":"Comput. Electron. Agric"},{"key":"B45","doi-asserted-by":"publisher","first-page":"789","DOI":"10.1109\/TASE.2022.3168621","article-title":"Unified automatic control of vehicular systems with reinforcement learning","volume":"20","author":"Yan","year":"2022","journal-title":"IEEE Trans. Autom. Sci. Eng"},{"key":"B46","doi-asserted-by":"publisher","first-page":"155939","DOI":"10.1016\/j.scitotenv.2022.155939","article-title":"UAV remote sensing applications in marine monitoring: knowledge visualization and review","volume":"838","author":"Yang","year":"2022","journal-title":"Sci. Total Environ"},{"key":"B47","doi-asserted-by":"publisher","first-page":"1105480","DOI":"10.3389\/fnbot.2022.1105480","article-title":"Research on reinforcement learning-based safe decision-making methodology for multiple unmanned aerial vehicles","volume":"16","author":"Yue","year":"2023","journal-title":"Front. Neurorobot"},{"key":"B48","doi-asserted-by":"publisher","first-page":"4091","DOI":"10.1109\/TIE.2016.2542134","article-title":"Data-driven optimal consensus control for discrete-time multi-agent systems with unknown dynamics using reinforcement learning method","volume":"64","author":"Zhang","year":"2016","journal-title":"IEEE Trans. Ind. Electron"},{"key":"B49","doi-asserted-by":"publisher","first-page":"1221","DOI":"10.3390\/rs13061221","article-title":"A review of unmanned aerial vehicle low-altitude remote sensing (UAV-LARS) use in agricultural monitoring in china","volume":"13","author":"Zhang","year":"2021","journal-title":"Remote Sens"},{"key":"B50","doi-asserted-by":"publisher","first-page":"108194","DOI":"10.1016\/j.asoc.2021.108194","article-title":"Autonomous navigation of UAV in multi-obstacle environments based on a deep reinforcement learning approach","volume":"115","author":"Zhang","year":"2022","journal-title":"Appl. Soft. Comput"},{"key":"B51","doi-asserted-by":"publisher","first-page":"1243174","DOI":"10.3389\/fnbot.2023.1243174","article-title":"MW-MADDPG: a meta-learning based decision-making method for collaborative UAV swarm","volume":"17","author":"Zhao","year":"2023","journal-title":"Front. Neurorobot"}],"container-title":["Frontiers in Neurorobotics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2023.1302898\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,1,22]],"date-time":"2024-01-22T04:24:16Z","timestamp":1705897456000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2023.1302898\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,22]]},"references-count":51,"alternative-id":["10.3389\/fnbot.2023.1302898"],"URL":"https:\/\/doi.org\/10.3389\/fnbot.2023.1302898","relation":{},"ISSN":["1662-5218"],"issn-type":[{"value":"1662-5218","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,22]]},"article-number":"1302898"}}