{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T07:24:31Z","timestamp":1778570671375,"version":"3.51.4"},"reference-count":27,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2021,3,23]],"date-time":"2021-03-23T00:00:00Z","timestamp":1616457600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the seed Foundation of Innovation and Creation for Graduate Students in Northwestern Polytechnical University","award":["ZZ2019021"],"award-info":[{"award-number":["ZZ2019021"]}]},{"name":"the Natural Science Basic Research Program of Shaanxi Program","award":["No. 2020JM-147"],"award-info":[{"award-number":["No. 2020JM-147"]}]},{"name":"the Key Laboratory Project Foundation","award":["6142504190105"],"award-info":[{"award-number":["6142504190105"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>How to operate an unmanned aerial vehicle (UAV) safely and efficiently in an interactive environment is challenging. A large amount of research has been devoted to improve the intelligence of a UAV while performing a mission, where finding an optimal maneuver decision-making policy of the UAV has become one of the key issues when we attempt to enable the UAV autonomy. In this paper, we propose a maneuver decision-making algorithm based on deep reinforcement learning, which generates efficient maneuvers for a UAV agent to execute the airdrop mission autonomously in an interactive environment. Particularly, the training set of the learning algorithm by the Prioritized Experience Replay is constructed, that can accelerate the convergence speed of decision network training in the algorithm. It is shown that a desirable and effective maneuver decision-making policy can be found by extensive experimental results.<\/jats:p>","DOI":"10.3390\/s21062233","type":"journal-article","created":{"date-parts":[[2021,3,23]],"date-time":"2021-03-23T23:59:41Z","timestamp":1616543981000},"page":"2233","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":14,"title":["A UAV Maneuver Decision-Making Algorithm for Autonomous Airdrop Based on Deep Reinforcement Learning"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0192-5534","authenticated-orcid":false,"given":"Ke","family":"Li","sequence":"first","affiliation":[{"name":"School of Electronics and Information, Northwestern Polytechnical University, Xi\u2019an 710072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kun","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Electronics and Information, Northwestern Polytechnical University, Xi\u2019an 710072, China"},{"name":"Science and Technology on Electro-Optic Control Laboratory, Luoyang 471009, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhenchong","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Electronics and Information, Northwestern Polytechnical University, Xi\u2019an 710072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zekun","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Electronics and Information, Northwestern Polytechnical University, Xi\u2019an 710072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuai","family":"Hua","sequence":"additional","affiliation":[{"name":"School of Electronics and Information, Northwestern Polytechnical University, Xi\u2019an 710072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianliang","family":"He","sequence":"additional","affiliation":[{"name":"Science and Technology on Electro-Optic Control Laboratory, Luoyang 471009, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,3,23]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1109\/MCOM.2017.1600238CM","article-title":"UAV-enabled intelligent transportation systems for the smart city: Applications and challenges","volume":"55","author":"Menouar","year":"2017","journal-title":"IEEE Commun. Mag."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"859","DOI":"10.1007\/s10514-020-09902-3","article-title":"Autonomous ballistic airdrop of objects from a small fixed-wing unmanned aerial vehicle","volume":"44","author":"Mathisen","year":"2020","journal-title":"Auton. Robot."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Klinkmueller, K., Wieck, A., Holt, J., Valentine, A., Bluman, J.E., Kopeikin, A., and Prosser, E. (2019, January 7\u201311). Airborne delivery of unmanned aerial vehicles via joint precision airdrop systems. Proceedings of the AIAA Scitech 2019 Forum, San Diego, CA, USA.","DOI":"10.2514\/6.2019-2285"},{"key":"ref_4","unstructured":"Yang, L., Qi, J., Xiao, J., and Yong, X. (July, January 29). A literature review of UAV 3D path planning. Proceedings of the 11th World Congress on Intelligent Control and Automation, Shenyang, China."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Huang, S., and Teo, R.S.H. (2019, January 11\u201314). Computationally efficient visibility graph-based generation of 3D shortest collision-free path among polyhedral obstacles for unmanned aerial vehicles. Proceedings of the 2019 International Conference on Unmanned Aircraft Systems (ICUAS), Atlanta, GA, USA.","DOI":"10.1109\/ICUAS.2019.8798322"},{"key":"ref_6","unstructured":"Cheng, X., Zhou, D., and Zhang, R. (2013, January 5\u20138). New method for UAV online path planning. Proceedings of the 2013 IEEE International Conference on Signal Processing, Communication and Computing (ICSPCC 2013), KunMing, China."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Sun, Q., Li, M., Wang, T., and Zhao, C. (2018, January 9\u201311). UAV path planning based on improved rapidly-exploring random tree. Proceedings of the 2018 Chinese Control and Decision Conference (CCDC), Shenyang, China.","DOI":"10.1109\/CCDC.2018.8408258"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"525","DOI":"10.1007\/s11633-013-0750-9","article-title":"Path planning in complex 3D environments using a probabilistic roadmap method","volume":"10","author":"Yan","year":"2013","journal-title":"Int. J. Autom. Comput."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Tseng, F.H., Liang, T.T., Lee, C.H., Der Chou, L., and Chao, H.C. (2014, January 27\u201329). A star search algorithm for civil UAV path planning with 3G communication. Proceedings of the 2014 Tenth International Conference on Intelligent Information Hiding and Multimedia Signal Processing, Kitakyushu, Japan.","DOI":"10.1109\/IIH-MSP.2014.236"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Meng, B.B., and Gao, X. (2010, January 11\u201312). UAV path planning based on bidirectional sparse A* search algorithm. Proceedings of the 2010 International Conference on Intelligent Computation Technology and Automation, Changsha, China.","DOI":"10.1109\/ICICTA.2010.235"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"122757","DOI":"10.1109\/ACCESS.2020.3007496","article-title":"A novel real-time penetration path planning algorithm for stealth UAV in 3D complex dynamic environment","volume":"8","author":"Zhang","year":"2020","journal-title":"IEEE Access"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1760008","DOI":"10.1142\/S0218213017600089","article-title":"Heuristic and genetic algorithm approaches for UAV path planning under critical situation","volume":"26","author":"Williams","year":"2017","journal-title":"Int. J. Artif. Intell. Tools"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"29","DOI":"10.2514\/2.4229","article-title":"Trajectory tracking for autonomous vehicles: An integrated approach to guidance and control","volume":"21","author":"Kaminer","year":"1998","journal-title":"J. Guid. Control. Dyn."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"281","DOI":"10.1007\/s12555-015-0289-3","article-title":"Trajectory tracking control of multirotors from modelling to experiments: A survey","volume":"15","author":"Lee","year":"2017","journal-title":"Int. J. Control Autom. Syst."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_16","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015). Continuous control with deep reinforcement learning. arXiv."},{"key":"ref_17","unstructured":"Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2015). Prioritized experience replay. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Van Hasselt, H., Guez, A., and Silver, D. (2016, January 12\u201317). Deep reinforcement learning with double q-learning. Proceedings of the AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA.","DOI":"10.1609\/aaai.v30i1.10295"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Hou, Y., Liu, L., Wei, Q., Xu, X., and Chen, C. (2017, January 5\u20138). A novel DDPG method with prioritized experience replay. Proceedings of the 2017 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Banff, AB, Canada.","DOI":"10.1109\/SMC.2017.8122622"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"2397","DOI":"10.1109\/TAES.2013.6621824","article-title":"UAV path planning in a dynamic environment via partially observable Markov decision process","volume":"49","author":"Ragi","year":"2013","journal-title":"IEEE Trans. Aerosp. Electron. Syst."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Fran\u00e7ois-Lavet, V., Henderson, P., Islam, R., Bellemare, M.G., and Pineau, J. (2018). An introduction to deep reinforcement learning. arXiv.","DOI":"10.1561\/9781680835397"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhang, K., Li, K., He, J., Shi, H., Wang, Y., and Niu, C. (2020, January 1\u20134). A UAV Autonomous Maneuver Decision-Making Algorithm for Route Guidance. Proceedings of the 2020 International Conference on Unmanned Aircraft Systems (ICUAS), Athens, Greece.","DOI":"10.1109\/ICUAS48674.2020.9213968"},{"key":"ref_23","unstructured":"Ng, A.Y., Harada, D., and Russell, S. (1999, January 27\u201330). Policy invariance under reward transformations: Theory and application to reward shaping. Proceedings of the Sixteenth International Conference on Machine Learning, Bled, Slovenia."},{"key":"ref_24","unstructured":"Badnava, B., and Mozayani, N. (2019). A new potential-based reward shaping for reinforcement learning agent. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"58","DOI":"10.1145\/203330.203343","article-title":"Temporal difference learning and TD-Gammon","volume":"38","author":"Tesauro","year":"1995","journal-title":"Commun. ACM"},{"key":"ref_26","unstructured":"Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014, January 22\u201324). Deterministic policy gradient algorithms. Proceedings of the International Conference on Machine Learning, Beijing, China."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"167","DOI":"10.1111\/1467-9868.00282","article-title":"Non-Gaussian Ornstein\u2013Uhlenbeck-based models and some of their uses in financial economics","volume":"63","author":"Shephard","year":"2001","journal-title":"J. R. Stat. Soc. Ser. B"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/6\/2233\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:39:42Z","timestamp":1760161182000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/6\/2233"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,23]]},"references-count":27,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2021,3]]}},"alternative-id":["s21062233"],"URL":"https:\/\/doi.org\/10.3390\/s21062233","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,3,23]]}}}