{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T04:16:12Z","timestamp":1781756172417,"version":"3.54.5"},"reference-count":42,"publisher":"MDPI AG","issue":"19","license":[{"start":{"date-parts":[[2020,10,1]],"date-time":"2020-10-01T00:00:00Z","timestamp":1601510400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Unmanned aerial vehicle (UAV) autonomous tracking and landing is playing an increasingly important role in military and civil applications. In particular, machine learning has been successfully introduced to robotics-related tasks. A novel UAV autonomous tracking and landing approach based on a deep reinforcement learning strategy is presented in this paper, with the aim of dealing with the UAV motion control problem in an unpredictable and harsh environment. Instead of building a prior model and inferring the landing actions based on heuristic rules, a model-free method based on a partially observable Markov decision process (POMDP) is proposed. In the POMDP model, the UAV automatically learns the landing maneuver by an end-to-end neural network, which combines the Deep Deterministic Policy Gradients (DDPG) algorithm and heuristic rules. A Modular Open Robots Simulation Engine (MORSE)-based reinforcement learning framework is designed and validated with a continuous UAV tracking and landing task on a randomly moving platform in high sensor noise and intermittent measurements. The simulation results show that when the moving platform is moving in different trajectories, the average landing success rate of the proposed algorithm is about 10% higher than that of the Proportional-Integral-Derivative (PID) method. As an indirect result, a state-of-the-art deep reinforcement learning-based UAV control method is validated, where the UAV can learn the optimal strategy of a continuously autonomous landing and perform properly in a simulation environment.<\/jats:p>","DOI":"10.3390\/s20195630","type":"journal-article","created":{"date-parts":[[2020,10,1]],"date-time":"2020-10-01T11:28:57Z","timestamp":1601551737000},"page":"5630","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":52,"title":["UAV Autonomous Tracking and Landing Based on Deep Reinforcement Learning Strategy"],"prefix":"10.3390","volume":"20","author":[{"given":"Jingyi","family":"Xie","sequence":"first","affiliation":[{"name":"Key Laboratory of Electronics and Information Technology for Space System, National Space Science Center, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaodong","family":"Peng","sequence":"additional","affiliation":[{"name":"Key Laboratory of Electronics and Information Technology for Space System, National Space Science Center, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haijiao","family":"Wang","sequence":"additional","affiliation":[{"name":"Alibaba Damo Academy, Hangzhou 311121, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1643-1705","authenticated-orcid":false,"given":"Wenlong","family":"Niu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Electronics and Information Technology for Space System, National Space Science Center, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiao","family":"Zheng","sequence":"additional","affiliation":[{"name":"Key Laboratory of Electronics and Information Technology for Space System, National Space Science Center, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2020,10,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Kim, J., Jung, Y., Lee, D., and Shim, D.H. (2014, January 27\u201330). Outdoor autonomous landing on a moving platform for quadrotors using an omnidirectional camera. Proceedings of the IEEE International Conference on Unmanned Aircraft Systems (ICUAS), Orlando, FL, USA.","DOI":"10.1109\/ICUAS.2014.6842381"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Li, K., Liu, P., Pang, T., Yang, Z., and Chen, B.M. (2015, January 9\u201312). Development of an unmanned aerial vehicle for rooftop landing and surveillance. Proceedings of the IEEE International Conference on Unmanned Aircraft Systems (ICUAS), Denver, CO, USA.","DOI":"10.1109\/ICUAS.2015.7152368"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"874","DOI":"10.1002\/rob.21858","article-title":"Autonomous landing on a moving vehicle with an unmanned aerial vehicle","volume":"36","author":"Baca","year":"2019","journal-title":"J. Field Robot."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"369","DOI":"10.1007\/s10846-016-0399-z","article-title":"Vision based autonomous landing of multirotor UAV on moving platform","volume":"85","author":"Araar","year":"2017","journal-title":"J. Intell. Robot. Syst."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"10488","DOI":"10.1016\/j.ifacol.2017.08.1980","article-title":"Autonomous landing of a multirotor micro air vehicle on a high velocity ground vehicle","volume":"50","author":"Borowczyk","year":"2017","journal-title":"IFAC PapersOnLine"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Xu, Y., Liu, Z., and Wang, X. (2018, January 25\u201327). Monocular Vision based Autonomous Landing of Quadrotor through Deep Reinforcement Learning. Proceedings of the 37th Chinese Control Conference (CCC), Wuhan, China.","DOI":"10.23919\/ChiCC.2018.8482830"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"989","DOI":"10.1002\/asjc.1758","article-title":"Nonlinear and adaptive intelligent control techniques for quadrotor uav\u2014A survey","volume":"21","author":"Mo","year":"2019","journal-title":"Asian J. Control."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Yang, T., Ren, Q., Zhang, F., Xie, B., Ren, H., Li, J., and Zhang, Y. (2018). Hybrid camera array-based uav auto-landing on moving ugv in gps-denied environment. Remote Sens., 10.","DOI":"10.3390\/rs10111829"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Yang, B., and Dutta, P. (2019, January 17\u201321). Cooperative Navigation for Small UAVs in GPS-Intermittent Environments. Proceedings of the AIAA Aviation 2019 Forum, Dallas, TX, USA.","DOI":"10.2514\/6.2019-3515"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Van den Meijdenberg, J., Totu, L., Schi\u00f8ler, H., and Leth, J. (2018, January 6\u20138). Stochastic Controller Design for multi-rotor UAV under Intermittent Localization. Proceedings of the IEEE Australian & New Zealand Control Conference (ANZCC), Melbourne, VIC, Australia.","DOI":"10.1109\/ANZCC.2018.8606617"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Arora, S., Jain, S., Scherer, S., Nuske, S., Chamberlain, L., and Singh, S. (2013, January 6\u201310). Infrastructure-free ship deck tracking for autonomous landing. Proceedings of the 2013 IEEE International Conference on Robotics and Automation (ICRA), Karlsruhe, Germany.","DOI":"10.1109\/ICRA.2013.6630595"},{"key":"ref_12","first-page":"199","article-title":"Vision analysis system for autonomous landing of micro drone","volume":"8","author":"Skoczylas","year":"2014","journal-title":"Actamech. Automat."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TAES.2017.2767958","article-title":"Optimal rendezvous trajectory for unmanned aerial-ground vehicles","volume":"54","author":"Rucco","year":"2017","journal-title":"IEEE Trans. Aerosp. Electron. Syst."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Vlantis, P., Marantos, P., Bechlioulis, C.P., and Kyriakopoulos, K.J. (2015, January 26\u201330). Quadrotor Landing on an Inclined Platform of a Moving Ground Vehicle. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Seattle, WA, USA.","DOI":"10.1109\/ICRA.2015.7139490"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1007\/s10846-010-9473-0","article-title":"Automatic take off, tracking and landing of a miniature UAV on a moving carrier vehicle","volume":"61","author":"Wenzel","year":"2011","journal-title":"J. Intell. Robot. Syst."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2627","DOI":"10.1007\/s13369-018-3330-z","article-title":"Fuzzy Logic-Based Robust and Autonomous Safe Landing for UAV Quadcopter","volume":"44","author":"Talha","year":"2019","journal-title":"Arab. J. Sci. Eng."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Prach, A., G\u00fcrsoy, G., and Yavrucuk, L. (2019, January 10\u201312). Nonlinear Controller for a Fixed-Wing Aircraft Landing. Proceedings of the IEEE American Control Confefrence (ACC), Philadelphia, PA, USA.","DOI":"10.23919\/ACC.2019.8814970"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Rossi, E., Bruschetta, M., Carli, R., Chen, Y., and Farina, M. (2019, January 25\u201328). Online nonlinear model predictive control for tethered uavs to perform a safe and constrained maneuver. Proceedings of the IEEE European Control Conference (ECC), Naples, Italy.","DOI":"10.23919\/ECC.2019.8796032"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Beul, M., Houben, S., Nieuwenhuisen, M., and Behnke, S. (2017, January 6\u20138). Fast autonomous landing on a moving target at MBZIRC. Proceedings of the European Conference on Mobile Robots, Paris, France.","DOI":"10.1109\/ECMR.2017.8098669"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Kamath, A.K., Tripathi, V.K., Yogi, S.C., and Behera, L. (2019, January 14\u201318). Vision-based Fast-terminal Sliding Mode Super Twisting Controller for Autonomous Landing of a Quadrotor on a Static Platform. Proceedings of the IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), New Delhi, India.","DOI":"10.1109\/RO-MAN46459.2019.8956302"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1830","DOI":"10.1049\/iet-cta.2017.0998","article-title":"Saturated adaptive sliding mode control for autonomous vessel landing of a quadrotor","volume":"12","author":"Huang","year":"2018","journal-title":"IET Control. Theor. Appl."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Fei, Q., Zhang, J., Wang, Z., and Huang, X. (2019). Sliding Mode Control with Uncertain Model for a Quadrotor UAV\u2019s Automatic Visual Landing Problem. Proceedings of the Chinese Intelligent Automation Conference, Springer.","DOI":"10.1007\/978-981-32-9050-1_26"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"He, S., Wang, H., and Zhang, S. (2019, January 3\u20135). Vision Based Autonomous Landing of the Quadrotor Using Fuzzy Logic Control. Proceedings of the IEEE Chinese Control and Decision Conference (CCDC), Nanchang, China.","DOI":"10.1109\/CCDC.2019.8832729"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"38407","DOI":"10.1109\/ACCESS.2019.2906345","article-title":"Autonomous moving target-tracking for a UAV quadcopter based on fuzzy-PI","volume":"7","author":"Rabah","year":"2019","journal-title":"IEEE Access"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"522","DOI":"10.1007\/s40313-019-00465-y","article-title":"Autonomous landing of UAV based on artificial neural network supervised by fuzzy logic","volume":"30","author":"Marcato","year":"2019","journal-title":"J. Control Automat. Electr. Syst."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Ananthakrishnan, U., Akshay, N., Manikutty, G., and Bhavani, R.R. (2017). Control of Quadrotors Using Neural Networks for Precise Landing Maneuvers. Artificial Intelligence and Evolutionary Computations in Engineering Systems, Springer.","DOI":"10.1007\/978-981-10-3174-8_10"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_28","unstructured":"Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014, January 21\u201326). Deterministic Policy Gradient Algorithms. Proceedings of the International Conference on Machine Learning, Beijing, China."},{"key":"ref_29","unstructured":"Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016, January 19\u201324). Asynchronous Methods for Deep Reinforcement Learning. Proceedings of the International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Polvara, R., Patacchiola, M., Sharma, S., Wan, J., Manning, A., Sutton, R., and Cangelosi, A. (2018, January 12\u201315). Toward End-to-End Control for UAV Autonomous Landing via Deep Reinforcement Learning. Proceedings of the IEEE International Conference on Unmanned Aircraft Systems (ICUAS), Dallas, TX, USA.","DOI":"10.1109\/ICUAS.2018.8453449"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"James, S., Wohlhart, P., Kalakrishnan, M., Kalashnikov, D., Irpan, A., Ibarz, J., Levine, S., Hadsell, R., and Bousmalis, K. (2019, January 16\u201320). Sim-to-Real via Sim-to-Sim: Data-efficient Robotic Grasping via Randomized-to-Canonical Adaptation Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Los Angeles, CA, USA.","DOI":"10.1109\/CVPR.2019.01291"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"351","DOI":"10.1007\/s10846-018-0891-8","article-title":"A deep reinforcement learning strategy for UAV autonomous landing on a moving platform","volume":"93","author":"Sampedro","year":"2019","journal-title":"J. Intell. Robot. Syst."},{"key":"ref_33","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015). Continuous control with deep reinforcement learning. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Falanga, D., Zanchettin, A., Simovic, A., Delmerico, J., and Scaramuzza, D. (2017, January 11\u201313). Vision-based autonomous quadrotor landing on a moving platform. Proceedings of the IEEE International Symposium on Safety, Security and Rescue Robotics, Shanghai, China.","DOI":"10.1109\/SSRR.2017.8088164"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"282","DOI":"10.1287\/opre.26.2.282","article-title":"The optimal control of partially observable Markov processes over the infinite horizon: Discounted costs","volume":"26","author":"Sondik","year":"1978","journal-title":"Oper. Res."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Gautam, A., Sujit, P., and Saripalli, S. (2014, January 27\u201330). A survey of autonomous landing techniques for UAVs. Proceedings of the IEEE 2014 International Conference on Unmanned Aircraft Systems (ICUAS), Orlando, FL, USA.","DOI":"10.1109\/ICUAS.2014.6842377"},{"key":"ref_37","first-page":"792","article-title":"Markov decision processes: Discrete stochastic dynamic programming","volume":"46","author":"Puterman","year":"1995","journal-title":"J. Oper. Res. Soc."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Tan, L., Wu, J., Yang, X., and Song, S. (2019). Research on Optimal Landing Trajectory Planning Method between an UAV and a Moving Vessel. Appl. Sci., 9.","DOI":"10.3390\/app9183708"},{"key":"ref_39","unstructured":"Forsmo, E.J. (2012). Optimal Path Planning for Unmanned Aerial Systems. [Master\u2019s Thesis, Norwegian University of Science and Technology]."},{"key":"ref_40","unstructured":"Kirk, D.E. (2004). Optimal Control Theory: An Introduction, Dover Publications. Courier Corporation."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Degris, T., Pilarski, P.M., and Sutton, R.S. (2012, January 27\u201329). Model-Free Reinforcement Learning with Continuous Action in Practice. Proceedings of the IEEE American Control Conference (ACC), Montreal, QC, Canada.","DOI":"10.1109\/ACC.2012.6315022"},{"key":"ref_42","first-page":"36","article-title":"Various techniques used in connection with random digits","volume":"12","year":"1951","journal-title":"Appl. Math. Ser."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/19\/5630\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:15:49Z","timestamp":1760177749000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/19\/5630"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,1]]},"references-count":42,"journal-issue":{"issue":"19","published-online":{"date-parts":[[2020,10]]}},"alternative-id":["s20195630"],"URL":"https:\/\/doi.org\/10.3390\/s20195630","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,10,1]]}}}