{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,14]],"date-time":"2026-01-14T16:05:46Z","timestamp":1768406746617,"version":"3.49.0"},"reference-count":30,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2024,3,20]],"date-time":"2024-03-20T00:00:00Z","timestamp":1710892800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Reinforcement learning (RL) is pivotal in empowering Unmanned Aerial Vehicles (UAVs) to navigate and make decisions efficiently and intelligently within complex and dynamic surroundings. Despite its significance, RL is hampered by inherent limitations such as low sample efficiency, restricted generalization capabilities, and a heavy reliance on the intricacies of reward function design. These challenges often render single-method RL approaches inadequate, particularly in the context of UAV operations where high costs and safety risks in real-world applications cannot be overlooked. To address these issues, this paper introduces a novel RL framework that synergistically integrates meta-learning and imitation learning. By leveraging the Reptile algorithm from meta-learning and Generative Adversarial Imitation Learning (GAIL), coupled with state normalization techniques for processing state data, this framework significantly enhances the model\u2019s adaptability. It achieves this by identifying and leveraging commonalities across various tasks, allowing for swift adaptation to new challenges without the need for complex reward function designs. To ascertain the efficacy of this integrated approach, we conducted simulation experiments within both two-dimensional environments. The empirical results clearly indicate that our GAIL-enhanced Reptile method surpasses conventional single-method RL algorithms in terms of training efficiency. This evidence underscores the potential of combining meta-learning and imitation learning to surmount the traditional barriers faced by reinforcement learning in UAV trajectory planning and decision-making processes.<\/jats:p>","DOI":"10.3390\/fi16030105","type":"journal-article","created":{"date-parts":[[2024,3,20]],"date-time":"2024-03-20T09:14:33Z","timestamp":1710926073000},"page":"105","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["UAV Control Method Combining Reptile Meta-Reinforcement Learning and Generative Adversarial Imitation Learning"],"prefix":"10.3390","volume":"16","author":[{"given":"Shui","family":"Jiang","sequence":"first","affiliation":[{"name":"College of Computer and Cyber Security, Fujian Normal University, Fuzhou 350007, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-1682-5123","authenticated-orcid":false,"given":"Yanning","family":"Ge","sequence":"additional","affiliation":[{"name":"College of Computer and Cyber Security, Fujian Normal University, Fuzhou 350007, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xu","family":"Yang","sequence":"additional","affiliation":[{"name":"College of Computer and Control Engineering, Minjiang University, Fuzhou 350108, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7800-2215","authenticated-orcid":false,"given":"Wencheng","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Mathematics, Physics and Computing, University of Southern Queensland, Darling Heights, QLD 4350, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5820-2233","authenticated-orcid":false,"given":"Hui","family":"Cui","sequence":"additional","affiliation":[{"name":"Department of Software Systems & Cybersecurity, Monash University, Melbourne, VIC 3800, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,3,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1109\/JRPROC.1961.287775","article-title":"Steps toward artificial intelligence","volume":"49","author":"Minsky","year":"1961","journal-title":"Proc. IRE"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Kendall, A., Hawke, J., Janz, D., Mazur, P., Reda, D., Allen, J.M., Lam, V.D., Bewley, A., and Shah, A. (2019, January 20\u201324). Learning to drive in a day. Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8793742"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1038\/nature24270","article-title":"Mastering the game of go without human knowledge","volume":"550","author":"Silver","year":"2017","journal-title":"Nature"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"129274","DOI":"10.1109\/ACCESS.2020.3009329","article-title":"Sample Efficient Reinforcement Learning Method via High Efficient Episodic Memory","volume":"8","author":"Yang","year":"2020","journal-title":"IEEE Access"},{"key":"ref_5","unstructured":"Eschmann, J. (2021). Reinforcement Learning Algorithms: Analysis and Applications, Springer."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Zhou, W., and Li, W. (March, January 22). Programmatic reward design by example. Proceedings of the AAAI Conference on Artificial Intelligence 2022, Virtual.","DOI":"10.1609\/aaai.v36i8.20910"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1471","DOI":"10.1109\/LRA.2021.3057046","article-title":"Model-based meta-reinforcement learning for flight with suspended payloads","volume":"6","author":"Belkhale","year":"2021","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"166","DOI":"10.23919\/JCC.2022.04.013","article-title":"Multi-agent few-shot meta reinforcement learning for trajectory design and channel selection in UAV-assisted networks","volume":"19","author":"Zhou","year":"2022","journal-title":"China Commun."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"768","DOI":"10.1007\/s42405-020-00254-x","article-title":"Vision-based obstacle avoidance for UAVs via imitation learning with sequential neural networks","volume":"21","author":"Park","year":"2020","journal-title":"Int. J. Aeronaut. Space Sci."},{"key":"ref_10","unstructured":"Beck, J., Vuorio, R., Liu, E.Z., Xiong, Z., Zintgraf, L., Finn, C., and Whiteson, S. (2023). A survey of meta-reinforcement learning. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Azar, A.T., Koubaa, A., Ali Mohamed, N., Ibrahim, H.A., Ibrahim, Z.F., Kazim, M., Ammar, A., Benjdira, B., Khamis, A.M., and Hameed, I.A. (2021). Drone deep reinforcement learning: A review. Electronics, 10.","DOI":"10.3390\/electronics10090999"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3054912","article-title":"Imitation learning: A survey of learning methods","volume":"50","author":"Hussein","year":"2017","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"729","DOI":"10.1109\/TWC.2019.2935201","article-title":"Multi-agent reinforcement learning-based resource allocation for UAV networks","volume":"19","author":"Cui","year":"2019","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_14","unstructured":"Osa, T., Pajarinen, J., Neumann, G., Bagnell, J.A., Abbeel, P., and Peters, J. (2024, March 01). An Algorithmic Perspective on Imitation Learning; Foundations and Trends\u00ae in Robotics. Available online: https:\/\/www.nowpublishers.com\/article\/Details\/ROB-053."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"3730","DOI":"10.1109\/TITS.2020.3023958","article-title":"Distributed learning for vehicle routing decision in software defined Internet of vehicles","volume":"22","author":"Lin","year":"2020","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_16","unstructured":"Ho, J., and Ermon, S. (2016). Generative adversarial imitation learning. Adv. Neural Inf. Process. Syst., 29."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3301273","article-title":"Reinforcement learning for UAV attitude control","volume":"3","author":"Koch","year":"2019","journal-title":"ACM Trans. Cyber-Phys. Syst."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Yijing, Z., Zheng, Z., Xiaoyi, Z., and Yang, L. (2017, January 26\u201328). Q learning algorithm based UAV path learning and obstacle avoidence approach. Proceedings of the 2017 36th Chinese control conference (CCC), Dalian, China.","DOI":"10.23919\/ChiCC.2017.8027884"},{"key":"ref_19","unstructured":"Pham, H., La, H., Feil-Seifer, D., and Nguyen, L. (2018). Autonomous uav navigation using reinforcement learning. arXiv."},{"key":"ref_20","unstructured":"He, L., Aouf, N., Whidborne, J.F., and Song, B. (2020). Deep reinforcement learning based local planner for UAV obstacle avoidance using demonstration data. arXiv."},{"key":"ref_21","unstructured":"Yang, B., Ma, C., and Xia, X. (2021, January 3\u20137). Drone formation control via belief-correlated imitation learning. Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems, Online."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Wang, T., and Chang, D.E. (June, January 30). Robust navigation for racing drones based on imitation learning and modularization. Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi\u2019an, China.","DOI":"10.1109\/ICRA48506.2021.9560743"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Hu, Y., Chen, M., Saad, W., Poor, H.V., and Cui, S. (2020, January 7\u201311). Meta-reinforcement learning for trajectory design in wireless UAV networks. Proceedings of the GLOBECOM 2020\u20142020 IEEE Global Communications Conference, Taipei, Taiwan.","DOI":"10.1109\/GLOBECOM42002.2020.9322414"},{"key":"ref_24","unstructured":"Prat, A., and Johns, E. (2021, January 4). PERIL: Probabilistic embeddings for hybrid meta-reinforcement and imitation learning. Proceedings of the International Conference on Learning Representations, Vienna, Austria."},{"key":"ref_25","unstructured":"Puterman, M.L. (2014). Markov Decision Processes: Discrete Stochastic Dynamic Programming, John Wiley & Sons."},{"key":"ref_26","unstructured":"Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S. (2020, January 16\u201318). Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. Proceedings of the Conference on Robot Learning PMLR, Virtual."},{"key":"ref_27","unstructured":"Nagabandi, A., Clavera, I., Liu, S., Fearing, R.S., Abbeel, P., Levine, S., and Finn, C. (2018). Learning to adapt in dynamic, real-world environments through meta-reinforcement learning. arXiv."},{"key":"ref_28","unstructured":"Nichol, A., Achiam, J., and Schulman, J. (2018). On first-order meta-learning algorithms. arXiv."},{"key":"ref_29","unstructured":"Finn, C., Abbeel, P., and Levine, S. (2017, January 6\u201311). Model-agnostic meta-learning for fast adaptation of deep networks. Proceedings of the International Conference on Machine Learning PMLR, Sydney, Australia."},{"key":"ref_30","unstructured":"Robak, M. (2024, March 01). Zastosowanie Uczenia ze Wzmocnieniem (Reinforcement Learning) do Stabilizacji Ruchu. Available online: https:\/\/ruj.uj.edu.pl\/xmlui\/handle\/item\/287455."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/16\/3\/105\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:16:41Z","timestamp":1760105801000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/16\/3\/105"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,20]]},"references-count":30,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2024,3]]}},"alternative-id":["fi16030105"],"URL":"https:\/\/doi.org\/10.3390\/fi16030105","relation":{},"ISSN":["1999-5903"],"issn-type":[{"value":"1999-5903","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,20]]}}}