{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T15:26:47Z","timestamp":1787066807376,"version":"build-2736575974"},"reference-count":169,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,3,31]]},"abstract":"<jats:p>In recent years, Unmanned Aerial Vehicles (UAVs) have attracted a lot of attention due to their flexibility and mobility. However, due to the increasingly complex environments faced by UAVs and the rising demands on UAV systems, traditional UAV control methods can no longer efficiently control the UAV under multi-constraint situations. Reinforcement Learning (RL), as an emerging robot control technology, is well suited to the needs of UAV systems in terms of its ability to interact with and learn from the environment. Therefore, RL-based UAV systems are gradually becoming a new trend in research. Nonetheless, as a new research field, it faces some challenges. To fully grasp the landscape of RL-based UAV systems, it is paramount to provide a comprehensive overview and analysis of the existing specific RL methods applied to UAV systems. In this survey, we first provide a comprehensive overview and summary of the application of RL in different UAV scenarios based on the classification of RL methods. After that, based on the existing relevant literature, we conduct a systematic analysis of the challenges and recent advancements when applying RL to UAV systems. Finally, we discuss the potential research directions for RL-based UAV systems.<\/jats:p>","DOI":"10.1145\/3769426","type":"journal-article","created":{"date-parts":[[2025,9,25]],"date-time":"2025-09-25T11:28:07Z","timestamp":1758799687000},"page":"1-37","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["A Survey on Reinforcement Learning Methods for UAV Systems"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-4830-8087","authenticated-orcid":false,"given":"Hengsheng","family":"Chen","sequence":"first","affiliation":[{"name":"Fuzhou University","place":["Fuzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6717-0950","authenticated-orcid":false,"given":"Yuanguo","family":"Lin","sequence":"additional","affiliation":[{"name":"Jimei University","place":["Xiamen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4920-7235","authenticated-orcid":false,"given":"Mingjian","family":"Fu","sequence":"additional","affiliation":[{"name":"College of Computer and Data Science, Fuzhou University","place":["Fuzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4149-839X","authenticated-orcid":false,"given":"Lina","family":"Yao","sequence":"additional","affiliation":[{"name":"University of New South Wales","place":["Sydney, Australia"]},{"name":"CSIRO's Data61","place":["Sydney, Australia"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3326-4147","authenticated-orcid":false,"given":"Michael","family":"Sheng","sequence":"additional","affiliation":[{"name":"Macquarie University","place":["Sydney, Australia"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,10,25]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2020.3039617"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.vehcom.2021.100391"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2020.3020220"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2022.105321"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2912306"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIV.2024.3414140"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSYST.2021.3082837"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2017.2743240"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMST.2023.3323344"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/ROBOT.2005.1570447"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCC.2024.3360443"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TWC.2022.3216049"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/WCNC51071.2022.9771867"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2020.3023733"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2021.3094273"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSAC.2019.2906789"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2023.3340177"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2022.3188473"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2023.3340669"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2021.3113052"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/WCNC51071.2022.9771588"},{"key":"e_1_3_1_23_2","series-title":"NIPS\u201923","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Das Devleena","year":"2023","unstructured":"Devleena Das, Sonia Chernova, and Been Kim. 2023. State2Explanation: Concept-based explanations to benefit agent learning and user understanding. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS\u201923). Curran Associates Inc., Red Hook, NY, USA, Article 2935, 27 pages."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2024.3357821"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1142\/S2301385019400065"},{"key":"e_1_3_1_26_2","volume-title":"Proceedings of the 40th International Conference on Machine Learning (ICML\u201923)","author":"Du Yuqing","year":"2023","unstructured":"Yuqing Du, Olivia Watkins, Zihan Wang, C\u00e9dric Colas, Trevor Darrell, Pieter Abbeel, Abhishek Gupta, and Jacob Andreas. 2023. Guiding pretraining in reinforcement learning with large language models. In Proceedings of the 40th International Conference on Machine Learning (ICML\u201923). JMLR.org, Article 346, 21 pages."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1002\/wcs.119"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/WCNC57260.2024.10570813"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1007\/11787006_1"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2020.2966989"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNSM.2024.3378677"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2024.3394235"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2024.3397007"},{"key":"e_1_3_1_34_2","unstructured":"Pu Feng Junkang Liang Size Wang Xin Yu Xin Ji Yiting Chen Kui Zhang Rongye Shi and Wenjun Wu. 2024. Hierarchical Consensus-Based Multi-Agent Reinforcement Learning for Multi-Robot Cooperation Tasks. arxiv:2407.08164 [cs.AI] https:\/\/arxiv.org\/abs\/2407.08164"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2023.3296769"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TWC.2023.3281812"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2021.3051370"},{"key":"e_1_3_1_38_2","series-title":"Proceedings of Machine Learning Research","first-page":"1587","volume-title":"Proceedings of the 35th International Conference on Machine Learning","volume":"80","author":"Fujimoto Scott","year":"2018","unstructured":"Scott Fujimoto, Herke van Hoof, and David Meger. 2018. Addressing function approximation error in actor-critic methods. In Proceedings of the 35th International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol. 80). PMLR, 1587\u20131596."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2024.3367624"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.5555\/2789272.2886795"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/LCOMM.2024.3377005"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2024.3406220"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3422622"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2023.3320796"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2023.3311484"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSAC.2023.3310065"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCOMM.2022.3222460"},{"key":"e_1_3_1_48_2","unstructured":"Tuomas Haarnoja Aurick Zhou Pieter Abbeel and Sergey Levine. 2018. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. arXiv:1801.01290. Retrieved from https:\/\/arxiv.org\/abs\/\/1801.01290"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.5555\/3016100.3016191"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ast.2021.107052"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2021.3088129"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jksuci.2023.101817"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2020.3003639"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2023.3283502"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2019.2952549"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2022.3172936"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICAIGE58321.2023.10346406"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2022.3181308"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/MNET.011.2000440"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2024.111604"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.15837\/ijccc.2023.6.5505"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.5555\/1622737.1622748"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2023.3240173"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-023-06419-4"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDS50568.2020.9268683"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNSM.2024.3391664"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jnca.2023.103813"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMST.2023.3312221"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/OJCOMS.2021.3075201"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2022.3161571"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2021.3113128"},{"key":"e_1_3_1_72_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCOMM.2023.3333880"},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2019.2945037"},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.1109\/LCOMM.2023.3265214"},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2020.2975749"},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.1109\/TWC.2022.3198665"},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","DOI":"10.1109\/WCNC45663.2020.9120668"},{"key":"e_1_3_1_78_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.adhoc.2023.103341"},{"key":"e_1_3_1_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASE.2024.3382730"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2021.3062659"},{"key":"e_1_3_1_81_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.comcom.2023.11.006"},{"key":"e_1_3_1_82_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2023.02.027"},{"key":"e_1_3_1_83_2","doi-asserted-by":"publisher","DOI":"10.1145\/3556973"},{"key":"e_1_3_1_84_2","unstructured":"Timothy P. Lillicrap Jonathan J. Hunt Alexander Pritzel Nicolas Heess Tom Erez Yuval Tassa David Silver and Daan Wierstra. 2019. Continuous control with deep reinforcement learning. (2019). arXiv:1509.02971. Retrieved from https:\/\/arxiv.org\/abs\/\/1509.02971"},{"key":"e_1_3_1_85_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2023.3280161"},{"key":"e_1_3_1_86_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSAC.2019.2904353"},{"key":"e_1_3_1_87_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2021.3110138"},{"key":"e_1_3_1_88_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIV.2022.3213703"},{"key":"e_1_3_1_89_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2019.2920284"},{"key":"e_1_3_1_90_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2022.3142269"},{"key":"e_1_3_1_91_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNSE.2023.3292570"},{"key":"e_1_3_1_92_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.abg5810"},{"key":"e_1_3_1_93_2","unstructured":"Scott M. Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach California USA) (NIPS\u201917). Curran Associates Inc. Red Hook NY USA 4768\u20134777."},{"key":"e_1_3_1_94_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSAC.2022.3228558"},{"key":"e_1_3_1_95_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMST.2019.2916583"},{"key":"e_1_3_1_96_2","doi-asserted-by":"publisher","DOI":"10.1109\/TWC.2021.3086503"},{"key":"e_1_3_1_97_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2017.10.009"},{"key":"e_1_3_1_98_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCCN.2024.3394859"},{"key":"e_1_3_1_99_2","doi-asserted-by":"publisher","DOI":"10.1145\/3616864"},{"key":"e_1_3_1_100_2","doi-asserted-by":"publisher","DOI":"10.5555\/3045390.3045594"},{"key":"e_1_3_1_101_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_1_102_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCCN.2023.3262242"},{"key":"e_1_3_1_103_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSAC.2021.3088718"},{"key":"e_1_3_1_104_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.iot.2024.101342"},{"key":"e_1_3_1_105_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2022.3150184"},{"key":"e_1_3_1_106_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2021.3134162"},{"key":"e_1_3_1_107_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2022.3172587"},{"key":"e_1_3_1_108_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2022.3165227"},{"key":"e_1_3_1_109_2","doi-asserted-by":"publisher","DOI":"10.1109\/OJCOMS.2023.3251297"},{"key":"e_1_3_1_110_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSYST.2023.3266769"},{"key":"e_1_3_1_111_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNSM.2023.3243543"},{"key":"e_1_3_1_112_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.comnet.2020.107148"},{"key":"e_1_3_1_113_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2021.3071531"},{"key":"e_1_3_1_114_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2020.3023861"},{"key":"e_1_3_1_115_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2020.3042925"},{"key":"e_1_3_1_116_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2020.2991326"},{"key":"e_1_3_1_117_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCCN.2019.2948324"},{"key":"e_1_3_1_118_2","first-page":"1889","volume-title":"Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37 (ICML\u201915)","author":"Schulman John","year":"2015","unstructured":"John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan, and Pieter Abbeel. 2015. Trust region policy optimization. In Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37 (ICML\u201915). JMLR.org, 1889\u20131897."},{"key":"e_1_3_1_119_2","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal Policy Optimization Algorithms. arXiv:1707.06347. Retrieved from https:\/\/arxiv.org\/abs\/\/1707.06347"},{"key":"e_1_3_1_120_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2024.124379"},{"key":"e_1_3_1_121_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ast.2023.108609"},{"key":"e_1_3_1_122_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2023.3246925"},{"key":"e_1_3_1_123_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.adhoc.2023.103371"},{"key":"e_1_3_1_124_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ast.2018.10.027"},{"key":"e_1_3_1_125_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.iot.2023.100867"},{"key":"e_1_3_1_126_2","series-title":"Proceedings of Machine Learning Research","first-page":"2256","volume-title":"Proceedings of the 32nd International Conference on Machine Learning","volume":"37","author":"Sohl-Dickstein Jascha","year":"2015","unstructured":"Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol. 37). PMLR, Lille, France, 2256\u20132265."},{"key":"e_1_3_1_127_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2022.3208457"},{"key":"e_1_3_1_128_2","unstructured":"Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement learning: An introduction. A Bradford Book Cambridge MA USA."},{"key":"e_1_3_1_129_2","unstructured":"Richard S. Sutton David McAllester Satinder Singh and Yishay Mansour. 1999. Policy gradient methods for reinforcement learning with function approximation. In Proceedings of the 13th International Conference on Neural Information Processing Systems (Denver CO) (NIPS\u201999). MIT Press Cambridge MA USA 1057\u20131063."},{"key":"e_1_3_1_130_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSAC.2021.3126073"},{"key":"e_1_3_1_131_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11798."},{"key":"e_1_3_1_132_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2017.8202133"},{"key":"e_1_3_1_133_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSAC.2021.3088633"},{"key":"e_1_3_1_134_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICUAS.2018.8453386"},{"key":"e_1_3_1_135_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCCN.2021.3063170"},{"key":"e_1_3_1_136_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2020.2973193"},{"key":"e_1_3_1_137_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCC52777.2021.9580299"},{"key":"e_1_3_1_138_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2023.110604"},{"key":"e_1_3_1_139_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2021.3059691"},{"key":"e_1_3_1_140_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3069908"},{"key":"e_1_3_1_141_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2022.3153585"},{"key":"e_1_3_1_142_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2024.3360078"},{"key":"e_1_3_1_143_2","doi-asserted-by":"publisher","DOI":"10.1109\/LWC.2023.3274535"},{"key":"e_1_3_1_144_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNSE.2020.3014385"},{"key":"e_1_3_1_145_2","doi-asserted-by":"publisher","DOI":"10.5555\/3045390.3045601"},{"key":"e_1_3_1_146_2","doi-asserted-by":"publisher","unstructured":"C. J. C. H. Watkins and P. Dayan. 1992. Q-learning. Mach Learn 8 (1992) 279\u2013292. 10.1007\/BF00992698","DOI":"10.1007\/BF00992698"},{"key":"e_1_3_1_147_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2022.3160887"},{"key":"e_1_3_1_148_2","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-016-0043-6"},{"key":"e_1_3_1_149_2","doi-asserted-by":"publisher","DOI":"10.1109\/WCNC51071.2022.9771555"},{"key":"e_1_3_1_150_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN60899.2024.10650538"},{"key":"e_1_3_1_151_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2024.127904"},{"key":"e_1_3_1_152_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2020.3027586"},{"key":"e_1_3_1_153_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCCN.2020.3027696"},{"key":"e_1_3_1_154_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11276-023-03488-1"},{"key":"e_1_3_1_155_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2022.3144277"},{"key":"e_1_3_1_156_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPSN61024.2024.00041"},{"key":"e_1_3_1_157_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIV.2023.3298292"},{"key":"e_1_3_1_158_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIEA.2016.7603784"},{"key":"e_1_3_1_159_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2022.3143175"},{"key":"e_1_3_1_160_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.dcan.2023.07.006"},{"key":"e_1_3_1_161_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCC49849.2020.9238842"},{"key":"e_1_3_1_162_2","doi-asserted-by":"publisher","DOI":"10.1109\/WCNC51071.2022.9771769"},{"key":"e_1_3_1_163_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3070908"},{"key":"e_1_3_1_164_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2023.3280848"},{"key":"e_1_3_1_165_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2020.3015578"},{"key":"e_1_3_1_166_2","doi-asserted-by":"publisher","DOI":"10.1109\/TWC.2021.3121584"},{"key":"e_1_3_1_167_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2020.3047800"},{"key":"e_1_3_1_168_2","doi-asserted-by":"publisher","DOI":"10.1109\/TWC.2022.3204794"},{"key":"e_1_3_1_169_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSAC.2021.3088669"},{"key":"e_1_3_1_170_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2021.3133278"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3769426","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T14:15:09Z","timestamp":1761401709000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3769426"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,25]]},"references-count":169,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,3,31]]}},"alternative-id":["10.1145\/3769426"],"URL":"https:\/\/doi.org\/10.1145\/3769426","relation":{"has-preprint":[{"id-type":"doi","id":"10.36227\/techrxiv.172893546.67426756\/v1","asserted-by":"object"},{"id-type":"doi","id":"10.36227\/techrxiv.172893546.67426756\/v2","asserted-by":"object"}]},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,25]]},"assertion":[{"value":"2024-12-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}