{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,15]],"date-time":"2026-03-15T15:32:10Z","timestamp":1773588730930,"version":"3.50.1"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62306088"],"award-info":[{"award-number":["62306088"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100005046","name":"Natural Science Foundation of Heilongjiang Province","doi-asserted-by":"crossref","award":["YQ2024F007"],"award-info":[{"award-number":["YQ2024F007"]}],"id":[{"id":"10.13039\/501100005046","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012130","name":"Aviation Science Foundation of China","doi-asserted-by":"crossref","award":["2023Z021077001"],"award-info":[{"award-number":["2023Z021077001"]}],"id":[{"id":"10.13039\/501100012130","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Auton. Adapt. Syst."],"published-print":{"date-parts":[[2026,3,31]]},"abstract":"<jats:p>Unmanned combat aerial vehicle (UCAV) within-visual-range (WVR) engagement, referring to a fight between two or more UCAVs at close quarters, plays a decisive role on the aerial battlefields. With the development of artificial intelligence, WVR engagement progressively advances toward intelligent and autonomous modes. However, autonomous WVR engagement policy learning is hindered by challenges such as weak exploration capabilities, low learning efficiency, and unrealistic simulated environments. To overcome these challenges, we propose a novel imitative reinforcement learning framework, which efficiently leverages expert data while enabling autonomous exploration. The proposed framework not only enhances learning efficiency through expert imitation but also ensures adaptability to dynamic environments via autonomous exploration with reinforcement learning. Therefore, the proposed framework can learn a successful policy of \u201cpursuit-lock-launch\u201d for UCAVs. To support data-driven learning, we establish an environment based on the Harfang3D sandbox. The extensive experimental results indicate that the proposed framework excels in this multistage task and significantly outperforms state-of-the-art reinforcement learning and imitation learning methods. Thanks to the ability of imitating experts and autonomous exploration, our framework can quickly learn the critical knowledge in complex aerial combat tasks, achieving up to a 100% success rate and demonstrating excellent robustness.<\/jats:p>","DOI":"10.1145\/3750733","type":"journal-article","created":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T17:12:41Z","timestamp":1753377161000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["An Imitative Reinforcement Learning Framework for Pursuit-Lock-Launch Missions"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7965-598X","authenticated-orcid":false,"given":"Siyuan","family":"Li","sequence":"first","affiliation":[{"name":"Faculty of Computing, Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-8034-8975","authenticated-orcid":false,"given":"Rongchang","family":"Zuo","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-1663-148X","authenticated-orcid":false,"given":"Bofei","family":"Liu","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-1224-3137","authenticated-orcid":false,"given":"Yaoyu","family":"He","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6568-1335","authenticated-orcid":false,"given":"Peng","family":"Liu","sequence":"additional","affiliation":[{"name":"Faculty of Computing, Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-5037-8813","authenticated-orcid":false,"given":"Yingnan","family":"Zhao","sequence":"additional","affiliation":[{"name":"Harbin Engineering University, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,3,10]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"Proceedings of the International Command and Control Research & Technology Symposium","author":"Van\u2019t Wout Magdalena C.","year":"2019","unstructured":"Magdalena C. Van\u2019t Wout, Shaun V. Ball, and Rudolph Oosthuizen. 2019. A framework for implementing a data science capability in a military intelligence system. In Proceedings of the International Command and Control Research & Technology Symposium. Track 3: Battlefields of the Future and the Internet of Intelligent Things."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/1869397.1869402"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.3390\/electronics11030467"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.adhoc.2020.102324"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.4324\/9781315243092-17"},{"key":"e_1_3_2_7_2","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton Richard S.","year":"2018","unstructured":"Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction. MIT Press."},{"key":"e_1_3_2_8_2","article-title":"Autonomous helicopter flight via reinforcement learning","volume":"16","author":"Kim H.","year":"2003","unstructured":"H. Kim, Michael Jordan, Shankar Sastry, and Andrew Ng. 2003. Autonomous helicopter flight via reinforcement learning. In Advances in Neural Information Processing Systems, Vol. 16.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICTC55111.2022.9778652"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1155\/2022\/4186303"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1049\/cit2.12195"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.3390\/machines10111033"},{"issue":"1","key":"e_1_3_2_13_2","first-page":"3657814","article-title":"Autonomous maneuver decision of UCAV air combat based on double deep q network algorithm and stochastic game theory","volume":"2023","author":"Cao Yuan","year":"2023","unstructured":"Yuan Cao, Ying-Xin Kou, Zhan-Wu Li, and An Xu. 2023. Autonomous maneuver decision of UCAV air combat based on double deep q network algorithm and stochastic game theory. International Journal of Aerospace Engineering 2023, 1 (2023), 3657814.","journal-title":"International Journal of Aerospace Engineering"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3653979"},{"issue":"6","key":"e_1_3_2_15_2","first-page":"1702","article-title":"Intelligent avoidance decision of unmanned aerial vehicles based on deep reinforcement learning algorithm","volume":"45","author":"Wu Fengguo","year":"2023","unstructured":"Fengguo Wu, Wei Tao, Hui Li, and Zhang Jianwei. 2023. Intelligent avoidance decision of unmanned aerial vehicles based on deep reinforcement learning algorithm. Systems Engineering and Electronic Technology 45, 6 (2023), 1702\u20131711.","journal-title":"Systems Engineering and Electronic Technology"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2018.8594352"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8793979"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2017.7989381"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2018.8460655"},{"key":"e_1_3_2_20_2","first-page":"1861","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Haarnoja Tuomas","year":"2018","unstructured":"Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the International Conference on Machine Learning. PMLR, 1861\u20131870."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.2514\/6.2004-4923"},{"key":"e_1_3_2_22_2","unstructured":"Muhammed Murat \u00d6zbek S\u00fcleyman Y\u0131ld\u0131r\u0131m Muhammet Aksoy Eric Kernin and Emre Koyuncu. 2022. Harfang3d dog-fight sandbox: A reinforcement learning research platform for the customized control tasks of fighter aircrafts. arXiv:2210.07282. Retrieved from https:\/\/arxiv.org\/abs\/2210.07282"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1002\/9781118557426.ch1"},{"key":"e_1_3_2_24_2","article-title":"Actor-critic algorithms","volume":"12","author":"Konda Vijay","year":"1999","unstructured":"Vijay Konda and John Tsitsiklis. 1999. Actor-critic algorithms. In Advances in Neural Information Processing Systems, Vol. 12.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2966237"},{"issue":"1","key":"e_1_3_2_26_2","first-page":"107","article-title":"A decision-making method for air combat maneuver based on hybrid deep learning network","volume":"31","author":"Li Bo","year":"2022","unstructured":"Bo Li, Shiyang Liang, Daqing Chen, and Xitong Li. 2022. A decision-making method for air combat maneuver based on hybrid deep learning network. Chinese Journal of Electronics 31, 1 (2022), 107\u2013115.","journal-title":"Chinese Journal of Electronics"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3054912"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIV.2020.3002505"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICUAS54217.2022.9836131"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2021.103500"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.3390\/electronics9071121"},{"key":"e_1_3_2_32_2","article-title":"Generative adversarial imitation learning","volume":"29","author":"Ho Jonathan","year":"2016","unstructured":"Jonathan Ho and Stefano Ermon. 2016. Generative adversarial imitation learning. In Advances in Neural Information Processing Systems, Vol. 29.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_33_2","first-page":"1928","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Mnih Volodymyr","year":"2016","unstructured":"Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016. Asynchronous methods for deep reinforcement learning. In Proceedings of the International Conference on Machine Learning, 1928\u20131937."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v30i1.10295"},{"key":"e_1_3_2_35_2","article-title":"OpenAI pieter abbeel, and Igor Mordatch. Multi-agent actor-critic for mixed cooperative-competitive environments","volume":"30","author":"Lowe Ryan","year":"2017","unstructured":"Ryan Lowe, Y. I. Wu, Aviv Tamar, and Jean Harb. 2017. OpenAI pieter abbeel, and Igor Mordatch. Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_36_2","unstructured":"Timothy P. Lillicrap Jonathan J. Hunt Alexander Pritzel Nicolas Heess Tom Erez Yuval Tassa David Silver and Daan Wierstra. 2015. Continuous control with deep reinforcement learning. arXiv:1509.02971. Retrieved from https:\/\/arxiv.org\/abs\/1509.02971"},{"issue":"6","key":"e_1_3_2_37_2","first-page":"1547","article-title":"Multi-dimensional decision-making for UAV air combat based on hierarchical reinforcement learning","volume":"44","author":"Zhang Jiandong","year":"2023","unstructured":"Jiandong Zhang, Dinghan Wang, Qiming Yang, Guoqing Shi, Yi Lu, and Yaozhong Zhang. Multi-dimensional decision-making for UAV air combat based on hierarchical reinforcement learning. Acta Armamentarii 44, 6 (2023), 1547.","journal-title":"Acta Armamentarii"},{"key":"e_1_3_2_38_2","first-page":"20132","article-title":"A minimalist approach to offline reinforcement learning","volume":"34","author":"Fujimoto Scott","year":"2021","unstructured":"Scott Fujimoto and Shixiang Shane Gu. 2021. A minimalist approach to offline reinforcement learning. In Advances in Neural Information Processing Systems, Vol. 34, 20132\u201320145.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_39_2","first-page":"1587","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Fujimoto Scott","year":"2018","unstructured":"Scott Fujimoto, Herke Hoof, and David Meger. 2018. Addressing function approximation error in actor-critic methods. In Proceedings of the International Conference on Machine Learning. PMLR, 1587\u20131596."},{"key":"e_1_3_2_40_2","first-page":"5999","volume-title":"Proceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI \u201924)","author":"Niu Hui","year":"2004","unstructured":"Hui Niu, Siyuan Li, Jiahao Zheng, Zhouchi Lin, Bo An, Jian Li, and Jian Guo. 2004. IMM: An imitative reinforcement learning approach with predictive representation learning for automatic market making. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI \u201924). Kate Larson (Ed.), International Joint Conferences on Artificial Intelligence Organization, 5999\u20136007."},{"key":"e_1_3_2_41_2","first-page":"6008","volume-title":"Proceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI \u201924)","author":"Niu Hui","year":"2004","unstructured":"Hui Niu, Siyuan Li, and Jian Li. 2004. Macmic: Executing iceberg orders via hierarchical reinforcement learning. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI \u201924). Kate Larson (Ed.), International Joint Conferences on Artificial Intelligence Organization, 6008\u20136016."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3511808.3557363"},{"key":"e_1_3_2_43_2","unstructured":"Jingliang Duan Wenxuan Wang Liming Xiao Jiaxin Gao and Shengbo Eben Li. 2023. DSAC-T: Distributional soft actor-critic with three refinements. arXiv:2310.05858. Retrieved from https:\/\/arxiv.org\/abs\/2310.05858"},{"key":"e_1_3_2_44_2","first-page":"15737","article-title":"Error bounds of imitating policies and environments","volume":"33","author":"Xu Tian","year":"2020","unstructured":"Tian Xu, Ziniu Li, and Yang Yu. 2020. Error bounds of imitating policies and environments. In Advances in Neural Information Processing Systems, Vol. 33, 15737\u201315749.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_45_2","first-page":"661","volume-title":"Proceedings of the 13th International Conference on Artificial Intelligence and Statistics","author":"Ross St\u00e9phane","year":"2010","unstructured":"St\u00e9phane Ross and Drew Bagnell. 2010. Efficient reductions for imitation learning. In Proceedings of the 13th International Conference on Artificial Intelligence and Statistics. JMLR Workshop and Conference Proceedings, 661\u2013668."},{"key":"e_1_3_2_46_2","first-page":"627","volume-title":"Proceedings of the 14th International Conference on Artificial Intelligence and Statistics","author":"Ross St\u00e9phane","year":"2011","unstructured":"St\u00e9phane Ross, Geoffrey Gordon, and Drew Bagnell. 2011. A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics. JMLR Workshop and Conference Proceedings, 627\u2013635."}],"container-title":["ACM Transactions on Autonomous and Adaptive Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3750733","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,15]],"date-time":"2026-03-15T14:10:44Z","timestamp":1773583844000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3750733"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,10]]},"references-count":45,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,3,31]]}},"alternative-id":["10.1145\/3750733"],"URL":"https:\/\/doi.org\/10.1145\/3750733","relation":{},"ISSN":["1556-4665","1556-4703"],"issn-type":[{"value":"1556-4665","type":"print"},{"value":"1556-4703","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,10]]},"assertion":[{"value":"2024-04-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-19","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}