{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T16:20:46Z","timestamp":1779898846612,"version":"3.53.1"},"reference-count":36,"publisher":"American Institute of Aeronautics and Astronautics (AIAA)","issue":"5","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["71871079"],"award-info":[{"award-number":["71871079"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["71971075"],"award-info":[{"award-number":["71971075"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["72271076"],"award-info":[{"award-number":["72271076"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["arc.aiaa.org"],"crossmark-restriction":true},"short-container-title":["Journal of Aerospace Information Systems"],"published-print":{"date-parts":[[2025,5]]},"abstract":"<jats:p> Deep reinforcement learning has proven highly effective in addressing sequential decision-making challenges, particularly in aerial combat involving unmanned aerial vehicles (UAVs). However, due to the prolonged decision-making process and uncertain opponent strategies in aerial combat, expert experience and manually crafted rules are often the preferred solutions. Consequently, the expert experience soft actor\u2013critic (EESAC) algorithm is proposed to tackle the problem of dynamic target assignment for UAVs. EESAC features a customized reinforcement learning framework that includes a variable-scale reward mechanism to address the challenge of sparse rewards. Furthermore, a suppression exploration algorithm is implemented to reduce frequent target reassignments by UAVs during the initial stages, ensuring more stable and efficient mission execution. EESAC and comparator algorithms were integrated into an aerial combat simulation platform for a comprehensive evaluation. The training process and game results indicate that, compared to the comparator algorithms, EESAC significantly improves both exploration efficiency and win rate. <\/jats:p>","DOI":"10.2514\/1.i011435","type":"journal-article","created":{"date-parts":[[2025,2,10]],"date-time":"2025-02-10T17:09:26Z","timestamp":1739207366000},"page":"379-390","update-policy":"https:\/\/doi.org\/10.2514\/aiaa_crossmarkpolicy","source":"Crossref","is-referenced-by-count":1,"title":["Expert Experience Soft Actor\u2013Critic for Unmanned Aerial Vehicle Dynamic Target Assignment"],"prefix":"10.2514","volume":"22","author":[{"given":"Yuxuan","family":"Chen","sequence":"first","affiliation":[{"name":"Hefei University of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"He","family":"Luo","sequence":"additional","affiliation":[{"name":"Hefei University of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2549-6212","authenticated-orcid":false,"given":"Guoqiang","family":"Wang","sequence":"additional","affiliation":[{"name":"Hefei University of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1387","reference":[{"key":"r1","doi-asserted-by":"publisher","DOI":"10.2514\/1.I010759"},{"key":"r2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cor.2005.02.039"},{"key":"r4","doi-asserted-by":"publisher","DOI":"10.1016\/j.vehcom.2022.100469"},{"key":"r5","doi-asserted-by":"publisher","DOI":"10.1007\/s10846-014-0088-8"},{"key":"r6","doi-asserted-by":"publisher","DOI":"10.1007\/s13042-015-0364-3"},{"key":"r7","doi-asserted-by":"publisher","DOI":"10.1007\/s00500-021-05675-8"},{"key":"r8","doi-asserted-by":"publisher","DOI":"10.1109\/TAES.2023.3323441"},{"key":"r9","doi-asserted-by":"publisher","DOI":"10.1016\/j.cie.2021.107717"},{"key":"r10","doi-asserted-by":"publisher","DOI":"10.1109\/JSEE.2015.00109"},{"key":"r11","doi-asserted-by":"publisher","DOI":"10.1007\/s11590-010-0259-x"},{"key":"r12","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2023.110445"},{"key":"r13","doi-asserted-by":"publisher","DOI":"10.1016\/j.cja.2017.09.005"},{"key":"r14","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2023.106404"},{"key":"r15","doi-asserted-by":"publisher","DOI":"10.2514\/1.I011041"},{"key":"r16","doi-asserted-by":"publisher","DOI":"10.1007\/s10846-013-9887-6"},{"key":"r17","doi-asserted-by":"publisher","DOI":"10.2514\/1.I010501"},{"key":"r18","doi-asserted-by":"publisher","DOI":"10.1007\/s10846-017-0493-x"},{"key":"r19","doi-asserted-by":"publisher","DOI":"10.1109\/TAES.2020.3029624"},{"key":"r20","doi-asserted-by":"publisher","DOI":"10.1007\/s10489-021-02502-3"},{"key":"r21","doi-asserted-by":"publisher","DOI":"10.1016\/j.cor.2021.105400"},{"key":"r22","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2019.04.056"},{"key":"r23","doi-asserted-by":"publisher","DOI":"10.2514\/1.I010961"},{"key":"r24","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2022.3183099"},{"key":"r25","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2022.117796"},{"key":"r26","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2020.09.018"},{"key":"r27","doi-asserted-by":"publisher","DOI":"10.1080\/0305215X.2020.1867120"},{"key":"r28","doi-asserted-by":"publisher","DOI":"10.3390\/drones6080215"},{"key":"r29","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2022.116995"},{"key":"r30","doi-asserted-by":"publisher","DOI":"10.1016\/j.tre.2022.102694"},{"key":"r31","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2943253"},{"key":"r32","doi-asserted-by":"publisher","DOI":"10.23919\/JSEE.2021.000121"},{"key":"r33","unstructured":"HaarnojaT.ZhouA.HartikainenK.TuckerG.HaS.TanJ.KumarV.ZhuH.GuptaA.AbbeelP.LevineS. \u201cSoft Actor-Critic Algorithms and Applications, arXiv preprint arXiv: 1812.05905, 2018."},{"key":"r34","unstructured":"ChristodoulouP. \u201cSoft Actor-Critic for Discrete Action Settings, arXiv preprint arXiv: 1910.07207, 2019."},{"key":"r35","doi-asserted-by":"publisher","DOI":"10.1016\/j.dt.2022.08.010"},{"key":"r36","doi-asserted-by":"publisher","DOI":"10.1126\/science.220.4598.671"},{"key":"r37","first-page":"1","author":"Piao H.","year":"2020","journal-title":"International Joint Conference on Neural Networks (IJCNN)"}],"container-title":["Journal of Aerospace Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/arc.aiaa.org\/doi\/pdf\/10.2514\/1.I011435","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,12]],"date-time":"2025-05-12T12:40:41Z","timestamp":1747053641000},"score":1,"resource":{"primary":{"URL":"https:\/\/arc.aiaa.org\/doi\/10.2514\/1.I011435"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5]]},"references-count":36,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2025,5]]}},"alternative-id":["10.2514\/1.I011435"],"URL":"https:\/\/doi.org\/10.2514\/1.i011435","relation":{},"ISSN":["1940-3151","2327-3097"],"issn-type":[{"value":"1940-3151","type":"print"},{"value":"2327-3097","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5]]},"assertion":[{"value":"2024-01-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-29","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-19","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}