{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T11:09:08Z","timestamp":1779275348137,"version":"3.51.4"},"reference-count":51,"publisher":"American Institute of Aeronautics and Astronautics (AIAA)","issue":"6","funder":[{"DOI":"10.13039\/501100005073","name":"Agency for Defense Development","doi-asserted-by":"publisher","award":["UD240002SD"],"award-info":[{"award-number":["UD240002SD"]}],"id":[{"id":"10.13039\/501100005073","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["arc.aiaa.org"],"crossmark-restriction":true},"short-container-title":["Journal of Aerospace Information Systems"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:p>This paper presents a novel multi-agent reinforcement learning (MARL) approach that incorporates agent priorities to address weapon\u2013target assignment (WTA) with constraints, such as heterogeneous engagement time windows. The proposed approach begins by defining the decentralized Markov decision process (Dec-MDP) formulation for WTA involving heterogeneous, multiple agents. Our approach employs a hierarchical structure for MARL training, comprising an agent selector and a target selector, which sequentially determine the order of agents for assignment, i.e., preferred shooter selection and target selection. Through experimental designs, the proposed model demonstrates its ability to generate high-quality assignment plans within a short execution time. The model demonstrates superior performance across various scenarios, achieving the lowest threat survivability with a clear advantage over other baseline methods, especially in tightly constrained scenarios. Ablation studies and qualitative analyses are conducted to illustrate the influence of key components on performance, and these qualitative studies reveal the learning mechanism in agent and target selection. Additionally, transferability tests confirm the model\u2019s applicability to unseen problem cases, where training and testing environments are different, indicating its potential for real-world adaptation in various scenarios.<\/jats:p>","DOI":"10.2514\/1.i011676","type":"journal-article","created":{"date-parts":[[2026,4,26]],"date-time":"2026-04-26T17:31:01Z","timestamp":1777224661000},"page":"464-479","update-policy":"https:\/\/doi.org\/10.2514\/aiaa_crossmarkpolicy","source":"Crossref","is-referenced-by-count":0,"title":["Multi-Agent Reinforcement Learning Considering Agent Priority for Weapon\u2013Target Assignment"],"prefix":"10.2514","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7687-2513","authenticated-orcid":false,"given":"Hyungho","family":"Na","sequence":"first","affiliation":[{"name":"Ulsan National Institute of Science and Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4971-5130","authenticated-orcid":false,"given":"Jaemyung","family":"Ahn","sequence":"additional","affiliation":[{"name":"Korea Advanced Institute of Science and Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1798-1306","authenticated-orcid":false,"given":"Il-Chul","family":"Moon","sequence":"additional","affiliation":[{"name":"Korea Advanced Institute of Science and Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1387","reference":[{"key":"r1","unstructured":"LloydS. P.WitsenhausenH. S. \u201cWeapons Allocation is NP-Complete,\u201d 1986 Summer Computer Simulation Conference, Soc. for Modelling and Simulation International (SCS), San Diego, CA, 1986, pp.\u00a01054\u20131058."},{"key":"r2","doi-asserted-by":"publisher","DOI":"10.1002\/9781119266235.ch5"},{"key":"r3","unstructured":"ShinM.K.LeeD.ChoiH.L. \u201cWeapon-Target Assignment Problem with Interference Constraints Using Mixed-Integer Linear Programming,\u201d arXiv preprint arXiv: 1911.12567, 2019. 10.48550\/arXiv.1911.12567"},{"key":"r4","doi-asserted-by":"publisher","DOI":"10.2514\/6.2020-0388"},{"key":"r5","doi-asserted-by":"publisher","DOI":"10.2514\/1.I011041"},{"key":"r6","doi-asserted-by":"publisher","DOI":"10.1287\/opre.1070.0440"},{"key":"r7","doi-asserted-by":"publisher","DOI":"10.5139\/JKSAS.2019.47.11.787"},{"key":"r8","doi-asserted-by":"publisher","DOI":"10.1080\/02533839.2002.9670703"},{"key":"r9","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCB.2003.808174"},{"key":"r10","doi-asserted-by":"crossref","unstructured":"BoZ.Feng-XingZ.Jia-HuaW. \u201cA Novel Approach to Solving Weapon-Target Assignment Problem Based on Hybrid Particle Swarm Optimization Algorithm,\u201d Proceedings of 2011 International Conference on Electronic & Mechanical Engineering and Information Technology, Vol.\u00a03, Inst. of Electrical and Electronics Engineers, New York, 2011, pp.\u00a01385\u20131387. 10.1109\/EMEIT.2011.6023352","DOI":"10.1109\/EMEIT.2011.6023352"},{"key":"r11","doi-asserted-by":"crossref","unstructured":"WangS.ChenW. \u201cSolving Weapon-Target Assignment Problems by Cultural Particle Swarm Optimization,\u201d 2012 4th International Conference on Intelligent Human-Machine Systems and Cybernetics, Vol.\u00a01, Inst. of Electrical and Electronics Engineers, New York, 2012, pp.\u00a0141\u2013144. 10.1109\/IHMSC.2012.41","DOI":"10.1109\/IHMSC.2012.41"},{"key":"r12","doi-asserted-by":"publisher","DOI":"10.1109\/TSMC.2017.2784187"},{"key":"r13","doi-asserted-by":"publisher","DOI":"10.5787\/39-2-115"},{"key":"r14","unstructured":"BelloI.PhamH.LeQ. V.NorouziM.BengioS. \u201cNeural Combinatorial Optimization with Reinforcement Learning,\u201d arXiv preprint arXiv: 1611.09940, 2016. 10.48550\/arXiv.1611.09940"},{"key":"r15","volume-title":"28th Annual Conference on Neural Information Processing Systems (NeurIPS)","volume":"28","author":"Vinyals O.","year":"2015"},{"key":"r16","volume-title":"Advances in Neural Information Processing Systems","volume":"31","author":"Nazari M.","year":"2018"},{"key":"r17","unstructured":"KoolW.van HoofH.WellingM. \u201cAttention, Learn to Solve Routing Problems!,\u201d International Conference on Learning Representations, New Orleans, LA, 2019. 10.48550\/arXiv.1803.08475"},{"key":"r18","doi-asserted-by":"crossref","unstructured":"LiX.LuoW.YuanM.WangJ.LuJ.WangJ.L\u00fcJ.ZengJ. \u201cLearning to Optimize Industry-Scale Dynamic Pickup and Delivery Problems,\u201d 2021 IEEE 37th International Conference on Data Engineering (ICDE), Inst. of Electrical and Electronics Engineers, New York, 2021, pp.\u00a02511\u20132522. 10.1109\/ICDE51399.2021.00283","DOI":"10.1109\/ICDE51399.2021.00283"},{"key":"r19","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2021.3056120"},{"key":"r20","unstructured":"YangY.LuoR.LiM.ZhouM.ZhangW.WangJ. \u201cMean Field Multi-Agent Reinforcement Learning,\u201d 35th International Conference on Machine Learning (ICML 2018), Stockholm, Sweden, Proceedings of Machine Learning Research, Vol.\u00a080, 2018, pp.\u00a05571\u20135580. 10.48550\/arXiv.1802.05438"},{"key":"r21","doi-asserted-by":"publisher","DOI":"10.9766\/KIMST.2020.23.4.337"},{"key":"r22","doi-asserted-by":"crossref","unstructured":"ChengQ.ChenD.GongJ. \u201cWeapon-Target Assignment of Ballistic Missiles Based on Q-Learning and Genetic Algorithm,\u201d 2021 IEEE International Conference on Unmanned Systems (ICUS), Inst. of Electrical and Electronics Engineers, New York, 2021, pp.\u00a0908\u2013912. 10.1109\/ICUS52573.2021.9641190","DOI":"10.1109\/ICUS52573.2021.9641190"},{"key":"r23","doi-asserted-by":"publisher","DOI":"10.1109\/TSMC.2021.3096997"},{"key":"r24","doi-asserted-by":"publisher","DOI":"10.1016\/j.dt.2022.04.001"},{"key":"r25","doi-asserted-by":"publisher","DOI":"10.2514\/1.I011150"},{"key":"r26","doi-asserted-by":"crossref","unstructured":"ByunM.NaH.MoonI.C. \u201cTime-Efficient Weapon-Target Assignment by Actor-Critic Reinforcement,\u201d 2023 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Inst. of Electrical and Electronics Engineers, New York, 2023, pp.\u00a01419\u20131424. 10.1109\/SMC53992.2023.10394214","DOI":"10.1109\/SMC53992.2023.10394214"},{"key":"r27","doi-asserted-by":"publisher","DOI":"10.1007\/s12530-024-09587-4"},{"key":"r28","doi-asserted-by":"crossref","unstructured":"WangY.DuG.LiuY.WuX. \u201cAnti-Missile Firepower Allocation Based on Multi-Agent Reinforcement Learning,\u201d International Conference on Autonomous Unmanned Systems, Vol.\u00a01171, Springer Nature Singapore, Singapore, 2023, pp.\u00a0160\u2013169. 10.1007\/978-981-97-1083-6_15","DOI":"10.1007\/978-981-97-1083-6_15"},{"key":"r29","doi-asserted-by":"crossref","unstructured":"DroriI.KharkarA.SickingerW. R.KatesB.MaQ.GeS.DolevE.DietrichB.WilliamsonD. P.UdellM. \u201cLearning to Solve Combinatorial Optimization Problems on Real-World Graphs in Linear Time,\u201d 2020 19th IEEE International Conference on Machine Learning and Applications (ICMLA), Inst. of Electrical and Electronics Engineers, New York, 2020, pp.\u00a019\u201324. 10.1109\/ICMLA51294.2020.00013","DOI":"10.1109\/ICMLA51294.2020.00013"},{"key":"r30","doi-asserted-by":"publisher","DOI":"10.1016\/j.cor.2021.105400"},{"key":"r31","unstructured":"YoungB. W. \u201cFuture Integrated Fire Control,\u201d 10th International Command and Control Research and Technology Symposium (ICCRTS 2005): The Future of Command and Control, McLean, VA, 2005, https:\/\/api.semanticscholar.org\/CorpusID:171089665."},{"key":"r32","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton R. S.","year":"2018"},{"key":"r33","doi-asserted-by":"publisher","DOI":"10.1613\/jair.2447"},{"key":"r34","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-28929-8"},{"issue":"178","key":"r35","first-page":"1","volume":"21","author":"Rashid T.","year":"2020","journal-title":"Journal of Machine Learning Research"},{"key":"r36","unstructured":"WangJ.RenZ.LiuT.YuY.ZhangC. \u201cQplex: Duplex Dueling Multi-Agent q-Learning,\u201d arXiv preprint arXiv: 2008.01062, 2020. 10.48550\/arXiv.2008.01062"},{"key":"r37","unstructured":"NaH.SeoY.MoonI.C. \u201cEfficient Episodic Memory Utilization of Cooperative Multi-Agent Reinforcement Learning,\u201d arXiv preprint arXiv: 2403.01112, 2024. 10.48550\/arXiv.2403.01112"},{"key":"r38","unstructured":"NaH.MoonI.C. \u201cLAGMA: LAtent Goal-Guided Multi-Agent Reinforcement Learning,\u201d arXiv preprint arXiv: 2405.19998, 2024. 10.48550\/arXiv.2405.19998"},{"key":"r39","unstructured":"NaH.LeeK.LeeS.MoonI.C. \u201cTrajectory-Class-Aware Multi-Agent Reinforcement Learning,\u201d 13th International Conference on Learning Representations, 2025. 10.48550\/arXiv.2503.01440"},{"key":"r40","unstructured":"FoersterJ.AssaelI. A.De FreitasN.WhitesonS. \u201cLearning to Communicate with Deep Multi-Agent Reinforcement Learning,\u201d Advances in Neural Information Processing Systems, Vol.\u00a029, Curran Assoc., Red Hook, NY, 2016. 10.48550\/arXiv.1605.06676"},{"key":"r41","unstructured":"DasA.GervetT.RomoffJ.BatraD.ParikhD.RabbatM.PineauJ. \u201cTarmac: Targeted Multi-Agent Communication,\u201d 36th International Conference on Machine Learning, Vol.\u00a097, PMLR, 2019, pp.\u00a01538\u20131546. 10.48550\/arXiv.1810.11187"},{"key":"r42","doi-asserted-by":"publisher","DOI":"10.1007\/s00500-022-06820-7"},{"key":"r43","doi-asserted-by":"publisher","DOI":"10.1002\/9781118557426"},{"key":"r44","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2021.03.091"},{"key":"r45","unstructured":"VaswaniA.ShazeerN.ParmarN.UszkoreitJ.JonesL.GomezA. N.Kaiser\u0141.PolosukhinI. \u201cAttention Is All You Need,\u201d 31st International Conference on Neural Information Processing Systems, Long Beach, CA, 2017, pp.\u00a06000\u20136010. 10.5555\/3295222.3295349"},{"key":"r46","doi-asserted-by":"crossref","unstructured":"BelloI.ZophB.VaswaniA.ShlensJ.LeQ. V. \u201cAttention Augmented Convolutional Networks,\u201d Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul Korea, 2019, pp.\u00a03286\u20133295. 10.1109\/ICCV.2019.00338","DOI":"10.1109\/ICCV.2019.00338"},{"key":"r47","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-022-0271-y"},{"key":"r48","unstructured":"KumarA.IrsoyO.OndruskaP.IyyerM.BradburyJ.GulrajaniI.ZhongV.PaulusR.SocherR. \u201cAsk Me Anything: Dynamic Memory Networks for Natural Language Processing,\u201d 33rd International Conference on Machine Learning, Vol.\u00a048, PMLR, New York, 2016, pp.\u00a01378\u20131387. 10.48550\/arXiv.1506.07285"},{"key":"r49","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.3019893"},{"key":"r50","unstructured":"NazariM.OroojlooyA.SnyderL. V.Tak\u00e1\u010dM. \u201cReinforcement Learning for Solving the Vehicle Routing Problem,\u201d arXiv preprint arXiv: 1802.04240, 2018. 10.48550\/arXiv.1802.04240"},{"key":"r51","doi-asserted-by":"publisher","DOI":"10.1109\/TAES.2019.2923331"}],"container-title":["Journal of Aerospace Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/arc.aiaa.org\/doi\/pdf\/10.2514\/1.I011676","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T10:44:21Z","timestamp":1779273861000},"score":1,"resource":{"primary":{"URL":"https:\/\/arc.aiaa.org\/doi\/10.2514\/1.I011676"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6]]},"references-count":51,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["10.2514\/1.I011676"],"URL":"https:\/\/doi.org\/10.2514\/1.i011676","relation":{},"ISSN":["1940-3151","2327-3097"],"issn-type":[{"value":"1940-3151","type":"print"},{"value":"2327-3097","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6]]},"assertion":[{"value":"2025-04-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-26","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}