{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T04:40:56Z","timestamp":1776400856482,"version":"3.51.2"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2022,11,9]],"date-time":"2022-11-09T00:00:00Z","timestamp":1667952000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2020AAA0106000"],"award-info":[{"award-number":["2020AAA0106000"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U1936217, 61971267, 61972223, 61941117, 61861136003"],"award-info":[{"award-number":["U1936217, 61971267, 61972223, 61941117, 61861136003"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Beijing Natural Science Foundation","award":["L182038"],"award-info":[{"award-number":["L182038"]}]},{"DOI":"10.13039\/501100017582","name":"Beijing National Research Center for Information Science and Technology","doi-asserted-by":"crossref","award":["20031887521"],"award-info":[{"award-number":["20031887521"]}],"id":[{"id":"10.13039\/501100017582","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Tsinghua University \u2014 Tencent Joint Laboratory for Internet Innovation Technology"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2023,2,28]]},"abstract":"<jats:p>With the advent of the COVID-19 pandemic, the shortage in medical resources became increasingly more evident. Therefore, efficient strategies for medical resource allocation are urgently needed. However, conventional rule-based methods employed by public health experts have limited capability in dealing with the complex and dynamic pandemic-spreading situation. In addition, model-based optimization methods such as dynamic programming (DP) fail to work since we cannot obtain a precise model in real-world situations most of the time. Model-free reinforcement learning (RL) is a powerful tool for decision-making; however, three key challenges exist in solving this problem via RL: (1) complex situations and countless choices for decision-making in the real world; (2) imperfect information due to the latency of pandemic spreading; and (3) limitations on conducting experiments in the real world since we cannot set up pandemic outbreaks arbitrarily. In this article, we propose a hierarchical RL framework with several specially designed components. We design a decomposed action space with a corresponding training algorithm to deal with the countless choices, ensuring efficient and real-time strategies. We design a recurrent neural network\u2013based framework to utilize the imperfect information obtained from the environment. We also design a multi-agent voting method, which modifies the decision-making process considering the randomness during model training and, thus, improves the performance. We build a pandemic-spreading simulator based on real-world data, serving as the experimental platform. We then conduct extensive experiments. The results show that our method outperforms all baselines, which reduces infections and deaths by 14.25% on average without the multi-agent voting method and up to 15.44% with it.<\/jats:p>","DOI":"10.1145\/3552436","type":"journal-article","created":{"date-parts":[[2022,7,30]],"date-time":"2022-07-30T11:04:32Z","timestamp":1659179072000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Hierarchical Multi-agent Model for Reinforced Medical Resource Allocation with Imperfect Information"],"prefix":"10.1145","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7109-3588","authenticated-orcid":false,"given":"Qianyue","family":"Hao","sequence":"first","affiliation":[{"name":"Beijing National Research Center for Information Science and Technology (BNRist), Department of Electronic Engineering, Tsinghua University, Haidian, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5720-4026","authenticated-orcid":false,"given":"Fengli","family":"Xu","sequence":"additional","affiliation":[{"name":"Knowledge Lab, Department of Sociology, University of Chicago, IL, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2605-749X","authenticated-orcid":false,"given":"Lin","family":"Chen","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Kowloon, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6026-1083","authenticated-orcid":false,"given":"Pan","family":"Hui","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Kowloon, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5617-1659","authenticated-orcid":false,"given":"Yong","family":"Li","sequence":"additional","affiliation":[{"name":"Beijing National Research Center for Information Science and Technology (BNRist), Department of Electronic Engineering, Tsinghua University, Haidian, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,11,9]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2021.103500"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.0906910106"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3481617"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1093\/jamia\/ocaa324"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1177\/0272989X12437247"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3470890"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.comcom.2019.12.054"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/d14-1179"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.idm.2020.04.001"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1056\/NEJMsb2005114"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/2882969"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.chaos.2020.110599"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1111\/bjhp.12439"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1056\/NEJMoa20020"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3461645"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41591-020-0869-5"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0140-6736(20)30183-5"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2020.2980158"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10479-020-03871-7"},{"key":"e_1_3_2_21_2","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations ICLR 2015 San Diego CA USA May 7-9 2015 Conference Track Proceedings Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1412.6980."},{"key":"e_1_3_2_22_2","unstructured":"Vijay R. Konda and John N. Tsitsiklis. 2000. Actor-critic algorithms. In Advances in Neural Information Processing Systems 12 [NIPS Conference Denver Colorado USA November 29 - December 4 1999] Sara A. Solla Todd K. Leen and Klaus-Robert M\u00fcller (Eds.). The MIT Press 1008\u20131014. http:\/\/papers.nips.cc\/paper\/1786-actor-critic-algorithms."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0251550"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313433"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1016\/0025-5564(95)92756-5"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1056\/NEJMoa2001316"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.trc.2021.103059"},{"key":"e_1_3_2_28_2","unstructured":"Timothy P. Lillicrap Jonathan J. Hunt Alexander Pritzel Nicolas Heess Tom Erez Yuval Tassa David Silver and Daan Wierstra. 2015. Continuous control with deep reinforcement learning. In 4th International Conference on Learning Representations ICLR 2016 San Juan Puerto Rico May 2-4 2016 Conference Track Proceedings Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1509.02971."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219993"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2020\/186"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1126\/sciadv.abf1374"},{"key":"e_1_3_2_32_2","series-title":"Proceedings of the 33rd International Conference on Machine Learning","first-page":"1928","volume":"48","author":"Mnih Volodymyr","year":"2016","unstructured":"Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016. Asynchronous methods for deep reinforcement learning. In Proceedings of the 33rd International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol. 48). New York, New York, 1928\u20131937."},{"key":"e_1_3_2_33_2","article-title":"Playing Atari with deep reinforcement learning","author":"Mnih Volodymyr","year":"2013","unstructured":"Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013. Playing Atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602","journal-title":"arXiv preprint arXiv:1312.5602"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3317572"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-020-79147-8"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1007\/s13042-020-01167-7"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0140-6736(09)60137-9"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1126\/sciadv.aap7885"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1056\/NEJMp2006141"},{"key":"e_1_3_2_41_2","unstructured":"Sara J. Rosenbaum. 2011. Ethical considerations for decision making regarding allocation of mechanical ventilators during a severe influenza pandemic or other public health emergency. (2011)."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2006.246725"},{"key":"e_1_3_2_43_2","unstructured":"Tom Schaul John Quan Ioannis Antonoglou and David Silver. 2015. Prioritized experience replay. In 4th International Conference on Learning Representations ICLR 2016 San Juan Puerto Rico May 2-4 2016 Conference Track Proceedings Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1511.05952."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41591-021-01379-6"},{"key":"e_1_3_2_45_2","article-title":"Proximal policy optimization algorithms","author":"Schulman John","year":"2017","unstructured":"John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. CoRR","journal-title":"CoRR"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1258\/147775006779151201"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature24270"},{"key":"e_1_3_2_48_2","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton Richard S.","year":"2018","unstructured":"Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction. MIT Press."},{"key":"e_1_3_2_49_2","unstructured":"Richard S. Sutton David A. McAllester Satinder Singh and Yishay Mansour. 1999. Policy gradient methods for reinforcement learning with function approximation. In Advances in Neural Information Processing Systems 12 [NIPS Conference Denver Colorado USA November 29 - December 4 1999] Sara A. Solla Todd K. Leen and Klaus-Robert M\u00fcller (Eds.). The MIT Press 1057\u20131063. http:\/\/papers.nips.cc\/paper\/1713-policy-gradient-methods-for-reinforcementlearning-with-function-approximation."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3469860"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2020.3017079"},{"key":"e_1_3_2_52_2","unstructured":"Hado Van Hasselt Arthur Guez and David Silver. 2016. Deep reinforcement learning with double q-learning. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence February 12-17 2016 Phoenix Arizona USA Dale Schuurmans and Michael P. Wellman (Eds.). AAAI Press 2094\u20132100. http:\/\/www.aaai.org\/ocs\/index.php\/AAAI\/AAAI16\/paper\/view\/12389"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.2807\/1560-7917.ES.2020.25.13.2000323"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-019-1724-z"},{"key":"e_1_3_2_55_2","unstructured":"Ziyu Wang Tom Schaul Matteo Hessel Hado Hasselt Marc Lanctot and Nando Freitas. 2016. Dueling network architectures for deep reinforcement learning. In Proceedings of the 33nd International Conference on Machine Learning ICML 2016 New York City NY USA June 19-24 2016 (JMLR Workshop and Conference Proceedings Vol. 48) Maria-Florina Balcan and Kilian Q. Weinberger (Eds.). JMLR.org 1995\u20132003. http:\/\/proceedings.mlr.press\/v48\/wangf16.html."},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992698"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/5.58337"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0140-6736(20)30260-9"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/MVT.2019.2903655"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1186\/1476-072X-10-26"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSG.2020.3034827"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICInfA.2018.8812452"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3552436","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3552436","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:47:41Z","timestamp":1750178861000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3552436"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,11,9]]},"references-count":61,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2023,2,28]]}},"alternative-id":["10.1145\/3552436"],"URL":"https:\/\/doi.org\/10.1145\/3552436","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,11,9]]},"assertion":[{"value":"2022-01-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-07-19","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-11-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}