{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,7]],"date-time":"2026-05-07T23:48:48Z","timestamp":1778197728072,"version":"3.51.4"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2024,3,28]],"date-time":"2024-03-28T00:00:00Z","timestamp":1711584000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100000923","name":"Australian Research Council","doi-asserted-by":"crossref","award":["DP220101360"],"award-info":[{"award-number":["DP220101360"]}],"id":[{"id":"10.13039\/501100000923","id-type":"DOI","asserted-by":"crossref"}]},{"name":"NSFC China","award":["62276196"],"award-info":[{"award-number":["62276196"]}]},{"name":"#x0201C;Medical Text Feature Representations based on Pre-trained Language Models\u201d","award":["871238"],"award-info":[{"award-number":["871238"]}]},{"name":"Faculty Research Grants","award":["DB24A4 and DB23B2"],"award-info":[{"award-number":["DB24A4 and DB23B2"]}]},{"name":"Lingnan University, Hong Kong"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2024,4,30]]},"abstract":"<jats:p>\n            Personalized clinical decision support systems are increasingly being adopted due to the emergence of data-driven technologies, with this approach now gaining recognition in critical care. The task of incorporating diverse patient conditions and treatment procedures into critical care decision-making can be challenging due to the heterogeneous nature of medical data. Advances in Artificial Intelligence (AI), particularly Reinforcement Learning (RL) techniques, enables the development of personalized treatment strategies for severe illnesses by using a learning agent to recommend optimal policies. In this study, we propose a Deep Reinforcement Learning (DRL) model with a tailored reward function and an LSTM-GRU-derived state representation to formulate optimal treatment policies for vasopressor administration in stabilizing patient physiological states in critical care settings. Using an ICU dataset and the Medical Information Mart for Intensive Care (MIMIC-III) dataset, we focus on patients with Acute Respiratory Distress Syndrome (ARDS) that has led to Sepsis, to derive optimal policies that can prioritize patient recovery over patient survival. Both the DDQN (\n            <jats:italic>RepDRL-DDQN<\/jats:italic>\n            ) and Dueling DDQN (\n            <jats:italic>RepDRL-DDDQN<\/jats:italic>\n            ) versions of the DRL model surpass the baseline performance, with the proposed model\u2019s learning agent achieving an optimal learning process across our performance measuring schemes. The robust state representation served as the foundation for enhancing the model\u2019s performance, ultimately providing an optimal treatment policy focused on rapid patient recovery.\n          <\/jats:p>","DOI":"10.1145\/3643856","type":"journal-article","created":{"date-parts":[[2024,2,1]],"date-time":"2024-02-01T11:56:39Z","timestamp":1706788599000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Optimal Treatment Strategies for Critical Patients with Deep Reinforcement Learning"],"prefix":"10.1145","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1716-7475","authenticated-orcid":false,"given":"Simi","family":"Job","sequence":"first","affiliation":[{"name":"School of Mathematics, Physics and Computing, University of Southern Queensland, Toowoomba, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0020-077X","authenticated-orcid":false,"given":"Xiaohui","family":"Tao","sequence":"additional","affiliation":[{"name":"School of Mathematics, Physics and Computing, University of Southern Queensland, Toowoomba, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7553-6916","authenticated-orcid":false,"given":"Lin","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science and Artificial Intelligence, Wuhan University of Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0965-3617","authenticated-orcid":false,"given":"Haoran","family":"Xie","sequence":"additional","affiliation":[{"name":"Department of Computing and Decision Sciences, Lingnan University, Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3624-6120","authenticated-orcid":false,"given":"Taotao","family":"Cai","sequence":"additional","affiliation":[{"name":"School of Mathematics, Physics and Computing, University of Southern Queensland, Toowoomba, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4111-1076","authenticated-orcid":false,"given":"Jianming","family":"Yong","sequence":"additional","affiliation":[{"name":"School of Business, University of Southern Queensland, Springfield, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3370-471X","authenticated-orcid":false,"given":"Qing","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Computing, the Hong Kong Polytechnic University, Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,3,28]]},"reference":[{"key":"e_1_3_1_2_2","article-title":"Optimal policy learning for disease prevention using reinforcement learning","volume":"2020","author":"Khan Zahid Alam","year":"2020","unstructured":"Zahid Alam Khan, Zhengyong Feng, M. Irfan Uddin, Noor Mast, Syed Atif Ali Shah, Muhammad Imtiaz, Mahmoud Ahmad Al-Khasawneh, and Marwan Mahmoud. 2020. Optimal policy learning for disease prevention using reinforcement learning. Scientific Programming 2020 (2020).","journal-title":"Scientific Programming"},{"key":"e_1_3_1_3_2","first-page":"arXiv\u20132202","article-title":"Optimizing warfarin dosing using deep reinforcement learning","author":"Zadeh Sadjad Anzabi","year":"2022","unstructured":"Sadjad Anzabi Zadeh, W. Nick Street, and Barrett W. Thomas. 2022. Optimizing warfarin dosing using deep reinforcement learning. arXiv e-prints (2022), arXiv\u20132202.","journal-title":"arXiv e-prints"},{"key":"e_1_3_1_4_2","article-title":"Distributed distributional deterministic policy gradients","author":"Barth-Maron Gabriel","year":"2018","unstructured":"Gabriel Barth-Maron, Matthew W. Hoffman, David Budden, Will Dabney, Dan Horgan, Dhruva Tb, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap. 2018. Distributed distributional deterministic policy gradients. arXiv preprint arXiv:1804.08617 (2018).","journal-title":"arXiv preprint arXiv:1804.08617"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.3390\/jcm9051548"},{"key":"e_1_3_1_6_2","unstructured":"Dheeru Dua and Casey Graff. 2017. UCI Machine Learning Repository. http:\/\/archive.ics.uci.edu\/ml"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.3389\/fdgth.2021.608893"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1002\/mp.13271"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1016\/S2213-2600(18)30156-5"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1161\/01.CIR.101.23.e215"},{"key":"e_1_3_1_11_2","first-page":"201","article-title":"Vasopressor and Inotropic support in ECMO patients with refractory shock","author":"Harper Michael D.","year":"2022","unstructured":"Michael D. Harper and Marc O. Maybauer. 2022. Vasopressor and Inotropic support in ECMO patients with refractory shock. Extracorporeal Membrane Oxygenation: An Interdisciplinary Problem-Based Learning Approach (2022), 201.","journal-title":"Extracorporeal Membrane Oxygenation: An Interdisciplinary Problem-Based Learning Approach"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TETCI.2023.3246559"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1038\/sdata.2016.35"},{"key":"e_1_3_1_14_2","article-title":"A conservative Q-learning approach for handling distribution shift in sepsis treatment strategies","author":"Kaushik Pramod","year":"2022","unstructured":"Pramod Kaushik, Sneha Kummetha, Perusha Moodley, and Raju S. Bapi. 2022. A conservative Q-learning approach for handling distribution shift in sepsis treatment strategies. arXiv preprint arXiv:2203.13884 (2022).","journal-title":"arXiv preprint arXiv:2203.13884"},{"key":"e_1_3_1_15_2","first-page":"139","volume-title":"Machine Learning for Health","author":"Killian Taylor W.","year":"2020","unstructured":"Taylor W. Killian, Haoran Zhang, Jayakumar Subramanian, Mehdi Fatemi, and Marzyeh Ghassemi. 2020. An empirical study of representation learning for reinforcement learning in healthcare. In Machine Learning for Health. PMLR, 139\u2013160."},{"issue":"1","key":"e_1_3_1_16_2","first-page":"1","article-title":"Computational medication regimen for Parkinson\u2019s disease using reinforcement learning","volume":"11","author":"Kim Yejin","year":"2021","unstructured":"Yejin Kim, Jessika Suescun, Mya C. Schiess, and Xiaoqian Jiang. 2021. Computational medication regimen for Parkinson\u2019s disease using reinforcement learning. Scientific Reports 11, 1 (2021), 1\u20139.","journal-title":"Scientific Reports"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41591-018-0213-5"},{"key":"e_1_3_1_18_2","first-page":"325","volume-title":"Web Intelligence","author":"Lafta Raid","year":"2016","unstructured":"Raid Lafta, Ji Zhang, Xiaohui Tao, Yan Li, Vincent S. Tseng, Yonglong Luo, and Fulong Chen. 2016. An intelligent recommender system based on predictive analysis in telehealthcare environment. In Web Intelligence, Vol. 14. IOS Press, 325\u2013336."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.radonc.2022.02.013"},{"key":"e_1_3_1_20_2","article-title":"G-Net: A deep learning approach to G-computation for counterfactual outcome prediction under dynamic treatment regimes","author":"Li Rui","year":"2020","unstructured":"Rui Li, Zach Shahn, Jun Li, Mingyu Lu, Prithwish Chakraborty, Daby Sow, Mohamed Ghalwash, and Li-wei H. Lehman. 2020. G-Net: A deep learning approach to G-computation for counterfactual outcome prediction under dynamic treatment regimes. arXiv preprint arXiv:2003.10551 (2020).","journal-title":"arXiv preprint arXiv:2003.10551"},{"key":"e_1_3_1_21_2","article-title":"Continuous control with deep reinforcement learning","author":"Lillicrap Timothy P.","year":"2015","unstructured":"Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971 (2015).","journal-title":"arXiv preprint arXiv:1509.02971"},{"key":"e_1_3_1_22_2","unstructured":"Tianlai Lin Xinjue Zhang Jianbing Gong Rundong Tan Xiang Xu Lijun Wang Yingxia Pan Weiming Li and Junhui Gao. 2022. A dosing strategy model of deep deterministic policy gradient algorithm for sepsis. (2022)."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.2991\/ijcis.d.200512.001"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TETC.2019.2896325"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICHI48887.2020.9374313"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2022.104165"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_1_28_2","article-title":"Patient level simulation and reinforcement learning to discover novel strategies for treating ovarian cancer","author":"Murphy Brian","year":"2021","unstructured":"Brian Murphy, Mustafa Nasir-Moin, Grace von Oiste, Viola Chen, Howard A. Riina, Douglas Kondziolka, and Eric K. Oermann. 2021. Patient level simulation and reinforcement learning to discover novel strategies for treating ovarian cancer. arXiv preprint arXiv:2110.11872 (2021).","journal-title":"arXiv preprint arXiv:2110.11872"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pdig.0000012"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2022.117932"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1186\/s12911-022-01789-7"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.12688\/f1000research.20411.1"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-19-0638-1"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0022-3476(97)70065-9"},{"key":"e_1_3_1_35_2","article-title":"The MIMIC III clinical database, version 1.4","author":"Pollard Tom J.","year":"2016","unstructured":"Tom J. Pollard and A. E. W. Johnson III. 2016. The MIMIC III clinical database, version 1.4. The MIMIC-III Clinical Database. PhysioNet (2016).","journal-title":"The MIMIC-III Clinical Database. PhysioNet"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.3390\/jcm11092390"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00134-018-5480-6"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1088\/1361-6560\/ac09a2"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1002\/sim.7144"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1177\/0165551520959798"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v30i1.10295"},{"issue":"10","key":"e_1_3_1_42_2","doi-asserted-by":"crossref","first-page":"e0224563","DOI":"10.1371\/journal.pone.0224563","article-title":"Cumulative fluid balance predicts mortality and increases time on mechanical ventilation in ARDS patients: An observational cohort study","volume":"14","author":"Mourik Niels van","year":"2019","unstructured":"Niels van Mourik, Hennie A. Metske, Jorrit J. Hofstra, Jan M. Binnekade, Bart F. Geerts, Marcus J. Schultz, and Alexander P. J. Vlaar. 2019. Cumulative fluid balance predicts mortality and increases time on mechanical ventilation in ARDS patients: An observational cohort study. PloS One 14, 10 (2019), e0224563.","journal-title":"PloS One"},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","unstructured":"J. L. Vincent Rui Moreno Jukka Takala Sheila Willatts Arnaldo De Mendon\u00e7a Hajo Bruining C. K. Reinhart Peter M. Suter and Lambertius G. Thijs. 1996. The SOFA (Sepsis-related Organ Failure Assessment) score to describe organ dysfunction\/failure: On behalf of the Working Group on Sepsis-Related Problems of the European Society of Intensive Care Medicine (see contributors to the project in the appendix).","DOI":"10.1007\/BF01709751"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3368555.3384469"},{"key":"e_1_3_1_45_2","first-page":"1995","volume-title":"International Conference on Machine Learning","author":"Wang Ziyu","year":"2016","unstructured":"Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas. 2016. Dueling network architectures for deep reinforcement learning. In International Conference on Machine Learning. PMLR, 1995\u20132003."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/EMBC44109.2020.9175311"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijmedinf.2020.104122"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414807"},{"key":"e_1_3_1_49_2","article-title":"Time-series generative adversarial networks","volume":"32","author":"Yoon Jinsung","year":"2019","unstructured":"Jinsung Yoon, Daniel Jarrett, and Mihaela Van der Schaar. 2019. Time-series generative adversarial networks. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"2","key":"e_1_3_1_50_2","first-page":"111","article-title":"Inverse reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units","volume":"19","author":"Yu Chao","year":"2019","unstructured":"Chao Yu, Jiming Liu, and Hongyi Zhao. 2019. Inverse reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units. BMC Medical Informatics and Decision Making 19, 2 (2019), 111\u2013120.","journal-title":"BMC Medical Informatics and Decision Making"},{"issue":"3","key":"e_1_3_1_51_2","first-page":"1","article-title":"Supervised-actor-critic reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units","volume":"20","author":"Yu Chao","year":"2020","unstructured":"Chao Yu, Guoqi Ren, and Yinzhao Dong. 2020. Supervised-actor-critic reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units. BMC Medical Informatics and Decision Making 20, 3 (2020), 1\u20138.","journal-title":"BMC Medical Informatics and Decision Making"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ymssp.2020.107322"},{"key":"e_1_3_1_53_2","article-title":"Near-optimal reinforcement learning in dynamic treatment regimes","volume":"32","author":"Zhang Junzhe","year":"2019","unstructured":"Junzhe Zhang and Elias Bareinboim. 2019. Near-optimal reinforcement learning in dynamic treatment regimes. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.3390\/s20185058"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643856","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3643856","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T23:57:34Z","timestamp":1750291054000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643856"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,28]]},"references-count":53,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,4,30]]}},"alternative-id":["10.1145\/3643856"],"URL":"https:\/\/doi.org\/10.1145\/3643856","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,28]]},"assertion":[{"value":"2023-05-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-15","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-03-28","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}