{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T19:51:08Z","timestamp":1784922668857,"version":"3.55.0"},"reference-count":74,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,2,26]],"date-time":"2025-02-26T00:00:00Z","timestamp":1740528000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,2,26]],"date-time":"2025-02-26T00:00:00Z","timestamp":1740528000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Hum-Cent Intell Syst"],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>In critical medicine, data-driven methods that assist in physician decisions often require accurate responses and controllable safety risks. Most recent reinforcement learning models developed for clinical research typically use fixed-length and very short time series data. Unfortunately, such methods generalize poorly on variable-length data that can be overlong. In such as case, a single final reward signal appears very sparse. Meanwhile, safety is often overlooked by many models, leading them to make excessively extreme recommendations. In this paper, we study how to recommend effective and safe treatments for critically ill septic patients. We develop an offline reinforcement learning model based on CQL (Conservative Q-Learning), which underestimates the expected rewards of rarely seen treatments in data, thus enjoying a high safety standard. We further enhance the model with intermediate rewards by particularly using the Apache II scoring system. This can effectively deal with variable-length episodes with sparse rewards. By performing extensive experiments on the MIMIC-III database, we demonstrated the enhanced performance and robustness in safety. Our code of data extraction, preprocessing, and modeling can be found at <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/OOPSDINOSAUR\/RL_safety_model\" ext-link-type=\"uri\">https:\/\/github.com\/OOPSDINOSAUR\/RL_safety_model<\/jats:ext-link>.<\/jats:p>","DOI":"10.1007\/s44230-025-00093-7","type":"journal-article","created":{"date-parts":[[2025,2,26]],"date-time":"2025-02-26T10:05:49Z","timestamp":1740564349000},"page":"63-76","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Offline Safe Reinforcement Learning for Sepsis Treatment: Tackling Variable-Length Episodes with Sparse Rewards"],"prefix":"10.1007","volume":"5","author":[{"given":"Rui","family":"Tu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4053-5443","authenticated-orcid":false,"given":"Zhipeng","family":"Luo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chuanliang","family":"Pan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhong","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jie","family":"Su","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yifan","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,2,26]]},"reference":[{"issue":"14","key":"93_CR1","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1007\/PL00003795","volume":"27","author":"I Matot","year":"2001","unstructured":"Matot I, Sprung CL. Definition of sepsis. Intensive Care Med. 2001;27(14):3\u20139.","journal-title":"Intensive Care Med."},{"issue":"12","key":"93_CR2","doi-asserted-by":"publisher","first-page":"1889","DOI":"10.1097\/CCM.0000000000003342","volume":"46","author":"CJ Paoli","year":"2018","unstructured":"Paoli CJ, Reynolds MA, Sinha M, Gitlin M, Crouser E. Epidemiology and costs of sepsis in the united states-an analysis based on timing of diagnosis and severity level. Crit Care Med. 2018;46(12):1889\u201397.","journal-title":"Crit Care Med."},{"issue":"3","key":"93_CR3","doi-asserted-by":"publisher","first-page":"286","DOI":"10.3390\/mi11030286","volume":"11","author":"A Teggert","year":"2020","unstructured":"Teggert A, Datta H, Ali Z. Biomarkers for point-of-care diagnosis of sepsis. Micromachines. 2020;11(3):286.","journal-title":"Micromachines."},{"issue":"12","key":"93_CR4","doi-asserted-by":"publisher","first-page":"1012","DOI":"10.1016\/j.amjmed.2007.01.035","volume":"120","author":"JM O\u2019Brien Jr","year":"2007","unstructured":"O\u2019Brien JM Jr, Ali NA, Aberegg SK, Abraham E. Sepsis. Am J Med. 2007;120(12):1012\u201322.","journal-title":"Am J Med."},{"key":"93_CR5","doi-asserted-by":"publisher","first-page":"102003","DOI":"10.1016\/j.artmed.2020.102003","volume":"112","author":"L Roggeveen","year":"2021","unstructured":"Roggeveen L, El Hassouni A, Ahrendt J, Guo T, Fleuren L, Thoral P, Girbes AR, Hoogendoorn M, Elbers PW. Transatlantic transferability of a new reinforcement learning model for optimizing haemodynamic treatment for critically ill patients with sepsis. Artif Intell Med. 2021;112:102003.","journal-title":"Artif Intell Med."},{"key":"93_CR6","doi-asserted-by":"crossref","unstructured":"Ju S, Kim YJ, Ausin MS, Mayorga ME, Chi M. To reduce healthcare workload: Identify critical sepsis progression moments through deep reinforcement learning. In: 2021 IEEE International Conference on Big Data (Big Data). IEEE; 2021. pp. 1640\u20131646.","DOI":"10.1109\/BigData52589.2021.9671407"},{"key":"93_CR7","doi-asserted-by":"crossref","unstructured":"Yang M, Li R, Hao T, Ma C, Li J, Liu C, Raising high-risk awareness in hemodynamic treatment with reinforcement learning for septic shock patients. In: 2022 Computing in Cardiology (CinC), vol. 498. IEEE; 2022. p. 1\u20134.","DOI":"10.22489\/CinC.2022.164"},{"key":"93_CR8","doi-asserted-by":"publisher","first-page":"766447","DOI":"10.3389\/fmed.2022.766447","volume":"9","author":"L Su","year":"2022","unstructured":"Su L, Li Y, Liu S, Zhang S, Zhou X, Weng L, Su M, Du B, Zhu W, Long Y. Establishment and implementation of potential fluid therapy balance strategies for ICU sepsis patients based on reinforcement learning. Front Med. 2022;9:766447.","journal-title":"Front Med."},{"issue":"11","key":"93_CR9","doi-asserted-by":"publisher","first-page":"1716","DOI":"10.1038\/s41591-018-0213-5","volume":"24","author":"M Komorowski","year":"2018","unstructured":"Komorowski M, Celi LA, Badawi O, Gordon AC, Faisal AA. The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care. Nat Med. 2018;24(11):1716\u201320.","journal-title":"Nat Med."},{"issue":"9","key":"93_CR10","doi-asserted-by":"publisher","first-page":"11034","DOI":"10.1007\/s10489-022-04099-7","volume":"53","author":"D Liang","year":"2023","unstructured":"Liang D, Deng H, Liu Y. The treatment of sepsis: an episodic memory-assisted deep reinforcement learning approach. Appl Intell. 2023;53(9):11034\u201344.","journal-title":"Appl Intell."},{"key":"93_CR11","doi-asserted-by":"crossref","unstructured":"Do TC, Yang HJ, Yoo SB, Oh I-J. Combining reinforcement learning with supervised learning for sepsis treatment. In: The 9th international conference on smart media and applications; 2020. pp. 219\u2013223.","DOI":"10.1145\/3426020.3426077"},{"issue":"2","key":"93_CR12","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1586\/eri.12.159","volume":"11","author":"N Adam","year":"2013","unstructured":"Adam N, Kandelman S, Mantz J, Chr\u00e9tien F, Sharshar T. Sepsis-induced brain dysfunction. Expert Rev Anti-Infective Ther. 2013;11(2):211\u201321.","journal-title":"Expert Rev Anti-Infective Ther."},{"issue":"12","key":"93_CR13","doi-asserted-by":"publisher","first-page":"2151","DOI":"10.1097\/CCM.0000000000005152","volume":"49","author":"SL Kimball","year":"2021","unstructured":"Kimball SL, Levy MM. Sepsis and the opioid crisis: integrating treatment for two public health emergencies. Crit Care Med. 2021;49(12):2151\u20133.","journal-title":"Crit Care Med."},{"issue":"1","key":"93_CR14","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/sdata.2016.35","volume":"3","author":"AE Johnson","year":"2016","unstructured":"Johnson AE, Pollard TJ, Shen L, Lehman L-H, Feng M, Ghassemi M, Moody B, Szolovits P, Anthony Celi L, Mark RG. MIMIC-III, a freely accessible critical care database. Sci Data. 2016;3(1):1\u20139.","journal-title":"Sci Data."},{"key":"93_CR15","first-page":"1179","volume":"33","author":"A Kumar","year":"2020","unstructured":"Kumar A, Zhou A, Tucker G, Levine S. Conservative q-learning for offline reinforcement learning. Adv Neural Inf Process Syst. 2020;33:1179\u201391.","journal-title":"Adv Neural Inf Process Syst."},{"key":"93_CR16","doi-asserted-by":"crossref","unstructured":"Van\u00a0Hasselt H, Guez A, Silver D. Deep reinforcement learning with double q-learning. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 30; 2016.","DOI":"10.1609\/aaai.v30i1.10295"},{"issue":"10","key":"93_CR17","doi-asserted-by":"publisher","first-page":"818","DOI":"10.1097\/00003246-198510000-00009","volume":"13","author":"WA Knaus","year":"1985","unstructured":"Knaus WA, Draper EA, Wagner DP, Zimmerman JE. Apache II: a severity of disease classification system. Crit Care med. 1985;13(10):818\u201329.","journal-title":"Crit Care med."},{"issue":"1","key":"93_CR18","doi-asserted-by":"publisher","first-page":"6085","DOI":"10.1038\/s41598-018-24271-9","volume":"8","author":"Z Che","year":"2018","unstructured":"Che Z, Purushotham S, Cho K, Sontag D, Liu Y. Recurrent neural networks for multivariate time series with missing values. Sci Rep. 2018;8(1):6085.","journal-title":"Sci Rep."},{"issue":"1","key":"93_CR19","doi-asserted-by":"publisher","first-page":"130","DOI":"10.1016\/j.cor.2006.02.017","volume":"35","author":"P Duchesne","year":"2008","unstructured":"Duchesne P, Pacurar M. Evaluating financial time series models for irregularly spaced data: a spectral density approach. Comput Oper Res. 2008;35(1):130\u201355.","journal-title":"Comput Oper Res."},{"key":"93_CR20","doi-asserted-by":"crossref","unstructured":"Yang J, Rahardja S, Rahardja S, Click fraud detection: Hk-index for feature extraction from variable-length time series of user behavior. In: 2022 IEEE 32nd International Workshop on Machine Learning for Signal Processing (MLSP). IEEE; 2022. pp. 1\u20136.","DOI":"10.1109\/MLSP55214.2022.9943422"},{"key":"93_CR21","doi-asserted-by":"crossref","unstructured":"Xu D, Cheng W, Zong B, Song D, Ni J, Yu W, Liu Y, Chen H, Zhang X. Tensorized lstm with adaptive shared memory for learning trends in multivariate time series. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34; 2020. pp. 1395\u20131402.","DOI":"10.1609\/aaai.v34i02.5496"},{"key":"93_CR22","doi-asserted-by":"crossref","unstructured":"Rajeh T M, Luo Z, Javed M H, et al. A clustering-based multi-agent reinforcement learning framework for finer-grained taxi dispatching[J]. IEEE Trans Intell Transport Syst. 2024.","DOI":"10.1109\/TITS.2024.3370820"},{"issue":"2","key":"93_CR23","doi-asserted-by":"publisher","first-page":"1180","DOI":"10.1214\/10-AOS864","volume":"39","author":"M Qian","year":"2011","unstructured":"Qian M, Murphy SA. Performance guarantees for individualized treatment rules. Ann Stat. 2011;39(2):1180.","journal-title":"Ann Stat."},{"key":"93_CR24","doi-asserted-by":"crossref","unstructured":"Wang Z, Zhao H, Ren P, Zhou Y, Sheng M. Learning optimal treatment strategies for sepsis using offline reinforcement learning in continuous space. In: International Conference on Health Information Science. Berlin: Springer; 2022. pp. 113\u2013124.","DOI":"10.1007\/978-3-031-20627-6_11"},{"issue":"6","key":"93_CR25","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3531228","volume":"13","author":"X Wu","year":"2022","unstructured":"Wu X, Huang C, Robles-Granda P, Chawla NV. Representation learning on variable length and incomplete wearable-sensory time series. ACM Trans Intell Syst Technol. 2022;13(6):1\u201321.","journal-title":"ACM Trans Intell Syst Technol."},{"issue":"1","key":"93_CR26","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3531326","volume":"14","author":"MA Morid","year":"2023","unstructured":"Morid MA, Sheng ORL, Dunbar J. Time series prediction using deep learning methods in healthcare. ACM Trans Manag Inf Syst. 2023;14(1):1\u201329.","journal-title":"ACM Trans Manag Inf Syst."},{"key":"93_CR27","doi-asserted-by":"crossref","unstructured":"Lea C, Flynn MD, Vidal R, Reiter A, Hager GD. Temporal convolutional networks for action segmentation and detection. In: Proceedings of the IEEE conference on computer vision and pattern recognition; 2017. pp. 156\u2013165.","DOI":"10.1109\/CVPR.2017.113"},{"key":"93_CR28","unstructured":"Ashish V. Attention is all you need. In: Advances in neural information processing systems. 2017;30."},{"issue":"2194","key":"93_CR29","doi-asserted-by":"publisher","first-page":"20200209","DOI":"10.1098\/rsta.2020.0209","volume":"379","author":"B Lim","year":"2021","unstructured":"Lim B, Zohren S. Time-series forecasting with deep learning: a survey. Philos Trans R Soc A. 2021;379(2194):20200209.","journal-title":"Philos Trans R Soc A."},{"issue":"2","key":"93_CR30","doi-asserted-by":"publisher","first-page":"174","DOI":"10.1007\/s41666-020-00069-1","volume":"4","author":"S Daberdaku","year":"2020","unstructured":"Daberdaku S, Tavazzi E, Di Camillo B. A combined interpolation and weighted k-nearest neighbours approach for the imputation of longitudinal icu laboratory data. J Healthc Inf Res. 2020;4(2):174\u201388.","journal-title":"J Healthc Inf Res."},{"issue":"3","key":"93_CR31","doi-asserted-by":"publisher","first-page":"295","DOI":"10.1007\/s41666-020-00073-5","volume":"4","author":"A Jazayeri","year":"2020","unstructured":"Jazayeri A, Liang OS, Yang CC. Imputation of missing data in electronic health records based on patients\u2019 similarities. J Healthc Inf Res. 2020;4(3):295\u2013307.","journal-title":"J Healthc Inf Res."},{"issue":"6","key":"93_CR32","doi-asserted-by":"publisher","first-page":"1322","DOI":"10.1093\/jamia\/ocae088","volume":"31","author":"FS Bashiri","year":"2024","unstructured":"Bashiri FS, Carey KA, Martin J, Koyner JL, Edelson DP, Gilbert ER, Mayampurath A, Afshar M, Churpek MM. Development and external validation of deep learning clinical prediction models using variable-length time series data. J Am Med Inf Assoc. 2024;31(6):1322\u201330.","journal-title":"J Am Med Inf Assoc."},{"key":"93_CR33","unstructured":"Haarnoja T, Zhou A, Abbeel P, Levine S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In: International Conference on Machine Learning. PMLR; 2018. pp. 1861\u20131870."},{"key":"93_CR34","unstructured":"Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347; 2017."},{"key":"93_CR35","doi-asserted-by":"crossref","unstructured":"Tamboli D, Chen J, Jotheeswaran K P, et al. Reinforced sequential decision-making for sepsis treatment: the posnegdm framework with mortality classifier and transformer[J]. IEEE J Biomed Health Inf. 2024.","DOI":"10.1109\/JBHI.2024.3377214"},{"key":"93_CR36","unstructured":"Peng X, Ding Y, Wihl D, Gottesman O, Komorowski M, Li-wei HL, Ross A, Faisal A, Doshi-Velez F. Improving sepsis treatment strategies by combining deep and kernel-based reinforcement learning. In: AMIA Annual Symposium Proceedings, vol. 2018. American Medical Informatics Association; 2018. p. 887."},{"key":"93_CR37","unstructured":"Sinha S, Mandlekar A, Garg A. S4rl: surprisingly simple self-supervision for offline reinforcement learning in robotics. In: Conference on Robot Learning. PMLR; 2022. pp. 907\u2013917."},{"issue":"6","key":"93_CR38","doi-asserted-by":"publisher","first-page":"4909","DOI":"10.1109\/TITS.2021.3054625","volume":"23","author":"BR Kiran","year":"2021","unstructured":"Kiran BR, Sobh I, Talpaert V, Mannion P, Al Sallab AA, Yogamani S, P\u00e9rez P. Deep reinforcement learning for autonomous driving: a survey. IEEE Trans Intell Transp Syst. 2021;23(6):4909\u201326.","journal-title":"IEEE Trans Intell Transp Syst."},{"issue":"3","key":"93_CR39","doi-asserted-by":"publisher","first-page":"653","DOI":"10.1109\/TNNLS.2016.2522401","volume":"28","author":"Y Deng","year":"2016","unstructured":"Deng Y, Bao F, Kong Y, Ren Z, Dai Q. Deep direct reinforcement learning for financial signal representation and trading. IEEE Trans Neural Netw Learn Syst. 2016;28(3):653\u201364.","journal-title":"IEEE Trans Neural Netw Learn Syst."},{"key":"93_CR40","doi-asserted-by":"publisher","first-page":"104376","DOI":"10.1016\/j.jbi.2023.104376","volume":"142","author":"H Emerson","year":"2023","unstructured":"Emerson H, Guy M, McConville R. Offline reinforcement learning for safer blood glucose control in people with type 1 diabetes. J Biomed Inf. 2023;142:104376.","journal-title":"J Biomed Inf."},{"key":"93_CR41","unstructured":"Amodei D, Olah C, Steinhardt J, Christiano P, Schulman J, Man\u00e9 D. Concrete problems in ai safety. arXiv preprint arXiv:1606.06565; 2016."},{"key":"93_CR42","doi-asserted-by":"publisher","DOI":"10.1201\/9781315140223","volume-title":"Constrained Markov decision processes","author":"E Altman","year":"2021","unstructured":"Altman E. Constrained Markov decision processes. Boca Raton: Routledge; 2021."},{"issue":"1","key":"93_CR43","doi-asserted-by":"publisher","first-page":"236","DOI":"10.1016\/0022-247X(85)90288-4","volume":"112","author":"FJ Beutler","year":"1985","unstructured":"Beutler FJ, Ross KW. Optimal policies for controlled Markov chains with a constraint. J Math Anal Appl. 1985;112(1):236\u201352.","journal-title":"J Math Anal Appl."},{"issue":"2","key":"93_CR44","doi-asserted-by":"publisher","first-page":"341","DOI":"10.2307\/1427303","volume":"18","author":"FJ Beutler","year":"1986","unstructured":"Beutler FJ, Ross KW. Time-average optimal constrained semi-Markov decision processes. Adv Appl Probab. 1986;18(2):341\u201359.","journal-title":"Adv Appl Probab."},{"key":"93_CR45","volume-title":"Linear programming and finite Markovian control problems","author":"LC Kallenberg","year":"1983","unstructured":"Kallenberg LC. Linear programming and finite Markovian control problems. Amsterdam: Mathematical Centre; 1983."},{"key":"93_CR46","doi-asserted-by":"crossref","unstructured":"Clouse JA, Utgoff PE. A teaching method for reinforcement learning. In: Machine Learning Proceedings 1992, pp. 92\u2013101. Elsevier, Amsterdam; 1992","DOI":"10.1016\/B978-1-55860-247-2.50017-6"},{"key":"93_CR47","unstructured":"Moldovan TM, Abbeel P. Safe exploration in markov decision processes. arXiv preprint arXiv:1205.4810; 2012."},{"key":"93_CR48","unstructured":"Khattar V, Ding Y, Sel B, Lavaei J, Jin M. A cmdp-within-online framework for meta-safe reinforcement learning. arXiv preprint arXiv:2405.16601; 2024."},{"key":"93_CR49","doi-asserted-by":"crossref","unstructured":"Kalagarla KC, Jain R, Nuzzo P. A sample-efficient algorithm for episodic finite-horizon MDP with constraints. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35; 2021. pp. 8030\u20138037.","DOI":"10.1609\/aaai.v35i9.16979"},{"key":"93_CR50","unstructured":"Xiong X, Wang J, Zhang F, Li K. Combining deep reinforcement learning and safety based control for autonomous driving. arXiv preprint arXiv:1612.00147; 2016."},{"issue":"9","key":"93_CR51","doi-asserted-by":"publisher","first-page":"9966","DOI":"10.1109\/TITS.2023.3271642","volume":"24","author":"Z Gu","year":"2023","unstructured":"Gu Z, Gao L, Ma H, Li SE, Zheng S, Jing W, Chen J. Safe-state enhancement method for autonomous driving via direct hierarchical reinforcement learning. IEEE Trans Intell Transp Syst. 2023;24(9):9966\u201383.","journal-title":"IEEE Trans Intell Transp Syst."},{"key":"93_CR52","doi-asserted-by":"publisher","first-page":"120297","DOI":"10.1016\/j.eswa.2023.120297","volume":"227","author":"SJ Yoo","year":"2023","unstructured":"Yoo SJ, Gu YH, et al. Safety AARL: weight adjustment for reinforcement-learning-based safety dynamic asset allocation strategies. Expert Syst Appl. 2023;227:120297.","journal-title":"Expert Syst Appl."},{"key":"93_CR53","doi-asserted-by":"crossref","unstructured":"Yuan Y, Shi J, Yang J, Li C, Cai Y, Tang B. Conservative q-learning for mechanical ventilation treatment using diagnose transformer-encoder. In: 2023 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). IEEE; 2023. pp. 2346\u20132351.","DOI":"10.1109\/BIBM58861.2023.10385663"},{"key":"93_CR54","unstructured":"Kondrup F, Jiralerspong T, Lau E, et al. Deep conservative reinforcement learning for personalization of mechanical ventilation treatment[J]."},{"key":"93_CR55","unstructured":"Kaushik P, Kummetha S, Moodley P, Bapi RS. A conservative q-learning approach for handling distribution shift in sepsis treatment strategies. arXiv preprint arXiv:2203.13884; 2022."},{"issue":"1","key":"93_CR56","doi-asserted-by":"publisher","first-page":"459","DOI":"10.1109\/JBHI.2023.3321099","volume":"28","author":"X Cai","year":"2023","unstructured":"Cai X, Chen J, Zhu Y, et al. Towards real-world applications of personalized anesthesia using policy Constraint Q Learning for Propofol Infusion Control[J]. IEEE J Biomed Health Inf. 2023;28(1):459\u201369.","journal-title":"IEEE Journal of Biomedical and Health Informatics"},{"key":"93_CR57","doi-asserted-by":"crossref","unstructured":"Mataric MJ. Reward functions for accelerated learning. In: Machine Learning Proceedings 1994. Elsevier, Amsterdam; 1994. pp. 181\u2013189.","DOI":"10.1016\/B978-1-55860-335-6.50030-1"},{"key":"93_CR58","doi-asserted-by":"crossref","unstructured":"Lin R, Stanley MD, Ghassemi MM, Nemati SA, deep deterministic policy gradient approach to medication dosing and surveillance in the ICU. In: 2018 40th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). IEEE; 2018. pp 4927\u201331.","DOI":"10.1109\/EMBC.2018.8513203"},{"key":"93_CR59","doi-asserted-by":"publisher","first-page":"608893","DOI":"10.3389\/fdgth.2021.608893","volume":"3","author":"N Eghbali","year":"2021","unstructured":"Eghbali N, Alhanai T, Ghassemi MM. Patient-specific sedation management via deep reinforcement learning. Front Dig Health. 2021;3:608893.","journal-title":"Front Dig Health."},{"issue":"8","key":"93_CR60","doi-asserted-by":"publisher","first-page":"801","DOI":"10.1001\/jama.2016.0287","volume":"315","author":"M Singer","year":"2016","unstructured":"Singer M, Deutschman CS, Seymour CW, Shankar-Hari M, Annane D, Bauer M, Bellomo R, Bernard GR, Chiche J-D, Coopersmith CM, et al. The third international consensus definitions for sepsis and septic shock (sepsis-3). JAMA. 2016;315(8):801\u201310.","journal-title":"JAMA."},{"issue":"1","key":"93_CR61","doi-asserted-by":"publisher","first-page":"8","DOI":"10.1097\/00005650-199801000-00004","volume":"36","author":"A Elixhauser","year":"1998","unstructured":"Elixhauser A, Steiner C, Harris DR, Coffey RM. Comorbidity measures for use with administrative data. Med Care. 1998;36(1):8\u201327.","journal-title":"Med Care."},{"key":"93_CR62","unstructured":"Hug C. Detecting hazardous intensive care patient episodes using real-time mortality models. PhD thesis; 2009."},{"issue":"5","key":"93_CR63","doi-asserted-by":"publisher","first-page":"710","DOI":"10.1109\/TIP.2004.826093","volume":"13","author":"T Blu","year":"2004","unstructured":"Blu T, Th\u00e9venaz P, Unser M. Linear interpolation revitalized. IEEE Trans Image Process. 2004;13(5):710\u20139.","journal-title":"IEEE Trans Image Process."},{"key":"93_CR64","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1016\/j.csda.2015.04.009","volume":"90","author":"G Tutz","year":"2015","unstructured":"Tutz G, Ramzan S. Improved methods for the imputation of missing data by nearest neighbor methods. Comput Stat Data Anal. 2015;90:84\u201399.","journal-title":"Comput Stat Data Anal."},{"key":"93_CR65","unstructured":"Hendrycks D, Gimpel K. A baseline for detecting misclassified and out-of-distribution examples in neural networks. arXiv preprint arXiv:1610.02136; 2016."},{"key":"93_CR66","doi-asserted-by":"crossref","unstructured":"Nitsch J, Itkina M, Senanayake R, Nieto J, Schmidt M, Siegwart R, Kochenderfer MJ, Cadena C, Out-of-distribution detection for automotive perception. In: 2021 IEEE International Intelligent Transportation Systems Conference (ITSC). IEEE; 2021. p. 2938\u201343.","DOI":"10.1109\/ITSC48978.2021.9564545"},{"key":"93_CR67","doi-asserted-by":"publisher","first-page":"279","DOI":"10.1007\/BF00992698","volume":"8","author":"CJ Watkins","year":"1992","unstructured":"Watkins CJ, Dayan P. Q-learning. Mach Learn. 1992;8:279\u201392.","journal-title":"Mach Learn."},{"issue":"7540","key":"93_CR68","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG, Graves A, Riedmiller M, Fidjeland AK, Ostrovski G, et al. Human-level control through deep reinforcement learning. Nature. 2015;518(7540):529\u201333.","journal-title":"Nature."},{"key":"93_CR69","unstructured":"Uehara M, Shi C, Kallus N. A review of off-policy evaluation in reinforcement learning. arXiv preprint arXiv:2212.06355; 2022."},{"key":"93_CR70","doi-asserted-by":"crossref","unstructured":"Jia Y, Burden J, Lawton T, Habli I. Safe reinforcement learning for sepsis treatment. In: 2020 IEEE International Conference on Healthcare Informatics (ICHI). IEEE; 2020. pp. 1\u20137.","DOI":"10.1109\/ICHI48887.2020.9374367"},{"key":"93_CR71","unstructured":"Tang S, Wiens J. Model selection for offline reinforcement learning: practical considerations for healthcare settings. In: Machine Learning for Healthcare Conference. PMLR; 2021. pp. 2\u201335"},{"key":"93_CR72","unstructured":"Le H, Voloshin C, Yue Y. Batch policy learning under constraints. In: International conference on machine learning. PMLR; 2019. pp. 3703\u20133712."},{"issue":"19","key":"93_CR73","doi-asserted-by":"publisher","first-page":"1368","DOI":"10.1056\/NEJMoa010307","volume":"345","author":"E Rivers","year":"2001","unstructured":"Rivers E, Nguyen B, Havstad S, Ressler J, Muzzin A, Knoblich B, Peterson E, Tomlanovich M. Early goal-directed therapy in the treatment of severe sepsis and septic shock. N Engl J Med. 2001;345(19):1368\u201377.","journal-title":"N Engl J Med."},{"issue":"17","key":"93_CR74","doi-asserted-by":"publisher","first-page":"1583","DOI":"10.1056\/NEJMoa1312173","volume":"370","author":"P Asfar","year":"2014","unstructured":"Asfar P, Meziani F, Hamel J-F, Grelon F, Megarbane B, Anguel N, Mira J-P, Dequin P-F, Gergaud S, Weiss N, et al. High versus low blood-pressure target in patients with septic shock. N Engl J Med. 2014;370(17):1583\u201393.","journal-title":"N Engl J Med."}],"container-title":["Human-Centric Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44230-025-00093-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44230-025-00093-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44230-025-00093-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,4,23]],"date-time":"2025-04-23T14:03:28Z","timestamp":1745417008000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44230-025-00093-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,26]]},"references-count":74,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,3]]}},"alternative-id":["93"],"URL":"https:\/\/doi.org\/10.1007\/s44230-025-00093-7","relation":{},"ISSN":["2667-1336"],"issn-type":[{"value":"2667-1336","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,26]]},"assertion":[{"value":"20 December 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 February 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 February 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics Approval and Consent to Participate"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for Publication"}}]}}