{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T14:59:30Z","timestamp":1783954770446,"version":"3.55.0"},"reference-count":75,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2023,10,11]],"date-time":"2023-10-11T00:00:00Z","timestamp":1696982400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,10,11]],"date-time":"2023-10-11T00:00:00Z","timestamp":1696982400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Interreg France-Suisse"},{"name":"University of Applied Sciences and Arts Western Switzerland"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2024,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Offline reinforcement learning (RL) has emerged as a promising paradigm for real-world applications since it aims to train policies directly from datasets of past interactions with the environment. The past few years, algorithms have been introduced to learn from high-dimensional observational states in offline settings. The general idea of these methods is to encode the environment into a latent space and train policies on top of this smaller representation. In this paper, we extend this general method to stochastic environments (i.e., where the reward function is stochastic) and consider a risk measure instead of the classical expected return. First, we show that, under some assumptions, it is equivalent to minimizing a risk measure in the latent space and in the natural space. Based on this result, we present Latent Offline Distributional Actor-Critic (LODAC), an algorithm which is able to train policies in high-dimensional stochastic and offline settings to minimize a given risk measure. Empirically, we show that using LODAC to minimize Conditional Value-at-Risk (CVaR) outperforms previous methods in terms of CVaR and return on stochastic environments.<\/jats:p>","DOI":"10.1007\/s00521-023-09029-3","type":"journal-article","created":{"date-parts":[[2023,10,11]],"date-time":"2023-10-11T10:04:40Z","timestamp":1697018680000},"page":"585-598","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Offline reinforcement learning in high-dimensional stochastic environments"],"prefix":"10.1007","volume":"36","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-3425-8189","authenticated-orcid":false,"given":"F\u00e9licien","family":"H\u00eache","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Oussama","family":"Barakat","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4585-5494","authenticated-orcid":false,"given":"Thibaut","family":"Desmettre","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1443-9423","authenticated-orcid":false,"given":"Tania","family":"Marx","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7326-2357","authenticated-orcid":false,"given":"Stephan","family":"Robert-Nicoud","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,10,11]]},"reference":[{"issue":"2","key":"9029_CR1","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3054912","volume":"50","author":"A Hussein","year":"2017","unstructured":"Hussein A, Gaber MM, Elyan E, Jayne C (2017) Imitation learning: a survey of learning methods. ACM Comput Surv (CSUR) 50(2):1\u201335. https:\/\/doi.org\/10.1145\/3054912","journal-title":"ACM Comput Surv (CSUR)"},{"issue":"7587","key":"9029_CR2","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver D, Huang A, Maddison CJ, Guez A, Sifre L, Van Den Driessche G, Schrittwieser J, Antonoglou I, Panneershelvam V, Lanctot M et al (2016) Mastering the game of go with deep neural networks and tree search. Nature 529(7587):484\u2013489. https:\/\/doi.org\/10.1038\/nature16961","journal-title":"Nature"},{"issue":"6419","key":"9029_CR3","doi-asserted-by":"publisher","first-page":"1140","DOI":"10.1126\/science.aar6404","volume":"362","author":"D Silver","year":"2018","unstructured":"Silver D, Hubert T, Schrittwieser J, Antonoglou I, Lai M, Guez A, Lanctot M, Sifre L, Kumaran D, Graepel T et al (2018) A general reinforcement learning algorithm that masters chess, shogi, and go through self-play. Science 362(6419):1140\u20131144. https:\/\/doi.org\/10.1126\/science.aar6404","journal-title":"Science"},{"key":"9029_CR4","unstructured":"Haarnoja T, Zhou A, Abbeel P, Levine S (2018) Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor. In: International conference on machine learning, pp 1861\u20131870. PMLR"},{"key":"9029_CR5","doi-asserted-by":"publisher","unstructured":"Johannink T, Bahl S, Nair A, Luo J, Kumar A, Loskyll M, Ojea JA, Solowjow E, Levine S (2019) Residual reinforcement learning for robot control. In: 2019 international conference on robotics and automation (ICRA), pp 6023\u20136029. https:\/\/doi.org\/10.1109\/ICRA.2019.8794127. IEEE","DOI":"10.1109\/ICRA.2019.8794127"},{"key":"9029_CR6","doi-asserted-by":"publisher","unstructured":"Wang L, Zhang W, He X, Zha H (2018) Supervised reinforcement learning with recurrent neural network for dynamic treatment recommendation. In: Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining, pp 2447\u20132456. https:\/\/doi.org\/10.1145\/3219819.3219961","DOI":"10.1145\/3219819.3219961"},{"issue":"1","key":"9029_CR7","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3477600","volume":"55","author":"C Yu","year":"2021","unstructured":"Yu C, Liu J, Nemati S, Yin G (2021) Reinforcement learning in healthcare: a survey. ACM Comput Surv (CSUR) 55(1):1\u201336. https:\/\/doi.org\/10.1145\/3477600","journal-title":"ACM Comput Surv (CSUR)"},{"key":"9029_CR8","doi-asserted-by":"publisher","unstructured":"Tassa Y, Doron Y, Muldal A, Erez T, Li Y, Casas DL, Budden D, Abdolmaleki A, Merel J, Lefrancq A et al (2018) Deepmind control suite. arXiv:1801.00690. https:\/\/doi.org\/10.48550\/arXiv.1801.00690","DOI":"10.48550\/arXiv.1801.00690"},{"key":"9029_CR9","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2023.3250269","author":"RF Prudencio","year":"2023","unstructured":"Prudencio RF, Maximo MR, Colombini EL (2023) A survey on offline reinforcement learning: taxonomy, review, and open problems. IEEE Trans Neural Netw Learn Syst. https:\/\/doi.org\/10.1109\/TNNLS.2023.3250269","journal-title":"IEEE Trans Neural Netw Learn Syst"},{"key":"9029_CR10","first-page":"4","volume":"3","author":"M Liu","year":"2021","unstructured":"Liu M, Zhao H, Yang Z, Shen J, Zhang W, Zhao L, Liu T-Y (2021) Curriculum offline imitating learning. Adv Neural Inf Process Syst 3:4","journal-title":"Adv Neural Inf Process Syst"},{"key":"9029_CR11","unstructured":"Kostrikov I, Nair A, Levine S (2021) Offline reinforcement learning with implicit q-learning. In: Deep RL workshop NeurIPS 2021"},{"key":"9029_CR12","first-page":"4085","volume":"35","author":"H Xu","year":"2022","unstructured":"Xu H, Jiang L, Jianxiong L, Zhan X (2022) A policy-guided imitation approach for offline reinforcement learning. Adv Neural Inf Process Syst 35:4085\u20134098","journal-title":"Adv Neural Inf Process Syst"},{"key":"9029_CR13","unstructured":"Xu H, Jiang L, Li J, Yang Z, Wang Z, Chan VWK, Zhan X (2022) Offline rl with no ood actions: In-sample learning via implicit value regularization. In: The eleventh international conference on learning representations"},{"key":"9029_CR14","unstructured":"Snell CV, Kostrikov I, Su Y, Yang S, Levine S (2022) Offline rl for natural language generation with implicit language q learning. In: The eleventh international conference on learning representations"},{"key":"9029_CR15","unstructured":"Zheng Q, Henaff M, Amos B, Grover A (2023) Semi-supervised offline reinforcement learning with action-free trajectories. In: International conference on machine learning, pp 42339\u201342362. PMLR"},{"key":"9029_CR16","unstructured":"Kumar A, Fu J, Soh M, Tucker G, Levine S (2019) Stabilizing off-policy q-learning via bootstrapping error reduction. In: Advances in neural information processing systems, vol 32"},{"key":"9029_CR17","unstructured":"Liu Y, Swaminathan A, Agarwal A, Brunskill E (2020) Off-policy policy gradient with stationary distribution correction. In: Uncertainty in artificial intelligence, pp. 1180\u20131190. PMLR"},{"key":"9029_CR18","unstructured":"Rashidinejad P, Zhu H, Yang K, Russell S, Jiao J (2022) Optimal conservative offline rl with general function approximation via augmented Lagrangian. In: The eleventh international conference on learning representations"},{"key":"9029_CR19","unstructured":"Rafailov R, Yu T, Rajeswaran A, Finn C (2021) Offline reinforcement learning from images with latent space models. In: Learning for dynamics and control, pp 1154\u20131168. PMLR"},{"key":"9029_CR20","doi-asserted-by":"publisher","unstructured":"Argenson A, Dulac-Arnold G (2020) Model-based offline planning. arXiv:2008.05556. https:\/\/doi.org\/10.48550\/arXiv.2008.05556","DOI":"10.48550\/arXiv.2008.05556"},{"key":"9029_CR21","unstructured":"Hong Z-W, Agrawal P, Combes RT, Laroche R (2022) Harnessing mixed offline reinforcement learning datasets via trajectory weighting. In: The eleventh international conference on learning representations"},{"key":"9029_CR22","first-page":"1179","volume":"33","author":"A Kumar","year":"2020","unstructured":"Kumar A, Zhou A, Tucker G, Levine S (2020) Conservative q-learning for offline reinforcement learning. Adv Neural Inf Process Syst 33:1179\u20131191","journal-title":"Adv Neural Inf Process Syst"},{"key":"9029_CR23","first-page":"28954","volume":"34","author":"T Yu","year":"2021","unstructured":"Yu T, Kumar A, Rafailov R, Rajeswaran A, Levine S, Finn C (2021) Combo: conservative offline model-based policy optimization. Adv Neural Inf Process Syst 34:28954\u201328967","journal-title":"Adv Neural Inf Process Syst"},{"key":"9029_CR24","unstructured":"Shi L, Li G, Wei Y, Chen Y, Chi Y (2022) Pessimistic q-learning for offline reinforcement learning: towards optimal sample complexity. In: International conference on machine learning, pp 19967\u201320025. PMLR"},{"key":"9029_CR25","doi-asserted-by":"publisher","unstructured":"Ha D, Schmidhuber J (2018) World models. arXiv:1803.10122. https:\/\/doi.org\/10.48550\/arXiv.1803.10122","DOI":"10.48550\/arXiv.1803.10122"},{"issue":"3","key":"9029_CR26","doi-asserted-by":"publisher","first-page":"4401","DOI":"10.1109\/LRA.2021.3068655","volume":"6","author":"AS Chen","year":"2021","unstructured":"Chen AS, Nam H, Nair S, Finn C (2021) Batch exploration with examples for scalable robotic reinforcement learning. IEEE Robot Autom Lett 6(3):4401\u20134408. https:\/\/doi.org\/10.1109\/LRA.2021.3068655","journal-title":"IEEE Robot Autom Lett"},{"key":"9029_CR27","first-page":"16931","volume":"35","author":"D Driess","year":"2022","unstructured":"Driess D, Schubert I, Florence P, Li Y, Toussaint M (2022) Reinforcement learning with neural radiance fields. Adv Neural Inf Process Syst 35:16931\u201316945","journal-title":"Adv Neural Inf Process Syst"},{"key":"9029_CR28","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.126455","author":"Q Yi","year":"2023","unstructured":"Yi Q, Zhang R, Peng S, Guo J, Hu X, Du Z, Guo Q, Chen R, Li L, Chen Y (2023) Learning controllable elements oriented representations for reinforcement learning. Neurocomputing. https:\/\/doi.org\/10.1016\/j.neucom.2023.126455","journal-title":"Neurocomputing"},{"key":"9029_CR29","doi-asserted-by":"publisher","unstructured":"Cui B, Chow Y, Ghavamzadeh M (2020) Control-aware representations for model-based reinforcement learning. arXiv:2006.13408. https:\/\/doi.org\/10.48550\/arXiv.2006.13408","DOI":"10.48550\/arXiv.2006.13408"},{"key":"9029_CR30","unstructured":"Laskin M, Srinivas A, Abbeel P (2020) Curl: cunsupervised representations for reinforcement learning. In: International conference on machine learning, pp 5639\u20135650. PMLR"},{"key":"9029_CR31","unstructured":"Ma G, Wang Z, Yuan Z, Wang X, Yuan B, Tao D (2022) A comprehensive survey of data augmentation in visual reinforcement learning. arXiv:2210.04561"},{"key":"9029_CR32","unstructured":"Nair AV, Pong V, Dalal M, Bahl S, Lin S, Levine S (2018) Visual reinforcement learning with imagined goals. In: Advances in neural information processing systems, vol 31"},{"key":"9029_CR33","unstructured":"Gelada C, Kumar S, Buckman J, Nachum O, Bellemare MG (2019) Deepmdp: learning continuous latent space models for representation learning. In: International conference on machine learning, pp 2170\u20132179. PMLR"},{"key":"9029_CR34","unstructured":"Hafner D, Lillicrap T, Ba J, Norouzi M (2019) Dream to control: learning behaviors by latent imagination. In: International conference on learning representations"},{"key":"9029_CR35","doi-asserted-by":"publisher","unstructured":"Hafez MB, Weber C, Kerzel M, Wermter S (2019) Efficient intrinsically motivated robotic grasping with learning-adaptive imagination in latent space. In: 2019 Joint IEEE 9th international conference on development and learning and epigenetic robotics (Icdl-Epirob). IEEE, pp 1\u20137. https:\/\/doi.org\/10.1109\/DEVLRN.2019.8850723","DOI":"10.1109\/DEVLRN.2019.8850723"},{"key":"9029_CR36","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2020.103630","volume":"133","author":"MB Hafez","year":"2020","unstructured":"Hafez MB, Weber C, Kerzel M, Wermter S (2020) Improving robot dual-system motor learning with intrinsically motivated meta-control and latent-space experience imagination. Robot Auton Syst 133:103630. https:\/\/doi.org\/10.1016\/j.robot.2020.103630","journal-title":"Robot Auton Syst"},{"key":"9029_CR37","doi-asserted-by":"publisher","unstructured":"Han D, Doya K, Tani J (2019) Variational recurrent models for solving partially observable control tasks. arXiv:1912.10703. https:\/\/doi.org\/10.48550\/arXiv.1912.10703","DOI":"10.48550\/arXiv.1912.10703"},{"key":"9029_CR38","first-page":"741","volume":"33","author":"AX Lee","year":"2020","unstructured":"Lee AX, Nagabandi A, Abbeel P, Levine S (2020) Stochastic latent actor-critic: deep reinforcement learning with a latent variable model. Adv Neural Inf Process Syst 33:741\u2013752","journal-title":"Adv Neural Inf Process Syst"},{"issue":"1","key":"9029_CR39","first-page":"1437","volume":"16","author":"J Garc\u0131a","year":"2015","unstructured":"Garc\u0131a J, Fern\u00e1ndez F (2015) A comprehensive survey on safe reinforcement learning. J Mach Learn Res 16(1):1437\u20131480","journal-title":"J Mach Learn Res"},{"key":"9029_CR40","unstructured":"Fei Y, Yang Z, Chen Y, Wang Z (2021) Exponential bellman equation and improved regret bounds for risk-sensitive reinforcement learning. In: Advances in neural information processing systems, vol 34"},{"key":"9029_CR41","doi-asserted-by":"crossref","unstructured":"Zhang K, Zhang X, Hu B, Basar T (2021) Derivative-free policy optimization for linear risk-sensitive and robust control design: implicit regularization and sample complexity. In: Advances in neural information processing systems, vol 34","DOI":"10.1007\/978-3-030-92270-2_1"},{"key":"9029_CR42","first-page":"32639","volume":"35","author":"I Greenberg","year":"2022","unstructured":"Greenberg I, Chow Y, Ghavamzadeh M, Mannor S (2022) Efficient risk-averse reinforcement learning. Adv Neural Inf Process Syst 35:32639\u201332652","journal-title":"Adv Neural Inf Process Syst"},{"issue":"7","key":"9029_CR43","doi-asserted-by":"publisher","first-page":"325","DOI":"10.3390\/a16070325","volume":"16","author":"T Th\u00e9ate","year":"2023","unstructured":"Th\u00e9ate T, Ernst D (2023) Risk-sensitive policy with distributional reinforcement learning. Algorithms 16(7):325. https:\/\/doi.org\/10.3390\/a16070325","journal-title":"Algorithms"},{"issue":"5","key":"9029_CR44","doi-asserted-by":"publisher","first-page":"1281","DOI":"10.1111\/1468-0262.00158","volume":"68","author":"M Rabin","year":"2000","unstructured":"Rabin M (2000) Risk aversion and expected-utility theory: a calibration theorem. Econometrica 68(5):1281\u20131292","journal-title":"Econometrica"},{"issue":"4","key":"9029_CR45","doi-asserted-by":"publisher","first-page":"297","DOI":"10.1007\/BF00122574","volume":"5","author":"A Tversky","year":"1992","unstructured":"Tversky A, Kahneman D (1992) Advances in prospect theory: cumulative representation of uncertainty. J Risk Uncertain 5(4):297\u2013323","journal-title":"J Risk Uncertain"},{"issue":"7","key":"9029_CR46","doi-asserted-by":"publisher","first-page":"1443","DOI":"10.1016\/S0378-4266(02)00271-6","volume":"26","author":"RT Rockafellar","year":"2002","unstructured":"Rockafellar RT, Uryasev S (2002) Conditional value-at-risk for general loss distributions. J Bank Finance 26(7):1443\u20131471","journal-title":"J Bank Finance"},{"key":"9029_CR47","doi-asserted-by":"publisher","unstructured":"Sarykalin S, Serraino G, Uryasev S (2008) Value-at-risk vs. conditional value-at-risk in risk management and optimization. In: State-of-the-art decision-making tools in the information-intensive age. Informs, Maryland, pp 270\u2013294. https:\/\/doi.org\/10.1287\/educ.1080.0052","DOI":"10.1287\/educ.1080.0052"},{"issue":"3","key":"9029_CR48","doi-asserted-by":"publisher","first-page":"203","DOI":"10.1111\/1467-9965.00068","volume":"9","author":"P Artzner","year":"1999","unstructured":"Artzner P, Delbaen F, Eber J-M, Heath D (1999) Coherent measures of risk. Math Finance 9(3):203\u2013228","journal-title":"Math Finance"},{"key":"9029_CR49","unstructured":"Pinto L, Davidson J, Sukthankar R, Gupta A (2017) Robust adversarial reinforcement learning. In: International conference on machine learning. PMLR, pp 2817\u20132826"},{"key":"9029_CR50","unstructured":"Chow Y, Ghavamzadeh M (2014) Algorithms for cvar optimization in mdps. In: Advances in neural information processing systems, vol 27"},{"key":"9029_CR51","unstructured":"Chow Y, Tamar A, Mannor S, Pavone M (2015) Risk-sensitive and robust decision-making: a cvar optimization approach. In: Advances in neural information processing systems, vol 28"},{"key":"9029_CR52","doi-asserted-by":"publisher","unstructured":"Ying C, Zhou X, Su H, Yan D, Chen N, Zhu J (2022) Towards safe reinforcement learning via constraining conditional value-at-risk. arXiv:2206.04436. https:\/\/doi.org\/10.48550\/arXiv.2206.04436","DOI":"10.48550\/arXiv.2206.04436"},{"key":"9029_CR53","doi-asserted-by":"publisher","unstructured":"Ma X, Xia L, Zhou Z, Yang J, Zhao Q (2020) Dsac: distributional actor critic for risk-sensitive reinforcement learning. arXiv:2004.14547. https:\/\/doi.org\/10.48550\/arXiv.2004.14547","DOI":"10.48550\/arXiv.2004.14547"},{"key":"9029_CR54","unstructured":"Armengol\u00a0Urp\u00ed N, Curi S, Krause A (2021) Risk-averse offline reinforcement learning. In: International conference on learning representations (ICLR 2021). OpenReview"},{"key":"9029_CR55","unstructured":"Fujimoto S, Meger D, Precup D (2019) Off-policy deep reinforcement learning without exploration. In: International conference on machine learning. PMLR, pp 2052\u20132062"},{"key":"9029_CR56","unstructured":"Ma Y, Jayaraman D, Bastani O (2021) Conservative offline distributional reinforcement learning. In: Advances in neural information processing systems, vol 34"},{"key":"9029_CR57","doi-asserted-by":"crossref","unstructured":"Rockafellar RT (2007) Coherent approaches to risk in optimization under uncertainty. In: OR tools and applications: glimpses of future technologies. Informs, Maryland, USA, pp 38\u201361","DOI":"10.1287\/educ.1073.0032"},{"key":"9029_CR58","doi-asserted-by":"publisher","unstructured":"Delbaen F (2002) Coherent risk measures on general probability spaces. In: Advances in finance and stochastics. Springer, Berlin, pp 1\u201337. https:\/\/doi.org\/10.1007\/978-3-662-04790-3_1","DOI":"10.1007\/978-3-662-04790-3_1"},{"key":"9029_CR59","doi-asserted-by":"crossref","unstructured":"Rockafellar RT, Uryasev SP, Zabarankin M (2002) Deviation measures in risk analysis and optimization. University of Florida, Department of Industrial & Systems Engineering Working Paper (2002\u20137)","DOI":"10.2139\/ssrn.365640"},{"key":"9029_CR60","doi-asserted-by":"publisher","DOI":"10.2307\/253675","author":"SS Wang","year":"2000","unstructured":"Wang SS (2000) A class of distortion operators for pricing financial and insurance risks. J Risk Insur. https:\/\/doi.org\/10.2307\/253675","journal-title":"J Risk Insur"},{"key":"9029_CR61","doi-asserted-by":"publisher","unstructured":"Ahmadi-Javid A (2011) An information-theoretic approach to constructing coherent risk measures. In: 2011 IEEE international symposium on information theory proceedings. IEEE, pp 2125\u20132127. https:\/\/doi.org\/10.1109\/ISIT.2011.6033932","DOI":"10.1109\/ISIT.2011.6033932"},{"key":"9029_CR62","doi-asserted-by":"publisher","first-page":"21","DOI":"10.21314\/JOR.2000.038","volume":"2","author":"RT Rockafellar","year":"2000","unstructured":"Rockafellar RT, Uryasev S et al (2000) Optimization of conditional value-at-risk. J Risk 2:21\u201342","journal-title":"J Risk"},{"issue":"1","key":"9029_CR63","doi-asserted-by":"publisher","first-page":"51","DOI":"10.1007\/s00780-005-0165-8","volume":"10","author":"RT Rockafellar","year":"2006","unstructured":"Rockafellar RT, Uryasev S, Zabarankin M (2006) Generalized deviations in risk analysis. Finance Stoch 10(1):51\u201374. https:\/\/doi.org\/10.1007\/s00780-005-0165-8","journal-title":"Finance Stoch"},{"key":"9029_CR64","doi-asserted-by":"publisher","unstructured":"Levine S (2018) Reinforcement learning and control as probabilistic inference: Tutorial and review. arXiv:1805.00909. https:\/\/doi.org\/10.48550\/arXiv.1805.00909","DOI":"10.48550\/arXiv.1805.00909"},{"key":"9029_CR65","doi-asserted-by":"publisher","DOI":"10.1007\/s11370-021-00402-6","author":"S Kim","year":"2022","unstructured":"Kim S, Jo H, Song J-B (2022) Object manipulation system based on image-based reinforcement learning. Intell Serv Robot. https:\/\/doi.org\/10.1007\/s11370-021-00402-6","journal-title":"Intell Serv Robot"},{"key":"9029_CR66","unstructured":"Hafner D, Lillicrap T, Fischer I, Villegas R, Ha D, Lee H, Davidson J (2019) Learning latent dynamics for planning from pixels. In: International conference on machine learning. PMLR, pp 2555\u20132565"},{"key":"9029_CR67","doi-asserted-by":"publisher","unstructured":"Odaibo S (2019) Tutorial: deriving the standard variational autoencoder (vae) loss function. arXiv:1907.08956. https:\/\/doi.org\/10.48550\/arXiv.1907.08956","DOI":"10.48550\/arXiv.1907.08956"},{"key":"9029_CR68","doi-asserted-by":"publisher","first-page":"73","DOI":"10.1214\/aoms\/1177703732","volume":"35","author":"PJ Huber","year":"1964","unstructured":"Huber PJ (1964) Robust estimation of a location parameter. Ann Math Stat 35:73\u2013101","journal-title":"Ann Math Stat"},{"issue":"7","key":"9029_CR69","doi-asserted-by":"publisher","first-page":"1505","DOI":"10.1016\/S0378-4266(02)00281-9","volume":"26","author":"C Acerbi","year":"2002","unstructured":"Acerbi C (2002) Spectral measures of risk: a coherent representation of subjective risk aversion. J Bank Finance 26(7):1505\u20131518. https:\/\/doi.org\/10.1016\/S0378-4266(02)00281-9","journal-title":"J Bank Finance"},{"key":"9029_CR70","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2021.3082568","author":"J Duan","year":"2021","unstructured":"Duan J, Guan Y, Li SE, Ren Y, Sun Q, Cheng B (2021) Distributional actor-critic: off-policy reinforcement learning for addressing value estimation errors. IEEE Trans Neural Netw Learn Syst. https:\/\/doi.org\/10.1109\/TNNLS.2021.3082568","journal-title":"IEEE Trans Neural Netw Learn Syst"},{"key":"9029_CR71","doi-asserted-by":"crossref","unstructured":"Dabney W, Ostrovski G, Silver D, Munos R (2018) Implicit quantile networks for distributional reinforcement learning. In: International conference on machine learning. PMLR, pp 1096\u20131105","DOI":"10.1609\/aaai.v32i1.11791"},{"key":"9029_CR72","unstructured":"Ba JL, Kiros JR, Hinton GE (2016) Layer normalization. arXiv:1607.06450"},{"key":"9029_CR73","unstructured":"Fujimoto S, Hoof H, Meger D (2018) Addressing function approximation error in actor-critic methods. In: International conference on machine learning. PMLR, pp 1587\u20131596"},{"key":"9029_CR74","doi-asserted-by":"publisher","unstructured":"Fu J, Kumar A, Nachum O, Tucker G, Levine S (2020) D4rl: datasets for deep data-driven reinforcement learning. arXiv:2004.07219. https:\/\/doi.org\/10.48550\/arXiv.2004.07219","DOI":"10.48550\/arXiv.2004.07219"},{"key":"9029_CR75","first-page":"29304","volume":"34","author":"R Agarwal","year":"2021","unstructured":"Agarwal R, Schwarzer M, Castro PS, Courville AC, Bellemare M (2021) Deep reinforcement learning at the edge of the statistical precipice. Adv Neural Inf Process Syst 34:29304\u201329320","journal-title":"Adv Neural Inf Process Syst"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-023-09029-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-023-09029-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-023-09029-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,1,5]],"date-time":"2024-01-05T07:12:50Z","timestamp":1704438770000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-023-09029-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,11]]},"references-count":75,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,1]]}},"alternative-id":["9029"],"URL":"https:\/\/doi.org\/10.1007\/s00521-023-09029-3","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,11]]},"assertion":[{"value":"18 April 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 September 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 October 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant financial or non-financial interest to disclose.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}