{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T12:36:06Z","timestamp":1781181366477,"version":"3.54.1"},"reference-count":53,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2023,10,5]],"date-time":"2023-10-05T00:00:00Z","timestamp":1696464000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,10,5]],"date-time":"2023-10-05T00:00:00Z","timestamp":1696464000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100002428","name":"Austrian Science Fund","doi-asserted-by":"publisher","award":["F65"],"award-info":[{"award-number":["F65"]}],"id":[{"id":"10.13039\/501100002428","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100007601","name":"Horizon 2020","doi-asserted-by":"publisher","award":["754411"],"award-info":[{"award-number":["754411"]}],"id":[{"id":"10.13039\/501100007601","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004955","name":"\u00d6sterreichische Forschungsf\u00f6rderungsgesellschaft","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100004955","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004955","name":"\u00d6sterreichische Forschungsf\u00f6rderungsgesellschaft","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100004955","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012306","name":"Universit\u00e0 degli Studi di Trieste","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100012306","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2024,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We investigate the potential of Multi-Objective, Deep Reinforcement Learning for stock and cryptocurrency single-asset trading: in particular, we consider a Multi-Objective algorithm which generalizes the reward functions and discount factor (i.e., these components are not specified a priori, but incorporated in the learning process). Firstly, using several important assets (BTCUSD, ETHUSDT, XRPUSDT, AAPL, SPY, NIFTY50), we verify the reward generalization property of the proposed Multi-Objective algorithm, and provide preliminary statistical evidence showing increased predictive stability over the corresponding Single-Objective strategy. Secondly, we show that the Multi-Objective algorithm has a clear edge over the corresponding Single-Objective strategy when the reward mechanism is sparse (i.e., when non-null feedback is infrequent over time). Finally, we discuss the generalization properties with respect to the discount factor. The entirety of our code is provided in open-source format.<\/jats:p>","DOI":"10.1007\/s00521-023-09033-7","type":"journal-article","created":{"date-parts":[[2023,10,5]],"date-time":"2023-10-05T04:01:30Z","timestamp":1696478490000},"page":"619-637","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["Multi-objective reward generalization: improving performance of Deep Reinforcement Learning for applications in single-asset trading"],"prefix":"10.1007","volume":"36","author":[{"given":"Federico","family":"Cornalba","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Constantin","family":"Disselkamp","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Davide","family":"Scassola","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christopher","family":"Helf","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,10,5]]},"reference":[{"key":"9033_CR1","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Graves A, Antonoglou I, Wierstra D, Riedmiller M (2013) Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602"},{"issue":"11","key":"9033_CR2","doi-asserted-by":"publisher","first-page":"1238","DOI":"10.1177\/0278364913495721","volume":"32","author":"J Kober","year":"2013","unstructured":"Kober J, Bagnell JA, Peters J (2013) Reinforcement learning in robotics: a survey. Int J Robot Res 32(11):1238\u20131274","journal-title":"Int J Robot Res"},{"key":"9033_CR3","doi-asserted-by":"crossref","unstructured":"Zheng G, Zhang F, Zheng Z, Xiang Y, Yuan NJ, Xie X, Li Z (2018) Drn: a deep reinforcement learning framework for news recommendation. In: Proceedings of the 2018 World Wide Web Conference, 167\u2013176","DOI":"10.1145\/3178876.3185994"},{"key":"9033_CR4","doi-asserted-by":"crossref","unstructured":"Mao H, Alizadeh M, Menache I, Kandula S (2016) Resource management with deep reinforcement learning. In: Proceedings of the 15th ACM Workshop on Hot Topics in Networks, 50\u201356","DOI":"10.1145\/3005745.3005750"},{"key":"9033_CR5","unstructured":"Friedman, E., Fontaine, F.: Generalizing across multi-objective reward functions in deep reinforcement learning. arXiv preprint arXiv:1809.06364 (2018)"},{"key":"9033_CR6","doi-asserted-by":"crossref","unstructured":"Castelletti A, Pianosi F, Restelli M (2012) Tree-based fitted q-iteration for multi-objective markov decision problems. In: The 2012 International Joint Conference on Neural Networks (IJCNN), 1\u20138. IEEE","DOI":"10.1109\/IJCNN.2012.6252759"},{"key":"9033_CR7","unstructured":"Ernst D, Geurts P, Wehenkel L (2005) Tree-based batch mode reinforcement learning. J Mach Learn Res 6"},{"key":"9033_CR8","doi-asserted-by":"crossref","unstructured":"Lee JW, Jangmin O (2002) A multi-agent q-learning framework for optimizing stock trading systems. In: International Conference on Database and Expert Systems Applications, 153\u2013162. Springer","DOI":"10.1007\/3-540-46146-9_16"},{"key":"9033_CR9","doi-asserted-by":"crossref","unstructured":"Bisht K, Kumar A (2020) Deep reinforcement learning based multi-objective systems for financial trading. In: 2020 5th IEEE International Conference on Recent Advances and Innovations in Engineering (ICRAIE), 1\u20136. IEEE","DOI":"10.1109\/ICRAIE51050.2020.9358319"},{"key":"9033_CR10","doi-asserted-by":"crossref","unstructured":"Si W, Li J, Ding P, Rao R (2017) A multi-objective deep reinforcement learning approach for stock index future\u2019s intraday trading. In: 2017 10th International Symposium on Computational Intelligence and Design (ISCID), 431\u2013436:2. IEEE","DOI":"10.1109\/ISCID.2017.210"},{"key":"9033_CR11","unstructured":"Sutton RS, Barto AG (2018) Reinforcement Learning: An Introduction. MIT press"},{"issue":"1","key":"9033_CR12","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s10458-019-09433-x","volume":"34","author":"R R\u0103dulescu","year":"2020","unstructured":"R\u0103dulescu R, Mannion P, Roijers DM, Now\u00e9 A (2020) Multi-objective multi-agent decision making: a utility-based analysis and survey. Auton Agent Multi-Agent Syst 34(1):1\u201352","journal-title":"Auton Agent Multi-Agent Syst"},{"key":"9033_CR13","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1613\/jair.3987","volume":"48","author":"DM Roijers","year":"2013","unstructured":"Roijers DM, Vamplew P, Whiteson S, Dazeley R (2013) A survey of multi-objective sequential decision-making. J Artif Intell Res 48:67\u2013113","journal-title":"J Artif Intell Res"},{"issue":"1","key":"9033_CR14","doi-asserted-by":"publisher","first-page":"26","DOI":"10.1007\/s10458-022-09552-y","volume":"36","author":"CF Hayes","year":"2022","unstructured":"Hayes CF, R\u0103dulescu R, Bargiacchi E, K\u00e4llstr\u00f6m J, Macfarlane M, Reymond M, Verstraeten T, Zintgraf LM, Dazeley R, Heintz F (2022) A practical guide to multi-objective reinforcement learning and planning. Auton Agent Multi-Agent Syst 36(1):26","journal-title":"Auton Agent Multi-Agent Syst"},{"key":"9033_CR15","doi-asserted-by":"crossref","unstructured":"Zitzler E, Knowles J, Thiele L (2008) Quality assessment of pareto set approximations. Multiobjective optimization, 373\u2013404","DOI":"10.1007\/978-3-540-88908-3_14"},{"key":"9033_CR16","unstructured":"Reymond M, Now\u00e9 A (2019) Pareto-dqn: Approximating the pareto front in complex multi-objective decision problems. In: Proceedings of the Adaptive and Learning Agents Workshop (ALA-19) at AAMAS"},{"key":"9033_CR17","doi-asserted-by":"crossref","unstructured":"Natarajan S, Tadepalli P (2005) Dynamic preferences in multi-criteria reinforcement learning. In: Proceedings of the 22nd International Conference on Machine Learning, 601\u2013608","DOI":"10.1145\/1102351.1102427"},{"key":"9033_CR18","doi-asserted-by":"crossref","unstructured":"Barrett L, Narayanan S (2008) Learning all optimal policies with multiple criteria. In: Proceedings of the 25th International Conference on Machine Learning, 41\u201347","DOI":"10.1145\/1390156.1390162"},{"key":"9033_CR19","unstructured":"Andrychowicz M, Wolski F, Ray A, Schneider J, Fong R, Welinder P, McGrew B, Tobin J, Pieter Abbeel O, Zaremba W (2017) Hindsight experience replay. Adv Neural Inf Proc Syst 30"},{"key":"9033_CR20","unstructured":"Abels A, Roijers D, Lenaerts T, Now\u00e9 A, Steckelmacher D (2019) Dynamic weights in multi-objective deep reinforcement learning. In: International Conference on Machine Learning, 11\u201320. PMLR"},{"key":"9033_CR21","unstructured":"K\u00e4llstr\u00f6m J, Heintz F (2019) Tunable dynamics in agent-based simulation using multi-objective reinforcement learning. In: Adaptive and Learning Agents Workshop (ALA-19) at AAMAS, Montreal, Canada, May 13-14, 2019, 1\u20137"},{"key":"9033_CR22","unstructured":"Mossalam H, Assael YM, Roijers DM, Whiteson S (2016) Multi-objective deep reinforcement learning. arXiv preprint arXiv:1610.02707"},{"key":"9033_CR23","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2020.103915","volume":"96","author":"TT Nguyen","year":"2020","unstructured":"Nguyen TT, Nguyen ND, Vamplew P, Nahavandi S, Dazeley R, Lim CP (2020) A multi-objective deep reinforcement learning framework. Eng Appl Artif Intell 96:103915","journal-title":"Eng Appl Artif Intell"},{"key":"9033_CR24","unstructured":"Tajmajer T (2017) Multi-objective deep q-learning with subsumption architecture. arXiv preprint arXiv:1704.06676"},{"key":"9033_CR25","doi-asserted-by":"crossref","unstructured":"Tajmajer T (2018) Modular multi-objective deep reinforcement learning with decision values. In: 2018 Federated Conference on Computer Science and Information Systems (FedCSIS), 85\u201393 . IEEE","DOI":"10.15439\/2018F231"},{"key":"9033_CR26","doi-asserted-by":"crossref","unstructured":"Dusparic I, Cahill V (2009) Distributed w-learning: Multi-policy optimization in self-organizing systems. In: 2009 Third IEEE International Conference on Self-adaptive and Self-organizing Systems, 20\u201329. IEEE","DOI":"10.1109\/SASO.2009.23"},{"key":"9033_CR27","unstructured":"Shelton C (2000) Balancing multiple sources of reward in reinforcement learning. Adv Neural Inf Proc Syst 13"},{"key":"9033_CR28","doi-asserted-by":"crossref","unstructured":"Hafez MB, Weber C, Kerzel M, Wermter S (2019) Efficient intrinsically motivated robotic grasping with learning-adaptive imagination in latent space. In: 2019 Joint Ieee 9th International Conference on Development and Learning and Epigenetic Robotics (Icdl-Epirob), 1\u20137. IEEE","DOI":"10.1109\/DEVLRN.2019.8850723"},{"key":"9033_CR29","unstructured":"Chen E, Hong Z-W, Pajarinen J, Agrawal P (2022) Redeeming intrinsic rewards via constrained optimization. arXiv preprint arXiv:2211.07627"},{"key":"9033_CR30","unstructured":"Fischer TG (2018) Reinforcement learning in financial markets-a survey. Technical report, FAU Discussion Papers in Economics"},{"key":"9033_CR31","unstructured":"Neuneier R (1995) Optimal asset allocation using adaptive dynamic programming. Adv Neural Inf Proc Syst 8"},{"key":"9033_CR32","doi-asserted-by":"crossref","unstructured":"Corazza M, Bertoluzzo F (2014) Q-learning-based financial trading systems with applications. Working Papers 2014:15, Department of Economics, University of Venice \"Ca\u2019 Foscari\". https:\/\/EconPapers.repec.org\/RePEc:ven:wpaper:2014:15","DOI":"10.2139\/ssrn.2507826"},{"key":"9033_CR33","unstructured":"Jin O, El-Saawy H (2016) Portfolio management using reinforcement learning. Stanford University"},{"issue":"4","key":"9033_CR34","doi-asserted-by":"publisher","first-page":"744","DOI":"10.1109\/72.935088","volume":"12","author":"MA Dempster","year":"2001","unstructured":"Dempster MA, Payne TW, Romahi Y, Thompson GW (2001) Computational learning techniques for intraday fx trading using popular technical indicators. IEEE Trans Neural Netw 12(4):744\u2013754","journal-title":"IEEE Trans Neural Netw"},{"key":"9033_CR35","unstructured":"Gu Y, Mabu S, Yang Y, Li J, Hirasawa K (2011) Trading rules on stock markets using genetic network programming-sarsa learning with plural subroutines. In: SICE Annual Conference 2011, 143\u2013148. IEEE"},{"issue":"5","key":"9033_CR36","doi-asserted-by":"publisher","first-page":"4741","DOI":"10.1016\/j.eswa.2010.09.001","volume":"38","author":"Z Tan","year":"2011","unstructured":"Tan Z, Quek C, Cheng PY (2011) Stock trading with cycles: A financial application of anfis and reinforcement learning. Expert Syst Appl 38(5):4741\u20134755","journal-title":"Expert Syst Appl"},{"key":"9033_CR37","doi-asserted-by":"publisher","first-page":"100","DOI":"10.1016\/j.dss.2014.04.011","volume":"64","author":"D Eilers","year":"2014","unstructured":"Eilers D, Dunis CL, Mettenheim H-J, Breitner MH (2014) Intelligent trading of seasonal effects: A decision support algorithm based on reinforcement learning. Decis Support Syst 64:100\u2013108","journal-title":"Decis Support Syst"},{"key":"9033_CR38","doi-asserted-by":"crossref","unstructured":"Sherstov AA, Stone P (2004) Three automated stock-trading agents: A comparative study. In: International Workshop on Agent-Mediated Electronic Commerce, 173\u2013187. Springer","DOI":"10.1007\/11575726_13"},{"key":"9033_CR39","doi-asserted-by":"crossref","unstructured":"Nevmyvaka Y, Feng Y, Kearns M (2006) Reinforcement learning for optimized trade execution. In: Proceedings of the 23rd International Conference on Machine Learning, 673\u2013680","DOI":"10.1145\/1143844.1143929"},{"key":"9033_CR40","unstructured":"Kaur S (2017) Algorithmic trading using reinforcement learning augmented with hidden markov model. Technical report, Working paper, Stanford University"},{"issue":"15","key":"9033_CR41","doi-asserted-by":"publisher","first-page":"2121","DOI":"10.1016\/j.ins.2005.10.009","volume":"176","author":"O Jangmin","year":"2006","unstructured":"Jangmin O, Lee J, Lee JW, Zhang B-T (2006) Adaptive stock trading with dynamic asset allocation using reinforcement learning. Inf Sci 176(15):2121\u20132147","journal-title":"Inf Sci"},{"key":"9033_CR42","unstructured":"Watts S (2015) Hedging basis risk using reinforcement learning. Technical report, Technical report, Working Paper, University of Oxford"},{"issue":"5\u20136","key":"9033_CR43","doi-asserted-by":"publisher","first-page":"441","DOI":"10.1002\/(SICI)1099-131X(1998090)17:5\/6<441::AID-FOR707>3.0.CO;2-#","volume":"17","author":"J Moody","year":"1998","unstructured":"Moody J, Wu L, Liao Y, Saffell M (1998) Performance functions and reinforcement learning for trading systems and portfolios. J Forecast 17(5\u20136):441\u2013470","journal-title":"J Forecast"},{"key":"9033_CR44","doi-asserted-by":"crossref","unstructured":"Gold C (2003) Fx trading via recurrent reinforcement learning. In: 2003 IEEE International Conference on Computational Intelligence for Financial Engineering, 2003. Proceedings., 363\u2013370. IEEE","DOI":"10.1109\/CIFER.2003.1196283"},{"issue":"3","key":"9033_CR45","doi-asserted-by":"publisher","first-page":"543","DOI":"10.1016\/j.eswa.2005.10.012","volume":"30","author":"MA Dempster","year":"2006","unstructured":"Dempster MA, Leemans V (2006) An automated fx trading system using adaptive reinforcement learning. Expert Syst Appl 30(3):543\u2013552","journal-title":"Expert Syst Appl"},{"issue":"3","key":"9033_CR46","doi-asserted-by":"publisher","first-page":"653","DOI":"10.1109\/TNNLS.2016.2522401","volume":"28","author":"Y Deng","year":"2016","unstructured":"Deng Y, Bao F, Kong Y, Ren Z, Dai Q (2016) Deep direct reinforcement learning for financial signal representation and trading. IEEE Trans Neural Netw Learn syst 28(3):653\u2013664","journal-title":"IEEE Trans Neural Netw Learn syst"},{"key":"9033_CR47","unstructured":"Jiang Z, Xu D, Liang J (2017) A deep reinforcement learning framework for the financial portfolio management problem. arXiv preprint arXiv:1706.10059"},{"key":"9033_CR48","doi-asserted-by":"crossref","unstructured":"Li H, Dagli CH, Enke D (2007) Short-term stock market timing prediction under reinforcement learning schemes. In: 2007 IEEE International Symposium on Approximate Dynamic Programming and Reinforcement Learning, 233\u2013240. IEEE","DOI":"10.1109\/ADPRL.2007.368193"},{"issue":"6","key":"9033_CR49","doi-asserted-by":"publisher","first-page":"1153","DOI":"10.1016\/j.jedc.2010.01.015","volume":"34","author":"SD Bekiros","year":"2010","unstructured":"Bekiros SD (2010) Heterogeneous trading strategies with adaptive fuzzy actor-critic reinforcement learning: A behavioral approach. J Econ Dyn Control 34(6):1153\u20131170","journal-title":"J Econ Dyn Control"},{"key":"9033_CR50","unstructured":"Chan NT, Shelton C (2001) An adaptive electronic market-maker. Computing in Economics and Finance 2001 146, Society for Computational Economics. https:\/\/EconPapers.repec.org\/RePEc:sce:scecf1:146"},{"issue":"6","key":"9033_CR51","doi-asserted-by":"publisher","first-page":"864","DOI":"10.1109\/TSMCA.2007.904825","volume":"37","author":"JW Lee","year":"2007","unstructured":"Lee JW, Park J, Jangmin O, Lee J, Hong E (2007) A multiagent approach to $$ q $$-learning for daily stock trading. IEEE Trans Syst, Man, Cybern-Part A: Syst Humans 37(6):864\u2013877","journal-title":"IEEE Trans Syst, Man, Cybern-Part A: Syst Humans"},{"key":"9033_CR52","unstructured":"Lee JW, Zhang B-T (2002) Stock trading system using reinforcement learning with cooperative agents. In: Proceedings of the Nineteenth International Conference on Machine Learning, 451\u2013458"},{"key":"9033_CR53","unstructured":"Fedus W, Ramachandran P, Agarwal R, Bengio Y, Larochelle H, Rowland M, Dabney W (2020) Revisiting fundamentals of experience replay. In: International Conference on Machine Learning, 3061\u20133071. PMLR"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-023-09033-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-023-09033-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-023-09033-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,30]],"date-time":"2024-10-30T01:57:53Z","timestamp":1730253473000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-023-09033-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,5]]},"references-count":53,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,1]]}},"alternative-id":["9033"],"URL":"https:\/\/doi.org\/10.1007\/s00521-023-09033-7","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,5]]},"assertion":[{"value":"4 May 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 September 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 October 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant financial or non-financial interests to disclose.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest statement"}}]}}