{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,8]],"date-time":"2025-09-08T06:39:56Z","timestamp":1757313596060,"version":"3.41.0"},"reference-count":49,"publisher":"Springer Science and Business Media LLC","issue":"19","license":[{"start":{"date-parts":[[2022,7,5]],"date-time":"2022-07-05T00:00:00Z","timestamp":1656979200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,7,5]],"date-time":"2022-07-05T00:00:00Z","timestamp":1656979200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"National University Ireland, Galway"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2025,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>In many real-world scenarios, the utility of a user is derived from a single execution of a policy. In this case, to apply multi-objective reinforcement learning, the expected utility of the returns must be optimised. Various scenarios exist where a user\u2019s preferences over objectives (also known as the utility function) are unknown or difficult to specify. In such scenarios, a set of optimal policies must be learned. However, settings where the expected utility must be maximised have been largely overlooked by the multi-objective reinforcement learning community and, as a consequence, a set of optimal solutions has yet to be defined. In this work, we propose first-order stochastic dominance as a criterion to build solution sets to maximise expected utility. We also define a new dominance criterion, known as expected scalarised returns (ESR) dominance, that extends first-order stochastic dominance to allow a set of optimal policies to be learned in practice. Additionally, we define a new solution concept called the ESR set, which is a set of policies that are ESR dominant. Finally, we present a new multi-objective tabular distributional reinforcement learning (MOTDRL) algorithm to learn the ESR set in multi-objective multi-armed bandit settings.<\/jats:p>","DOI":"10.1007\/s00521-022-07334-x","type":"journal-article","created":{"date-parts":[[2022,7,5]],"date-time":"2022-07-05T11:09:06Z","timestamp":1657019346000},"page":"13079-13099","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Expected scalarised returns dominance: a new solution concept for multi-objective decision making"],"prefix":"10.1007","volume":"37","author":[{"given":"Conor F.","family":"Hayes","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Timothy","family":"Verstraeten","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Diederik M.","family":"Roijers","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Enda","family":"Howley","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Patrick","family":"Mannion","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,7,5]]},"reference":[{"doi-asserted-by":"publisher","unstructured":"Ali MM (1975) Stochastic dominance and portfolio analysis. J Finan Econ 2(2): 205\u2013229. https:\/\/doi.org\/10.1016\/0304-405X(75)90005-7. https:\/\/www.sciencedirect.com\/science\/article\/pii\/0304405X75900057","key":"7334_CR1","DOI":"10.1016\/0304-405X(75)90005-7"},{"issue":"2","key":"7334_CR2","doi-asserted-by":"publisher","first-page":"183","DOI":"10.2307\/2297269","volume":"49","author":"AB Atkinson","year":"1982","unstructured":"Atkinson AB, Bourguignon F (1982) The comparison of multi-dimensioned distributions of economic status. Rev Econ Stud 49(2):183\u2013201. https:\/\/doi.org\/10.2307\/2297269","journal-title":"Rev Econ Stud"},{"unstructured":"Auer P, Chiang CK, Ortner R, Drugan M (2016) Pareto front identification from stochastic bandit feedback. In: Gretton A, Robert CC (eds) Proceedings of the 19th international conference on artificial intelligence and statistics, proceedings of machine learning research, vol\u00a051, pp 939\u2013947. PMLR, Cadiz, Spain. http:\/\/proceedings.mlr.press\/v51\/auer16.html","key":"7334_CR3"},{"doi-asserted-by":"publisher","unstructured":"Bawa VS (1975) Optimal rules for ordering uncertain prospects. J Finan Econ 2(1): 95\u2013121. https:\/\/doi.org\/10.1016\/0304-405X(75)90025-2. http:\/\/www.sciencedirect.com\/science\/article\/pii\/0304405X75900252","key":"7334_CR4","DOI":"10.1016\/0304-405X(75)90025-2"},{"doi-asserted-by":"crossref","unstructured":"Bawa VS (1978) Safety-first, stochastic dominance, and optimal portfolio choice. J Finan Quant Anal 13(2): 255\u2013271. http:\/\/www.jstor.org\/stable\/2330386","key":"7334_CR5","DOI":"10.2307\/2330386"},{"issue":"6","key":"7334_CR6","doi-asserted-by":"publisher","first-page":"698","DOI":"10.1287\/mnsc.28.6.698","volume":"28","author":"VS Bawa","year":"1982","unstructured":"Bawa VS (1982) Research bibliography-stochastic dominance: a research bibliography. Manage Sci 28(6):698\u2013712. https:\/\/doi.org\/10.1287\/mnsc.28.6.698","journal-title":"Manage Sci"},{"unstructured":"Bellemare MG, Dabney W, Munos R (2017) A distributional perspective on reinforcement learning. In: International conference on machine learning, pp. 449\u2013458. PMLR, Sydney","key":"7334_CR7"},{"doi-asserted-by":"publisher","unstructured":"Choi E, Johnson S (1988) Stochastic dominance and uncertain price prospects. Center for agricultural and rural development (CARD) at Iowa State University, Center for Agricultural and Rural Development (CARD) Publications 55. https:\/\/doi.org\/10.2307\/1059583","key":"7334_CR8","DOI":"10.2307\/1059583"},{"key":"7334_CR9","doi-asserted-by":"publisher","DOI":"10.2514\/6.2018-0665","author":"L Cook","year":"2018","unstructured":"Cook L, Jarrett J (2018) Using stochastic dominance in multi-objective optimizers for aerospace design under uncertainty. Am Instit Aeronaut Astronaut J. https:\/\/doi.org\/10.2514\/6.2018-0665","journal-title":"Am Instit Aeronaut Astronaut J"},{"doi-asserted-by":"crossref","unstructured":"Darling DA (1957) The kolmogorov\u2013smirnov, cramer\u2013von mises tests. Ann Math Stat 28(4): 823\u2013838. http:\/\/www.jstor.org\/stable\/2237048","key":"7334_CR10","DOI":"10.1214\/aoms\/1177706788"},{"doi-asserted-by":"publisher","unstructured":"Drugan MM, Nowe A (2013) Designing multi-objective multi-armed bandits algorithms: a study. In: The 2013 international joint conference on neural networks (IJCNN), pp 1\u20138. https:\/\/doi.org\/10.1109\/IJCNN.2013.6707036","key":"7334_CR11","DOI":"10.1109\/IJCNN.2013.6707036"},{"key":"7334_CR12","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-021-05961-4","author":"G Dulac-Arnold","year":"2021","unstructured":"Dulac-Arnold G, Levine N, Mankowitz DJ, Li J, Paduraru C, Gowal S, Hester T (2021) Challenges of real-world reinforcement learning: definitions, benchmarks and analysis. Mach Learn. https:\/\/doi.org\/10.1007\/s10994-021-05961-4","journal-title":"Mach Learn"},{"issue":"1","key":"7334_CR13","doi-asserted-by":"publisher","first-page":"51","DOI":"10.1007\/BF01763120","volume":"7","author":"PC Fishburn","year":"1978","unstructured":"Fishburn PC (1978) Non-cooperative stochastic dominance games. Int J Game Theory 7(1):51\u201361","journal-title":"Int J Game Theory"},{"unstructured":"Hadar J, Russell WR (1969) Rules for ordering uncertain prospects. Am Econ Rev 59(1): 25\u201334. http:\/\/www.jstor.org\/stable\/1811090","key":"7334_CR14"},{"unstructured":"Hayes CF, Reymond M, Roijers DM, Howley E, Mannion P (2021) Distributional Monte Carlo tree search for risk-aware and multi-objective reinforcement learning. In: Proceedings of the 20th international conference on autonomous agents and multiagent systems, vol. 2021. IFAAMAS (2021 In Press)","key":"7334_CR15"},{"unstructured":"Hayes CF, Reymond M, Roijers DM, Howley E, Mannion P (2021) Risk-aware and multi-objective decision making with distributional Monte Carlo tree search. In: Proceedings of the adaptive and learning agents workshop at AAMAS 2021","key":"7334_CR16"},{"unstructured":"Hayes CF, Verstraeten T, Roijers DM, Howley E, Mannion P (2021) Dominance criteria and solution sets for the expected scalarised returns. In: Proceedings of the adaptive and learning agents workshop at AAMAS 2021 (2021)","key":"7334_CR17"},{"issue":"1","key":"7334_CR18","doi-asserted-by":"publisher","first-page":"26","DOI":"10.1007\/s10458-022-09552-y","volume":"36","author":"CF Hayes","year":"2022","unstructured":"Hayes CF, R\u0103dulescu R, Bargiacchi E, K\u00e4llstr\u00f6m J, Macfarlane M, Reymond M, Verstraeten T, Zintgraf LM, Dazeley R, Heintz F, Howley E, Irissappane AA, Mannion P, Now\u00e9 A, Ramos G, Restelli M, Vamplew P, Roijers DM (2022) A practical guide to multi-objective reinforcement learning and planning. Auton Agent Multi-Agent Syst 36(1):26. https:\/\/doi.org\/10.1007\/s10458-022-09552-y","journal-title":"Auton Agent Multi-Agent Syst"},{"doi-asserted-by":"crossref","unstructured":"Levhari D, Paroush J, Peleg B (1975) Efficiency analysis for multivariate distributions. Rev Econ Stud 42(1): 87\u201391. http:\/\/www.jstor.org\/stable\/2296822","key":"7334_CR19","DOI":"10.2307\/2296822"},{"doi-asserted-by":"crossref","unstructured":"Levy H (1992) Stochastic dominance and expected utility: survey and analysis. Manag Sci 38(4): 555\u2013593. http:\/\/www.jstor.org\/stable\/2632436","key":"7334_CR20","DOI":"10.1287\/mnsc.38.4.555"},{"unstructured":"Malerba F, Mannion P (2021) Evaluating tunable agents with non-linear utility functions under expected scalarised returns. In: Multi-objective decision making workshop (MODeM 2021)","key":"7334_CR21"},{"unstructured":"Martin J, Lyskawinski M, Li X, Englot B (2020) Stochastically dominant distributional reinforcement learning. In: International conference on machine learning, pp 6745\u20136754. PMLR","key":"7334_CR22"},{"key":"7334_CR23","volume-title":"Microeconomic theory","author":"A Mas-Colell","year":"1995","unstructured":"Mas-Colell A, Whinston MD, Green JR et al (1995) Microeconomic theory, vol 1. Oxford University Press, New York"},{"key":"7334_CR24","first-page":"3663","volume":"15","author":"KV Moffaert","year":"2014","unstructured":"Moffaert KV, Nowe A (2014) Multi-objective reinforcement learning using sets of pareto dominating policies. J Mach Learn Res 15:3663\u20133692","journal-title":"J Mach Learn Res"},{"doi-asserted-by":"publisher","unstructured":"Nakayama H, Tanino T, Sawaragi Y (1981) Stochastic dominance for decision problems with multiple attributes and\/or multiple decision-makers. IFAC proceedings volumes 14(2), 1397\u20131402. https:\/\/doi.org\/10.1016\/S1474-6670(17)63673-5.http:\/\/www.sciencedirect.com\/science\/article\/pii\/S1474667017636735. 8th IFAC World Congress on Control Science and Technology for the Progress of Society, Kyoto, Japan, 24-28 August 1981","key":"7334_CR25","DOI":"10.1016\/S1474-6670(17)63673-5."},{"unstructured":"O\u2019Callaghan D, Mannion P (2021) Exploring the impact of tunable agents in sequential social dilemmas. arXiv preprint: arXiv:2101.11967","key":"7334_CR26"},{"unstructured":"\u00d6ner D, Karakurt A, Ery\u0131lmaz A, Tekin C (2018) Combinatorial multi-objective multi-armed bandit problem","key":"7334_CR27"},{"unstructured":"Pareto V (1896) Manuel d\u2019Economie Politique, vol 1. Giard, Paris","key":"7334_CR28"},{"doi-asserted-by":"crossref","unstructured":"R\u0103dulescu R, Mannion P, Roijers DM, Now\u00e9 A (2020) Multi-objective multi-agent decision making: a utility-based analysis and survey. Auton Agents Multi-Agent Syst 34(10)","key":"7334_CR29","DOI":"10.1007\/s10458-019-09433-x"},{"doi-asserted-by":"crossref","unstructured":"R\u0103dulescu R, Mannion P, Zhang Y, Roijers DM, Now\u00e9 A (2020) A utility-based analysis of equilibria in multi-objective normal-form games. Knowl Eng Rev 35 (2020)","key":"7334_CR30","DOI":"10.1017\/S0269888920000351"},{"unstructured":"Reymond M, Hayes C, Roijers DM, Steckelmacher D, Now\u00e9 A (2021) Actor-critic multi-objective reinforcement learning for non-linear utility functions. In: Multi-objective decision making workshop (MODeM 2021)","key":"7334_CR31"},{"doi-asserted-by":"crossref","unstructured":"Richard SF (1975) Multivariate risk aversion, utility independence and separable utility functions. Manag Sci 22(1): 12\u201321. http:\/\/www.jstor.org\/stable\/2629784","key":"7334_CR32","DOI":"10.1287\/mnsc.22.1.12"},{"unstructured":"Roijers DM, Steckelmacher D, Now\u00e9 A (2018) Multi-objective reinforcement learning for the expected utility of the return. In: Proceedings of the adaptive and learning agents workshop at FAIM 2018","key":"7334_CR33"},{"unstructured":"Roijers DM, Whiteson S, Oliehoek FA (2014) Linear support for multi-objective coordination graphs. In: Proceedings of the 2014 international conference on autonomous agents and multi-agent systems, AAMAS \u201914, pp 1297\u20131304. International foundation for autonomous agents and multiagent systems, Richland, SC","key":"7334_CR34"},{"doi-asserted-by":"crossref","unstructured":"Roijers DM, Zintgraf LM, Now\u00e9 A (2017) Interactive thompson sampling for multi-objective multi-armed bandits. In: International conference on algorithmic decisiontheory, pp 18\u201334. Springer, New York","key":"7334_CR35","DOI":"10.1007\/978-3-319-67504-6_2"},{"key":"7334_CR36","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1613\/jair.3987","volume":"48","author":"DM Roijers","year":"2013","unstructured":"Roijers DM, Vamplew P, Whiteson S, Dazeley R (2013) A survey of multi-objective sequential decision-making. J Artif Intell Res 48:67\u2013113","journal-title":"J Artif Intell Res"},{"doi-asserted-by":"crossref","unstructured":"Scarsini, M.: Dominance conditions for multivariate utility functions. Manag Sci 34(4): 454\u2013460 (1988). http:\/\/www.jstor.org\/stable\/2631934","key":"7334_CR37","DOI":"10.1287\/mnsc.34.4.454"},{"issue":"1","key":"7334_CR38","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1007\/BF03024818","volume":"18","author":"N Schappacher","year":"1996","unstructured":"Schappacher N (1996) Beppo levi and the arithmetic of elliptic curves. Math Intell 18(1):57\u201369","journal-title":"Math Intell"},{"doi-asserted-by":"publisher","unstructured":"Sriboonchitta S, Wong WK, Dhompongsa s, Nguyen H (2009) Stochastic dominance and applications to finance, risk and economics. Chapman and Hall\/CRC, New York. https:\/\/doi.org\/10.1201\/9781420082678","key":"7334_CR39","DOI":"10.1201\/9781420082678"},{"key":"7334_CR40","volume-title":"Reinforcement learning: an introduction","author":"RS Sutton","year":"2018","unstructured":"Sutton RS, Barto AG (2018) Reinforcement learning: an introduction. A Bradford Book, Cambridge, MA, USA"},{"key":"7334_CR41","doi-asserted-by":"publisher","first-page":"372","DOI":"10.1007\/978-3-540-89378-3_37","volume-title":"AI 2008: advances in artificial intelligence","author":"P Vamplew","year":"2008","unstructured":"Vamplew P, Yearwood J, Dazeley R, Berry A (2008) On the limitations of scalarisation for multi-objective reinforcement learning of pareto fronts. In: Wobcke W, Zhang M (eds) AI 2008: advances in artificial intelligence. Springer, Berlin Heidelberg, pp 372\u2013378"},{"key":"7334_CR42","doi-asserted-by":"publisher","first-page":"51","DOI":"10.1007\/s10994-010-5232-5","volume":"84","author":"P Vamplew","year":"2011","unstructured":"Vamplew P, Dazeley R, Berry A, Issabekov R, Dekker E (2011) Empirical evaluation methods for multiobjective reinforcement learning algorithms. Mach Learn 84:51\u201380. https:\/\/doi.org\/10.1007\/s10994-010-5232-5","journal-title":"Mach Learn"},{"key":"7334_CR43","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-021-05859-1","author":"P Vamplew","year":"2021","unstructured":"Vamplew P, Foale C, Dazeley R (2021) The impact of environmental stochasticity on value-based multiobjective reinforcement learning. Neural Comput Appl. https:\/\/doi.org\/10.1007\/s00521-021-05859-1","journal-title":"Neural Comput Appl"},{"unstructured":"Vamplew P, Smith BJ, Kallstrom J, Ramos G, Radulescu R, Roijers DM, Hayes CF, Heintz F, Mannion P, Libin PJ, et\u00a0al. (2021) Scalar reward is not enough: a response to silver, singh, precup and sutton. arXiv preprint arXiv:2112.15422","key":"7334_CR44"},{"unstructured":"Wang W, Sebag M (2012) Multi-objective Monte-Carlo tree search. In: Hoi SCH, Buntine W (eds) Proceedings of machine learning research, vol\u00a025, pp 507\u2013522. PMLR, Singapore","key":"7334_CR45"},{"doi-asserted-by":"publisher","unstructured":"Wolfstetter E (1999) Topics in microeconomics: industrial organization, auctions, and incentives. Cambridge University Press, Cambridge. https:\/\/doi.org\/10.1017\/CBO9780511625787","key":"7334_CR46","DOI":"10.1017\/CBO9780511625787"},{"unstructured":"Yahyaa S, Manderick B (2015) Thompson sampling for multi-objective multi-armed bandits problem. In: Proceedings, p\u00a047. Presses universitaires de Louvain, Elsevier","key":"7334_CR47"},{"unstructured":"Yang R, Sun X, Narasimhan K (2019) A generalized algorithm for multi-objective reinforcement learning and policy adaptation. In: Wallach H, Larochelle H, Beygelzimer A, d\u2019 Alch\u00e9-Buc F, Fox E, Garnett R (eds) Advances in neural information processing systems, vol.\u00a032. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper\/2019\/file\/4a46fbfca3f1465a27b210f4bdfe6ab3-Paper.pdf","key":"7334_CR48"},{"unstructured":"Zintgraf LM, Kanters TV, Roijers DM, Oliehoek F, Beau P (2015) Quality assessment of morl algorithms: a utility-based approach. In: Benelearn 2015: proceedings of the 24th annual machine learning conference of Belgium and the Netherlands","key":"7334_CR49"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-022-07334-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-022-07334-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-022-07334-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,27]],"date-time":"2025-06-27T08:27:41Z","timestamp":1751012861000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-022-07334-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,5]]},"references-count":49,"journal-issue":{"issue":"19","published-print":{"date-parts":[[2025,7]]}},"alternative-id":["7334"],"URL":"https:\/\/doi.org\/10.1007\/s00521-022-07334-x","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"type":"print","value":"0941-0643"},{"type":"electronic","value":"1433-3058"}],"subject":[],"published":{"date-parts":[[2022,7,5]]},"assertion":[{"value":"15 May 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 April 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 July 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declaration"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}