{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T17:20:06Z","timestamp":1763400006263,"version":"3.37.3"},"reference-count":20,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2021,11,2]],"date-time":"2021-11-02T00:00:00Z","timestamp":1635811200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,11,2]],"date-time":"2021-11-02T00:00:00Z","timestamp":1635811200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100011033","name":"Agencia Estatal de Investigaci\u00f3n","doi-asserted-by":"publisher","award":["TIN2016-80774-R (AEI\/FEDER, UE)"],"award-info":[{"award-number":["TIN2016-80774-R (AEI\/FEDER, UE)"]}],"id":[{"id":"10.13039\/501100011033","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Glob Optim"],"published-print":{"date-parts":[[2022,3]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper addresses the problem of approximating the set of all solutions for Multi-objective Markov Decision Processes. We show that in the vast majority of interesting cases, the number of solutions is exponential or even infinite. In order to overcome this difficulty we propose to approximate the set of all solutions by means of a limited precision approach based on White\u2019s multi-objective value-iteration dynamic programming algorithm. We prove that the number of calculated solutions is tractable and show experimentally that the solutions obtained are a good approximation of the true Pareto front.<\/jats:p>","DOI":"10.1007\/s10898-021-01096-x","type":"journal-article","created":{"date-parts":[[2021,11,2]],"date-time":"2021-11-02T01:02:28Z","timestamp":1635814948000},"page":"595-614","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Multi-objective dynamic programming with limited precision"],"prefix":"10.1007","volume":"82","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8141-498X","authenticated-orcid":false,"given":"L.","family":"Mandow","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7743-323X","authenticated-orcid":false,"given":"J. L.","family":"Perez-de-la-Cruz","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"N.","family":"Pozas","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,11,2]]},"reference":[{"key":"1096_CR1","doi-asserted-by":"publisher","first-page":"591","DOI":"10.1057\/jors.1980.114","volume":"31","author":"H Daellenbach","year":"1980","unstructured":"Daellenbach, H., De Kluyver, C.: Note on multiple objective dynamic programming. J. Oper. Res. Soc. 31, 591\u2013594 (1980)","journal-title":"J. Oper. Res. Soc."},{"key":"1096_CR2","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.neucom.2017.06.020","volume":"263","author":"M Drugan","year":"2017","unstructured":"Drugan, M., Wiering, M., Vamplew, P., Chetty, M.: Special issue on multi-objective reinforcement learning. Neurocomputing 263, 1\u20132 (2017)","journal-title":"Neurocomputing"},{"key":"1096_CR3","doi-asserted-by":"publisher","first-page":"228","DOI":"10.1016\/j.swevo.2018.03.011","volume":"44","author":"MM Drugan","year":"2019","unstructured":"Drugan, M.M.: Reinforcement learning versus evolutionary computation: a survey on hybrid algorithms. Swarm Evol. Comput. 44, 228\u2013246 (2019)","journal-title":"Swarm Evol. Comput."},{"key":"1096_CR4","doi-asserted-by":"crossref","unstructured":"Etessami, K., Kwiatkowska, M.Z., Vardi, M.Y., Yannakakis, M.: Multi-objective model checking of markov decision processes. Log Methods Comput. Sci. 4(4) (2008)","DOI":"10.2168\/LMCS-4(4:8)2008"},{"key":"1096_CR5","unstructured":"Forejt, V., Kwiatkowska, M.Z., Norman, G., Parker, D., Qu, H.: Quantitative multi-objective verification for probabilistic systems. In: Abdulla PA, Leino KRM (eds) Tools and Algorithms for the Construction and Analysis of Systems - 17th International Conference, TACAS 2011, Held as Part of the Joint European Conferences on Theory and Practice of Software, ETAPS 2011, Saarbr\u00fccken, Germany, March 26-April 3, 2011. Proceedings, Springer, Lecture Notes in Computer Science, vol 6605, pp 112\u2013127, (2011)"},{"key":"1096_CR6","doi-asserted-by":"crossref","unstructured":"Hansen, P.: Bicriterion path problems. In: Lecture Notes in Economics and Mathematical Systems, vol. 177, pp. 109\u2013127. Springer, Berlin (1980)","DOI":"10.1007\/978-3-642-48782-8_9"},{"key":"1096_CR7","unstructured":"Lizotte, D.J., Bowling, M., Murphy, S.A.: Efficient reinforcement learning with multiple reward functions for randomized controlled trial analysis. In: Proceedings of the 27th International Conference on Machine Learning, pp. 695\u2013702 (2010)"},{"key":"1096_CR8","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s10898-019-00829-3","volume":"75","author":"K Miettinen","year":"2019","unstructured":"Miettinen, K., Ruiz, F.: Preface on the special issue global optimization with multiple criteria: theory, methods and applications. J. Glob. Optim. 75, 1\u20132 (2019)","journal-title":"J. Glob. Optim."},{"key":"1096_CR9","first-page":"969","volume":"2010","author":"P Perny","year":"2010","unstructured":"Perny, P., Weng, P.: On finding compromise solutions in multiobjective markov decision processes. ECAI 2010, 969\u2013970 (2010)","journal-title":"ECAI"},{"key":"1096_CR10","doi-asserted-by":"publisher","DOI":"10.1002\/9780470316887","volume-title":"Markov Decision Processes: Discrete Stochastic Dynamic Programming","author":"ML Puterman","year":"1994","unstructured":"Puterman, M.L.: Markov Decision Processes: Discrete Stochastic Dynamic Programming. Wiley, New Jersey (1994)"},{"issue":"1","key":"1096_CR11","first-page":"1","volume":"11","author":"DM Roijers","year":"2017","unstructured":"Roijers, D.M., Whiteson, S.: Multi-objective decision making. Synth. Lect. Artif. Intell. Mach. Learn. 11(1), 1\u2013129 (2017)","journal-title":"Synth. Lect. Artif. Intell. Mach. Learn."},{"key":"1096_CR12","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1613\/jair.3987","volume":"48","author":"DM Roijers","year":"2013","unstructured":"Roijers, D.M., Vamplew, P., Whiteson, S., Dazeley, R.: A survey of multi-objective sequential decision-making. J. Artif. Intell. Res. (JAIR) 48, 67\u2013113 (2013)","journal-title":"J. Artif. Intell. Res. (JAIR)"},{"key":"1096_CR13","doi-asserted-by":"publisher","first-page":"15","DOI":"10.1016\/j.neucom.2016.10.100","volume":"263","author":"M Ruiz-Montiel","year":"2017","unstructured":"Ruiz-Montiel, M., Mandow, L., P\u00e9rez-de-la-Cruz, J.: A temporal difference method for multi-objective reinforcement learning. Neurocomputing 263, 15\u201325 (2017)","journal-title":"Neurocomputing"},{"key":"1096_CR14","volume-title":"Reinforcement learning: an introduction","author":"R Sutton","year":"2018","unstructured":"Sutton, R., Barto, A.: Reinforcement learning: an introduction, 2nd edn. The MIT Press, Cambridge (2018)","edition":"2"},{"key":"1096_CR15","doi-asserted-by":"publisher","first-page":"372","DOI":"10.1007\/978-3-540-89378-3_37","volume-title":"AI 2008: Advances in Artificial Intelligence","author":"P Vamplew","year":"2008","unstructured":"Vamplew, P., Yearwood, J., Dazeley, R., Berry, A.: On the limitations of scalarisation for multi-objective reinforcement learning of pareto fronts. In: Wobcke, W., Zhang, M. (eds.) AI 2008: Advances in Artificial Intelligence, pp. 372\u2013378. Springer, Berlin Heidelberg (2008)"},{"key":"1096_CR16","doi-asserted-by":"crossref","unstructured":"Vamplew, P., Dazeley, R., Barker, E., Kelarev, A.: Constructing stochastic mixture policies for episodic multiobjective reinforcement learning tasks. In: AI 2009: Proceedings of the Twenty-Second Australasian Joint Conference on Artificial Intelligence, pp. 340\u2013349 (2009)","DOI":"10.1007\/978-3-642-10439-8_35"},{"key":"1096_CR17","first-page":"3663","volume":"15","author":"K Van Moffaert","year":"2014","unstructured":"Van Moffaert, K., Now\u00e9, A.: Multi-objective reinforcement learning using sets of pareto dominating policies. J. Mach. Learn. Res. 15, 3663\u20133692 (2014)","journal-title":"J. Mach. Learn. Res."},{"key":"1096_CR18","doi-asserted-by":"publisher","first-page":"639","DOI":"10.1016\/0022-247X(82)90122-6","volume":"89","author":"DJ White","year":"1982","unstructured":"White, D.J.: Multi-objective infinite-horizon discounted Markov decision processes. J. Math. Anal. Appl. 89, 639\u2013647 (1982)","journal-title":"J. Math. Anal. Appl."},{"key":"1096_CR19","doi-asserted-by":"crossref","unstructured":"Wray, K.H., Zilberstein, S., Mouaddib, A.: Multi-objective mdps with conditional lexicographic reward preferences. In: Twenty-Ninth AAAI Conference on Artificial Intelligence (2015)","DOI":"10.1609\/aaai.v29i1.9647"},{"issue":"2","key":"1096_CR20","doi-asserted-by":"publisher","first-page":"117","DOI":"10.1109\/TEVC.2003.810758","volume":"7","author":"E Zitzler","year":"2003","unstructured":"Zitzler, E., Thiele, L., Laumanns, M., Fonseca, C.M., da Fonseca, V.G.: Performance assessment of multiobjective optimizers: an analysis and review. IEEE Trans. Evol. Comput. 7(2), 117\u2013132 (2003)","journal-title":"IEEE Trans. Evol. Comput."}],"container-title":["Journal of Global Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10898-021-01096-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10898-021-01096-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10898-021-01096-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,14]],"date-time":"2023-01-14T08:26:18Z","timestamp":1673684778000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10898-021-01096-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,11,2]]},"references-count":20,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,3]]}},"alternative-id":["1096"],"URL":"https:\/\/doi.org\/10.1007\/s10898-021-01096-x","relation":{},"ISSN":["0925-5001","1573-2916"],"issn-type":[{"type":"print","value":"0925-5001"},{"type":"electronic","value":"1573-2916"}],"subject":[],"published":{"date-parts":[[2021,11,2]]},"assertion":[{"value":"30 November 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 September 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 November 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}