{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T07:10:42Z","timestamp":1780384242665,"version":"3.54.1"},"reference-count":78,"publisher":"Wiley","issue":"1","license":[{"start":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T00:00:00Z","timestamp":1777420800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T00:00:00Z","timestamp":1777420800000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["510629371"],"award-info":[{"award-number":["510629371"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Networks"],"published-print":{"date-parts":[[2026,7]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>With rising customer expectations and increasing computational potential, many transport, manufacturing, and production operations face real\u2010time decision making in stochastic dynamic environments. Decision makers must find and adapt complex plans that are effective now but also flexible with respect to future developments. The challenges of searching a high\u2010dimensional constrained decision space for effective and flexible decisions are reflected in the three parts of the Bellman equation: the reward function, the value function, and the decision space. In the literature, reinforcement learning (RL) has shown potential to quickly evaluate the reward\u2010 and value function for a limited number of decisions but struggles to search a constrained decision space present in most planning problems. The question of how to combine the thorough search of the complex decision space with RL\u2010evaluation techniques is still open. We propose two RL\u2010based solution methods and detail a third one to search for and evaluate decisions in an integrated manner. Each method is inspired by one component of the Bellman equation. The first two methods dynamically shape the reward function or decision space to encourage effective and flexible decisions or prohibit inflexible decisions. The third method models the Bellman equation as a mixed\u2010integer linear programming formulation in which the value function is approximated by a neural network. We compare our proposed solution methods in a structured analysis for carefully designed problem classes. We demonstrate the effectiveness of our methods compared to prominent benchmark methods and highlight how the methods' performances depend not only on the problem classes but also on the instances' parameterizations.<\/jats:p>","DOI":"10.1002\/net.70039","type":"journal-article","created":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T12:53:01Z","timestamp":1777467181000},"page":"138-157","update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Shaping Decision Models for Stochastic Dynamic Optimization Problems via Reinforcement Learning"],"prefix":"10.1002","volume":"88","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0298-275X","authenticated-orcid":false,"given":"Florentin D.","family":"Hildebrandt","sequence":"first","affiliation":[{"name":"Chair of Management Science Otto\u2010von\u2010Guericke\u2010Universit\u00e4t  Magdeburg Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alexander","family":"Bode","sequence":"additional","affiliation":[{"name":"Institute for Decision Support Technische Universit\u00e4t Braunschweig  Braunschweig Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marlin W.","family":"Ulmer","sequence":"additional","affiliation":[{"name":"Chair of Management Science Otto\u2010von\u2010Guericke\u2010Universit\u00e4t  Magdeburg Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dirk C.","family":"Mattfeld","sequence":"additional","affiliation":[{"name":"Institute for Decision Support Technische Universit\u00e4t Braunschweig  Braunschweig Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2026,4,29]]},"reference":[{"key":"e_1_2_14_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2021.07.014"},{"key":"e_1_2_14_3_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2023.02.013"},{"key":"e_1_2_14_4_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2022.08.049"},{"key":"e_1_2_14_5_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2023.08.003"},{"key":"e_1_2_14_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2020.01.033"},{"key":"e_1_2_14_7_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.2020.1000"},{"key":"e_1_2_14_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2018.05.032"},{"key":"e_1_2_14_9_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature16961"},{"key":"e_1_2_14_10_1","unstructured":"D.Silver T.Hubert J.Schrittwieser et al. \u201cMastering Chess and Shogi by Self\u2010Play With a General Reinforcement Learning Algorithm \u201d(2017) arXiv preprint arXiv:1712.01815."},{"key":"e_1_2_14_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cor.2022.106071"},{"key":"e_1_2_14_12_1","doi-asserted-by":"publisher","DOI":"10.1002\/9781119815068"},{"key":"e_1_2_14_13_1","first-page":"15931","article-title":"Learning to Utilize Shaping Rewards: A New Approach of Reward Shaping","volume":"33","author":"Hu Y.","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_14_14_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cor.2009.01.014"},{"key":"e_1_2_14_15_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.1050.0133"},{"key":"e_1_2_14_16_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.33.4.381"},{"key":"e_1_2_14_17_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.34.4.426.12325"},{"key":"e_1_2_14_18_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.1030.0068"},{"key":"e_1_2_14_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/3468.668962"},{"key":"e_1_2_14_20_1","doi-asserted-by":"publisher","DOI":"10.1287\/opre.1040.0170"},{"key":"e_1_2_14_21_1","first-page":"593","volume-title":"Proceedings of the Annual Allerton Conference on Communication Control and Computing, the University, 1998","author":"Keslassy I.","year":"2001"},{"key":"e_1_2_14_22_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.trb.2003.09.002"},{"key":"e_1_2_14_23_1","doi-asserted-by":"publisher","DOI":"10.3138\/infor.46.3.165"},{"key":"e_1_2_14_24_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.1060.0183"},{"key":"e_1_2_14_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2022.02.050"},{"key":"e_1_2_14_26_1","doi-asserted-by":"crossref","unstructured":"C.Riley P.Van Hentenryck andE.Yuan \u201cReal\u2010Time Dispatching of Large\u2010Scale Ride\u2010Sharing Systems: Integrating Optimization Machine Learning and Model Predictive Control \u201d(2020) arXiv preprint arXiv:2003.10942.","DOI":"10.24963\/ijcai.2020\/609"},{"key":"e_1_2_14_27_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2020.03.037"},{"key":"e_1_2_14_28_1","doi-asserted-by":"publisher","DOI":"10.1287\/opre.1040.0124"},{"key":"e_1_2_14_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2016.09.040"},{"key":"e_1_2_14_30_1","doi-asserted-by":"publisher","DOI":"10.1287\/opre.49.5.796.10608"},{"key":"e_1_2_14_31_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.2017.0767"},{"key":"e_1_2_14_32_1","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton R. S.","year":"2018"},{"key":"e_1_2_14_33_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.2018.0840"},{"key":"e_1_2_14_34_1","volume-title":"A Comparative Analysis of Neural Networks in Anticipatory Transportation Planning","author":"Akkerman F.","year":"2022"},{"key":"e_1_2_14_35_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.tre.2021.102496"},{"key":"e_1_2_14_36_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijpe.2017.10.028"},{"key":"e_1_2_14_37_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2021.06.021"},{"key":"e_1_2_14_38_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2022.12.009"},{"key":"e_1_2_14_39_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.trpro.2019.05.006"},{"key":"e_1_2_14_40_1","first-page":"23609","article-title":"A Hierarchical Reinforcement Learning Based Optimization Framework for Large\u2010Scale Dynamic Pickup and Delivery Problems","volume":"34","author":"Ma Y.","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_14_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/ETFA.2018.8502508"},{"key":"e_1_2_14_42_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12599-019-00582-7"},{"key":"e_1_2_14_43_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.2019.0958"},{"key":"e_1_2_14_44_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.2016.0719"},{"key":"e_1_2_14_45_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2018.02.038"},{"key":"e_1_2_14_46_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2019.04.029"},{"key":"e_1_2_14_47_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.trc.2023.104401"},{"key":"e_1_2_14_48_1","doi-asserted-by":"publisher","DOI":"10.65109\/KWHR3408"},{"key":"e_1_2_14_49_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.simpat.2004.12.003"},{"key":"e_1_2_14_50_1","first-page":"507","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Shah S.","year":"2020"},{"key":"e_1_2_14_51_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.trc.2020.102861"},{"key":"e_1_2_14_52_1","first-page":"609","article-title":"Reinforcement Learning With Combinatorial Actions: An Application to Vehicle Routing","volume":"33","author":"Delarue A.","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_14_53_1","first-page":"17295","volume-title":"Proceedings of the 39th International Conference on Machine Learning","author":"Papalexopoulos T. P.","year":"2022"},{"key":"e_1_2_14_54_1","doi-asserted-by":"publisher","DOI":"10.1287\/moor.1060.0208"},{"key":"e_1_2_14_55_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.2022.1164"},{"key":"e_1_2_14_56_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.tre.2016.09.002"},{"key":"e_1_2_14_57_1","doi-asserted-by":"crossref","unstructured":"M.BiggsandG.Perakis \u201cDynamic Routing With Tree Based Value Function Approximations \u201d(2020) available at SSRN 3680162.","DOI":"10.2139\/ssrn.3680162"},{"key":"e_1_2_14_58_1","doi-asserted-by":"publisher","DOI":"10.1287\/msom.2022.0617"},{"key":"e_1_2_14_59_1","unstructured":"W.vanHeeswijkandH.La Poutr\u00e9 \u201cApproximate Dynamic Programming With Neural Networks in Linear Discrete Action Spaces \u201d(2019) arXiv preprint arXiv:1902.09855."},{"key":"e_1_2_14_60_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejtl.2023.100105"},{"key":"e_1_2_14_61_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.2024.0510"},{"key":"e_1_2_14_62_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.2022.0434"},{"key":"e_1_2_14_63_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2023.05.017"},{"key":"e_1_2_14_64_1","doi-asserted-by":"publisher","DOI":"10.1287\/opre.46.1.17"},{"key":"e_1_2_14_65_1","doi-asserted-by":"publisher","DOI":"10.1287\/opre.49.1.26.11185"},{"key":"e_1_2_14_66_1","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.48.4.550.208"},{"key":"e_1_2_14_67_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2014.04.040"},{"key":"e_1_2_14_68_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2016.05.031"},{"key":"e_1_2_14_69_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cor.2021.105357"},{"key":"e_1_2_14_70_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.1040.0105"},{"key":"e_1_2_14_71_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2013.08.028"},{"key":"e_1_2_14_72_1","doi-asserted-by":"publisher","DOI":"10.1287\/trsc.2019.0927"},{"key":"e_1_2_14_73_1","unstructured":"J.Schulman F.Wolski P.Dhariwal A.Radford andO.Klimov \u201cProximal Policy Optimization Algorithms \u201d(2017) arXiv preprint arXiv:1707.06347."},{"key":"e_1_2_14_74_1","volume-title":"3rd International Conference on Learning Representations","author":"Kingma D.","year":"2015"},{"key":"e_1_2_14_75_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1009642405419"},{"key":"e_1_2_14_76_1","doi-asserted-by":"publisher","DOI":"10.1287\/opre.5.2.266"},{"key":"e_1_2_14_77_1","article-title":"Algorithms for Hyper\u2010Parameter Optimization","volume":"24","author":"Bergstra J.","year":"2011","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_14_78_1","first-page":"486","volume-title":"Proceedings of the 2nd Conference on Learning for Dynamics and Control","author":"Fan J.","year":"2020"},{"key":"e_1_2_14_79_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-020-01474-5"}],"container-title":["Networks"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/net.70039","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1002\/net.70039","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/net.70039","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T06:58:05Z","timestamp":1780383485000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/net.70039"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,29]]},"references-count":78,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,7]]}},"alternative-id":["10.1002\/net.70039"],"URL":"https:\/\/doi.org\/10.1002\/net.70039","archive":["Portico"],"relation":{},"ISSN":["0028-3045","1097-0037"],"issn-type":[{"value":"0028-3045","type":"print"},{"value":"1097-0037","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,29]]},"assertion":[{"value":"2024-07-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-26","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-29","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}