{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,14]],"date-time":"2026-02-14T03:25:16Z","timestamp":1771039516223,"version":"3.50.1"},"reference-count":85,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2024,2,15]],"date-time":"2024-02-15T00:00:00Z","timestamp":1707955200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,2,15]],"date-time":"2024-02-15T00:00:00Z","timestamp":1707955200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100000084","name":"Directorate for Engineering","doi-asserted-by":"publisher","award":["CBET-1554018"],"award-info":[{"award-number":["CBET-1554018"]}],"id":[{"id":"10.13039\/100000084","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Glob Optim"],"published-print":{"date-parts":[[2024,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In this paper, we address the difficulty of solving large-scale multi-dimensional knapsack instances (MKP), presenting a novel deep reinforcement learning (DRL) framework. In this DRL framework, we train different agents compatible with a discrete action space for sequential decision-making while still satisfying any resource constraint of the MKP. This novel framework incorporates the decision variable values in the 2D DRL where the agent is responsible for assigning a value of 1 or 0 to each of the variables. To the best of our knowledge, this is the first DRL model of its kind in which a 2D environment is formulated, and an element of the DRL solution matrix represents an item of the MKP. Our framework is configured to solve MKP instances of different dimensions and distributions. We propose a K-means approach to obtain an initial feasible solution that is used to train the DRL agent. We train four different agents in our framework and present the results comparing each of them with the CPLEX commercial solver. The results show that our agents can learn and generalize over instances with different sizes and distributions. Our DRL framework shows that it can solve medium-sized instances at least 45 times faster in CPU solution time and at least 10 times faster for large instances, with a maximum solution gap of 0.28% compared to the performance of CPLEX. Furthermore, at least 95% of the items are predicted in line with the CPLEX solution. Computations with DRL also provide a better optimality gap with respect to state-of-the-art approaches.<\/jats:p>","DOI":"10.1007\/s10898-024-01364-6","type":"journal-article","created":{"date-parts":[[2024,2,15]],"date-time":"2024-02-15T09:03:06Z","timestamp":1707987786000},"page":"655-685","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["A K-means Supported Reinforcement Learning Framework to Multi-dimensional Knapsack"],"prefix":"10.1007","volume":"89","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0433-692X","authenticated-orcid":false,"given":"Sabah","family":"Bushaj","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8928-2638","authenticated-orcid":false,"given":"\u0130. Esra","family":"B\u00fcy\u00fcktahtak\u0131n","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,2,15]]},"reference":[{"key":"1364_CR1","unstructured":"Afshar, R.R., Zhang, Y., Firat, M., Kaymak, U.: A state aggregation approach for solving knapsack problem with deep reinforcement learning. In: Asian Conference on Machine Learning, pp. 81\u201396. PMLR (2020)"},{"issue":"1","key":"1364_CR2","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1007\/s10479-006-0150-4","volume":"150","author":"Y Ak\u00e7ay","year":"2007","unstructured":"Ak\u00e7ay, Y., Li, H., Xu, S.H.: Greedy algorithm for the general multidimensional knapsack problem. Ann. Oper. Res. 150(1), 17\u201329 (2007)","journal-title":"Ann. Oper. Res."},{"issue":"1","key":"1364_CR3","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1287\/mnsc.26.1.86","volume":"26","author":"E Balas","year":"1980","unstructured":"Balas, E., Martin, C.H.: Pivot and complement-a heuristic for 0\u20131 programming. Manag. Sci. 26(1), 86\u201396 (1980)","journal-title":"Manag. Sci."},{"issue":"1","key":"1364_CR4","doi-asserted-by":"crossref","first-page":"63","DOI":"10.1016\/j.ejor.2006.02.058","volume":"186","author":"S Balev","year":"2008","unstructured":"Balev, S., Yanev, N., Fr\u00e9ville, A., Andonov, R.: A dynamic programming based reduction procedure for the multidimensional 0\u20131 knapsack problem. Eur. J. Oper. Res. 186(1), 63\u201376 (2008)","journal-title":"Eur. J. Oper. Res."},{"key":"1364_CR5","doi-asserted-by":"crossref","unstructured":"Barrett, T., Clements, W., Foerster, J., Lvovsky, A.: Exploratory combinatorial optimization with reinforcement learning. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34(04), pp. 3243\u20133250 (2020)","DOI":"10.1609\/aaai.v34i04.5723"},{"key":"1364_CR6","unstructured":"Bello, I., Pham, H., Le, Q.V., Norouzi, M., Bengio, S.: Neural combinatorial optimization with reinforcement learning. CoRR arXiv:1611.09940 (2016)"},{"issue":"4","key":"1364_CR7","doi-asserted-by":"crossref","first-page":"550","DOI":"10.1287\/mnsc.48.4.550.208","volume":"48","author":"D Bertsimas","year":"2002","unstructured":"Bertsimas, D., Demir, R.: An approximate dynamic programming approach to multidimensional knapsack problems. Manag. Sci. 48(4), 550\u2013565 (2002)","journal-title":"Manag. Sci."},{"issue":"3","key":"1364_CR8","doi-asserted-by":"crossref","first-page":"658","DOI":"10.1016\/j.ejor.2007.06.068","volume":"199","author":"V Boyer","year":"2009","unstructured":"Boyer, V., Elkihel, M., El Baz, D.: Heuristics for the 0\u20131 multidimensional knapsack problem. Eur. J. Oper. Res. 199(3), 658\u2013664 (2009)","journal-title":"Eur. J. Oper. Res."},{"issue":"3","key":"1364_CR9","doi-asserted-by":"crossref","first-page":"1094","DOI":"10.1016\/j.ejor.2021.08.035","volume":"299","author":"S Bushaj","year":"2022","unstructured":"Bushaj, S., B\u00fcy\u00fcktahtak\u0131n, \u0130E., Haight, R.G.: Risk-averse multi-stage stochastic optimization for surveillance and operations planning of a forest insect infestation. Eur. J. Oper. Res. 299(3), 1094\u20131110 (2022)","journal-title":"Eur. J. Oper. Res."},{"issue":"1","key":"1364_CR10","doi-asserted-by":"crossref","DOI":"10.1111\/nrm.12267","volume":"34","author":"S Bushaj","year":"2020","unstructured":"Bushaj, S., B\u00fcy\u00fcktahtak\u0131n, \u0130E., Yemshanov, D., Haight, R.G.: Optimizing surveillance and management of emerald ash borer in urban environments. Nat. Resour. Model. 34(1), e12267 (2020)","journal-title":"Nat. Resour. Model."},{"issue":"1","key":"1364_CR11","doi-asserted-by":"crossref","first-page":"245","DOI":"10.1007\/s10479-022-04926-7","volume":"328","author":"S Bushaj","year":"2023","unstructured":"Bushaj, S., Yin, X., Beqiri, A., Andrews, D., B\u00fcy\u00fcktahtak\u0131n, \u0130E.: A simulation-deep reinforcement learning (sirl) approach for epidemic control optimization. Ann. Oper. Res. 328(1), 245\u2013277 (2023)","journal-title":"Ann. Oper. Res."},{"issue":"1","key":"1364_CR12","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s10479-021-04388-3","volume":"309","author":"\u0130E B\u00fcy\u00fcktahtak\u0131n","year":"2022","unstructured":"B\u00fcy\u00fcktahtak\u0131n, \u0130E.: Stage-t scenario dominance for risk-averse multi-stage stochastic mixed-integer programs. Ann. Oper. Res. 309(1), 1\u201335 (2022)","journal-title":"Ann. Oper. Res."},{"key":"1364_CR13","doi-asserted-by":"crossref","DOI":"10.1016\/j.cor.2023.106149","volume":"153","author":"\u0130E B\u00fcy\u00fcktahtak\u0131n","year":"2023","unstructured":"B\u00fcy\u00fcktahtak\u0131n, \u0130E.: Scenario-dominance to multi-stage stochastic lot-sizing and knapsack problems. Comput. Oper. Res. 153, 106149 (2023)","journal-title":"Comput. Oper. Res."},{"issue":"2","key":"1364_CR14","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1016\/S0377-2217(99)00261-1","volume":"123","author":"A Caprara","year":"2000","unstructured":"Caprara, A., Kellerer, H., Pferschy, U., Pisinger, D.: Approximation algorithms for knapsack problems with cardinality constraints. Eur. J. Oper. Res. 123(2), 333\u2013345 (2000)","journal-title":"Eur. J. Oper. Res."},{"key":"1364_CR15","unstructured":"Chen, W., Xu, Y., Wu, X.: Deep reinforcement learning for multi-resource multi-machine job scheduling. arXiv preprint arXiv:1711.07440 (2017)"},{"issue":"1","key":"1364_CR16","doi-asserted-by":"crossref","first-page":"63","DOI":"10.1023\/A:1009642405419","volume":"4","author":"PC Chu","year":"1998","unstructured":"Chu, P.C., Beasley, J.E.: A genetic algorithm for the multidimensional knapsack problem. J. Heurist. 4(1), 63\u201386 (1998)","journal-title":"J. Heurist."},{"key":"1364_CR17","unstructured":"Dai, H., Dai, B., Song, L.: Discriminative embeddings of latent variable models for structured data. CoRR arXiv:1603.05629 (2016)"},{"key":"1364_CR18","unstructured":"Dai, H., Khalil, E.B., Zhang, Y., Dilkina, B., Song, L.: Learning combinatorial optimization algorithms over graphs. CoRR arXiv:1704.01665 (2017)"},{"key":"1364_CR19","unstructured":"Delarue, A., Anderson, R., Tjandraatmadja, C.: Reinforcement learning with combinatorial actions: an application to vehicle routing. arXiv preprint arXiv:2010.12001 (2020)"},{"issue":"4","key":"1364_CR20","doi-asserted-by":"crossref","first-page":"515","DOI":"10.1287\/moor.7.4.515","volume":"7","author":"G Dobson","year":"1982","unstructured":"Dobson, G.: Worst-case analysis of greedy heuristics for integer programming with nonnegative data. Math. Oper. Res. 7(4), 515\u2013531 (1982)","journal-title":"Math. Oper. Res."},{"key":"1364_CR21","doi-asserted-by":"crossref","unstructured":"Etheve, M., Al\u00e8s, Z., Bissuel, C., Juan, O., Kedad-Sidhoum, S.: Reinforcement learning for variable selection in a branch and bound algorithm. In: International Conference on Integration of Constraint Programming, Artificial Intelligence, and Operations Research, pp. 176\u2013185. Springer (2020)","DOI":"10.1007\/978-3-030-58942-4_12"},{"key":"1364_CR22","unstructured":"Eysenbach, B., Gupta, A., Ibarz, J., Levine, S.: Diversity is all you need: Learning skills without a reward function. arXiv preprint arXiv:1802.06070 (2018)"},{"issue":"4","key":"1364_CR23","doi-asserted-by":"crossref","first-page":"613","DOI":"10.1002\/nav.3800320408","volume":"32","author":"GE Fox","year":"1985","unstructured":"Fox, G.E., Scudder, G.D.: A heuristic with tie breaking for certain 0\u20131 integer programming models. Nav. Res. Logist. Q. 32(4), 613\u2013623 (1985)","journal-title":"Nav. Res. Logist. Q."},{"issue":"3","key":"1364_CR24","doi-asserted-by":"crossref","first-page":"413","DOI":"10.1016\/0377-2217(93)90197-U","volume":"68","author":"A Fr\u00e9ville","year":"1993","unstructured":"Fr\u00e9ville, A., Plateau, G.: An exact search for the solution of the surrogate dual of the 0\u20131 bidimensional knapsack problem. Eur. J. Oper. Res. 68(3), 413\u2013421 (1993)","journal-title":"Eur. J. Oper. Res."},{"issue":"1","key":"1364_CR25","doi-asserted-by":"crossref","first-page":"100","DOI":"10.1016\/0377-2217(84)90053-5","volume":"15","author":"A Frieze","year":"1984","unstructured":"Frieze, A., Clarke, M.: Approximation algorithms for the m-dimensional 0\u20131 knapsack problem: Worst-case and probabilistic analyses. Eur. J. Oper. Res. 15(1), 100\u2013109 (1984)","journal-title":"Eur. J. Oper. Res."},{"issue":"4","key":"1364_CR26","doi-asserted-by":"crossref","first-page":"330","DOI":"10.1504\/IJMHEUR.2020.111600","volume":"7","author":"D Gaspar","year":"2020","unstructured":"Gaspar, D., Lu, Y., Song, M.S., Vasko, F.J.: Simple population-based metaheuristics for the multiple demand multiple-choice multidimensional knapsack problem. Int. J. Metaheurist. 7(4), 330\u2013351 (2020)","journal-title":"Int. J. Metaheurist."},{"issue":"1","key":"1364_CR27","doi-asserted-by":"crossref","first-page":"78","DOI":"10.1007\/BF02591863","volume":"31","author":"B Gavish","year":"1985","unstructured":"Gavish, B., Pirkul, H.: Efficient algorithms for solving multiconstraint zero-one knapsack problems to optimality. Math. Program. 31(1), 78\u2013105 (1985)","journal-title":"Math. Program."},{"issue":"7","key":"1364_CR28","doi-asserted-by":"crossref","first-page":"583","DOI":"10.1109\/TC.1986.1676799","volume":"35","author":"B Gavish","year":"1986","unstructured":"Gavish, B., Pirkul, H.: Computer and database location in distributed computer systems. IEEE Trans. Comput. 35(7), 583\u2013590 (1986)","journal-title":"IEEE Trans. Comput."},{"key":"1364_CR29","doi-asserted-by":"crossref","unstructured":"Glover, F., Kochenberger, G.A.: Critical event Tabu search for multidimensional knapsack problems. In: Meta-heuristics, pp. 407\u2013427. Springer (1996)","DOI":"10.1007\/978-1-4613-1361-8_25"},{"key":"1364_CR30","volume-title":"Deep Learning","author":"I Goodfellow","year":"2016","unstructured":"Goodfellow, I., Bengio, Y., Courville, A., Bengio, Y.: Deep Learning, vol. 1. MIT Press, Cambridge (2016)"},{"key":"1364_CR31","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.neucom.2019.06.111","volume":"390","author":"S Gu","year":"2020","unstructured":"Gu, S., Hao, T., Yao, H.: A pointer network based deep learning algorithm for unconstrained binary quadratic programming problem. Neurocomputing 390, 1\u201311 (2020)","journal-title":"Neurocomputing"},{"issue":"2\u20133","key":"1364_CR32","doi-asserted-by":"crossref","first-page":"659","DOI":"10.1016\/S0377-2217(97)00296-8","volume":"106","author":"S Hanafi","year":"1998","unstructured":"Hanafi, S., Freville, A.: An efficient tabu search approach for the 0\u20131 multidimensional knapsack problem. Eur. J. Oper. Res. 106(2\u20133), 659\u2013675 (1998)","journal-title":"Eur. J. Oper. Res."},{"key":"1364_CR33","doi-asserted-by":"crossref","unstructured":"Haul, C., Voss, S.: Using surrogate constraints in genetic algorithms for solving multidimensional knapsack problems. In: Advances in Computational and Stochastic Optimization, Logic Programming, and Heuristic Search, pp. 235\u2013251. Springer (1998)","DOI":"10.1007\/978-1-4757-2807-1_9"},{"issue":"4","key":"1364_CR34","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1287\/opre.17.4.600","volume":"17","author":"FS Hillier","year":"1969","unstructured":"Hillier, F.S.: Efficient heuristic procedures for integer linear programming with an interior. Oper. Res. 17(4), 600\u2013637 (1969)","journal-title":"Oper. Res."},{"key":"1364_CR35","unstructured":"Hu, H., Zhang, X., Yan, X., Wang, L., Xu, Y.: Solving a new 3d bin packing problem with deep reinforcement learning method. arXiv preprint arXiv:1708.05930 (2017)"},{"key":"1364_CR36","unstructured":"Hubbs, C.D., Perez, H.D., Sarwar, O., Sahinidis, N.V., Grossmann, I.E., Wassick, J.M.: Or-gym: A reinforcement learning library for operations research problem. arXiv preprint arXiv:2008.06319 (2020)"},{"issue":"2","key":"1364_CR37","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1111\/j.1469-8137.1912.tb05611.x","volume":"11","author":"P Jaccard","year":"1912","unstructured":"Jaccard, P.: The distribution of the flora in the alpine zone. 1. New Phytol. 11(2), 37\u201350 (1912)","journal-title":"New Phytol."},{"key":"1364_CR38","doi-asserted-by":"crossref","unstructured":"Kellerer, H., Pferschy, U., Pisinger, D.: Multidimensional knapsack problems. In: Knapsack Problems, pp. 235\u2013283. Springer (2004)","DOI":"10.1007\/978-3-540-24777-7_9"},{"key":"1364_CR39","unstructured":"Kong, W., Liaw, C., Mehta, A., Sivakumar, D.: A new dog learns old tricks: Rl finds classic optimization algorithms. In: Proceedings of International Conference on Learning Representations, pp. 1\u201325 (2019)"},{"key":"1364_CR40","unstructured":"Kool, W., Van Hoof, H., Welling, M.: Attention, learn to solve routing problems! Proceedings of International Conference on Learning Representations 3499, 3508 (2019)"},{"key":"1364_CR41","first-page":"21188","volume":"33","author":"Y-D Kwon","year":"2020","unstructured":"Kwon, Y.-D., Choo, J., Kim, B., Yoon, I., Gwon, Y., Min, S.: Pomo: Policy optimization with multiple optima for reinforcement learning. Adv. Neural. Inf. Process. Syst. 33, 21188\u201321198 (2020)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"issue":"3","key":"1364_CR42","first-page":"402","volume":"34","author":"JS Lee","year":"1988","unstructured":"Lee, J.S., Guignard, M.: Note-an approximate algorithm for multidimensional zero-one knapsack problems-a parametric approach. Manag. Sci. 34(3), 402\u2013410 (1988)","journal-title":"Manag. Sci."},{"key":"1364_CR43","doi-asserted-by":"crossref","unstructured":"Li, F., Hu, B.: Deepjs: Job scheduling based on deep reinforcement learning in cloud data center. In: Proceedings of the 2019 4th International Conference on Big Data and Computing, pp. 48\u201353 (2019)","DOI":"10.1145\/3335484.3335513"},{"key":"1364_CR44","unstructured":"Li, Y.: Deep reinforcement learning: an overview. arXiv preprint arXiv:1701.07274 (2017)"},{"key":"1364_CR45","doi-asserted-by":"crossref","unstructured":"Liao, H., Zhang, W., Dong, X., Poczos, B., Shimada, K., Burak\u00a0Kara, L.: A deep reinforcement learning approach for global routing. J. Mech. Des. 142(6) (2020)","DOI":"10.1115\/1.4045044"},{"issue":"2","key":"1364_CR46","doi-asserted-by":"crossref","first-page":"129","DOI":"10.1109\/TIT.1982.1056489","volume":"28","author":"S Lloyd","year":"1982","unstructured":"Lloyd, S.: Least squares quantization in pcm. IEEE Trans. Inf. Theory 28(2), 129\u2013137 (1982)","journal-title":"IEEE Trans. Inf. Theory"},{"key":"1364_CR47","doi-asserted-by":"crossref","first-page":"229","DOI":"10.1086\/294081","volume":"28","author":"JH Lorie","year":"1955","unstructured":"Lorie, J.H., Savage, L.J.: Three problems in rationing capital. J. Bus. 28, 229\u2013229 (1955)","journal-title":"J. Bus."},{"issue":"6","key":"1364_CR48","doi-asserted-by":"crossref","first-page":"1101","DOI":"10.1287\/opre.27.6.1101","volume":"27","author":"R Loulou","year":"1979","unstructured":"Loulou, R., Michaelides, E.: New greedy-like heuristics for the multidimensional 0\u20131 knapsack problem. Oper. Res. 27(6), 1101\u20131114 (1979)","journal-title":"Oper. Res."},{"key":"1364_CR49","unstructured":"Ma, Q., Ge, S. He, D., Thaker, D., Drori, I.: Combinatorial optimization by graph pointer networks and hierarchical reinforcement learning. arXiv preprint arXiv:1911.04936 (2019)"},{"issue":"3","key":"1364_CR50","doi-asserted-by":"crossref","first-page":"319","DOI":"10.1016\/0377-2217(84)90286-8","volume":"16","author":"M Magazine","year":"1984","unstructured":"Magazine, M., Oguz, O.: A heuristic algorithm for the multidimensional zero-one knapsack problem. Eur. J. Oper. Res. 16(3), 319\u2013326 (1984)","journal-title":"Eur. J. Oper. Res."},{"issue":"3","key":"1364_CR51","doi-asserted-by":"crossref","first-page":"399","DOI":"10.1287\/ijoc.1110.0460","volume":"24","author":"R Mansini","year":"2012","unstructured":"Mansini, R., Speranza, M.G.: Coral: An exact algorithm for the multidimensional knapsack problem. INFORMS J. Comput. 24(3), 399\u2013415 (2012)","journal-title":"INFORMS J. Comput."},{"key":"1364_CR52","unstructured":"Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., Kavukcuoglu, K.: Asynchronous methods for deep reinforcement learning. In: International Conference on Machine Learning, pp. 1928\u20131937. PMLR (2016)"},{"key":"1364_CR53","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., Riedmiller, M.: Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602 (2013)"},{"key":"1364_CR54","unstructured":"Nazari, M., Oroojlooy, A., Snyder, L., Tak\u00e1c, M.: Reinforcement learning for solving the vehicle routing problem. In: Advances in Neural Information Processing Systems, pp. 9839\u20139849 (2018)"},{"key":"1364_CR55","doi-asserted-by":"crossref","first-page":"224200","DOI":"10.1109\/ACCESS.2020.3044005","volume":"8","author":"HA Nomer","year":"2020","unstructured":"Nomer, H.A., Alnowibet, K.A., Elsayed, A., Mohamed, A.W.: Neural knapsack: a neural network based solver for the knapsack problem. IEEE Access 8, 224200\u2013224210 (2020)","journal-title":"IEEE Access"},{"issue":"2","key":"1364_CR56","doi-asserted-by":"crossref","first-page":"161","DOI":"10.1002\/1520-6750(198704)34:2<161::AID-NAV3220340203>3.0.CO;2-A","volume":"34","author":"H Pirkul","year":"1987","unstructured":"Pirkul, H.: A heuristic solution procedure for the multiconstraint zero-one knapsack problem. Nav. Res. Logist. 34(2), 161\u2013172 (1987)","journal-title":"Nav. Res. Logist."},{"issue":"5","key":"1364_CR57","doi-asserted-by":"crossref","first-page":"758","DOI":"10.1287\/opre.45.5.758","volume":"45","author":"D Pisinger","year":"1997","unstructured":"Pisinger, D.: A minimal algorithm for the 0\u20131 knapsack problem. Oper. Res. 45(5), 758\u2013767 (1997)","journal-title":"Oper. Res."},{"issue":"6","key":"1364_CR58","doi-asserted-by":"crossref","first-page":"1299","DOI":"10.1080\/00207540110118640","volume":"40","author":"P Pontrandolfo","year":"2002","unstructured":"Pontrandolfo, P., Gosavi, A., Okogbaa, O.G., Das, T.K.: Global supply chain management: a reinforcement learning approach. Int. J. Prod. Res. 40(6), 1299\u20131317 (2002)","journal-title":"Int. J. Prod. Res."},{"issue":"268","key":"1364_CR59","first-page":"1","volume":"22","author":"A Raffin","year":"2021","unstructured":"Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., Dormann, N.: Stable-baselines3: Reliable reinforcement learning implementations. J. Mach. Learn. Res. 22(268), 1\u20138 (2021)","journal-title":"J. Mach. Learn. Res."},{"key":"1364_CR60","unstructured":"Schaul, T., Quan, J., Antonoglou, I., Silver, D.: Prioritized experience replay. arXiv preprint arXiv:1511.05952 (2015)"},{"key":"1364_CR61","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., Klimov, O.: Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017)"},{"key":"1364_CR62","doi-asserted-by":"crossref","unstructured":"Senju, S., Toyoda, Y.: An approach to linear programming with 0-1 variables. Manag. Sci. B196\u2013B207 (1968)","DOI":"10.1287\/mnsc.15.4.B196"},{"key":"1364_CR63","doi-asserted-by":"crossref","unstructured":"Shehab, M., Khader, A.T., Alia, M.A.: Enhancing cuckoo search algorithm by using reinforcement learning for constrained engineering optimization problems. In 2019 IEEE Jordan international joint conference on electrical engineering and information technology (JEEIT), pp. 812\u2013816. IEEE (2019)","DOI":"10.1109\/JEEIT.2019.8717366"},{"issue":"6419","key":"1364_CR64","doi-asserted-by":"crossref","first-page":"1140","DOI":"10.1126\/science.aar6404","volume":"362","author":"D Silver","year":"2018","unstructured":"Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al.: A general reinforcement learning algorithm that masters chess, shogi, and go through self-play. Science 362(6419), 1140\u20131144 (2018)","journal-title":"Science"},{"key":"1364_CR65","unstructured":"Tang, Y., Agrawal, S., Faenza, Y.: Reinforcement learning for integer programming: Learning to cut. In International Conference on Machine Learning, pp. 9367\u20139376. PMLR (2020)"},{"key":"1364_CR66","unstructured":"Thesen, A.: Scheduling of computer programs in a multiprogramming environment (1974)"},{"issue":"2","key":"1364_CR67","doi-asserted-by":"crossref","first-page":"341","DOI":"10.1002\/nav.3800220210","volume":"22","author":"A Thesen","year":"1975","unstructured":"Thesen, A.: A recursive branch and bound algorithm for the multidimensional knapsack problem. Nav. Res. Logist. Q. 22(2), 341\u2013353 (1975)","journal-title":"Nav. Res. Logist. Q."},{"issue":"12","key":"1364_CR68","doi-asserted-by":"crossref","first-page":"1417","DOI":"10.1287\/mnsc.21.12.1417","volume":"21","author":"Y Toyoda","year":"1975","unstructured":"Toyoda, Y.: A simplified algorithm for obtaining approximate solutions to zero-one programming problems. Manag. Sci. 21(12), 1417\u20131427 (1975)","journal-title":"Manag. Sci."},{"key":"1364_CR69","unstructured":"Vasquez, M., Hao, J.-K.: A hybrid approach for the 0-1 multidimensional knapsack problem. In: IJCAI, pp. 328\u2013333 (2001)"},{"issue":"1","key":"1364_CR70","doi-asserted-by":"crossref","first-page":"70","DOI":"10.1016\/j.ejor.2004.01.024","volume":"165","author":"M Vasquez","year":"2005","unstructured":"Vasquez, M., Vimont, Y.: Improved results on the 0\u20131 multidimensional knapsack problem. Eur. J. Oper. Res. 165(1), 70\u201381 (2005)","journal-title":"Eur. J. Oper. Res."},{"key":"1364_CR71","volume-title":"Approximation Algorithms","author":"VV Vazirani","year":"2013","unstructured":"Vazirani, V.V.: Approximation Algorithms. Springer, Berlin (2013)"},{"key":"1364_CR72","unstructured":"Verma, R., Singhal, A., Khadilkar, H., Basumatary, A., Nayak, S., Singh, H.V., Kumar, S., Sinha, R.: A generalized reinforcement learning algorithm for online 3d bin-packing. arXiv preprint arXiv:2007.00463 (2020)"},{"key":"1364_CR73","unstructured":"Vinyals, O., Fortunato, M., Jaitly, N.: Pointer networks. arXiv preprint arXiv:1506.03134 (2015)"},{"issue":"7","key":"1364_CR74","doi-asserted-by":"crossref","first-page":"485","DOI":"10.1287\/mnsc.12.7.485","volume":"12","author":"HM Weingartner","year":"1966","unstructured":"Weingartner, H.M.: Capital budgeting of interrelated projects: survey and synthesis. Manag. Sci. 12(7), 485\u2013516 (1966)","journal-title":"Manag. Sci."},{"issue":"1","key":"1364_CR75","doi-asserted-by":"crossref","first-page":"83","DOI":"10.1287\/opre.15.1.83","volume":"15","author":"HM Weingartner","year":"1967","unstructured":"Weingartner, H.M., Ness, D.N.: Methods for the solution of the multidimensional 0\/1 knapsack problem. Oper. Res. 15(1), 83\u2013103 (1967)","journal-title":"Oper. Res."},{"key":"1364_CR76","doi-asserted-by":"crossref","unstructured":"Woeginger, G.J.: Exact algorithms for np-hard problems: a survey. In: Combinatorial Optimization-Eureka, You Shrink!, pp. 185\u2013207. Springer (2003)","DOI":"10.1007\/3-540-36478-1_17"},{"key":"1364_CR77","first-page":"5279","volume":"30","author":"Y Wu","year":"2017","unstructured":"Wu, Y., Mansimov, E., Grosse, R.B., Liao, S., Ba, J.: Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation. Adv. Neural. Inf. Process. Syst. 30, 5279\u20135288 (2017)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"issue":"5","key":"1364_CR78","first-page":"1291","volume":"40","author":"Yan Yang","year":"2020","unstructured":"Yang, Yan, Shengjian Liu, Y.Z.: Greedy binary lion swarm optimization algorithm for solving multidimensional knapsack problem. J. Comput. Appl. 40(5), 1291\u20131294 (2020)","journal-title":"J. Comput. Appl."},{"issue":"1","key":"1364_CR79","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1016\/S0377-2217(99)00448-8","volume":"131","author":"M-H Yang","year":"2001","unstructured":"Yang, M.-H.: An efficient algorithm to allocate shelf space. Eur. J. Oper. Res. 131(1), 107\u2013118 (2001)","journal-title":"Eur. J. Oper. Res."},{"key":"1364_CR80","unstructured":"Yang, Y., Rajgopal, J.: Learning combined set covering and traveling salesman problem. arXiv preprint arXiv:2007.03203 (2020)"},{"key":"1364_CR81","doi-asserted-by":"publisher","unstructured":"Yilmaz, D., B\u00fcy\u00fcktahtak\u0131n, \u0130.E.: An expandable learning-optimization framework for sequentially dependent decision-making. Eur. J. Oper. Res. 314(1), 280\u2013296 (2024). https:\/\/doi.org\/10.1016\/j.ejor.2023.10.045","DOI":"10.1016\/j.ejor.2023.10.045"},{"key":"1364_CR82","doi-asserted-by":"crossref","unstructured":"Yilmaz, D., B\u00fcy\u00fcktahtak\u0131n, \u0130.E.: Learning optimal solutions via an LSTM-optimization framework. Oper. Res. Forum 4(2), 28 (2023)","DOI":"10.1007\/s43069-023-00224-5"},{"issue":"1","key":"1364_CR83","doi-asserted-by":"crossref","first-page":"52","DOI":"10.1080\/24725579.2021.1938298","volume":"12","author":"X Yin","year":"2022","unstructured":"Yin, X., B\u00fcy\u00fcktahtak\u0131n, \u0130E.: Risk-averse multi-stage stochastic programming to optimizing vaccine allocation and treatment logistics for effective epidemic response. IISE Trans. Healthc. Syst. Eng. 12(1), 52\u201374 (2022)","journal-title":"IISE Trans. Healthc. Syst. Eng."},{"issue":"1","key":"1364_CR84","doi-asserted-by":"crossref","first-page":"255","DOI":"10.1016\/j.ejor.2021.11.052","volume":"304","author":"X Yin","year":"2023","unstructured":"Yin, X., B\u00fcy\u00fcktahtak\u0131n, \u0130E., Patel, B.: COVID-19: Data-driven optimal allocation of ventilator supply under uncertainty and risk. Eur. J. Oper. Res. 304(1), 255\u2013275 (2023)","journal-title":"Eur. J. Oper. Res."},{"key":"1364_CR85","doi-asserted-by":"crossref","unstructured":"Yilmaz, Dogacan and B\u00fcy\u00fcktahtak\u0131n, \u0130Esra.: A deep reinforcement learning framework for solving two-stage stochastic programs. Optimization Letters, 1\u201328 (2023)","DOI":"10.1007\/s11590-023-02009-5"}],"container-title":["Journal of Global Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10898-024-01364-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10898-024-01364-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10898-024-01364-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,15]],"date-time":"2024-06-15T04:14:16Z","timestamp":1718424856000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10898-024-01364-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,15]]},"references-count":85,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,7]]}},"alternative-id":["1364"],"URL":"https:\/\/doi.org\/10.1007\/s10898-024-01364-6","relation":{},"ISSN":["0925-5001","1573-2916"],"issn-type":[{"value":"0925-5001","type":"print"},{"value":"1573-2916","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,2,15]]},"assertion":[{"value":"26 September 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 January 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 February 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}