{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,30]],"date-time":"2026-05-30T04:00:17Z","timestamp":1780113617262,"version":"3.54.0"},"reference-count":69,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2025,9,6]],"date-time":"2025-09-06T00:00:00Z","timestamp":1757116800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,9,6]],"date-time":"2025-09-06T00:00:00Z","timestamp":1757116800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100009534","name":"Universit\u00e4t Stuttgart","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100009534","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Intell Manuf"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>In modern industrial and logistics environments, the rapid expansion of fast delivery services has heightened the demand for storage systems that combine high efficiency with increased density. Multi-deep autonomous vehicle storage and retrieval systems present a viable solution for achieving greater storage density. However, these systems encounter significant challenges during retrieval operations due to lane blockages. A conventional approach to mitigate this issue involves storing items with homogeneous characteristics in a single lane, but this strategy restricts the flexibility and adaptability of multi-deep storage systems. Building on this background, this work presents a deep reinforcement learning-based framework to optimize the retrieval process in multi-deep storage systems with heterogeneous item configurations. Each item is associated with a specific due date, and objective function is total tardiness minimization. To effectively capture the system\u2019s topology, we introduce a graph-based state representation that integrates both item attributes and the local topological structure of the multi-deep warehouse. For processing this representation, we design a novel neural network architecture that combines a Graph Neural Network (GNN) with a Transformer model. The GNN encodes topological and item-specific information into embeddings for all directly accessible items, while the Transformer maps these embeddings into global priority assignments. The Transformer\u2019s strong generalization capability further allows our approach to be applied to storage systems with diverse layouts. Extensive numerical experiments, including comparisons with heuristic methods, demonstrate the superiority of the proposed neural network architecture and the effectiveness of the trained agent in optimizing retrieval tardiness.<\/jats:p>","DOI":"10.1007\/s10845-025-02654-w","type":"journal-article","created":{"date-parts":[[2025,9,6]],"date-time":"2025-09-06T08:49:43Z","timestamp":1757148583000},"page":"2503-2536","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Topology-aware and highly generalizable deep reinforcement learning for efficient retrieval in multi-deep storage systems"],"prefix":"10.1007","volume":"37","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4455-6997","authenticated-orcid":false,"given":"Funing","family":"Li","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4521-2096","authenticated-orcid":false,"given":"Yuan","family":"Tian","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2863-2216","authenticated-orcid":false,"given":"Ruben","family":"Noortwyck","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3910-6300","authenticated-orcid":false,"given":"Jifeng","family":"Zhou","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-2853-6031","authenticated-orcid":false,"given":"Liming","family":"Kuang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0091-1632","authenticated-orcid":false,"given":"Robert","family":"Schulz","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,9,6]]},"reference":[{"key":"2654_CR1","first-page":"53007","volume":"26","author":"A Aazami","year":"2019","unstructured":"Aazami, A., & Saidi-Mehrabad, M. (2019). Benders decomposition algorithm for robust aggregate production planning considering pricing decisions in competitive environment: A case study. Scientia Iranica, 26, 53007\u20133031.","journal-title":"Scientia Iranica"},{"key":"2654_CR2","doi-asserted-by":"publisher","first-page":"223","DOI":"10.1016\/j.jmsy.2020.12.001","volume":"58","author":"A Aazami","year":"2021","unstructured":"Aazami, A., & Saidi-Mehrabad, M. (2021). A production and distribution planning of perishable products with a fixed lifetime under vertical competition in the seller-buyer systems: A real-world application. J. Manuf. Syst., 58, 223\u2013247.","journal-title":"J. Manuf. Syst."},{"key":"2654_CR3","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1007\/s00170-016-9962-9","volume":"92","author":"R Accorsi","year":"2017","unstructured":"Accorsi, R., Baruffaldi, G., & Manzini, R. (2017). Design and manage deep lane storage system layout. an iterative decision-support model. The International Journal of Advanced Manufacturing Technology, 92, 57\u201367.","journal-title":"The International Journal of Advanced Manufacturing Technology"},{"key":"2654_CR4","doi-asserted-by":"crossref","unstructured":"Azad, N., Aazami, A., Papi, A., Jabbarzadeh, A. (2019). A two-phase genetic algorithm for incorporating environmental considerations with production, inventory and routing decisions in supply chain networks. Proceedings of the Genetic and Evolutionary Computation Conference Companion (PP 41\u201342). Association for Computing Machinery.","DOI":"10.1145\/3319619.3326781"},{"key":"2654_CR5","doi-asserted-by":"publisher","first-page":"4917","DOI":"10.1287\/trsc.2018.0873","volume":"53","author":"K Azadeh","year":"2019","unstructured":"Azadeh, K., De Koster, R., & Roy, D. (2019). Robotized and automated warehouse systems: Review and recent developments. Transportation Science, 53, 4917\u2013945.","journal-title":"Transportation Science"},{"key":"2654_CR6","doi-asserted-by":"publisher","DOI":"10.1002\/9781119262602","volume-title":"Principles of sequencing and scheduling","author":"KR Baker","year":"2018","unstructured":"Baker, K. R., & Trietsch, D. (2018). Principles of sequencing and scheduling. Hoboken: John Wiley & Sons."},{"issue":"3","key":"2654_CR7","doi-asserted-by":"publisher","first-page":"1699","DOI":"10.1111\/itor.12454","volume":"27","author":"F Ballest\u00edn","year":"2020","unstructured":"Ballest\u00edn, F., P\u00e9rez, \u00c1., & Quintanilla, S. (2020). A multistage heuristic for storage and retrieval problems in a warehouse with random storage. International Transactions in Operational Research, 27(3), 1699\u20131728.","journal-title":"International Transactions in Operational Research"},{"key":"2654_CR8","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijpe.2021.108342","volume":"243","author":"R Braune","year":"2022","unstructured":"Braune, R., Benda, F., Doerner, K. F., & Hartl, R. F. (2022). A genetic programming learning approach to generate dispatching rules for flexible shop scheduling problems. International Journal of Production Economics, 243, Article 108342.","journal-title":"International Journal of Production Economics"},{"key":"2654_CR9","doi-asserted-by":"publisher","first-page":"107221","DOI":"10.1016\/j.cie.2021.107221","volume":"156","author":"B Cals","year":"2021","unstructured":"Cals, B., Zhang, Y., Dijkman, R., & van Dorst, C. (2021). Solving the online batching problem using deep reinforcement learning. Computers & Industrial Engineering, 156, 107221.","journal-title":"Computers & Industrial Engineering"},{"key":"2654_CR10","doi-asserted-by":"crossref","unstructured":"Cebi, C., Atac, E., Sahingoz, O.K. (2020). Job shop scheduling problem and solution algorithms: a review. 2020 11th International Conference on Computing, Communication and Networking Technologies (PP 1\u20137). IEEE.","DOI":"10.1109\/ICCCNT49239.2020.9225581"},{"key":"2654_CR11","doi-asserted-by":"publisher","first-page":"109053","DOI":"10.1016\/j.cie.2023.109053","volume":"177","author":"Z Chen","year":"2023","unstructured":"Chen, Z., Zhang, L., Wang, X., & Wang, K. (2023). Cloud-edge collaboration task scheduling in cloud manufacturing: An attention-based deep reinforcement learning approach. Computers & Industrial Engineering, 177, 109053.","journal-title":"Computers & Industrial Engineering"},{"key":"2654_CR12","first-page":"10707","volume":"33","author":"F Christianos","year":"2020","unstructured":"Christianos, F., Sch\u00e4fer, L., & Albrecht, S. (2020). Shared experience actor-critic for multi-agent reinforcement learning. Advances in neural information processing systems, 33, 10707\u201310717. Curran Associates, Inc.","journal-title":"Advances in neural information processing systems"},{"issue":"7897","key":"2654_CR13","doi-asserted-by":"publisher","first-page":"414","DOI":"10.1038\/s41586-021-04301-9","volume":"602","author":"J Degrave","year":"2022","unstructured":"Degrave, J., Felici, F., Buchli, J., Neunert, M., Tracey, B., Carpanese, F., & Riedmiller, M. (2022). Magnetic control of tokamak plasmas through deep reinforcement learning. Nature, 602(7897), 414\u2013419.","journal-title":"Nature"},{"issue":"1","key":"2654_CR14","doi-asserted-by":"publisher","first-page":"269","DOI":"10.1007\/BF01386390","volume":"1","author":"EW Dijkstra","year":"1959","unstructured":"Dijkstra, E. W. (1959). A note on two problems in connexion with graphs. Numerical Mathematics, 1(1), 269\u2013271.","journal-title":"Numerical Mathematics"},{"issue":"2","key":"2654_CR15","doi-asserted-by":"publisher","first-page":"630","DOI":"10.1016\/j.ejor.2023.10.006","volume":"314","author":"W Dong","year":"2024","unstructured":"Dong, W., & Jin, M. (2024). Automated storage and retrieval system design with variant lane depths. European Journal of Operational Research, 314(2), 630\u2013646.","journal-title":"European Journal of Operational Research"},{"issue":"1","key":"2654_CR16","doi-asserted-by":"publisher","first-page":"111","DOI":"10.1007\/s10479-021-03967-8","volume":"302","author":"W Dong","year":"2021","unstructured":"Dong, W., Jin, M., Wang, Y., & Kelle, P. (2021). Retrieval scheduling in crane-based 3d automated retrieval and storage systems with shuttles. Annals of Operations Research, 302(1), 111\u2013135.","journal-title":"Annals of Operations Research"},{"key":"2654_CR17","doi-asserted-by":"publisher","first-page":"92","DOI":"10.1016\/j.neucom.2022.06.111","volume":"503","author":"SR Dubey","year":"2022","unstructured":"Dubey, S. R., Singh, S. K., & Chaudhuri, B. B. (2022). Activation functions in deep learning: A comprehensive survey and benchmark. Neurocomputing, 503, 92\u2013108.","journal-title":"Neurocomputing"},{"key":"2654_CR18","doi-asserted-by":"publisher","first-page":"1919","DOI":"10.1007\/s00170-019-03985-8","volume":"104","author":"G D\u2019Antonio","year":"2019","unstructured":"D\u2019Antonio, G., & Chiabert, P. (2019). Analytical models for cycle time and throughput evaluation of multi-shuttle deep-lane AVS\/RS. The International Journal of Advanced Manufacturing Technology, 104, 1919\u20131936.","journal-title":"The International Journal of Advanced Manufacturing Technology"},{"issue":"1","key":"2654_CR19","doi-asserted-by":"publisher","first-page":"859","DOI":"10.1007\/s00170-019-04831-7","volume":"107","author":"M Eder","year":"2020","unstructured":"Eder, M. (2020). An approach for a performance calculation of shuttle-based storage and retrieval systems with multiple-deep storage. The International Journal of Advanced Manufacturing Technology, 107(1), 859\u2013873.","journal-title":"The International Journal of Advanced Manufacturing Technology"},{"issue":"16","key":"2654_CR20","doi-asserted-by":"publisher","first-page":"5772","DOI":"10.1080\/00207543.2022.2104180","volume":"61","author":"A Esteso","year":"2023","unstructured":"Esteso, A., Peidro, D., Mula, J., & D\u00edaz-Madro\u00f1ero, M. (2023). Reinforcement learning applied to production planning and control. International Journal of Production Research, 61(16), 5772\u20135789.","journal-title":"International Journal of Production Research"},{"key":"2654_CR21","doi-asserted-by":"publisher","first-page":"3386","DOI":"10.1109\/ACCESS.2023.3347047","volume":"12","author":"AE-S Ezugwu","year":"2024","unstructured":"Ezugwu, A.E.-S. (2024). Metaheuristic optimization for sustainable unrelated parallel machine scheduling: A concise overview with a proof-of-concept study. IEEE Access, 12, 3386\u20133416.","journal-title":"IEEE Access"},{"issue":"7930","key":"2654_CR22","doi-asserted-by":"publisher","first-page":"47","DOI":"10.1038\/s41586-022-05172-4","volume":"610","author":"A Fawzi","year":"2022","unstructured":"Fawzi, A., Balog, M., Huang, A., Hubert, T., Romera-Paredes, B., Barekatain, M., & Kohli, P. (2022). Discovering faster matrix multiplication algorithms with reinforcement learning. Nature, 610(7930), 47\u201353.","journal-title":"Nature"},{"issue":"3","key":"2654_CR23","doi-asserted-by":"publisher","first-page":"245","DOI":"10.1080\/07408179108963859","volume":"23","author":"M Goetschalckx","year":"1991","unstructured":"Goetschalckx, M., & Donald Ratldff, H. (1991). Optimal lane depths for single and multiple products in block stacking storage systems. IIE Transaction, 23(3), 245\u2013258.","journal-title":"IIE Transaction"},{"issue":"1","key":"2654_CR24","first-page":"88","volume":"16","author":"A Goli","year":"2018","unstructured":"Goli, A., Aazami, A., & Jabbarzadeh, A. (2018). Accelerated cuckoo optimization algorithm for capacitated vehicle routing problem in competitive conditions. International Journal of Artificial Intelligence, 16(1), 88\u2013112.","journal-title":"International Journal of Artificial Intelligence"},{"key":"2654_CR25","doi-asserted-by":"publisher","first-page":"287","DOI":"10.1016\/S0167-5060(08)70356-X","volume":"5","author":"RL Graham","year":"1979","unstructured":"Graham, R. L., Lawler, E. L., Lenstra, J. K., & Kan, A. R. (1979). Optimization and approximation in deterministic sequencing and scheduling: A survey. Annals of Discrete Mathematics, 5, 287\u2013326.","journal-title":"Annals of Discrete Mathematics"},{"key":"2654_CR26","unstructured":"Grand View Research (2024). Same day delivery market size, share & trends analysis report, 2025\u20132030. https:\/\/www.grandviewresearch.com\/industry-analysis\/same-day-delivery-market."},{"key":"2654_CR27","unstructured":"Hamilton, W.L., Ying, R., Leskovec, J. (2017). Inductive representation learning on large graphs. Advances in Neural Information Processing Systems (VOL\u00a030). Curran Associates, Inc."},{"issue":"7","key":"2654_CR28","doi-asserted-by":"publisher","first-page":"787","DOI":"10.1038\/s42256-024-00861-3","volume":"6","author":"L Han","year":"2024","unstructured":"Han, L., Zhu, Q., Sheng, J., Zhang, C., Li, T., Zhang, Y., & Zhang, Z. (2024). Lifelike agility and play in quadrupedal robots using reinforcement learning and generative pre-trained models. Nature Machine Intelligence, 6(7), 787\u2013798.","journal-title":"Nature Machine Intelligence"},{"issue":"2","key":"2654_CR29","doi-asserted-by":"publisher","first-page":"820","DOI":"10.1016\/j.ejor.2022.03.042","volume":"305","author":"J He","year":"2023","unstructured":"He, J., Liu, X., Duan, Q., Chan, W. K. V., & Qi, M. (2023). Reinforcement learning for multi-item retrieval in the puzzle-based storage system. European Journal of Operational Research, 305(2), 820\u2013837.","journal-title":"European Journal of Operational Research"},{"issue":"2","key":"2654_CR30","first-page":"134","volume":"11","author":"M Heydari","year":"2018","unstructured":"Heydari, M., & Aazami, A. (2018). Minimizing the maximum tardiness and makespan criteria in a job shop scheduling problem with sequence dependent setup times. Journal of Industrial and Systems Engineering, 11(2), 134\u2013150.","journal-title":"Journal of Industrial and Systems Engineering"},{"key":"2654_CR31","doi-asserted-by":"publisher","DOI":"10.1016\/j.cor.2023.106474","volume":"162","author":"K Hu","year":"2024","unstructured":"Hu, K., Che, Y., Ng, T. S., & Deng, J. (2024). Unrelated parallel batch processing machine scheduling with time requirements and two-dimensional packing constraints. Computers & Operations Research, 162, Article 106474.","journal-title":"Computers & Operations Research"},{"key":"2654_CR32","unstructured":"IHL Group 2023. Retail inventory distortion study\u2013 the good, the bad, the ugly\u2013 2023. https:\/\/www.ihlservices.com\/product\/retail-inventory-distortion-study-the-good-the-bad-the-ugly-2023\/."},{"key":"2654_CR33","doi-asserted-by":"publisher","first-page":"105091","DOI":"10.1016\/j.cor.2020.105091","volume":"125","author":"Z-Z Jiang","year":"2021","unstructured":"Jiang, Z.-Z., Wan, M., Pei, Z., & Qin, X. (2021). Spatial and temporal optimization for smart warehouses with fast turnover. Computers & Operations Research, 125, 105091.","journal-title":"Computers & Operations Research"},{"key":"2654_CR34","doi-asserted-by":"publisher","first-page":"109088","DOI":"10.1016\/j.ijpe.2023.109088","volume":"267","author":"I Kaynov","year":"2024","unstructured":"Kaynov, I., van Knippenberg, M., Menkovski, V., van Breemen, A., & van Jaarsveld, W. (2024). Deep reinforcement learning for one-warehouse multi-retailer inventory management. International Journal of Production Economics, 267, 109088.","journal-title":"International Journal of Production Economics"},{"key":"2654_CR35","doi-asserted-by":"publisher","DOI":"10.1016\/j.tre.2020.102083","volume":"143","author":"MA Klapp","year":"2020","unstructured":"Klapp, M. A., Erera, A. L., & Toriello, A. (2020). Request acceptance in same-day delivery. Transportation Research Part E: Logistics and Transportation Review, 143, Article 102083.","journal-title":"Transportation Research Part E: Logistics and Transportation Review"},{"key":"2654_CR36","doi-asserted-by":"crossref","unstructured":"Krnjaic, A., Steleac, R.D., Thomas, J.D., Papoudakis, G., Sch\u00e4fer, L., To, A.W.K.,V.Albrecht, S. (2024). Scalable multi-agent reinforcement learning for warehouse logistics with robotic and human co-workers. 2024 IEEE\/RSJ International Conference on Intelligent Robots and Systems (PP 677\u2013684). IEEE.","DOI":"10.1109\/IROS58592.2024.10802813"},{"key":"2654_CR37","doi-asserted-by":"publisher","first-page":"121","DOI":"10.1016\/S0167-5060(08)70821-5","volume":"4","author":"JK Lenstra","year":"1979","unstructured":"Lenstra, J. K., & Kan, A. R. (1979). Computational complexity of discrete optimization problems. Annals of Discrete Mathematics, 4, 121\u2013140.","journal-title":"Annals of Discrete Mathematics"},{"key":"2654_CR38","doi-asserted-by":"crossref","unstructured":"Lerher, T., Marolt, J., Sgarbossa, F., Ekren, B.Y., Dukic, G. 2024. Design and operation of single-and multi-deep shuttle-based storage and retrieval systems (SBS\/RS). Warehousing and material handling systems for the digital industry: The new challenges for the digital circular economy (PP 333\u2013375). Springer.","DOI":"10.1007\/978-3-031-50273-6_13"},{"issue":"3","key":"2654_CR39","doi-asserted-by":"publisher","first-page":"1107","DOI":"10.1007\/s10845-023-02094-4","volume":"35","author":"F Li","year":"2024","unstructured":"Li, F., Lang, S., Hong, B., & Reggelin, T. (2024). A two-stage RNN-based deep reinforcement learning approach for solving the parallel machine scheduling problem with due dates and family setups. Journal of Intelligent Manufacturing, 35(3), 1107\u20131140.","journal-title":"Journal of Intelligent Manufacturing"},{"key":"2654_CR40","doi-asserted-by":"crossref","unstructured":"Li, F., Lang, S., Tian, Y., Hong, B., Rolf, B., Noortwyck, R.,Reggelin, T. (2024). A Transformer-based deep reinforcement learning approach for dynamic parallel machine scheduling problem with family setups. Journal of Intelligent Manufacturing 1\u201334,","DOI":"10.1007\/s10845-024-02470-8"},{"issue":"2","key":"2654_CR41","doi-asserted-by":"publisher","first-page":"131","DOI":"10.23919\/CSMS.2021.0013","volume":"1","author":"L Luo","year":"2021","unstructured":"Luo, L., Zhao, N., & Lodewijks, G. (2021). Scheduling storage process of shuttle-based storage and retrieval systems based on reinforcement learning. Complex System Modeling and Simulation, 1(2), 131\u2013144.","journal-title":"Complex System Modeling and Simulation"},{"issue":"3","key":"2654_CR42","doi-asserted-by":"publisher","first-page":"1449","DOI":"10.1007\/s00170-024-13160-3","volume":"131","author":"G Lupi","year":"2024","unstructured":"Lupi, G., Accorsi, R., Battarra, I., Manzini, R., & Sirri, G. (2024). Space efficiency and throughput performance in AVS\/RS under variant lane depths. The International Journal of Advanced Manufacturing Technology, 131(3), 1449\u20131466.","journal-title":"The International Journal of Advanced Manufacturing Technology"},{"key":"2654_CR43","doi-asserted-by":"publisher","first-page":"34","DOI":"10.1016\/j.cie.2016.08.005","volume":"100","author":"A Makui","year":"2016","unstructured":"Makui, A., Heydari, M., Aazami, A., & Dehghani, E. (2016). Accelerating benders decomposition approach for robust aggregate production planning of products with a very limited expiration date. Computers & industrial engineering, 100, 34\u201351.","journal-title":"Computers & industrial engineering"},{"key":"2654_CR44","unstructured":"MarketsandMarkets (2024). Warehouse management system market size, share, statistics and industry growth analysis report by offering (software, services), deployment (on-premises, cloud-based), tier (advanced, intermediate, basic), end user (automotive, e-commerce, electricals & electronics) and region - global growth driver and industry forecast to 2029. https:\/\/www.marketsandmarkets.com\/Market-Reports\/warehouse-management-system-market-41614951.html."},{"issue":"15","key":"2654_CR45","doi-asserted-by":"publisher","first-page":"4991","DOI":"10.1080\/00207543.2022.2087568","volume":"61","author":"J Marolt","year":"2023","unstructured":"Marolt, J., \u0160inko, S., & Lerher, T. (2023). Model of a multiple-deep automated vehicles storage and retrieval system following the combination of depth-first storage and depth-first relocation strategies. International Journal of Production Research, 61(15), 4991\u20135008.","journal-title":"International Journal of Production Research"},{"key":"2654_CR46","unstructured":"Ng, A.Y., Harada, D., Russell, S. (1999). Policy invariance under reward transformations: Theory and application to reward shaping. Proceedings of the 16th International Conference on Machine Learning (PP 278\u2013287). Morgan Kaufmann Publishers Inc."},{"key":"2654_CR47","unstructured":"Oono, K., & Suzuki, T. (2019). Graph neural networks exponentially lose expressive power for node classification. arXiv preprint arXiv:1905.10947."},{"key":"2654_CR48","unstructured":"Papoudakis, G., Christianos, F., Sch\u00e4fer, L., Albrecht, S.V. (2021). Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks. arXiv preprint arXiv:2006.07869."},{"key":"2654_CR49","doi-asserted-by":"publisher","first-page":"31491","DOI":"10.3390\/app12031491","volume":"12","author":"J Para","year":"2022","unstructured":"Para, J., Del Ser, J., & Nebro, A. J. (2022). Energy-aware multi-objective job shop scheduling optimization with metaheuristics in manufacturing industries: a critical survey, results, and perspectives. Applied Science, 12, 31491.","journal-title":"Applied Science"},{"key":"2654_CR50","doi-asserted-by":"publisher","first-page":"2395","DOI":"10.1016\/j.ejor.2019.01.063","volume":"280","author":"R Pellerin","year":"2020","unstructured":"Pellerin, R., Perrier, N., & Berthaut, F. (2020). A survey of hybrid metaheuristics for the resource-constrained project scheduling problem. European Journal of Operational Research, 280, 2395\u2013416.","journal-title":"European Journal of Operational Research"},{"key":"2654_CR51","first-page":"2363","volume":"33","author":"CN Potts","year":"1985","unstructured":"Potts, C. N., & Van Wassenhove, L. N. (1985). A branch and bound algorithm for the total weighted tardiness problem. Operertion Research, 33, 2363\u2013377.","journal-title":"Operertion Research"},{"issue":"11","key":"2654_CR52","doi-asserted-by":"publisher","first-page":"13187","DOI":"10.1007\/s10462-023-10470-y","volume":"56","author":"K Rajwar","year":"2023","unstructured":"Rajwar, K., Deep, K., & Das, S. (2023). An exhaustive review of the metaheuristic algorithms for search and optimization: taxonomy, applications, and open challenges. Artificial Intelligence Review, 56(11), 13187\u201313257.","journal-title":"Artificial Intelligence Review"},{"key":"2654_CR53","first-page":"3140","volume":"10","author":"M Saeedi Mehrabad","year":"2017","unstructured":"Saeedi Mehrabad, M., Aazami, A., & Goli, A. (2017). A location-allocation model in the multi-level supply chain with multi-objective evolutionary approach. Journal of Industrial and Systems Engineering, 10, 3140\u2013160.","journal-title":"Journal of Industrial and Systems Engineering"},{"key":"2654_CR54","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347."},{"key":"2654_CR55","doi-asserted-by":"publisher","first-page":"1140","DOI":"10.1126\/science.aar6404","volume":"3626419","author":"D Silver","year":"2018","unstructured":"Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., & Hassabis, D. (2018). A general reinforcement learning algorithm that masters chess, shogi, and go through self-play. Science, 3626419, 1140\u20131144.","journal-title":"Science"},{"key":"2654_CR56","unstructured":"Sutton, R.S., & Barto, A.G. (2018). Reinforcement learning: An introduction (2nd ed). MIT Press."},{"key":"2654_CR57","doi-asserted-by":"publisher","first-page":"165","DOI":"10.1016\/j.neucom.2022.10.045","volume":"517","author":"Y Tian","year":"2023","unstructured":"Tian, Y., Kladny, K.-R., Wang, Q., Huang, Z., & Fink, O. (2023). Multi-agent actor-critic with time dynamical opponent model. Neurocomputing, 517, 165\u2013172.","journal-title":"Neurocomputing"},{"key":"2654_CR58","doi-asserted-by":"crossref","unstructured":"Tian, Y., Wang, Q., Huang, Z., Li, W., Dai, D., Yang, M.,Fink, O. (2020). Off-policy reinforcement learning for efficient and effective gan architecture search. Computer Vision\u2013ECCV 2020: 16th European Conference (PP 175\u2013192). Springer.","DOI":"10.1007\/978-3-030-58571-6_11"},{"key":"2654_CR59","first-page":"13930","volume":"16","author":"Y Tian","year":"2025","unstructured":"Tian, Y., Zhou, W., Viscione, M., Dong, H., Kammer, D. S., & Fink, O. (2025). Interactive symbolic regression with co-design mechanism through offline reinforcement learning. Nature Communication, 16, 13930.","journal-title":"Nature Communication"},{"key":"2654_CR60","unstructured":"United Nations Conference on Trade and Development (UNCTAD) 2024. Digital economy report 2024: Shaping an environmentally sustainable and inclusive digital future. https:\/\/unctad.org\/publication\/digital-economy-report-2024."},{"key":"2654_CR61","first-page":"5998","volume":"30","author":"A Vaswani","year":"2017","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998\u20136008. Curran Associates, Inc.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"7782","key":"2654_CR62","doi-asserted-by":"publisher","first-page":"350","DOI":"10.1038\/s41586-019-1724-z","volume":"575","author":"O Vinyals","year":"2019","unstructured":"Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., & Silver, D. (2019). Grandmaster level in starcraft ii using multi-agent reinforcement learning. Nature, 575(7782), 350\u2013354.","journal-title":"Nature"},{"key":"2654_CR63","doi-asserted-by":"publisher","first-page":"452","DOI":"10.1016\/j.jmsy.2022.08.013","volume":"65","author":"X Wang","year":"2022","unstructured":"Wang, X., Zhang, L., Liu, Y., Zhao, C., & Wang, K. (2022). Solving task scheduling problems in cloud manufacturing via attention mechanism and deep reinforcement learning. Journal of Manufacturing Systems, 65, 452\u2013468.","journal-title":"Journal of Manufacturing Systems"},{"key":"2654_CR64","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1016\/j.ejor.2016.07.030","volume":"257","author":"H Xiong","year":"2017","unstructured":"Xiong, H., Fan, H., Jiang, G., & Li, G. (2017). A simulation-based study of dispatching rules in a dynamic job shop scheduling problem with batch release and extended technical precedence constraints. European Journal of Operational Research, 257, 113\u201324.","journal-title":"European Journal of Operational Research"},{"issue":"21","key":"2654_CR65","doi-asserted-by":"publisher","first-page":"7333","DOI":"10.1080\/00207543.2022.2148011","volume":"61","author":"Z Xu","year":"2023","unstructured":"Xu, Z., Xu, L., Ling, X., & Zhang, B. (2023). Data-driven hierarchical learning and real-time decision-making of equipment scheduling and location assignment in automatic high-density storage systems. International Journal of Production Research, 61(21), 7333\u20137352.","journal-title":"International Journal of Production Research"},{"key":"2654_CR66","doi-asserted-by":"publisher","DOI":"10.1016\/j.tre.2022.102712","volume":"162","author":"Y Yan","year":"2022","unstructured":"Yan, Y., Chow, A. H., Ho, C. P., Kuo, Y.-H., Wu, Q., & Ying, C. (2022). Reinforcement learning for logistics and supply chain management: Methodologies, state of the art, and future opportunities. Transportation Research Part E: Logistics and Transportation Review, 162, Article 102712.","journal-title":"Transportation Research Part E: Logistics and Transportation Review"},{"key":"2654_CR67","doi-asserted-by":"publisher","first-page":"2696","DOI":"10.1016\/j.ejor.2022.11.037","volume":"308","author":"J Yang","year":"2023","unstructured":"Yang, J., de Koster, R. B., Guo, X., & Yu, Y. (2023). Scheduling shuttles in deep-lane shuttle-based storage systems. European Journal of Operational Research, 308, 2696\u2013708.","journal-title":"European Journal of Operational Research"},{"key":"2654_CR68","doi-asserted-by":"publisher","first-page":"269","DOI":"10.1080\/0740817X.2011.575441","volume":"44","author":"Y Yu","year":"2012","unstructured":"Yu, Y., & De Koster, R. B. (2012). Sequencing heuristics for storing and retrieving unit loads in 3d compact automated warehousing systems. IIE Transaction, 44, 269\u201387.","journal-title":"IIE Transaction"},{"key":"2654_CR69","doi-asserted-by":"publisher","first-page":"6705","DOI":"10.3390\/su16156705","volume":"1615","author":"M Zarreh","year":"2024","unstructured":"Zarreh, M., Khandan, M., Goli, A., Aazami, A., & Kummer, S. (2024). Integrating perishables into closed-loop supply chains: A comprehensive review. Sustainability, 1615, 6705.","journal-title":"Sustainability"}],"container-title":["Journal of Intelligent Manufacturing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10845-025-02654-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10845-025-02654-w","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10845-025-02654-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,14]],"date-time":"2026-05-14T09:05:39Z","timestamp":1778749539000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10845-025-02654-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,6]]},"references-count":69,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["2654"],"URL":"https:\/\/doi.org\/10.1007\/s10845-025-02654-w","relation":{},"ISSN":["0956-5515","1572-8145"],"issn-type":[{"value":"0956-5515","type":"print"},{"value":"1572-8145","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,6]]},"assertion":[{"value":"11 February 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 July 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 September 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}