{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T17:08:47Z","timestamp":1785431327415,"version":"3.56.0"},"reference-count":38,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T00:00:00Z","timestamp":1780876800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100004329","name":"Slovenian Research and Innovation Agency","doi-asserted-by":"publisher","award":["P5-0018"],"award-info":[{"award-number":["P5-0018"]}],"id":[{"id":"10.13039\/501100004329","id-type":"DOI","asserted-by":"publisher"}]},{"award":["P5-0018"],"award-info":[{"award-number":["P5-0018"]}],"id":[{"id":"https:\/\/ror.org\/059bp8k51","id-type":"ROR","asserted-by":"publisher"}]},{"name":"Ministry of Higher Education, Science, and Innovation of the Republic of Slovenia","award":["3330-22-3515"],"award-info":[{"award-number":["3330-22-3515"]}]},{"name":"Ministry of Higher Education, Science, and Innovation of the Republic of Slovenia","award":["C3330-22-953012"],"award-info":[{"award-number":["C3330-22-953012"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Systems"],"abstract":"<jats:p>Order picking is one of the most resource-intensive warehouse operations; therefore, improving routing efficiency remains an important challenge. Deep reinforcement learning (DRL) has shown promise in complex optimization problems. However, its application to warehouse order picking is still limited, and graph-based representation learning using graph neural networks (GNNs) in this context remains largely unexplored. This paper proposes a GNN-based DRL method that models warehouse layouts as graphs to optimize order-picking paths while simultaneously learning graph-based structural embeddings of storage locations. The approach is evaluated against exact optimal solutions for smaller instances and against classical heuristic baselines, including the Lin\u2013Kernighan algorithm, in simulated warehouse environments of different scales. The results show that the proposed GNN\u2013DRL approach produces routing solutions with low optimality gaps across different order sizes and remains effective across different warehouse sizes when fine-tuned. In addition, a preliminary small-scale multi-picker experiment illustrates that the proposed framework could be extended toward more complex warehouse optimization settings. Moreover, the learned node embeddings capture meaningful structural properties of warehouse layouts and adapt to different operational contexts, highlighting the potential of integrating GNNs and DRL as a flexible foundation for advanced warehouse optimization.<\/jats:p>","DOI":"10.3390\/systems14060659","type":"journal-article","created":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T10:04:14Z","timestamp":1780913054000},"page":"659","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Graph Neural Networks and Deep Reinforcement Learning for Warehouse Order Picking and Representation Learning"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-3799-5083","authenticated-orcid":false,"given":"Nejc","family":"\u010celik","sequence":"first","affiliation":[{"name":"Cybernetics & Decision Support Systems Laboratory, Faculty of Organizational Sciences, University of Maribor, Kidri\u010deva Cesta 55a, 4000 Kranj, Slovenia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4471-6743","authenticated-orcid":false,"given":"Andrej","family":"\u0160kraba","sequence":"additional","affiliation":[{"name":"Cybernetics & Decision Support Systems Laboratory, Faculty of Organizational Sciences, University of Maribor, Kidri\u010deva Cesta 55a, 4000 Kranj, Slovenia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,6,8]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"107564","DOI":"10.1016\/j.ijpe.2019.107564","article-title":"Order Picker Routing in Warehouses: A Systematic Literature Review","volume":"224","author":"Masae","year":"2020","journal-title":"Int. J. Prod. Econ."},{"key":"ref_2","first-page":"171","article-title":"Robotic Warehousing Operations: A Learn-Then-Optimize Approach to Large-Scale Neighborhood Search","volume":"7","author":"Barnhart","year":"2025","journal-title":"Inf. J. Optim."},{"key":"ref_3","first-page":"266","article-title":"Algorithm for Robotic Picking in Amazon Fulfillment Centers Enables Humans and Robots to Work Together Effectively","volume":"53","author":"Allgor","year":"2023","journal-title":"Inf. J. Appl. Anal."},{"key":"ref_4","first-page":"1671","article-title":"How to Deploy Robotic Mobile Fulfillment Systems","volume":"57","author":"Zhen","year":"2023","journal-title":"Transp. Sci."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1002\/net.21628","article-title":"Dynamic Vehicle Routing Problems: Three Decades and Counting","volume":"67","author":"Psaraftis","year":"2016","journal-title":"Networks"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"185","DOI":"10.1287\/trsc.2017.0767","article-title":"Offline\u2013Online Approximate Dynamic Programming for Dynamic Vehicle Routing with Stochastic Requests","volume":"53","author":"Ulmer","year":"2019","journal-title":"Transp. Sci."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"507","DOI":"10.1287\/opre.31.3.507","article-title":"Order-Picking in a Rectangular Warehouse: A Solvable Case of the Traveling Salesman Problem","volume":"31","author":"Ratliff","year":"1983","journal-title":"Oper. Res."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"105400","DOI":"10.1016\/j.cor.2021.105400","article-title":"Reinforcement Learning for Combinatorial Optimization: A Survey","volume":"134","author":"Mazyavkina","year":"2021","journal-title":"Comput. Oper. Res."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"107221","DOI":"10.1016\/j.cie.2021.107221","article-title":"Solving the Online Batching Problem Using Deep Reinforcement Learning","volume":"156","author":"Cals","year":"2021","journal-title":"Comput. Ind. Eng."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"107112","DOI":"10.1016\/j.cor.2025.107112","article-title":"Deep Reinforcement Learning for Dynamic Order Picking in Warehouse Operations","volume":"182","author":"Mahmoudinazlou","year":"2025","journal-title":"Comput. Oper. Res."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"102959","DOI":"10.1016\/j.rcim.2025.102959","article-title":"Dynamic Multi-Tour Order Picking in an Automotive-Part Warehouse Based on Attention-Aware Deep Reinforcement Learning","volume":"94","author":"Wang","year":"2025","journal-title":"Robot. Comput.-Integr. Manuf."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"106207","DOI":"10.1016\/j.neunet.2024.106207","article-title":"A Comprehensive Survey on Deep Graph Representation Learning","volume":"173","author":"Ju","year":"2024","journal-title":"Neural Netw."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1109\/TNNLS.2020.2978386","article-title":"A Comprehensive Survey on Graph Neural Networks","volume":"32","author":"Wu","year":"2021","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_14","first-page":"89","article-title":"Clustering-Based Optimisation of Multiple Traveling Salesman Problem","volume":"8","year":"2019","journal-title":"Prod. Syst. Inf. Eng."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"6450","DOI":"10.1080\/00207543.2018.1456692","article-title":"A Cross-Entropy Method for Optimising Robotic Automated Storage and Retrieval Systems","volume":"56","author":"Foumani","year":"2018","journal-title":"Int. J. Prod. Res."},{"key":"ref_16","unstructured":"Begnardi, L., Baier, H., van Jaarsveld, W., and Zhang, Y. (2024). Deep Reinforcement Learning for Two-Sided Online Bipartite Matching in Collaborative Order Picking. Proceedings of the 15th Asian Conference on Machine Learning, PMLR."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Cho, S.-H., Shin, W.-J., Ahn, J., Joo, S., and Kim, H.-J. (2024). Dynamic Crane Scheduling with Reinforcement Learning for a Steel Coil Warehouse. Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), IEEE.","DOI":"10.1109\/ICRA57147.2024.10610858"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"102250","DOI":"10.1016\/j.inffus.2024.102250","article-title":"MACNS: A Generic Graph Neural Network Integrated Deep Reinforcement Learning Based Multi-Agent Collaborative Navigation System for Dynamic Trajectory Planning","volume":"105","author":"Xiao","year":"2024","journal-title":"Inf. Fusion"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Ansel, J., Yang, E., He, H., Gimelshein, N., Jain, A., Voznesensky, M., Bao, B., Bell, P., Berard, D., and Burovski, E. (2024). PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation. Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems ASPLOS: 24, Association for Computing Machinery.","DOI":"10.1145\/3620665.3640366"},{"key":"ref_20","unstructured":"Fey, M., Sunil, J., Nitta, A., Puri, R., Shah, M., Stojanovi\u010d, B., Bendias, R., Barghi, A., Kocijan, V., and Zhang, Z. (2025). PyG 2.0: Scalable Learning on Real World Graphs. arXiv."},{"key":"ref_21","unstructured":"Fey, M., and Lenssen, J.E. (2019). Fast Graph Representation Learning with PyTorch Geometric. arXiv."},{"key":"ref_22","unstructured":"(2026, March 10). GitHub-Networkx\/Networkx: Network Analysis in Python. GitHub. Available online: https:\/\/github.com\/networkx\/networkx."},{"key":"ref_23","unstructured":"(2026, March 10). GitHub-Fillipe-Gsm\/Python-Tsp: Library to Solve Traveling Salesperson Problems with Pure Python Code. GitHub. Available online: https:\/\/github.com\/fillipe-gsm\/python-tsp."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1126\/science.153.3731.34","article-title":"Dynamic Programming","volume":"153","author":"Bellman","year":"1966","journal-title":"Science"},{"key":"ref_25","unstructured":"Christofides, N. (1976). Worst-Case Analysis of a New Heuristic for the Travelling Salesman Problem, Springer International Publishing."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"791","DOI":"10.1287\/opre.6.6.791","article-title":"A Method for Solving Traveling-Salesman Problems","volume":"6","author":"Croes","year":"1958","journal-title":"Oper. Res."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"498","DOI":"10.1287\/opre.21.2.498","article-title":"An Effective Heuristic Algorithm for the Traveling-Salesman Problem","volume":"21","author":"Lin","year":"1973","journal-title":"Oper. Res."},{"key":"ref_28","unstructured":"\u010celik, N., and \u0160kraba, A. (2025). Graph-Based Modeling of Warehouse Layouts Based on Data from Relation Database for Optimizing Order Picking Path, Slovenian Society Informatika, Section for Operational Research."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"269","DOI":"10.1007\/BF01386390","article-title":"A Note on Two Problems in Connexion with Graphs","volume":"1","author":"Dijkstra","year":"1959","journal-title":"Numer. Math."},{"key":"ref_30","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal Policy Optimization Algorithms. arXiv."},{"key":"ref_31","unstructured":"Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016). Asynchronous Methods for Deep Reinforcement Learning. Proceedings of the 33rd International Conference on Machine Learning, PMLR."},{"key":"ref_32","unstructured":"Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P. (2015). High-Dimensional Continuous Control Using Generalized Advantage Estimation. arXiv."},{"key":"ref_33","unstructured":"Loshchilov, I., and Hutter, F. (2017). Decoupled Weight Decay Regularization. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Shi, Y., Huang, Z., Feng, S., Zhong, H., Wang, W., and Sun, Y. (2020). Masked Label Prediction: Unified Message Passing Model for Semi-Supervised Classification. arXiv.","DOI":"10.24963\/ijcai.2021\/214"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"80","DOI":"10.2307\/3001968","article-title":"Individual Comparisons by Ranking Methods","volume":"1","author":"Wilcoxon","year":"1945","journal-title":"Biom. Bull."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"McInnes, L., Healy, J., and Melville, J. (2018). UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv.","DOI":"10.21105\/joss.00861"},{"key":"ref_37","first-page":"281","article-title":"Some Methods for Classification and Analysis of Multivariate Observations","volume":"Volume 5.1","author":"MacQueen","year":"1967","journal-title":"Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Statistics"},{"key":"ref_38","unstructured":"M\u00fcllner, D. (2011). Modern Hierarchical, Agglomerative Clustering Algorithms. arXiv."}],"container-title":["Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2079-8954\/14\/6\/659\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T10:57:54Z","timestamp":1780916274000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2079-8954\/14\/6\/659"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,8]]},"references-count":38,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2026,6]]}},"alternative-id":["systems14060659"],"URL":"https:\/\/doi.org\/10.3390\/systems14060659","relation":{},"ISSN":["2079-8954"],"issn-type":[{"value":"2079-8954","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,8]]}}}