{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T16:21:56Z","timestamp":1775665316702,"version":"3.50.1"},"reference-count":57,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T00:00:00Z","timestamp":1775606400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Natural Science Research Key Program of the Anhui Provincial Department of Education","award":["2024AH050611"],"award-info":[{"award-number":["2024AH050611"]}]},{"name":"Natural Science Research Key Program of the Anhui Provincial Department of Education","award":["2025AHGXZK30698"],"award-info":[{"award-number":["2025AHGXZK30698"]}]},{"name":"Natural Science Research Key Program of the Anhui Provincial Department of Education","award":["2024AH050608"],"award-info":[{"award-number":["2024AH050608"]}]},{"name":"Anhui Xinhua University\u2019s Quality Engineering Project","award":["2024jy011"],"award-info":[{"award-number":["2024jy011"]}]},{"name":"Anhui Xinhua University\u2019s Quality Engineering Project","award":["2024hhkcx01"],"award-info":[{"award-number":["2024hhkcx01"]}]},{"name":"Anhui Xinhua University\u2019s Quality Engineering Project","award":["2024jy016"],"award-info":[{"award-number":["2024jy016"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>With the rapid development of e-commerce and smart manufacturing, automated warehouse systems have become critical infrastructure for modern logistics. In China\u2019s vast market, the dynamic scheduling of Rail-Guided Vehicles (RGVs) faces significant challenges due to complex task uncertainties, hierarchical supply chain structures, and real-time collision avoidance requirements. Traditional rule-based methods and static optimization models often fail to adapt to such dynamic environments. To address these issues, this paper proposes a novel hybrid deep reinforcement learning framework integrating a Dynamic Graph Neural Network (DGNN) and a Transformer model. The DGNN captures the spatiotemporal dependencies of the warehouse network topology, while the Transformer mechanism enhances long-range feature extraction for task prioritization. Furthermore, we design a centralized Deep Q-network (DQN) framework with parameterized action spaces to coordinate multiple RGVs collaboratively. While the system manages multiple physical vehicles, the learning architecture employs a single-agent global scheduler to avoid the non-stationarity issues inherent in multi-agent reinforcement learning. Experimental results based on real-world data from a large-scale electronics manufacturing warehouse demonstrate that our method reduces average task completion time by 18.5% and improves system throughput by 22.3% compared to state-of-the-art baselines. The proposed approach demonstrates potential for intelligent warehouse management in dynamic industrial scenarios.<\/jats:p>","DOI":"10.3390\/a19040289","type":"journal-article","created":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T15:26:39Z","timestamp":1775661999000},"page":"289","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Intelligent Scheduling of Rail-Guided Shuttle Cars via Deep Reinforcement Learning Integrating Dynamic Graph Neural Networks and Transformer Model"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-1143-4728","authenticated-orcid":false,"given":"Fang","family":"Zhu","sequence":"first","affiliation":[{"name":"Department of Mathematics, Ministry of General Education, Anhui Xinhua University, Hefei 230088, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-9089-3994","authenticated-orcid":false,"given":"Shanshan","family":"Peng","sequence":"additional","affiliation":[{"name":"School of Business, Anhui Xinhua University, Hefei 230088, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,4,8]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"892","DOI":"10.1016\/j.jclepro.2019.04.134","article-title":"Energy-cyber-physical system enabled management for energy-intensive manufacturing industries","volume":"226","author":"Ma","year":"2019","journal-title":"J. Clean. Prod."},{"key":"ref_2","first-page":"159","article-title":"Research on the Key Technology Scheme of Intelligent Sorting Winder Based on Mechanical Automation","volume":"9","author":"Zhou","year":"2021","journal-title":"Instrum. Equip."},{"key":"ref_3","first-page":"555684","article-title":"Advancements in Automated Storage and Retrieval Systems: A Comprehensive Review","volume":"6","author":"Khan","year":"2025","journal-title":"Robot. Autom. Eng. J."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"693","DOI":"10.1631\/FITEE.2000156","article-title":"Warehouse automation by logistic robotic networks: A cyber-physical control approach","volume":"21","author":"Cai","year":"2020","journal-title":"Front. Inf. Technol. Electron. Eng."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"801","DOI":"10.1080\/00207543.2018.1483587","article-title":"Scheduling of smart intra-factory material supply operations using mobile robots","volume":"57","author":"Kousi","year":"2019","journal-title":"Int. J. Prod. Res."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"101849","DOI":"10.1016\/j.rcim.2019.101849","article-title":"A proactive material handling method for CPS enabled shop-floor","volume":"61","author":"Wang","year":"2020","journal-title":"Robot. Comput.-Integr. Manuf."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"3899","DOI":"10.1007\/s00170-019-03941-6","article-title":"Augmented reality application to support the assembly of highly customized products and to adapt to production re-scheduling","volume":"105","author":"Mourtzis","year":"2019","journal-title":"Int. J. Adv. Manuf. Technol."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Mousavi, M., Yap, H.J., Musa, S.N., Tahriri, F., and Dawal, S.Z. (2017). Multi-objective AGV scheduling in an FMS using a hybrid of genetic algorithm and particle swarm optimization. PLoS ONE, 12.","DOI":"10.1371\/journal.pone.0169817"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"109111","DOI":"10.1016\/j.compchemeng.2025.109111","article-title":"Leveraging graph neural networks and multi-agent reinforcement learning for inventory control in supply chains","volume":"199","author":"Kotecha","year":"2025","journal-title":"Comput. Chem. Eng."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Verma, V., Qu, M., Kawaguchi, K., Lamb, A., Bengio, Y., Kannala, J., and Tang, J. (2021, January 2\u20139). Graph Mix: Improved training of GNNs for semi-supervised learning. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual.","DOI":"10.1609\/aaai.v35i11.17203"},{"key":"ref_11","unstructured":"Siems, J., Schambach, M., Schulze, S., and Otterbach, J.S. (2023, January 23\u201329). Interpretable reinforcement learning via neural additive models for inventory management. Proceedings of the 40th International Conference on Machine Learning, Honolulu, HI, USA."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"108518","DOI":"10.1016\/j.compchemeng.2023.108518","article-title":"Constrained continuous-action reinforcement learning for supply chain inventory management","volume":"181","author":"Burtea","year":"2024","journal-title":"Comput. Chem. Eng."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Rangel-Martinez, D., and Ricardez-Sandoval, L. (2023, January 3\u20136). Application of reinforcement learning with recurrent neural networks for optimal scheduling of flow-shop systems under uncertainty. Proceedings of the 2023 9th International Conference on Control, Decision and Information Technologies (CoDIT), Rome, Italy.","DOI":"10.1109\/CoDIT58514.2023.10284480"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"108748","DOI":"10.1016\/j.compchemeng.2024.108748","article-title":"A recurrent reinforcement learning strategy for optimal scheduling of partially observable job-shop and flow-shop batch chemical plants under uncertainty","volume":"188","year":"2024","journal-title":"Comput. Chem. Eng."},{"key":"ref_15","unstructured":"Zhang, R., Fu, H., Miao, Y., and Konidaris, G. (2024). Model-based reinforcement learning for parameterized action spaces. arXiv."},{"key":"ref_16","unstructured":"Stranieri, F., and Stella, F. (2022). A deep reinforcement learning approach to supply chain inventory management. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1836","DOI":"10.1177\/10591478241305863","article-title":"Multi-agent deep reinforcement learning for multi-echelon inventory management","volume":"34","author":"Liu","year":"2025","journal-title":"Prod. Oper. Manag."},{"key":"ref_18","unstructured":"Liu, Q., Chung, A., Szepesv\u00e1ri, C., and Jin, C. (2022, January 2\u20135). When is partially observable reinforcement learning not scary?. Proceedings of the Conference on Learning Theory, London, UK."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Khirwar, M., Gurumoorthy, K.S., Jain, A.A., and Manchenahally, S. (2023). Cooperative multi-agent reinforcement learning for inventory management. arXiv.","DOI":"10.1007\/978-3-031-43427-3_37"},{"key":"ref_20","unstructured":"Sultana, N.N., Meisheri, H., Baniwal, V., Nath, S., Ravindran, B., and Khadilkar, H. (2020). Reinforcement learning for multi-product multi-node inventory management in supply chains. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"108783","DOI":"10.1016\/j.compchemeng.2024.108783","article-title":"An analysis of multi-agent reinforcement learning for decentralized inventory control systems","volume":"188","author":"Mousa","year":"2024","journal-title":"Comput. Chem. Eng."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Tampuu, A., Matiisen, T., Kodelja, D., Kuzovkin, I., Korjus, K., Aru, J., Aru, J., and Vicente, R. (2017). Multiagent cooperation and competition with deep reinforcement learning. PLoS ONE, 12.","DOI":"10.1371\/journal.pone.0172395"},{"key":"ref_23","unstructured":"Nekoei, H., Badrinaaraayanan, A., Sinha, A., Amini, M., Rajendran, J., Mahajan, A., and Chandar, S. (2023, January 22\u201325). Dealing with non-Stationarity in decentralized cooperative multi-agent deep reinforcement learning via multi-timescale learning. Proceedings of the Conference on Lifelong Learning Agents, Montr\u00e9al, QC, Canada."},{"key":"ref_24","first-page":"24611","article-title":"The surprising effectiveness of PPO in cooperative multi-agent games","volume":"Volume 35","author":"Yu","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_25","unstructured":"Nayak, S., Choi, K., Ding, W., Dolan, S., Gopalakrishnan, K., and Balakrishnan, H. (2023, January 23\u201329). Scalable multi-agent reinforcement learning through intelligent information aggregation. Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA."},{"key":"ref_26","unstructured":"Hu, J., Hu, S., and Liao, S.-W. (2021). Policy regularization via noisy advantage values for cooperative multi-agent actor-critic methods. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"104816","DOI":"10.1016\/j.knosys.2019.06.024","article-title":"Dyngraph2vec: Capturing network dynamics using dynamic graph representation learning","volume":"187","author":"Goyal","year":"2020","journal-title":"Knowl.-Based Syst."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"105400","DOI":"10.1016\/j.cor.2021.105400","article-title":"Reinforcement learning for combinatorial optimization: A survey","volume":"134","author":"Mazyavkina","year":"2021","journal-title":"Comput. Oper. Res."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"15051","DOI":"10.1109\/TNNLS.2023.3283523","article-title":"Challenges and opportunities in deep reinforcement learning with graph neural networks: A comprehensive review of algorithms and applications","volume":"35","author":"Munikoti","year":"2024","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_30","unstructured":"Zhang, C., Song, W., Cao, Z., Zhang, J., Tan, P.S., and Chi, X. (2020, January 6\u201312). Learning to dispatch for job shop scheduling via deep reinforcement learning. Proceedings of the 34th Conference on Neural Information Processing Systems, Virtual."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Zhou, L., Yang, Y., Ren, X., Wu, F., and Zhuang, Y. (2018, January 2\u20137). Dynamic network embedding by modeling triadic closure process. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11257"},{"key":"ref_32","unstructured":"Zhang, C., Song, D., Chen, Y., and Feng, X. (February, January 27). A deep neural network for unsupervised anomaly detection and diagnosis in multivariate time series data. Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-19), Honolulu, HI, USA."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Jin, W., Qu, M., Jin, X., and Ren, X. (2020). Recurrent Event Network: Auto-regressive structure inference over temporal knowledge graphs. arXiv.","DOI":"10.18653\/v1\/2020.emnlp-main.541"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Li, Z., Jin, X., Li, W., Guan, S., Guo, J., Shen, H., Wang, Y., and Cheng, X. (2021). Temporal knowledge graph reasoning based on evolutional representation learning. arXiv.","DOI":"10.1145\/3404835.3462963"},{"key":"ref_35","unstructured":"Wang, Y., Chang, Y., Liu, Y., Leskovec, J., and Li, P. (2022). Inductive representation learning in temporal networks via causal anonymous Walks. arXiv."},{"key":"ref_36","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Chen, H., Wang, Y., Guo, T., Xu, C., Deng, Y., Liu, Z., Ma, S., Xu, C., Xu, C., and Gao, W. (2021, January 20\u201325). Pre-trained image processing transformer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01212"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Liang, J., Cao, J., Sun, G., Zhang, K., Van Gool, L., and Timofte, R. (2021, January 11\u201317). SwinIR: Image restoration using Swin Transformer. Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops (ICCVW), Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00210"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., Fu, Y., Feng, J., Xiang, T., and Torr, P.H.S. (2021, January 20\u201325). Rethinking semantic segmentation from a sequence-to-sequence perspective with Transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00681"},{"key":"ref_40","unstructured":"Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J.M., and Luo, P. (2021, January 6\u201314). SegFormer: Simple and efficient design for semantic segmentation with Transformers. Proceedings of the Neural Information Processing Systems, Virtual."},{"key":"ref_41","unstructured":"Lin, L., Fan, H., Zhang, Z.G., Xu, Y., and Ling, H. (December, January 28). SwinTrack: A simple and strong baseline for Transformer tracking. Proceedings of the Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"106749","DOI":"10.1016\/j.cie.2020.106749","article-title":"Deep reinforcement learning based AGVs real-time scheduling with mixed rule for flexible shop floor in industry 4.0","volume":"149","author":"Hu","year":"2020","journal-title":"Comput. Ind. Eng."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"109088","DOI":"10.1016\/j.ijpe.2023.109088","article-title":"Deep reinforcement learning for one-warehouse multi-retailer inventory management","volume":"267","author":"Kaynov","year":"2024","journal-title":"Int. J. Prod. Econ."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"110767","DOI":"10.1016\/j.cie.2024.110767","article-title":"Joint optimization of storage assignment and order batching in robotic mobile fulfillment system with dynamic storage depth and surplus items","volume":"200","author":"Liu","year":"2025","journal-title":"Comput. Ind. Eng."},{"key":"ref_45","first-page":"1349","article-title":"A deep reinforcement learning approach for dynamic inventory optimization","volume":"24","author":"Boute","year":"2022","journal-title":"Manuf. Serv. Oper. Manag."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"5294","DOI":"10.12677\/MOS.2023.126481","article-title":"Multi-AGV path planning in automated warehouses based on multi-agent deep reinforcement learning","volume":"12","author":"Wang","year":"2023","journal-title":"Model. Simul."},{"key":"ref_47","first-page":"102481","article-title":"Deep reinforcement learning for robotic warehouse sorting systems","volume":"79","author":"Chen","year":"2023","journal-title":"Robot. Comput.-Integr. Manuf."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"107221","DOI":"10.1016\/j.cie.2021.107221","article-title":"Solving the online batching problem using deep reinforcement learning","volume":"156","author":"Cals","year":"2021","journal-title":"Comput. Ind. Eng."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"107112","DOI":"10.1016\/j.cor.2025.107112","article-title":"Deep reinforcement learning for dynamic order picking in warehouse operations","volume":"182","author":"Mahmoudinazlou","year":"2025","journal-title":"Comput. Oper. Res."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Krnjaic, A., Steleac, R.D., Thomas, J.D., Papoudakis, G., Sch\u00e4fer, L., To, A.W.K., Lao, K.-H., Cubuktepe, M., Haley, M., and B\u00f6rsting, P. (2023). Scalable Multi-Agent Reinforcement Learning for Warehouse Logistics with Robotic and Human Co-Workers. arXiv.","DOI":"10.1109\/IROS58592.2024.10802813"},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"130184","DOI":"10.1016\/j.neucom.2025.130184","article-title":"Dynamic and prioritized task scheduling of heterogeneous multi-robot systems using deep reinforcement learning","volume":"638","author":"Zhang","year":"2025","journal-title":"Neurocomputing"},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"166","DOI":"10.1007\/s10015-021-00713-y","article-title":"Using sim-to-real transfer learning to close gaps between simulation and real environments through reinforcement learning","volume":"27","author":"Ushida","year":"2022","journal-title":"Artif. Life Robot."},{"key":"ref_53","first-page":"103587","article-title":"Energy-efficient warehouse operations via deep reinforcement learning","volume":"61","author":"Zhao","year":"2024","journal-title":"Sustain. Energy Technol. Assess."},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"128259","DOI":"10.1016\/j.eswa.2025.128259","article-title":"Reinforcement learning-based simulation optimization for an integrated manufacturing-warehouse system: A two-stage approach","volume":"290","author":"Hosseini","year":"2025","journal-title":"Expert Syst. Appl."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Cs\u00e1nyi, G., and Varga, L.Z. (2025). Clustered reverse resumable A* algorithm for warehouse robot path finding. Machines, 13.","DOI":"10.3390\/machines13121127"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Lin, S., Liu, A., and Wang, J. (2023). A dual-layer weight-leader-vicsek model for multi-AGV path planning in warehouse. Biomimetics, 8.","DOI":"10.3390\/biomimetics8070549"},{"key":"ref_57","first-page":"100351","article-title":"Deep reinforcement learning for dynamic vehicle routing with demand and traffic uncertainty","volume":"15","author":"Kadyrov","year":"2025","journal-title":"Oper. Res. Perspect."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/4\/289\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T15:32:33Z","timestamp":1775662353000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/4\/289"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,8]]},"references-count":57,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2026,4]]}},"alternative-id":["a19040289"],"URL":"https:\/\/doi.org\/10.3390\/a19040289","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,8]]}}}