{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T03:04:24Z","timestamp":1781924664935,"version":"3.54.5"},"reference-count":40,"publisher":"MDPI AG","issue":"15","license":[{"start":{"date-parts":[[2023,8,3]],"date-time":"2023-08-03T00:00:00Z","timestamp":1691020800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100003052","name":"Ministry of Trade, Industry and Energy","doi-asserted-by":"publisher","award":["20016343"],"award-info":[{"award-number":["20016343"]}],"id":[{"id":"10.13039\/501100003052","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003052","name":"Ministry of Trade, Industry and Energy","doi-asserted-by":"publisher","award":["RS-2022-00155911"],"award-info":[{"award-number":["RS-2022-00155911"]}],"id":[{"id":"10.13039\/501100003052","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Korea government (MSIT)","award":["20016343"],"award-info":[{"award-number":["20016343"]}]},{"name":"Korea government (MSIT)","award":["RS-2022-00155911"],"award-info":[{"award-number":["RS-2022-00155911"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Manufacturing systems need to be resilient and self-organizing to adapt to unexpected disruptions, such as product changes or rapid order, in supply chain changes while increasing the automation level of robotized logistics processes to cope with the lack of human experts. Deep Reinforcement Learning is a potential solution to solve more complex problems by introducing artificial neural networks in Reinforcement Learning. In this paper, a game engine was used for Deep Reinforcement Learning training, which allows visualization of view learning and result processes more intuitively than other tools, as well as a physical engine for a more realistic problem-solving environment. The present research demonstrates that a Deep Reinforcement Learning model can effectively address the real-time sequential 3D bin packing problem by utilizing a game engine to visualize the environment. The results indicate that this approach holds promise for tackling complex logistical challenges in dynamic settings.<\/jats:p>","DOI":"10.3390\/s23156928","type":"journal-article","created":{"date-parts":[[2023,8,3]],"date-time":"2023-08-03T11:23:03Z","timestamp":1691061783000},"page":"6928","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["BoxStacker: Deep Reinforcement Learning for 3D Bin Packing Problem in Virtual Environment of Logistics Systems"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-4420-3350","authenticated-orcid":false,"given":"Shokhikha Amalana","family":"Murdivien","sequence":"first","affiliation":[{"name":"Department of Industrial and Management System Engineering, Kyung Hee University, 1732 Deogyeong-daero, Yongin-si 17104, Republic of Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8040-6144","authenticated-orcid":false,"given":"Jumyung","family":"Um","sequence":"additional","affiliation":[{"name":"Department of Industrial and Management System Engineering, Kyung Hee University, 1732 Deogyeong-daero, Yongin-si 17104, Republic of Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,8,3]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Wang, F., and Hauser, K. (2019, January 20\u201324). Stable bin packing of non-convex 3D objects with a robot manipulator. Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8794049"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"287","DOI":"10.1111\/itor.12094","article-title":"A comparative review of 3D container loading algorithms","volume":"23","author":"Zhao","year":"2016","journal-title":"Int. Trans. Oper. Res."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Tanaka, T., Kaneko, T., Sekine, M., Tangkaratt, V., and Sugiyama, M. (2020\u201324, January 24). Simultaneous Planning for Item Picking and Placing by Deep Reinforcement Learning. Proceedings of the 2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA.","DOI":"10.1109\/IROS45743.2020.9340929"},{"key":"ref_4","unstructured":"Levin, M.S. (2016). Towards bin packing (preliminary problem survey, models with multiset estimates). arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zuo, Q., Liu, X., and Chan, W.K.V. (2022, January 2\u20134). A Constructive Heuristic Algorithm for 3D Bin Packing of Irregular Shaped Items. Proceedings of the INFORMS International Conference on Service Science, Beijing, China.","DOI":"10.1007\/978-3-031-15644-1_29"},{"key":"ref_6","unstructured":"Hu, H., Zhang, X., Yan, X., Wang, L., and Xu, Y. (2017). Solving a new 3d bin packing problem with Deep Reinforcement Learning method. arXiv."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1140","DOI":"10.1016\/j.ejor.2017.10.050","article-title":"A new load balance methodology for container loading problem in road transportation","volume":"266","author":"Ramos","year":"2018","journal-title":"Eur. J. Oper. Res."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Kolhe, P., and Christensen, H. (2010, January 28\u201330). Planning in Logistics: A survey. Proceedings of the 10th Performance Metrics for Intelligent Systems Workshop, Baltimore, MD, USA.","DOI":"10.1145\/2377576.2377586"},{"key":"ref_9","unstructured":"Den Boef, E., Korst, J., Martello, S., Pisinger, D., and Vigo, D. (2003). A Note on Robot-Packable and Orthogonal Variants of the Three-Dimensional Bin Packing Problem, Department of Computer Science, University of Copenhagen. Technical Report 03\/02."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"368","DOI":"10.1287\/ijoc.1070.0250","article-title":"Extreme point-based heuristics for three-dimensional bin packing","volume":"20","author":"Crainic","year":"2008","journal-title":"Informs J. Comput."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"500","DOI":"10.1016\/j.ijpe.2013.04.019","article-title":"A biased random key genetic algorithm for 2D and 3D bin packing problems","volume":"145","author":"Resende","year":"2013","journal-title":"Int. J. Prod. Econ."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"De Andoin, M.G., Osaba, E., Oregi, I., Villar-Rodriguez, E., and Sanz, M. (2022, January 9\u201313). Hybrid quantum-classical heuristic for the bin packing problem. Proceedings of the Genetic and Evolutionary Computation Conference Companion, Boston, MA, USA.","DOI":"10.1145\/3520304.3533986"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"De Andoin, M.G., Oregi, I., Villar-Rodriguez, E., Osaba, E., and Sanz, M. (2022, January 4\u20137). Comparative Benchmark of a Quantum Algorithm for the Bin Packing Problem. Proceedings of the 2022 IEEE Symposium Series on Computational Intelligence (SSCI), Singapore.","DOI":"10.1109\/SSCI51031.2022.10022156"},{"key":"ref_14","unstructured":"Bozhedarov, A., Boev, A., Usmanov, S., Salahov, G., Kiktenko, E., and Fedorov, A. (2023). Quantum and quantum-inspired optimization for solving the minimum bin packing problem. arXiv."},{"key":"ref_15","unstructured":"Ross, P., Schulenburg, S., Mar\u00edn-Bl\u00e4zquez, J.G., and Hart, E. (2002, January 9\u201313). Hyper-heuristics: Learning to combine simple heuristics in bin packing problems. Proceedings of the 4th Annual Conference on Genetic and Evolutionary Computation, New York, NY, USA."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.ejor.2012.12.006","article-title":"Constraints in container loading\u2013A state-of-the-art review","volume":"229","author":"Bortfeldt","year":"2013","journal-title":"Eur. J. Oper. Res."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Kundu, O., Dutta, S., and Kumar, S. (2019, January 14\u201318). Deep-pack: A vision-based 2d online bin packing algorithm with deep Reinforcement Learning. Proceedings of the 2019 28th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), New Delhi, India.","DOI":"10.1109\/RO-MAN46459.2019.8956393"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Le, T.P., Lee, D., and Choi, D. (2021, January 12\u201314). A Deep Reinforcement Learning-based Application Framework for Conveyor Belt-based Pick-and-Place Systems using 6-axis Manipulators under Uncertainty and Real-time Constraints. Proceedings of the 2021 18th International Conference on Ubiquitous Robots (UR), Gangneung-si, Republic of Korea.","DOI":"10.1109\/UR52253.2021.9494631"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"107518","DOI":"10.1016\/j.cie.2021.107518","article-title":"Multi-objective 3D bin packing problem with load balance and product family concerns","volume":"159","author":"Erbayrak","year":"2021","journal-title":"Comput. Ind. Eng."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Jia, J., Shang, H., and Chen, X. (2022, January 15\u201318). Robot Online 3D Bin Packing Strategy Based on Deep Reinforcement Learning and 3D Vision. Proceedings of the 2022 IEEE International Conference on Networking, Sensing and Control (ICNSC), Shanghai, China.","DOI":"10.1109\/ICNSC55942.2022.10004170"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"81","DOI":"10.1109\/LRA.2022.3222996","article-title":"Planning irregular object packing via hierarchical reinforcement learning","volume":"8","author":"Huang","year":"2022","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Oucheikh, R., L\u00f6fstr\u00f6m, T., Ahlberg, E., and Carlsson, L. (2021). Rolling Cargo Management Using a Deep Reinforcement Learning Approach. Logistics, 5.","DOI":"10.3390\/logistics5010010"},{"key":"ref_23","unstructured":"Vinyals, O., Fortunato, M., and Jaitly, N. (2015). Pointer networks. arXiv."},{"key":"ref_24","unstructured":"Bello, I., Pham, H., Le, Q.V., Norouzi, M., and Bengio, S. (2016). Neural combinatorial optimization with Reinforcement Learning. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Bo, A., Lu, J., and Zhao, C. (2022, January 25\u201327). Deep Reinforcement Learning in POMDPs for 3-D palletization problem. Proceedings of the 2022 China Automation Congress (CAC), Xiamen, China.","DOI":"10.1109\/CAC57257.2022.10054950"},{"key":"ref_26","unstructured":"Mower, C., Stouraitis, T., Moura, J., Rauch, C., Yan, L., Behabadi, N.Z., Gienger, M., Vercauteren, T., Bergeles, C., and Vijayakumar, S. (2022, January 14\u201318). ROS-PyBullet Interface: A framework for reliable contact simulation and human-robot interaction. Proceedings of the Conference on Robot Learning, PMLR, Auckland, New Zealand."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1059","DOI":"10.1016\/j.procir.2023.03.149","article-title":"Benchmark of the Physics Engine MuJoCo and Learning-based Parameter Optimization for Contact-rich Assembly Tasks","volume":"119","author":"Salteris","year":"2023","journal-title":"Procedia CIRP"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Kempka, M., Wydmuch, M., Runc, G., Toczek, J., and Ja\u015bkowski, W. (2016, January 20\u201323). Vizdoom: A doom-based ai research platform for visual reinforcement learning. Proceedings of the 2016 IEEE Conference on Computational Intelligence and Games (CIG), Santorini, Greece.","DOI":"10.1109\/CIG.2016.7860433"},{"key":"ref_29","first-page":"1","article-title":"Potentials of game engines for wind power digital twin development: An investigation of the Unreal Engine","volume":"5","author":"Ma","year":"2022","journal-title":"Energy Inform."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"3711","DOI":"10.1007\/s10462-022-10253-x","article-title":"A review of platforms for simulating embodied agents in 3D virtual environments","volume":"56","author":"Kaur","year":"2023","journal-title":"Artif. Intell. Rev."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Wang, S., Mao, Z., Zeng, C., Gong, H., Li, S., and Chen, B. (2010, January 18\u201320). A new method of virtual reality based on Unity3D. Proceedings of the 2010 18th International Conference on Geoinformatics, Beijing, China.","DOI":"10.1109\/GEOINFORMATICS.2010.5567608"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"13677","DOI":"10.1007\/s10489-022-04105-y","article-title":"A review of cooperative multi-agent Deep Reinforcement Learning","volume":"53","author":"Oroojlooy","year":"2023","journal-title":"Appl. Intell."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhao, H., She, Q., Zhu, C., Yang, Y., and Xu, K. (2021, January 2\u20139). Online 3D bin packing with constrained Deep Reinforcement Learning. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual.","DOI":"10.1609\/aaai.v35i1.16155"},{"key":"ref_34","unstructured":"Verma, R., Singhal, A., Khadilkar, H., Basumatary, A., Nayak, S., Singh, H.V., Kumar, S., and Sinha, R. (2020). A generalized Reinforcement Learning algorithm for online 3d bin packing. arXiv."},{"key":"ref_35","unstructured":"Duan, L., Hu, H., Qian, Y., Gong, Y., Zhang, X., Xu, Y., and Wei, J. (2018). A multi-task selected learning approach for solving 3D flexible bin packing problem. arXiv."},{"key":"ref_36","unstructured":"Gleave, A., Dennis, M., Legg, S., Russell, S., and Leike, J. (2020). Quantifying differences in reward functions. arXiv."},{"key":"ref_37","first-page":"20118","article-title":"Explicable reward design for Reinforcement Learning agents","volume":"34","author":"Devidze","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_38","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal policy optimization algorithms. arXiv."},{"key":"ref_39","unstructured":"Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015, January 7\u20139). Trust region policy optimization. Proceedings of the International Conference on Machine Learning, PMLR, Lille, France."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Xu, C., Zhu, R., and Yang, D. (2021, January 17\u201319). Karting racing: A revisit to PPO and SAC algorithm. Proceedings of the 2021 International Conference on Computer Information Science and Artificial Intelligence (CISAI), Kunming, China.","DOI":"10.1109\/CISAI54367.2021.00066"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/15\/6928\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:25:29Z","timestamp":1760127929000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/15\/6928"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,3]]},"references-count":40,"journal-issue":{"issue":"15","published-online":{"date-parts":[[2023,8]]}},"alternative-id":["s23156928"],"URL":"https:\/\/doi.org\/10.3390\/s23156928","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,3]]}}}