{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,28]],"date-time":"2025-11-28T12:34:50Z","timestamp":1764333290582,"version":"build-2065373602"},"reference-count":35,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2023,7,24]],"date-time":"2023-07-24T00:00:00Z","timestamp":1690156800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Beijing Municipal Natural Science Foundation","award":["4224092","2021DJ7704"],"award-info":[{"award-number":["4224092","2021DJ7704"]}]},{"name":"Scientific Research and Technology Development Project","award":["4224092","2021DJ7704"],"award-info":[{"award-number":["4224092","2021DJ7704"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Crude oil resource scheduling is one of the critical issues upstream in the crude oil industry chain. It aims to reduce transportation and inventory costs and avoid alerts of inventory limit violations by formulating reasonable crude oil transportation and inventory strategies. Two main difficulties coexist in this problem: the large problem scale and uncertain supply and demand. Traditional operations research (OR) methods, which rely on forecasting supply and demand, face significant challenges when applied to the complicated and uncertain short-term operational process of the crude oil supply chain. To address these challenges, this paper presents a novel hierarchical optimization framework and proposes a well-designed hierarchical reinforcement learning (HRL) algorithm. Specifically, reinforcement learning (RL), as an upper-level agent, is used to select the operational operators combined by various sub-goals and solving orders, while the lower-level agent finds a viable solution and provides penalty feedback to the upper-level agent based on the chosen operator. Additionally, we deploy a simulator based on real-world data and execute comprehensive experiments. Regarding the alert number, maximum alert penalty, and overall transportation cost, our HRL method outperforms existing OR and two RL algorithms in the majority of time steps.<\/jats:p>","DOI":"10.3390\/a16070354","type":"journal-article","created":{"date-parts":[[2023,7,25]],"date-time":"2023-07-25T01:27:32Z","timestamp":1690248452000},"page":"354","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Hierarchical Reinforcement Learning for Crude Oil Supply Chain Scheduling"],"prefix":"10.3390","volume":"16","author":[{"given":"Nan","family":"Ma","sequence":"first","affiliation":[{"name":"Key Laboratory of Oil & Gas Business Chain Optimization, CNPC, Petrochina Planning and Engineering Institute, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ziyi","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zeyu","family":"Ba","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinran","family":"Li","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ning","family":"Yang","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1458-2828","authenticated-orcid":false,"given":"Xinyi","family":"Yang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Oil & Gas Business Chain Optimization, CNPC, Petrochina Planning and Engineering Institute, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haifeng","family":"Zhang","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"},{"name":"School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,7,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"115618","DOI":"10.1016\/j.ces.2020.115618","article-title":"Simultaneous scheduling of multi-product pipeline distribution and depot inventory management for petroleum refineries","volume":"220","author":"Yu","year":"2020","journal-title":"Chem. Eng. Sci"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"100478","DOI":"10.1016\/j.segan.2021.100478","article-title":"Risk-constrained non-probabilistic scheduling of coordinated power-to-gas conversion facility and natural gas storage in power and gas based energy systems","volume":"26","author":"Ma","year":"2021","journal-title":"Sustain. Energy Grids Netw."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"107827","DOI":"10.1016\/j.ress.2021.107827","article-title":"A taxonomy of railway track maintenance planning and scheduling: A review and research trends","volume":"215","author":"Sedghi","year":"2021","journal-title":"Reliab. Eng. Syst. Saf."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1016\/j.tre.2015.09.004","article-title":"Modeling downstream petroleum supply chain: The importance of multi-mode transportation to strategic planning","volume":"83","author":"Kazemi","year":"2015","journal-title":"Transport. Res. Part E-Logist."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"66","DOI":"10.2516\/ogst\/2018056","article-title":"A robust crude oil supply chain design under uncertain demand and market price: A case study","volume":"73","author":"Beiranvand","year":"2018","journal-title":"Oil Gas Sci. Technol."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Yang, X., Wang, Z., Zhang, H., Ma, N., Yang, N., Liu, H., Zhang, H., and Yang, L. (2022). A Review: Machine Learning for Combinatorial Optimization Problems in Energy Areas. Algorithms, 15.","DOI":"10.3390\/a15060205"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.ejor.2006.12.006","article-title":"A survey on the continuous nonlinear resource allocation problem","volume":"185","author":"Patriksson","year":"2008","journal-title":"Eur. J. Oper. Res."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1109\/MWC.2018.1700099","article-title":"Resource allocation for downlink NOMA systems: Key techniques and open issues","volume":"25","author":"Islam","year":"2018","journal-title":"IEEE Wirel Commun"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"5307","DOI":"10.1007\/s11269-021-03004-0","article-title":"Sustainable water supply and demand management in semi-arid regions: Optimizing water resources allocation based on RCPs scenarios","volume":"35","author":"Mirdashtvan","year":"2021","journal-title":"Water Resour. Manag."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"183","DOI":"10.1016\/j.compchemeng.2014.05.024","article-title":"Scheduling and energy\u2013Industrial challenges and opportunities","volume":"72","author":"Merkert","year":"2015","journal-title":"Comput. Chem. Eng."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"871","DOI":"10.1016\/j.compchemeng.2003.09.018","article-title":"A general modeling framework for the operational planning of petroleum supply chains","volume":"28","author":"Neiro","year":"2004","journal-title":"Comput. Chem. Eng."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"2696","DOI":"10.1039\/C8EE01419A","article-title":"Review of electrical energy storage technologies, materials and systems: Challenges and prospects for large-scale grid storage","volume":"11","year":"2018","journal-title":"Energy Environ. Sci."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"142","DOI":"10.1561\/2200000058","article-title":"Non-convex optimization for machine learning","volume":"10","author":"Jain","year":"2017","journal-title":"Found. Trends Mach. Learn."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"956","DOI":"10.1016\/j.conb.2012.05.008","article-title":"Hierarchical reinforcement learning and decision making","volume":"22","author":"Botvinick","year":"2012","journal-title":"Curr. Opin. Neurobiol."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"855","DOI":"10.1016\/j.compchemeng.2003.09.013","article-title":"Challenges of strategic supply chain planning and modeling","volume":"28","author":"Shapiro","year":"2004","journal-title":"Comput. Chem. Eng."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"643","DOI":"10.1016\/j.cie.2018.11.003","article-title":"Mathematical programming and solution approaches for minimizing tardiness and transportation costs in the supply chain scheduling problem","volume":"127","author":"Tamannaei","year":"2019","journal-title":"Comput. Ind. Eng."},{"key":"ref_17","first-page":"249","article-title":"Two meta-heuristic algorithms for optimizing a multi-objective supply chain scheduling problem in an identical parallel machines environment","volume":"12","author":"Farmand","year":"2021","journal-title":"Int. J. Ind. Eng. Comput."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"106375","DOI":"10.1016\/j.cie.2020.106375","article-title":"Dynamic coordinated scheduling for supply chain under uncertain production time to empower smart production for Industry 3.5","volume":"142","author":"Jamrus","year":"2020","journal-title":"Comput. Ind. Eng."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"103485","DOI":"10.1016\/j.autcon.2020.103485","article-title":"Integrated scheduling of suppliers and multi-project activities for green construction supply chains under uncertainty","volume":"122","author":"RezaHoseini","year":"2021","journal-title":"Autom. Constr."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1106","DOI":"10.1016\/j.ejor.2020.09.052","article-title":"A data-driven optimization approach for multi-period resource allocation in cholera outbreak control","volume":"291","author":"Du","year":"2021","journal-title":"Eur. J. Oper. Res."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"106003","DOI":"10.1016\/j.cie.2019.106003","article-title":"Multi-agent supply chain scheduling problem by considering resource allocation and transportation","volume":"137","author":"Aminzadegan","year":"2019","journal-title":"Comput. Ind. Eng."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"107694","DOI":"10.1016\/j.cie.2021.107694","article-title":"A multi-objective modeling approach to harvesting resource scheduling: Decision support for a more sustainable Thai sugar industry","volume":"162","author":"Jarumaneeroj","year":"2021","journal-title":"Comput. Ind. Eng."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"6103","DOI":"10.1109\/TII.2020.2974875","article-title":"Dynamical resource allocation in edge for trustable Internet-of-Things systems: A reinforcement learning method","volume":"16","author":"Deng","year":"2020","journal-title":"IEEE Trans. Ind. Inf."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"131","DOI":"10.1109\/JSAC.2020.3036962","article-title":"Multi-agent reinforcement learning based resource management in MEC-and UAV-assisted vehicular networks","volume":"39","author":"Peng","year":"2020","journal-title":"IEEE J. Sel. Areas Commun."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"485","DOI":"10.1016\/j.comcom.2019.12.054","article-title":"Intelligent resource allocation management for vehicles network: An A3C learning approach","volume":"151","author":"Chen","year":"2020","journal-title":"Comput. Commun."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"103244","DOI":"10.1016\/j.compind.2020.103244","article-title":"Machine learning for predictive scheduling and resource allocation in large scale manufacturing systems","volume":"120","author":"Morariu","year":"2020","journal-title":"Comput. Ind."},{"key":"ref_27","first-page":"3303","article-title":"Data-efficient hierarchical reinforcement learning","volume":"31","author":"Nachum","year":"2018","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_28","unstructured":"Ma, Q., Ge, S., He, D., Thaker, D., and Drori, I. (2019). Combinatorial optimization by graph pointer networks and hierarchical reinforcement learning. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"297","DOI":"10.1049\/iet-its.2019.0317","article-title":"Hierarchical reinforcement learning for self-driving decision-making without reliance on labelled driving data","volume":"14","author":"Duan","year":"2020","journal-title":"IET Intell. Transp. Syst."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Bacon, P.L., Harb, J., and Precup, D. (2017, January 4\u20139). The option-critic architecture. Proceedings of the AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.10916"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"7095","DOI":"10.1109\/JIOT.2021.3071531","article-title":"Enabling Efficient Scheduling in Large-Scale UAV-Assisted Mobile-Edge Computing via Hierarchical Reinforcement Learning","volume":"9","author":"Ren","year":"2021","journal-title":"IEEE Internet Things J."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"3495","DOI":"10.1109\/TVT.2022.3146439","article-title":"Meta-Hierarchical Reinforcement Learning (MHRL)-based Dynamic Resource Allocation for Dynamic Vehicular Networks","volume":"71","author":"He","year":"2022","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"121703","DOI":"10.1016\/j.energy.2021.121703","article-title":"Hierarchical reinforcement learning based energy management strategy for hybrid electric vehicle","volume":"238","author":"Qi","year":"2022","journal-title":"Energy"},{"key":"ref_34","unstructured":"(2023, June 14). Gurobi Optimization, LLC. Gurobi Optimizer Reference Manual. Available online: http:\/\/www.gurobi.com."},{"key":"ref_35","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/16\/7\/354\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:18:06Z","timestamp":1760127486000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/16\/7\/354"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,24]]},"references-count":35,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2023,7]]}},"alternative-id":["a16070354"],"URL":"https:\/\/doi.org\/10.3390\/a16070354","relation":{},"ISSN":["1999-4893"],"issn-type":[{"type":"electronic","value":"1999-4893"}],"subject":[],"published":{"date-parts":[[2023,7,24]]}}}