{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T08:48:01Z","timestamp":1782809281070,"version":"3.54.5"},"reference-count":0,"publisher":"ECMS","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026,6,23]]},"abstract":"<jats:p>Modern supply chain management requires balancing economic performance, operational efficiency, and environmental sustainability among dynamic demand uncertainties. Traditional analytical models often fail to capture the stochastic nature of these complex, multi-echelon networks, driving the need for advanced data-driven methodologies. This study investigates the application of Reinforcement Learning to address the gap in jointly optimizing procurement, production, and replenishment decisions while explicitly minimizing carbon emissions and maintaining high service levels. A simulation-based approach was developed to model a three-echelon juice supply chain, comprising two suppliers, one manufacturing factory, and five regional warehouses facing stochastic seasonal demand. Two algorithms, Proximal Policy Optimization and Advantage Actor-Critic, were trained and evaluated within this continuous, multi-objective environment. The computational results reveal a significant divergence in policy stability and performance. The Proximal Policy Optimization algorithm demonstrated superiority, converging to a low-variance strategy that generated approximately 35% more profit and reduced total accumulated operational costs by 33%. Furthermore, it successfully optimized the distribution of goods, achieving a superior customer fulfilment rate while simultaneously operating with lower inventory levels. On the other hand, the Advantage Actor-Critic agent achieved lower carbon emissions.<\/jats:p>","DOI":"10.7148\/2026-0063","type":"proceedings-article","created":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T08:36:46Z","timestamp":1782808606000},"page":"63-69","source":"Crossref","is-referenced-by-count":0,"title":["Towards efficient and sustainable operations in supply chain management: a reinforcement learning approach"],"prefix":"10.7148","author":[{"given":"Djonathan","family":"Quadras","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Romeo","family":"Bandinelli","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Virginia","family":"Fani","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"4144","published-online":{"date-parts":[[2026,6,23]]},"event":{"name":"40th ECMS International Conference on Modelling and Simulation"},"container-title":["ECMS 2026 Proceedings edited by Filippo Sanfilippo, Florenc Demrozi, Fabio Sgarbossa, Mohammad Poursina"],"original-title":[],"deposited":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T08:36:49Z","timestamp":1782808609000},"score":1,"resource":{"primary":{"URL":"http:\/\/www.scs-europe.net\/dlib\/2026\/ecms2026acceptedpapers\/0063_bpmi_ecms2026_0039.pdf"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,23]]},"references-count":0,"URL":"https:\/\/doi.org\/10.7148\/2026-0063","relation":{},"subject":[],"published":{"date-parts":[[2026,6,23]]}}}