Carbon-Aware Offline Reinforcement Learning for Joint Inventory Replenishment and Transportation Scheduling in Multi-Echelon Supply Chains

Authors

  • Pan Li University of Hull, UK
  • Menglin Bian University of Missouri, USA
  • Yizhen Lin Stanford University, USA

DOI:

https://doi.org/10.54097/p40ecd94

Keywords:

Offline reinforcement learning, conservative Q-learning, multi-echelon inventory, transportation scheduling, carbon budget, supply-chain control.

Abstract

Joint replenishment and transportation decisions couple service, inventory, vehicle utilization, and carbon emissions over time. Offline reinforcement learning (RL) is attractive when operational logs exist but online exploration is unsafe; however, distribution shift and cumulative carbon constraints make direct offline Q-learning unreliable. This paper proposes budget-conditioned carbon-aware conservative Q-learning (CA-CQL), a discrete-action offline method that combines a remaining-budget state, carbon-shaped conservative value learning, a learned behavior-support prior, a state-dependent carbon shadow price, and a pathwise carbon-feasibility screen. A reproducible supplier–distribution-center–three-retailer benchmark is constructed because public sales datasets lack synchronized replenishment actions, vehicle schedules, loads, and emissions. Transport emissions are calibrated with the official UK Government 2026 vehicle-kilometer conversion factors. Across five independently generated offline datasets and 250 evaluation episodes per policy, CA-CQL achieves an operating cost of 16,865 ± 404, emissions of 14,119 ± 57 kg CO₂e, a 92.18 ± 0.88% fill rate, and zero carbon-cap violations. Relative to a support-regularized CQL baseline, it reduces emissions by 110.8 kg and the violation rate by 13.6 percentage points, with a 1.36% cost increase and a 0.63-point fill-rate decrease. Ablations show that the behavior prior prevents severe under-replenishment and that the carbon screen is necessary for hard pathwise compliance. The results establish a transparent compliance–service trade-off rather than claiming universal dominance.

Downloads

Download data is not yet available.

References

[1] Clark, A. J., & Scarf, H. (1960). Optimal policies for a multi echelon inventory problem. Management Science, 6(4), 475–490. https://doi.org/10.1287/mnsc.6.4.475

[2] Campbell, A. M., Clarke, L., Kleywegt, A. J., & Savelsbergh, M. W. P. (1998). The inventory routing problem. In T. G. Crainic & G. Laporte (Eds.), Fleet management and logistics (pp. 95–112). Springer. https://doi.org/10.1007/978 1 4615 5755 5_4

[3] Trudeau, P., & Dror, M. (1992). Stochastic inventory routing: Route design with stockouts and route failures. Transportation Science, 26(3), 171–184. https://doi.org/10.1287/trsc.26.3.171

[4] Kleywegt, A. J., Nori, V. S., & Savelsbergh, M. W. P. (2002). The stochastic inventory routing problem with direct deliveries. Transportation Science, 36(1), 94–118. https://doi.org/10.1287/trsc.36.1.94

[5] Kleywegt, A. J., Nori, V. S., & Savelsbergh, M. W. P. (2004). Dynamic programming approximations for a stochastic inventory routing problem. Transportation Science, 38(1), 42–70. https://doi.org/10.1287/trsc.1030.0041

[6] Campbell, A. M., & Savelsbergh, M. W. P. (2004). A decomposition approach for the inventory routing problem. Transportation Science, 38(4), 488–502. https://doi.org/10.1287/trsc.1030.0054

[7] Cheng, C., Qi, M., Wang, X., & Zhang, Y. (2016). Multi period inventory routing problem under carbon emission regulations. International Journal of Production Economics, 182, 263–275. https://doi.org/10.1016/j.ijpe.2016.09.001

[8] Bektaş, T., & Laporte, G. (2011). The pollution routing problem. Transportation Research Part B: Methodological, 45(8), 1232–1250. https://doi.org/10.1016/j.trb.2011.02.004

[9] Demir, E., Bektaş, T., & Laporte, G. (2014). A review of recent research on green road freight transportation. European Journal of Operational Research, 237(3), 775–793. https://doi.org/10.1016/j.ejor.2013.12.033

[10] World Resources Institute, & World Business Council for Sustainable Development. (2011). Corporate value chain (Scope 3) accounting and reporting standard.

[11] Smart Freight Centre. (2023). Global logistics emissions council framework for logistics emissions accounting and reporting (Version 3.0).

[12] UK Department for Energy Security and Net Zero, & Department for Environment, Food & Rural Affairs. (2026). Greenhouse gas reporting: Conversion factors 2026.

[13] Altman, E. (1999). Constrained Markov decision processes. CRC Press.

[14] Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.

[15] Achiam, J., Held, D., Tamar, A., & Abbeel, P. (2017). Constrained policy optimization. In Proceedings of the 34th International Conference on Machine Learning (pp. 22–31).

[16] Fujimoto, S., Meger, D., & Precup, D. (2019). Off policy deep reinforcement learning without exploration. In Proceedings of the 36th International Conference on Machine Learning (pp. 2052–2062).

[17] Kumar, A., Fu, J., Soh, M., Tucker, G., & Levine, S. (2019). Stabilizing off policy Q learning via bootstrapping error reduction. In Advances in Neural Information Processing Systems 32.

[18] Kumar, A., Zhou, A., Tucker, G., & Levine, S. (2020). Conservative Q learning for offline reinforcement learning. In Advances in Neural Information Processing Systems 33 (pp. 1179–1191).

[19] Fu, J., Kumar, A., Nachum, O., Tucker, G., & Levine, S. (2021). D4RL: Datasets for deep data driven reinforcement learning. In Advances in Neural Information Processing Systems 34 (pp. 4933–4946).

Downloads

Published

10-09-2026

How to Cite

Li, P., Bian , M., & Lin, Y. (2026). Carbon-Aware Offline Reinforcement Learning for Joint Inventory Replenishment and Transportation Scheduling in Multi-Echelon Supply Chains. Highlights in Business, Economics and Management, 69, 251-262. https://doi.org/10.54097/p40ecd94