This paper embeds a differentiable convex optimization module into a deep reinforcement learning policy for constrained inventory control. A neural network proposes continuous action targets, a quadratic program projects them onto a relaxed feasible set, and a dual-informed integer mapping restores integrality while preserving feasibility. The method reports an average optimality gap below 1% on small instances, improvements of up to 9.75% over echelon base-stock policies and at least 7.7% over a rolling-horizon multistage stochastic program on larger networks, plus up to 3.22% cost reduction in an ASML industry-scale case study.
No heat snapshots are available in the last 24 hours.