Hybrid Framework
AI EngineeringPredictive & Adaptive Supply Chain Management
A system that doesn't just forecast, it acts on the forecast. Three CatBoost models predict sales, delivery time, and profit per order; a PPO reinforcement-learning agent, implemented from scratch in NumPy, turns those forecasts into reorder-quantity decisions.
01
Problem
What problem were you solving?
Static forecasting models tell you what's likely to happen, but leave the operational response, what to restock, prioritize, or reroute, to manual rules that don't scale with order volume.
02
Solution
What did you build?
A two-stage pipeline: CatBoost handles accurate multi-metric forecasting, and a PPO reinforcement-learning agent handles what forecasting can't: sequential, reward-driven operational decisions.
03
Key Features
Forecasting
Three CatBoost models predict sales, delivery time, and profit per order.
Reorder decisions
PPO agent outputs a continuous reorder quantity from live forecasts.
Custom PPO
Policy network, value network, and GAE implemented from scratch in NumPy.
Live simulation
A try-it panel projects outcomes from your own inventory inputs.
04
Architecture
How does it work?
PPO was chosen over value-based methods like DQN for stability on a mixed discrete/continuous action space: its clipped objective keeps policy updates from destabilizing training when the reward landscape is noisy.
05
Tech Stack & Tools
| Component | Purpose |
|---|---|
| CatBoost | 3 regressors: sales, delivery time, profit |
| Gymnasium | Custom InventoryEnv reinforcement-learning environment |
| PPO (NumPy) | Hand-implemented policy net, value net, GAE, Adam |
| Python | Pipeline and training implementation |
| Pandas / NumPy | Data processing |
| DataCo dataset | 180,519-row evaluation set |
06
Technical Decisions
Why these technologies?
Why CatBoost?
Native handling of categorical features (shipping mode, region, category) without heavy preprocessing, and strong out-of-the-box accuracy on tabular data.
Why PPO over DQN?
More stable on a mixed discrete/continuous action space, with a clipped objective that tolerates a noisy, shaped reward.
Why split forecasting from decisions?
Lets a strong supervised model do what it's good at, while the RL agent focuses purely on sequential decision-making.
Why a shaped reward?
Scoring on service level, cost, and delay reduction, not forecast accuracy alone, keeps the agent optimizing for real operational outcomes.
Why implement PPO from scratch?
Stable-Baselines3 needed a multi-gigabyte PyTorch/CUDA install unreachable in a network-locked build environment, so the policy net, value net, GAE, and Adam were hand-implemented in plain NumPy instead.
07
Results
Measured against a fixed reorder-point baseline, 200 held-out episodes each
The trained agent takes bigger, more deliberate reorder swings rather than the baseline's flat fixed-quantity response, which improves mean reward at the cost of higher variance (std. dev. 1567.4 vs. 1063.2).
Forecasting accuracy per target
| Forecast | RMSE | MAE | R² |
|---|---|---|---|
| Sales per customer | 1.56 | 0.84 | 0.9998 |
| Delivery time (days) | 1.24 | 0.97 | 0.417 |
| Profit per order | 101.61 | 54.22 | −0.004 |
Sales is close to deterministic from price, quantity and discount, so the near-perfect Rยฒ is expected rather than suspicious. Delivery time has real, moderate predictability from shipping mode and region. Profit's Rยฒ near zero is an honest result, not a bug: profit per order in this dataset carries enormous unexplained variance, a pattern also noted in other public analyses of the same data.
08
Challenges
What difficulties did you face?
- Balancing cost, service level and delay in a single reward function
- Keeping PPO training stable on a mixed action space
- Validating that forecast errors didn't compound into bad decisions
09
Limitations
What could be better
- Trained and evaluated on a single historical dataset, not live data
- Reward weights are hand-tuned rather than learned
- No online retraining as conditions drift
10
What I Learned
- Reward shaping is the real design work, more than model architecture
- Forecast quality directly bounds decision quality
- A clear baseline is what makes a reward number mean anything
11
Future Improvements
- Online/continual retraining as new order data arrives
- Learned reward weighting instead of hand-tuned coefficients
- Extend the action space to supplier selection