โ† Back to Projects Live Demo โ†—

Hybrid Framework

AI Engineering

Predictive & Adaptive Supply Chain Management

A system that doesn't just forecast, it acts on the forecast. Three CatBoost models predict sales, delivery time, and profit per order; a PPO reinforcement-learning agent, implemented from scratch in NumPy, turns those forecasts into reorder-quantity decisions.

CatBoostPPO / RLPythonTime-series
Hybrid supply chain framework interface

01

Problem

What problem were you solving?

Static forecasting models tell you what's likely to happen, but leave the operational response, what to restock, prioritize, or reroute, to manual rules that don't scale with order volume.

02

Solution

What did you build?

A two-stage pipeline: CatBoost handles accurate multi-metric forecasting, and a PPO reinforcement-learning agent handles what forecasting can't: sequential, reward-driven operational decisions.

03

Key Features

๐Ÿ“ˆ

Forecasting

Three CatBoost models predict sales, delivery time, and profit per order.

๐Ÿ“ฆ

Reorder decisions

PPO agent outputs a continuous reorder quantity from live forecasts.

๐Ÿงฎ

Custom PPO

Policy network, value network, and GAE implemented from scratch in NumPy.

๐ŸŽ›๏ธ

Live simulation

A try-it panel projects outcomes from your own inventory inputs.

04

Architecture

How does it work?

Order Data180,519 historical rows
โ†’
CatBoost3 forecast targets
โ†’
State Vector+ inventory, demand
โ†’
PPO AgentNumPy, from scratch
โ†’
Reorder QuantityContinuous, 0–50 units

PPO was chosen over value-based methods like DQN for stability on a mixed discrete/continuous action space: its clipped objective keeps policy updates from destabilizing training when the reward landscape is noisy.

05

Tech Stack & Tools

ComponentPurpose
CatBoost3 regressors: sales, delivery time, profit
GymnasiumCustom InventoryEnv reinforcement-learning environment
PPO (NumPy)Hand-implemented policy net, value net, GAE, Adam
PythonPipeline and training implementation
Pandas / NumPyData processing
DataCo dataset180,519-row evaluation set

06

Technical Decisions

Why these technologies?

Why CatBoost?

Native handling of categorical features (shipping mode, region, category) without heavy preprocessing, and strong out-of-the-box accuracy on tabular data.

Why PPO over DQN?

More stable on a mixed discrete/continuous action space, with a clipped objective that tolerates a noisy, shaped reward.

Why split forecasting from decisions?

Lets a strong supervised model do what it's good at, while the RL agent focuses purely on sequential decision-making.

Why a shaped reward?

Scoring on service level, cost, and delay reduction, not forecast accuracy alone, keeps the agent optimizing for real operational outcomes.

Why implement PPO from scratch?

Stable-Baselines3 needed a multi-gigabyte PyTorch/CUDA install unreachable in a network-locked build environment, so the policy net, value net, GAE, and Adam were hand-implemented in plain NumPy instead.

07

Results

Measured against a fixed reorder-point baseline, 200 held-out episodes each

1843.0mean reward, PPO agent
1342.7mean reward, fixed reorder-point baseline
+37%improvement over baseline
180,519rows evaluated

The trained agent takes bigger, more deliberate reorder swings rather than the baseline's flat fixed-quantity response, which improves mean reward at the cost of higher variance (std. dev. 1567.4 vs. 1063.2).

Forecasting accuracy per target

ForecastRMSEMAER²
Sales per customer1.560.840.9998
Delivery time (days)1.240.970.417
Profit per order101.6154.22−0.004

Sales is close to deterministic from price, quantity and discount, so the near-perfect Rยฒ is expected rather than suspicious. Delivery time has real, moderate predictability from shipping mode and region. Profit's Rยฒ near zero is an honest result, not a bug: profit per order in this dataset carries enormous unexplained variance, a pattern also noted in other public analyses of the same data.

08

Challenges

What difficulties did you face?

  • Balancing cost, service level and delay in a single reward function
  • Keeping PPO training stable on a mixed action space
  • Validating that forecast errors didn't compound into bad decisions

09

Limitations

What could be better

  • Trained and evaluated on a single historical dataset, not live data
  • Reward weights are hand-tuned rather than learned
  • No online retraining as conditions drift

10

What I Learned

  • Reward shaping is the real design work, more than model architecture
  • Forecast quality directly bounds decision quality
  • A clear baseline is what makes a reward number mean anything

11

Future Improvements

  • Online/continual retraining as new order data arrives
  • Learned reward weighting instead of hand-tuned coefficients
  • Extend the action space to supplier selection