Manual demand forecasting in Excel leads to MAPE of 20–30%. For a retailer with 50,000 SKUs, this translates to millions of rubles in losses monthly due to write-offs or stockouts. We automate ML pipelines that reduce error to 8–12% through promo modeling, hierarchical reconciliation, and automatic feature engineering. Regardless of vertical — manufacturing, FMCG, e-commerce, or services — the key problems are the same: data quality, promo effects, and scaling. Our solutions are adapted to business specifics: for retail — short horizons and promo lifts, for manufacturing — long-term forecasts and macro indicators. Reducing MAPE from 20% to 8% cuts write-offs by 15–20 million rubles per year for a chain of 500 stores. Savings on logistics costs reach 10–12 million rubles. A common mistake is ignoring the promo calendar and seasonal patterns, leading to skewed forecasts. We use automatic feature selection and model ensembles to increase robustness.
Vertical-Specific Challenges
Manufacturing
Horizon 3–6 months due to lead time, key KPI is accuracy for raw material planning. Data: historical orders, macro indicators.
FMCG / Retail
Horizon 1–4 weeks, promo lifts up to +300%. Separating baseline demand and incremental is the main difficulty.
E-commerce
Horizon 1–7 days, SKU-level, extreme seasonality (Black Friday). Global models for 100,000+ SKUs are required.
Services
No physical inventory, but capacity (call center operators, servers). Load is forecasted, not goods.
How ML Models Handle Promo Effects?
Promotions are the largest source of error. Decomposing Total Demand = Baseline + Incremental Lift allows modeling the lift separately. Example: a LightGBM regressor for lift takes features like discount_pct, mechanic, display_flag, brand_strength and outputs a coefficient — e.g., 1.85 (+85%). Cross-SKU effects (cannibalization and halo) adjust forecasts across the category via a cross-elasticities matrix. LightGBM outperforms Prophet by 1.15–1.2 times in MAPE for short-term forecasts with promotions.
lift_features = { 'discount_pct': 20.0, 'mechanic': '2+1', 'display_flag': 1, 'leaflet_flag': 0, 'competitor_promo': 0, 'category': 'soft_drinks', 'brand_strength': 0.8, 'seasonality_index': 1.2 } predicted_lift = lift_model.predict([lift_features]) Why Hierarchical Forecasting Is Critical for 10,000 SKUs?
Hierarchy: Total → Category → Brand → SKU → Location. Manual reconciliation of 500,000 forecasts daily is impossible. MinT (Minimum Trace) provides unbiased estimates at all levels. Comparison of methods:
| Method | Accuracy (WMAPE) | Speed | Interpretability |
|---|---|---|---|
| Bottom-up | Medium | High | High |
| Top-down | Low | High | Medium |
| MinT | High | Medium | Low |
| Optimal Combination | High | Low | Low |
We use a hybrid: statistical methods for long horizons and ML for short, aggregated via MinT.
New Product Introduction (NPI)
New SKUs with no history are a common pain. Three approaches:
- Analog-based: forecast based on sales of similar products at launch.
- Attribute-based: regression on characteristics (brand, category, price).
- Bayesian prior: initial forecast = analog, updated with first sales.
| Method | Start Accuracy | Adaptability | Required Data |
|---|---|---|---|
| Analog-based | Medium | Low | History of analogs |
| Attribute-based | Low | Medium | Product characteristics |
| Bayesian prior | High | High | Sales of 1–4 weeks |
Forecasting System Architecture
Full pipeline: from data to forecasts
Data Sources → Feature Engineering → Model Training → Forecast → Activation Data Sources: ├── Internal: ERP sales, WMS, CRM ├── External: macro data, weather, search trends └── Promotional: trade calendar, planned campaigns Feature Engineering (dbt / Spark): ├── Temporal lags: t-1, t-7, t-28, t-52 (weeks) ├── Rolling aggregations: 4w, 13w, 52w ├── Promotional features: lift estimation, channel flags └── External features: weather index, macro indicators Model Training (MLflow): ├── Baseline: Seasonal Naive, ETS ├── Statistical: Prophet, SARIMA ├── ML: LightGBM, DeepAR └── Ensemble: Stacking / Weighted Average Forecast Generation: └── Hierarchical reconciliation → SKU × Location prognoses How to Build a Baseline Forecast for New SKUs?
- Collect data on launches of similar SKUs over the last 2 years.
- Extract attributes: brand, category, price segment, launch season.
- Build a regression model to predict the first 4 weeks of sales.
- Use Bayesian update: adjust the forecast weekly based on actual sales.
This approach gives start accuracy WMAPE 15–18% vs. 30% for naive average.
Forecast Integration and Activation
The system exports forecasts to S&OP (SAP IBP, Anaplan) via API and generates purchase orders for VMI. Accuracy tracking: 1 – WMAPE on a dashboard. We guarantee transparency: model card, SHAP reports, pipeline documentation. Thanks to forecast accuracy, our clients reduce storage costs by 20–30%.
What's Included in the Work (Deliverables)
- Audit of data sources and cleaning.
- Feature engineering pipeline (dbt/Spark).
- Model training and validation (MLflow).
- Architectural documentation and model card.
- Integration with ERP/S&OP via API.
- Team training and 3 months of support.
We will assess your project in 2–3 days — get in touch with us. 5+ years of experience, 30+ implemented demand forecasting systems for retail and manufacturing. Typical timelines: basic system with LightGBM for 1000+ SKUs — 6–8 weeks, full hierarchical system with NPI and reconciliation — 4–6 months. Get a consultation to discuss the details.







