The marketing around machine learning for business forecasting tends to outrun the evidence by a considerable distance. The M4 competition in 2018 - the most rigorous independent benchmarking exercise in forecasting history, covering 100,000 time series - produced results that are worth reading carefully before signing any software contract.
What M4 actually showed
Pure ML methods did not win. The top-performing submission used a hybrid approach combining exponential smoothing with a recurrent neural network, and it outperformed the best pure statistical method by 9.4% on the symmetric mean absolute percentage error metric. That is a real improvement. It is also far smaller than vendor claims typically suggest.
The implementation reality
- ML models require substantially more data to train reliably - typically 3 to 5 years of clean, granular history at minimum.
- Model maintenance costs are higher. A well-specified ARIMA model can run without retraining for 12 to 18 months; most ML forecasting models need quarterly revalidation.
- Interpretability is genuinely limited. When an ML model produces an anomalous forecast, identifying the cause takes significantly longer than with a transparent statistical model.
For organisations with fewer than 200 SKUs or limited data infrastructure, the evidence does not support ML as a priority investment. For those with rich, clean historical data and dedicated analytical capacity, the 9% to 14% accuracy improvement documented in independent studies is worth the complexity.