The debate between quantitative models and human judgment in forecasting has been running since at least the 1970s, when Meehl documented that actuarial models outperformed clinical judgment in 19 out of 20 studies he reviewed. Decades later, the pattern holds in business contexts, but with important caveats that are frequently glossed over.
What the meta-analyses show
Armstrong's research across 32 forecasting studies found that structured models reduced error by an average of 16% compared to unaided expert opinion. That is meaningful. But those same studies showed that experts consistently outperformed models during structural breaks - periods when the underlying data relationships changed due to events like supply chain disruptions or regulatory shifts.
Four specific failure modes worth knowing
- Time-series models trained on pre-2020 data were off by more than 30% for 7 out of 10 retail categories during the 2021 demand surge.
- Expert panels anchored to recent performance underestimated recovery speed in 4 out of 6 post-recession periods studied by Makridakis.
- Combination forecasts - averaging model output with adjusted expert estimates - reduced error in 68% of cases reviewed in the M4 competition dataset.
- The M4 competition itself, with 100,000 time series, showed no single method dominated across all horizons.
The evidence suggests neither approach is reliably superior. The more useful question is which failure mode your organisation can better absorb.