Illustration, not data
One forecast, three different numbers
- 1The figure you beat nine times in ten
- 2Midpoint: beaten half the time
- 3Average outcome
Every forecast is a range: pick the point each decision needs from what a miss costs, and judge forecasters by whether their ranges held as often as they claimed.
What it means for the business
Ask what the number is. The average? The midpoint, which you beat half the time? Or the figure you beat nine times in ten? A hiring plan, a budget and a sales plan can each need a different point of the same forecast. They should not be three different forecasts.
Price the miss. For each use, write down what a forecast that turns out too high costs and what one that turns out too low costs. If too high costs nine times as much, as it can when you hire against it, use the figure you beat nine times in ten. If too low is the expensive one, as when running out of stock loses sales, go above the midpoint by the same rule. If the two costs are equal, use the midpoint.
Illustration, not data
The cost of a miss decides which point of the range to use
Beaten 9 times in 10
Midpoint: beaten half the time
Beaten 1 time in 10
Keep every forecast with its range. Being close once is luck. File every forecast together with the range given at the time, and a year later you can see whose ranges held as often as they claimed, and whose were narrowest while doing so. That is a fair measure, and it costs nothing but filing.
Illustration, not data
A range should hold as often as it claims, and be as narrow as it can while doing so
- Range given at the time (claimed: 9 in 10)
- Outcome inside the range
- Outcome outside the range
Watch the width. When conditions change, the range around a good forecast widens until it holds again. A forecast whose range looked the same through a crisis was not telling you how much less it knew.
Illustration, not data
When conditions change, a range that learns from its misses gets wider
Adapts: the range learns from its misses
1 of 12 outcomes fell outside the range
Never adapts: the range stays the same
6 of 12 outcomes fell outside the range
- Range given at the time
- Outcome inside
- Outcome outside
Three questions to ask your data team:
- Which point of the range is this number, and what is the figure we beat nine times in ten?
- For each decision we make from it, what does a forecast that is too high cost us, and one that is too low?
- How did our past forecasts do against the ranges they gave, over as long a record as we have?
The theory, for the technical reader
Which point: the cost of a miss decides
When a forecast that is too low costs u per unit and one that is too high costs o per unit, the best single number to report is the quantile at u ÷ (u + o): the figure the outcome falls below that share of the time. With o nine times u that is the 10% quantile, the figure you beat nine times in ten; with equal costs it is the median. This is old decision theory, the newsvendor's critical ratio (Arrow, Harris and Marschak, 1951), and quantile regression (Koenker and Bassett, 1978) estimates such points directly. Gneiting showed the other side in 2011: a point forecast can only be judged fairly when the scoring rule matches the point it was asked for. Score a median forecast by squared error and the comparison can reward the wrong forecaster. The peak of the curve, the most likely outcome, is missing from the list on purpose: for a continuous quantity no scoring rule makes the most likely value the best number to report (Heinrich, 2014).
Calibration and sharpness
A forecast range has two separate qualities. Calibration means that a range claimed to hold nine times in ten holds nine times in ten. Sharpness means the range is narrow enough to act on. The aim, in Gneiting, Balabdaoui and Raftery's phrase, is to maximise sharpness subject to calibration. Neither is any good alone: a forecaster who always gives the usual spread of past months, whatever is known about next month, can be calibrated and still tell you nothing new. Sharpness can be read off the forecasts themselves; calibration only from a record of forecasts and what then happened. Twelve forecasts are enough to catch a range that is badly wrong, not to tell a range that holds 85% of the time from one that holds 90%, so the longer the record, the better.
Ranges you can trust without trusting the model
A method called conformal prediction takes the record of how far off past forecasts were and sets the range wide enough to have held nine times in ten over that record. The guarantee is then about the range, not the model: whatever produced the forecast, the range holds on average at least nine times in ten, as long as the coming months are like the past ones. A growing or seasonal business often is not, which is why an adaptive version exists that widens the range as misses pile up and keeps the long-run rate under shifting conditions (Gibbs and Candès, 2021). The nine in ten also holds across all forecasts together, not for each group: the range can hold nine times in ten overall and only seven in ten for your largest customers. Setting the range separately for each group that matters fixes that; a guarantee for every single case is impossible without further assumptions.
Averaging forecasts
The average of several reasonable forecasts is hard to beat. Fifty years of studies across economics, weather, demand and more say so, starting with Bates and Granger in 1969. Different forecasts make different mistakes, and in an average the mistakes partly cancel. For the usual measures of error the average is never further off than the forecasts are on average, and when two of them miss in opposite directions it is often closer than both. Weighting the forecasts that have been more accurate should do better still, but the weights are usually estimated from so little history that the estimate is mostly noise, and the plain average tends to win; this is the forecast combination puzzle (Smith and Wallis, 2009). It is a tendency, not a law: in the M4 competition a learned weighting came second overall and beat the competition's simple-average benchmark (Montero-Manso and colleagues, 2020). One condition matters in a business: the forecasts have to be made in good faith. A number shaped by a bonus or a quota is a target, and averaging it in imports its bias.
Illustration, not data
Averaging two forecasts is never further off than the two are on average, and often closer than both
Forecasts that have to add up
Regions sum to the company; months sum to the year. If each region and the total are forecast separately, the regional numbers rarely sum exactly to the total. A common fix is to scale the regions at the end so that they do, which gives every region the same correction, whether or not it was the one that was off. The total shows patterns that are lost in noise at the regional level, and each region shows local patterns that wash out in the total, so a better way is to combine the levels according to how reliable each has been (Wickramasuriya, Athanasopoulos and Hyndman, 2019). The combined forecasts add up by construction and are on average more accurate taken together; in practice most levels improve, although an individual series can get worse. With little history, the weights are kept simple, for the same reason a plain average often beats fitted weights.
Illustration, not data
Regions forecast on their own rarely add up to the company forecast
The usual fix: scale the regions
Each region is cut by about 5% so the sum is 100, including the regions that were right.The better way: combine the levels
Each level counts by how reliable it has been. The result adds up, and each level borrows what the others show.Sources
- Arrow, K. J., Harris, T. and Marschak, J. (1951). Optimal inventory policy. Econometrica 19, 250–272.
- Koenker, R. and Bassett, G. (1978). Regression quantiles. Econometrica 46, 33–50.
- Gneiting, T. (2011). Making and evaluating point forecasts. JASA 106(494), 746–762.
- Heinrich, C. (2014). The mode functional is not elicitable. Biometrika 101(1), 245–251.
- Gneiting, T., Balabdaoui, F. and Raftery, A. E. (2007). Probabilistic forecasts, calibration and sharpness. JRSS B 69(2), 243–268.
- Vovk, V., Gammerman, A. and Shafer, G. (2005). Algorithmic Learning in a Random World. Springer.
- Gibbs, I. and Candès, E. (2021). Adaptive conformal inference under distribution shift. NeurIPS.
- Bates, J. M. and Granger, C. W. J. (1969). The combination of forecasts. Operational Research Quarterly 20(4), 451–468.
- Larrick, R. P. and Soll, J. B. (2006). Intuitions about combining opinions: misappreciation of the averaging principle. Management Science 52(1), 111–127.
- Smith, J. and Wallis, K. F. (2009). A simple explanation of the forecast combination puzzle. Oxford Bulletin of Economics and Statistics 71(3), 331–355.
- Makridakis, S., Spiliotis, E. and Assimakopoulos, V. (2020). The M4 Competition: 100,000 time series and 61 forecasting methods. International Journal of Forecasting 36(1), 54–74.
- Wang, X., Hyndman, R. J., Li, F. and Kang, Y. (2023). Forecast combinations: an over 50-year review. International Journal of Forecasting 39(4), 1518–1547.
- Montero-Manso, P., Athanasopoulos, G., Hyndman, R. J. and Talagala, T. S. (2020). FFORMA: feature-based forecast model averaging. International Journal of Forecasting 36(1), 86–92.
- Wickramasuriya, S. L., Athanasopoulos, G. and Hyndman, R. J. (2019). Optimal forecast reconciliation for hierarchical and grouped time series through trace minimization. JASA 114(526), 804–819.
- Hyndman, R. J. and Athanasopoulos, G. (2021). Forecasting: Principles and Practice, 3rd ed. OTexts.
Mathias Lau Nielsen
Freelance data and AI engineer
I build and fix data platforms and set up AI coding agents for development teams. Technical responsibility for the whole data platform at two companies.
