Forecast Accuracy Benchmarks: What's Good Enough?
There is no forecast accuracy benchmark worth planning against. What to measure instead, at which level, what bias tells you that error does not, and how to read your own number when it moves.
·
8 min read"Our forecasts are off." That sentence is always true. Forecasts are estimates of an uncertain future and they are never exactly right, so the useful question was never whether yours are wrong. It is whether they are wrong in a way that changes what you should have bought — and whether the number you use to judge them is measuring the forecast or measuring the way you measure.
What is a good forecast accuracy percentage?
There is no figure to give you. Accuracy is not a property of a forecast on its own; it is a property of a forecast, a level, a horizon and a period, and changing any of those four moves the number without anything about the forecast changing at all.
That is not a hedge. It is arithmetic:
| Measured at | Why the number looks the way it does | What it is actually good for |
|---|---|---|
| Total company, monthly | Individual SKU misses cancel out against each other, so the number looks strong | Cash planning and board reporting |
| Category, monthly | Some cancellation survives, especially across variants of the same product | Production and co-packer capacity |
| SKU, monthly | A whole month of demand smooths over the weeks you were short | Long-lead-time buying decisions |
| SKU × channel, weekly | Nothing cancels, so this number is the harshest one you will see | The purchase order you are about to place |
| New SKU, first quarter | No history exists to forecast from, so error is high whatever the method | Deciding how small the first order should be |
Every row above can describe the same catalogue in the same month. A brand quoting a total-company monthly figure and a brand quoting a SKU-week figure are not comparable, and neither is comparable to a benchmark that does not say which one it measured. Almost none of them do.
Why this page does not quote a benchmark
This page used to publish accuracy ranges by industry and by SKU tier. They have been removed, and it is worth being explicit about why rather than quietly editing them out.
Every range we could find traced back to another article quoting a third, with no study, methodology or sample at the end of the chain. None stated the aggregation level, which is the single input that moves an accuracy number most. Publishing a number we could not source would have been the one mistake that costs a reader real money: they would have measured at SKU-week level, compared against a figure quietly derived at category level, concluded their forecasting was broken, and spent the next quarter buying tools instead of cleaning data.
What demand forecasting is takes the same position on the same question, for the same reason. If you want a target, the honest one is your own number from last quarter.
How do you measure forecast accuracy?
Before comparing anything, fix a method and keep it. All four of the measures below are worth having, and they answer different questions.
Error in units
The plainest measure there is: actual units minus forecast units, per SKU, per period. Positive means you sold more than you planned; negative means you sold less. Keep this one even after you have percentages, because it is the only version denominated in the thing you actually order.
MAPE
Mean absolute percentage error is the most common industry metric. Take the absolute difference between actual and forecast, divide by actual, and average across SKUs. Accuracy is then reported as one hundred minus MAPE.
For example: you forecast 100 units and sold 80. The absolute error is 20 units, which against actual sales of 80 is an error of 25%, and an accuracy of 75%. Do that for every SKU and average it.
MAPE has two known weaknesses and both bite in consumable CPG. It is undefined when actual sales are zero, which happens constantly in a long tail, and it treats a large percentage miss on a ten-unit SKU as equal to the same percentage miss on a ten-thousand-unit one.
Weighted MAPE
Weighted MAPE fixes the second weakness by dividing the sum of the absolute errors by the sum of actual units, rather than averaging the individual percentages. High-volume products then carry the weight they carry in the business. This is usually the more honest headline number for a catalogue with a long tail.
Bias
Bias is the measure most brands do not track and the one that changes decisions. Sum the forecasts, sum the actuals, and look at the direction of the difference. A consistently positive bias means you are over-forecasting and building inventory you did not need. A consistently negative bias means you are under-forecasting and paying for it in stockouts, missed retail commitments and a sales history that now understates real demand.
Deliberate bias is legitimate — running slightly long to protect a retail fill rate is a real choice. Accidental bias is not, and the two look identical on a report that only shows error.
Which level and period should you measure at?
Measure at the level of the decision. If you order by SKU from one supplier, SKU-level accuracy is the number that matters, and a category-level figure is a comfort blanket.
- Level: the level you place orders at, usually SKU or SKU × channel. Roll up for reporting if you like, but never plan against the roll-up.
- Period: the period you re-order on. Weekly versus monthly forecasting is the more consequential choice here, and it should be made before you pick a metric rather than after.
- Horizon: measured at the distance you actually need to see, which is your supplier lead time plus your order cycle. A forecast that is accurate one week out and useless twelve weeks out cannot support a buy that has to be placed twelve weeks ahead.
- Stockout weeks: decide once whether you exclude them, then never change it mid-comparison. Weeks when you were out of stock record constrained supply as absent demand, and leaving them in makes a forecast look better than it was.
What should you compare the number against?
Your own history, held still. Three comparisons are worth running, and none of them needs an external figure:
- The same measure last period. Same level, same horizon, same stockout treatment. If it improved, something you changed worked.
- The same measure across SKUs of similar shape. Group products by the shape of their history — steady, seasonal, trending, promotion-driven — and compare within the group. A seasonal variant losing to a hero SKU tells you nothing; a seasonal variant losing to your other seasonal variants tells you where to look.
- The decision, not the metric. Did you stock out? Did the buffer get touched and hold? Did you write anything off? Accuracy is a leading indicator of those outcomes, and when the two disagree, the outcomes win.
That third one is the important one. A forecast exists to produce an order, and a page of accuracy metrics that never gets compared to what happened on the shelf is a report rather than a control.
When is a better forecast the wrong thing to buy?
Improving accuracy has real costs — data cleanup, tooling, analyst time — and there is a point where a different lever is cheaper for the same business outcome.
Buffer stock is often cheaper than precision
Sizing safety stock against your measured variability produces the same service outcome as a tighter forecast, and it is usually cheaper to hold a little more of one product than to raise accuracy across a catalogue. The comparison to run is the carrying cost of the extra buffer against the cost of the tools and the time.
Shorter lead times beat better forecasts
If you can reorder in two weeks instead of twelve, you need to see less of the future accurately. Negotiating a shorter lead time, a smaller minimum order quantity or a second supplier reduces how much forecast you have to be right about in the first place.
Some demand is not forecastable
Weather, a competitor going out of stock, a video that lands. No model anticipates these from sales history, which is why there is a floor under error that no amount of modelling gets below. Adjustments for the events you do know about — promotions, launches, retail resets — belong on top of the statistical baseline as separate, labelled changes, so that afterwards you can tell which part of the miss was the model and which part was you. Adjusting a forecast for promotions covers that mechanic.
The uncomfortable case
Sometimes accuracy is fine and the plan is still wrong, because the forecast was accurate at a level nobody buys at, or over a horizon shorter than the lead time. That failure reads as success on every dashboard until the order goes out.
Where Planster fits, and where it doesn't
Planster pulls sales and inventory from 150+ systems, tests several statistical models — Prophet, ETS, SARIMA and linear regression — against each SKU's own history and keeps the one that fits, re-selecting it as the pattern changes. It detects seasonality at SKU level and lets you exclude stockout periods so they stop distorting the fit. That forecast is netted against what is on hand, on order and already committed, and turned into a reorder point, an order quantity and an order-by date per SKU. That is the engine, and the measurement sits on top of it: each SKU's actuals are tracked against its forecast and shown as a variance, a percentage above or below, at the level you buy at.
On top of that, the overnight run re-checks every SKU, order and channel and brings you a ranked list in the morning with the purchase orders already drafted. You open Planster, change what you want changed, and approve. Nothing reaches a supplier until you do. It is flat $1,000/month, built for consumable CPG brands between $10M and $50M — food, beverage, supplements, beauty, household goods.
Where it does not help: it does not publish an industry benchmark for you to measure against, because we do not have a defensible one. It will not forecast a product with no sales history — that first order stays a judgement call. And if you sell a handful of SKUs through one channel from one supplier, the arithmetic on this page is genuinely a spreadsheet job. See the forecasting engine for how model selection works in the product, statistical forecasting methods for the mechanics of each model, and pricing for the whole number on one page.