Forecast Accuracy Benchmarks: What's Good Enough?

There is no forecast accuracy benchmark worth planning against. What to measure instead, at which level, what bias tells you that error does not, and how to read your own number when it moves.

·

8 min read

"Our forecasts are off." That sentence is always true. Forecasts are estimates of an uncertain future and they are never exactly right, so the useful question was never whether yours are wrong. It is whether they are wrong in a way that changes what you should have bought — and whether the number you use to judge them is measuring the forecast or measuring the way you measure.

What is a good forecast accuracy percentage?

There is no figure to give you. Accuracy is not a property of a forecast on its own; it is a property of a forecast, a level, a horizon and a period, and changing any of those four moves the number without anything about the forecast changing at all.

That is not a hedge. It is arithmetic:

Measured atWhy the number looks the way it doesWhat it is actually good for
Total company, monthlyIndividual SKU misses cancel out against each other, so the number looks strongCash planning and board reporting
Category, monthlySome cancellation survives, especially across variants of the same productProduction and co-packer capacity
SKU, monthlyA whole month of demand smooths over the weeks you were shortLong-lead-time buying decisions
SKU × channel, weeklyNothing cancels, so this number is the harshest one you will seeThe purchase order you are about to place
New SKU, first quarterNo history exists to forecast from, so error is high whatever the methodDeciding how small the first order should be

Every row above can describe the same catalogue in the same month. A brand quoting a total-company monthly figure and a brand quoting a SKU-week figure are not comparable, and neither is comparable to a benchmark that does not say which one it measured. Almost none of them do.

Why this page does not quote a benchmark

This page used to publish accuracy ranges by industry and by SKU tier. They have been removed, and it is worth being explicit about why rather than quietly editing them out.

Every range we could find traced back to another article quoting a third, with no study, methodology or sample at the end of the chain. None stated the aggregation level, which is the single input that moves an accuracy number most. Publishing a number we could not source would have been the one mistake that costs a reader real money: they would have measured at SKU-week level, compared against a figure quietly derived at category level, concluded their forecasting was broken, and spent the next quarter buying tools instead of cleaning data.

What demand forecasting is takes the same position on the same question, for the same reason. If you want a target, the honest one is your own number from last quarter.

How do you measure forecast accuracy?

Before comparing anything, fix a method and keep it. All four of the measures below are worth having, and they answer different questions.

Error in units

The plainest measure there is: actual units minus forecast units, per SKU, per period. Positive means you sold more than you planned; negative means you sold less. Keep this one even after you have percentages, because it is the only version denominated in the thing you actually order.

MAPE

Mean absolute percentage error is the most common industry metric. Take the absolute difference between actual and forecast, divide by actual, and average across SKUs. Accuracy is then reported as one hundred minus MAPE.

For example: you forecast 100 units and sold 80. The absolute error is 20 units, which against actual sales of 80 is an error of 25%, and an accuracy of 75%. Do that for every SKU and average it.

MAPE has two known weaknesses and both bite in consumable CPG. It is undefined when actual sales are zero, which happens constantly in a long tail, and it treats a large percentage miss on a ten-unit SKU as equal to the same percentage miss on a ten-thousand-unit one.

Weighted MAPE

Weighted MAPE fixes the second weakness by dividing the sum of the absolute errors by the sum of actual units, rather than averaging the individual percentages. High-volume products then carry the weight they carry in the business. This is usually the more honest headline number for a catalogue with a long tail.

Bias

Bias is the measure most brands do not track and the one that changes decisions. Sum the forecasts, sum the actuals, and look at the direction of the difference. A consistently positive bias means you are over-forecasting and building inventory you did not need. A consistently negative bias means you are under-forecasting and paying for it in stockouts, missed retail commitments and a sales history that now understates real demand.

Deliberate bias is legitimate — running slightly long to protect a retail fill rate is a real choice. Accidental bias is not, and the two look identical on a report that only shows error.

Which level and period should you measure at?

Measure at the level of the decision. If you order by SKU from one supplier, SKU-level accuracy is the number that matters, and a category-level figure is a comfort blanket.

  • Level: the level you place orders at, usually SKU or SKU × channel. Roll up for reporting if you like, but never plan against the roll-up.
  • Period: the period you re-order on. Weekly versus monthly forecasting is the more consequential choice here, and it should be made before you pick a metric rather than after.
  • Horizon: measured at the distance you actually need to see, which is your supplier lead time plus your order cycle. A forecast that is accurate one week out and useless twelve weeks out cannot support a buy that has to be placed twelve weeks ahead.
  • Stockout weeks: decide once whether you exclude them, then never change it mid-comparison. Weeks when you were out of stock record constrained supply as absent demand, and leaving them in makes a forecast look better than it was.

What should you compare the number against?

Your own history, held still. Three comparisons are worth running, and none of them needs an external figure:

  1. The same measure last period. Same level, same horizon, same stockout treatment. If it improved, something you changed worked.
  2. The same measure across SKUs of similar shape. Group products by the shape of their history — steady, seasonal, trending, promotion-driven — and compare within the group. A seasonal variant losing to a hero SKU tells you nothing; a seasonal variant losing to your other seasonal variants tells you where to look.
  3. The decision, not the metric. Did you stock out? Did the buffer get touched and hold? Did you write anything off? Accuracy is a leading indicator of those outcomes, and when the two disagree, the outcomes win.

That third one is the important one. A forecast exists to produce an order, and a page of accuracy metrics that never gets compared to what happened on the shelf is a report rather than a control.

When is a better forecast the wrong thing to buy?

Improving accuracy has real costs — data cleanup, tooling, analyst time — and there is a point where a different lever is cheaper for the same business outcome.

Buffer stock is often cheaper than precision

Sizing safety stock against your measured variability produces the same service outcome as a tighter forecast, and it is usually cheaper to hold a little more of one product than to raise accuracy across a catalogue. The comparison to run is the carrying cost of the extra buffer against the cost of the tools and the time.

Shorter lead times beat better forecasts

If you can reorder in two weeks instead of twelve, you need to see less of the future accurately. Negotiating a shorter lead time, a smaller minimum order quantity or a second supplier reduces how much forecast you have to be right about in the first place.

Some demand is not forecastable

Weather, a competitor going out of stock, a video that lands. No model anticipates these from sales history, which is why there is a floor under error that no amount of modelling gets below. Adjustments for the events you do know about — promotions, launches, retail resets — belong on top of the statistical baseline as separate, labelled changes, so that afterwards you can tell which part of the miss was the model and which part was you. Adjusting a forecast for promotions covers that mechanic.

The uncomfortable case

Sometimes accuracy is fine and the plan is still wrong, because the forecast was accurate at a level nobody buys at, or over a horizon shorter than the lead time. That failure reads as success on every dashboard until the order goes out.

Where Planster fits, and where it doesn't

Planster pulls sales and inventory from 150+ systems, tests several statistical models — Prophet, ETS, SARIMA and linear regression — against each SKU's own history and keeps the one that fits, re-selecting it as the pattern changes. It detects seasonality at SKU level and lets you exclude stockout periods so they stop distorting the fit. That forecast is netted against what is on hand, on order and already committed, and turned into a reorder point, an order quantity and an order-by date per SKU. That is the engine, and the measurement sits on top of it: each SKU's actuals are tracked against its forecast and shown as a variance, a percentage above or below, at the level you buy at.

On top of that, the overnight run re-checks every SKU, order and channel and brings you a ranked list in the morning with the purchase orders already drafted. You open Planster, change what you want changed, and approve. Nothing reaches a supplier until you do. It is flat $1,000/month, built for consumable CPG brands between $10M and $50M — food, beverage, supplements, beauty, household goods.

Where it does not help: it does not publish an industry benchmark for you to measure against, because we do not have a defensible one. It will not forecast a product with no sales history — that first order stays a judgement call. And if you sell a handful of SKUs through one channel from one supplier, the arithmetic on this page is genuinely a spreadsheet job. See the forecasting engine for how model selection works in the product, statistical forecasting methods for the mechanics of each model, and pricing for the whole number on one page.

Common questions

What is a good forecast accuracy percentage?

Forecast accuracy has no benchmark percentage worth planning against, and the ranges published online are almost never traceable to a study anyone can read. Accuracy depends so heavily on the level you measure at, the length of the horizon and the volatility of the category that one figure describes nothing. A brand measuring at total-company level and a brand measuring at SKU-week level can report the same percentage while running completely different businesses. Measure your own, the same way every period, and judge it by whether it is improving.

How do you calculate forecast accuracy?

Forecast accuracy starts with one subtraction per SKU per period: actual units minus forecast units. Divide that difference by actual units to get the error as a proportion, and average the absolute values across the SKUs you care about to get mean absolute percentage error, usually written as MAPE. Accuracy is then reported as one hundred minus MAPE. Do the arithmetic in units rather than revenue, and at the level you actually place orders at, because both choices change the answer more than the formula does.

What is the difference between forecast error and forecast bias?

Forecast error measures how far off a forecast was in either direction, and it sizes the buffer a product needs. Forecast bias measures whether the misses land consistently on the same side of the forecast, and it points at an assumption that is wrong rather than at ordinary variability. A product that comes in above forecast eleven periods out of twelve does not have an error problem, it has a broken assumption, and holding more inventory will not correct it. Bias is the more actionable of the two and the more commonly ignored.

How often should you measure forecast accuracy?

Measure forecast accuracy on the cadence you place orders on, because that is the cadence at which being wrong costs money. Most consumable CPG brands land on a weekly measurement for ordering decisions and a monthly roll-up for cash and capacity planning. What matters more than the frequency is holding the method fixed: the same level, the same horizon and the same treatment of stockout weeks every time, so that a change in the number is a change in the forecast rather than a change in the measurement.

Should you measure forecast accuracy in units or in dollars?

Measure in units at the level you place orders at, because units are what you buy and what runs out. Revenue-weighted accuracy is useful for a finance conversation about cash, but it hides the failure mode that hurts operations: a cheap component stocking out and stopping a build, or a low-price SKU missing badly while an expensive one carries the average. Two products at the same revenue can have entirely different lead times, minimum order quantities and shelf lives, so a dollar figure cannot be ordered against.