In short: Forecast value add takes the forecast as it stood after each stage, scores every one of them against what actually happened, and reports the difference between consecutive stages. A negative stage means that step made the forecast worse, which is a finding rather than a verdict, and the first explanation people reach for is usually wrong. Four measurement mistakes account for most broken value add reports, and each of them can be checked before the ladder is published. Value add tells you whether a stage improved accuracy and says nothing about whether the improvement was worth anything.
The number that leaves an S&OP cycle is produced once and then rewritten five or six times.
A naive baseline gets computed. A statistical model improves on it. A key account team supplies a customer read. Sales adds an overlay. A demand planner overrides three hundred lines because they know something the model does not. A consensus meeting adjusts the total to something the room can live with. Then the plan ships.
Every planning system on the market reports the last number in that chain and calls it the forecast. Almost none of them report whether any given step in the chain made it better. Which means a sales team that pads every number toward quota, or a planner who overrides out of habit rather than evidence, is invisible in every metric the business tracks.
Forecast value add is the audit that makes each step visible. Mike Gilliland at SAS formalised the idea and gave it a name, setting it out at length in The Business Forecasting Deal in 2010, and it has aged into one of the few genuinely settled ideas in demand planning.
The mechanic is simpler than the politics
Take the forecast as it stood after each stage. Score every one of them against what actually happened. Report the difference between consecutive stages.
That is it. The naive forecast is the floor, the statistical output is stage two, and every human touch after that is a stage you evaluate on the same footing as the model.
The output is a ladder that looks like this, using weighted absolute percentage error so that volume counts:
| Stage | Error | Value add |
|---|---|---|
| Naive | 34.1% | baseline |
| Statistical model | 21.6% | +12.5 |
| Customer collaboration input | 20.9% | +0.7 |
| Sales overlay | 24.8% | -3.9 |
| Planner override | 22.2% | +2.6 |
| Consensus adjustment | 23.5% | -1.3 |
Two of those six stages are destroying accuracy. Both survive year after year in most organisations, because nobody has ever put a number next to them.
The arithmetic takes an afternoon once the data is stored properly. Storing it properly is the actual project, and it is where most attempts stall. You need every intermediate version of the forecast, stamped with which stage produced it, retained past the point where the actuals arrive. If your system overwrites the forecast in place, value add is not computable and no amount of analysis will recover it. Check that before you promise anyone a report.
The check is a query. Pick a period that has already closed and count distinct forecast rows per item, period and stage in your history table. One row per item and period, with no stage dimension at all, means the intermediate versions were overwritten and the only thing scoreable is the final number. A stage dimension that exists but carries rows for only the last two stages means somebody turned on snapshotting partway along the chain, which happens more often than you would expect and produces a ladder that starts in the middle of the story.
Reading a negative stage without firing anybody
A negative stage is a finding rather than a verdict, and the first instinct is usually wrong.
When a sales overlay scores negative, the reflex is to conclude that sales is optimistic. Sometimes true. Often the real cause is structural. Sales is being asked to forecast in a system that also sets their quota, which makes the exercise a negotiation with a spreadsheet in the middle rather than a forecast. The overlay is doing precisely what the incentive designed it to do. Changing the incentive fixes more than a training session will.
When a planner override scores negative in aggregate, split it before acting. Override behaviour is almost always bimodal. A small number of large, well-reasoned interventions, usually tied to an event the planner knew about and the model did not, plus a long tail of small habitual adjustments applied out of routine. Score those separately by size and you typically find the large ones add substantial value and the small ones destroy a little, consistently, forever. The intervention that follows is a guardrail rather than a ban: overrides below a threshold require a reason code, above a threshold require nothing at all because those are the ones that work.
The arithmetic on a single month makes the case better than the principle does. Say 3,000 lines were overridden. Split them at 20% of the base forecast: 120 large overrides covering 45% of the overridden volume, and 2,880 small ones covering the other 55%.
Score the two groups separately. The large overrides take weighted absolute percentage error from 24.0% to 18.5%, worth 5.5 points on their share of volume. The small ones take it from 21.0% to 22.3%, costing 1.3 points on theirs. Blended, that is 0.45 times 5.5 less 0.55 times 1.3, or plus 1.8 points, and the stage reports as a positive contributor. Nobody investigates a positive stage.
Underneath that plus 1.8 sits a month of work making the forecast worse. At ninety seconds an override, 2,880 small adjustments come to 72 hours, close to half a planner's month, spent moving numbers in the wrong direction. The 120 large ones took three hours and carried the entire stage.
That result shows up often enough to be worth expecting. Fildes, Goodwin, Lawrence and Nikolopoulos examined judgemental adjustment across four supply chain companies in the International Journal of Forecasting in 2009 and found the same shape: large adjustments made in response to real information improved accuracy, small routine ones frequently did not, and adjustments in the two directions behaved differently from each other.
When the consensus stage scores negative, you have found something more awkward. It usually means the meeting is adjusting the number toward what the business wants rather than toward what it expects, which is a governance problem wearing a forecasting costume. The fix is upstream of the forecast entirely, in whether the room is allowed to hold two numbers, a plan and a target, without pretending they are the same one.
Where teams get the measurement wrong
Four mistakes account for most of the broken value add reports I have seen.
Measuring at the wrong grain. Value add computed at national monthly total will look flat, because aggregation absorbs the damage. Compute it at the grain where the decisions get made, then aggregate the scores upward. If overrides happen at item-location, that is where the scoring belongs. The absorption is arithmetic rather than a reporting quirk: two item-location overrides of plus 400 and minus 380 in the same region net to plus 20 at the regional total, so a stage that moved 780 units in the wrong direction reports as having moved 20.
Using a percentage error metric on a mixed catalogue. Mean absolute percentage error explodes on near-zero actuals and penalises over-forecasting harder than under-forecasting, so it will quietly reward a stage that biases downward. Kolassa and Siemsen make the case for weighted absolute percentage error as the sensible default when volumes vary, and this is exactly the situation it was recommended for.
Ignoring bias. A stage can have identical error before and after while shifting the whole distribution in one direction, and directional drift is what actually fills warehouses. Track bias alongside the error for every stage, or you will approve a change that leaves accuracy untouched and adds three weeks of coverage.
Evaluating on too few periods. One quarter is not enough to separate a stage that adds value from one that got lucky. Run it over at least a year, and treat a small margin as unproven rather than as a result.
One more thing corrupts a ladder quietly, and it is worth ruling out before acting on a negative stage. The scores are computed against actuals, and actuals on a stocked-out item understate demand. A planner or a customer team who raised a forecast ahead of a shortage was right and gets scored as wrong, because the sales that would have proved them right never happened. The pattern to look for is a negative stage whose damage concentrates on item-periods with low availability. Unconstraining the history before scoring is a separate piece of work with its own methods, and the minimum defence is to exclude item-periods where availability fell below a threshold and to say in the report that you did.
What changes once people can see it
The first cycle after publishing a value add ladder is uncomfortable and it is worth pushing through, because three things tend to follow.
Meeting time reallocates. When the statistical stage is contributing twelve points and the consensus stage is contributing minus one, the argument about the consensus number stops being the centre of the cycle. Teams that publish this consistently end up spending review time on exceptions and assumptions rather than on relitigating the total.
Override volume falls without anyone banning overrides. Planners can see which of their own interventions worked. Most people, given a scoreboard, stop doing the thing that scores badly, and they do it faster than a policy would have made them.
And it becomes possible to evaluate an automation claim honestly, which is increasingly the point. If you are being sold an agent that adjusts forecasts continuously, value add is the only instrument that tells you whether it is helping. Insert it as a stage, score it like any other pair of hands, and require it to earn its place against the stage it replaced. That test is more informative than any demonstration, and it is the reason this metric is getting attention again after twenty years of steady, unglamorous usefulness.
The limit
Value add tells you whether a stage improved accuracy. It does not tell you whether the improvement was worth anything.
Those come apart more often than the metric's fans admit. A planner override that improves accuracy by half a point on a slow-moving item with plenty of cover has produced a better number and changed no decision. An override that corrects a forecast on a constrained, high-margin line by the same half point might have prevented a service failure worth a great deal. Same value add score, completely different value.
The honest way to close that gap is to weight the ladder by something economic rather than by units, so that error on items with real exposure counts more. That is harder than it sounds because it requires a cost of error per item that most businesses have never estimated. In the absence of it, weight by margin contribution as a rough proxy and be clear internally that it is a proxy.
The other limit is that value add says nothing about the ceiling. A stage adding two points might be doing excellently against a hard baseline or failing badly against an easy one, and the ladder alone cannot distinguish those. Measure the achievable accuracy on your own history separately, and read the two together.
Start with the storage question, since without retained intermediate forecasts nothing here is computable. Then score one cycle, on one segment, and show the ladder to the people whose stages are on it before you show it to anyone else.