In short: Planning assumptions management means keeping a ledger of the beliefs a plan depends on that are not derived from the model, each with an owner and a date. Scoring every entry next cycle against what actually happened is what turns a register into a mechanism rather than a document. Plenty of misses attributed to forecast error are assumption failures, and without a ledger the two cannot be separated after the fact. Registers die from bloat, so keep the ledger to entries that would change a decision if they turned out to be wrong.
A plan misses by eleven percent. The review asks what went wrong, and the answer that comes back is that the forecast was wrong.
It usually was not. The model performed about as well as it always does. What happened is that the plan assumed a competitor would not respond to the price move, and they did, within three weeks. That belief was stated once in a meeting in January, absorbed into the numbers, and never written down. By the time the plan missed, nobody could reconstruct it, so the miss got attributed to forecasting and a project was launched to improve the model.
The model was not the problem, and the project will not help.
What belongs in the ledger
An assumption is any belief about the future that the plan depends on and that is not derived from the model. Four categories cover most of them.
Commercial. A listing lands in March. A competitor holds price. A customer's promotional calendar looks like last year's. The new pack achieves distribution in sixty percent of stores.
Supply. A supplier recovers capacity by Q3. The line upgrade completes on schedule. Ocean freight normalises after the peak.
Macro. Category growth continues at trend. The currency stays within a band. Input costs move with the published index.
Internal. The marketing budget is approved at the level requested. The salesforce restructure completes without disruption.
Each entry needs four fields and no more: what is assumed, who owns it, when it should be known, and what the plan does if it is wrong. The fourth is the one most often skipped and it is the one that makes the ledger operational rather than documentary.
Field four has a quality test attached. An entry whose response reads "we would revisit the plan" has not been filled in, because that describes what would happen anyway. A usable response names the lever and its size: hold the March promotion and recover roughly nine hundred cases, or bring the second supplier forward at a cost of about forty thousand. When somebody cannot write that sentence, the useful information is that the plan has no prepared response to a belief it depends on, and that is worth surfacing at the point the assumption is offered rather than at the point it fails.
Field one has a test too, and it is stricter than it sounds. Write the assumption so that a stranger holding next quarter's actuals could score it as held or not held without asking anyone what was meant. That rules out most of what gets said in planning meetings. "Distribution builds well in the first half" is unscoreable. "The new pack reaches sixty percent weighted distribution by the end of May" is scoreable, and writing it in that form is usually the moment somebody realises they were assuming something considerably more optimistic than they would have defended out loud.
Scoring is what makes it work
A register of assumptions is a document. What turns it into a mechanism is scoring each entry next cycle against what actually happened.
Three outcomes: held, did not hold, still open. Record which, and for the ones that did not hold, record the plan impact.
Two things follow from doing this consistently.
The first is that people become noticeably more careful about the assumptions they offer. An assumption that will be scored in eight weeks with your name on it gets more thought than one that disappears into a discussion. This effect is larger than any process improvement and it appears within two or three cycles.
The second is that the post-mortem becomes tractable. When a plan misses, you can separate a modelling failure from a wrong assumption, and those call for completely different responses. Confusing them is how organisations end up investing in forecasting technology to solve a problem that was a commercial judgement.
Over time the record also shows which categories of assumption are reliably wrong in your business, which is genuinely useful. If supply recovery dates hold at a rate of about one in three, that is a pattern worth knowing about when the next one is offered.
Four cycles of fourteen entries gives you fifty-six scored beliefs, which is enough to see the pattern by category even though it is nowhere near enough to be precise about any one of them. A record might come out as thirty-four commercial assumptions with twenty-one held, twelve supply recovery dates with four held, and ten macro assumptions with eight held. Those are hit rates of sixty-two percent, thirty-three percent and eighty percent, and the interesting one is the gap between them rather than any individual figure. Commercial and macro beliefs in that business are roughly calibrated. Supply recovery dates are systematically optimistic by a wide margin, and they have been for a year, and nobody had noticed because each one was assessed on its own merits at the time it was offered.
Kahneman and Lovallo described the general version of this in 1993 as the difference between the inside view and the outside view: people forecast a specific case by reasoning about its particulars, and they get closer to the truth by asking how cases of this type have historically turned out. An assumption ledger is a cheap way of building the outside view for your own business, on the specific classes of belief your plans keep resting on. The record is doing the thing an individual estimate cannot do, which is telling you the base rate.
Keeping it small enough to survive
The failure mode of assumption registers is bloat. Every stated belief gets logged, the ledger reaches four hundred entries, nobody reviews it, and it becomes a compliance artefact.
Three constraints keep it useful.
Materiality threshold. Log an assumption only if being wrong about it would change the plan by more than a stated amount. Everything below the line is noise. Set the amount by arithmetic rather than by feel: on a plan of 240 million, one percent is 2.4 million, and that will usually produce a list in the right range for a business unit. Run it backwards to check. Take last cycle's stated beliefs, work out which of them could have moved the plan by 2.4 million, and count. Sixty survivors means the threshold is set too low for the size of the business. Three means it is set too high, or that the plan is resting on fewer things than anyone thinks, which is worth knowing either way.
A cap per cycle. Somewhere between ten and twenty entries for a business unit. Forcing the choice of what makes the list is itself a useful discipline, because it requires someone to decide what the plan actually depends on.
Expiry. Entries resolve or they get closed. An assumption still open after three cycles is a permanent uncertainty, and it belongs in a risk register rather than in a planning ledger.
Two failure modes show up reliably once a ledger has been running for a few cycles, and both have a symptom you can check for.
The first is drift toward the unfalsifiable. Once people understand that entries get scored, some of them start writing entries that cannot lose: the category remains broadly supportive, the supply position stays manageable, competitive intensity is similar to last year. The symptom is a ledger with a hit rate above ninety percent. A register scoring that well is measuring how vague its own entries have become, since a plan whose every stated dependency holds was never uncertain enough to need a ledger. Somewhere around two thirds is what a useful register looks like, because assumptions worth writing down are ones that could plausibly go either way. The fix is the scoreability test above, applied by somebody other than the author.
The second is a swelling open bucket. Entries that were supposed to resolve by a date arrive at that date unresolved, get rolled, and the ledger fills with things nobody can score. The symptom is the share of entries resolved by their own stated date falling below about half, and the open ones being disproportionately the large ones, because those are the beliefs everybody is least keen to settle. Track that share as a number, alongside the hit rate. It measures whether the dates in field three were ever real.
Where it connects
The ledger is most useful when it is wired into two other things rather than maintained alongside them.
Scenarios. Each scenario should name which assumptions it varies. A scenario that is not traceable to a changed assumption is a sensitivity, and one that is traceable becomes a stated position about what might be different.
Gap plays. When the plan falls short of the target, the proposed plays carry their own assumptions, usually optimistic ones about how much a lever will deliver. Logging those alongside the play means that next cycle you can score whether the play delivered what it promised, which is how a business learns which of its levers actually work.
That second connection is the one that produces the most surprising results. Most organisations have a standard set of gap-closing plays that get proposed every cycle, and very few have ever checked which of them delivered.
Once you have the hit rates, they can be applied rather than admired. Suppose a gap plan closes 12,000 cases and rests on three supply recoveries worth 4,000 cases each. If supply recovery dates in your business have held one time in three, the expected contribution of each is about 1,300 cases and the plan's honest expectation is a little over 4,000, against a gap of 12,000. The plan does not close. Nothing in the meeting says so, because each of the three recoveries was presented by someone with a reason to believe it and none of them is individually unreasonable.
That calculation takes a minute and it changes what the room does next, because the response to a gap that closes on paper is to approve it, and the response to a gap that closes to a third is to find two more plays. Apply the haircut in the review rather than afterwards, and state which hit rate you used, so that the argument is about the base rate rather than about optimism.
The limit
An assumption ledger records beliefs that were stated. It cannot record the ones nobody thought to state, and the assumptions that do the most damage are frequently the unexamined ones: that the channel mix stays broadly stable, that the supply base continues to exist in its current shape, that demand is generated by the same mechanism it was last year.
Those are not going to appear in a ledger because nobody experiences them as assumptions. Periodic structured challenge, where somebody is tasked with asking what the plan takes for granted, is the only partial answer, and it works better with an outsider than with the people who built the plan.
The most practical version of that challenge is the premortem, which Gary Klein set out in Harvard Business Review in 2007. The room is told to assume the plan has already failed badly, twelve months on, and each person writes down independently why. Stating the failure as accomplished rather than possible gets people past the social cost of raising a doubt, and it reliably surfaces beliefs that nobody would have volunteered when the question was phrased as a risk. Run it once a year against the annual plan, take the two or three items that appear on several people's lists, and put those in the ledger as entries with dates. That converts the exercise from a discussion into rows, which is the only form in which any of this survives to the next cycle.
There is also a limit on what scoring can fairly do. An assumption can be well reasoned and wrong, and treating the score as a performance measure will produce assumptions chosen for their scoreability rather than their importance. The record is diagnostic, and using it as an appraisal input destroys the thing it was measuring. That distinction should be stated explicitly when the ledger is introduced, because people will assume the opposite.
Start with ten entries in the next cycle and score them in the one after. The mechanism proves itself faster than the argument for it does.