In short: A consensus forecasting process works when the room argues about the assumptions underneath the numbers rather than about the numbers themselves. The negotiation is unproductive because participants disagree about premises while arguing over a conclusion, which is how a middle figure emerges that nobody would defend on its own merits. Sales padding is a measurement problem before it is a behaviour problem, so measure the bias by person and by cycle before trying to correct it. The most useful output of the review is a decision, an owner and a dated assumption, and a signed number on its own is the weakest thing the meeting can produce.
The consensus meeting has a statistical forecast, a sales number, a marketing view and a finance expectation on the table. They disagree, sometimes by a lot.
What happens next in most organisations is a negotiation. Sales defends their number, supply argues it is unbuildable, finance points at the target, and after forty minutes a figure emerges that is roughly in the middle and that nobody in the room would defend on its own merits.
Everyone signs off. Nobody plans against it. Supply builds to their own view of what will actually happen, sales continues to work to their quota, and the consensus number goes into a system where it becomes the official record of a decision that was never really made.
Argue about the assumptions
The reason the negotiation is unproductive is that the participants are arguing about a conclusion when they disagree about a premise.
Sales says 12,000. Supply says 9,500. Neither number is an argument. Underneath them sit different beliefs: sales expects a listing at a new account to land in March, supply has heard the listing has slipped, and neither has said so out loud because the conversation is happening in units.
Structuring the meeting around assumptions rather than totals changes what gets discussed. Each contributor states what they are assuming, and the assumptions are the object of debate. When the listing question surfaces, it has an owner and a resolution path, and the number follows from resolving it rather than from splitting a difference.
This is a small change in meeting design and it is the largest single improvement available to most consensus processes.
There is a test for whether your meeting has this problem. Take the minutes or the recording from the last cycle and mark every minute spent on one of three activities: establishing what a number means, arguing about what a number should be, and deciding what to do. The first category is a data problem wearing a meeting's clothes. The second is where the assumptions are hiding, unexamined. Most consensus reviews spend the overwhelming majority of their time in the first two, and the third is where the value was supposed to come from.
The disciplines that hold
Four things separate a process that works from one that produces a signed number nobody uses.
One grid. Everyone works in the same view with the same hierarchy, calendar and units. When sales works in revenue by account, marketing in share by brand and supply in cases by plant, most of the meeting is spent establishing whether two numbers are the same number. A quick way to find out how bad it is: pick one item, ask each function to state next quarter's figure in the other functions' units before the meeting, and see how many of them can. The conversions that nobody can do in advance are the ones being improvised in the room.
Named ownership per input. Each contribution has a person attached rather than a function. A number owned by sales is owned by nobody.
Deadlines that bite. Inputs arriving during the meeting cannot be examined before the meeting, which means the meeting becomes the first time anyone looks at them. A cutoff, enforced, with a default applied when an input is late, changes behaviour within two cycles. The default should be the statistical forecast, since it makes lateness costly to the person who was late rather than to everyone else.
Disputes settled by the record. When two views conflict and the assumptions have been examined without resolving it, the tiebreak should be evidence about who has been right before rather than seniority. Scoring each contributor's historical accuracy makes that possible, and it is the mechanism that changes the character of the meeting most.
Padding is a measurement problem before it is a behaviour problem
Sales forecasts are frequently biased, and the direction depends on the incentive.
Where the forecast feeds a quota, the incentive is to forecast low, so the target is achievable. Where it feeds a supply commitment and being short is painful, the incentive is to forecast high, so stock is available. Where both apply, which is common, the behaviour is inconsistent in a way that looks like noise.
Treating this as an integrity issue is a mistake. People respond to how they are measured, and the response is rational. What changes it is measurement.
The research on human adjustment to statistical forecasts is worth knowing before you design the process around it. Fildes, Goodwin, Lawrence and Nikolopoulos studied four supply chain companies in 2009 and found that adjustments were applied to a large majority of forecasts, that small adjustments were frequently harmful, and that downward adjustments improved accuracy considerably more reliably than upward ones. That last asymmetry is the one to carry into the room, because it says the optimistic revision and the cautious revision are not equally trustworthy and should not be treated as though they are.
Track bias by contributor over time. Not accuracy, bias, because that is the directional signal. A contributor who is consistently fifteen percent high is providing a usable input once you know they are consistently fifteen percent high, and the correction can be applied mechanically while the underlying incentive is dealt with separately.
The size of the prize here surprises people, so work an example. A regional manager submits 1,150, then 1,320, then 1,080, then 1,265 across four quarters, against actuals of 1,000, 1,200, 900 and 1,100. The percentage errors are plus fifteen, plus ten, plus twenty and plus fifteen, giving a mean absolute error of fifteen percent and an average bias of the same fifteen percent, since every error runs the same way. Now divide each submission by 1.15 and rerun it: 1,000, 1,148, 939, 1,100, against the same actuals. The errors become zero, minus four point three, plus four point three and zero, and the mean absolute error falls to a little over two percent.
Nothing about that person's judgement changed. Almost all of what looked like a bad forecaster was a stable offset, and the offset was removable with one division. The residual two percent is their actual contribution, and on that measure they are one of the better inputs in the room.
The condition attached to that correction is that the offset has to be stable, so track the bias of the bias. A contributor whose offset sits between twelve and eighteen percent for eight quarters can be corrected mechanically. One whose offset swings from plus twenty to minus five with the sales cycle cannot, and applying a fixed correction to them makes the forecast worse than leaving it alone.
Publishing the bias record does two things. It stops the pattern being invisible, and it gives the person a reason to change that does not require anyone to accuse them of anything.
There is a failure mode here worth anticipating, because it appears in almost every implementation. Once contributors know a correction factor is being applied, some of them start submitting a number chosen to survive the correction, which means the correction is now chasing a target that moves in response to it. The symptom is distinctive: a contributor's measured bias shrinks over two cycles and then crosses zero and grows in the opposite direction, which is what compensating for a compensation looks like. The protection is to publish the correction openly and to score contributors on their raw submission rather than on the corrected one, so that the honest input is the one that scores well. A correction applied silently produces exactly the behaviour it was meant to remove, one cycle later and harder to see.
What the meeting should produce
A consensus review that ends with a number has produced the least valuable of the available outputs. Three better ones.
The assumptions, with owners and dates. This is the durable record. Next cycle it becomes the scorecard.
The disagreements that were not resolved. A forecast where two functions still disagree is more useful when the disagreement is documented than when it is buried in a compromise. Supply can plan for the range rather than for a midpoint nobody believes.
The plays. Where the consensus falls short of the target, the output is the actions proposed to close the gap, each with a cost and an owner. This is what makes the meeting a decision forum rather than a reporting one.
Where consensus is the wrong tool
Two situations where the process should be shortened rather than improved.
For the segment where the statistical forecast reliably beats every human touch, consensus is a cost with no return. That segment is identifiable by scoring the stages, and once identified it should flow through untouched. Most catalogues have a large one, and freeing the meeting from it leaves time for the items where judgement genuinely helps.
For items with essentially unforecastable demand, consensus produces an agreed number that is no better than any other number. Those items should be managed by policy rather than by forecast, and putting them on a consensus agenda wastes the room's attention on a question that has no answer.
The meeting should cover the middle: items where the model is uncertain and humans have information the model does not. That is a much shorter list than most agendas, and shortening it is usually the single change that makes the cycle sustainable.
The arithmetic of that change is more interesting than it first looks. Reviewing 2,400 active items at an average of forty seconds each consumes about twenty-seven hours of team time a cycle, which is where the review effort in a mid-sized business typically goes. Cut the agenda to the 300 items where judgement plausibly helps and give each of them three minutes instead, and you have spent fifteen hours. The total time roughly halves, and the attention paid to each item that mattered goes up by a factor of four and a half. The saving is the smaller half of the benefit, and the reason to do it is that forty seconds is not enough time to bring any information to bear on anything.
Reviewing 2,400 items at forty seconds each is also, in practice, a description of a process where nobody is really looking. That is worth stating plainly when the reduction gets resisted on coverage grounds, because coverage of that kind is a record of items having been opened.
The limit
None of this survives a structure where the forecast and the target are the same field. If the number the room agrees to is also the number people are measured against, the process will produce a target dressed as a forecast, and every discipline above will be applied faithfully to a fiction.
Separating them is a prerequisite rather than an improvement, and it is not a change a planning function can usually make alone. Where it cannot be made, the honest position is that the consensus process is a commitment ritual rather than a forecasting one, and the planning team should maintain its own unbiased view alongside it for supply purposes. That is a compromise and it is better than pretending the agreed number is a prediction.
Start by adding an assumptions column to the existing grid and requiring one line per input. It changes the first meeting it is used in.