In short: Collaborative planning, forecasting and replenishment needs four parts to function: a shared forecast at a granularity both sides plan at, agreed exception thresholds, a named resolution process with a rule for deadlock, and a commitment about what the agreed number obliges each party to do. Most programmes build the data exchange and stop, because the remaining three are commercial concessions rather than process design. The unresolved question under a dead exception list is whose forecast wins when the two disagree, and sharing data does not answer it. Checking a quarter of flagged variances against what was actually produced and actually ordered will tell you in an afternoon whether the arrangement changes any decision.
There is usually a shared view somewhere. Your forecast for the customer in one column, theirs in the next, a variance column between them that turns amber past a threshold somebody set during implementation and nobody has touched since. It refreshes on Monday mornings. Ask either side to name a decision that came out differently last quarter because of it, and the answer takes a while to arrive.
The method has been written down and publicly available for a long time. The Voluntary Interindustry Commerce Standards association published the first voluntary guidelines in 1998, revised them in 2004 into a shorter model built around four collaboration activities, and the material moved to GS1 US when the two organisations merged in 2012. That is nearly thirty years of a documented, standardised way of working that most trading relationships still do not run. The reasons have very little to do with software.
What the arrangement actually requires
Strip the guidelines back and a working arrangement has four parts. All four have to be present for the thing to function.
A shared forecast at an agreed granularity. One number for a defined scope, at a level of detail both sides can act on. Weekly, by item, by distribution centre is the usual shape in grocery. The level matters, because a forecast agreed at a level neither party plans at is a document rather than an input.
Agreed exception thresholds. A rule saying which variances are worth a human being. Absolute and percentage together, since either one alone misbehaves at an end of the volume range.
A resolution process. Named people on both sides, a clock, and a rule for what happens when they cannot agree. The last part is the one that gets skipped, and it decides whether any of the rest has teeth.
A commitment about what the agreed number does. The supplier plans production and stock against it, the customer orders against it, and both accept a consequence for departing from it without notice.
Most implementations get the first, sometimes the second, and stop. Infrastructure to move data between two companies is straightforward now and it obliges nobody. The other three are commercial concessions dressed as process design, and they get negotiated by people who did not realise they were negotiating.
Where it stalls
The visible symptom is an exception list that nobody works. It comes in two versions. Thresholds set wide enough that almost nothing triggers, which everyone quietly prefers because the alternative is a phone call. Or thresholds set at implementation defaults, so several hundred items flag every week, the list becomes unworkable within a month, and it gets ignored in a way nobody has to admit to.
Underneath both is the same unresolved question: when the two forecasts disagree, whose number wins. Sharing data does not answer it. A joint decision means one side accepts the other's view on some items and plans against it, which means accepting the cost when that view turns out to be wrong. Nobody wants to sign that without knowing how the cost gets shared, and most programmes never put the question on the table because raising it makes the whole initiative look adversarial at exactly the moment both sides are describing it as a partnership.
There is a diagnostic worth running before you invest in another portal. Take the last quarter and find the items where the two forecasts differed by more than the threshold. For each, check what was actually produced and what was actually ordered. If the supplier produced to their own number and the customer ordered to theirs, the collaboration is a display. This test takes an afternoon and it is more informative than any maturity assessment.
The asymmetry nobody prices
The costs and the benefits of these arrangements land on different balance sheets, and usually not the same one.
In the common grocery version, the retailer contributes point of sale data and a forecast they already produce. That is close to free for them. The supplier contributes analyst time, holds the buffer that absorbs the residual error, funds the service level, and carries the obsolescence. The benefit shows up mostly as on-shelf availability, which converts into retailer sales. Both parties gain something. One of them gains considerably more per unit of effort, and after eighteen months the side doing the work starts asking what it is for.
The same asymmetry runs the other way in categories where the supplier holds the category knowledge and the retailer is buying insight they cannot generate themselves.
Cachon and Fisher's 2000 paper in Management Science is worth reading in this context, because it studied exactly this and found the gains from sharing demand information to be modest next to the gains from shortening lead times and reducing order batch sizes. Sharing data is the cheapest thing on the list and the least valuable in isolation. The value arrives when the physical policy changes, which is the part that costs somebody money.
The practical response is to name the split before you start. Who pays, who gains, and what moves across to balance it. That might be a service term in the trade agreement, or something reciprocal and non-financial: the retailer commits to a longer order lead time, or to giving the supplier the promotional calendar earlier, or to a range decision the supplier can plan against. An arrangement with no reciprocity is one party doing the other party's planning at their own expense, and it has a predictable life expectancy.
A narrow version that survives
The full arrangement across a whole business is a large programme with a long payback and a lot of ways to fail. The narrow version that survives is one category with one customer, collaborating on the promotional calendar and nothing else, and it is small enough that neither side needs a business case to try it.
There are reasons that particular scope holds up when broader ones do not. Promotional volume is usually the largest single source of forecast error and it is the error both parties directly cause, so there is something real to fix. The calendar is a genuinely joint object, since neither side can produce it alone. The data volume is small enough to manage in a spreadsheet if it has to be, which removes the systems dependency from the first year. And the feedback cycle is weeks rather than quarters, so both sides find out whether it worked while they still remember agreeing to it.
What you exchange is narrower than a forecast. Event dates and mechanics, expected volume with a range around it rather than a single figure, store or outlet coverage, and the display or feature support that goes with it. The supplier returns what they can supply against that plan and where the constraints are. Measuring what the promotion actually did to demand is a separate method with its own literature, and it is covered elsewhere on this site.
Set exception thresholds on the event rather than the week. Actual sell-out landing outside the agreed range triggers a call within two working days, with a defined decision at the end of it: adjust the remaining events, release additional stock, or accept the miss and record why. Three or four of those calls in a quarter is a healthy programme. Zero means the ranges are too wide.
Move to base demand after this has run for a year and both sides can point at decisions it changed. Starting with base demand across the full range is how most of these programmes get their reputation.
The measurement that keeps both sides honest
Forecast value add is the right instrument, applied to the joint stage. Score the collaborative forecast against what actually happened, and score both parties' independent forecasts against the same actual over the same period. If the joint number does not beat both, the process is costing two organisations time to produce something worse than either could manage alone, and that happens more often than anyone reports. The methodology itself is covered in its own post; the point here is where you attach it.
Two properties make the scoreboard work. It has to be symmetric, meaning both sides' inputs are scored on the same basis and both sides see the result. A scorecard that only measures the supplier is a vendor audit with a friendly name, and suppliers respond to it the way anyone responds to a one sided audit. And the baseline has to be set honestly, using rolling origin cross validation on history rather than a naive comparison chosen after the fact, otherwise the programme can show improvement that came from the baseline being weak.
Two measures beyond accuracy are worth carrying. Exception closure, meaning the share of triggered exceptions that reached a documented decision and how long that took. An exception process with no closure metric decays within about two quarters. And plan adherence, meaning the share of the agreed plan that was actually produced and actually ordered. That last one catches the failure mode nobody talks about, where both sides agree the number in the meeting and then go back to their own systems and do something else.
Expect the accuracy improvement to be small and concentrated. Most of it will be in promoted weeks and in items with high promotional dependence. Base demand for a stable item is already forecast about as well as it can be by either party, and the shared version will not beat it by much. That is the correct result and it should not be presented as a disappointment.
The limit
None of this survives a commercial relationship where one side routinely uses forecast information as negotiating leverage.
The mechanism is simple enough to predict. A supplier shares a forecast showing strong sell-out and growing category share. At the next range review, that same figure appears in a slide justifying a margin request or a bigger promotional contribution. It happens once, the supplier's commercial team hears about it, and from then on the shared forecast is a negotiating position with a chart attached. The numbers stay accurate enough to be defensible and stop being the actual planning number, and the planning function on the supplier side quietly maintains a second one.
The tell is a timing pattern. Pull the history of forecast revisions and mark the dates of commercial reviews. If revisions cluster in the weeks before those meetings, the forecast has become a commercial instrument and no amount of process design will recover it.
Partial protection exists. Write down what the collaborative data may be used for and who may see it, keep the joint planning contact separate from the trading contact, and keep the shared number out of the terms conversation as a matter of policy rather than goodwill. None of that is enforceable and all of it helps, because it makes a breach visible when it happens.
Two other conditions matter. These arrangements need relationship stability, and where the buyer changes every eighteen months you rebuild the trust from zero each time, which is longer than the payback. And some categories do not justify it at all, because stable demand on a low value item with no promotional activity gives two parties nothing to collaborate about.
Take your largest customer and the category with the heaviest promotional volume, put one quarter of agreed event plans next to what actually sold, and let the size of that gap decide whether the conversation is worth opening.