In short: Shipment history is a record of ordering behaviour rather than of consumption, so a forecast built on sell-in learns your customers' ordering habits. Three distortions live inside that history and their severity varies enormously by channel, which is why the first step is checking for the signature rather than rebuilding anything. The structural fix is to forecast consumption and derive shipments through an inventory balance, rather than forecasting shipments directly. Full point of sale coverage is rare outside the largest retail relationships, so work at the level of coverage you have rather than waiting for all of it.
Most demand forecasts in consumer goods are built on shipment history. What left our warehouse, by week, by customer. It is the cleanest data in the building: it is in the ERP, it ties to revenue, and finance already trusts it.
It is also a record of ordering behaviour rather than of consumption, and three distortions live inside it.
Order batching. A distributor does not reorder every time a case sells. They wait until they hit a reorder point, a truckload minimum, or a delivery day, and then place a larger order. Steady consumer demand arrives at your warehouse as lumpy orders, and the lumpiness is a property of their replenishment policy rather than of the market.
Forward buy. A feature deal or a price increase announcement pulls future ordering into the present. Three months of demand compresses into two weeks, then the following weeks are empty because the channel is full. A model trained on that sees a genuine seasonal pattern and predicts it again.
Trade loading. Quarter-end pressure moves stock into the channel to hit a number. The pattern is regular enough that models learn it as seasonality, which then becomes a self-fulfilling forecast that justifies the next quarter's load.
Lee, Padmanabhan and Whang set out the mechanics of this in Management Science in 1997, in the paper that named the bullwhip effect. Two of the four causes they identified are order batching and price fluctuation, which is forward buy under a more formal name. Their argument was that each party in the chain behaves rationally on its own information, so the amplification survives everybody doing their job properly. A third cause on their list, shortage gaming, turns up in shipment history whenever you have been on allocation. Customers who were short-shipped inflate their next orders to protect their share, and those inflated orders stay in the record long after the shortage cleared, teaching the model a demand level that never existed.
The result is a series where a meaningful share of the variance is caused by decisions made inside your own commercial organisation and inside your customers' purchasing departments. You can forecast it accurately, and doing so tells you very little about whether anyone is buying the product.
The signature to look for
Before rebuilding anything, check whether you have the problem, because the severity varies enormously by channel.
Two comparisons do most of the diagnostic work.
Plot cumulative sell-in against cumulative sell-out over eighteen months for a category where you have both. If the lines track with a constant offset, your channel inventory is stable and sell-in is a reasonable proxy. If the gap widens and narrows, that gap is channel stock going up and down, and it is inside your forecast pretending to be demand.
Then compute channel cover in days: distributor inventory divided by daily sell-out velocity. Published rules of thumb put fifteen to twenty days in the normal operating range for many consumer categories, with cover climbing past forty or fifty as a fairly reliable sign that the channel has been loaded. Track it by customer and by quarter. If it spikes every third month and drains in the following one, you have found your trade loading calendar.
A third check catches the distortion that gets misread most often. Count orders per customer per month and put average order size next to it. When the order count halves and the average size doubles while monthly volume holds flat, the customer has changed their ordering cadence or their truckload minimum. The symptom in your own metrics is a jump in weekly forecast error on that customer with monthly error unchanged, and no statistical method fixes it, because the model is being asked to predict a delivery schedule.
The same query catches the more expensive version. Order count steady, average size creeping up across two quarters, monthly volume rising: that is the channel filling, and it usually ends in a quarter where the orders stop.
The financial version of the same tell is a divergence between revenue and cash flow, with returns rising afterwards. That one tends to get noticed by other people before it gets noticed by planning.
Building the bridge
The structural fix is to forecast consumption and derive shipments, rather than forecasting shipments directly. The relationship is an inventory balance and it is not complicated:
Sell-in for a period equals sell-out for that period, plus the change in channel inventory, plus returns and adjustments.
Rearranged, that means if you can forecast sell-out and you can model how your customers manage their stock, you can produce a sell-in forecast that separates consumer demand from channel behaviour. Three components.
Forecast sell-out. This is the demand signal proper. Point of sale data where you can get it, distributor secondary sales where you cannot, syndicated panel data as a fallback with the understanding that panel coverage is partial and the levels will not tie to your own numbers.
Model the channel inventory target. Most customers replenish to a policy, whether or not they describe it as one. Estimate the days of cover each customer holds and how it moves seasonally. This does not need to be sophisticated. A per-customer target cover, estimated from history and updated as it drifts, gets you most of the way.
Derive the shipment plan. Given a sell-out forecast and a cover target, the required shipment is whatever brings channel stock to the target given expected consumption. Deviations then have a cause you can name: the customer is destocking, or building ahead of a promotion, or has changed policy.
Worth doing the arithmetic once, because the size of the effect surprises people.
Take a customer whose consumption runs at 1,050 units a week, steady, holding twenty days of cover. Twenty days at 150 units a day is 3,000 units on hand. Ahead of a feature they decide to hold thirty days, or 4,500 units, and they build it over four weeks.
Sell-in during the build is consumption plus the inventory change: 4,200 units of sell-out plus 1,500 units of stock build, so 5,700 units, or 1,425 a week. That runs 36% above the underlying rate. After the event they revert to twenty days and release the same 1,500 units across the following four weeks, so sell-in is 4,200 minus 1,500, or 675 a week. A 36% fall.
Consumer demand was flat at 1,050 units a week for the whole eight weeks. The eight-week shipment total is 8,400, which is exactly right. Every weekly number inside it is wrong, and the series your model sees has a mean of 1,050 and a standard deviation of 375, a coefficient of variation of 36% manufactured by one customer moving one policy parameter.
That last figure is the one to carry around. Safety stock sized on the shipment series is buffering against your customer's cover target, and it scales with a variability you could have anticipated from a monthly stock declaration.
The returns and adjustments term is the one that gets dropped, and it is the one that breaks the bridge when it matters. Damaged stock, price protection credits and end-of-life buybacks all land there, and they cluster around exactly the events you are trying to measure. When a customer's bridge refuses to reconcile, check whether credits are posted against a different period from the shipments they relate to. A two-week posting lag makes a stable channel look like it is oscillating, and the oscillation will have the same period as your credit run.
The payoff is diagnosis rather than accuracy. When sell-in comes in low, you can tell whether consumption fell or the channel is working down inventory, and those two situations call for opposite responses. Reading a destock as a demand decline and cutting production is a common and expensive mistake.
When you only have partial visibility
Full point of sale coverage across every customer is rare outside the largest retail relationships. Most businesses have a mix, and the practical approach is to work at the level of coverage you have rather than waiting.
Where you have direct data, usually the top few accounts, build the bridge properly. In many consumer goods businesses those accounts carry more than half of volume, so the bridge covers most of the revenue even when it covers a minority of customers.
Where you have distributor secondary sales, treat them as sell-out with a caveat, since secondary sales are shipments from the distributor to the retailer, which is one step closer to consumption but still a shipment.
Where you have nothing, infer channel behaviour from your own order patterns. The order cadence, quantities and gaps let you estimate a reorder policy, and a customer whose orders arrive every four weeks in similar quantities is telling you their cycle even without sharing data. Deconvolving a batched order series into an implied consumption rate is imperfect and considerably better than assuming the orders are the demand.
Inferring the policy is more tractable than it sounds. For a customer ordering on a roughly fixed cycle, their stock swings between a floor and the floor plus one order quantity, so their average cover is the floor plus half the cycle. A customer taking 4,200 units every four weeks against consumption of 150 a day is carrying the floor plus fourteen days. The floor stays invisible from outside, and its movements show up clearly. A step change in order size with the cadence unchanged means the floor shifted. A step change in cadence with annual volume unchanged means the cycle shifted.
A pragmatic move that gets skipped: ask. Distributor inventory reporting is a common item in trade terms, and a monthly stock declaration from your top twenty customers costs a conversation rather than a project.
What it changes downstream
Three effects worth planning for.
Production stabilises. When you plan against consumption plus a modelled channel movement, you stop amplifying the batching. This is the bullwhip effect operating in reverse, and it usually shows up as reduced changeover frequency and lower finished goods swings before it shows up in a forecast accuracy metric.
Promotional planning improves, because the pull-forward becomes visible as a channel inventory movement rather than as incremental demand. The event's true incremental contribution is measured on sell-out, and the sell-in spike is correctly recognised as a timing shift.
Commercial conversations change. When you can show a customer their own cover trend against category velocity, the discussion about an unusually large order stops being a negotiation and starts being a look at a shared number. This is more valuable than it sounds and it is the part that survives changes in planning systems.
The limits
Sell-out is not consumption either. Point of sale records the transaction at the till, and consumer stockpiling means purchases and usage diverge, particularly during price events. For most planning purposes this distinction is second order, and it stops being second order during unusual periods when consumers buy ahead.
Panel and syndicated data carry their own problems. Coverage is partial, the store universe is projected, and the levels will not match your shipments. Use them for direction and shape rather than level, and do not attempt to reconcile them to your own totals exactly, because the difference is definitional rather than an error.
There is also a latency cost. Sell-out data usually arrives later than your own shipment data, sometimes by weeks, which means the freshest information you have about the very near term is still your own orders. The sensible arrangement uses consumption for the tactical horizon and your own order signal for the immediate week or two, blended rather than switched.
Start with the cover chart for your top ten customers. It takes an afternoon if the data exists, and it will tell you quickly whether this is a real problem for you or a theoretical one.