In short: A promotion does three separate things to demand: it creates incremental volume, it pulls forward purchases that would have happened anyway, and it takes volume from other items in the range. Only the incremental part has commercial value, and separating the three requires a baseline, which means estimating a world you cannot observe. Run the decomposition consistently across a promotional calendar and the resulting pattern is uncomfortable and fairly universal. Cannibalisation estimated from observational data is an inference rather than a measurement, so it should be reported with a range.
A four week feature deal runs on a 500ml pack. Sales during the period are 180% of the preceding four weeks. The event report says uplift of 80%, everyone agrees it worked, and the same deal goes into next year's calendar.
Three separate things happened in those four weeks and only one of them is worth anything.
Incremental volume. Consumers who bought because of the deal and would not otherwise have bought. This is the part you were paying for.
Cannibalisation. Consumers who would have bought the 750ml pack, or the sister brand, and switched to the promoted one. The category total barely moved. You discounted a sale you already had.
Pull-forward. Consumers who would have bought in three weeks anyway and stocked up instead. The volume is real and it is borrowed from the weeks after the event, which will now look weak.
Reported as one number, all three look like success. Measured separately, a large share of promotional calendars turn out to be moving volume around rather than creating it.
The counterfactual problem
The difficulty is that you need to know what would have happened without the promotion, and that world is not observable. Everything else is technique for estimating it.
Three approaches, in increasing order of rigour and cost.
Baseline decomposition. Model the non-promoted demand pattern using seasonality, trend and price, then treat the gap between actual and modelled baseline during the event as uplift. Cheap, works reasonably when promotions are infrequent, and degrades badly when the item is promoted most of the year, because there are not enough clean weeks to estimate a baseline from.
Control groups. Compare promoted stores or regions against comparable ones that did not run the event. Substantially more credible when the control is genuinely comparable, which is a stronger condition than it sounds. Retailers rarely randomise which stores get a feature, so the promoted set is usually selected on characteristics that also predict sales.
Designed variation. Deliberately hold out a set of stores or a region from an event. This gives a clean read and it costs revenue in the holdout, which is why almost nobody does it. In my experience the businesses that run two or three deliberate holdouts a year learn more about their own price response than those running elaborate models on observational data.
Baseline decomposition has a longer pedigree than most people using it realise. Abraham and Lodish built PROMOTER on it at Information Resources in 1987, and the structure has survived largely intact: estimate what the series would have done from its unpromoted periods, then read the event as the residual. The reason it degrades on heavily promoted items is visible in that description, since the estimator needs unpromoted weeks to learn from.
Measuring each of the three effects
Once you have a baseline you trust, the decomposition becomes tractable.
Gross uplift is actual minus baseline over the event window. The easy part.
Cannibalisation requires looking at the products the promoted item competes with over the same window. Compute their actual against their own baselines. Where they fall short, that shortfall is a candidate for volume that transferred rather than disappeared.
The scope of the comparison set matters more than the method. Too narrow, and you miss transfer to a different pack size. Too wide, and you attribute unrelated noise to the promotion. A sensible default is the same category and need state, weighted by pack size proximity, because a 400ml bottle catches far more of a promoted 500ml than a 2 litre does. That structure reflects how substitution actually works and it is worth encoding explicitly rather than treating every item in the category as an equal substitute.
Pull-forward shows up in the weeks after the event, as a dip below baseline. Measure the post-event window until demand returns to baseline, and the cumulative shortfall is the borrowed volume. The window needs to be long enough, which for a stockpileable product with a long consumption cycle can be six to eight weeks. Cutting the measurement window short is the most common way pull-forward gets hidden, and it always flatters the event.
Net incremental volume is gross uplift, minus cannibalised volume, minus pulled-forward volume. That is the number that belongs on the event report.
Worth putting numbers on it once, because the gap between the two reports is wider than most people expect.
Take the 500ml pack from the opening. Its modelled baseline is 1,000 units a week and it sells 1,800 a week through the four week event, so 7,200 units against a baseline of 4,000. Gross uplift is 3,200 units, which is the 80% everyone agreed on.
The 750ml runs a baseline of 600 units a week and delivers 460, a shortfall of 140 a week, or 560 units across the event. In litres that is 420, which converts to 840 units of 500ml equivalent. The sister brand falls 50 units a week short, another 200 equivalent units. Cannibalisation totals 1,040.
The post-event dip on the promoted item runs six weeks before demand returns to baseline, with weekly shortfalls of 220, 180, 140, 90, 50 and 20. That is 700 units of pull-forward.
Net incremental volume is 3,200 minus 1,040 minus 700, or 1,460 units.
Now the money. The pack lists at 2.00 against a variable cost of 1.20, so a normal unit earns 0.80. The deal price is 1.60, so a promoted unit earns 0.40. Of the 7,200 units that shipped on deal, 4,000 were baseline and 700 were pulled forward from later weeks, and all 4,700 of those would have sold at 0.80 without any discount. That is 1,880 of margin handed over. The 1,460 genuinely incremental units earned 0.40 each, or 584.
So the event bought 584 of new margin with 1,880 of discount applied to volume you already had, before counting the margin lost on the 750ml it took sales from. The four components tie back to the 7,200 units sold, which is a useful check that the decomposition is complete.
The result nobody enjoys
Do this consistently across a promotional calendar and a pattern emerges that is uncomfortable and fairly universal.
A small number of events are strongly incremental, usually those that bring new buyers into the category or land at a moment of genuine seasonal demand.
A large middle group is roughly neutral, generating gross uplift that is mostly cannibalisation and pull-forward. These events are not disasters. They also are not returning their discount.
A tail is actively negative, where the discount cost exceeds any incremental margin, typically deep discounts on products with low price elasticity and high stockpiling potential.
The proportions vary by category and business, and the shape recurs. Abraham and Lodish, writing in Harvard Business Review in 1990 from scanner data across a large body of events, reported that only a small minority of the trade promotions they examined were profitable. Published estimates since have ranged widely, commonly in a ten to twenty percent band, though the widely repeated sixteen percent traces to AMR Research rather than to that article, and have pointed in the same unflattering direction for decades. The specific number matters less than the fact that the distribution is skewed and most organisations cannot tell which of their events sit where.
There is also a practical way to stop guessing at the post-event window. For each past event, compute the cumulative shortfall against baseline week by week after the event ends, and find where that cumulative curve flattens. Across a category it tends to flatten at a consistent point, and that point is your window. Businesses using four weeks because the reporting period is four weeks are cutting the curve wherever the calendar happens to fall, which flatters some events and penalises others at random.
When the promotion data is incomplete
A practical problem that stops a lot of this work before it starts: promotional calendars are frequently unreliable. Events recorded at the wrong date, mechanics not captured, in-store execution differing from what was planned, and historical calendars that were never maintained beyond the current year.
There is a workable fallback that uses price rather than flags. Compute a rolling median base price per item, detect periods where the actual price dipped meaningfully below it, and treat those as detected events. Then estimate response by regressing log quantity on log price relative to base, with the extreme observations winsorised so a single data error does not set the elasticity.
This recovers a price elasticity and an event list without any calendar at all. It cannot distinguish a feature from a display from a straight temporary price reduction, since it only sees the price. As a way to get started on a dataset with no reliable promotional history, it works, and it frequently reveals that the recorded calendar and the actual price history disagree substantially.
The price view also exposes a failure mode that a calendar-based measurement hides completely. When an item is on deal for most of the year, the rolling median base price drifts down toward the promoted price, and so does any statistical baseline estimated from the series. The uplift you measure then shrinks year after year while the underlying consumer response has not changed at all, because the comparison point has been sliding.
The diagnostic is one ratio. Count the weeks where the actual price sat below 95% of the rolling median base price, and divide by total weeks. Above about half, baseline decomposition has stopped being viable on that item, and the honest treatment is a price response model with the reference price set by a policy decision rather than estimated from history. Items above 70% are on permanent promotion, and the useful question about them is the list price rather than the event calendar.
What to do with the answer
Two decisions follow directly from a decomposed event history.
The forecast improves, because the model now has a promotional uplift estimate per mechanic rather than a blended average that overstates the effect. Forecasting a planned event using a gross uplift figure that included cannibalisation will over-order the promoted item and under-order everything it steals from, and both errors cost money.
The measurement of the commercial function changes. When the reported metric is net incremental margin rather than gross uplift, the incentive to run volume-shifting events disappears. That change is more consequential than any modelling improvement, and it is also the one that will be resisted, because a lot of people's numbers get worse when the measurement gets honest.
The limits
Cannibalisation estimated from observational data is an inference rather than a measurement. When two products move together, you cannot fully separate substitution from a shared cause such as a seasonal shift or a competitor action. The estimate is directionally useful and it should carry a range rather than a point.
Pull-forward measurement assumes the post-event period is otherwise normal. If another promotion follows within the recovery window, which happens constantly in heavily promoted categories, the windows overlap and clean attribution becomes impossible. In those categories the practical unit of analysis is a quarter of promotional activity rather than an individual event, and the honest report says so.
There is also an effect that none of this captures. A promotion can build long-term penetration by getting a household to try a product they then repeat-buy at full price. That is genuine value and it appears months later, outside every measurement window described here. Panel data can see it and shipment data cannot, which is a reason to be careful about condemning trial-generating events on a four-week read.
Start with one category and one quarter, and measure the post-event window properly. The pull-forward number alone usually changes how the calendar gets built.