In short: Returns forecasting is a convolution: units sold in a week, multiplied by a return rate, spread forward across a lag distribution fitted from your own transaction history. Fit and size, damage, and remorse on high-consideration purchases carry different rates, lags and dispositions, so a single blended rate absorbs a movement in any one of them and hides it for a quarter. Grading is a supply lead time with a throughput rate and a queue, and ungraded stock in the building is invisible supply that a planner will buy again. Reason code data is the binding limit in most operations, and until the customer's stated reason and the warehouse's observed condition are held as separate fields the model has no explanatory structure under it.
It is the second week of January and the returns dock has a queue. Trailers of December's sales are coming back across six weeks, the goods-in team was staffed against inbound purchase orders, and the grading bench is four days behind. Meanwhile a planner is raising a replenishment order for a size 12 in a line that has eleven hundred units sitting unsorted forty metres away, invisible to the system because nothing has been graded yet.
Everybody in that building knows returns are coming. Nobody has a number for them, and nothing in the plan has a line for them. The returns rate gets quoted as an annual percentage in a finance pack, which is the least useful form the information can take, because it carries no timing, no location and no disposition.
Returns behave well enough to forecast. They lag sales with a distribution that is stable within a category and channel, and that distribution is recoverable from history you already have.
Why returns are forecastable
A return is a sale that comes back after a delay. If you know the sales by week and you know the distribution of the delay, the returns forecast is a convolution: units sold in week t, multiplied by the return rate, spread across weeks t plus one through t plus n according to the lag distribution.
Fitting the lag is a distributed lag estimation on your own transaction data. For each cohort of sales, count what fraction came back one week later, two weeks later and so on. Toktay, Wein and Zenios used exactly this structure in Management Science in 2000 to estimate the return flow for a remanufactured product from sales history, and the broader modelling family was surveyed by Fleischmann and colleagues in the European Journal of Operational Research in 1997. The mechanics have been settled for a long time; what is rare is anyone running them on their own data.
Three properties make the estimation easier than it sounds. The returns window is a policy you set, so the lag distribution is truncated at a date you already know, and the truncation point is a hard boundary rather than an estimate. The distribution is usually strongly right-skewed with a mode inside the first two weeks, so most of the mass arrives early and the tail matters mainly for capacity rather than for inventory. And the rate and the lag are separable, which means you can update the rate quickly when a range changes while keeping a lag shape estimated on years of history.
The one thing that breaks the convolution is a policy change. Extending a returns window from thirty days to ninety, or introducing free returns on a channel that previously charged, moves both the rate and the shape, and it does so immediately rather than gradually. Treat policy changes as a break point in the series in the same way you would treat a change in list price.
The National Retail Federation, in the annual returns studies it published with Appriss Retail in 2023 and with Happy Returns in 2024, reported total United States retail returns in the mid-teens as a percentage of sales in its 2023 and 2024 editions, with online return rates running consistently above the overall figure. Whatever your own number is, it is large enough that a demand plan ignoring the reverse flow is planning against the wrong net requirement.
Three causes that behave differently
The aggregate return rate is a weighted average of causes that have almost nothing in common, which is why a single rate forecasts badly the moment mix moves.
Fit and size. The dominant cause in apparel and footwear, and the reason apparel sits at the top of every published category ranking. The rate is high, the lag is short because the customer knows within minutes of opening the parcel, and the units come back saleable. Bracketing behaviour, where a customer orders three sizes intending to keep one, produces a return rate that is partly a function of your own size guidance and partly a function of how the checkout is designed. The signal that matters here is the rate by style and size, because a single badly graded size in a new range will show up as a return spike four days after the first deliveries land.
Damage. Concentrated in fragile goods, large items and anything with a long final mile. The rate is lower, the lag is short, and the units come back unsaleable in their original form. The useful correlations are with carrier, lane and packaging specification rather than with the product, and a damage rate that moves without a product change is usually telling you something about a depot or a route.
Buyer remorse in high-consideration purchases. Electronics, furniture, jewellery, anything expensive enough that the customer thinks about it after buying. The rate is moderate, the lag is long because the item sits in a hallway for three weeks, and the value at stake per unit is high. These returns are the most likely to come back after the selling window has closed, which makes them an inventory problem rather than a capacity one.
Forecast each separately where the volume supports it. The rate for one can move sharply while the others are flat, and a single blended rate will absorb the movement and hide it for a quarter.
Disposition and what it does to the inventory position
A unit arriving on the returns dock has four possible futures, and the split between them is a forecastable quantity in its own right.
Restock. Back into sellable stock at full value. The most common outcome in apparel and the one with the largest planning consequence, because those units are supply that a planner will otherwise buy twice.
Refurbish or repackage. Recoverable with work, which means a lead time and a labour cost, and often a different item code with its own demand stream.
Liquidate. Sold through a secondary channel at a fraction of value, which is a revenue decision with a cannibalisation risk attached.
Destroy. Sometimes the cheapest option once processing cost is counted, and always the one that most needs an explicit threshold rather than a case-by-case judgement.
The planning point sits in the first of those. Restocked units have to re-enter the available-to-promise position, and they have to enter it with a date. A unit that is physically in the building but ungraded is invisible supply, and invisible supply produces exactly the behaviour in the opening scene: a purchase order raised against demand that a returns bin already covers. The grading operation is therefore a supply lead time, and it should be modelled as one, with a throughput rate and a queue, rather than treated as an administrative step that happens eventually.
Speed has value beyond the visibility. Blackburn, Guide, Souza and Van Wassenhove argued in California Management Review in 2004 that the marginal value of time in a reverse chain is high for products that lose value quickly, and that the reverse network should be designed for response rather than for cost where that is true. A returned smartphone or a seasonal garment loses value every week it sits in a queue. A returned industrial fitting does not. The design implication is that the grading decision belongs as far forward in the network as you can put it for fast-depreciating goods, and as centralised as possible for everything else. Guide and Van Wassenhove traced how that literature developed in Operations Research in 2009, and the split between responsive and efficient reverse chains has held up well since.
Capacity for the returns operation
The returns dock is a warehouse with a demand plan nobody wrote. It has receiving, inspection, grading, repackaging and put-away, all of which consume hours, and it is usually staffed off a rolling average with overtime absorbing the difference.
Once you have a returns forecast by week, that operation can be planned with the same driver-based method used on the outbound side: units arriving, multiplied by minutes per unit for each disposition path, converted into hours and laid against shift capacity. The complication specific to returns is that the mix of dispositions drives the hours as much as the volume does, because a restock is quick and a refurbishment is not. A week with flat volume and a shift toward damage will need materially more hours than the unit count suggests.
The second complication is that returns processing capacity has an inventory consequence rather than only a service one. Rather than losing a sale directly, underprocessing converts sellable stock into a queue, extends the effective lead time on that stock, and pushes replenishment orders that were avoidable. That makes it easy to underfund, because the cost lands in someone else's budget.
The spike after peak
The most plannable thing in the whole area is the seasonal one, and it is the one most often absorbed as unbudgeted overtime.
Returns peak after the selling peak, at a distance determined by the lag distribution and by the returns policy. Two effects stack in December. Gift purchases have a longer lag than self-purchases, because the item is not opened until the gift is given, so the December cohort's lag distribution is shifted right relative to the rest of the year and needs to be estimated separately. And extended Christmas returns windows, where anything bought from November can come back until late January, push the tail further out and raise the total, because a longer window increases the rate as well as the delay.
Estimating that properly is a matter of fitting the lag on the December cohort alone across several years and treating it as its own distribution. The output is a weekly arrival profile from late December into February, which is what the returns operation needs to hire against, and what the demand planner needs in order to stop buying inventory that is already on its way back.
The limit
All of this rests on returns reason codes, and reason code data in most operations is not good enough to carry it.
The failure is well known to anyone who has looked. A customer-facing dropdown with twelve options where the first one is selected forty percent of the time. Warehouse graders working to a throughput target choosing whichever code clears the screen fastest. A catch-all "other" absorbing a third of volume. Codes that mix cause with disposition, so that "damaged" and "refund issued" sit in the same list. When the data looks like that, the model cannot separate a fit problem from a carrier problem, and since those two behave differently in rate, lag and disposition, a blended forecast will be wrong in a way that is stable enough to look like it is working until mix moves.
Fixing it is unglamorous and cheap. Shorten the customer-facing list to the smallest set that distinguishes causes which behave differently, which is usually five or six options. Keep the customer's stated reason and the warehouse's observed condition as two separate fields, because they answer different questions and merging them destroys both. Randomise the option order or remove the default so the first item stops collecting non-answers. Then audit a sample of a few hundred units against their codes and measure the agreement rate before trusting anything built on top.
Until that is in place, forecast at the aggregate level and be honest that the model is a rate and a lag with no explanatory structure underneath. That is still considerably better than nothing, and it will still tell the returns operation when to hire.
Take one category, pull last year's weekly sales alongside weekly returns with their original sale dates, and fit the return rate by weeks-since-sale; the curve takes an afternoon to produce and it will usually be stable enough across quarters to plan the January dock against.