In short: A Kalman filter carries a level estimate and a variance for it, and each week it splits the difference between what it expected and what arrived, in proportion to which it trusts more. Two numbers control that split: the variance of the underlying level's movement and the variance of the measurement noise. Their ratio fixes the steady-state gain, so a ratio of 0.11 gives a gain near 0.28 and takes seven weeks to absorb a permanent step, while a ratio of 0.44 gives 0.48 and takes three and a half at the cost of roughly forty percent more wobble on a stable item. Estimate both by maximum likelihood rather than by eye, and read the standardised innovations weekly to find out whether the model you assumed is the one generating your data.
An item picked up a fourth retailer in March and its weekly baseline stepped up by about thirty percent overnight. The exponential smoothing model took most of a quarter to close the gap. The planner overrode it every week for eleven weeks, and after that she stopped looking at the statistical forecast at all, which is the real cost of the episode.
Alongside that, the same item had three sources of truth. Point of sale data from two of the four retailers, weekly shipments out of the distribution centre, and a monthly sell-out file from the distributor covering the third. They disagreed, they arrived on different calendars, and the standard answer was to pick one and ignore the rest.
Both of those are state estimation problems, and the machinery for state estimation has been sitting in the literature since Kalman published it in the Journal of Basic Engineering in 1960.
What happens in the weekly update
The idea is to treat the forecast as a belief about a hidden quantity, carried forward from week to week and corrected whenever evidence arrives, rather than as a formula applied to a window of history.
Take the simplest useful version, the local level model. There is a true underlying demand level that drifts week to week, and there is what you observe, which is that level plus measurement noise. Two equations: the level today equals the level last week plus a random disturbance with variance Q, and the observation equals the level plus a separate disturbance with variance R.
The filter carries two things: its current estimate of the level, and the variance of that estimate. Work one week by hand. Suppose it comes into the week believing the level is 1,000 with an estimate variance of 400, and R is 900. The variance it expects for the incoming observation is 400 plus 900, which is 1,300. Then 1,150 units arrive. The innovation, meaning the part of the observation the filter did not predict, is 150.
The gain is the ratio of the filter's own uncertainty to the total predicted uncertainty, 400 divided by 1,300, which is 0.308. The updated level is 1,000 plus 0.308 times 150, or 1,046. The estimate variance falls to 400 times 0.692, which is 277, and then grows again by Q before the next week starts.
That is the whole mechanism. Every week the filter asks how much of the surprise to believe, and answers by comparing how uncertain it is about the level against how noisy it thinks the measurement is. Where the measurement is precise and the level is volatile, the gain runs high and the filter follows the data. Where the measurement is noisy and the level is stable, the gain runs low and the filter holds its ground.
The two variances decide everything
For the local level model, the steady-state gain depends only on the ratio q of the two variances, through a closed-form expression: the gain settles at the positive root of the relation, which works out as minus q plus the square root of q squared plus four q, all over two.
Run two cases. With Q of 100 and R of 900, q is 0.111 and the gain settles at 0.28. With Q of 400 and R of 900, q is 0.444 and the gain settles at 0.48. Those two numbers behave very differently in front of the March event above. After a permanent step, the remaining gap shrinks by one minus the gain each week, so at a gain of 0.28 it takes seven weeks to close ninety percent of the step, and at 0.48 it takes three and a half.
Faster is not free. On an item whose level is genuinely stable, the noise in the level estimate scales as the gain divided by two minus the gain, times R. That works out at 0.16 R for the low gain and 0.32 R for the high one, so the responsive setting carries about forty percent more wobble in the level it reports every week. The choice between them is a real trade and it should be made per segment, since an item on stable distribution and an item mid-rollout want different answers.
Anyone who has set a smoothing constant will recognise the shape of this, and the equivalence is exact: the local level filter at steady state reproduces simple exponential smoothing with the gain as its parameter, a correspondence set out in full by Hyndman, Koehler, Ord and Snyder in Forecasting with Exponential Smoothing (2008). The parameterisation of the wider smoothing family, including how the seasonal component is handled, is J2's subject.
What the state space form adds is a principled way to get the parameter. Rather than a grid search over the smoothing constant against a chosen error measure, the filter's own one-step-ahead prediction errors and their variances give you the likelihood directly, through the prediction error decomposition, and both Q and R can be estimated by maximising it. Harvey's Forecasting, Structural Time Series Models and the Kalman Filter (1989) and Durbin and Koopman's Time Series Analysis by State Space Methods (2012) both set out the mechanics. The practical consequence is that the responsiveness of every item is fitted from that item's own history instead of being set by a planner who once picked 0.3 because it looked sensible.
Missing weeks, closed depots and short months
The filter handles a missing observation by skipping the update step. It runs the prediction, has nothing to correct with, and carries the level forward with a variance that has grown by Q. Nothing else changes, and the forecast is still available.
That sounds like a small technical convenience and it removes a persistent source of damage. A depot shut for three weeks over a national holiday, a retailer whose file failed to arrive, an item on temporary listing suspension: in every case the usual treatment is to write zeros into the history, and zeros are demand statements. A smoothing model fed three zeros will pull the level down by a third or more and then take a quarter to recover, which is the same failure the March episode caused, running in the opposite direction.
The filter also copes with observations on different calendars, which matters when monthly distributor files and weekly shipments describe the same underlying series. You model the level weekly and connect the monthly observation to the sum of the four or five weeks it covers, and the filter reconciles them. Trading calendars with four and five week months stop being a data cleaning problem and become part of the observation equation.
Blending sources that disagree
The observation equation can carry more than one measurement, each with its own noise variance, and the filter weights them by precision automatically. Give point of sale a small R because it is close to the consumer and arrives clean, give the distributor's monthly file a large R because it is late and lumpy, and the filter will lean on the first while still extracting what information the second contains.
This is the honest version of what gets sold as signal blending. There is no weighting committee and no set of coefficients somebody tuned. Each source contributes in proportion to how precisely it measures the state, and those precisions are estimated from the data. Where a source has a persistent offset rather than noise, for instance shipments running systematically above consumption because of channel loading, you model that offset as part of the observation equation rather than pretending it is random. What a fast downstream signal can and cannot fix at the planning level is D10's subject.
The same structure extends to drivers. Regression coefficients can sit in the state vector and be allowed to move, so the price elasticity the model uses is re-estimated as evidence accumulates rather than fixed at its historical average. Scott and Varian's 2014 paper in the International Journal of Mathematical Modelling and Numerical Optimisation combines that with variable selection, and West and Harrison's Bayesian Forecasting and Dynamic Models (1997) covers the wider family.
Reading the innovations
The most useful output of a filter in production is the series of standardised innovations, meaning each week's surprise divided by the standard deviation the filter predicted for it. If the model is right, those numbers are uncorrelated, centred on zero, and have a standard deviation of one. Every way they fail tells you something specific.
Positive autocorrelation at lag one means the filter is consistently surprised in the same direction two weeks running, which is what happens when Q is too small and the level is moving faster than the model allows. A Ljung-Box test, from Ljung and Box in Biometrika in 1978, applied to the innovations of each item, turns this into a weekly report.
A standard deviation persistently above one means the filter's predicted variance is too small, so R or Q is understated and every interval it publishes is too narrow. This is worth checking before anyone trusts a service level computed from those intervals.
A single large innovation followed by a return to normal is an outlier in the measurement. A large innovation followed by a run of same-signed smaller ones is a break in the level. Harvey and Koopman's 1992 paper in the Journal of Business and Economic Statistics gives the formal separation through auxiliary residuals, which point at the irregular disturbance for the first case and the level disturbance for the second. That distinction is the difference between correcting one week of history and accepting a new baseline, and it is the single most valuable diagnostic on this list.
Three ways this goes wrong in a planning system
Someone tunes the variances by eye. A planner asks why the forecast is slow, an analyst raises Q, and the model starts chasing promotional spikes as though they were baseline shifts. Gardner's 2006 review of exponential smoothing in the International Journal of Forecasting examined the adaptive parameter methods that try to solve this automatically and found they had not delivered consistent gains. Fit the variances by maximum likelihood, refit on a schedule, and put the resulting gain on a report so a change in it is visible.
The filter is fed shipments and told they are demand. It will track whatever it is given with equal diligence. Shipment series carry order batching and channel loading, and the amplification of demand variability as you move upstream, described by Lee, Padmanabhan and Whang in Management Science in 1997, is present in the very series most planning systems have easiest access to. Weeks where you were out of stock are worse still, since they record what you could ship. Recovering demand from censored sales is D1's subject.
Diffuse initialisation is skipped. At the start of a series the filter has no information about the level, and the correct treatment gives the initial state an infinite variance, handled exactly through the method Koopman published in the Journal of the American Statistical Association in 1997. Implementations that instead seed the level with the first observation and an arbitrary small variance produce confident nonsense for the first several periods. On a new item with eight weeks of history that is the whole of the forecast, and how to get a defensible forecast from a short series is J24's territory.
Where this stops
The standard filter assumes linear dynamics and Gaussian disturbances, and demand for slow-moving items is a count with a floor at zero. Snyder, Ord and Beaumont made the case in the International Journal of Forecasting in 2012 for modelling intermittent series with distributions that respect that, and the whole intermittent question, including which method suits which pattern, sits with D5. Do not run a Gaussian filter across the tail of your catalogue because it worked on the fast movers.
The filter is a real-time estimator, and that is a constraint as much as a feature. It cannot see a level break until enough evidence has arrived to distinguish it from noise, and the estimate that makes the March step look obvious is the smoothed one, computed backwards over the whole sample once later weeks are available. Any demonstration where a filter appears to detect a shift the week it happened is usually a smoother being run on complete data, which is a different thing from what runs on a Wednesday morning.
The intervals it produces are correct conditional on the model being right, which is a stronger condition than it sounds. Check their empirical coverage before using them to set stock, and where coverage is poor, calibrating intervals against realised errors is J26's subject.
There is an operational cost too. A filter keeps state per item and location, so a hundred thousand series means a hundred thousand state vectors that have to be versioned, backfilled when history is restated, and reconstructed when someone reloads two years of corrected data. On stable fast movers a well-fitted smoothing model gets close enough that this engineering commitment is hard to justify. The filter earns it where observations are missing or irregular, where several sources of different quality describe the same series, and where the level genuinely moves.
Pick twenty items that gained or lost distribution in the last year, fit a local level model to each, and plot the standardised innovations against the date of the change to see how many weeks the filter took to notice.