In short: Service failures blamed on forecast accuracy are frequently caused by replenishment parameters calculated during implementation and never recalculated since. The parameters that do the damage are reorder point, order quantity, safety stock, planning lead time and the planning calendar, and each can drift away from operating reality independently of the others. Ownership has to be explicit per parameter, since a value anyone may edit and nobody maintains ends up recording the last emergency rather than current conditions. Reviewing on a trigger, meaning a demand rate or a lead time that has moved outside a stated band, catches drift that an annual calendar review will not.
Service on a category has been sliding for two quarters. The review pulls up the forecast accuracy for those items, which is fine, roughly where it has always been. Somebody suggests the model needs retuning.
Then a planner opens one of the failing items and looks at the reorder point. It was calculated in a migration spreadsheet during the implementation, against a weekly demand rate from the year before go-live, using a lead time the supplier quoted at the time. Demand on that item has grown by half since then and the supplier moved production to a different site eighteen months ago. The number has never been recalculated, because nothing in the system asks it to be.
That item stocks out on time, every time, exactly as configured.
The parameters that actually exist
Most planning conversations treat replenishment as one setting. It is seven or eight, they interact, and each one goes stale in its own way with its own symptom.
Reorder point. The level at which supply gets triggered. It should equal expected demand over the exposure window plus the buffer, and the exposure window is lead time plus review period rather than lead time alone. It drifts whenever demand rate or lead time moves, which is constantly. The symptom is stockouts on items where the order was raised correctly and simply too late, and where the forecast for the period was within its normal error.
Order quantity. How much gets ordered when the trigger fires. Usually derived from an economic order quantity, which is Harris's 1913 model and depends on an ordering cost and a holding rate that were entered once. The symptom of a stale figure is order frequency that makes no sense against value, typically expensive items ordered in large batches because the holding rate is set too low.
Minimum order quantity. The supplier's floor, which is a commercial term living in a planning field. It goes stale when volumes change and nobody reopens the contract. The symptom is a slow-moving item that receives one delivery covering three years of demand, then sits.
Order multiple. The rounding increment: layer, pallet, case, full container. It goes wrong when packaging changes and the planning field does not follow. The symptom is systematic overshoot on small-demand items, where the rounding is a larger share of the order than the order was.
Lot sizing rule. How requirements get grouped into orders over time. Lot-for-lot, fixed period, periods of supply, or an algorithm in the Wagner-Whitin family from 1958 or the Silver-Meal heuristic from 1973. It goes stale when the demand pattern shifts underneath the rule. The symptom is plan instability, orders appearing and disappearing between cycles, which Blackburn, Kropp and Millen analysed in 1986 as nervousness and which is still usually diagnosed as a system fault rather than a parameter choice.
Review period. How often the item is looked at for replenishment. It goes wrong quietly, because the configured period and the actual ordering calendar diverge without anyone changing a field. A weekly review period on an item the buyer actually orders monthly means the buffer is sized against a four-week exposure and the item is exposed for seven.
Safety stock. Either a calculated figure or a static number typed in during implementation. The symptom of the static kind is easy to spot: convert every safety stock into days of cover and look at the spread within a segment. Where items with similar demand and lead time carry wildly different cover, the numbers were entered rather than computed.
Lead time sits underneath several of these as an input rather than being a parameter you set, and it is the one that does the most damage when stale, because it enters both the trigger and the buffer.
Why the parameters are usually the cause and the forecast usually is not
When service drops, the forecast gets investigated first. It is the most visible part of the process, it has a number attached, and there is always someone willing to say the model needs work.
The argument for looking at parameters first is about the shape of the two errors. Forecast error is bounded, roughly symmetric and already accounted for, since the buffer exists precisely to absorb it. A mis-set parameter is a systematic offset that points the same direction on every cycle and that nothing in the design compensates for. A reorder point that is thirty percent too low does not average out over a year, it fails on every replenishment where demand lands anywhere near the mean.
There is a specific test that separates them, and it takes an afternoon. Take the last several service failures. For each one, pull the demand that actually occurred during the exposure window, and compare it against the forecast for that window plus the safety stock that was supposed to cover it. Where actual demand landed inside that envelope and you still stocked out, the forecast did its job and something else did not: the trigger level, the order quantity, the review timing or the execution. Where actual demand blew through the envelope, you have a genuine forecast or buffer sizing question.
Run that across twenty failures and the split is usually informative enough to stop the argument. The sizing question, when you do get one, belongs to a different calculation with its own literature, and the placement question of which node should hold the buffer at all is different again.
Who is allowed to change what
Most parameter drift is a governance gap rather than a technical one. Fields that anyone can edit and nobody owns end up holding whatever the last person under pressure typed.
Three tiers is usually enough.
System calculated, not editable. Reorder point and safety stock where a method is in place. If a planner disagrees with the output, the argument is about the inputs or the service target, and it should be resolved there rather than by typing over the answer. This tier is the one that generates the most resistance and it is the one that holds the design together.
Editable with evidence and an expiry. Planner overrides on the calculated fields, requiring a reason code and, more importantly, an end date. An override without an expiry becomes permanent by default, and permanent by default is how a temporary cover increase during a supplier problem in 2023 is still inflating stock today. On expiry the field returns to the calculated value and the planner is told it did.
Owned outside planning. Minimum order quantity, order multiple, supplier lead time and the cost inputs behind the order quantity. These describe commercial and physical facts, planning consumes them, and the evidence required to change them is a contract, a packaging specification or a receipt history rather than a planner's judgement.
Whatever the tiers, the field-level change log has to exist and be readable. Who changed what, from what, when, and why. Without it the parameter audit below cannot run, and more immediately, nobody can answer the question of when a number was last known to be right.
Review on a trigger rather than on a calendar
The standard answer is an annual or quarterly parameter review. In practice these become large low-yield exercises where a team works through thousands of items, most of which needed nothing, and abandons the exercise partway through the second one.
Triggers work better because they concentrate effort where something changed.
A demand rate that has moved more than a set percentage from the value used at the last calculation. A lead time whose trailing actuals have drifted from the parameter by more than a set margin. A service failure or an expedite on the item. A sourcing change, a new supplier contract, a packaging change, a status change. A segment reassignment, since the policy that governs the item just changed.
Each trigger produces a recalculation, and the recalculation either applies automatically inside a defined tolerance or lands in a queue as a proposal. The calendar review still has a role, and its role is a sample audit to confirm the triggers are firing rather than a line-by-line pass.
The audit that finds the worst offenders in an afternoon
Six queries will find most of the damage in a catalogue of any size. Rank every result by exposure rather than by the size of the discrepancy, because a two-hundred-percent error on a dead item is worth nothing.
Compare each reorder point against a freshly computed value using current demand rate, current lead time and current buffer. Sort by the ratio. Anything below 0.7 is under-triggering and is generating service failures now, anything above 1.5 is carrying cover nobody chose.
Convert every safety stock into days of demand and look at the distribution inside each segment. The outliers at both ends are almost always typed values rather than computed ones.
Express minimum order quantity as a share of annual demand. Anything above one year of supply is a commercial term nobody has revisited, and the list is usually short enough to take to procurement in a single conversation.
Express order quantity as months of supply and look at the top of that list against item value. High value with high months of supply is where the cash is sitting.
Pull the last-changed date on every planning parameter and join it to the demand trend. Items whose parameters have not changed since creation and whose demand has moved by more than half in either direction are the highest-yield population in the whole audit.
Look for suspicious uniformity. Several hundred items sharing an identical safety stock figure, an identical lead time or a round-number order quantity means a default was applied at load and never revisited, and finding the default value tells you exactly which population to work.
The limit
Parameter tuning at scale only works with a rollback path, and this is the part that gets skipped because the update itself is easy.
A bulk recalculation applied across ten thousand item-locations is a single click and is close to impossible to unwind by hand if the segment definition behind it was wrong. The damage also arrives on a delay, as a wave of purchase orders over the following weeks, by which time the original values are gone, the supply is committed, and the only record of what the parameters used to be is whatever your change log happened to capture.
The requirements are unglamorous. Snapshot every affected field before the update, in a form you can restore from selectively. Apply to one segment at a time rather than to the catalogue. Run the change through a simulation first and look at the projected order pattern for the next two months, since that is where a mis-specified segment shows itself. Hold the first application to a population small enough that you could fix it manually if you had to, which is a good working definition of how much you should trust the change.
The second-order effect deserves a mention as well. A large parameter change alters your ordering behaviour immediately, and your suppliers experience that as a step change in demand with no explanation. If you are moving a lot of items at once, the supply side needs to hear about it before the orders arrive rather than after.
Run the reorder point comparison on your top two hundred items by value at risk this week. The ratio distribution will tell you within an hour whether you have a forecasting problem or a maintenance problem.