In short: Service parts demand comes from a population of machines failing rather than from customers deciding to buy, so the forecast input is the installed base by age cohort plus whatever the preventive maintenance calendar consumes. The two turning points that cause most of the damage, growth to plateau and the retirement decline, are invisible in the part's own demand history and visible in the installed base. The all-time buy at end of production is a newsvendor problem with a very asymmetric cost pair, treated formally by Fortuin as far back as 1980. The number worth reporting is fill rate at the point of use, meaning the share of jobs where the engineer had every part at the first visit, rather than fill rate at the central warehouse.
An engineer is standing in front of a stopped machine in Leeds waiting for a pressure sensor that costs forty pounds. There are two hundred of them in the central warehouse in Northampton. The machine stays down for three days, the customer has a four hour response contract, and the credit note is worth more than the annual holding cost of stocking that sensor in every van in the region.
The monthly report says central fill rate is 98.4 percent. Nobody reports the number that mattered, which is whether the part was on the van.
Service parts planning goes wrong in a specific and repeatable way: the methods that work for finished goods get applied to a demand stream that is generated by a completely different mechanism, and the metrics that work for a distribution network get applied to a network whose last echelon is a vehicle.
The demand driver is a population of machines
Finished goods demand comes from customers deciding to buy. Service parts demand comes from machines you have already sold deciding to fail. That difference has practical consequences, and the first one is that the forecast input is a population rather than a trend.
The structure is straightforward. For each part, demand in a period is the sum across the installed base of the probability that each installed unit needs that part in that period. Split the base into age cohorts, attach a failure rate per unit per period to each cohort, multiply and sum. Add the deterministic component, which is often larger than people expect: preventive maintenance schedules generate part consumption on a known calendar, and any part on a scheduled replacement interval is forecastable from the service calendar with almost no statistics involved.
What this changes in practice is where you look when demand moves. A part whose consumption rose thirty percent this year has either an ageing population, a growing one, a reliability problem, or a change in service policy, and those four have different responses. A time series model fitted to the part's own history will register the rise and extrapolate it, with no view on which of the four caused it or whether it continues.
It also changes the lead indicator. Equipment sales tell you about parts demand several years out rather than this year, because a machine sold this quarter consumes almost nothing under warranty and starts generating real demand once it ages into the part of the curve where wear-out begins. The installed base register, and specifically its age profile, is the forecast input. Equipment sales are a forecast of the installed base.
The demand series themselves are usually lumpy, with long runs of zeros punctuated by ones and twos, and the estimation methods appropriate to that pattern are a separate subject covered elsewhere. The point here is that even a well chosen intermittent method is fitting a series whose generating process you could have modelled directly.
The lifecycle shape
Parts demand for a given part number follows the population that consumes it, and that population has a life.
Early on, the installed base is small and young, warranty absorbs much of what fails, and demand is low and noisy. Through the middle years the base grows and ages at the same time, and demand rises on both counts. It plateaus when new installations roughly balance retirements. Then it falls as machines are decommissioned, and the fall is usually slower than the rise because the surviving population skews old and old machines consume more parts per unit.
Two turning points cause most of the damage. The first is the transition from growth to plateau, where a model fitted on the growth period keeps extrapolating and you buy into a plateau. The second is the retirement decline, where a model fitted on the plateau keeps forecasting flat while the population shrinks underneath it, and the resulting stock becomes an obsolescence problem two or three years later, long after the decision that caused it.
Neither turn is visible in the part's own demand history until after it has happened. Both are visible in the installed base, if you can see it: installations by year, retirements by year, and the resulting age distribution. Retirement is the harder half to observe, because customers rarely tell you they scrapped a machine. The workable proxy is service activity cessation. A machine that has generated no service call, no part and no contract renewal in a defined window is probably gone, and calibrating that window against the machines you know were decommissioned gives you a retirement curve good enough to plan against.
The all-time buy
At end of production, the supplier stops making the part and you have a support obligation that runs for years past that date. Somewhere in that period a decision gets made about a final order, and it is one of the few genuinely one-shot decisions in planning.
The sizing question has been treated formally since Fortuin wrote about the all-time requirement for service parts in the International Journal of Operations and Production Management in 1980, and Teunter and Fortuin took it further with an end-of-life service case study in the European Journal of Operational Research in 1998, followed by their end-of-life service paper in the International Journal of Production Economics in 1999. The structure is a newsvendor with an unusually asymmetric cost pair. Order too many and you write off the surplus at the end of the support period, at full cost, having held it for years. Order too few and you face emergency remanufacture at a specialist price, contractual penalties, or the reputational cost of telling a customer that a machine they bought eight years ago is now unsupportable.
Sizing it needs three inputs and only one of them is a demand forecast. The remaining demand is the integral of the failure rate over the declining installed base across the support period, which is the lifecycle curve above extended to the horizon. The overage cost is the unit cost plus the holding cost across an average of several years, which is often close to the full value of the part. The underage cost is the one people skip, and it has to be a real number, obtained by asking what the business actually does when a part runs out: a remanufacture quote, a contractual liability, or a machine buy-back.
Two moves reduce the exposure without changing the arithmetic. Sizing in stages, where you take an initial quantity and negotiate an option on a second run at a stated price, converts a one-shot decision into two smaller ones and is frequently available if you ask early. And harvesting parts from decommissioned machines is a supply source that grows exactly as the installed base retires, which is the same period when the final order is running down. Both are worth pricing before the final order is placed, because neither is available afterwards.
Criticality sets the availability target
Setting one availability target across a service parts catalogue produces the specific failure in the opening scene: high fill rates on cheap consumables that were easy to stock, and gaps on the low-volume parts that stop machines.
The differentiation that matters is what happens when the part is missing. A machine-down part on a contract with a four hour response commitment has a stockout cost measured in penalties and lost renewals. A consumable that the engineer can fetch next visit has a stockout cost close to zero. A cosmetic part has less than that. Classifying parts on consequence rather than on value or volume is the oldest idea in this discipline and it is still the one most often skipped, because ABC by value is already sitting in the system and criticality has to be assigned by someone who knows the machine.
Once parts are classified, the target has to be attached to the combination of criticality and contract rather than to the part alone. The same sensor might justify van stock for customers on a four hour response and regional stock for customers on next business day, and holding it everywhere for everyone is the expensive way to avoid making that distinction.
The allocation across the resulting classes is a constrained optimisation with a concave return, so the marginal analysis used to allocate a buffer budget across a catalogue applies here directly. What changes is the objective: the thing being bought with the last pound is avoided downtime on a critical part rather than units of fill rate on an average one.
The network ends at a van
Service networks have more echelons than distribution networks and the last one moves. Central warehouse, regional hub, forward stocking locations near clusters of installed base, and then van stock, which is inventory that drives around.
The formal treatment of this structure is older than most planning software. Sherbrooke published METRIC in Operations Research in 1968 for multi-echelon recoverable item control, Muckstadt extended it to hierarchical parts in Management Science in 1973, and Graves published an alternative approximation in Management Science in 1985. The framing they established is the one that still matters: the objective is expected backorders at the point of use, weighted by what a backorder costs there, rather than fill rate measured at each location independently. Optimising each echelon to its own service target reliably produces the Leeds outcome, because every location hits its number and the engineer still has nothing.
Two features of service networks change the stocking answer relative to a standard multi-echelon problem, which has its own treatment elsewhere. Lateral supply is real: a part on another engineer's van two hours away is a genuine source, and a network with reliable lateral transfer needs less total stock than one without, though it needs visibility of van inventory that many operations do not have. And returns of failed units feed repair loops, so for repairables the stock in the system is largely fixed and the planning question is where it sits and how fast the repair loop turns rather than how much to buy.
The measurement follows from that. The number worth reporting is the fill rate at the point where the work happens, meaning the proportion of jobs where the engineer had every part needed at the first visit. First time fix rate is the same measurement from the customer's side, and it correlates with contract renewal in a way that central fill rate does not. Cohen, Agrawal and Agrawal made the commercial case for taking this seriously in Harvard Business Review in 2006, pointing out that the aftermarket is often a larger profit pool than the equipment sale it follows.
The limit
Everything above is driven by the installed base register, and installed base data is usually worse than anyone wants to admit.
Machines get sold through distributors who do not report the end customer. Equipment moves site without anyone updating a record. Configurations change in the field, so the register says one variant and the machine in the room is another with different parts. Serial numbers are captured at manufacture and lost at installation. Decommissioning is almost never reported, so the register accumulates ghosts that inflate the base and pull forecasts upward for years.
The consequence is specific and worth naming, because it is easy to misdiagnose. A forecast driven by a population you cannot see accurately produces errors that look like demand volatility. The part appears erratic, someone concludes the demand is unforecastable, a larger buffer gets applied, and the underlying data problem stays invisible because the symptom has been treated. The distinguishing test is whether the errors are structured: consistently high in regions where the register is stale, consistently low where installations were captured well.
The reconciliation that finds it is cheap. Take a year of service calls and match them to the installed base register by serial number. Count machines that generated work but do not appear in the register, and count machines in the register that have generated nothing for three years. The first number is base you cannot see, the second is base that probably no longer exists, and together they size the error sitting inside every parts forecast you produce.
There is also a floor below which none of this pays. For a small installed base of a few hundred machines with low failure rates, the forecast for many parts is genuinely a judgement about whether to hold one or none, and no amount of modelling improves that decision. What modelling does there is tell you which parts deserve the judgement.
Take one part family and reconcile last year's service calls against the installed base register by serial number; the count of machines that generated work while absent from the register is the size of the error currently sitting inside your forecast, and it usually takes a day to produce.