In short: A calendar inspection interval is a proxy for a degradation rate nobody measured, so it runs too short for most items in a population and too long for a few, and opening them is how you find out which is which. Rotating and static equipment are different prediction problems, since rotating machines produce dense vibration and process data that supports a data driven model, while static equipment has sparse readings and needs a corrosion mechanism model. A single remaining useful life figure carries less information than a probability of failure before a stated date, because the decision it feeds is whether an item can wait for the next planned shutdown. Turnaround scope is where the money sits, since moving work out of an unplanned outage and into a planned one is worth more than marginal accuracy on any single prediction.
The inspection plan for the year lists a few hundred vessels and several thousand piping circuits, each with a due date derived from a code interval and a previous thickness reading. Crews work the list. Most of what they open is in the condition the last inspection predicted. Then a line that was not on this year's list develops a leak, the unit comes down for four days, and the review afterwards concludes that the inspection programme was followed correctly.
Both halves of that are true at once, and the reason is that a calendar interval is a proxy for a degradation rate that nobody measured. The interval was set to be conservative across a population. Applied to an individual item it is too short for most of them and too long for a few, and you find out which is which by opening things.
The four steps, and what each one actually costs to take
The maturity ladder in this field is well worn and worth restating in terms of what each rung demands rather than what it delivers.
Reactive. Run to failure, fix on breakdown. It requires nothing and it is the correct policy for a genuinely non-critical item with a cheap consequence and a short repair. The mistake is applying it by default rather than by decision.
Calendar based. Fixed intervals from a code, a manufacturer recommendation, or history. This requires an asset register that is accurate, which is a larger ask than it sounds, and a work management system that can schedule against it. For pressure equipment the intervals come from the American Petroleum Institute inspection codes: API 510 for pressure vessels, API 570 for piping, API 653 for storage tanks, each with a maximum interval and a mechanism for extending it on evidence.
Condition based. Act on a measured indicator crossing a threshold. This requires instrumentation on the thing you are monitoring, a baseline for what normal looks like, and somebody who owns the alarm. The step from calendar to condition is mostly a capital and wiring problem, and it is the step most organisations are actually on rather than the one they describe themselves as being on.
Predictive. Estimate the future state and act ahead of it. This requires the condition monitoring above plus a history long enough to relate an indicator's trajectory to an outcome, which is where the requirement changes in kind. The first three rungs need instruments and process. The fourth needs a data history that includes the outcomes you want to predict, and most integrity programmes have been run precisely so that those outcomes did not happen.
Between calendar and predictive sits risk-based inspection, which is worth naming because it is often the higher-value move and it needs no new sensors. API RP 580 and the quantitative methodology in API RP 581 set inspection effort by combining a probability of failure with a consequence of failure, so that a vessel in benzene service near a control room gets attention that a nitrogen receiver does not. Reallocating a fixed inspection budget by risk usually beats spending more on the same allocation.
Rotating and static equipment are different prediction problems
Grouping them under one programme name hides the fact that they behave nothing alike.
Rotating equipment gives you a signal. A pump, a compressor, a turbine or a fan produces vibration, and the spectrum of that vibration carries interpretable structure: imbalance at running speed, misalignment at twice running speed, bearing defect frequencies determined by geometry, blade pass frequencies. Thermal, acoustic and lubrication oil analysis add channels. The degradation happens over days to months, which means a monitored machine will usually announce itself with enough time to plan a change out. The prediction problem is a signal processing and pattern recognition problem, and it is the one that machine learning approaches genuinely address, because the sampling rate is high and the time between installation and failure is short enough to observe repeatedly across a fleet of similar machines.
Static equipment gives you almost nothing in real time. A vessel or a piping circuit degrades through corrosion and cracking mechanisms that evolve over years: internal corrosion under a given fluid and temperature, corrosion under insulation, chloride stress corrosion cracking, hydrogen damage, creep in high temperature service, fatigue at supports. The observable is a wall thickness measured at a handful of locations every few years, a time series with perhaps five points in it over two decades, taken by different technicians at locations that may not be exactly the same spot.
That difference in data density changes the method entirely. For static equipment the useful model is a corrosion rate model grounded in the damage mechanism, informed by process conditions the plant already records: temperature, water content, chloride and sulphur levels, velocity, injection point locations. API RP 571 catalogues the mechanisms and the conditions that drive them, which gives you a physical basis for saying which circuits should be corroding faster rather than waiting for measurements to tell you. Online corrosion monitoring, meaning ultrasonic thickness sensors permanently mounted at known locations, turns a five-point series into a continuous one, and it is worth installing where a mechanism model says the rate is uncertain and the consequence is high.
Treating both classes with one method produces a programme that works acceptably on pumps and produces confident nonsense on piping.
A remaining useful life number is worse than a probability
Condition monitoring vendors like to output remaining useful life as a single figure. This bearing has 47 days left. The number is easy to display and the wrong shape for the decision it feeds.
Nobody acts on 47 days as such. The question a reliability engineer is actually asking is whether the item will survive until the next planned opportunity to work on it, which is a specific date already in the plan. The right output is therefore a probability of failure before that date, with an interval around it, and that probability is what supports the decision to add scope to the shutdown, to change out now at unplanned cost, or to accept the risk and monitor.
Framing it this way has a second benefit: a probability composes. With twenty items each carrying a probability of failing before the turnaround, you can compute an expected number of unplanned events and an expected cost, and compare that against the cost of pulling scope forward. Twenty separate remaining useful life figures do not combine into anything.
It also forces an honest statement of uncertainty. A point estimate of 47 days invites the reader to treat it as known. A statement that failure before the March shutdown is somewhere between 15 and 40 percent invites the correct next question, which is what would narrow the range and whether narrowing it changes the decision.
The turnaround is where the money is
In a process plant, predictive maintenance pays for itself through the turnaround rather than through individual work orders.
The arithmetic on a major unit shutdown has two parts that are worth separating on your own numbers. The direct cost is contractors, materials, rented equipment, scaffolding and inspection resources, and it is large. The deferred margin is the unit's contribution over the outage duration, and on a well-utilised unit it is frequently the larger of the two. Both scale with duration, which means the entire value of an integrity programme flows through two decisions: when the shutdown happens and what goes into its scope.
Predictive output changes both. Scope is built months in advance from a list of items where somebody believes work is needed. Every item on that list that turns out to be in acceptable condition consumed direct cost and duration for nothing, and every item that should have been on the list becomes either an unplanned event later or a discovery during the outage, which is the most expensive category because it extends a critical path already committed.
Interval extension is the other lever. Where evidence supports running an item longer, the inspection codes provide a route to a longer interval, and pushing a shutdown from a four-year to a six-year cycle removes an entire outage from a decade. That decision needs evidence a regulator or a certifying authority will accept, which is a documentation problem as much as a technical one, and it is the single highest value output an integrity data programme can produce.
The unplanned trip has a tail that is easy to miss in the cost case. A compressor trip pushes gas to the flare, and the volume from an unplanned shutdown is the part of a flaring number that no gathering investment ever fixes (N10). If you are building the value case for reliability work, that volume belongs in it.
The cost of being wrong in the safe direction
False positives get treated as harmless because the failure mode is doing extra work. They are not harmless, for three reasons that compound.
The first is that intervention itself introduces risk. Nowlan and Heap's 1978 study for United Airlines, prepared for the United States Department of Defense, examined failure modes across a large aircraft fleet and found that only about eleven percent showed the wear-out pattern that an age-based overhaul is designed to address. The largest single group showed high failure rates immediately after installation or intervention followed by a constant rate afterwards. Opening a vessel, breaking a flange, disturbing a bearing and reassembling it puts an item back on the steep part of that curve. A programme that inspects everything more often has manufactured new infant mortality.
The second is credibility. An alarm that produces a work order that finds nothing, repeated for six months, ends with the alarm being ignored. Alarm fatigue in a reliability programme has the same shape as it does in a control room, and once the operators have decided the system cries wolf, the true positive it eventually produces gets treated the same way as the rest.
The third is the opportunity cost inside a fixed shutdown window. Turnaround duration is bounded by critical path work, and scope added on a false positive competes for the same crews, the same crane time and the same confined space entries as scope that mattered. The item you did not do because the schedule was full is the real cost of the item you did not need to do.
The practical response is to price both error types before setting a threshold. A model with a tunable decision boundary should be tuned against the cost of a missed failure and the cost of an unnecessary intervention, not against an accuracy figure, and those two costs differ by orders of magnitude across an asset register.
The limit
Predictive maintenance learns from failures, and a well-run asset integrity programme produces very few of them. That is the central data problem in this field and there is no clever way around it.
A refinery might run a critical service compressor for fifteen years with two significant unplanned events. Those two events are the entire training set for that machine, they occurred under different process conditions, and one of them was preceded by a sensor that has since been replaced with a different model. Any method that requires labelled run-to-failure sequences will not find them in your own history.
Three responses are worth the effort, and none of them fully solves it. Pooling across a fleet gets you more instances of a machine type, which works where the machines are genuinely comparable and misleads where duty and service differ. Industry data collection is the formalised version of this: ISO 14224 defines how reliability and maintenance data should be recorded for petroleum, petrochemical and natural gas equipment, and the OREDA joint industry project has been pooling failure data across operators on that basis for decades, which gives population failure rates where your own records give none. Anomaly detection sidesteps labels entirely by learning what normal looks like and flagging departures from it, at the cost of telling you that something changed without telling you what will happen. Physics and mechanism based models replace the missing failure history with engineering knowledge, which is why the static equipment side of integrity has always been model driven rather than data driven and will remain so.
The honest framing for a vendor conversation is that a predictive claim on rotating equipment is testable against your own history and a predictive claim on static equipment usually is not, because there is no history dense enough to test it. In the static case what you are buying is a mechanism model with a monitoring layer, which is valuable, and it should be described that way.
Take next year's turnaround scope list and mark each item with the evidence that put it there. The proportion sitting on a code interval alone, with no measurement and no mechanism argument behind it, is the size of the opportunity and it usually surprises the people who built the list.