In short: Planning implementations fail without an outage or an escalation, through planners who hit a few obviously wrong recommendations and quietly move the work back into spreadsheets. Trust in a planning system is lost far faster than it is rebuilt, so the first weeks after go-live decide adoption. Most of the underlying faults sit in master data, clustered in lead times that describe an old sourcing arrangement, bills of material carrying substituted components, and history holding one-off orders the model reads as seasonality. Override rate over time and export volume to spreadsheet are both in the system logs, and they measure adoption in a way login counts never will.
Planning implementations do not fail loudly. There is no outage, no escalation, no post-mortem.
What happens is that a planner opens the new system, works through their category, and finds three items where the recommendation is obviously wrong. A lead time that has not been right since 2022. A bill of materials with a component that was substituted last year. An item whose demand history includes a one-off order that the model has learned as seasonality.
They fix those three by hand. Next week there are four more. Within a month they have a spreadsheet that holds their real numbers, and the system has become a place they export from. Nobody tells anyone, because from their point of view they are doing their job.
Six months later, adoption metrics look fine because people are logging in. The plan the business actually runs on is in a spreadsheet on somebody's laptop, and it will stay there.
The mechanism is trust, and it is asymmetric
The thing that makes this hard to manage is that trust in a planning system is lost much faster than it is built.
A planner who sees ten good recommendations and one absurd one remembers the absurd one, and reasonably so, because the absurd one is evidence that the system can be confidently wrong in ways they cannot predict. A number that is quietly wrong is worse than no number, because they have to check everything to find the ones that need checking.
This has been measured outside supply chain and the result is stronger than most implementation plans assume. Dietvorst, Simmons and Massey ran a series of experiments published in 2015 in which people chose between their own judgement and a statistical model. Participants who had seen the model make a mistake abandoned it far more readily than participants who had never seen it perform at all, and they did so even when they had also seen that the model outperformed them on average. Seeing the error mattered more than seeing the score. A planning rollout puts every planner in exactly that condition within the first fortnight.
The same authors published a partial remedy in 2018. People were considerably more willing to use an imperfect model when they were allowed to adjust its output, and the effect held even when the permitted adjustment was small. That points at a design choice most implementations get backwards. A system that locks its recommendations to protect them from human interference produces planners who route around it entirely, while one that allows a bounded, recorded adjustment keeps the work inside the system where it can be measured. The adjustment is also the data you need later, because an override you can see is a signal and an override made in a spreadsheet is invisible.
That asymmetry has a direct implication for sequencing. Getting the data right before go-live is worth more than getting the model right, because model quality is a gradual improvement and data quality is a threshold effect. Below the threshold, nobody uses the output regardless of how good the model is.
The threshold is a real one and you can locate it with arithmetic. A planner reviewing 200 recommendations a week, taking about ninety seconds to verify one against the source systems, spends five hours to check the lot. If three percent of the underlying records are wrong, those five hours catch about six errors, which is fifty minutes of effort per error found. That is defensible, and a planner who does the sum implicitly will keep checking, which means they are not using the system so much as auditing it.
Now cut the error rate to three in a thousand. The same five hours now catch about half an error a week, which is over eight hours of checking per error found. Checking stops being worth anyone's time, and the rational move flips from verify to trust. One order of magnitude in data quality moved the behaviour completely, and nothing about the model changed in either case. That is what a threshold effect looks like in practice, and it is why data work sequenced after go-live arrives too late to matter.
What actually breaks
The failure is nearly always in master data, and it clusters in a small number of places.
Lead times that describe an old supplier arrangement. Nobody updates the planning parameter when sourcing changes, because nothing breaks immediately.
Bills of material with substituted components. Engineering makes a change, the substitution happens on the shop floor, and the planning BOM keeps the original.
Duplicate item and supplier records. The same supplier under three spellings, so spend and risk analysis fragment and nothing aggregates correctly.
Units of measure inconsistencies. Cases against eaches, pallets against layers, and a conversion factor that is right for most items and wrong for the ones that were set up in a hurry.
Demand history containing one-off events. A single large order to a customer who no longer exists, sitting in the history as evidence of a seasonal peak.
Calendars that were correct when configured. Shift patterns, holidays and shutdown windows that have moved.
Each is individually trivial. Together they produce a system whose output is wrong often enough that checking becomes rational, and once checking is rational, the spreadsheet has won.
Run the checks continuously
The productive move is to treat data health as a continuously running diagnostic rather than a one-off cleanse during implementation.
A useful check set is not exotic. Items with no demand history but active status. Items with demand and no lead time. Lead times outside a plausible band for their supplier and mode. Bills of material with components that have no supply source. Duplicate suspects by string similarity across item and supplier masters. Demand history containing single observations more than several standard deviations from the rest of the series. Units of measure conversions that fail a round trip. Calendar entries that have not been updated in over a year.
Each check produces a list, each list has an owner, and the count over time is the metric. The count going down is the leading indicator of adoption, and it moves before any adoption metric does.
Two things make this work rather than becoming another report nobody reads.
Rank by consequence. A wrong lead time on a high-volume constrained item matters enormously; the same error on a discontinued line does not. Sorting by exposure means the finite remediation effort goes where it changes outputs. Exposure here is a specific quantity rather than a judgement: annual demand multiplied by unit cost, multiplied by how sensitive the recommendation is to the field in question. Sort the failed-check list that way, take the top fifty, and the first pass of remediation is a week of somebody's time covering the great majority of the value at stake. Publishing that ranking also settles the argument about whether data quality is a project or a permanent function, because the list refills.
Show the check result next to the recommendation. When a planner sees a suggestion accompanied by a note that this item's lead time has not been updated in three years, two things happen. They know how much to trust it, and the fix gets raised by the person who noticed rather than sitting in a queue.
The other half is the process mismatch
Data explains most of the failures and not all of them. The second cause is configuring a system around a future-state process while the business continues to run the current one.
This happens for understandable reasons. The implementation is an opportunity to improve the process, the target operating model is designed, and the system is built to support it. Then the process change does not happen at the pace the system assumed, and planners are left operating a tool that expects a way of working the organisation has not adopted.
The symptom is a system that is technically working and operationally awkward: steps that assume an input nobody produces, approvals that assume a role nobody holds, a cycle that assumes a cadence nobody runs.
There is a query that finds this in about ten minutes. Look at the workflow or approval tables for items sitting in a pending state, and compare the age of each against the planning horizon it was meant to influence. A healthy queue clears well inside the cycle. A queue where a substantial number of items have been pending longer than the horizon they affect is telling you those approvals are being routed to a role the organisation never staffed, and the planners downstream have long since worked out how to proceed without them. Every one of those items represents a step in the designed process that the real process has already deleted. Deciding deliberately whether to staff the role or remove the step is a half-day of work, and leaving the queue in place teaches everyone that parts of the system can be ignored, which is a lesson that generalises fast.
The practical protection is to build for the current process with the future one as a configuration option, rather than the reverse. It feels like a compromise and it means the system is usable on day one, which is when trust is decided.
What to measure
Login counts are not adoption. Three better signals, all available from system logs.
Override rate over time. Some overriding is healthy. A rate that stays flat or rises after go-live means the system's output is not being trusted, and the value add score on those overrides will tell you whether the planners are right to distrust it.
Export volume. A high and sustained rate of exports to spreadsheet is the clearest available signal that the real work is happening elsewhere. This is easy to measure and rarely measured, possibly because the answer is uncomfortable.
Time from recommendation to action. If recommendations sit unactioned, the system is producing output nobody is using. Distinguish between rejected and ignored, because they mean different things.
The limit
Not every failure is a data problem or a process problem. Sometimes the tool genuinely cannot do what the business needs, and planners recreate logic in spreadsheets because the system has no way to express it.
The distinguishing signal is what the spreadsheet contains. If it holds corrected inputs, it is a data problem. If it holds logic the system does not have, the system is the wrong shape for the business, and no amount of change management fixes that.
There is also a limit to what any of this can do about an implementation that was scoped to a budget rather than to a requirement. If the configuration was cut back to fit a number, the gaps were decided at that point, and the planners are being asked to compensate for a decision made before they saw the system. That is worth naming honestly in a review rather than treating as an adoption failure, because the people being blamed are the ones working around it.
Start by counting exports to spreadsheet last month. It is a two minute query and it tells you where the plan actually lives.