In short: Which lines should flow through untouched is a measurable question, answered by scoring touched against untouched forecasts on actuals by segment rather than by writing a policy. The guardrails are the product, meaning the per-line data preconditions, the value and variability limits, and the rules that return a line to a human when its conditions change. A staged rollout with exit criteria agreed before it starts is what keeps the programme survivable, since the failure mode is one bad week that ends the initiative. Automatic release depends on the data feeding it, so the defect position on the fields behind the decision has to be established and the material gaps closed before the population is extended.
A demand planner opens the cycle with nine hundred item-locations in front of them. They adjust something on about three hundred. Most of those adjustments are small, a few percent up or down, applied while scrolling, on items whose demand has done nothing unusual for two years.
Ask why any particular one of the three hundred was changed and the answer is usually reasonable in the moment and hard to defend in aggregate. The number looked slightly low. The customer mentioned something. Last year it came in higher.
The useful question about that desk is which of those nine hundred lines should never have arrived in front of a person at all. Speeding the planner up leaves the same work in front of them, and the question of which lines belong there has an answer you can compute rather than argue about.
Find the segment by measurement rather than by policy
The common approach is to decide by rule. Automate the C items. Automate anything below a value threshold. Automate the stable ones. Each of those rules is picking a proxy for the thing you care about, and each picks the wrong set.
The thing you care about is narrow and directly measurable: the population where the untouched statistical output performs at least as well against actuals as the released output that a human adjusted. That is a comparison between two forecasts you already have, scored against outcomes you already know, and the method for scoring stage by stage is a settled subject covered elsewhere on this site.
Set it up carefully in four respects.
Score at the grain where the touch happened. Adjustments made at item-location and evaluated at national monthly will look harmless, because aggregation absorbs them.
Score over at least a full year, ideally including a promotional season and whatever passes for peak in your business. A quarter is not long enough to separate a stage that adds value from one that had a quiet three months.
Split by adjustment size as well as by segment. This is where the result gets interesting. Fildes, Goodwin, Lawrence and Nikolopoulos published an evaluation in 2009 covering tens of thousands of forecasts across four supply chain companies, and found that the large adjustments made in response to real information tended to help while small routine ones tended not to, with the direction of the adjustment mattering as well. Any desk that has never looked will usually find the same shape in its own numbers, and the small habitual tier is where most of the volume sits.
Track bias separately from error. A touch that leaves error unchanged while shifting the whole distribution upward is filling a warehouse, and an error-only comparison will call it neutral.
What comes out is a map of the catalogue with a boundary drawn on it by evidence. Segmentation gives you the frame to describe that boundary in, since the eligible population usually clusters in particular cells rather than scattering, but the segmentation does not determine it. A high-value, stable, high-volume line is frequently the single most automatable item on a desk, and no policy written around value would ever have selected it.
The guardrails are the product
Deciding which lines can flow untouched is the easy half. Making automatic release safe is the half that determines whether it survives its first bad week.
A value ceiling on the individual decision, and a second one on the run. One line can only commit so much, and the whole automated batch can only move so much in a cycle. The second limit is the one people forget, and it is what stops a model fault from becoming a purchasing event.
A deviation limit against the previously released plan. Absolute and percentage, both, since a percentage limit is useless on small numbers and an absolute limit is useless on large ones. Anything outside goes to a human. This single rule catches most of what actually goes wrong, because a model failing tends to produce a number that is obviously different rather than subtly different.
Lifecycle and event restrictions. No items inside their launch window, no items in phase-out, no items in a live promotion period, no items under allocation or supply constraint. Each of these is a state where history stops describing the future, and each is knowable from data you already hold.
Data health preconditions on the specific line. The item's lead time has been confirmed inside a defined window, there is no open master data defect against it, there is enough clean history to have fitted anything. A line that fails its preconditions leaves the automated set until it passes, without anyone deciding.
A circuit breaker on aggregate behaviour. If the automated population as a whole moves more than a set amount in one cycle, or if the recent accuracy of that population degrades past a threshold, the release holds and a person is told. Individual guardrails catch individual faults. This one catches the drift that looks reasonable line by line and is wrong in total.
An escalation path with a named owner and a response time. An escalation that goes to a shared mailbox is a suppression with extra steps, and the queue it lands in has to be designed to be finishable rather than appended to whatever the desk already fails to clear.
One structural point applies to all of them. The guardrail has to live in the code path around the decision rather than in a policy document, a training deck or a model instruction. A limit that depends on the system choosing to respect it is not a limit, and the same argument applies to the agent conversation more broadly, which is covered elsewhere.
Roll it out in three stages, with exit criteria set in advance
The failure mode here is a pilot that runs indefinitely because nobody defined what success looked like before it started. Write the exit criteria for each stage first, in numbers, and agree them with the people who will have to live with the result.
Stage one, recommend. The system produces the answer, the planner sees it, the planner decides. Nothing changes operationally. What changes is that you start recording what the planner did to each recommendation and why. Exit criterion: on the candidate population, the planner's changes have been shown over a defined period to add nothing measurable against actuals.
Stage two, release with review. The system releases automatically, and a person reviews after the fact rather than before. Review a random sample plus every escalation and every guardrail trip, and record specifically what the reviewer would have changed. Exit criterion: the reviewer changes nothing on a large majority of the sample, and the changes they do make cluster in an identifiable subset you can move back out of scope.
Stage three, automatic with sampling. Full release, with a continuing random sample reviewed for drift and a control group held back deliberately.
Each stage should run through at least one seasonal or promotional cycle before it advances. A touchless configuration that has never met a peak has not been tested, and the fastest way to lose the whole programme is to advance a stage in September and discover the gap in November.
What to watch once it is on
Rolling performance against a control group. Keep a randomly selected slice of the eligible population on the manual process permanently. It costs a little planner time and it is the only thing that answers the question of whether the automation is still earning its place, as opposed to whether it was earning it when you switched it on. Businesses that skip the control group cannot distinguish a degrading model from a changing market, and they usually find out during a bad quarter.
Guardrail trip rate, by rule. Every rule should trip sometimes. A rule that has never fired is not protecting anything and should be tightened or removed, since its presence in the design is providing false comfort. A rule that fires on a large share of lines is mis-set and is quietly returning the population to manual.
Escalation volume and time to resolve. Rising escalation volume is the earliest sign the eligible boundary has drifted. Rising resolution time means the escalations are landing somewhere with no capacity, which converts your safety mechanism into a delay.
Membership churn. Items enter and leave the eligible set as their behaviour changes. A stable membership is suspicious; it usually means the eligibility test is running once and being cached.
Service and inventory outcomes on the automated population. Accuracy is the intermediate measure. What you promised the business was service held at lower cost and lower effort, and that is the pair to report.
The reluctance is the real constraint
Two things are consistently true when a desk runs this analysis honestly. The eligible population is much larger than anyone expected, often a substantial share of the lines and a smaller share of the value. And the organisation does not want to release it.
The stated objection is usually risk. The real objection is usually accountability, and it deserves a straight answer rather than a change management programme. A planner who has been personally answerable for a number is being asked to be answerable for a process that produces the number without them, and that is a genuine change in what their job is.
The answer that works is to move the ownership explicitly rather than letting it evaporate. The planner owns the eligibility boundary, the guardrail settings, the escalation queue and the monitoring. They are accountable for the design of the automated set rather than for each line inside it, and that is a defensible position in a review in a way that "the system did it" is not.
It also helps to be honest about what the manual process was actually delivering on those lines. A touch applied in four seconds while scrolling was never a considered decision, and the measurement usually shows it. Framing automation as removing scrutiny is only accurate if the scrutiny was real.
Where this stops
Touchless is defensible only where the inputs underneath it are trustworthy, and that makes a data health baseline a prerequisite rather than a workstream running alongside.
The reason is about speed rather than severity. A planner working manually acts as an accidental filter on bad master data, because a recommendation built on a lead time that has been wrong since 2022 looks wrong to somebody who knows the item, and they fix it quietly and never report it. Remove the person and remove the filter, and the systematic input error now flows straight through into orders at full rate. The error was always there. Automation changes how fast it arrives and how long it runs before anyone notices.
The practical sequence is to establish the defect position on the fields that feed the decision, close the material ones, and put the standing checks in place before extending automatic release to a population, with the per-line data preconditions above as the ongoing enforcement. This is a real dependency and it is worth resisting the pressure to run the two in parallel, because the parallel version delivers automation on top of unverified data and produces exactly the bad week that ends the programme.
Take one segment where you believe the statistical output is already good, score the touched and untouched forecasts against actuals for the last twelve months, and count how many lines the human hand did not improve.