In short: The number of exceptions a queue produces is a design parameter rather than a property of the business, and a list longer than the time available to work it detects problems without resolving any of them. Thresholds set from statistical deviation generate volume, while thresholds set from consequence, meaning what happens if nobody acts, generate a list somebody can finish. Twelve exceptions caused by one supplier delay are one exception, so grouping by cause shortens the queue without lowering the bar. Ageing is what kills a queue, since unworked lines accumulate until the planner abandons the list for the shortage report, which makes clearance rate the measure that matters rather than detection count.
The exception list opens on Monday with four hundred and eleven lines on it. The planner has maybe an hour before the supply review, and the list is sorted by item number, so they work down from the top and get through eleven or twelve. The rest sit there.
On Tuesday the list has four hundred and thirty. The twelve worked lines have dropped off, thirty new ones have arrived, and the four hundred that nobody read are still there, one day older. By Thursday the planner has stopped opening it at all and is working from the shortage report and their inbox, which is a smaller and more honest version of the same queue.
Nobody involved is doing anything wrong. The queue was designed to detect problems, it detects them, and detection was never the constraint.
The number of exceptions is a design parameter
The queue length is treated as an output of the thresholds. It should be treated as an input, and the thresholds should be derived from it.
Work the arithmetic from the wrong end deliberately. A planner covering a desk has some number of hours a week that are genuinely free of meetings, data fixing and the rest, and on most desks that is a good deal less than people assume. Take the free hours, decide how many minutes a real exception deserves, and you have a weekly capacity in lines. Multiply by the number of planners and you have the queue you can afford. Anything the system generates beyond that number is not being worked, whatever the report says.
Then set the thresholds to hit the number. This feels like cheating, since you are choosing how many problems to see. What you are actually choosing is which problems to see, because the alternative is a queue that chooses for you by putting the important ones on page nine.
The process industries settled this argument decades ago and wrote it down. EEMUA Publication 191, first issued in 1999 and now in its fourth edition, sets a benchmark for control room operators of roughly one alarm per ten minutes in steady operation, and treats anything above ten alarms in ten minutes as a flood in which the operator cannot function. ANSI/ISA-18.2, published in 2009 and revised in 2016, builds the standard around rationalisation: every alarm has to have a defined consequence and a defined operator response, and an alarm with no response defined for it gets removed rather than kept for information.
That last rule is the one worth importing. If you cannot say what the planner is supposed to do about a given exception type, you have a data point rather than an exception, and it belongs on a dashboard where it costs nobody a decision.
Set the threshold from consequence rather than deviation
Most exception configurations trigger on statistical distance. Demand deviated more than three standard deviations from forecast. Inventory fell below the reorder point. Forecast changed more than twenty percent versus last cycle.
Those are measures of surprise. What the planner needs is a measure of harm, and the two come apart constantly.
A three-sigma demand spike on an item carrying twelve weeks of cover requires nothing from anyone. A one-sigma drift on a constrained line running into a promotion with two weeks of cover is the most important thing on the desk. A statistical trigger ranks the first above the second and will do so every week, and after a few weeks the planner learns that the top of the list is not where the problem is.
Build the trigger from a projection instead. Run the plan forward, and raise an exception where the projection crosses something that costs money: a stockout inside the exposure window, an excess position that will not clear before shelf life or the disposition horizon, a supply commitment that cannot be met, a purchase order that will arrive after the demand it was raised for. Then rank by the size of the consequence rather than by the size of the deviation.
This is more work to configure and it changes the character of the queue completely. The list stops being a list of things that look odd and becomes a list of things that will cost money if left alone, which is a list people finish.
One practical note on the projection. It has to run against the current plan including everything in flight, not against a snapshot from the overnight batch, or you will raise exceptions for problems that were solved yesterday afternoon. Those are the fastest way to teach a planner that the queue is not worth trusting.
Twelve exceptions from one supplier delay are one exception
A single event fans out. A supplier confirms a four-week delay on one component, and the queue produces an exception for every finished item that uses it, every location expecting stock, and every forecast period the shortage touches. The planner sees forty lines. The decision in front of them is one decision.
Grouping is the highest-return change available in most exception configurations, and it has three parts.
Group by cause. Where several exceptions share a root event, present the event with its consequences attached rather than presenting the consequences separately. The planner should be reading a supplier delay with a list of what it hits, and taking one action.
Deduplicate across the horizon. The same item flagged in week three and again in week five of the same projection is one problem with a start and an end, and showing it twice doubles the queue without adding information.
Suppress where an action is already in flight. An item with an expedite raised, a transfer in transit or a supplier commitment logged against it does not need re-raising every night until the stock lands. Suppress it, hold it against the expected resolution date, and bring it back loudly if that date passes without the problem clearing.
Suppression makes people nervous, and the nervousness is reasonable, since a suppressed exception that never comes back is a silent failure. The safeguard is that suppression must always carry an expiry and a re-raise condition. Nothing gets suppressed permanently, and nothing gets suppressed without something that will bring it back.
Ageing is what actually kills the queue
An exception that is not worked in the period it was raised stays in the queue, and the sitting is the problem.
The mechanism is straightforward. Once the queue contains a large body of old lines that nobody has ever worked, the planner cannot distinguish the new and urgent from the old and ignored by looking at it. The rational response is to stop reading the list and work from whatever channel is producing real signals, usually email and phone calls from the people affected. At that point the queue has become a compliance artefact, and it will keep generating lines and consuming licence cost for years.
Three mechanics prevent it.
Give every exception type an explicit life. If a line has not been worked within its window, it either escalates to a named person with the reason it was missed, or it closes with a recorded outcome of no action taken. Both are better than sitting.
Re-score rather than re-raise. An exception raised last week may be more urgent now or entirely irrelevant, and the queue should reflect the current projection rather than the state at the moment of first detection.
Watch the age distribution as a first-class metric. The mean age of open exceptions rising over consecutive weeks is the earliest available signal that the queue has gone past what the desk can absorb, and it moves well before anyone complains.
Measure clearance rather than detection
Most exception reporting counts what was raised. That number measures the sensitivity of the configuration and nothing else, and it rises when the queue gets worse.
Four measures are worth putting on a page.
Clearance rate. The share of exceptions raised in a period that were worked to an outcome within that period. Below a reasonable rate, everything else you do to the queue is decoration.
Age profile of the open set. Look at the distribution rather than the average, because the tail is where the abandoned lines hide.
Action rate. The share of worked exceptions that resulted in a change to the plan. An exception type that is repeatedly examined and closed without action is a threshold that is set wrong, and it should be retuned or retired.
Miss rate. The service failures, expedites and write-offs that had no exception raised against them beforehand. This is the only measure that catches thresholds set too loosely, and it is the one nobody builds, because it requires joining outcomes back to the queue after the fact.
Action rate and miss rate have to be read together. Tightening thresholds to raise the action rate will eventually raise the miss rate, and the point of holding both is that you can see the trade rather than discovering it in a review six months later. The alarm literature has the same pair under different names, and The Joint Commission's 2013 alert on medical device alarm safety put the share of hospital alarm signals not requiring clinical intervention somewhere in the range of 85 to 99 percent, which is what an action rate looks like when only the detection side has ever been tuned.
Where this stops
A well-designed queue surfaces the right problems in the right order. It cannot create the capacity to fix them.
That distinction matters when the tuning stops working. If you have set thresholds from consequence, grouped by cause, suppressed what is in flight and aged out what is dead, and the queue is still not clearable, you have finished the tuning problem and arrived at a resourcing one. The honest answer at that point is that the desk is carrying more decisions than the people on it can take, and the options are more people, fewer decisions through automation, or an explicit choice to leave a defined segment unmanaged and accept the service that produces.
That third option is legitimate and it is almost never made explicitly. Every desk already leaves part of its catalogue unmanaged. Making it a decision, with a named segment and an accepted service level, is better than leaving it to whichever items happen to sort low.
There is also a limit on what any of this fixes upstream. A queue full of exceptions caused by wrong lead times and stale parameters is a symptom, and retuning the queue makes the symptom quieter without touching the cause. When one exception type dominates the volume, the useful question is what is generating it rather than what threshold would show less of it.
Pull last month's exception log, count how many lines were raised, and count how many were closed with a recorded action. The gap between those two numbers is your real starting point.