In short: Parasuraman, Sheridan and Wickens split a human-machine system into four stages in 2000, and most products sold as a supply chain control tower do the first two, acquiring and analysing data before handing a coloured icon to a person. The deliverable is a ranked exception queue where each item carries what it affects, what doing nothing costs, two or three priced options, an owner and a decide-by time. Ranking should ask whether an event consumes buffer that something depends on, since in most networks the majority of late inbound shipments break nothing. Sendelbach and Funk reported in 2013 that between roughly 72 and 99 percent of clinical alarms were false or clinically insignificant, and the same desensitisation is what empties a logistics exception queue.
The screen on the wall of the logistics office is showing 47 red items. At the 09:00 call somebody reads out the worst eight, everyone agrees they look bad, and three people write the same container numbers into three different notebooks. By 11:00 two of the eight have resolved themselves, one has become a real problem that nobody worked on because it was ranked ninth, and the screen is showing 51 red items.
The tower cost a lot of money and it did what it promised, which was to make things visible. Nobody's job changed, because visibility was never the part that was missing.
Visibility and control sit at different stages of the same model
Parasuraman, Sheridan and Wickens published a model in IEEE Transactions on Systems, Man and Cybernetics in 2000 that breaks any human-machine system into four stages: information acquisition, information analysis, decision and action selection, and action implementation. It was written for cockpits and process control, and it describes the control tower market with uncomfortable precision.
Almost every product sold as a control tower does the first two stages well. It acquires data from carriers, ports, warehouses and ERPs, normalises it, and analyses it into a status. Then it hands a coloured icon to a human and stops. Stages three and four, choosing the response and executing it, remain entirely manual.
The difference is visible in what an alert contains. Compare an alert that says container MSKU4471203 is running two days late with an alert that says the same container now arrives Thursday, which leaves line 2 short of a component on Friday, and here are three responses: expedite the remaining quantity by air at a stated cost and a Wednesday arrival, pull four days of cover from the Rotterdam DC at a stated transfer cost leaving them thin for a week, or reschedule line 2 to Monday and push two customer orders with named accounts and revised promise dates. The second version can be approved or rejected in a minute. The first version starts an hour of research by whoever picks it up.
Building the second kind is harder in a specific and boring way. It requires the links between the shipment, the material, the production order, the finished goods and the customer commitment to actually exist in the model. Where those links exist, the staged response falls out of them. Where they do not, no amount of visualisation will produce one.
The exception queue is the product
Everything else in a control tower is packaging around a ranked list of things a person should do today. Treat that list as the deliverable and the design questions become concrete.
Each item on the queue needs a small set of attributes, and most implementations are missing at least two of them.
What happened, in one line, with the source and the time it was detected.
What it affects, traced through the model rather than asserted. This shipment, that material, those two production orders, these customer commitments, at this value.
What it costs if nobody acts. The do-nothing option is a real option and it needs a number, because a delay that consumes slack costs nothing and a delay that breaks a promise costs a great deal, and they look identical on a map.
The options, each with a cost, a lead time and a consequence. Two or three is enough. The point is that somebody has already done the arithmetic.
An owner and a decide-by time. The second half of that is the attribute almost nobody implements, and it matters because the option set decays. Air freight booked on day one is cheaper and more likely to arrive useful than air freight booked on day four, when the same recovery also needs a partial shipment and an expedite fee. An exception with no expiry sits in a queue behaving like a task when it is really a perishable choice.
Ranking by consequence rather than by event severity
Most towers rank by the severity of the event. A typhoon warning generates a high-severity alert whether or not it touches anything you own. A carrier's two-day roll on a shipment of packaging into a DC with six weeks of cover generates the same amber icon as a two-day roll on a sole-source component feeding tomorrow's build.
The ranking that works asks a different question: does this event consume buffer that something depends on? Run the projected balance at the consuming node with the event applied, and see whether it breaches the buffer inside the horizon. If it does not, the event is a fact and belongs in a log. If it does, the size of the breach and the value of what breaks give you the rank.
That test filters aggressively, and it should. In most networks the majority of late inbound shipments break nothing, because inventory exists precisely to absorb them. A tower that surfaces all of them is reporting the normal functioning of the system as a series of emergencies.
The same consequence logic applies to supplier risk signals, which have their own scoring problem and their own literature (R2), and the planning-desk version of queue design has enough differences to be worth treating separately (P2). On the execution side the events are more mechanical: missed dock appointments, port dwell, short shipments, refused deliveries, temperature alarms, tender rejections, and detention accruing on a container nobody has scheduled to unload.
Alert fatigue is the dominant failure mode
The medical literature has already run this experiment at scale. The Joint Commission issued Sentinel Event Alert 50 in 2013 on medical device alarm safety after collecting alarm-related patient deaths, and the mechanism it described is the same one operating in every logistics office: clinicians exposed to constant alarms become desensitised, and the desensitisation is rational because most alarms are wrong. Sendelbach and Funk, reviewing the evidence in AACN Advanced Critical Care in 2013, reported that between roughly 72 and 99 percent of clinical alarms were false or clinically insignificant across the studies they examined.
Parasuraman and Riley named the pattern disuse in Human Factors in 1997, in a paper on how people misuse and abandon automation. Operators who learn that a system cries wolf stop responding to it, and the abandonment is fastest among the operators with the most work, which are exactly the ones you built it for.
Designing against this takes a few specific commitments. Set a precision target before launch, meaning a minimum share of alerts that lead to an action, and raise thresholds until you meet it, accepting that you will miss things. A queue with a 30 percent action rate that people work beats a comprehensive queue that people have stopped opening.
Suppress anything already covered by an open action, so one problem produces one item rather than one item per affected order. Group by cause, so a port closure is a single exception with 340 affected shipments attached instead of 340 exceptions. Cap the queue by design, at something like a working day's capacity per person, which forces the ranking to be honest, since an uncapped queue lets the system avoid deciding what matters by showing everything. Give people a button that says this is not an exception, and feed those clicks back into thresholds, per alert type, visibly.
The loop has to close, which means the tower has to write
If the outcome of every exception is a phone call, what you own is a notification service with good graphics. Closing the loop means the option a person selects becomes a transaction in the system that executes it. An expedite becomes a booking in the TMS. A transfer becomes a stock transfer order. A reschedule becomes a change to the production plan. A revised promise becomes an updated date on the sales order, with the change recorded against the exception that caused it.
This is where most implementations stop, and the reason is rarely technical ambition. Reading is easy and writing needs authority, validation and a reversal path. Every write-back needs a value cap, a permitted scope, an audit record, and a way to undo it, which are the same requirements that apply to anything acting on your systems without a human in the loop for each instance (E1).
The staging that works is narrow and early. Pick one action type that is high volume and low value, a dock appointment reschedule or a transfer order under a cap, and build the whole write path for it: permission model, audit record, reversal, exception attribution. Prove the loop closes on something boring, then widen the scope. Teams that defer all write-back to phase two generally do not reach phase two, because the business case was consumed by the visibility phase.
What to measure, and what to stop measuring
Two numbers tell you whether the thing is working. The first is time from signal to decision, measured from the event occurring rather than from the alert appearing. Detection latency is easy to improve and mostly a question of data feeds. The interval that matters runs from the world changing to a person committing to a response, and in most operations it is measured in days for reasons that have nothing to do with data availability.
The second is the share of alerts that resulted in an action. Track it per alert type and retire any type that stays below a floor for two months, because that alert type is training people to ignore the queue. This number is also the honest way to evaluate a tower after a year: if nine in ten items are dismissed unread, the tower is generating work.
Behind those, two supporting measures. Queue age distribution, because items sitting for five days are either not exceptions or have no owner. And cost avoided per exception, calculated against the do-nothing option that was priced in the alert, which is only credible if the option set was stored at decision time rather than reconstructed afterwards.
Stop reporting alert volume, feed latency in seconds and dashboard views. All three improve when the tower gets worse.
Where this stops
A control tower sitting on top of systems that cannot execute a change is an expensive way to watch things go wrong. If the TMS will not accept a re-plan after tender, if the plant schedule is frozen three weeks out by policy, if the supplier only takes changes by email and answers in two days, then the decision the tower produces has nowhere to go and the loop stays open no matter how good the ranking is.
That is testable before you buy anything. Take five exceptions from last month, trace what would have had to change in which system and at what point in time, and check whether that system accepts a change at that point. If three of the five end at a person sending an email, the constraint is execution capability and the tower will not touch it.
The other dependency is the model underneath. Consequence ranking needs the chain from shipment to material to order to commitment. Where those links are missing, a tower can only rank on event attributes, which returns you to sorting by how loud the event was.
There is an organisational version of the same limit. A tower that prices the do-nothing option makes the cost of hesitation visible and attributes it, which changes who is accountable for a delay. That is the point of it, and it is also the reason some implementations quietly settle back into being a screen on a wall.
Pull last month's exceptions and measure what share of them ended in a recorded change to a system of record, because that ratio is the honest current state of your loop.