In short: There is no benchmark number of items per planner, and the same person can reasonably own a few hundred item locations in one configuration and several thousand in another. Span of control is an equation with two levers, the share of lines a planner actually opens in a cycle and the average time spent on each, so capacity moves by changing either one. Reviewing and deciding are different activities with different costs, and most measured planner time goes on reaching a view rather than on making the call. Touch rate, the count of distinct item locations opened divided by the number on the desk, is what the capacity question turns on, and it comes straight out of the audit log.
The question arrives in a headcount conversation, usually late. Somebody wants to know whether the team can absorb the new region without another head, and somebody else says that a planner can handle around two thousand SKUs.
Nobody in the room knows where that number came from. It gets used anyway, because the alternative is admitting that the team has never measured what a planner's week contains, and the number sounds like the kind of thing that ought to be a benchmark.
There is no benchmark. Span of control on a planning desk is a function of a few things you can measure, and the same planner can reasonably own three hundred item-locations in one configuration and eight thousand in another.
Where the time goes when you measure it
Four categories account for nearly all of a planner's week, and the labels matter because each one has a different route to removal.
Data correction. Fixing inputs the system should have had right. Overriding a lead time the planner knows is wrong, adjusting a history point that contains a one-off order, correcting a unit of measure before the calculation runs. This work is invisible in every process map, because officially it does not exist.
Exception chasing. Working the queue, plus the part of the queue that is really communication: emailing a supplier for a confirmed date, calling a site to ask what actually arrived, waiting for the answer, and coming back to it twice.
Preparation. Building the pack for the review, reconciling two numbers that came from two systems, and constructing the explanation for a variance that will be discussed for ninety seconds.
Judgement. The part where the planner knows something the model does not, and acts on it. A customer's tender result, a competitor delisting, a launch date that moved.
In every desk where I have seen this measured properly, the judgement category is the smallest of the four and often by a wide margin. That result is uncomfortable and it is also the useful one, because the other three categories are all removable in ways that judgement is not.
The corollary is worth stating plainly. When a planning team is under-resourced, the thing that gets squeezed is judgement, since the other three arrive with deadlines attached. The desk keeps functioning and quietly stops doing the only part of the job that required a person.
Span of control is an equation with two levers
Write it out and the drivers become obvious.
A planner has some quantity of genuinely available minutes in a week, after meetings, holidays and the rest. Of the item-locations on their desk, some proportion requires a human touch in any given cycle. Each touch costs some number of minutes. Capacity is available minutes divided by minutes per touch, divided again by the touch rate.
Put illustrative numbers on it, and use your own rather than these. Twelve hundred available minutes a week. Four minutes per touched line, which is generous for a routine review and mean for anything requiring a phone call. That is three hundred touches. At a touch rate of forty percent, the desk can be seven hundred and fifty item-locations. At a touch rate of five percent, the same planner with the same minutes can own six thousand.
That single sensitivity is the whole argument. There are exactly two levers on span of control, the touch rate and the minutes per touch, and everything people usually propose either moves one of those or moves nothing.
The touch rate is set by automation level, by data quality and by the demand character of the desk. Automation removes lines that need no decision. Data quality removes the correction work and, more subtly, removes the checking, since a planner who has been burned by bad inputs checks everything. Demand character sets a floor, because a desk that is mostly promoted, intermittent or newly launched items will never have a low touch rate no matter how good the system is.
Minutes per touch is set by the tooling and by how much context has to be assembled before the decision can be made. A planner who has to open three screens and a spreadsheet to understand one exception is paying most of the cost in assembly rather than in deciding.
Neither lever is moved by working harder, and neither is moved by a reorganisation.
Reviewing and deciding are different activities
A planner who looks at two thousand lines and changes forty has performed two thousand reviews and forty decisions. The forty decisions are the job. The two thousand reviews are the cost, and they are where the week went.
The design goal is to cut reviews without cutting decisions, and the two get conflated constantly, usually by someone proposing that planners should "look at everything" as a matter of diligence. Looking at everything is how nothing gets looked at properly.
The measurement that makes this visible is a change rate per reviewed line, available from the audit log in most systems. If a planner opens six hundred items in a week and modifies twelve of them, the review is producing a change two percent of the time, and the other ninety-eight percent is an inspection with no output. Some of that inspection has value as assurance. Most of it does not, and the share that does can be replaced by a sample.
What makes a review feel mandatory is usually trust rather than policy, which is a separate subject with its own post. A planner who has seen the system be confidently wrong will check everything, and they are behaving rationally. Cutting the review rate therefore depends on the data underneath being demonstrably sound rather than on instructing people to check less.
Fildes, Goodwin, Lawrence and Nikolopoulos, in their 2009 evaluation across four supply chain companies, found that the great majority of statistical forecasts were adjusted before release. That is a review rate approaching one, and it is the condition most demand desks are still operating in.
What to remove first
The order below is by ease of removal rather than by size, because a capacity programme that starts with the hardest item never reaches the second one.
Start with the lines where the human touch has already been scored as adding nothing. This is the largest single block on most demand desks and it has the cleanest evidence behind it, and the operating decision about which lines flow through untouched has its own post.
Then take the data correction the planner is doing on behalf of another function. Every lead time a planner overrides by hand is a field somebody else owns, and the correction is being applied at the point of use instead of at the source, which means it will be applied again next week by somebody else on a different item. Moving that work to the owner is a governance question and it removes real hours.
Then remove the exceptions that were never actionable. A queue tuned on detection generates volume that consumes attention and produces no decisions, and retuning it is cheap relative to what it returns.
Then attack preparation. Most review packs are assembly rather than analysis, and most of the assembly is a report that could be standing. This is the least glamorous item on the list and frequently the fastest win, because a single recurring pack can cost a planner half a day a month with nobody having ever priced it.
Last, and hardest, look at overrides that have been scored as neutral or negative and are being made out of routine. This is the one that touches how people see their own job, and it should be handled with the evidence in front of them rather than as a policy.
Structuring the desk around segments
Most planning teams are organised by brand, category or geography, because that is how the commercial organisation is arranged and because it gives the sales director one name to call.
The cost is that a brand desk contains a stable high-volume line and an erratic slow mover in the same head, and the planner switches method, cadence and buffer philosophy dozens of times a day. Every switch has a cost, and the working method never becomes routine because it is never the same twice.
A desk organised around segments lets the method be constant. One person owns the automatable tail across every brand and spends their week on guardrails, monitoring and the exceptions that escape, which is a different job from the one their colleague is doing on the volatile, high-consequence lines that need judgement and relationships.
This works where the catalogue is large, the segmentation is stable, and a meaningful share of the tail is genuinely automatable. It works badly in small teams, where a segment desk means one person doing one thing and no cover when they are away. It also works badly where the commercial relationship is the primary demand signal, since a planner who never speaks to the account team loses the information that made them useful.
A workable middle position is to keep commercial ownership arranged by brand and move only the automatable population onto a single shared desk. The planners keep their relationships and their judgement work, and the volume that needed neither stops consuming their attention.
The limit
Measuring planner time changes planner behaviour while you are measuring it. This has been known since the Hawthorne studies at Western Electric between 1924 and 1932, and it applies with full force to any exercise where people record their own activity for a week while their manager watches.
Self-reported time also runs high in a specific direction. Robinson and Godbey's time-use research in the 1990s found that self-reported working hours exceed diary-based measurement systematically, and the gap grows the longer the reported week. A planner asked how long they spend on data correction will give you an honest answer that is wrong, usually low on the tedious parts and high on the interesting ones.
Use system telemetry for the mechanical categories. Records edited and by whom. Exceptions opened and closed, with timestamps. Session duration by screen. Override counts and sizes. Export events. All of it already exists in the audit log, none of it needs anyone to fill in a form, and it will not change what people do because they are not aware of it happening.
Telemetry has its own blind spot, which is that it counts events and not thinking. The most valuable half hour a planner spends in a week might be a phone call that generates no records at all, and a capacity model built purely from click data will under-count exactly the work you are trying to protect. Use telemetry for the three removable categories, and accept self-report or observation for the judgement remainder, knowing what it is worth.
There is also an honest limit on what any of this settles. A capacity number derived this way tells you what the desk can carry under its current configuration. It does not tell you what the right configuration is, and a team that measures itself carefully and changes nothing has produced a very well documented account of a problem.
Pull the audit log for last month, count the distinct item-locations each planner actually opened, and divide by the number on their desk. That ratio is your touch rate, and it is the number the whole capacity question turns on.