In short: Most planning hierarchies are inherited from finance or merchandising, which grouped items by margin structure or by how a shop is laid out rather than by how demand behaves. The grouping that helps a forecast puts items that substitute for one another in the same node, because the total across substitutes is stable in weeks when the members are not. Storing attributes on the item and generating groupings as views handles the questions that cut across the tree, which a single fixed hierarchy cannot answer without a data migration. Reclassification quietly restates history, so membership wants an effective date and a place on the calendar. The level to forecast at is settled by running the same model at two levels and scoring both after disaggregation at the level you actually order at.
The category review asks why the forecast for household cleaning dropped fourteen percent against last year when sales have been flat. Nobody in the room can reproduce it. Two hours later it turns out that sixty items moved out of household cleaning and into a new laundry subcategory during the last quarterly refresh, the item master carries only the current parent, and the comparison is between this year's smaller category and last year's larger one. Sales did not move. The boundary did.
Everyone accepts the explanation and nobody fixes anything, because the hierarchy is owned by merchandising, the comparison is owned by finance, and the forecast that reads both is owned by planning.
The hierarchies you already have
By the time planning arrives, the catalogue has usually been classified three or four times over. Finance has a product line structure built around margin reporting and legal entity. Merchandising has a category tree shaped by how the shop or the catalogue is laid out, which is a statement about shopper navigation. The ERP has a material group that exists mainly to drive account determination and tax codes. Marketing has brand and sub-brand, which follows the advertising budget. Larger businesses also carry an external standard somewhere, UNSPSC for procurement classification or the GS1 Global Product Classification for trading partner data exchange, usually populated once for a compliance reason and never revisited.
Each of those was built well for its own purpose. None was built around how demand behaves, and that is the property a planning hierarchy needs, because the whole reason to aggregate is that the total is easier to predict than the parts.
The default move is to adopt whichever tree looks most complete and plan on it. That works when the categories happen to line up with demand drivers, which happens more often in industrial distribution than in consumer goods, where a merchandising category can contain a premium seasonal line and a private label staple that share a shelf and nothing else.
What a planning hierarchy is actually for
Three jobs, and they pull in different directions.
Aggregation for forecasting. You aggregate to find a level where the signal is stronger than the noise, then bring the result back down. The grouping helps when the members are correlated with each other in a way the model can use.
Grouping for policy. Service targets, review periods, segmentation classes and planner ownership all attach to nodes. This job wants stable groups of manageable size, which is a different requirement from the first one.
Reconciliation. Once you forecast at more than one level, the levels have to agree, and the statistical machinery for that is a separate subject with its own literature.
The first job has a design rule that is easy to state and rarely applied: items that substitute for one another belong in the same node. Three colourways of the same shirt are individually erratic, because a customer who wanted navy will take black. The total across the three is smooth. Put them in one node and the aggregate forecast has real signal in it. Split them across nodes by colour, which some merchandising trees genuinely do, and you have taken a predictable total and turned it into three intermittent series that no model can help you with.
The same logic runs the other way. A node holding one item that sells forty thousand a week and eleven that sell twelve a week is dominated by the large item, and the aggregate forecast tells you nothing about the other eleven. Balance within a node matters more than tidiness of the tree.
Fixed levels break, attributes survive
A single tree forces one parent per item at every level. Most of the questions planning actually asks cut across it: every item sourced from the plant that just lost a line, every item with a six month shelf life, everything on promotion in week 12, everything whose primary component comes from one country. In a fixed hierarchy each of those is either a new level, a new tree, or a spreadsheet somebody maintains privately.
The alternative is to store the properties on the item and generate hierarchies as views over them. Sourcing location, pack format, shelf life, brand, lifecycle stage, seasonality profile, price band and substitution group are each a column. A grouping becomes a query, so a new one costs an afternoon rather than a migration, and two teams can hold different views of the same catalogue without arguing about which tree is correct.
The forecasting literature has a name for structures where items belong to several crossed groupings at once. Hyndman, Lee and Wang set out the computation for grouped time series in 2016, covering exactly the case where a catalogue can be cut by category and by region and by channel without those cuts nesting inside one another. Grouped structures are more work to reconcile than a clean tree and they describe most real businesses.
Two practical cautions. Attribute-driven grouping needs the attributes populated, and a column that is null on a third of the catalogue produces a view that is worse than the tree it replaced, because the missing items land in a residual bucket that nobody looks at. And attributes need controlled values. A shelf life column holding 180, 6M, six months and 0.5yr is not a column you can group on.
Ragged trees and the node called Other
Two structural defects turn up in almost every catalogue that has been running for more than a few years.
Unbalanced depth, where one branch runs three levels and another runs six because two businesses merged and neither tree was rebuilt. Reporting rollups still work. Anything that assumes a consistent number of levels, including most reconciliation code and most level-based policy assignment, quietly misbehaves.
Then the residual node. It is called Other, Miscellaneous, Unassigned or 999, and it is where items go when the person setting them up did not know the answer and the field was mandatory. It grows monotonically because nothing ever gets reviewed out of it. Measure its share of items and its share of value, and if the second number is material, that is the highest return classification work available to you, since those items are currently being forecast against an aggregate of unrelated products.
New items are the usual source. The record gets created under time pressure with the fields that block the purchase order filled in properly and the classification set to whatever passes validation.
Reclassification quietly restates history
Move an item to a new parent and you have changed every historical number that rolls up through that parent, unless the system stores membership with an effective date. Most do not. The item master carries the current parent, every report joins history to it, and last year suddenly reads differently from how it read last month.
Kimball and Ross laid out the general treatments in The Data Warehouse Toolkit in 2013. Overwrite the attribute and all history follows the new value. Keep a dated version of the record and history stays with the value in force at the time. Both are defensible and they answer different questions, which is why the useful step is to keep the dated version and then choose per use case.
Forecasting generally wants current membership applied across all history, because the model needs a consistent series and a category that changed definition halfway through is two series pretending to be one. Financial comparison generally wants as-was, because the prior year number should not move after it has been reported. Holding effective dated membership lets both be produced from one source. Holding only the current parent forces an argument every time the two teams meet.
The operational half of this is calendar discipline. Reclassification during an active planning cycle changes the numbers under a review that has already started. Batch the changes, run them at a known point between cycles, and publish what moved, so the person who spots a fourteen percent shift can check the change log before booking two hours of investigation.
Choosing the level you forecast at
Most of the argument about hierarchy design is really an argument about the level to forecast at, and that is answerable with a test rather than an opinion.
Set up the comparison honestly. Take the two candidate levels, item-location and whatever sits above it. Forecast at each with the same model family and the same history. Disaggregate the higher level result down to item-location using historical proportions. Score both at item-location, because that is where the order is placed, over a rolling origin so you get more than one window. Whichever wins, wins on that catalogue, and the answer will differ by segment.
The literature is old and the findings hold up. Gross and Sohl compared disaggregation methods for product line forecasting in the Journal of Forecasting in 1990 and found the choice of proportion matters as much as the choice of level. Fliedner's 2001 review in Industrial Management and Data Systems set out when aggregation helps, and the condition it turns on is correlation between the series being pooled. Where items in a node move together, aggregating buys you a stronger signal. Where they move independently, the aggregate is smoother by arithmetic and tells you nothing extra about any member.
A rough guide before you run the test. Items with weekly volumes in single figures, high intermittency, or heavy substitution within the node tend to forecast better from above. Items with strong individual seasonality, distinct promotional calendars or their own supply constraints tend to forecast better from below. A catalogue usually contains both, and a single global answer is the compromise, not the result.
The level you forecast at also does not have to be the level you plan at. The decision level is fixed by the replenishment decision, which is item-location almost everywhere. The forecast level is a modelling choice, and the disaggregation step is what connects them.
Where this stops
A hierarchy cannot repair a catalogue whose item codes are wrong underneath it. If one code covers two physically different products, or the same product carries three codes across regions, every level above inherits the confusion and grouping it more cleverly does not help.
Attribute-driven design has a real cost that gets understated. Somebody has to define the controlled values, populate them for existing items, and keep them populated at creation, and that is ongoing work with an owner and a queue rather than a project with an end date. A business that cannot sustain that is better off with a well-maintained single tree than with a rich attribute model that is forty percent empty.
There is also a ceiling on what the structure can do. When demand is driven by something outside the catalogue, a price change, a competitor going out of stock, weather, no arrangement of nodes recovers it. Hierarchy design makes the aggregation sensible and the policy assignment coherent. It does not add information that was never in the history.
Run one query this week: the share of items and the share of annual value sitting in the residual node at the level you forecast at, alongside the fill rate of the three attributes you would group by if you could.