In short: A long requirements matrix does not discriminate in demand planning software selection, because every established vendor answers yes to more than ninety percent of it. The lines that would separate the products describe behaviour under conditions the evaluation team has not encountered yet, which is why nobody thought to ask about them. Insist on a demonstration that runs on your data with tasks you specify, and put that in the invitation rather than raising it at the meeting. Reference calls are worth having when the questions are framed so that an unhappy answer is socially acceptable.
The requirements matrix has 412 lines. Five vendors have responded. Four of them score above 94 percent and the fifth is at 91 because they answered honestly on two lines the others fudged. The evaluation team is now trying to choose between four systems that are, according to the document they spent three months building, effectively identical.
That outcome is close to guaranteed by the method. A requirements list assembled by asking stakeholders what they need produces a list of things planning systems do, and every established product does them. The lines that would separate the products are the ones nobody thought to ask, because they describe behaviour under conditions you have not encountered yet.
I sell a planning platform, which means you should read this with the appropriate suspicion. I have tried to write it so that it would be equally useful to somebody who ends up buying from a competitor, and the test of that is whether any requirement below is written in a way that only one product could satisfy. I do not think any of them is.
Why the long list does not discriminate
Run the arithmetic on your own matrix before the responses come back. Of 412 lines, how many would you expect a serious established vendor to answer yes to? Usually somewhere above 90 percent. If 380 lines get a yes from everyone, the decision rests on the remaining 32, and those 32 were never weighted, because the weighting was applied across all 412 lines equally or by a stakeholder's opinion of importance rather than by discriminating power.
There is a second, subtler problem. A requirements list written by a team that has used one system for a decade will use that system's vocabulary, and the vendor whose product shares that vocabulary will score better without being better. Watch for terms that are product names rather than concepts.
The correction is straightforward. Before issuing anything, mark each line with your honest expectation of how many vendors will satisfy it. Delete or downgrade anything you expect all of them to meet. What remains is your actual decision, and it is usually short enough to fit on two pages.
The requirements that separate
These are the areas where products genuinely differ, in my experience of both sides of the table.
Censored history handling. What the system does with a period where you could not supply. Whether it can distinguish an out-of-stock week from a genuinely low-demand week, and whether the unconstraining method is stated. D1 covers why this matters.
Intermittent and lumpy series. Which methods are available for items that sell in scattered periods, and how the system decides which items get them. Ask what proportion of your catalogue their segmentation would route to intermittent methods.
The level at which models are fitted. Whether models are fitted at the level you plan at, at a higher level and disaggregated, or at several levels and reconciled. Ask how reconciliation is done and what happens to the reconciled numbers when a hierarchy changes.
Promotional representation. Whether a promotion is a flag, a set of features, or an explicit uplift model with baseline separation. Ask what happens to the baseline in a promoted week.
New product method. What the system does for an item with no history, and how an analogue is chosen and weighted.
Forecast versioning and audit. Whether you can retrieve the forecast exactly as published on a past date, with the inputs and the overrides that produced it. This one is frequently answered yes and delivered as a snapshot table that loses the reasoning.
What an override does next cycle. Whether a planner adjustment is a one-off, a persistent bias correction, or an input the model learns from. All three are defensible; the answer tells you a great deal about the product's design philosophy.
Behaviour on a mid-year hierarchy change. What happens to history, models and reconciliation when a category is reorganised in July.
Extensibility. Whether you can add a method the vendor does not ship, in what language, with what support boundary, and whether it survives an upgrade.
Run time at your volume. Not their benchmark. Yours, at your item-location count and history depth, measured during evaluation.
Concurrency. How many planners can work simultaneously and what happens to the others when one runs a large recalculation.
Data quality and exception surfacing. What the system does when an input looks wrong, as distinct from what it does when an input is missing.
That is twelve. Add the handful specific to your business, which usually includes something about a channel, a calendar or a regulatory obligation, and you arrive at something near twenty. Twenty lines that discriminate beat four hundred that do not.
Weight them by exposure rather than by enthusiasm. The way to do that is to attach a volume to each requirement. If intermittent items are 8 percent of your catalogue and 1 percent of your volume, that requirement is worth less than it feels like in a workshop where the planner responsible for them is present. If censored history affects the 300 item-locations that carry 40 percent of your revenue, it outranks nine of the others on its own. Working out those shares takes a morning against your own sales data and it changes the weighting more than any amount of discussion does.
The demonstration script to insist on
A vendor-controlled demonstration shows you a rehearsed path through prepared data. The only version worth watching runs on your data with tasks you specify, and you should say so in the invitation rather than at the meeting.
The script that produces the most information is a set of tasks with a stopwatch running.
Load a supplied extract of our history and produce a statistical forecast for these 5,000 item-locations. Show us how long each step took.
Here are twelve item-locations. For each, tell us which model was chosen and why, in a form a planner could explain to a sales director.
Here is an item that was out of stock for six weeks last year. Show us what the system does with that history and what setting controls it.
Change the product hierarchy by moving these four items to a different category, and show us what happens to the forecasts and to last year's comparison.
A planner disagrees with the forecast for this item and enters 4,000. Show us where that override is recorded, what it does to the parent level, and what happens to it next cycle.
Show us the forecast this system published for this item eight weeks ago, and the inputs that produced it.
Two planners edit the same product family at the same time. Show us what each of them sees.
Seven tasks, roughly two hours, and the information density is higher than a full day of slides. Whoever cannot do a task live should be asked to say so rather than to explain the roadmap, and a clear answer that a capability is coming is more useful than a demonstration of something adjacent.
Scoring accuracy against your own history is a separate exercise with its own design, and E4 covers how to run it so that the result means something.
Reference calls, and the questions that work
Vendors provide references who are happy. That is fine, and you can still get a great deal out of the call by asking questions where an unhappy answer is socially acceptable.
Ask what they would do differently if they were starting the implementation again. Everybody has an answer and it is usually specific.
Ask how long it took from contract signature to the first forecast a planner actually used. The gap between that and what you were told in the sales process is the most useful single number you will get.
Ask who from the original project team is still involved, and what happened to the others. Turnover in the customer's own team is a strong signal about how the project went.
Ask what broke at the last upgrade and how long it took to resolve.
Ask which of the capabilities they bought are still unused, and why. There is always a list.
Ask for a reference at your data volume and in your industry, and treat a refusal as information. A vendor with no comparable customer may still be the right choice, and you should know that going in rather than discovering it during implementation.
Score it so the answer survives scrutiny
Three practices make the decision defensible six months later when somebody asks why.
Weight the discriminating requirements explicitly and agree the weights before seeing any responses. Weights set afterwards get adjusted, usually unconsciously, towards the preferred answer.
Score independently and then discuss. A group scoring together converges on whoever spoke first.
Record the evidence for each score, in a sentence, naming what you saw. A score of four with no evidence is an opinion with a number attached.
Keep a written record of what each vendor could not do. That list is what you will need when the project hits the thing they could not do, and it converts a surprise into a known limitation you accepted deliberately.
One arithmetic check on the scoring itself is worth doing before you present it. Take the winning score and the runner-up, and work out how many of the twenty weighted lines would have to flip to reverse the result. If the answer is one or two, the process has not separated them and you should say so rather than dressing a coin toss as an analysis. In that situation the honest tie-breakers are commercial terms, implementation team quality and the reference calls, and choosing on those openly is more defensible than adjusting a weight until the spreadsheet agrees with you.
What I would not oversell
A selection process cannot tell you whether the product will work in your organisation, because most implementation failures are about data, process and adoption rather than about product capability. E5 covers what that failure looks like. The strongest selection process in the world will not save a deployment where nobody owns the master data.
The demonstration script also has a limit. Two hours on your data with seven tasks tells you a great deal about the product and very little about the team who will implement it, and the implementation team is at least as large a factor in the outcome. Ask who specifically will be assigned, and try to meet them.
And I should be honest about the position I am writing from. Every vendor prefers a selection process shaped around what they do well, and this list is not exempt from that. Test it by taking these requirements to two vendors, mine and one other, and seeing whether both can engage with them properly. If a requirement only one vendor can discuss is on your list, that requirement was written by that vendor.
Start by going through your existing requirements document and marking each line with how many of your shortlisted vendors you expect to satisfy it. The lines where you expect all of them to say yes can come out of the scoring entirely, and what remains is the decision you are actually making.