In short: Hydrocarbon demand breaks several assumptions ordinary demand planning relies on, beginning with joint production, since the volume of one product leaving a refinery is not chosen independently of the others. Volume committed under a term contract is an obligation rather than demand, and modelling it as a market response mixes two different quantities in one series. A pipeline or gas network has no inventory buffer in the usual sense, so balancing happens through rate and price, and price is a driver and an outcome at the same time. Beyond a couple of years the dominant uncertainties are policy, technology adoption and macroeconomic structure rather than anything in the historical series, which makes long horizon work scenario construction rather than statistical forecasting.
Most demand planning literature assumes a discrete product. A case, a unit, a stock keeping number, sold to a customer who either buys it or does not.
Hydrocarbons break several of those assumptions at once, and the differences are structural rather than cosmetic. A refinery does not choose how much diesel to sell independently of how much gasoline it produces. A gas network cannot hold inventory in any conventional sense. Half the volume is committed under term contracts signed years ago, and the price is set against a published marker rather than by the seller.
Anyone importing consumer goods planning methods directly into this environment finds they do not fit. Four differences do most of the damage, and each has a reasonable treatment.
Joint production means demand is not separable
A barrel of crude yields a slate of products in proportions set by the crude's properties and the configuration of the plant. Push more of one and you get less of another, within limits set by conversion units.
This means the demand forecast for any single product is not an independent quantity to be optimised against. The useful object is a demand vector across the whole slate, and the planning question is which combination of crude selection and unit operation best matches the vector you expect.
The practical consequence for forecasting is that accuracy on individual products matters less than accuracy on the ratios between them. A forecast that is five percent high on every product is a scaling issue and easy to absorb. One that has the gasoline to distillate ratio wrong by five percent forces an operating change or a distress sale, and that costs considerably more.
Numbers make the point faster than the argument does. A plant runs 200 thousand barrels a day, and the plan expects a slate of 90 gasoline, 70 distillate, 40 heavy and other. Demand arrives at 84 gasoline and 76 distillate, with the heavy stream where it was forecast.
On the level metrics this looks like a decent month. Each of the two light streams moved by 6 thousand barrels a day, so the absolute percentage errors are just under 7 and just under 9, the total volume is exactly right, and a volume-weighted scorecard records 94 percent.
The ratio says something else. Planned gasoline to distillate was 90 over 70, or 1.286. Actual was 84 over 76, or 1.105. The plan is out by 14 percent on the one quantity the plant has to physically produce, and closing that gap means changing severity on the conversion units, buying distillate in, or placing 6 thousand barrels a day of gasoline somewhere it was not intended to go. Take the last of those at a three dollar discount to the marker and a month of it is 180 thousand barrels and 540 thousand dollars, given away by a forecast the scorecard called 94 percent.
Report ratio accuracy alongside level accuracy. It is the metric that maps to the decision.
Contracted volume is not demand
A large share of volume moves under term contracts with agreed quantities, tolerance bands and lifting schedules. The customer has committed to a range, and their behaviour inside that range is what varies.
Treating contracted volume as a forecasting problem is a category error. The contract tells you the quantity. What needs forecasting is nomination behaviour: where within the tolerance band the customer will actually lift, and when.
That is a much narrower and more tractable question. It has strong predictors, chiefly the customer's own downstream demand, their inventory position, the relationship between the contract price and the spot market, and seasonality in their operations. A customer whose contract price is above spot will lift at the bottom of the band. One whose contract is favourable will lift at the top. This is not subtle and it is frequently unmodelled, with the plan assuming mid-band nomination as a default.
The size of what that default hides is worth computing once. A term contract for one million barrels a month with a ten percent operational tolerance carries one million in the plan. A customer paying above the prompt market lifts at the floor and takes 900 thousand. One with a favourable formula takes 1.1 million. That is a 200 thousand barrel spread on a single contract, sitting inside a number the plan treats as agreed. Ten similar contracts put two million barrels a month of swing inside the contracted book, which is larger than any plausible improvement in the statistical forecast of the spot book, and it moves with something you can observe.
The diagnostic fits in a spreadsheet. For every contract and every month of history, express the nomination as a position within the band, running from minus one at the floor to plus one at the ceiling, then regress that position on the spread between the contract price formula and the prompt market, lagged by the nomination notice period. A significant slope with the sign you would expect means the mid-band default is a known error nobody has corrected. A flat relationship means the driver is somewhere else, usually in the customer's own operating calendar, and that is worth knowing too.
Separate the book into contracted and spot before forecasting anything. They are different processes and pooling them produces a series that describes neither.
There is no inventory buffer in the usual sense
Consumer goods planning leans on inventory to absorb forecast error. Get the forecast wrong, hold more stock, accept the carrying cost.
Gas systems cannot do this. Line pack provides hours of flexibility, not weeks. Storage where it exists is expensive, limited and often committed. For a pipeline network, supply and demand have to balance more or less continuously, which means forecast error converts into an operational action rather than into an inventory movement.
Liquid products have more storage flexibility and it is still constrained by tank capacity, product segregation requirements and the fact that a tank full of one grade cannot hold another.
The planning consequence is that the value of forecast accuracy is much higher at short horizons than the equivalent consumer goods case, because there is no buffer to absorb the error. It also means a probabilistic forecast is more useful than a point forecast, since the operational question is usually about the tail rather than the expectation: what is the plausible high demand case that the system has to be able to serve.
The machinery for that is settled outside this sector. Hong and Fan's tutorial review of probabilistic load forecasting, published in the International Journal of Forecasting in 2016, sets out the quantile and density approaches along with the scoring rules used to compare them, and the electricity work generalises to any system where the operational question sits in the upper tail. What changes for hydrocarbons is which quantile you care about and how far ahead it has to be committed.
The failure mode worth watching for is a plan built at monthly granularity for a system that balances daily. A month that nets to zero can contain a week of oversupply followed by a week where the network cannot serve nominations, and the monthly plan shows neither. The symptom is a planning cycle reporting on target while operations spends that same month issuing curtailment notices and buying prompt cargoes. Where both of those are true, the granularity of the plan is wrong before anything about its accuracy comes into it.
Price is a driver and an outcome at once
For most planning contexts, price is set by the business and demand responds to it. In commodity energy, the price is set by a market and both your volume and your realised value depend on it.
That makes the forecast jointly determined with the price path in a way that a single-equation demand model handles badly. Two practical treatments.
Model the volume against forward curves rather than against a price forecast of your own. The forward market has already aggregated a large amount of information and using it as an input is more defensible than substituting your own view, unless you have a specific reason to differ and are willing to state it.
Run the demand plan across a small number of price scenarios rather than a single path, and look at which decisions change across them. Many operating decisions turn out to hold across a wide price range, and identifying those lets you commit to them early while leaving the price-sensitive decisions open. That separation is more valuable than a more accurate central case.
Testing for that invariance is quicker than it sounds. Solve the plan under a low, a central and a high price path, then compare the decision variables rather than the objective values. What usually falls out is that the near-month lifting schedule, the crude nominations already inside their notice periods and the maintenance windows are identical under all three, and the differences concentrate in a handful of discretionary cargoes plus how much term volume to commit at the far end of the horizon. The meeting is then about the short list that actually moved.
What transfers from ordinary demand planning
Several things do transfer, and they are worth stating because the differences above can make the domain sound entirely separate.
Hierarchical coherence still applies. Product totals have to reconcile with regional totals and with the system total, and the reconciliation machinery works identically.
Bias tracking still matters and is frequently neglected. A demand plan that is persistently two percent low on a distillate stream produces a chronic operating tension that gets attributed to execution rather than to the plan.
Value add scoring still applies to the human touches. Term marketing teams adjust nomination forecasts, and whether those adjustments help is measurable in exactly the standard way.
So does evaluation methodology. Rolling-origin cross validation, set out in Hyndman and Athanasopoulos, Forecasting: Principles and Practice, is how you establish whether a change to the model is an improvement rather than a lucky window, and it is the step most often skipped in favour of a single holdout period that happened to be quiet.
The weather sensitivity is stronger than in most sectors and is handled with the same techniques: degree day variables, non-linear response above and below comfort thresholds, and a distinction between the weather effect on volume and the weather effect on timing.
The limit
The demand-side treatment here stops at the plant gate. What a refinery should actually run, given a demand vector and a set of available crudes, is an optimisation problem with a completely different structure, and it is not solved by better demand forecasting.
There is also a horizon issue worth being clear about. Beyond a couple of years, energy demand forecasting becomes scenario work rather than statistical forecasting, because the dominant uncertainties are policy, technology adoption and macroeconomic structure rather than anything present in the historical series. A statistical model extrapolating twenty years of consumption is answering a question it cannot answer, and dressing that up with confidence intervals makes it worse rather than better. For those horizons, a small number of internally consistent scenarios with explicit assumptions is the honest instrument, and the value is in the assumptions rather than in the numbers attached to them.