In short: Power demand forecasting covers four problems under one name, and the horizon that matters is the one at which your position commits. Temperature explains more of the variance than everything else combined in any system with meaningful electric heating or cooling, and the relationship is non-linear. Scoring a model with the actual weather already known flatters it, because the operational forecast has to carry weather forecast error as well as load model error. Most operational questions about load concern the upper tail, such as a one-in-twenty winter peak or how much reserve to hold, and a point forecast answers none of them.
The monthly forecasting pack says the day-ahead model ran at 2.1 percent mean absolute percentage error, down from 2.4. The imbalance line in the same month's management accounts is up by a fifth, and nobody in either meeting has the other meeting's number.
Both are describing the same forecast. The gap between them is where most of the useful work sits, because a load forecast is bought for the decisions it commits, and the decisions commit at particular horizons, under particular price conditions, with an error that costs different amounts depending on its sign.
Four horizons wearing the same name
Load forecasting is one phrase covering four problems with different drivers, different data and different evaluation.
Very short term, minutes to a few hours, is dominated by persistence and by the autocorrelation in the recent series. Weather barely moves in that window, and the models that win are the ones that track the current level accurately.
Day ahead, meaning the horizon at which most wholesale purchasing commits, is dominated by the weather forecast and the calendar. This is where the money is for a supplier or a trading desk.
Weeks to months out is a mix. Calendar structure and the seasonal temperature climatology carry it, and the useful output is a distribution rather than a path, because no weather forecast at that range is informative about a specific day.
Years out is a different discipline again, driven by economic activity, electrification of heat and transport, and connected new load. Extrapolating a statistical model over that horizon produces a number with a confidence interval that means nothing, and the honest instrument is a small set of internally consistent scenarios.
The common failure is a single model, tuned and scored at one horizon, used to answer all four questions. The tell is a forecast accuracy report with one number on it.
The temperature response is most of the model
For any system with meaningful electric heating or cooling, temperature explains more of the variance than everything else combined, and it does so non-linearly.
The response has a comfort band where load is insensitive, and slopes on either side of it. The two slopes are usually different, the cooling slope is often steeper, and both interact with the hour of the day, because a hot afternoon and a hot midnight do very different things to load. A model that enters temperature linearly, or that uses a single degree-day variable without an hour interaction, throws away most of the available signal.
Put a number on the sensitivity for your own system, because it converts a meteorological error into a megawatt error directly. Take a distribution area peaking at 3,000 MW with a heating slope of 60 MW per degree Celsius below the balance point. A day-ahead temperature forecast that misses by 1.5 degrees, which is an ordinary miss rather than a bad one, is 90 MW of load error before the load model has made any mistake of its own. That is three percent of peak, arriving from outside the modelling team entirely.
Two things follow. The weather input deserves the same attention you give the load model, including which station or gridded product you use and how you weight several of them; Hong, Wang and White treated station selection directly in the International Journal of Forecasting in 2015 and found it moves accuracy materially. And the numerical weather prediction vendor's own error statistics belong in your error budget, because they set the floor.
Scoring with the weather already known flatters everything
Here is the version of the accuracy number that gets reported, and the version that describes what you will experience.
Backtests are commonly run by feeding the model actual observed temperature for the target period and comparing its output against actual load. That measures the load model in isolation, which is a reasonable thing to want when you are choosing between two model specifications.
The accuracy available to you on the day is lower, because on the day you have a temperature forecast rather than a temperature. The difference between those two figures is often larger than the difference between the model you have and the model you are considering buying.
Measure both, and report them separately. Run the backtest twice, once with observed weather and once with the archived forecast weather that was actually available at the decision time, and the gap between the two numbers is your weather error contribution. If that gap is a point of MAPE and the model comparison you are agonising over is worth a tenth of a point, you know where the next month of effort should go.
Archived forecast weather is the awkward input here, because most organisations keep the observations and discard the forecasts. Start retaining them now; a year from now the analysis becomes possible and today it is not.
What a point of accuracy is worth at settlement
The reason to care is settlement, and the arithmetic is worth doing rather than asserting.
Take a supply book averaging 500 MW. A mean absolute percentage error of 2.0 percent on the day-ahead position is 10 MW of average absolute imbalance. Across 8,760 hours that is 87,600 MWh of imbalance volume. At an average absolute spread of 40 currency units per MWh between the imbalance price and the day-ahead price you cleared at, the annual cost is about 3.5 million. Improving to 1.6 percent MAPE removes a fifth of that, or roughly 700 thousand a year.
Two corrections make the honest number worse than that.
Imbalance pricing is frequently asymmetric, so being short when the system is short costs considerably more than being long when the system is long earns. Multiplying a symmetric error by an average spread understates the cost whenever your error is skewed.
And your error correlates with everyone else's. When an unexpected cold snap arrives, every supplier in the market is short at the same time, the system is short, and the imbalance price is at its worst exactly when your volume exposure is at its largest. The cost is a product of two correlated quantities, so its expectation exceeds the product of the expectations. This is why a book with a well-behaved error distribution can still produce a fat imbalance line in one month of the year.
The published work supports the direction if not the specific figure. Hobbs and colleagues analysed the value of load forecast improvement for unit commitment in IEEE Transactions on Power Systems in 1999 and found the savings from a one point accuracy gain to be substantial relative to the cost of achieving it, which remains the shape of the answer three decades later even though the market structures have changed.
The version of this calculation that will convince your finance function uses your own settlement data. Take last year's half-hourly imbalance volumes and imbalance prices, recompute what the cost would have been under a forecast error scaled down by twenty percent, and you have the value of the improvement in the currency the business reports in.
The structure that usually gets left out
Three pieces of structure are cheap to add and commonly absent.
Recency in the temperature response. Load today depends on the temperature over the preceding hours and days, because buildings have thermal mass and behaviour adapts with a lag. Wang, Liu and Hong showed in the International Journal of Forecasting in 2016 that including lagged temperatures and moving averages of temperature improves accuracy consistently across utilities. The implementation is a handful of extra regressors.
Holiday and bridging day handling. A public holiday has a load shape closer to a Sunday than to the weekday it falls on, and the days around it have their own behaviour. Treating holidays as a single dummy variable pools genuinely different effects. Where holidays move against the solar calendar, the treatment needed is the one covered in D6.
Embedded generation behind the meter. Metered load in a distribution area is gross demand minus whatever rooftop solar produced, and if that fleet is growing, your load series has a non-stationary generation component inside it. Modelling the net series directly means the temperature and solar irradiance effects are entangled, and the coefficients drift as installations accumulate. Where the data exists, estimate gross load and the behind-the-meter contribution separately and recombine them, using installed capacity registers to scale the generation component rather than letting a fitted coefficient absorb it. The generation side of that estimate is a different forecasting problem with its own methods, covered in CC3.
Forecasting the tail rather than the middle
Most operational questions about load are about the upper tail. Whether the system can serve a one-in-twenty winter peak, how much reserve to hold, whether a constrained feeder needs reinforcement this year or in three years. A point forecast answers none of them.
Probabilistic load forecasting has a settled methodology, developed substantially through the Global Energy Forecasting Competitions; the 2014 edition and the methods that came out of it are documented by Hong, Pinson, Fan, Zareipour, Troccoli and Hyndman in the International Journal of Forecasting in 2016. The output is a set of quantiles or a full predictive density, scored with the pinball loss or with the continuous ranked probability score, both of which are proper in the sense set out by Gneiting and Raftery in the Journal of the American Statistical Association in 2007.
The practical distinction that matters for peak work is between the uncertainty in the load model and the uncertainty in the weather. For a one-in-twenty peak you want the distribution of load under the distribution of weather, which means running your load model across many historical or simulated weather years rather than widening an interval around a central case. Those two procedures give different answers and the second one is the right one.
The limit
Everything above assumes the relationship between weather and load is stable enough to estimate from history. Heat pump and electric vehicle adoption are changing that relationship while you are estimating it, and the change is not uniform across the day. A heating slope estimated on five years of data describes a building stock that no longer exists by the end of the estimation window, and the direction of the bias is that you will under-forecast winter evening peaks in an electrifying system.
There is no clean statistical fix, because the data on the new load simply has not accumulated. What helps is holding the new load out as an explicit additive component with its own assumptions, sized from connection records and vehicle registrations, so that the assumption is visible and arguable rather than buried inside a fitted coefficient. When the assumption turns out wrong you can correct it in one place.
The other limit is a data one. Settlement data, metering data and the forecasts themselves usually sit in three systems with three different time conventions, and the daylight saving transitions are wrong in at least one of them. Every analysis in this piece dies on that, quietly, producing plausible numbers that are an hour out for half the year.
Take one week in October and one in March, and reconcile load, price and forecast timestamps across all three systems by hand.