In short: Renewable generation forecasting is scored on average error and consumed by markets that price the extremes, which is how a good day by the monthly report can carry the month's largest imbalance charge. Below rated wind speed turbine power rises roughly with the cube of wind speed, so a small error in wind speed becomes a large error in power. Ramps, meaning large output changes over a short window, drive balancing cost and reserve requirement, and average error metrics score them badly. Aggregating sites before forecasting is the cheapest accuracy improvement available, because the variance arithmetic costs nothing.
The day-ahead forecast for the portfolio said 240 MW for the evening peak. Output held at 250 until quarter past six, then fell to 60 inside forty minutes as a front came through earlier than modelled. The day's mean absolute error came out at 6 percent of capacity, which is a good day by the monthly report's standards. The imbalance charge for those two settlement periods was larger than the whole of the previous week.
That combination, a decent average error and a painful outcome, is the normal shape of the problem. Renewable output forecasting is scored on averages and consumed by markets that price the extremes, and closing that gap is mostly about which quantities you forecast and how you evaluate them.
Two different physical problems under one heading
Wind and solar get discussed together because they share a market position. They are different forecasting problems.
Solar output is a deterministic astronomical quantity multiplied by a stochastic atmospheric one. The geometry, meaning sun position, panel orientation and theoretical clear-sky irradiance, is known exactly for every minute of the year. What has to be forecast is the clear-sky index, the ratio of actual to clear-sky irradiance, and that is a cloud problem. Forecasting the index rather than the power removes the entire diurnal and seasonal cycle from the target series before you start, which is the single most useful thing you can do to a solar model. For horizons out to a few hours, satellite cloud motion beats numerical weather prediction; beyond that the numerical model takes over.
Two solar-specific corrections belong on top of the index. Inverter clipping caps output when the panel array is oversized relative to the inverter, which truncates the upper tail and makes clear-sky summer days easier to forecast than the irradiance alone suggests. And panel temperature reduces efficiency as ambient temperature rises, so the hottest hour of a hot day produces less than the irradiance implies. Both are deterministic given the weather and both are commonly left to a fitted model to discover.
Wind output is a wind speed forecast pushed through a power curve, and the power curve is what makes it hard.
The power curve turns a small wind error into a large power error
Below rated wind speed, turbine power rises roughly with the cube of wind speed. Above rated it is flat. Below cut-in and above cut-out it is zero.
Work the cube. At a forecast of 8 metres per second, an error of 1 metre per second gives a power ratio of nine cubed over eight cubed, which is 1.42. A wind speed forecast that is 12.5 percent high produces a power forecast 42 percent high. The same 1 metre per second error at a forecast of 13 metres per second, above rated, produces almost no power error at all, because the curve is flat there. And an error either side of cut-out is the difference between full output and nothing.
Three consequences follow directly and all of them are structural.
The error distribution is heteroskedastic. Uncertainty in power depends on where on the curve you are, so a constant prediction interval is wrong everywhere. Pinson set this out clearly in Statistical Science in 2013, along with the reason the standard Gaussian machinery misbehaves here.
The distribution is bounded and skewed. Output cannot go below zero or above capacity, and near either bound the predictive density piles up against it. A symmetric interval centred on a point forecast will place probability mass outside the physically possible range, which is a visible defect in any calibration plot.
And the mapping is not linear, so the expected power is not the power of the expected wind speed. Running a single deterministic weather forecast through a power curve gives a biased answer whenever the curve is curved, which is most of the operating range. The fix is to push an ensemble or a distribution of wind speeds through the curve and take the mean of the outputs.
Ramps, and why they are scored badly
A ramp is a large change in output over a short window, and it is the event that drives balancing cost, reserve requirement and, in a constrained system, the risk of an intervention.
Standard error metrics handle ramps poorly because of the double penalty. Forecast a 160 MW drop at 18:00 when it actually happens at 18:40, and the metric charges you twice: once at 18:00 for predicting a fall that had not happened, and again at 18:40 for not predicting the fall that did. A forecast that missed the ramp entirely and held a flat mid-level line all evening can score better on mean absolute error than one that got the magnitude exactly right and the timing forty minutes wrong, even though the second forecast is far more useful operationally.
So ramps need their own evaluation alongside the level metrics. Define a ramp event for your assets, something like a change exceeding a threshold percentage of capacity within a defined window, then score detection and timing separately: how many actual ramp events were forecast at all, what the distribution of timing error looks like, and what the magnitude error was conditional on detection. Those three numbers tell an operations team something that a single MAE figure never will.
Put the money on it. Take a 200 MW farm ramping from 180 MW to 20 MW. If the timing is thirty minutes early in the forecast, you are long by roughly 160 MW for one half-hour period and short by a comparable amount when the ramp is finally delivered. That is about 80 MWh of imbalance volume in each direction, and ramps of this kind tend to happen when the whole region's wind fleet is doing the same thing, so the imbalance price is at its least forgiving in exactly that period.
What the balancing exposure actually is
The arithmetic that connects a forecast to a profit and loss line is short enough to do in a meeting.
Take a 300 MW portfolio at a 35 percent capacity factor, so about 920 GWh a year. Assume a day-ahead normalised mean absolute error of 12 percent of installed capacity, which is a reasonable single-region starting point, and substitute your own measured figure. That is 36 MW of average absolute error, or roughly 315 GWh of imbalance volume across the year.
At an average absolute spread of 30 currency units per MWh between the imbalance price and the day-ahead price, that is around 9.5 million a year, against a gross revenue on 920 GWh that might be 55 million at 60 per MWh. The balancing exposure is a sixth of the revenue line, which is why the forecast is a commercial asset rather than an engineering nicety.
Two corrections, both in the same direction. Imbalance prices are usually asymmetric, so the cost of being short exceeds the benefit of being long. And your error is correlated with every other wind operator in the region, so the periods when you are most exposed are the periods when the price is worst. The realised cost exceeds the product of average error and average spread for the same reason described in CC2 for load.
Aggregation is the cheapest accuracy improvement available
Before spending anything on models, check whether you are forecasting at the right level of aggregation, because the variance arithmetic is generous and free.
For a portfolio of n sites each with error standard deviation sigma and average pairwise correlation rho, the standard deviation of the portfolio error as a fraction of the single-site figure is the square root of (n plus n times (n minus 1) times rho) divided by n. Ten sites at a pairwise correlation of 0.3 gives the square root of 37 over 100, which is 0.61. The portfolio error is 39 percent smaller than a single site's, purely from geographic spread.
That result cuts both ways operationally. It means a portfolio forecast produced by summing individually optimised site forecasts is usually worse than one modelled at portfolio level, because the site models are tuned to minimise their own errors rather than the error of the sum. It also means the correlation structure is worth measuring rather than assuming; sites in the same wind regime behave close to a single large site and buy you almost nothing.
Where the same physical quantity has to be reported coherently at site, portfolio and market level, the reconciliation machinery covered in D8 applies unchanged.
Making the probabilistic version usable
Point forecasts are the wrong output for most of the decisions this feeds. Reserve procurement, storage scheduling, the decision to buy back a position intraday and the assessment of curtailment risk are all questions about a distribution.
The methods are settled. Quantile regression estimated directly on the predictors, quantile forests, and ensemble weather forecasts propagated through the power curve all produce calibrated densities, and the wind track of the 2014 Global Energy Forecasting Competition established a reasonable public benchmark for how well they do. Score with the pinball loss across the quantile grid, and check calibration and sharpness separately as Gneiting, Balabdaoui and Raftery set out in the Journal of the Royal Statistical Society in 2007. A forecast can be perfectly calibrated and useless if it is wide, and sharp and dangerous if it is not calibrated.
The operational point that gets missed is that trajectories matter as well as marginals. A set of hourly quantiles tells you the distribution of output in each hour separately, and says nothing about whether a low hour is followed by another low hour. Any decision spanning several periods, notably storage scheduling and reserve holding, depends on the joint behaviour across time, which means you need scenario paths rather than a fan chart.
The limit
The forecast is bounded by the weather model underneath it, and for the horizons where balancing cost is largest, one to twelve hours ahead, the numerical weather prediction is often the binding constraint. Statistical post-processing on your own site data recovers a useful amount, particularly by correcting systematic terrain and wake effects that the meteorological model resolves poorly. It cannot manufacture information about a front whose timing the atmospheric model has wrong.
Availability is a second limit and it is frequently the larger error. A forecast of what the wind will do is a forecast of what a fully available fleet would produce. Turbines are out for maintenance, curtailed by a network operator, derated for icing or noise, or offline for reasons the forecasting system knows nothing about. If your outage information reaches the forecasting pipeline late or not at all, a large share of the error attributed to meteorology is really a data plumbing problem. That one is cheap to fix and it is usually fixed last.
And a caution about backtests. A model tuned and evaluated on the same two years of weather has been fitted to those years' particular sequence of storms and blocking patterns. Rolling-origin evaluation across several years, in the form set out in Hyndman and Athanasopoulos, is the minimum defensible standard before believing that a change is an improvement.
Take last year's largest twenty ramp events at your sites, look up what the day-ahead forecast said for each, and count how many were present at all.