In short: Crop yield forecasting is a sequence of different problems wearing one name, because the information available changes shape three or four times between planting and harvest. Tonnage is area multiplied by yield, and area receives roughly a tenth of the forecasting effort while supplying half the answer. Regional aggregation cancels far less error than the arithmetic suggests, because weather is the dominant driver and it makes field level errors strongly correlated. A point estimate fails here because the cost of being short differs from the cost of being long, so the width of the distribution has to survive into the decision.
The commercial director wants a tonnage number in April for a crop that will be cut in September. The agronomy team gives him 8.4 tonnes a hectare. He multiplies by contracted area, gets a figure, and puts it in the plan.
Nobody involved thinks 8.4 is going to be the answer. The agronomists know the distribution around it runs from about 6 to about 10, with a longer tail downward than upward. The commercial director knows it too, in the sense that he has seen five years of the same number being wrong. What neither of them has is a way to carry that width into the decisions being made off the back of it, so the width gets discarded at the first arithmetic operation and the plan proceeds as though 8.4 were a measurement.
The forecast improves in steps, and the steps have dates
Yield forecasting is a sequence of different problems wearing one name, because the information available changes shape three or four times through a season.
At planting you have area, variety mix, soil, the previous crop, and the long-run trend. Nothing about this season's weather has happened yet, so the forecast is essentially the climatological distribution conditioned on what you planted. The width at this point is irreducible with any method.
Through vegetative growth you accumulate weather that has occurred: rainfall, thermal time, water balance. The forecast narrows as the season converts from unknown to known.
Around flowering and grain fill the critical window passes, and for most field crops this is where the largest single reduction in uncertainty happens, because yield formation in that window dominates the outcome.
In the last few weeks before harvest you are close to measuring rather than forecasting, and the residual uncertainty is mostly about harvest losses and the weather during the cut itself.
The useful thing to build is a skill curve for your own crop and region: forecast error by week of season, measured over as many past years as you have. A curve that shows a root mean square error of 1.35 tonnes a hectare at planting, 0.85 at flowering and 0.35 three weeks out tells every decision owner in the business when their decision can wait and when it cannot. Most businesses have never plotted it, and consequently argue about the forecast rather than about the calendar.
Half the tonnage answer is area, and it gets almost no attention
Tonnage is area multiplied by yield, and forecasting effort divides between them at roughly ninety to ten. That split made sense when area was declared, stable and knowable. It is less defensible now.
Contracted area and planted area diverge. Growers switch between crops late when relative prices move, they leave ground unsown after a wet autumn, and some of what was drilled gets written off in spring and redrilled with something else. In a bad establishment year the gap between the area you contracted and the area standing in June can run to a tenth of the book.
Two sources close most of it. Subsidy and land parcel declarations give a legally-filed statement of crop by parcel in many jurisdictions, with a known filing date. Satellite crop type classification gives an independent read that improves through the season as the canopies of different crops separate. Neither is perfect and they disagree usefully, because the disagreements point at the parcels worth a phone call.
An error of five percent in area is exactly as damaging as an error of five percent in yield, and it is considerably cheaper to remove.
The three method families, and why they get combined
Statistical models on weather and trend. Regress historical yield on a technology trend plus weather aggregates for defined windows. Cheap, transparent, and surprisingly hard to beat in a region with a long yield record. The weakness is extrapolation: a model fitted on observed conditions has no basis for a year outside them, and those are the years the business most needs a number for.
Process-based crop simulation. Models that step through the season daily, tracking phenology, water and nitrogen balance and biomass accumulation. The DSSAT cropping system model, described by Jones and colleagues in the European Journal of Agronomy in 2003, and APSIM, described by Keating and colleagues in the same journal and year, are the two long-established families. They carry causal structure, so they behave sensibly outside the historical range, and they need soil, management and cultivar parameters that are often unavailable at the resolution you want.
Remote sensing. Canopy indices from satellite imagery track how the crop is actually developing, at field resolution, without anybody visiting the field. The European Commission's Joint Research Centre has run the MARS crop monitoring programme on this basis since the late 1980s, and the GEOGLAM Crop Monitor has provided a coordinated global assessment since 2013. The limitation is that an index measures canopy rather than yield, and the relationship between them varies by variety, by season and by what happens after the canopy peaks.
The combination that works better than any of the three is assimilation: run the process model, and use remote sensing observations through the season to correct its state. Lobell and Asseng compared process-based and statistical approaches in Environmental Research Letters in 2017 and found the choice matters most exactly where the answers diverge, which is at the extremes. Basso and Liu's review in Advances in Agronomy in 2019 is the practical survey of what has been achieved operationally.
Aggregation helps far less than the arithmetic suggests
The intuition is that field-level forecasts are noisy and regional ones are fine, because errors cancel. They cancel only to the extent the errors are independent, and weather is the dominant driver, which makes them strongly correlated.
Work it. Field-level error of 1.3 tonnes a hectare, 200 fields in a catchment, correlation between field errors of 0.4. Independent errors would give a regional error per hectare of 1.3 divided by the square root of 200, which is 0.09. With correlation, the aggregate standard deviation scales by the square root of one plus 199 times 0.4, over the square root of 200, giving 0.63, so the regional error is 0.83 tonnes a hectare rather than 0.09.
Aggregating 200 fields removed about a third of the error and not nine tenths of it. That single number explains a great deal about this industry: regional yield risk does not diversify, which is why a bad year is bad for everyone at once, why forward positions and physical positions move together, and why a processor's supply risk and its price risk are the same risk.
The correlation is measurable from your own intake records. Take field or farm level yields over the years you have, compute the pairwise correlation, and use it before you promise anyone that aggregation will solve their problem.
Point estimates fail because the decisions are asymmetric
If the cost of being wrong were symmetric, the mean would be the right number and the width would be an interesting footnote. Almost nothing in this business is symmetric.
Intake and drying capacity sized to the mean is short in every above-average year, and the cost of being short is queues, grower dissatisfaction and a crop that arrives wet with nowhere to go. Contracted sales volume set at the mean is unfulfillable half the time, and the cost of buying in to cover a shortfall lands in a year when the market is already high, because your shortfall and the market's shortfall have the same cause.
The treatment is to carry the distribution to the decision rather than the point estimate, and to solve each decision at its own quantile. Capacity that is cheap to add and expensive to lack belongs somewhere in the upper part of the distribution. A sales commitment with a penalty clause belongs low. The general machinery for producing and using distributional forecasts is covered elsewhere (D9); what is specific here is that the distribution is wide, skewed, and available essentially free, because a process model or a weather-conditioned statistical model can be run across many historical weather years to generate one.
That last point is worth being concrete about. Take this season's planting, soil and management, run the crop model against each of the last thirty years' weather from today's date forward, and you have thirty yield outcomes conditioned on what has already happened this season. That ensemble is the distribution, it updates every week as more of the season becomes observed, and it requires no new data collection.
Trend is a modelling decision that quietly sets your answer
Yields rise over time from genetics, agronomy and equipment, at a rate that differs by crop and region. Any model fitted across twenty years has to decide what to do about that.
Treating trend as linear and extrapolating is the standard approach and it works until the trend flattens, which it has in several crops and regions. Fitting trend and weather simultaneously risks the trend absorbing a run of favourable years and overstating the underlying gain. Detrending first and modelling deviations is cleaner and makes the trend assumption explicit rather than buried.
Whichever you choose, the validation has to respect time. Fitting on all years and reporting fit statistics tells you nothing about forecasting performance, and a model evaluated that way will look excellent and fail in use. Rolling-origin evaluation, in which the model is refitted at each historical origin using only prior data and scored on what came next, is the standard, and Tashman set out the case for it in the International Journal of Forecasting in 2000. With twenty years of history you get a small number of test cases, which is a real constraint and an argument for pooling across regions rather than for pretending the in-sample fit means something.
Where this stops
The largest limit is the one that no method removes. Seasonal weather forecasts at a three to six month horizon have modest skill in most temperate regions, so a forecast made at planting is conditioned on climatology rather than on a prediction of this season. The uncertainty at that point is a property of the world and not a deficiency in your model, and a vendor offering a narrow yield forecast in April is either using information nobody else has or misrepresenting the width.
The second limit is management. Crop models assume a management regime: sowing date, nitrogen, irrigation, protection. Growers change these in response to conditions and to prices, and their responses are frequently adaptive in ways that partially offset the weather. A model run on planned management in a difficult year will usually be pessimistic, because it does not know that half the growers changed their nitrogen plan in May.
The third is that yield is one of two numbers you need and the easier one. Quality varies with the same weather that drives yield, often in the opposite direction, and a high-yield year with poor protein or high contamination can be commercially worse than a moderate one. Very few forecasting programmes model quality at all, and the ones that do generally find it harder than yield.
Plot forecast error against week of season for your main crop over the last ten years, and take that curve to whoever is asking for a number in April.