In short: Most reported methane figures are calculations rather than measurements, an activity count multiplied by an emission factor taken from an industry average. Factor based inventories miss the failure mode that dominates actual emissions, because a small number of large intermittent releases never appear in an average rate per device. A working measurement stack layers continuous site monitoring, periodic aerial or drone survey and satellite screening, and the substantive work is reconciling those layers against each other and against the inventory. Time from detection to repair is the operational number that changes emissions, since a leak found quickly and closed slowly keeps emitting for the whole gap.
Most reported methane emissions are calculations rather than measurements: an activity count multiplied by an emission factor, where the factor comes from an industry average established at some point in the past.
Count the pneumatic controllers, multiply by the assumed leak rate per controller per year, add it up. The number that comes out is a defensible estimate under the reporting standard, it is auditable in the sense that the arithmetic can be checked, and it has one significant weakness: it cannot see the failure that produces most of the actual emissions.
Aerial and satellite surveys have repeatedly found measured emissions substantially above inventory estimates in producing basins, and the reason is structural rather than a matter of the factors being slightly wrong.
Why factor-based inventories miss
Methane emissions are extremely skewed. A large share of total volume comes from a small number of large releases, usually equipment operating outside its design state: a stuck valve, a tank hatch left open, a flare that has gone out and is venting unburned gas, a compressor seal that has failed.
An emission factor describes average behaviour of equipment operating normally. It has no term for a component that is broken, because the average was computed across a population where most components were working. Multiply a factor by a count and you get the emissions of a facility where nothing is wrong.
The consequence is that the estimate is biased low rather than merely imprecise, in a way that grows with the size of the operation, and the error concentrates in exactly the events that could have been fixed if anyone had known.
Two published results establish both halves of that. Brandt, Heath and Cooley, writing in Environmental Science and Technology in 2016, showed that measured leak sizes across natural gas systems follow heavy-tailed distributions rather than anything symmetric, with a small minority of sources carrying most of the volume. Alvarez and co-authors, in Science in 2018, put United States oil and gas supply chain methane emissions for 2015 at roughly sixty percent above the national inventory estimate, and attributed much of the gap to equipment sitting in abnormal operating conditions that the inventory method has no way to represent.
The mechanism reproduces on any population you care to construct. Suppose a site has 400 components of a given type and the emission factor assigns each one unit a year. The inventory reads 400. Now suppose the real population is 396 components emitting 0.6 units because they are working properly, and 4 that have failed and are venting at 60 units each. Measured total is 238 plus 240, or 478, about twenty percent above the inventory, with the entire excess sitting in four components out of four hundred. Move the failure count to 8 and the measured total goes to 715, close to eighty percent above, while the component count and the factor have not changed at all. No adjustment to the factor recovers this, because the factor is describing the 396.
What a measurement stack looks like
Direct measurement has become practical over the last several years, and the useful architecture is layered because no single technology covers the whole range.
Satellite covers large areas at low cost with a high detection threshold. It finds large releases and it will not see a moderate leak. Its value is screening at basin scale and catching the very large events quickly.
Aircraft and drone survey covers a site or a field at moderate cost with a much lower threshold, typically able to attribute an emission to a specific facility or piece of equipment. Periodic rather than continuous.
Continuous site monitors provide fixed sensors around a facility perimeter, inferring source and rate from concentration readings combined with wind data. Continuous coverage, meaningful capital and maintenance cost, and an inversion problem to solve in going from concentration to emission rate.
Ground survey covers component-level detection with optical gas imaging or sniffers. High confidence, labour intensive, and the traditional compliance approach.
The layering matters because each technology has a different detection threshold and a different revisit frequency, and the gaps between them are where emissions hide. A programme with satellite screening and annual ground survey will miss a moderate release that starts the week after the survey and runs for eleven months.
How large those gaps are is arithmetic rather than judgement. A leak beginning at a uniformly random point between two surveys runs on average for half the survey interval before anyone can see it. Annual survey therefore carries an expected detection latency around 26 weeks. Quarterly brings it to about 6 and a half, monthly to a little over 2, and continuous perimeter monitoring with a next-day alert takes detection out of the equation and moves the whole problem into the repair queue. Running that calculation on your own survey plan prices the frequency decision in the same units as the emissions, which is the only way to compare it against the cost of the technology.
It also shows why layering beats any single tier. Satellite screening catches only the very large events, and those are precisely the ones you cannot afford to leave sitting in a 26 week latency window, so it removes the worst part of the tail cheaply and leaves the ground survey to find what it cannot see. The reporting frameworks have converged on the same structure. The Oil and Gas Methane Partnership 2.0, run under the UN Environment Programme since 2020, grades reporting across five levels, with the upper levels requiring source-level measurement, site-level measurement, and an explicit reconciliation between the two.
Reconciliation is the actual work
Once you have measurements, you have a problem you did not have before: two numbers that disagree. The inventory says one thing and the measurement says another, and reporting either alone is unsatisfactory.
Reconciliation is the discipline of explaining the difference, and it is where a measurement programme either becomes useful or becomes an argument.
The components of a reconciliation are straightforward to enumerate and require real work to quantify. Detected events that the inventory does not model, which is usually the largest term. Sources within the measurement footprint that belong to somebody else, which matters in dense producing areas. Temporal mismatch, where a survey caught a maintenance event that is not representative of the annual average. And genuine factor error, where the assumed rate for a normally operating component is simply wrong for your equipment population.
The output that makes this credible is a bridge: inventory estimate, plus and minus each explained component, arriving at the measured figure, with a residual that is stated rather than hidden. A programme reporting a measured number with no bridge back to the inventory has not finished the analysis, and it will be challenged on exactly that point.
Over time the bridge should get shorter as the inventory model absorbs what the measurements taught it. That convergence is the real deliverable, since the goal is an inventory that is right rather than a permanent parallel measurement operation.
Worth running before the programme scales: take one site that already has both an inventory figure and a measurement, and try to write the bridge using only data you hold today. Most teams find at that point that the survey date and the inventory period do not line up, that detected events were logged as findings without a rate estimate attached, that nobody stored wind data alongside the concentration readings, and that ownership of the sources inside the footprint was never recorded. Those gaps are cheap to close for one site and expensive to close for four hundred, and the exercise costs an afternoon.
Detection to repair is the metric that matters
Detecting a leak creates value only when it is fixed, and the interval between those two events is the number worth managing.
That interval decomposes into detection latency, which is a function of survey frequency and threshold, then attribution, which is working out which piece of equipment is responsible, then work order and repair, which is a maintenance scheduling problem.
Most programmes measure the first part and neglect the rest. A site that detects a leak within a week and repairs it fourteen weeks later has captured a small fraction of the available benefit, and the fourteen weeks is usually a prioritisation problem rather than a resource one. Ranking repairs by estimated emission rate rather than by work order date is a small change with a large effect, because the skew means the top few leaks dominate the total.
Put the arithmetic in front of whoever runs the maintenance queue. Take a leak measured at 40 kilograms an hour. Detected in week 2 and repaired in week 16, it ran for 14 weeks. Fourteen weeks is 2,352 hours, so the release is 94,080 kilograms, about 94 tonnes of methane. Cut the repair interval to two weeks and the same leak, found at the same moment by the same technology, releases 13,440 kilograms. The detection programme did not improve by a single kilogram. The work order queue did.
The volume avoided is also worth something directly. Methane that reaches the atmosphere was product, and at plausible gas prices a large sustained leak has a recoverable value that frequently exceeds the repair cost by a wide margin. Methane at standard conditions runs about 0.68 kilograms per cubic metre, so those 94 tonnes are roughly 138 thousand cubic metres, a little under 4.9 million standard cubic feet of gas. Apply your own netback and put the resulting figure on the work order, because a maintenance planner reading a number in the same units as everything else in the queue will rank it differently from one reading an emission rate in kilograms per hour.
The uncertainty that has to be stated
Every measurement carries meaningful uncertainty and the honest treatment reports it rather than quoting a point.
Satellite retrievals depend on surface reflectance, cloud cover and wind fields, and the rate estimate from a single overpass can have a wide band. Perimeter monitors depend on an atmospheric dispersion model to invert concentration into rate, and that inversion is sensitive to wind data quality. Aircraft surveys are snapshots extrapolated to annual figures on an assumption about persistence.
None of that is a reason to prefer the factor-based estimate, which has larger and less quantified error in a known direction. It is a reason to report ranges, to state the method behind each figure, and to be careful about comparing numbers produced by different technologies as though they were the same measurement.
Where this stops
Measurement tells you what is being emitted. It does not tell you what to do about it, and the abatement decision is a separate optimisation with its own economics: cost per tonne avoided, capital versus operating solutions, and the interaction with production.
There is also a boundary question that measurement cannot resolve. Emissions attributable to your operation depend on where the boundary is drawn, and a facility measurement includes whatever is inside the footprint regardless of who owns it. In areas with mixed operatorship this is a genuine attribution problem, and resolving it requires coordination between operators rather than a better sensor.
Finally, a caution about targets. A programme that improves measurement will frequently report higher emissions than the year before, because it is seeing more, and that looks like deterioration to anyone reading the headline. Managing that expectation before the first measured number lands is worth more than any technical decision in the programme.