Insights & Data

New Framework Turns Corporate AI Emissions From Guesswork Into Auditable Management Data

New Framework Turns Corporate AI Emissions From Guesswork Into Auditable Management Data
Share

Corporate AI use is expanding faster than the systems designed to measure its climate impact.

Published estimates can differ by orders of magnitude because providers disclose little and researchers count different parts of the computing stack.

A Watershed-led framework proposes a three-tier method to move companies from spend-based estimates to token-level provider data while keeping electricity and emissions visible as separate management metrics.

AI’s small footprint is rapidly scaling

Artificial intelligence may still represent a small share of many corporate carbon footprints, but the planning horizon is changing quickly. Data centres used about 5% of United States electricity in 2025 and could consume between 9% and 17% by 2030, according to the estimates cited in a new Watershed white paper.

Corporate use now stretches across employee assistants, coding tools, direct model APIs, cloud-hosted applications and AI embedded inside enterprise software.

However, sustainability teams often see only vendor invoices. They rarely receive the model, token, energy, grid and hardware data needed to produce an auditable Scope 3 estimate.

The paper, led by John Bistline with contributors from Watershed and academia, proposes a common framework built around a comprehensive boundary, emissions per million tokens and three calculation tiers.

Its main achievement is to turn uncertainty into a managed data-quality journey rather than a reason to report nothing.

Corporate accounting cannot measure opaque compute

The same corporate AI workload can produce sharply different carbon numbers depending on the method.

  • In the paper’s hypothetical financial-services company, annual spending of $100,000 and 17.6 billion tokens generated an estimate of 13.4 tonnes of CO2e using a spend factor.
  • More detailed activity data reduced the estimate to 3.7 - 5.4 tonnes, while provider-style data produced 3.2 - 4.5 tonnes.

That approximately fourfold range does not mean one number is automatically true and the others false. Each tier sees a different amount of the system.

  • Published studies may count only active accelerators, while a complete boundary also includes host CPUs and memory, idle capacity, cooling and power overhead, amortised hardware emissions and training.
  • Grid assumptions can also move the result materially.

Workload design adds another layer.

  • A short chatbot prompt is not comparable with image generation, extended reasoning or an agent that triggers repeated model calls.

The framework notes that reasoning models can be around 30 times more energy-intensive than smaller production variants for equivalent queries, as the United States grid subregions vary more than fivefold in carbon intensity.

Three tiers turn uncertainty into estimates

The proposed system boundary covers the full-service stack:

  • Active accelerators, host systems, reserved capacity, data-centre overhead, amortised accelerator hardware and training emissions reported separately.
  • User devices stay outside the boundary because their electricity belongs in the customer’s own Scope 2 account.

The distinction protects comparability and avoids hiding uncertain training allocations inside a single operational factor.

For inference, the functional unit is kilograms of CO2e per million tokens, with input and output tokens separated where possible.

  • Output generation is more energy-intensive than input processing, and the paper’s Activity Tier defaults use a 3:1 decode-to-prefill ratio.
  • Electricity is reported alongside carbon so a company can distinguish computing efficiency from a cleaner grid or renewable-energy claim.

The Spend Tier multiplies AI expenditure by an economic emissions factor, currently 0.134 kgCO2e per 2023 US dollar for the relevant data-processing sector.

  • It is audit-ready but imprecise because price, margins and compute do not move together.

The Activity Tier replaces money with token volumes and modelled energy intensity.

  • The Provider Tier uses supplier-reported per-token figures and is the preferred destination.
  • Provider data remain the binding constraint.

Google publishes a fleet-wide per-prompt figure for Gemini Apps, Meta provides relatively detailed training disclosures for Llama, and major cloud companies offer broader carbon dashboards.

However, model-, customer- and region-specific inference data remain uneven, leaving most buyers dependent on assumptions.

Training is the largest uncertainty in the worked example because no frontier provider discloses the lifetime token volume used to allocate a model’s upfront footprint.

Changing that denominator can move training from 5% to 83% of total estimated emissions.

The framework therefore reports training separately, allowing readers to see which portion is measured, modelled or highly assumption-dependent.

Better disclosure rewards cleaner, efficient providers

Transparent accounting can benefit suppliers and customers.

  • External benchmarks may overstate production inference energy by four to 20 times because they cannot see batching, caching, speculative decoding and hardware-software optimisation at scale.
  • A provider that publishes defensible figures can replace inflated third-party assumptions with data reflecting its actual efficiency.

The framework also connects climate management to cost control.

  • Companies already monitor tokens because tokens drive bills. Shorter context windows, fewer redundant calls, caching and routing simple tasks to smaller models can reduce both expense and electricity.
  • The commercial and environmental incentives therefore reinforce each other when engineering teams measure the same functional unit.

For African and other emerging-market companies, the opportunity is to build good practice before AI becomes material.

  • The paper’s defaults and worked example are US-centred, so local adoption will require credible country or subnational grid factors and clear information about where providers serve African traffic.

Without regional transparency, a precise token count can still be multiplied by the wrong electricity intensity.

Companies can measure and cut together

Companies should begin with an inventory by access channel: hosted assistants, developer tools, direct APIs, cloud-mediated APIs and AI embedded in software.

  • For each channel, procurement, finance and engineering teams should capture provider, model family, spend, input tokens, output tokens, cache activity and serving region where available.
  • The inventory may mix tiers; the largest channels should be upgraded first.

Procurement teams should request a standard disclosure set covering model family, blended or split energy intensity, location- and market-based operational emissions, serving region, PUE, training and embodied emissions, behind-the-meter generation, reference period and assurance.

  • Contracts can require regular refreshes without demanding commercially sensitive details such as batch sizes or proprietary architecture.

Sustainability teams should report electricity and CO2e together, disclose the tier and major assumptions, and keep training as a separate line.

  • A methodology change from Spend to Activity or Provider data may require a rebaseline; that is an improvement in data quality, rather than evidence that operational performance suddenly changed.

Engineering teams then have four practical levers:

  • Optimise prompts, choose the smallest capable model, select lower-carbon regions and prefer efficient data centres with credible clean-electricity matching.

Governance should also watch for rebound effects.

  • Falling energy per query can still raise total consumption if agentic and multimodal use expands faster than efficiency improves.

Carbon is not the whole sustainability account.

  • Data-centre growth can affect local air quality, water demand, grid investment and electricity prices.
  • About 30% of planned US capacity cited in the paper expects behind-the-meter generation, with nearly three-quarters of that capacity using natural gas.

Companies should therefore pair the corporate inventory with supplier questions about water, on-site generation and community impacts where those issues are material.

Build the ledger before emissions surge

Companies do not need perfect data to begin. They need a declared boundary, a transparent starting tier, an inventory of material AI channels and a plan to replace assumptions with provider evidence.

The goal is decision-useful accounting: data that steer model choice, prompt design, region selection and clean electricity procurement.

Establishing those flows now will make future climate disclosures more credible and help prevent AI growth from outrunning corporate control.

More Insights & Data

Start typing to search...