Watt If

Does asking an AI to do a task really cost more carbon than doing it yourself? This page won't tell you. It'll let you argue about it with your own numbers.

brought to you by element30

Every version of this argument, in both directions, comes down to four numbers. One is measured well. One is measured well but depends on which basis you think is fair. One is measured, moving fast, and runs out past about sixteen hours. The fourth isn't really a measurement at all — it's a value judgment wearing a unit.

So this page puts all four on sliders, does the two multiplications, and gets out of the way. There's no answer here it's trying to sell you, mostly because the defensible range of inputs spans about five orders of magnitude, and anything more confident than a slider would be lying.

If you land on a number you like, be suspicious of it. That's more or less the point.

At these settings the datacenter emits 1.9% of the CO₂ that the displaced human hours would have.

Log scale — an agentic coding session runs many model calls, not one query.
Hydro or nuclear ≈ 0.10. US average ≈ 0.37. Marginal US ≈ 0.64. Coal-heavy ≈ 0.80.
Log scale. This is the slider that decides the answer.
Laptop at home ≈ 0.2. Office plus an average commute ≈ 1.4. Fully loaded ≈ 3+.
Human hours avoided
11.91 kg
AI datacenter
226 g
+11.69 kg Saved by using the AI
52.8×
26.48 kWh Where the two sides tie, on this grid

Where these numbers come from

Four defaults, four arguments for them, and the sources to check each one against. Where the honest answer is that nobody knows, that's what it says.

Energy per completed task

The best-measured number in this calculator is also the smallest one. Google published a production measurement in 2025: the median Gemini text prompt drew 0.24 Wh — about nine seconds of television. The default here is some two thousand times that, because a task is not a prompt. An agentic session runs hundreds of calls, with long contexts, reasoning tokens, retries, and tool output fed back in. Nobody has published a per-completed-task figure for that. The log scale is an admission that reasonable people differ here by three orders of magnitude, not a flourish.

Worth knowing before you do your own arithmetic: Google found the accelerator itself is only 58% of the energy — host CPU and memory, idle capacity, and datacenter overhead are the rest. Estimates built from GPU draw alone undercount by roughly half.

Datacenter grid intensity

The US annual average is 0.37 kg/kWh. The default sits above it deliberately. Average is arguably the wrong basis for a datacenter that adds new load to a grid: what new demand actually causes is the marginal generator, and the EPA puts the US non-baseload rate at 0.64 kg/kWh — nearly double. The default splits the difference, closer to average. That is a judgment call, not a measurement, and it is the first thing to move if you disagree.

Grid intensity is the one input here you can often just look up. If you know which region the compute runs in, use that number instead of this one.

Hours a human would take

And it is the one with the least defensible default. METR measures the 50%-task-completion time horizon — the length of task, measured by how long a human expert takes, that a model finishes half the time. It has roughly doubled every seven months for six years. Rather than bake in a figure that goes stale in a season, this app points you at the live tracker.

METR notes that measurements above 16 hours are unreliable with their current task suite. The top of this slider is therefore past the point where anyone can currently measure — treat it as extrapolation you are doing yourself.

Human CO₂ per work-hour

Mostly the commute. The EPA puts an average passenger vehicle at 393 g CO₂e per mile. The average US commute runs around 15 miles each way, so a round trip is roughly 11 kg — about 1.4 kg per hour across an eight-hour day. That is where the default comes from, and you can redo it with your own mileage in your head. Slide down toward 0.2 for someone on a laptop at home; slide up past 3 if you want to load in a share of everything else about employing a person.

The genuinely contested part, and the reason this whole page is a toy rather than an answer: the human does not stop existing when the AI does the task. If the displaced hours get spent on something else that emits, the avoided column is fiction. No slider settles that.

What this leaves out

Any one of these can move the answer further than all four sliders put together.

  • Building the hardware The carbon spent manufacturing the GPUs and servers, before any of them run a single token.
  • Training Amortized over however many inferences you think is fair — and that denominator is doing a lot of work.
  • Whether the hours actually disappear The big one. A human whose task got automated does not stop existing; they do something else, which may emit more or less. If those hours get reallocated rather than freed, the avoided column is fiction. Nobody has a good answer to this, and economists have been arguing the general version of it since 1865.
  • Water Cooling has a footprint that carbon accounting does not capture at all.
  • Failure and retries The phrase "energy per completed task" quietly assumes you know how many attempts it took. In agentic work that number is frequently not one.

If that list makes the whole exercise look shaky — yes. This is a slider toy, not a study, and it is most useful for finding out which of your own assumptions the answer actually hangs on.

Got a number that's wrong, or a source better than one of these? That's the genuinely useful outcome — info@element30.com.

Everything here runs in your browser. Nothing you drag is recorded, and nothing is sent anywhere. Brought to you by element30.