Does asking an AI to do a task really cost more carbon than doing it yourself? This page won't tell you. It'll let
you argue about it with your own numbers.
Every version of this argument, in both directions, comes down to four numbers. One is measured well. One is
measured well but depends on which basis you think is fair. One is measured, moving fast, and runs out past about
sixteen hours. The fourth isn't really a measurement at all — it's a value judgment wearing a unit.
So this page puts all four on sliders, does the two multiplications, and gets out of the way. There's no answer
here it's trying to sell you, mostly because the defensible range of inputs spans about five orders of magnitude,
and anything more confident than a slider would be lying.
If you land on a number you like, be suspicious of it. That's more or less the point.
At these settings the datacenter emits 1.9% of the CO₂ that the displaced human hours would have.
AI task
Log scale — an agentic coding session runs many model calls, not one query.
Hydro or nuclear ≈ 0.10. US average ≈ 0.37. Marginal US ≈ 0.64. Coal-heavy ≈ 0.80.
Displaced human work
Log scale. This is the slider that decides the answer.
Laptop at home ≈ 0.2. Office plus an average commute ≈ 1.4. Fully loaded ≈ 3+.
Per task
Human hours avoided
11.91 kg
AI datacenter
226 g
Net CO₂ per task+11.69 kgSaved by using the AI
AI is cleaner by52.8×
Break-even energy26.48 kWhWhere the two sides tie, on this grid
Where these numbers come from
Four defaults, four arguments for them, and the sources to check each one against. Where the honest answer is that
nobody knows, that's what it says.
Energy per completed task
The best-measured number in this calculator is also the smallest one. Google published a production measurement in 2025: the median Gemini text prompt drew 0.24 Wh — about nine seconds of television. The default here is some two thousand times that, because a task is not a prompt. An agentic session runs hundreds of calls, with long contexts, reasoning tokens, retries, and tool output fed back in. Nobody has published a per-completed-task figure for that. The log scale is an admission that reasonable people differ here by three orders of magnitude, not a flourish.
Worth knowing before you do your own arithmetic: Google found the accelerator itself is only 58% of the energy — host CPU and memory, idle capacity, and datacenter overhead are the rest. Estimates built from GPU draw alone undercount by roughly half.
The US annual average is 0.37 kg/kWh. The default sits above it deliberately. Average is arguably the wrong basis for a datacenter that adds new load to a grid: what new demand actually causes is the marginal generator, and the EPA puts the US non-baseload rate at 0.64 kg/kWh — nearly double. The default splits the difference, closer to average. That is a judgment call, not a measurement, and it is the first thing to move if you disagree.
Grid intensity is the one input here you can often just look up. If you know which region the compute runs in, use that number instead of this one.
EPA eGRIDEmission rates for all 26 US subregions; eGRID2023rev2 (June 2025) runs ~3% below the figures above
Hours a human would take
And it is the one with the least defensible default. METR measures the 50%-task-completion time horizon — the length of task, measured by how long a human expert takes, that a model finishes half the time. It has roughly doubled every seven months for six years. Rather than bake in a figure that goes stale in a season, this app points you at the live tracker.
METR notes that measurements above 16 hours are unreliable with their current task suite. The top of this slider is therefore past the point where anyone can currently measure — treat it as extrapolation you are doing yourself.
Mostly the commute. The EPA puts an average passenger vehicle at 393 g CO₂e per mile. The average US commute runs around 15 miles each way, so a round trip is roughly 11 kg — about 1.4 kg per hour across an eight-hour day. That is where the default comes from, and you can redo it with your own mileage in your head. Slide down toward 0.2 for someone on a laptop at home; slide up past 3 if you want to load in a share of everything else about employing a person.
The genuinely contested part, and the reason this whole page is a toy rather than an answer: the human does not stop existing when the AI does the task. If the displaced hours get spent on something else that emits, the avoided column is fiction. No slider settles that.
Any one of these can move the answer further than all four sliders put together.
Building the hardwareThe carbon spent manufacturing the GPUs and servers, before any of them run a single token.
TrainingAmortized over however many inferences you think is fair — and that denominator is doing a lot of work.
Whether the hours actually disappearThe big one. A human whose task got automated does not stop existing; they do something else, which may emit more or less. If those hours get reallocated rather than freed, the avoided column is fiction. Nobody has a good answer to this, and economists have been arguing the general version of it since 1865.
WaterCooling has a footprint that carbon accounting does not capture at all.
Failure and retriesThe phrase "energy per completed task" quietly assumes you know how many attempts it took. In agentic work that number is frequently not one.
If that list makes the whole exercise look shaky — yes. This is a slider toy, not a study, and it is most useful for
finding out which of your own assumptions the answer actually hangs on.