How the GPU cluster power calculator works
Somewhere between "how much power does one GPU use" and "why is my state building a nuclear plant" sits a question people actually type: what does it take to run 100,000 H100s? This calculator answers it with public spec sheets and four honest multipliers. Start with the GPU's rated TDP (the number NVIDIA prints), multiply by how many you have, then by a node overhead factor (because every GPU drags CPUs, memory, network cards, and fans along with it), then by sustained utilization, then by PUE (because the building itself eats power for cooling and conversion). That gives facility power at the utility meter, and from there a year of energy, the electricity bill, and the CO2 on your grid are just multiplication.
The satisfying part: run the defaults with 100,000 H100s and this page lands at 152.9 MW. xAI's Colossus in Memphis, the real cluster that made that number famous, started life as 100,000 H100s drawing about 150 MW from the grid (the Tennessee Valley Authority approved exactly that much for it in November 2024). The calculator knows nothing about Memphis. It just multiplies spec sheets, and the answer comes out where reality did.
The formula
annual kWh = facility kW × 8,760 hours
TDP is the per-GPU rated power in watts. Overhead is node power divided by GPU power (default 1.82, derived below). Utilization is the share of rated power the cluster actually sustains (100 percent for training). PUE is total building power divided by IT power (default 1.2 for a modern AI build). Cost is annual kWh times your price; CO2 is annual kWh times the grid's intensity.
Worked example: the 100,000 H100 question
100,000 H100 SXMs, node overhead 1.82, sustained at 100 percent, PUE 1.2, electricity at 8 cents per kWh, US average grid:
100,000 × 700 W = 70 MW of bare GPU. × 1.82 node overhead = 127.4 MW of IT load. Training pins the rated power, so that is the sustained draw. × PUE 1.2 = 152.9 MW at the meter.
Over a year: 152.9 MW × 8,760 hours = 1,339,228,800 kWh, about 1.34 TWh. At $0.08 per kWh the power bill is $107,138,304 a year. On the US average grid that carries 466,052 metric tons of CO2, and the same electricity would run about 127,546 average US homes. At 2.5 people per household, that is roughly every home in a city of 300,000 people, powering arithmetic instead of dishwashers.
TDP is the sticker, the meter reads higher
TDP (thermal design power) is the maximum sustained power a GPU is designed to draw and its cooling designed to remove. It is a rating, not a measurement: your GPU can sit far below it (an idle H100 draws well under 100 W) and can brush against it all day under load. The reason this page defaults utilization to 100 percent is that large training runs genuinely do pin it. Modern clusters are tuned to maximize model FLOPs utilization, which in electrical terms means keeping the silicon saturated for weeks. Inference fleets are different: traffic breathes, batches vary, and average draw lands meaningfully below rated power, which is what the utilization field is for.
But the GPU is never the whole bill, and here the spec sheets do the arguing for us. An NVIDIA DGX H100 holds 8 H100s, which is 8 × 700 W = 5.6 kW of GPU TDP. The system's rated maximum is 10.2 kW. The other 4.6 kW is everything else in the box: two Xeon CPUs, 2 TB of RAM, eight 400 Gb network cards, NVMe storage, and the fans fighting to move all that heat. Divide and you get 10.2 / 5.6 = 1.82, our default overhead factor. Two honest caveats. First, that is nameplate against nameplate: a real training workload does not slam every component to its maximum simultaneously, so sustained overhead usually lands below 1.82, and the field is editable for exactly that reason. Second, the factor varies by design: a GB200 NVL72 rack runs about 120 kW for 72 GPUs at 1,200 W each, an overhead closer to 1.4, because rack-scale liquid cooling strips out most of the fans. Pleasingly, if you enter 8 GPUs at our defaults, the IT line reads 10.2 kW: the calculator reproduces the DGX datasheet, because it was built from it.
Why AI clusters forced liquid cooling
A decade ago a full server rack drew 5 to 10 kW and air conditioning handled it. An H100 node draws 10.2 kW by itself, and a GB200 NVL72 packs about 120 kW into a single rack, more than ten times the old ceiling. Air physically cannot carry heat away that fast at sane velocities, which is why every serious AI build since 2024 pipes liquid straight to the chips. This is also why PUE for AI facilities is improving even as total power explodes: liquid cooling is more efficient than chillers and fans, so a purpose-built AI hall can run a PUE near 1.2 while the industry average across all data centers still sits at 1.54 (Uptime Institute's 2025 survey). If you want the facility view of that story, sizing racks, rooms, and cooling from the building down, that is our data center power calculator. If you want the opposite altitude, what one chatbot prompt costs in watt-hours, that is the AI energy calculator. This page is the middle altitude: the cluster itself.
The grid decides the carbon
A cluster is exactly as clean as the wires feeding it. The same 100,000 H100s from the worked example emit 260,276 metric tons of CO2 a year on California's grid, 466,052 tons on the US average, and 753,158 tons on a coal-heavy Midwest grid: a 2.9x spread from location alone, with not one setting on the cluster changed. This is why the siting announcements you read always lead with power contracts, and why the same headline cluster can be described as a climate problem or a solved problem depending on which reporter called which utility. The grid factors here are EPA eGRID 2023 data, the same constants our EV pages use, so the two halves of this site can never quote you different grids.
What this page does not claim
This calculator prices power: what the meter reads while the cluster runs. It deliberately does not estimate what training a specific model costs or how much energy a training run consumed, because those depend on two numbers only you have: how long the run lasted and what utilization it actually sustained. Multiply this page's facility power by your own duration and you have that answer; we will not pretend to know your duration for you. It also excludes the embodied energy of manufacturing the chips and building the hall, and the water story (evaporative cooling is its own honest topic). One page, one claim, priced carefully.
Sources and method
- GPU TDPs: NVIDIA datasheets and spec pages: A100 SXM 400 W, H100 SXM up to 700 W configurable, H200 SXM up to 700 W, RTX 4090 450 W total graphics power. B200 at about 1,000 W (HGX air-cooled configuration, widely reported, including Dell's COO on the record). GB200 NVL72 at 1,200 W per Blackwell GPU (SemiAnalysis hardware analysis; note the 120 kW figure often quoted is the whole 72-GPU rack, not a GPU).
- Node overhead 1.82: NVIDIA DGX H100 datasheet, 10.2 kW max system power against 5.6 kW of GPU TDP. Nameplate-to-nameplate, editable because sustained real-world overhead runs lower.
- PUE 1.2 default: modern hyperscale AI builds target 1.2 or below; Google reports a fleet-wide 1.10 and Microsoft 1.12, while Uptime Institute's 2025 global survey puts the industry average at 1.54, its sixth flat year.
- Electricity price: EIA average retail price for industrial customers, about 8 to 9 cents per kWh through 2025 and 2026. Big AI campuses negotiate their own contracts, so the field is editable.
- Grid CO2: EPA eGRID 2023 data (Rev 2, June 2025), pounds of CO2 per MWh by region, held in lockstep with our EV calculators.
- Scale anchors: xAI's Colossus, 100,000 H100s stood up in Memphis in 2024 (about 122 days from bare building to running cluster, then doubled to 200,000 GPUs in 92 more), with a TVA-approved 150 MW grid connection. Average US household electricity about 10,500 kWh per year (EIA).
Figures as of August 2026. TDPs are ratings, overheads are design-dependent, and utility contracts are private, so treat results as honest engineering estimates, not an audit.