Skip to content
ServerCalc ServerCalc
Power & Energy

AI Server Rack Power Consumption Calculator

Start from the accelerator count, not the rack. Get racks, megawatts, cooling method and cost per GPU.

Inputs

Accelerators

How many GPUs the deployment requires, not how many racks.

W
90%
Node shape

GPUs per node and how many rack units it takes.

W

CPUs, memory, boot media and chassis, before accelerators.

18%

NICs, fabric, fans and PSU loss, as a share of accelerator power.

Cooling & rack

Sets a hard ceiling on watts per rack, and the facility PUE.

kW

Capped at what the cooling method can actually remove.

U
Facility & cost

Defaults follow the cooling method. Override if you know yours.

$
  • Racks fill on rack units at 5 nodes (40.5 kW). There is cooling headroom going unused, so a denser chassis would cut the rack count.
Results update live as you type.

Results

IT power
0.24 MW
0.26 MW at the facility, PUE 1.10.
Racks needed
7
5 nodes per rack at 37.6 kW.
Nodes
32
256 accelerators deployed.
Power per node
7.53 kW
8.11 kW peak; accelerators are 68% of it.
Rack fill limited by
Rack units
Heat output
821,807 BTU/hr
Cooling required
68 tons
Liquid / air split
85% / 15%
205 kW through coolant, 36 kW to room air.
Annual energy (facility)
2,321 MWh
Energy cost per year
$278,497
Cost per accelerator per year
$1,088
Cost per accelerator-hour
$0.124
Carbon footprint
891 t CO2e/yr
Floor area
196 sq ft

Cooling decides the rack count

Each cooling method puts a hard ceiling on watts per rack, and that ceiling sets how many racks the same accelerators need. Air cooling does not fail gracefully at AI density: it just needs many more racks, each mostly empty, until the numbers stop making sense.

Same accelerator count, different silicon

Accelerator TDP has roughly doubled in two generations. At a fixed GPU count that shows up as megawatts and as racks, long before it shows up on the invoice for the hardware.

AMD Instinct MI355X
1400 W · 40 GPU/rack
0.43 MW$1,959/GPU/yr
NVIDIA B300 (Blackwell Ultra)
1200 W · 40 GPU/rack
0.38 MW$1,710/GPU/yr
NVIDIA B200
1000 W · 40 GPU/rack
0.32 MW$1,461/GPU/yr
AMD Instinct MI325X
1000 W · 40 GPU/rack
0.32 MW$1,461/GPU/yr
Intel Gaudi 3
900 W · 40 GPU/rack
0.30 MW$1,337/GPU/yr
AMD Instinct MI300X
750 W · 40 GPU/rack
0.25 MW$1,150/GPU/yr
NVIDIA H200 SXM5
700 W · 40 GPU/rack
0.24 MW$1,088/GPU/yr
NVIDIA H100 SXM5
700 W · 40 GPU/rack
0.24 MW$1,088/GPU/yr
NVIDIA RTX PRO 6000 Blackwell SE
600 W · 40 GPU/rack
0.21 MW$963/GPU/yr
NVIDIA H200 NVL
600 W · 40 GPU/rack
0.21 MW$963/GPU/yr
NVIDIA A100 80GB SXM
400 W · 40 GPU/rack
0.16 MW$715/GPU/yr
NVIDIA H100 PCIe
350 W · 40 GPU/rack
0.14 MW$652/GPU/yr
NVIDIA L40S
350 W · 40 GPU/rack
0.14 MW$652/GPU/yr
NVIDIA A100 80GB PCIe
300 W · 40 GPU/rack
0.13 MW$590/GPU/yr
NVIDIA A40
300 W · 40 GPU/rack
0.13 MW$590/GPU/yr
NVIDIA Tesla V100 SXM2
300 W · 40 GPU/rack
0.13 MW$590/GPU/yr
AMD Instinct MI210
300 W · 40 GPU/rack
0.13 MW$590/GPU/yr
NVIDIA Tesla V100 PCIe
250 W · 40 GPU/rack
0.12 MW$528/GPU/yr
NVIDIA A30
165 W · 40 GPU/rack
0.09 MW$422/GPU/yr
NVIDIA A10
150 W · 40 GPU/rack
0.09 MW$403/GPU/yr

About this calculator

AI infrastructure is never specified as racks. It is specified as accelerators: 256 H100s, 1,024 B200s, whatever the model needs. The rack count, the megawatts and the cooling method are consequences of that number, and they are usually worked out too late, after the hardware is ordered and before anyone has checked what the building can deliver.

So this calculator runs in that direction. Enter the accelerator count and the node shape, and it returns nodes, racks, kilowatts per rack, total IT and facility megawatts, heat load, annual energy and cost per accelerator per year.

The pivot is the cooling method, and it is a harder constraint than most people expect. Each cooling class puts a ceiling on watts per rack, and that ceiling — not rack units, not your PDU — decides how many racks you need:

  • Air, no containment: about 10 kW per rack.
  • Air with hot/cold aisle containment: about 30 kW.
  • Rear-door heat exchanger: about 40 kW.
  • Direct-to-chip liquid: about 130 kW.
  • Immersion: about 200 kW.

Above roughly 40 kW per rack, liquid stops being an optimisation and becomes a requirement. That is why a single 8-GPU Blackwell node, at around 11 kW, will not share an air-cooled cabinet with anything: one node consumes an entire conventional rack's cooling budget.

The same 256 H100s need 32 racks on plain air cooling and 7 on direct liquid. Air cooling does not fail at AI density; it just multiplies the rack count, each one mostly empty, until the deployment stops making sense.

The formula

Node power is accelerators plus everything that carries them:

accelerator watts = GPUs × TDP × (0.12 + 0.88u)
node watts = accelerator watts × (1 + aux factor) + host base

The aux factor covers what scales with accelerator count rather than being fixed: one high-speed NIC per GPU, the fabric or NVSwitch layer, the fan and pump share, and the power supplies' conversion loss. The default of 18% with a 1.5 kW host base puts an 8 × H100 SXM node at 8.1 kW peak, which matches what this site's independently-built server model reports for a DGX H100 to within 20 watts.

The 0.12 is the accelerator idle floor. A GPU with a driver loaded and no work queued still draws roughly 12% of its TDP, which is why an idle cluster is expensive rather than free.

Racks

nodes = ceil(accelerators ÷ GPUs per node)
nodes per rack = min(rack U ÷ node U, rack budget ÷ node peak watts)
racks = ceil(nodes ÷ nodes per rack)

The rack budget is itself capped by the cooling method: power you cannot remove as heat is not capacity, however willing the electrical feed. The calculator reports which of the two terms bound the rack, because the fix differs. Space-limited means a denser chassis helps. Power-limited means only better cooling does.

Facility

IT power = nodes × node watts
facility power = IT power × PUE
BTU/hr = watts × 3.412, tons = BTU/hr ÷ 12,000

PUE follows the cooling method rather than being independent of it, because it is largely a consequence of it: roughly 1.60 for uncontained air, 1.45 contained, 1.25 with rear-door heat exchangers, 1.10 for direct-to-chip and 1.05 for immersion. Meta published a real example of this, going from 1.28 on advanced air to 1.09 after adding direct-to-chip liquid for GPUs at the same site.

Cost per accelerator

$/GPU/year = facility kWh × price ÷ accelerators

This is the number worth carrying into a business case. It converts a megawatt figure nobody can intuit into a per-GPU running cost that sits next to the per-GPU capital cost and the per-GPU-hour rate a cloud provider would quote you.

Common use cases

  • Turning an accelerator count into racks, megawatts and a facility requirement
  • Deciding whether a deployment needs liquid cooling or can stay on air
  • Comparing H100, H200, B200, B300, MI300X and MI355X on power rather than FLOPS
  • Checking whether an existing air-cooled hall can host any AI at all
  • Costing a GPU cluster per year and per accelerator-hour
  • Sizing the electrical service and cooling plant for an AI build-out
  • Estimating the carbon footprint of a training cluster
  • Explaining to finance why the power bill scales with the GPU count

Frequently Asked Questions

How much power does an AI server rack use?
Between about 10 kW and 130 kW, and the cooling method decides where in that range you land. An air-cooled rack tops out near 10 kW without containment and about 30 kW with it, which is one or three 8-GPU nodes. A rear-door heat exchanger takes you to roughly 40 kW. Direct-to-chip liquid supports about 130 kW, which is where rack-scale systems like the GB200 NVL72 sit. Immersion goes higher still. Enter your accelerator count above and the calculator shows the rack count under every method side by side.
How many GPUs fit in a rack?
Power decides, not rack units. An 8-GPU H100 node draws about 8.1 kW, so plain air cooling fits one node, eight GPUs, per rack. Containment fits three nodes, 24 GPUs. Direct liquid fits five 8U nodes, 40 GPUs, or considerably more with a denser liquid-cooled chassis. Blackwell makes it tighter: an 8-GPU B200 node is around 11 kW, so an uncontained air-cooled rack cannot hold even one. Rack-scale systems solve this differently, putting 72 GPUs in a 48U frame on a shared busbar with mandatory liquid cooling.
How many megawatts does a GPU cluster need?
As a rough guide, an H100 cluster runs about 1 kW of IT load per accelerator once you count host overhead, so 1,000 H100s is roughly 1 MW of IT and about 1.1 MW of facility power on direct liquid cooling. Blackwell is closer to 1.3 kW per accelerator. Multiply by PUE for the facility figure and remember that grid connection lead times, not hardware lead times, are increasingly what sets the delivery date on anything past a few megawatts.
When does an AI deployment need liquid cooling?
Above roughly 40 kW per rack it stops being optional. Below about 30 kW, air with proper hot and cold aisle containment works. Between 30 and 40 kW, rear-door heat exchangers extend air cooling by capturing exhaust before it re-enters the hall. Past that, direct-to-chip cold plates are the only practical route, and they are mandatory on current rack-scale systems. The threshold in accelerator terms: a single 8-GPU Blackwell node already exceeds what an uncontained air-cooled rack can remove.
Does liquid cooling reduce AI power consumption?
It reduces facility power, not IT power. The GPUs draw exactly the same either way. What changes is the overhead: PUE falls from around 1.45 on contained air to roughly 1.10 on direct-to-chip, because you stop spending large amounts of energy moving air. On a 1 MW IT load that is about 350 kW of facility power saved, which at $0.12 per kWh is around $370,000 a year. Liquid also removes most of the server's own fan power, which is IT-side, so there is a smaller second saving there too.
How much does it cost to run an AI GPU for a year?
At $0.12 per kWh, an H100 running training workloads costs roughly $1,100 a year in energy including facility overhead, or about $0.12 per GPU-hour. A B200 is nearer $1,500 a year. Those figures move with your tariff and PUE, and they are worth comparing against cloud GPU-hour rates: energy alone is typically a small fraction of what a cloud provider charges, with the rest being hardware amortisation, networking, staff and margin.
Why does accelerator TDP matter so much for facility planning?
Because it has roughly doubled in two generations and everything downstream scales with it. An A100 is 400 W, an H100 700 W, a B200 1,000 W and a B300 1,200 W. At a fixed accelerator count that is a threefold increase in megawatts, cooling tonnage and annual energy cost, and it arrives as a facility problem long before it arrives as a capital one. A hall built for A100 density typically cannot host Blackwell at all without a cooling retrofit.
What utilization should I assume for an AI cluster?
Far higher than for general-purpose servers, which is the most common sizing mistake. A training run pins every accelerator for days or weeks, so 85 to 95% sustained is realistic. Fine-tuning runs nearer 75%, batch inference 60 to 70%, interactive inference 40 to 50% because it follows demand, and shared development clusters 30 to 40%. Sizing an AI facility off enterprise averages of 20 to 40% under-provisions the electrical and cooling plant badly.
How much heat does an AI rack produce and where does it go?
All of it, at 3.412 BTU/hr per watt. A 40 kW rack produces about 136,500 BTU/hr, roughly 11 tons of cooling from one cabinet; a 130 kW liquid-cooled rack produces about 444,000 BTU/hr, 37 tons. The split matters for plant design: rear-door heat exchangers capture around 80% of it into water, direct-to-chip around 85%, and immersion effectively all of it. The remainder still leaves as air, so a liquid-cooled hall still needs air handling, just far less of it.
Is it cheaper to buy fewer, more powerful accelerators?
On power, usually not by as much as expected. Newer accelerators deliver more work per watt, so at a fixed amount of work fewer of them use less energy. But at a fixed accelerator count, which is how most deployments are actually specified, more powerful parts simply use more power and force denser cooling. The comparison table above holds the count constant deliberately, because that is the shape of the decision most people are actually making, and it shows the facility consequence rather than the performance one.
What limits how fast an AI cluster can be deployed?
Increasingly, power rather than hardware. Anything past a few megawatts needs a grid connection, and utility lead times for new capacity now run to years in many markets. Cooling plant, particularly the water side, is the next constraint, followed by the building itself: floor loading, ceiling height for liquid distribution, and the delivery route for racks that weigh over a tonne. The GPUs themselves are often the shortest pole once you are past the initial allocation.

Spot an error? Have feedback?

Tell us what is wrong with the math, what is missing, or which server model you would like added. We read everything.