Data Center Power: The 4 Hidden Limits on AI Compute

For two years the binding constraint on AI infrastructure was chip supply. Allocation decided who could build.

That has changed, and the reason is a mismatch in clock speeds.

Chip supply chains scale in months. Grid infrastructure scales in years. Interconnection queues, transformer manufacturing and utility capital planning all run on multi-year cycles, and none of them accelerated to match the demand curve.

The practical consequence reverses the old procurement logic. A facility with confirmed power and a later chip delivery date comes online sooner than one with chips in hand and no substation access. Deployment timelines are now set by interconnection dates and equipment delivery schedules.

The Uptime Institute has identified power as the single defining constraint on data centre growth globally. Gartner projects power shortages will restrict 40% of AI data centres by 2027.

One structural shift made this worse than the training-era forecasts assumed. Training is bursty; inference is continuous. As workloads shifted toward serving rather than training, data centre load moved from intermittent peaks to sustained high-wattage draw — a fundamentally harder ask of a grid.

Key Takeaways
  • Of roughly 16 GW of US data centre capacity targeted for 2026, only about 5 GW entered active construction. The gap is not funding and not chips.
  • ERCOT’s large-load interconnection queue grew from 63 GW to 226 GW in a single year. Queue position, not procurement, now sets deployment dates.
  • Power transformers average 128-week lead times and generator step-up units 144 weeks. Transformers are under 10% of project cost and close to 100% of the blockage.
  • Tokens per watt improved roughly a millionfold across six GPU generations. Aggregate demand rose faster, because efficiency creates demand rather than absorbing it.
  • Within a fixed power envelope, efficiency stops being a cost optimization and becomes the only remaining growth lever.

Quick Navigation


The Numbers Behind the Data Center Power Gap

The 2026 figures are stark enough that they need no framing.

MetricValue
US capacity targeted for 2026~16 GW
Actually under active construction~5 GW
Share of remaining pipeline expected to slip30–50%
Large-scale projects tracked~140
Share of those under construction~1 in 3
ERCOT large-load queue growth63 GW → 226 GW in one year
Typical interconnection wait3–7 years
Power transformer lead time~128 weeks
Generator step-up unit lead time~144 weeks
2026 AI capex, four largest hyperscalers>$650 billion

Set the last two rows against each other. More than $650 billion of committed capital, and the binding constraint is a piece of electrical equipment with a two-and-a-half-year queue.

The demand curve underneath is not slowing. Goldman Sachs Research projects US data centre power demand rising from 31 GW in 2025 to 66 GW by 2027. The IEA projects global data centre electricity consumption rising from 415 TWh in 2024 to 945 TWh by 2030.

Individual sites now approach 1 GW, with rack densities exceeding 100 kW for the newest training clusters. These are industrial loads arriving at distribution grids designed for something else.


The 4 Data Center Power Limits, Ranked

Four distinct constraints get compressed into the phrase “power shortage.” They have different causes, different timelines and different workarounds, so separating them is the useful move.

  • Limit 1 — Interconnection queue position. A regulatory and study-process constraint. You cannot connect until the utility has studied your load and the transmission upgrades it requires.
  • Limit 2 — Electrical equipment. A manufacturing constraint. Transformers, switchgear and batteries have multi-year lead times that no amount of capital shortens.
  • Limit 3 — Generation capacity. A physics and permitting constraint. Even with a connection and equipment, the electricity must exist.
  • Limit 4 — Delivery losses inside the facility. An engineering constraint. Power that arrives at the fence does not all reach the accelerators.

Limits 1 and 2 bind hardest right now. Limit 3 becomes dominant if the first two ease. Limit 4 is the only one an individual operator fully controls.


Data Center Power Limits 1 and 2: Queues and Equipment

The interconnection queue is the constraint most often misdescribed as a shortage. Nothing is physically absent; the process is saturated.

ERCOT’s large-load queue growing from 63 GW to 226 GW in a year is not a demand signal so much as a congestion signal. Lawrence Berkeley National Laboratory data shows median interconnection times having doubled since 2008, and analysts assess FERC Order 2023 reforms as unlikely to resolve the underlying physical capacity deficit before 2029.

Data Center Power

Some markets have simply closed. Dominion Energy has stated it cannot accommodate additional large-load interconnection requests in Northern Virginia through 2030 — the densest data centre market in the world, effectively full for four years. PJM, the largest grid operator in North America, has already failed to procure adequate capacity in a recent auction.

Typical waits run 3–7 years against a data centre build cycle of 2–3 years. The queue is longer than the construction project it gates.

Electrical equipment is a genuine physical shortage. Transformers at roughly 128 weeks and generator step-up units at roughly 144 weeks, with some large-transformer lead times quoted at four years as of May 2026.

Domestic production expansion from Hitachi Energy and Siemens Energy is projected to come online no earlier than 2028, which means the shortage persists for at least two more years on current trajectories.


Limits 3 and 4: Generation and Delivery Loss

Generation capacity is the constraint waiting behind the other two. Interconnection reform and transformer capacity would move the bottleneck rather than remove it, because the electricity still has to be generated.

This is where the multi-year nature of the problem becomes unavoidable. New generation — gas, nuclear, renewable with storage — takes years to permit and build. The conditional small modular reactor pipeline grew from 25 GW at the end of 2024 to 45 GW by April 2026, which signals intent rather than delivered capacity, since none of it is producing electricity yet.

Delivery losses are the limit operators can actually act on. NVIDIA’s own analysis notes that at gigawatt scale, up to 40% of power can be lost before it reaches compute — through cooling inefficiency, conversion losses and traditional overprovisioning.

That figure deserves to sit next to the interconnection numbers. A site fighting for four years to secure an extra 100 MW may have comparable headroom available inside its own fence, obtainable through cooling and power-delivery engineering rather than a utility negotiation.

There is a tension worth naming: running closer to thermal and electrical limits recovers capacity and increases fault risk. Recovering that 40% is an engineering programme with real reliability trade-offs, not free capacity.


Why Cheap Parts Block Expensive Data Center Power Builds

Here is the disproportion that makes this era strange, and the single most quotable fact in the whole picture.

The binding constraint set — transformers, switchgear, batteries — represents less than 10% of project cost and close to 100% of the blockage.

Capital is abundant. More than $650 billion of 2026 AI infrastructure spend is committed across four companies. Semiconductors are available. Land is available. What is scarce is the unglamorous electrical equipment that converts capital into energized megawatts.

Two things follow that change how you read industry announcements.

  1. Announced capacity is not deliverable capacity. A press release describes intent. Only the fraction with secured interconnection and equipment on order describes a plant that will exist. Roughly one in three tracked projects is under construction.
  2. Money cannot compress the timeline. In most markets, capital shortens delivery schedules. A 128-week transformer queue does not respond to a higher bid, because the constraint is manufacturing throughput rather than price discovery.

This is where the physical layer meets the economic one — the stack of dependencies from silicon up to serving is mapped in the AI compute stack.


Does Efficiency Solve the Data Center Power Problem?

The obvious rebuttal: chips are getting dramatically more efficient. Does that not resolve this?

The efficiency gains are real and enormous. NVIDIA reports roughly a millionfold improvement in tokens per megawatt across six architecture generations, from Kepler in 2012 to Rubin in 2026 — from roughly one token per megawatt to near 900,000.

Aggregate demand still grew faster.

Google’s disclosed token volume ran from roughly 9.7 trillion per month in May 2024 to 480 trillion by I/O 2025, 1.3 quadrillion by October 2025, and 3.2 quadrillion by May 2026 — about 7× year over year. China reported roughly 140 trillion daily token calls by March 2026, around 1,000× early-2024 levels.

This is Jevons paradox operating at industrial scale. When the cost per unit of useful output falls, total consumption of the input rises, because demand responds to price. Every order-of-magnitude improvement in token cost opens a demand class that did not previously pencil.

Efficiency gains do not moderate aggregate power demand. They enable it.

But the individual-operator conclusion is the opposite of the macro one, and this is the part worth internalizing.

NVIDIA frames it as: Revenue = Tokens per Watt × Available Gigawatts.

If your available gigawatts are fixed by an interconnection queue you cannot jump, then the second term is a constant and tokens per watt is your entire growth curve. A chip that doubles tokens per watt doubles your output within an unchanged power envelope.

That reframes efficiency from a cost optimization into the only available growth lever — which is precisely why accelerator leadership has shifted from raw FLOPS to performance per watt. The hardware side of that shift is covered in memory bandwidth and the limits of AI chips.


How Operators Are Routing Around Data Center Power

Four strategies are visible in 2026, with different risk profiles.

  1. Behind-the-meter generation. On-site gas turbines, fuel cells or dedicated renewable plus storage, bypassing the interconnection queue entirely. Fastest route to energized megawatts and the reason hybrid power deals are rising sharply. The trade-off is that you have become a power generation company.
  2. Geographic arbitrage. Building where interconnection is available rather than where latency is optimal. Viable for training and batch inference, less so for latency-sensitive serving.
  3. Acquiring position rather than building it. Buying sites with existing interconnection rights, or brownfield industrial locations with legacy heavy-load connections. Turns a four-year queue into a transaction.
  4. Squeezing the existing envelope. Liquid cooling, higher voltage distribution, reduced overprovisioning, and accelerator generations with better tokens per watt. The only strategy with no external dependency.

A useful way to read the market: the first three compete for a scarce external resource, and the fourth does not. Operators that treat efficiency as an infrastructure strategy rather than a procurement detail have an advantage that does not require anyone’s permission.


What Would Ease the Data Center Power Constraint

A constraint worth taking seriously deserves an honest account of what would relieve it. Four things could, on different timescales.

  • Permitting and queue reform. Federal legislation achieving substantial permitting reform and cluster-study acceleration would compress the process side of Limit 1. Analysts rate this low-confidence, because process fixes cannot substitute for physical grid expansion — but the queue is partly administrative, so partly addressable.
  • Transformer manufacturing capacity. Hitachi Energy and Siemens Energy expansions are projected to come online no earlier than 2028. Trade arrangements unlocking additional imports could move this sooner. This is the most predictable of the four, because factory build-outs have published timelines.
  • Demand moderation. If token growth slowed materially, existing supply would catch up. Nothing in current data suggests this. Google’s disclosed volumes are running near 7× year over year, and there is no sign of the curve bending.
  • A shift in the binding constraint itself. If interconnection and equipment ease, the constraint moves to generation, which has its own multi-year timeline. Relief in one layer relocates the bottleneck rather than removing it.

The realistic read is that the equipment constraint eases from roughly 2028 and the interconnection constraint persists to around 2029, with generation becoming dominant after that. This is a decade-shaped problem rather than a cycle-shaped one.

Two caveats belong on all of it. Forecasts in this area have a poor track record, and several figures here — announced capacity, queue volumes, projected demand — are estimates from parties with a commercial interest in the number being large. And the constraint is regional rather than national: a market with headroom and a market that is full share a country and almost nothing else.


What Data Center Power Limits Mean for Buyers

Most readers are not building data centres. Four consequences reach anyone buying compute.

  • GPU rental prices will not fall the way chip prices do. Supply is gated by energized capacity rather than manufacturing output. Falling per-token costs have so far increased total spend rather than reducing it, and anyone forecasting cheaper GPU-hours from cheaper tokens has the causality backwards.
  • Capacity commitments are worth more than they look. Reserved capacity is a claim on a genuinely scarce resource. Priced against a market where roughly half of announced 2026 capacity may not materialize on schedule, reservations look different.
  • Utilisation matters more, not less. If capacity is scarce and priced accordingly, an idle GPU wastes something with a four-year replacement lead time. The economics of that are set out in what inference actually costs per token.
  • Regional availability will diverge. With Northern Virginia effectively closed to new large loads through 2030 and ERCOT’s queue at 226 GW, where you can buy compute will increasingly depend on which grids have headroom. Treat region as a capacity question, not only a latency one.

Primary sources

Capacity and queue figures reflect reporting as of mid-2026 and change quickly. Lead times vary by equipment class and supplier; ranges are shown where sources differ.


Frequently Asked Questions

Is power really a bigger constraint than GPU supply?

For deployment timelines, yes. Chip supply chains scale in months while interconnection queues run 3–7 years and transformers average 128-week lead times. Of roughly 16 GW targeted for 2026 in the US, about 5 GW entered active construction.

Why can’t money solve the transformer shortage?

Because the constraint is manufacturing throughput rather than price. Domestic capacity expansions from major manufacturers are projected to come online no earlier than 2028, so the shortage persists regardless of willingness to pay.

Do efficiency improvements fix data center power problems?

Not in aggregate. Tokens per megawatt improved roughly a millionfold across six GPU generations while total demand grew faster, consistent with Jevons paradox. For an individual operator with a fixed power allocation, efficiency is the only growth lever available.

How much power is lost before reaching the chips?

Up to 40% at gigawatt scale, through cooling inefficiency, conversion losses and overprovisioning. Recovering it is an engineering programme with genuine reliability trade-offs rather than free capacity.

Where is data centre capacity still available?

It varies sharply by grid. Northern Virginia’s largest utility has said it cannot accommodate additional large-load requests through 2030, while ERCOT’s queue stands at 226 GW. Availability now depends on regional grid headroom rather than land or capital.


Keep reading

Data Center Power

Data Center Power: The 4 Hidden Limits on AI Compute

For two years the binding constraint on AI infrastructure was chip supply. Allocation decided who could build. That has changed, and the reason is a …

Read more

Self-Hosted LLM Cost

Self-Hosted LLM Cost: The 5 Hidden Fees in Your Bill

The seductive number is the hourly rental rate. An H200 rents for roughly $3.10 to $3.80 per GPU-hour from the cheaper providers, which works out …

Read more

Egress Control

Egress Control: The 7 Hidden Paths Out of Your Agent

There is one structural argument for this control, and it is worth stating precisely because everything else follows from it. Input filtering must recognize the …

Read more

Agent Observability

Agent Observability: The 4 Signals Your Stack Must Emit

Agent observability makes an agentic system legible after the fact. State, decisions, tool calls — captured, replayable, auditable. The vocabulary is borrowed from distributed systems: …

Read more

Advertisement

Leave a Comment