Blackwell Ultra vs MI450 vs TPU v7: What Now Wins

Three rack-scale AI accelerator platforms are now shipping, and every vendor claims the lead. All three claims are true, because each measures something different.

Nvidia’s GB300 NVL72 has shipped since January. AMD’s Helios entered full production this quarter. Google’s Ironwood reached general availability on April 22. So the question is no longer which AI accelerator is fastest. It is which one you can actually get, run, and afford to leave.

This AI accelerator guide answers that, and flags every number that does not compare.

Key Takeaways on the 2026 AI Accelerator Choice

  • Blackwell Ultra ships today. Rubin arrives in the second half of 2026, which makes timing the hardest part of the decision.
  • Helios offers the most memory per rack at 31TB of HBM4, against 20.7TB for GB300.
  • TPU v7 Ironwood is capable and rented only. Google still has no public price months after launch.
  • Vendor rack figures use different precisions and different comparison baselines. They are not interchangeable.
  • Software maturity, not peak FLOPS, decides most real deployments.

Quick Navigation

Why This AI Accelerator Comparison Is Not About Specs

Every buyer starts by comparing petaflops. Almost none decide that way. The reason is simple. Every AI accelerator here is fast enough that other things decide it: delivery date, team skills, site power, and exit cost.

So read the spec table as a bar to clear, not a ranking. Every AI accelerator here clears it for frontier work.

Four questions that decide it

Availability. A rack you can install in Q4 beats a faster one arriving in Q3 2027.

Fabric. Proprietary interconnect versus open standards is a ten-year commitment, not a purchase.

Software. CUDA depth against ROCm maturity against a Google-only toolchain.

Exit cost. How much work does it take to move off this AI accelerator in three years?

Blackwell Ultra: What Ships Today

Nvidia shipped the B300 in January 2026. This AI accelerator is a refresh, not a new generation.

Each B300 carries 288GB of HBM3e at 8 TB/s and delivers roughly 15 petaFLOPS of dense FP4. The memory jump comes from 12-high stacks replacing 8-high, not from a new die. Most frontier buyers take the GB300 NVL72 rack, not single GPUs. That rack holds 72 B300 GPUs and 36 Grace CPUs in a 48U liquid-cooled enclosure.

Nvidia rates it at 1.1 exaFLOPS of dense FP4. It holds 20.7TB of unified HBM3e and 130 TB/s of NVLink across the GPU domain. It draws about 120 kW and needs full liquid cooling. Against the GB200 rack, Blackwell Ultra adds 1.5x FP4 compute, double the attention performance, and 50% more memory per GPU.

The timing problem

Here is the awkward part of choosing this AI accelerator now. Rubin, the actual next generation, is expected to reach cloud providers in the second half of 2026.

Buying Blackwell Ultra today means buying the last refresh of an AI accelerator generation. That is fine if you need capacity now, and expensive if you can wait two quarters.

AMD MI450 and Helios: The Open Alternative

AMD launched Helios at its Advancing AI 2026 conference. It is the most credible non-Nvidia AI accelerator rack yet built.

A Helios rack connects 72 MI455X GPUs with 31TB of unified HBM4 and 2.9 exaFLOPS of dense FP4. Each GPU carries up to 432GB of HBM4 at 19.6 TB/s, built on CDNA 5.

Why the open fabric matters

Helios uses Ethernet and open standards throughout. It follows Meta’s Open Rack Wide spec. Scale-up runs on UALink over Ethernet, and scale-out runs on Ultra Ethernet. That is a real architectural difference rather than marketing. NVLink is closed, so an AI accelerator fleet built on it takes one supplier for the wires as well as the chips.

Customer commitments for this AI accelerator are unusually concrete. Anthropic signed for up to 2 gigawatts of MI450-series capacity, with AMD taking up to $5 billion in equity. That sits alongside its existing TPU commitment. OpenAI holds warrants for up to 10% of AMD stock under a 6-gigawatt agreement, and Oracle ordered 50,000 units. AMD claims up to 30% more tokens per dollar than Nvidia’s Vera Rubin NVL72. Note the baseline: that comparison targets Rubin, not Blackwell Ultra.

Shipments began at the end of the third quarter and ramp through 2027. So availability, not capability, is the constraint here.

TPU v7 Ironwood: Capable, but Rented

Google’s seventh-generation TPU reached general availability on April 22, 2026. It is the first such AI accelerator built explicitly for serving rather than training, a shift we traced across the whole inference market here.

Each Ironwood chip pairs 192GB of HBM3e at 7.37 TB/s with about 4.6 petaFLOPS of FP8. It draws roughly 600W, and it is the first TPU with native FP8 hardware. A single superpod links 9,216 chips into 42.5 FP8 exaFLOPS, with 1.77 petabytes of directly addressable HBM connected through optical circuit switches.

No other AI accelerator offers a shared memory domain that large. For big mixture-of-experts serving, that changes what is possible, not just what is fast.

The pricing gap nobody mentions

Now the part that should give any buyer pause. Google has published no public chip-hour rate for Ironwood, months after general availability. The only figure circulating is roughly $1.60 per TPU-hour, which SemiAnalysis estimates Anthropic pays under a large negotiated deal. That is not a list price, and treating it as one would be a mistake.

You also cannot buy the hardware. Ironwood exists inside Google Cloud only, so this AI accelerator comes with a cloud contract attached.

The AI Accelerator Specs Side by Side

AI accelerator figures at rack level, as published by each vendor.

Why the AI Accelerator Numbers Do Not Compare

Read that AI accelerator table skeptically, because four things make direct comparison unsound.

Different units. A GB300 rack and a Helios rack each hold 72 chips. A TPU pod holds 9,216 chips. Comparing a rack to a pod is a category error.

Different precisions. Vendors quote peak figures at whichever format flatters them. Nvidia leads with NVFP4, AMD with dense FP4, Google with FP8.

Different baselines. AMD tests Helios against Vera Rubin NVL72. That is Nvidia’s next generation, not the shipping one. So the claim is tough on AMD and useless here.

Vendor-run tests. AMD reports up to 34x higher token throughput on DeepSeek-V4-Flash against its own last generation. Impressive, and self-measured.

One figure disagrees with itself. AMD’s own page lists 1.7 PB/s of aggregate HBM bandwidth for Helios, while earlier releases said 1.4 PB/s. Check the current datasheet before quoting either.

Availability Is the Real AI Accelerator Constraint

If you take one thing from this AI accelerator guide, take this.

Blackwell Ultra installs now, though Rubin lands in months. Helios is in full production but ramps through 2027, and early slots already belong to Anthropic, OpenAI, Oracle, Microsoft, and Meta. Ironwood ships today, if you accept Google Cloud. Those three AI accelerator positions suit three different buyers, and no benchmark changes that.

What the hyperscaler commitments mean for you

Large deals eat supply. When one deal covers 2 gigawatts, ordinary orders queue behind it.

So ask any AI accelerator vendor for a delivery date in writing before comparing performance. A quoted lead time is worth more than a petaflops figure.

What an AI Accelerator Actually Costs You

Sticker price is the least useful AI accelerator number, and the one everyone asks for.

Three costs stack on top of the hardware. Site work comes first: power, cooling loops, and floor loading. Then engineering time to port and tune. Then the use rate you hit, which swings cost per token more than any spec.

A rack at 30% use costs about three times per token what the same rack costs at 90%. No AI accelerator upgrade gives you a swing that big.

So a slower AI accelerator your team can saturate often beats a faster one they cannot. That is dull advice, and it is usually right.

Owning an AI accelerator means capital, site risk, and a depreciation schedule. Renting means no capital and a price you do not set.

Ironwood forces the rented path. Blackwell Ultra and Helios allow either. For a three-year horizon that difference usually outweighs a 20% performance gap.

Common Mistakes in AI Accelerator Comparisons

Four AI accelerator errors show up repeatedly, and each costs real money.

Comparing peak numbers across precisions. FP4 against FP8 against BF16 tells you nothing. Fix the precision first, then compare.

Ignoring the comparison baseline. A vendor benchmarking against its own last generation is measuring progress, not competitiveness.

Treating a rack as a unit. Rack size, power, and chip count all differ between vendors.

Skipping the delivery date. The fastest AI accelerator you cannot install for eighteen months is slower than the one arriving next month.

Software Is the AI Accelerator Switching Cost

Silicon is the easy part. The toolchain is where AI accelerator projects stall.

CUDA remains the deepest ecosystem, and most published kernels assume it. ROCm has improved a lot, and the AMD-Anthropic deal funds more work on it. Still, the gap is real.

Google’s stack differs again. JAX and XLA are excellent, and they are Google-only. Code tuned for TPU v7 moves nowhere else.

Sizing the switching cost

Count your custom kernels. Teams running stock inference servers move between platforms in weeks. Teams with hand-written attention kernels measure the move in quarters.

That single question predicts migration cost better than any hardware spec. Our glossary covers the underlying terms if the vocabulary is unfamiliar.

Power and Cooling per AI Accelerator Rack

Every AI accelerator here demands facility changes, and the numbers differ enough to matter.

GB300 NVL72 draws about 120 kW in a 48U frame, fully liquid-cooled. Schneider Electric published a 246 kW design for Helios, a double-wide chassis under the Open Rack Wide spec.

That gap is not small. A site wired for one may not take the other without electrical work.

Ironwood sidesteps the question, since Google runs the site. For buyers with no liquid cooling or spare grid capacity, that AI accelerator model is an advantage rather than a compromise.

How to Choose Your AI Accelerator

Four short paths, depending on your situation.

You need capacity this quarter. Blackwell Ultra, and accept that Rubin follows. Availability beats waiting for the next tier.

You are building a multi-year fleet. Look hard at Helios. An open fabric cuts long-term supplier risk, and the memory lead per rack is real.

You serve very large MoE models. TPU v7 pod-scale shared memory is architecturally distinct. Just negotiate pricing hard, since there is no list to anchor against.

You have deep CUDA investment. Stay on Nvidia unless the cost gap is huge. Rewriting kernels costs more than most teams guess, and the bill lands as delay rather than spend.

One rule cuts through all four paths. Pick the AI accelerator you can install, staff, and power this year. A plan that needs none of those things is not a plan.

What Comes After This AI Accelerator Generation

AI Accelerator

All three AI accelerator roadmaps are public, which makes the wait-versus-buy math unusually clear.

Rubin reaches cloud providers in the second half of 2026 and pairs with HBM4. AMD ramps Helios through 2027, and the first gigawatt of the Anthropic build starts in the first half of that year.

Google previewed an eighth generation split into two chips: a Broadcom-designed training part and a MediaTek-designed inference part, both on TSMC’s 2nm process and both slated for late 2027.

That split is the best signal in the whole AI accelerator comparison. Google is the only vendor here dropping one general chip for two specialized ones.

If it works, everyone follows. If not, the general-purpose rack lasts another round. Either way you will know by 2028.

What to Ask Every AI Accelerator Vendor

Before any demo, send the same five questions to all three. The answers sort the field faster than a benchmark does.

First, what is the written delivery date for my order size? Second, what does a full rack draw at sustained load, not peak? Third, which of my frameworks ship day-one support? Then, what does the price look like in year three, not year one? And finally, what breaks if I move this workload elsewhere?

Vendors answer the first four readily. The fifth one tells you the most, because a reluctant answer is itself the answer.

So run that list before you compare a single AI accelerator specification. Most shortlists collapse to one option once the delivery dates arrive.

Conclusion: The 2026 AI Accelerator Decision Is About Terms

Every AI accelerator vendor here builds good silicon. What shapes your next three years is commercial, not technical.

Nvidia sells availability and ecosystem depth, at the cost of a proprietary fabric and a generation about to turn over. AMD sells open standards and memory capacity, at the price of a ramp that runs into 2027. Google sells scale and a simple operating model, at the price of renting rather than owning.

Pick the constraint you can live with, then choose the AI accelerator that fits it. Nobody gets all three. The wider infrastructure picture sits here, and it explains why memory keeps deciding these comparisons.

One AI accelerator prediction worth holding lightly. Google has already previewed an eighth generation split into separate training and inference chips for late 2027. That split says more about where this market is heading than any current benchmark does.

FAQ About the 2026 AI Accelerator Options

Which AI accelerator is fastest in 2026?

No single AI accelerator wins. Vendors publish peak figures at different precisions and scales. A Helios rack lists 2.9 exaFLOPS of dense FP4 against 1.1 for GB300 NVL72, while a TPU v7 pod reaches 42.5 FP8 exaFLOPS across 9,216 chips. Those are not comparable units.

Can you buy TPU v7 Ironwood?

No. This AI accelerator is available only through Google Cloud, and Google has not published a public chip-hour price months after general availability. The one figure in circulation, roughly $1.60 per TPU-hour, is an outside estimate of a negotiated enterprise rate rather than a list price.

Is MI450 better than Blackwell Ultra?

On published AI accelerator specifications AMD leads on memory, with 31TB of HBM4 against 20.7TB of HBM3e. But AMD benchmarks Helios against Nvidia’s next-generation Vera Rubin rather than Blackwell Ultra, and Helios shipments ramp through 2027 while Blackwell Ultra is available now.

How much power does each AI accelerator rack need?

GB300 NVL72 draws about 120 kW in a 48U liquid-cooled frame. Schneider Electric’s published design for AMD Helios is 246 kW in a double-wide chassis. Google gives no per-rack figure for Ironwood, since it runs the sites itself.

Should I wait for Nvidia Rubin?

It depends on your delivery pressure more than on the AI accelerator itself. Rubin is expected at cloud providers in the second half of 2026, so buying Blackwell Ultra now means acquiring the last refresh of the current generation. If you need capacity this quarter, that trade is usually worth making.

How hard is it to switch AI accelerator platforms?

AI accelerator migration scales with how much custom code you wrote. Teams running standard inference servers typically move in weeks. Teams with hand-written kernels tuned to one architecture measure the migration in quarters, which is why ecosystem depth matters more than peak performance for most buyers.

Keep reading

The AI Attack Surface: Securing LLM Systems End to End

5 Hidden Layers of the AI Attack Surface Exposed

OWASP released the 2026 edition of its Top 10 for LLM Applications on August 6. The AI attack surface is highlighted by this update for …

Read more

AI accelerator: Blackwell Ultra vs MI450 vs TPU v7

Blackwell Ultra vs MI450 vs TPU v7: What Now Wins

Three rack-scale AI accelerator platforms are now shipping, and every vendor claims the lead. All three claims are true, because each measures something different. Nvidia’s …

Read more

The AI Compute Stack: Chips, Memory, Power and Cost

The AI Compute Stack: 5 Layers That Now Break First

Every AI story eventually becomes an AI compute stack story. A model launch is really a memory story. Behind a funding round sits a power …

Read more

Inkling 975B: What the Open Weights Now Really Change

Inkling 975B: What the Open Weights Now Really Change

Thinking Machines Lab shipped Inkling on July 15, 2026. It runs 975 billion parameters. It ships under Apache 2.0. And it is the best open …

Read more

Leave a Comment