Why Inference Chips Now Win a Much Narrower Race

Inference Chips
Inference chips

Inference Chips – Two things happened in the last eight months, and together they tell the whole story.

Nvidia paid roughly $20 billion for Groq’s inference chips and most of its team. Then Cerebras went public on Nasdaq and raised $5.5 billion. So the two loudest challengers to Nvidia both cashed out big.

Neither maker of inference chips beat Nvidia, though. That is the part most coverage skips.

Inference chips did not win the AI hardware war. They won a narrow, valuable corner of it. And the shape of that corner explains both exits.

Key Takeaways: Inference Chips vs Training Chips
  • Training builds a model once. Serving it runs forever, so the spend never stops.
  • Estimates put running costs at 60–70% of a roughly $400 billion accelerator market in 2026, up from about 40% in 2023.
  • Inference chips chase latency, not raw throughput. That single choice reshapes the silicon.
  • Nvidia paid about $20 billion in December 2025 for Groq’s stack, hired its founder, and shut a door.
  • Cerebras booked $510 million in 2025 revenue, but 86% came from two UAE-linked buyers.
  • Nvidia still holds roughly 80% of accelerator revenue. So the narrow race stays narrow.

Quick Navigation

Why Inference Chips and Training Chips Solve Different Problems

Start with the workload, because the hardware follows from it.

Training is a batch job, and training silicon reflects that. You feed a cluster huge amounts of data, run it for weeks, and nobody waits on any single answer. So the goal is throughput per dollar and per watt.

Serving a model is the opposite, which is why inference chips exist. One user types a prompt and waits. Every token must arrive fast, and the job repeats billions of times a day.

The memory wall

Here is the technical crux. Generating each token means reading the model’s weights again.

That step is memory-bound, not compute-bound. So the chip sits idle waiting on data. GPUs solve this with High Bandwidth Memory stacked beside the die, which is fast but still off-chip.

Inference chips attack the same wall differently. Groq’s LPU keeps weights in on-chip SRAM. Cerebras packs 44GB of SRAM onto one wafer. Both skip the trip to external memory entirely.

Batching hides the problem that inference chips were built to solve. Group 256 requests together, and one weight read serves them all.

But batching adds delay. A user waiting on an agent that chains ten model calls feels every millisecond, ten times over. So the trick that saves money on paper costs you the product.

That gap is the whole opening for inference chips.

How Inference Chips Beat GPUs on Latency

The speed numbers for inference chips are real, and they are not small.

GroqCloud ran Llama 4 Scout at over 460 tokens per second. The same model on H100 hardware lands nearer 100–150. Cerebras has claimed roughly 15x GPU speed on comparable work.

Groq built its inference chips small and deterministic. Its Tensor Streaming Processor drops the cache and thread scheduling that GPUs rely on. Instead the compiler plans every cycle in advance, which makes latency predictable rather than merely fast.

Cerebras built its inference chips the other way. The WSE-3 uses a whole 300mm wafer as one chip: 46,225 square millimetres, 4 trillion transistors, 900,000 cores. It is about 57 times the area of an H100.

Why build a chip that big? Because a model spread across 64 GPUs pays a latency tax at every hop between them. Keep the model on one piece of silicon, and that tax disappears.

The catch nobody advertises

Both families of inference chips trade capacity for speed. SRAM is fast but small, so a single Groq card holds very little of a large model. You chain many together, and the rack gets expensive.

Cerebras hits a version of the same limit. Even 44GB of on-chip memory is tight for frontier models. The Register noted that SRAM capacity remains the constraint on what the design can hold.

So inference chips win on speed and lose on flexibility. That trade defines the narrow race.

The Narrower Race: What Inference Chips Actually Compete For

Zoom out and the market for inference chips splits into layers, not a single prize.

Training frontier models is roughly a two-vendor game. Nvidia dominates it, and Google’s TPUs handle the rest at scale. We covered how Trainium and TPU pressure Nvidia’s margins here, and that fight runs on different rules.

Bulk serving is a price fight that inference chips rarely win. Nvidia, AMD, and hyperscaler silicon all chase cost per million tokens, and CUDA matters less than it used to.

Where inference chips actually live

The real territory is thinner than the pitch decks suggest. Inference chips own the slice where latency is the product, not a feature.

Inference chips suit voice agents, live code completion, and reasoning models that chain many calls. In those cases a two-second wait kills the use case outright.

That slice is growing fast. Still, inference chips live in a slice.

Here is the honest framing. Inference chips did not take the accelerator market. They took the part of it where Nvidia’s architecture is structurally weakest, then sold that position at a premium.

What Groq’s $20 Billion Exit Really Says

The Groq story reads like a win for inference chips, and it was. But read the structure.

On December 24, 2025, Groq licensed its inference chips to Nvidia on a non-exclusive basis. CNBC reported the figure at about $20 billion in cash, citing the CEO of Disruptive, which led Groq’s last round.

The people went too. Founder Jonathan Ross joined Nvidia. So did president Sunny Madra and much of the engineering team. Simon Edwards, the finance chief, took over as CEO.

Three months earlier, Groq had raised $750 million at a $6.9 billion valuation. Nvidia paid roughly three times that, ninety days later.

One caveat matters. DCD noted that neither company confirmed the price, and the $20 billion number has not been independently verified.

Ross built the original Google TPU before founding Groq in 2016. So Nvidia bought the person who has now designed two credible answers to its own architecture.

The structure drew attention too. Senators Warren and Blumenthal asked the FTC whether a licensing deal plus a mass hire sidesteps merger review. That inquiry was still open as of May 2026.

Groq lives on as a company. It kept GroqCloud, distributed $7.6 billion to shareholders in February 2026, and raised $650 million more in June. Yet the architecture team now works at Nvidia.

That is what winning a narrow race looks like. You do not displace the incumbent. You become too useful to leave outside.

What Cerebras Proved, and What It Did Not

Cerebras sells inference chips too, but it took the other road and stayed independent.

The company filed its S-1 on April 17, 2026, then listed on Nasdaq as CBRS in May. It raised about $5.5 billion. Demand ran hot enough that the price range moved up mid-process.

Revenue from its inference chips reached $510 million in 2025, up 76%. Non-GAAP net income came in at $237.8 million, a swing from heavy losses the year before.

The concentration problem

Now the uncomfortable line. Two UAE-linked buyers supplied 86% of that revenue: MBZUAI at 62% and G42 at 24%.

A maker of inference chips with two customers is not really a market yet. So the OpenAI agreement carried enormous weight in the filing.

That deal commits over $20 billion for 750 megawatts of low-latency capacity, expandable toward 2 gigawatts by 2030. Futurum’s teardown treats it as the commercial proof of the architecture.

OpenAI reportedly uses it for a code-generating model. That fits: code assistants are latency-sensitive and run constantly.

What the filing does not settle

Three risks stay open for its inference chips. Wafer-scale yields are harder than normal chip yields, so gross margin under volume is unproven. TSMC capacity has to stretch. And a third anchor customer has not appeared.

Investing.com framed the listing as a direct market test of the serving thesis. That framing is right, and the test is not finished.

The Hyperscaler Squeeze on Inference Chips

A third force presses on inference chips, and it rarely gets named as a rival.

Amazon, Google, and Microsoft all build their own serving silicon. Trainium and Inferentia run inside AWS. TPUs run inside Google. Maia runs inside Azure.

These parts do not have to beat merchant inference chips on benchmarks. They only have to be good enough while costing the parent company less than a purchased GPU.

That is a brutal bar for a startup to clear. A hyperscaler can bundle serving capacity with storage, networking, and billing that customers already use. So switching away carries friction that no token-per-second chart erases.

AMD squeezes inference chips from another angle. Its chiplet design carries more memory per package, and memory is usually the binding constraint when serving models. Analysts put AMD near 5–10% of the accelerator market, mostly on serving work.

What that leaves for inference chips

Squeezed between Nvidia above and in-house silicon below, inference chips need a defensible edge. Ultra-low latency is that edge, and it is narrow by design.

Note the pattern in recent deals, though. AWS paired Trainium3 with Cerebras systems rather than replacing either. So the likely endgame is not displacement. It is inference chips working as a decode accelerator inside a larger fleet.

Where Inference Chips Still Lose to Nvidia

Speed benchmarks for inference chips make a clean headline. Buying decisions are messier.

CUDA locks buyers in less than it once did. vLLM, SGLang, and ONNX Runtime all abstract the hardware underneath.

But “less of a lock” is not “no lock.” Every custom chip needs its own kernels, its own quantization path, and its own debugging story. Teams pay that cost in engineer-months.

A GPU cluster trains on Monday and serves on Tuesday. Inference chips cannot do that, so the hardware sits idle whenever demand dips.

A GPU cluster trains on Monday and serves on Tuesday, while inference chips cannot switch. Fleet flexibility is worth real money to anyone running mixed workloads. Benchmark charts never show it.

The incumbent adapts

Nvidia did not stand still. It licensed Groq’s design and shipped an integrated product pairing LPU-style silicon with its Vera Rubin generation.

Meanwhile AWS paired Trainium3 with Cerebras systems to accelerate its own serving platform. So the specialists increasingly ship inside somebody else’s stack.

Nvidia still books roughly 80% of accelerator revenue. Any honest read of inference chips has to start there.

How to Choose Between Inference Chips and GPUs

Skip the token-per-second charts for a moment. Ask four questions about inference chips instead.

If a user waits on every response, speed is the product. Voice agents, live coding tools, and multi-step reasoning all qualify.

If you run overnight batch jobs, none of this matters. Buy on cost per token and move on.

How big is the model?

Small and mid-sized models fit SRAM-heavy inference chips well. Very large models fight them, because capacity is the binding constraint.

Quantization helps. So does routing: send the latency-critical calls to specialist hardware and the rest to GPUs.

What does a token actually cost you?

Speed and cost pull apart here. A specialist rack can be faster per request and still cost more per million tokens.

Work out the full picture. Count hardware or hourly rate, utilisation, and the engineering time to port your stack. Idle capacity is the silent line item, since a rack that only serves one workload sits unused whenever traffic dips.

Then weigh that against revenue. If faster responses lift conversion or retention, the premium pays for itself. If not, the cheaper fleet wins.

Can you tolerate one vendor?

Concentration cuts both ways here. Cerebras depends on a few buyers, and its buyers depend on a single supplier with novel manufacturing.

Second-sourcing inference chips is expensive. Yet a rack you cannot replace is a risk you carry on someone else’s balance sheet. The Anthropic and Google TPU arrangement shows how large buyers hedge that exposure.

What Comes Next for Inference Chips

Two roadmaps are worth watching, and they point in opposite directions.

The obvious moves are more SRAM and lower precision. Much of the industry has settled on FP8 and FP4 formats, and 3D chip stacking could raise on-wafer memory.

Capacity is the real question. If WSE-4 holds meaningfully more of a frontier model on one wafer, the design gets far more useful. If not, inference chips from Cerebras stay a specialist tool for mid-sized models.

Nvidia and the absorbed architecture

Nvidia already ships a combined product pairing Groq-style inference chips with its Vera Rubin generation. That aims squarely at agentic workloads, where long context meets tight latency budgets.

So the strongest challenger architecture now ships under the incumbent’s brand. Anyone tracking inference chips should read that as consolidation, not competition.

One caution on every speed claim made for inference chips. Vendor benchmarks pick favourable model sizes, batch settings, and context lengths.

We wrote about how saturated benchmarks stop separating systems, and hardware marketing has the same disease. Test on your own prompts, at your own context length, before believing any figure.

One more signal is worth tracking: who else signs. Cerebras needs a second anchor customer to prove the OpenAI deal was not a one-off, and Groq needs GroqCloud to grow without its founding architects.

Watch those two questions through 2027. They will settle the thesis faster than any benchmark round.

Conclusion: The Narrow Race Was the Right Race

Groq and Cerebras bet on inference chips years apart, and made the same call. Both refused to build a general-purpose accelerator, and both aimed at one workload instead.

That looked like a limitation in 2023. It turned out to be the strategy.

Serving models now consumes most accelerator spending, which is why inference chips found buyers, and the fastest-growing part of that spending punishes delay. So a chip that does one thing brilliantly found a buyer willing to pay a premium.

But keep the scoreboard honest. Nvidia holds the market, owns Groq’s architecture, and ships wafer-scale silicon inside partner platforms. Cerebras trades publicly on the strength of one enormous contract.

Winning a narrow race is still winning. It is just not the same as winning the race everyone was watching.

FAQ About Inference Chips

What is the difference between inference chips and training chips?

Training hardware optimizes for throughput across huge batch jobs that run for weeks. Inference chips optimize for latency on single requests that repeat billions of times. The main split shows up in memory: training parts lean on High Bandwidth Memory, while specialist serving parts keep weights in fast on-chip SRAM.

Are inference chips faster than Nvidia GPUs?

On specific workloads, inference chips are faster. GroqCloud has served Llama 4 Scout at over 460 tokens per second against roughly 100–150 on H100 hardware, and Cerebras claims about 15x on comparable tasks. Those gains apply to latency-sensitive serving, not to training or to every model size.

Did Nvidia buy Groq?

Not exactly. Groq licensed its inference chips to Nvidia on a non-exclusive basis in December 2025, and CNBC reported a figure near $20 billion. Founder Jonathan Ross and much of the team joined Nvidia, while Groq stayed independent under a new CEO and kept GroqCloud running.

Is Cerebras profitable?

Cerebras reported $510 million in 2025 revenue from its inference chips, up 76%, with non-GAAP net income of $237.8 million. Two UAE-linked customers accounted for 86% of that revenue, so the profit rests on a very narrow base.

Will inference chips replace GPUs?

Unlikely, and the market is not moving that way. Nvidia still books around 80% of accelerator revenue, and GPUs handle most serving today. The realistic outcome is a split fleet, where specialist hardware takes the latency-critical calls and general-purpose silicon handles everything else.

How large is the market for inference chips?

Estimates for inference chips vary. Analysts put total accelerator revenue near $400 billion in 2026, with serving workloads at 60–70% of it, up from roughly 40% in 2023. Treat those ranges as directional, since definitions of “inference spending” differ between sources.

Keep reading

Inference Chips

Why Inference Chips Now Win a Much Narrower Race

Inference Chips – Two things happened in the last eight months, and together they tell the whole story. Nvidia paid roughly $20 billion for Groq’s …

Read more

Anthropic S-1: What an AI Lab Must Now Disclose

Anthropic S-1: What an AI Lab Must Now Disclose

On June 1, 2026, the company behind Claude said something short and carefully lawyered. The Anthropic S-1 arrived as a confidentially submitted draft registration statement …

Read more

Why Composite Benchmarks Now Fail: The Kimi K3 Proof

Why Composite Benchmarks Now Fail: The Kimi K3 Proof

Two labs. Two completely different models. Nearly identical scores on MMLU, ARC, and HellaSwag. That happened this month, and composite benchmarks are the reason nobody …

Read more

Figure, Agility, Apptronik: Who's Winning the $39B Race?

The Ultimate Truth About the $39B Humanoid Robot Race

The race to build commercially viable humanoid robots is burning through billions in venture capital, and the numbers have moved fast enough to catch most …

Read more