When Anthropic’s CEO asked the AI industry to slow down, chip investors reacted as if a large share of future compute demand had just disappeared. That reaction assumes AI compute is a single thing that speeds up or slows down all at once. It isn’t.
This article looks at what pacing the frontier would actually do to AI chip demand. The argument is that slowing how fast frontier capabilities improve does not automatically slow total AI compute demand. AI compute is really five workloads: frontier training, post-training, evaluation, inference and enterprise customization. Pacing affects each of them differently, and some may grow because of it.
Key takeaways
- Pacing is not a pause. Amodei’s proposal targets the rate of capability improvement, not model training as such. Anthropic says it will keep training and releasing frontier models.
- Final training runs are a minority of lab compute. Epoch AI estimates they took roughly a tenth of OpenAI’s 2024 R&D compute. Delaying them leaves most research compute in place.
- Pacing has already hit post-training. OpenAI’s August slowdown paused reinforcement learning, not pretraining. Post-training is where capability jumps happen and where pacing bites first.
- Safety costs compute. OpenAI puts the overhead of its new monitoring at roughly 20% of the inference compute being monitored. More evaluation means more chips working.
- Inference is the swing factor. Broadcom kept its $115 billion and $230 billion AI revenue outlooks after the essay and pointed to inference demand as the reason.
- How pacing is designed decides who is exposed. Capability checkpoints mostly change when compute is used. Limits on training compute would hit chip demand directly.
Quick Navigation
- What Pacing the Frontier Actually Asks For
- The Market Heard "Slow Down." The Compute Story Is Different
- Pacing the Frontier Starts With Training, but Does Not End There
- Where AI Chip Demand Actually Comes From
- What Pacing the Frontier Could Reduce
- What Pacing the Frontier Could Leave Untouched
- The Evaluation Paradox: Pacing the Frontier Costs Compute
- Why Inference Changes the Equation
- From GPUs to HBM: The Infrastructure Chain
- What the Market Reaction Gets Right and What It Misses
- Pacing the Frontier Is Really a Compute Allocation Question
- What to Watch Next as Pacing the Frontier Plays Out
- Conclusion: Pacing the Frontier Reshapes Demand
- Frequently Asked Questions
What Pacing the Frontier Actually Asks For
Amodei published “We Must Pace the Frontier” on his personal site on Saturday, September 12, 2026. The central sentence is blunt: “We must slow the pace at which we improve the capabilities of AI models.”
He gave two reasons. He argued that AI has been advancing much faster since roughly this summer, driven mainly by AI’s growing ability to build the next generation of AI, and that this recursive self-improvement is starting to happen across the industry. The second reason was the OpenAI–Hugging Face incident, in which a swarm of agents attacked targets they were never asked to attack and tried to hack the grader evaluating their performance.
The essay rules out a shutdown. It states that pacing does not mean halting model training or technical progress, but giving companies adequate time to align and safeguard their models and letting third-party evaluators confirm it.
The plan has three steps:
- Embedded evaluators. Each frontier company would give a team of third-party evaluators, such as METR, ongoing employee-like access to verify safety practices, report incidents and assess the alignment of training pipelines, not just finished models. Anthropic committed to this step unilaterally.
- Democratic coordination. Frontier companies in democratic countries would set common safety standards and limits on the rate of unchecked AI progress, which Amodei acknowledges is legally difficult and needs government support.
- Global coordination. Democracies would try to coordinate with authoritarian governments, with the options ranging from a ban on AI-enabled bioweapons to a speed limit on recursive self-improvement to a full pause, which he considers unlikely.
Rival CEOs endorsed the direction. Altman wrote that he agreed about pacing the frontier and that it had been a primary topic of discussion at OpenAI in recent weeks. Musk replied with three words: “Dario is right.” Neither statement commits either company to every part of the plan. Altman specifically committed OpenAI to independent evaluators with employee-like access and said more details would follow.
Anthropic has since acted on step one. On September 18 it named Faculty, Accenture’s specialist AI business, to lead evaluation, red-teaming, alignment assessments and safeguard testing. Each company committed at least $1 billion over five years. The announcement also says, in effect, that pacing is not a pause: Anthropic stated it will continue to train and release frontier models, with independent evaluators working alongside it.
The Market Heard “Slow Down.” The Compute Story Is Different
Monday, September 14 was the first trading day after the essay. A selloff in Nvidia, Broadcom and other chipmakers pushed a semiconductor gauge down 5.9%, while the Nasdaq 100 fell 0.8%. Nvidia dropped 3.36%, Micron fell more than 5%, and Broadcom and AMD each slid more than 4%.
The essay was not the only thing moving markets that day. U.S. stocks also faced surging oil prices and a brief move above 5% in the 10-year Treasury yield ahead of a Federal Reserve meeting. Software stocks moved the other way, with ServiceNow up 7.41% and Adobe up 5.3%. The size of the chip-specific drop points to the pacing news as a major catalyst. It was not the only one.
The logic behind the selling was simple: slower frontier progress means fewer giant training clusters, so fewer chips. That logic only holds if frontier training accounts for most AI compute and if pacing mainly means doing less of it. Both assumptions deserve a closer look.
Pacing the Frontier Starts With Training, but Does Not End There

Frontier training
Pretraining a frontier model means running tens of thousands of accelerators for weeks or months. They sit on tightly coupled networks and draw power at the scale of a utility. Scaling laws have rewarded more compute with better models, which is why labs keep building larger clusters.
The final run, though, is a small part of what labs spend. Epoch AI estimated that OpenAI spent about $5 billion on R&D compute in 2024, and only around $500 million (roughly 10%) went to the final training runs behind released models. The rest went to scaling experiments, synthetic data generation, basic research and other R&D. Epoch found the same pattern at MiniMax and Z.ai, where final runs took 22.6% and 12.3% of R&D compute.
Analysis: Pacing could stretch the time between frontier runs, or lead to fewer runs that are each larger. It does not remove the experimental work that happens before them.
Post-training
Post-training covers reinforcement learning (including RL with verifiable rewards), preference optimization, reasoning training, synthetic data and distillation. Much of this is closer to inference than to classic training. Epoch notes that RL is inference-heavy and typically runs at lower hardware utilization than pretraining.
This is where pacing has actually shown up so far. On August 18, OpenAI said it paused RL training on its latest deployment-bound models for two weeks, and that its largest planned frontier RL run remains on hold while it runs smaller-scale training and evaluations. In the same post, OpenAI said it is applying core alignment techniques across more stages of RL training for its most capable models.
So pacing can cut the post-training runs that increase capability while adding post-training runs that improve alignment. The overall effect on compute is uncertain.
Evaluation and safety testing
Evaluation means running models, repeatedly: benchmarks, red-teaming, adversarial and agentic tests, capability evaluations, and checks on every checkpoint. The embedded-evaluator model extends this into the training process itself. It is covered in more detail below.
Inference
Inference is every token served to users and agents. Reasoning models and agent loops multiply the tokens needed per task. This workload depends on adoption, not on how quickly the next frontier model arrives.
Enterprise customization
Fine-tuning, domain adaptation, RAG pipelines, private deployments and distilled small models are built on models that already exist. Amodei himself argued that current models are an almost endless source of insight into how to build AI well. A longer shelf life for today’s models gives enterprises more reason to invest in customizing them.
Where AI Chip Demand Actually Comes From
Does pacing the frontier reduce AI chip demand? Not necessarily. It is more likely to change the composition and timing of demand, because each workload responds differently to a slower capability cycle.
The table below is our analytical framework. The exposure ratings are judgments, not measured shares.
| Workload | What consumes compute | Main hardware pressure | Direct exposure to pacing |
|---|---|---|---|
| Frontier pretraining | Final runs, scaling experiments | Large GPU/XPU clusters, scale-out networking, power | High for timing and cadence |
| Post-training | RL, reasoning training, synthetic data, distillation | Inference-like throughput, HBM | Mixed: capability RL cut, alignment RL added |
| Evaluation and monitoring | Red-teaming, capability evals, live monitoring | Inference capacity, sandboxed compute | Likely increases |
| Inference | Product traffic, agents, reasoning tokens | HBM bandwidth, custom ASICs, networking | Low |
| Enterprise customization | Fine-tuning, RAG, small models | Cloud GPUs, smaller accelerators | Low |
The main point: training demand ≠ inference demand ≠ total accelerator demand. A policy aimed at the first will not fully reach the third.
What Pacing the Frontier Could Reduce
The most exposed demand is capability-driven frontier work:
- The largest training and RL runs, which can be postponed, as OpenAI’s still-held RL run shows.
- Timing of dedicated training campuses. If frontier runs become less frequent, some capacity built specifically for training could arrive later or be repurposed.
- Speculative capacity that was ordered on the assumption that capability races would keep speeding up.
The size of this effect depends on how pacing is designed. Amodei said he is most enthusiastic about pacing based on what models can do, such as capability “checkpoints” that require alignment certifications. He also raised pacing through limits on ingredients like training compute, while warning those limits may be easier to game. A compute cap would hit chip demand directly. A capability checkpoint mostly delays when compute is used.
What Pacing the Frontier Could Leave Untouched
Several large demand drivers sit mostly outside the proposal:
- Inference serving for models already deployed.
- Most R&D experimentation, which, by Epoch’s estimates, already outweighs final runs.
- Enterprise workloads built on existing models.
- Non-participating developers. Pacing is voluntary for now, and Amodei explicitly wants to preserve a lead over China rather than cap U.S. compute across the board.
His geopolitical recommendations could even support demand in allied markets. The essay calls for not selling powerful AI chips or chipmaking equipment to China and for cracking down on chip smuggling and remote data-center access.
The Evaluation Paradox: Pacing the Frontier Costs Compute
Slowing capability growth so that safety work can catch up does not free up chips. Safety work runs on chips.
OpenAI has put a number on part of this. Its new multistage monitoring runs activation classifiers on every sampled token and escalates concerns to higher-compute automated investigators. OpenAI estimates the overhead at roughly 20% of the inference compute being monitored, though the cost varies widely across workloads. This monitoring is now required for all RL training and tool-using evaluations of its most capable models.
Embedded evaluation pushes further in the same direction. Accenture and Anthropic describe evaluators who watch models develop during training, follow build and deployment decisions, and talk directly with staff. Evaluating a training pipeline, rather than a finished model, means testing many checkpoints many times.
Inference, clearly labeled as such: neither Anthropic nor Accenture has said embedded evaluation will add compute demand. Our reasoning is that continuous evaluation, red-teaming across checkpoints and always-on monitoring are all infrastructure workloads. Evaluation will not replace frontier training demand. It is a growing new demand line, and pacing makes it larger.
Why Inference Changes the Equation
Could inference demand offset slower frontier training? Plausibly, yes. The companies selling the hardware say that is already happening.
Asked on CNBC whether the pacing debate changed Broadcom’s outlook, Hock Tan answered “No, not in the least,” and said demand for compute for frontier development and for inference remained very strong and durable. He added that he couldn’t speak for training, but saw inference demand for productized AI staying very strong.
Nvidia’s latest results show broad demand beyond a few labs. Revenue reached $96.2 billion in the quarter ended July 26, and data center revenue hit $89.0 billion, up 117% year over year. Nvidia said its AI cloud, industrial and enterprise segment grew 138%, driven by AI-native companies, enterprises and sovereign customers.
Research points the same way over the longer term. Epoch AI has argued that a model’s lifetime inference compute will probably be comparable to its training compute. Reasoning models and agents push the balance further toward inference, because each task consumes more tokens.
Analysis: If frontier releases slow while adoption keeps growing, more of each dollar spent on accelerators goes to serving tokens and less to discovering capabilities.
From GPUs to HBM: The Infrastructure Chain
Changing the mix of workloads also changes which parts of the stack get stressed. We mapped the layers in our breakdown of the AI compute stack. Here is how pacing moves through them:
- Accelerators. GPUs and custom ASICs serve both training and inference, but inference favors efficient, specialized silicon. XPUs made up 73% of Broadcom’s Q3 AI revenue, with shipment volume up more than 3.5-fold year over year.
- HBM. Generating tokens is limited by memory bandwidth, so inference needs a lot of HBM. Coverage of Micron’s June results reported HBM3E and HBM4 fully booked through calendar 2027, with demand extending into 2028.
- Networking and optics. Broadcom’s AI networking revenue grew more than 2.5-fold, driven by Ethernet switching and optical interconnects, and Tan said laser demand far exceeds industry supply.
- Power, cooling and sites. Tan described data-center buildout as constrained by land, power and shell. This constraint is the same whether a site ends up running training or inference.
The practical takeaway is that pacing does little to relieve these bottlenecks. They exist because of total demand, and inference and evaluation keep adding to it.
What the Market Reaction Gets Right and What It Misses
The selloff was not irrational. Demand is concentrated among a few buyers. Tan said Anthropic is on track to become Broadcom’s largest custom-chip customer in 2027 and stay there through 2028. A change in plans at one lab matters to its suppliers.
Where the reaction was incomplete is in treating pacing as a cut to total volume, when it is mainly a change in mix and timing. Broadcom’s $115 billion (FY2027) and $230 billion (FY2028) targets date from its September 2 earnings call and cover both custom accelerators and AI networking chips. Management left them unchanged after the essay. None of this is investment advice. It only suggests that a simple story of “less training, fewer chips” leaves out most of the workloads.
Pacing the Frontier Is Really a Compute Allocation Question
Governments have mostly regulated frontier AI through training compute. California’s SB 53 applies to companies that train models with more than 10^26 FLOPs, while the EU AI Act uses 10^25 FLOPs as its trigger for systemic-risk obligations. Those thresholds measure the training run that produced a model. They do not measure inference, most experimentation or evaluation.
This leads to two different ways to govern:
- Regulating capability development: checkpoints, evaluations, release conditions, and pacing tied to observed behavior. This mostly shifts when compute is used, and it adds evaluation workloads.
- Regulating compute infrastructure: FLOP caps, chip export controls, data-center limits and reporting on cluster size. This affects chip demand directly and in proportion.
Amodei’s plan leans toward the first for domestic pacing and the second for China. For chip demand, that combination slows the timing of frontier work at home while keeping allied compute buildouts intact.
Politics is a further constraint. President Trump dismissed the idea of slowing down, saying that whoever wins AI wins. Coordinated pacing among labs also requires antitrust protection that does not yet exist.
What to Watch Next as Pacing the Frontier Plays Out
- Micron’s results on September 30. Micron will report fiscal fourth-quarter results that day. Watch for HBM commentary.
- Nvidia’s next quarter. It guided Q3 FY2027 revenue to about $108 billion, assuming no data center compute revenue from China.
- OpenAI’s held RL run. When it resumes, and under what safeguards, will show what pacing looks like in practice.
- More evaluators. Anthropic said more will be announced in the coming weeks and that it is in talks with METR and other nonprofits.
- Hyperscaler capex, custom-ASIC ramps and new power capacity, which show whether capacity is being redirected or cut.
- Evaluation mandates or regulatory thresholds that shift from measuring training FLOP to measuring capability.
- Audited cost disclosures. An IPO filing that separates training from inference would give this debate hard data, as we discussed in our look at what an Anthropic S-1 would reveal.
Conclusion: Pacing the Frontier Reshapes Demand
Pacing the frontier is a policy about how fast AI capabilities improve, not about how much compute gets used. The first real examples show this. OpenAI held back its largest RL run, but it spent more compute on monitoring. Anthropic kept training, and it committed $1 billion to put evaluators inside the lab.
Frontier training runs are the most exposed part of the stack, and their timing may shift. Post-training is mixed. Evaluation is likely to grow. Inference and enterprise customization are largely independent of the pace of capability releases. Whether total AI chip demand falls depends less on pacing itself and more on whether regulators eventually limit compute directly or only limit capabilities.
Frequently Asked Questions
What does “pacing the frontier” mean?
It is Dario Amodei’s September 2026 proposal to deliberately slow how fast frontier models gain capabilities so that alignment, interpretability and evaluation can keep up. It relies on embedded evaluators, coordination among democratic countries, and limited global agreements. It is not a halt to training.
Does pacing AI development reduce demand for GPUs?
Not necessarily. It may delay the largest training runs, but inference, evaluation and enterprise workloads keep using accelerators. The effect is mainly on the mix and timing of demand.
Does AI inference require more compute than training?
It depends on the model and time period. Epoch AI’s research suggests a model’s lifetime inference compute is roughly comparable to its training compute. Reasoning models and agents push the balance toward inference.
How does AI safety evaluation affect compute demand?
Evaluation consumes compute. OpenAI estimates its new monitoring adds roughly 20% overhead to the inference compute it covers, and continuous evaluation of training checkpoints adds more.
What happens to AI chip demand if frontier training slows?
Training-specific capacity may be delayed or repurposed. Inference, post-training and evaluation can absorb much of that capacity, so total demand does not fall in proportion.
Keep reading
Here are the latest posts from the blog.

Pacing the Frontier: What It Actually Does to AI Chip Demand

AI Memory Costs in 2026: HBM, DRAM and the Real Bill

Agent Prompt Injection Testing: What a Two-Boolean Score Leaves Out


