AMD’s Hybrid AI Math: What the 40–60% Savings Model Assumes

Hybrid AI TCO

AMD says a fleet of 500 AI PCs running half their AI work locally can cost 40 to 60% less over three years than doing all of it through cloud APIs. The company built a calculator to show the arithmetic, which is more than most vendors do.

So the number is checkable. The question is what had to be true to produce it.

Run AMD’s Tokenomics Calculator at its defaults and three assumptions carry most of the weight: a daily token volume that only heavy agentic users reach, prompt caching switched off on the cloud side, and a local model whose quality the tool explicitly declines to compare. Change any one and the percentage moves substantially.

None of that makes the model dishonest. AMD discloses its assumptions more thoroughly than the headline suggests, including a frank exclusions list. But a modeled result under stated assumptions is not a property of hybrid AI, and the useful exercise for a buyer is working out which of those assumptions describe their own workload.

Key takeaways
  • The 40–60% figure is a modeled outcome for one scenario: 500 users at the Medium workload tier, 50% hybrid mix, three-year horizon, July 2026 cloud prices. It is not a general property of hybrid deployment.
  • The Medium tier assumes 5.74 million input and 574,000 output tokens per user per day — roughly 200 input tokens per second sustained across an eight-hour day. That is continuous agent activity, not assisted office work.
  • Prompt caching is supported in the calculator but off by default. Enabling it on a repetitive agentic workload can cut cloud input costs dramatically, which is the single biggest threat to the headline result.
  • At a 50% hybrid mix the calculator buys hardware for every user and runs half the tokens locally. The capital cost is full; the local utilisation is not.
  • AMD states plainly that the calculator does not consider performance needs and excludes model-quality differences, software licensing, IT management, migration, taxes, financing and egress.
  • Local throughput is measured on a quantised 35B open model. Comparing its cost against Claude Opus or GPT 5.5 pricing compares two different things.
  • The defensible conclusion: hybrid AI TCO can produce real savings for high-volume, repetitive, quality-tolerant workloads. Whether it does for you is a measurement question, not a percentage you can inherit.

Quick Navigation


What AMD’s Calculator Actually Models

Hybrid AI TCO is the total cost of running AI workloads across a mix of local hardware and cloud APIs, counting hardware, electricity and token charges over a defined period. AMD’s Tokenomics Calculator models three paths — cloud only, local on AMD devices, and a hybrid split — and reports three-year totals, monthly run rates and break-even months.

The inputs it takes:

InputOptions
Team size25, 250, 1,000 or custom
Analysis period1, 3, 4 or 5 years
Workload tierLight, Medium, Heavy or custom, as tokens per user per day
Cloud model mixGemini Pro $2/$12, Claude Sonnet $3/$15, Claude Opus $5/$25, GPT 5.5 $5/$30, or custom, per million tokens
Hybrid mix0% to 100% local
HardwareRyzen AI 9 HX 470, Ryzen AI Max+ 395, Radeon AI PRO R9700, with editable price, throughput, power and hours per day
EnergyElectricity rate and powered days per month
Cloud modifiersPrompt caching, batch discount, regional multiplier, long-context premium, monthly API surcharge

That is a more complete model than the headline implies, and the advanced panel exposes nearly every assumption for editing. Credit where it is due: the caching, batch and long-context controls exist, and most vendor calculators do not have them.

Three mechanics in the fine print matter more than anything on the main panel.

  1. Cloud cost is API tokens only. AMD states there is no seat, subscription or flat monthly price in the cloud figure. For an organisation actually paying per-seat for a coding assistant, this models a different cost structure than the one on their invoice — sometimes higher, sometimes lower.
  2. Hybrid buys hardware for everyone. At 50% mix, all users get a device and half the token volume runs on it. The hardware line is identical to a full local deployment.
  3. Hardware life equals the analysis period. Choose three years and the full device cost lands inside three years, adjusted by an end-of-life value percentage. There is no refresh cycle and no residual beyond that field.

AMD also prices local throughput from a specific measurement: Qwen 3.6 35B A3B at Q4_K_M quantisation, 262,000-token context prefilled to 128,000, sustained throughput under Vulkan llama.cpp. That is a real benchmark configuration, disclosed, and it is the quiet centre of the whole comparison — because it is the model whose cost is being compared against Claude Opus pricing.

The Token Assumption Does Most of the Work

5.74 million input tokens per user per day sounds abstract. Convert it and it stops being abstract.

Our calculations from AMD’s stated tier, using an eight-hour working day and 260 working days.

UnitInput tokensOutput tokens
Per day5,740,000574,000
Per hour (8h day)717,50071,750
Per second sustained~199~20
Per year~1.49bn~149m
Per 500-user fleet, per year~746bn~75bn

Two hundred input tokens per second, sustained, for every second of every working hour. That is not a person typing prompts. That is an agent running continuously, reading a repository, re-sending context and looping through tool calls.

For scale, 5.74 million tokens is in the region of four million words of input per person per day. No human reads or writes that. Only a machine loop produces it.

Three workload profiles, to show the spread. These are illustrative sketches, not measured data.

ProfileWhat it looks likeRough daily input tokens
Light office AIDrafting, summarising, occasional questionsTens of thousands to low hundreds of thousands
Knowledge worker with AI assistanceRegular use, document context, some retrievalHundreds of thousands to low millions
Continuous agentic codingCoding agent with large repo context, long sessionsMillions to tens of millions

AMD’s own Light tier sits at 574K input, a tenth of Medium. The gap between those two tiers is roughly the gap between a typical office deployment and a developer running a coding agent all day.

Why this dominates the result: cloud cost scales linearly with tokens while hardware cost is fixed. At the Medium tier on Claude Sonnet pricing, one user’s input tokens alone come to roughly $4,470 a year, with output adding about $2,240 — call it $6,700 per user per year, or $20,000 over three years. That is our arithmetic at AMD’s stated tier and prices. Against an AI PC costing one to three thousand dollars, the comparison is not close.

Drop to the Light tier and the same arithmetic yields roughly $670 per user per year. Now a $2,000 device takes about three years to pay for itself before counting electricity, management or anything AMD excludes, and the 40–60% headline does not survive.

The lesson is not that AMD picked an unfair number. It is that tier selection, more than hybrid mix, determines the answer. A buyer starting from the default tier rather than their own metered token volume is reading somebody else’s workload.

Where Prompt Caching Changes the Cloud Side

Here is the tension at the centre of the model. The workload that justifies the Medium tier — a continuous agent re-sending large context — is precisely the workload that prompt caching was designed for.

An agent turn is mostly repetition: the same system instructions, the same tool schemas, the same repository or document context, plus a small amount of genuinely new input. Providers price reused prefixes far below fresh input, and the discount is steep. On current frontier pricing, cached reads land near $0.20 per million tokens against input rates of $2 to $5. We worked through what that does to real bills in our analysis of the cache line that decides agentic costs.

AMD’s calculator supports this. The advanced panel has cache write share, cache read share, a write multiplier (1.25x for a five-minute window, 2x for an hour) and a read multiplier. The controls are there and they are correctly specified.

They are also off by default.

What that means in practice: the following is our illustrative arithmetic, not a calculator output. Take the Medium tier’s 5.74M daily input tokens on Claude Sonnet at $3 per million.

Cache-hit rateEffective input cost per user per dayAnnual input cost per user
0% (calculator default)$17.22~$4,477
50%$9.47~$2,462
80%$4.82~$1,253
90%$3.27~$850

At a 90% hit rate, the cloud input bill falls by roughly 80%. Since input dominates this workload’s token volume, that reshapes the entire comparison — and it reshapes it in the direction that hurts the local case.

The honest counterweight: high cache-hit rates are not automatic. They depend on prefix stability, on how the agent structures its context, on cache window duration against request frequency, and on provider implementation. An agent that rebuilds its context every turn gets little benefit and pays cache-write premiums for the privilege. Assuming 90% is as wrong as assuming zero.

So the correct criticism of AMD’s default is narrow and fair: for the specific workload profile that makes the Medium tier plausible, zero caching is the least likely assumption, and it is the one the tool ships with. Any buyer running this calculator should turn caching on and enter their own measured hit rate before reading the savings number.

The Utilisation Problem on the Local Side

The cloud side of this comparison is perfectly elastic: you pay for tokens consumed and nothing else. The local side is the opposite. You buy a device, and whether it earns its price depends entirely on how much work passes through it.

The calculator handles this more carefully than most. It has an “hrs/day running AI” field per device, it sizes capacity from throughput times that window, and it warns when fleet capacity falls short of token demand. Those are the right mechanics.

What the default scenario still implies is worth making explicit. At a 50% hybrid mix, every user gets a device and half the token volume runs on it. The device’s economic output is therefore half what the same hardware could produce — while its cost is unchanged. In utilisation terms, a 50% hybrid mix is a 50%-utilised fleet, at best.

Hybrid AI TCO

The gap between purchased and productive capacity has several sources, and they compound:

  • AI-active hours are not working hours. A device powered for eight hours may run inference for two.
  • One device per user caps sharing. The calculator states a hard ceiling of one device per user. A heavy user cannot borrow a colleague’s idle capacity.
  • Nights and weekends are free capacity nobody uses. Unless batch or background agent work is scheduled into them, two thirds of each day is idle.
  • Provisioning follows the peak. Hardware sized for the busiest user is over-provisioned for everyone else.
  • Adoption is uneven. Fleet-wide deployment assumes fleet-wide usage. Actual AI adoption inside organisations is famously lumpy.

This is the same economics we set out in our analysis of accelerator utilisation: once capacity is bought rather than rented, idle time stops being an efficiency question and becomes a cost-per-unit question. Every hour the device does not run inference raises the effective cost of the hours it does.

The flip side deserves equal weight, because it is the strongest argument for the local case. A device that runs background agents overnight, processes batch jobs on weekends, or serves a developer who genuinely saturates it can reach utilisation a cloud buyer pays dearly for. Local capacity has no marginal token cost. If you can fill it, you should.

The practical test is simple: measure AI-active hours per user, not headcount. A fleet where ten per cent of staff generate ninety per cent of token volume should not be deployed uniformly, and uniform deployment is exactly what the calculator’s hybrid model assumes.

The Missing Variable: Model Quality

AMD says this plainly, twice, and it deserves to be quoted rather than discovered: the calculator does not consider your performance needs, and model inference quality differences are excluded.

That exclusion is doing a lot of work, because of what is on each side of the comparison. The cloud column is priced from Claude Opus, GPT 5.5, Claude Sonnet or Gemini Pro. The local column is throughput-measured on Qwen 3.6 35B A3B at four-bit quantisation. Both are real. They are not the same capability, and the calculator compares their costs without comparing their outputs.

The economic principle is the one that matters here, and it generalises well beyond AMD:

Cost per token is not cost per completed task.

A cheaper model that needs more attempts can cost more in total. Work the chain through:

  • A task that takes two attempts instead of one doubles the token volume for that task
  • A task that needs human review consumes the most expensive resource in the building
  • A wrong output that reaches a downstream system costs far more than either
  • A longer reasoning path to the same answer burns more tokens on cheaper hardware

The arithmetic that follows is unforgiving. If a local model completes 70% of tasks that a frontier model completes at 95%, the effective cost per completed task rises by roughly a third before counting retries — and that is before anyone prices the human time spent on the failures.

None of this says local models are inadequate. Open models at 30B-class sizes have become genuinely capable, and for classification, extraction, summarisation, routing and structured transformation they are often indistinguishable from frontier models at a fraction of the cost. That is a real and growing category of work.

The point is narrower: a TCO model that excludes quality is answering “what does this cost?” rather than “what does this deliver per pound spent?” For the tasks where the two models are equivalent, AMD’s comparison is fair. For the tasks where they are not, the cost column is measuring the wrong thing — and the calculator cannot tell you which tasks are which. Only your own evaluation can.

Why 50/50 Is a Modeling Choice, Not a Strategy

The hybrid mix slider defaults to 50%. It is a reasonable midpoint for a demonstration and a poor description of how workloads actually divide.

Real routing is decided by task characteristics, not by percentages:

Route locallyRoute to cloud
Classification, extraction, taggingComplex multi-step reasoning
Summarisation of routine documentsNovel problem-solving
Privacy-sensitive or regulated dataTasks needing the newest model capabilities
High-volume repetitive transformationBurst workloads above local capacity
Offline or air-gapped workAnything where a wrong answer is expensive
Draft generation before human editingFinal output that ships unreviewed

Run that split against a real workload and the local share lands wherever it lands — 15%, 40%, 80%. The number is an output of the routing policy, not an input to it. An organisation doing mostly structured data work might route 80% locally. One doing mostly open-ended analysis might manage 20%.

This also explains why the calculator’s “max potential savings at 100% hybrid” figure should be read carefully. Full local means no access to frontier capability at all, which is a product decision rather than a cost optimisation.

What AMD’s Model Gets Right

Credit where it is earned, because the tool is better than the headline it generated.

  • It discloses its assumptions. Token tiers, model prices with a date, electricity rate, throughput test configuration, device specs — all visible and editable.
  • It publishes an exclusions list. Model quality, licensing, IT management, migration, taxes, financing, egress, volume discounts. Most vendor calculators omit this entirely.
  • It includes the modifiers that matter. Caching, batch discounts, long-context premiums and regional multipliers are all present, even if off by default.
  • It flags capacity shortfalls. If the fleet cannot serve the token demand, it says so and routes the overflow to cloud.
  • It states its own limitation on performance. “This calculator does not consider your performance needs” is an unusually honest sentence to put in a sales tool.
  • The underlying economics are sound. For sufficiently high, sufficiently repetitive token volumes, local inference genuinely is cheaper. That is not marketing; it is arithmetic.

What It Leaves Out That Could Matter

The exclusions are disclosed, but not all of them are equally consequential. Ranked by how much they could move a real result:

ExclusionWhy it mattersDirection
Model qualityChanges cost per completed task, not cost per tokenUsually favours cloud
Caching, when left offOverstates cloud input costs on repetitive workloadsFavours local as shipped
IT management and supportManaging local inference across a fleet is real workFavours cloud
Software licensingLocal serving stacks, management tooling, model licencesFavours cloud
Migration and deploymentOne-off but substantial for a 500-device rolloutFavours cloud
Hardware refreshLife is set to the analysis period; a 5-year window assumes 5-year devicesFavours local as shipped
Volume discountsEnterprise cloud agreements rarely pay list priceFavours cloud
Seat and subscription pricingMany organisations pay per seat, not per tokenDepends
Utilisation below 100%Capacity bought is not capacity usedFavours cloud

Notice the pattern. Most of the exclusions push in the same direction, and it is not the direction that makes local hardware look better.

A Better Hybrid AI TCO Calculation

Ten steps. The first four are measurement, and skipping them is how buyers end up with somebody else’s answer.

  1. Meter your actual token consumption. Pull it from provider dashboards or your gateway logs, not from headcount multiplied by an assumption.
  2. Separate input from output. They price differently and they scale differently. Input dominates agentic work; output dominates reasoning work.
  3. Measure your cache-hit rate. Most providers report cached-token usage. This single number may move your cloud bill more than any hardware decision.
  4. Measure AI-active hours per user, not working hours, and look at the distribution rather than the average.
  5. Classify tasks by quality requirement. Which actually need a frontier model? Test, do not assume.
  6. Calculate realistic local utilisation. Include nights, weekends and the users who barely touch AI.
  7. Add the omitted costs. Management, support, licensing, migration, refresh.
  8. Compute cost per completed task for both paths, including retries and human review.
  9. Test several routing mixes — 0%, 25%, 50%, 75%, 100% — rather than accepting the default.
  10. Run sensitivity on your two or three largest variables and see how wide the range gets.

Sensitivity: Which Assumptions Move the Answer

Directional, based on the model’s structure rather than on outputs we cannot reproduce exactly.

If this changesLocal and hybrid TCOWhy
Token volume risesImproves sharplyCloud scales linearly, hardware does not
Token volume fallsDeteriorates sharplyFixed hardware cost spread over less work
Cache-hit rate risesDeterioratesCloud input cost falls, sometimes by 80%+
Local utilisation risesImprovesSame capital, more output
Cloud API prices fallDeterioratesThe savings baseline shrinks
Electricity price risesDeteriorates mildlyPower is a small share at these volumes
Hardware price fallsImprovesLower capital to amortise
Quality requirements riseDeterioratesMore work must route to cloud regardless of cost
Refresh cycle shortensDeterioratesCapital amortised over fewer years
Hybrid mix rises toward 100%Improves on cost, narrows on capabilityFewer tokens billed, less frontier access

The two largest levers are token volume and cache-hit rate, and they pull in opposite directions on the same workload. That is the central tension in AMD’s model: the workloads with enough tokens to justify local hardware are the workloads most likely to benefit from caching.

FAQ

What is hybrid AI TCO?

The total cost of running AI workloads across both local hardware and cloud APIs over a defined period, counting hardware purchase, electricity and token charges. It differs from cloud-only TCO because part of the cost becomes fixed capital rather than metered usage.

Can hybrid AI really reduce AI costs by 40–60%?

AMD’s calculator produces that result for a specific scenario: 500 users at 5.74 million input and 574,000 output tokens per user per day, a 50% local mix, a three-year horizon and July 2026 cloud prices with prompt caching off. Under those assumptions the arithmetic holds. Whether it holds for a given organisation depends on whether its workload resembles that profile.

Is local AI cheaper than cloud AI?

For high-volume, repetitive, quality-tolerant workloads on well-utilised hardware, frequently yes. For light or intermittent usage, usually no, because fixed hardware cost is spread across too little work. The crossover point is set by token volume, utilisation and the cache-hit rate on the cloud side.

How does prompt caching affect hybrid AI TCO?

Significantly, and against the local case. Cached input tokens are priced far below fresh ones, so a repetitive agentic workload with a high hit rate can cut its cloud input bill by most of its value. AMD’s calculator supports caching but ships with it disabled, which prices every input token at full rate.

How much does AI PC utilisation affect TCO?

Proportionally. A device used half as much has roughly double the effective cost per unit of work, since the purchase price does not change. At a 50% hybrid mix the calculator buys hardware for every user and runs half their tokens locally, which is a 50%-utilised fleet by construction.

What costs are missing from AMD’s AI cost calculator?

AMD lists them: model inference quality differences, software licensing, IT management, migration effort, taxes, financing fees, network and egress costs unless entered manually, and provider-specific volume discounts. Most of these, if included, would favour the cloud side.

Is a 50/50 local and cloud split optimal?

It is a default, not an optimum. The right mix follows from task routing — which work genuinely needs frontier capability and which does not — and lands anywhere from 15% to 80% local depending on the workload.

How should businesses calculate AI TCO?

Start by metering actual token consumption split into input and output, measure the cache-hit rate and AI-active hours per user, classify which tasks need frontier models, then compute cost per completed task for several routing mixes rather than cost per token for one.

Does AMD’s calculator compare equivalent models?

No, and it says so. Local throughput is measured on Qwen 3.6 35B A3B at four-bit quantisation, while cloud costs are priced from Claude, GPT and Gemini models. AMD states explicitly that the tool does not consider performance needs.


Keep reading

Hybrid AI TCO

AMD’s Hybrid AI Math: What the 40–60% Savings Model Assumes

AMD says a fleet of 500 AI PCs running half their AI work locally can cost 40 to 60% less over three years than doing …
Gemini Free Plan

Gemini Free Plan Drops to Flash-Lite Only on October 9

If you use Gemini without paying, the model picker is about to get shorter. On October 9, 2026, the Gemini Free Plan stops offering Flash …
Agent Identity

Agent Identity: Why AI Agents Shouldn’t Share a Service Account

Most teams authenticate AI agents the way they authenticate cron jobs: one API key, one service account, shared by everything that calls the CRM. That …
Anthropic compute commitments

Anthropic’s $518B Compute Commitment: The Obligation Behind It

The number everyone repeated last week was $518 billion. The number that actually matters is 80%. That is the share of Anthropic’s future compute and …