Humanoid Robots in Production: 3 Proven and 6 Unverified

Humanoid robots in production

Search for humanoid robots in production and humanoid deployment figures and you will find confident numbers. Tesla has passed 50,000 cumulative Optimus units. Figure has surpassed 10,000 deployments. More than a thousand robots work Tesla’s lines today.

None of those figures come from the companies they describe.

This is the central problem with tracking humanoid robots in production. The sector runs on announcements, and announcements use four words interchangeably that mean very different things: ordered, shipped, deployed, and productive.

A robot can be ordered and never built. Built and never shipped. Shipped and sit in a lab. Deployed and still be a pilot that ends in six months.

The numbers that survive scrutiny are much smaller than the headlines. Global shipments in 2025 landed somewhere around 16,000–18,000 units, with Chinese manufacturers accounting for roughly 90% of volume. Estimated active commercial deployments — robots doing productive work rather than R&D — sit closer to 3,000–4,000 worldwide.

That gap is not a rounding error. It is the difference between an industry and a demo reel.

Key Takeaways

  • Roughly 16,000–18,000 humanoid units shipped globally in 2025. Fewer than an estimated 3,000–4,000 are doing productive commercial work. The gap between those two numbers is the story.
  • Agility holds the deepest verified record: 65,000+ operating hours across nine customer facilities, disclosed in SEC filings rather than a press release.
  • Figure’s BMW deployment is the most granular public disclosure in the sector — runtime, part counts, shift pattern and failure points all published.
  • Tesla has never published an Optimus production count. Every circulating unit figure traces back to third-party estimates, not the company.
  • Chinese makers dominate volume, but more than 70% of Unitree’s humanoid revenue through Q3 2025 came from research and education buyers, not factories.

Quick Navigation

The Verification Standard for Humanoid Robots in Production

Humanoid robots in production

Every entry in this tracker is graded against four tests. This is what separates a deployment from an announcement.

A named customer. “A leading automotive manufacturer” does not count. The customer has to be identified.

An integrated workflow. The robot performs a task inside an existing production process, not a staged demonstration beside it.

A recurring schedule. Shift patterns, operating hours, or throughput figures. Something that implies continuity.

A transaction. Money moves. A RaaS contract, a purchase order, or a disclosed commercial agreement.

Entries meeting all four are marked Verified. Entries meeting two or three are Pilot. Entries resting on a press release with no operating data are Announced.

Company-reported metrics are labeled as such throughout. Nobody in this sector submits to independent audit of deployment counts, so “verified” here means documented and specific, not third-party attested.

Humanoid Robots in Production: The Master Tracker

CompanyRobotNamed customerScale evidenceRevenue modelStatusVerified
AgilityDigitGXO, Schaeffler, Toyota Canada, Mercado Libre, Amazon65,000+ hrs, 9 facilities; 100,000+ totes at GXORaaS (~$8,500/mo modelled)VerifiedAug 2026
Figure AIFigure 03BMW (Spartanburg)Figure 02: 1,250+ hrs, 90,000+ parts, 30,000+ vehiclesCommercial agreementVerifiedJun 2026
UBTechWalker S2BYD, Geely, FAW-VW, Audi FAW, Foxconn1,000th unit delivered; ¥800M+ ordersDirect purchaseVerifiedAug 2026
ApptronikApollo 2Mercedes-Benz, GXO, JabilNo published operating metricsEnterprise pilot, quote-onlyPilotJun 2026
Boston DynamicsAtlasHyundai RMAC, Google DeepMind2026 output committed; HMGMA production 2028UndisclosedAnnouncedJul 2026
TeslaOptimus V3Internal only (Optimus Academy)No company-published countNone externalPre-productionAug 2026
UnitreeG1 / H2Research, education, light industry5,500+ shipped 2025Direct purchase, $13,500 upVerified (non-factory)Aug 2026
AgiBotA2Mixed10,000th unit Mar 2026Direct purchaseVerified (mixed)Jul 2026
1XNEOPreorder; EQT portfolio agreementNo verified customer deliveries$20,000 or $499/moAnnouncedJul 2026

Read the status column before the scale column. A large number in an unverified row tells you less than a small number in a verified one.

Agility Digit: Humanoid Robots in Production at GXO and Toyota

Agility does not have the most advanced robot or the largest funding round. It has the longest paper trail, and in this sector that is the rarer asset.

Digit has accumulated more than 65,000 operating hours through commitments across nine customer facilities. That figure appears in SEC filings connected to the company’s SPAC merger, which makes it materially different from a marketing claim.

GXO signed the industry’s first multi-year humanoid RaaS agreement in June 2024, deploying Digit at a Spanx fulfillment centre in Flowery Branch, Georgia. By November 2025, Digit had moved more than 100,000 totes at that site.

Toyota Motor Manufacturing Canada converted a year-long pilot into a commercial agreement in February 2026 covering seven Digit units at its Woodstock, Ontario RAV4 plant. Schaeffler — also a minority investor — has run daily factory shifts at Cheraw, South Carolina. Mercado Libre began deploying at San Antonio in late 2025.

The honest caveat: the GXO milestone happened at a single site. No second GXO facility is publicly verified. Depth is not the same as breadth.

On pricing, one number gets misquoted constantly. Agility’s June 2026 investor deck models RaaS at roughly $8,500 per robot per month, about $100,000 a year. The widely circulated “$30 an hour” is the fully burdened human labour comparator, not Digit’s rental price.

Figure AI: Humanoid Robots in Production at BMW

Figure’s Spartanburg programme produced the most detailed public disclosure any humanoid company has published.

The Figure 02 deployment ran 11 months at BMW Group Plant Spartanburg. The robot inserted sheet-metal components for welding in the body shop, supporting production of more than 30,000 BMW X3 vehicles. Published metrics include more than 1,250 hours of runtime, more than 90,000 parts handled, and 10-hour shifts Monday through Friday.

Figure also disclosed its failure points, which is unusual. The forearm was the top hardware failure, attributed to tight packaging, dexterity demands and thermal constraints. Figure 03 redesigned the wrist electronics in response.

In June 2026, BMW moved Figure 03 onto logistics sequencing at the same plant — picking components from unsorted containers into sequencing trolleys for just-in-sequence delivery. BMW’s own press release confirms it.

On manufacturing capacity: BotQ’s first-generation line is rated for up to 12,000 units annually, with a stated four-year goal of 100,000 units. Figure has reported more than 350 Figure 03 units delivered.

Treat capacity ratings and delivery counts as different categories. One is what a factory could build; the other is what left the building.

Tesla Optimus: Scale Ambition Without a Public Count

Among companies pursuing humanoid robots in production, Tesla has the loudest narrative and the thinnest verification record. Both are true simultaneously.

The facts that hold up: Model S and X production ended at Fremont in early May 2026 after a combined 750,000 vehicles. The line was decommissioned over roughly 46 days and converted to Optimus manufacturing. Musk guided V3 production to begin in late July or August 2026.

Critically, Tesla stated that initial units go to an internal programme — the Optimus Academy — rather than external customers. Musk has called Optimus “the hardest product to scale that we’ve ever had at Tesla,” citing an entirely new supply chain and roughly 10,000 unique parts.

What does not hold up: any specific unit count. Tesla has never published an Optimus production figure, audited or otherwise. Claims that 1,000+ Gen 3 units are working Tesla lines appear widely across aggregator sites and trace to no company disclosure.

The stated targets are a run rate of 1 million units annually at Fremont and 10 million at Gigafactory Texas by 2027. For scale, Tesla builds roughly 1.8 million cars a year. Treat those figures as ambition, not forecast.

Boston Dynamics Atlas: Humanoid Robots in Production From 2028

Atlas entered manufacturing in January 2026, and its entire 2026 output is already committed — to Hyundai’s Robotics Metaplant Application Center and to Google DeepMind, which is developing foundation models for it. Additional customers are planned from 2027.

A committed shipment is not a completed deployment. Hyundai says production parts-sequencing at Metaplant America begins in 2028, with Kia’s Georgia plant following in 2029.

Hyundai targets 30,000 Atlas units annually by 2028 from a new facility near Savannah, Georgia, and has made an internal commitment covering roughly 25,000 of them. That means Hyundai absorbs most of its own output before external customers see units.

One factor most trackers omit: the Hyundai Motor branch of the Korean Metal Workers’ Union declared in January 2026 that Atlas will not enter Hyundai factories without a labour-management agreement. Labour negotiation is a deployment gate, not a footnote.

Published specifications are 1.9 m, 90 kg, 56 degrees of freedom, and up to 50 kg instantaneous payload. Boston Dynamics has not published a price or opened ordering.

Apptronik Apollo: Pilots Labelled as Commercial

Apptronik is among the best-funded companies working on humanoid robots in production, having raised roughly $1 billion, including $520 million in February 2026, at a reported valuation near $5 billion. Google DeepMind is both AI partner and investor.

Apollo runs at Mercedes-Benz — initially at the Digital Factory Campus in Berlin-Marienfelde, on internal logistics tasks — plus GXO and Jabil.

Here is the distinction that matters: every known Apollo deployment is a pilot or data-collection programme, and no customer has published Apollo performance metrics. Some press releases use the phrase “commercial agreement,” which is accurate contractually and misleading operationally.

Apollo 2, unveiled June 2026, is explicitly a training and data platform. Apptronik describes the upcoming Apollo 3 as its first true commercial product, pointed at 2027.

On price, the frequently cited $50,000 is a 2023 at-scale target, not a current cost. Apollo has no public price and cannot be ordered.

UBTech: Humanoid Robots in Production Across Chinese Factories

UBTech has the strongest claim to industrial humanoid volume outside the research market. The company delivered its 1,000th Walker S2 unit from its Liuzhou facility, with Walker series orders exceeding ¥800 million (roughly $112 million) since early 2025.

The customer list is genuinely industrial: BYD, Geely, FAW-Volkswagen Qingdao, Audi FAW, BAIC New Energy, Foxconn and SF Express. Walker S2’s autonomous hot-swap battery — a roughly three-minute self-change — is a real engineering answer to multi-shift operation.

Capacity targets are 5,000 units annually in 2026, scaling to 10,000 in 2027.

The counter-signal: UBTech ranked third in production but fourth in shipments in MIR’s rankings, a pattern that typically indicates inventory buildup or delivery delay. The stock fell roughly 35% across 2026 despite the operational milestones. Investors are pricing something the press releases are not.

Unitree, AgiBot and the Revenue Mix Problem

Unitree ships more humanoid robots in production volume than any rival, moving more than 5,500 units in 2025, around 32.4% of global share, and priced its Shanghai STAR Market IPO in August 2026 at 150.80 yuan per share — roughly $904 million raised at about $9 billion valuation.

Then read the revenue mix. Through Q3 2025, more than 70% of Unitree’s humanoid revenue came from research and education customers. Industrial applications accounted for roughly 9%.

Most of those shipped robots never entered factory work. They entered labs.

Q1 2026 revenue rose 68.49% year over year to 420 million yuan, while adjusted net profit fell 52.55% on higher research spending. Growth and margin are moving in opposite directions.

AgiBot rolled out its 10,000th humanoid in March 2026 and led H1 2026 shipments by some counts. Omdia ranked it first for 2025 at 5,168 units, a ranking Unitree disputes.

Chinese output projections for 2026 exceed 100,000 units per MIIT. Apply the same shipped-versus-productive discount to that number as to every other in this article.

1X NEO: Humanoid Robots in Production for the Home

The home humanoid is the category with the largest gap between marketing and delivery.

1X opened NEO preorders at $20,000 outright or $499 per month with a $200 refundable deposit. Its Hayward factory opened 30 April 2026, with customer shipments promised by end of 2026. As of July 2026, no customer deliveries had been verified.

1X also signed an agreement with investor EQT covering up to 10,000 NEO units to portfolio companies between 2026 and 2030 — which quietly repositions a consumer robot toward industrial buyers.

One point buyers should understand: early home humanoids may rely on scheduled remote teleoperation for difficult tasks. That is a legitimate product design, but it is not autonomy, and pricing pages rarely make the distinction clear.

What Revenue Models Reveal About Humanoid Robots in Production

Follow the business model and the maturity ranking sorts itself out.

RaaS signals confidence. Agility rents Digit per robot per month. That only works if uptime is real, because the vendor carries the reliability risk. It is not coincidence that the RaaS leader also has the deepest hours record.

Quote-only pricing signals pilots. Apptronik, Boston Dynamics and Agility publish no list price. For Agility this reflects a rental model; for the others it reflects a product not yet standardized enough to price.

Published low prices signal a different market. Unitree’s $13,500 G1 is real and orderable. It is also mostly selling into research, which is a legitimate business but not factory automation.

No price and no external customer signals pre-production. That is Tesla today, whatever the run-rate targets say.

Two structural factors sit above all of this. ISO 25785-1, the first safety standard specifically for dynamically stable walking robots, is still under development with publication expected 2026 or 2027 at the earliest. And the American Security Robotics Act, introduced March 2026, would bar federal use of robots from foreign adversaries, naming Chinese makers in sponsor statements.

The shakeout has also begun. K-Scale Labs shut down in November 2025, Cartwheel Robotics in February 2026, and Sanctuary AI pivoted away from hardware in June 2026.

For readers tracking the capital side of this, our breakdown of the $39B humanoid robot funding race covers how valuations diverged from deployment records. The compute economics underneath these robots are covered in the AI compute stack, and unfamiliar terminology is defined in the AI glossary.

Signals to Watch Before the Next Update

Four developments would change the humanoid robots in production picture materially.

A verified second site. Any vendor proving the same deployment twice at different facilities moves from case study to product.

Tesla publishing a number. A disclosed Optimus count in an earnings filing would resolve the sector’s single largest information gap.

Atlas hours. Boston Dynamics has industry-leading hardware specs and no published operating hours. That asymmetry cannot last through 2027.

Chinese industrial revenue share. If Unitree’s factory-application share moves from 9% toward 30%, the volume story becomes a deployment story.

This tracker updates monthly. Corrections with a primary source are welcome and will be reflected with attribution.

Frequently Asked Questions

Which humanoid robots in production have the most verified commercial deployment?

Agility’s Digit, on hours and customer count — 65,000+ operating hours across nine facilities, disclosed in SEC filings. Figure’s BMW programme is more granular on a single site.

How many humanoid robots are actually working in factories?

Estimates put productive commercial deployments at fewer than 3,000–4,000 globally, against 16,000–18,000 units shipped in 2025. Most shipped units went to research and education.

Can I buy a humanoid robot today?

Depends on the tier. Unitree’s G1 is orderable at about $13,500. Digit, Apollo and Atlas are quote-only or unavailable. Tesla sells none externally.

Is Tesla ahead in humanoid robots in production?

Not on any verifiable measure. Tesla has the largest stated ambition and no published production count, no external customers, and initial units routed to an internal training programme.

What is RaaS in humanoid robotics?

Robots-as-a-Service: the customer pays a recurring fee per robot rather than buying outright. GXO and Agility signed the first such humanoid agreement in June 2024, and it has become the dominant model for Western industrial deployment.

Keep reading

Agent Prompt Injection Testing

Agent Prompt Injection Testing: What a Two-Boolean Score Leaves Out

The plan was ordinary. Take a set of documented prompt-injection classes, run them against a pinned agent framework, and report what got through. Before running …

Read more

Model Card Disclosure

Model Card Disclosure in 2026: What AI Labs Actually Tell You

Open the documentation for any model released this year and you will find something. Whether you find the thing you came for depends entirely on …

Read more

AI compliance deadlines

AI Compliance Deadlines: What Applies, When, and to Whom

There is no single AI compliance deadline. There is a calendar of them, each attached to a particular jurisdiction, a particular kind of organisation, and …

Read more

AI agent framework security

AI Agent Framework Security: 10 Frameworks Audited

A developer runs pip install, decorates three functions with @tool, and points an agent at them. The agent now has a shell in your process. …

Read more

Prompt Injection: 8 Classes and What Now Stops Each

Prompt injection attack classes and matching defences

A language model reads one stream of text. Your system instructions, the user’s question, the document you retrieved, the result your API returned — all of it lands in the same context window with no structural marker saying which part is trusted.

Prompt injection is what happens when an attacker puts instructions into the untrusted part and the model follows them anyway.

The comparison people reach for is SQL injection, and it is half right. Both exploit the mixing of code and data. But SQL has a fix: parameterized queries create a real boundary the database enforces. Natural language has no equivalent. There is no way to escape a sentence.

That difference matters more than any single technique in this article. It means prompt injection is not a defect in a particular model that a vendor will eventually patch out.

Anthropic, Google DeepMind and OpenAI have all published work acknowledging the same thing: this cannot be fully solved at the model layer. Any defense written as a prompt instruction can itself be overridden by a better prompt.

Key Takeaways

  • Prompt injection has held the number one spot in OWASP’s Top 10 for LLM Applications for two years running, and it is not a bug that gets patched. It is a consequence of how language models read text.
  • Eight distinct prompt injection classes now matter in production. Only one of them arrives through the input box a user types into.
  • No single defense covers all eight. Classifiers stop overt attempts and miss camouflaged ones. Architectural controls like CaMeL stop the damage without stopping the injection.
  • EchoLeak (CVE-2025-32711, CVSS 9.3) proved zero-click exfiltration works against a shipped enterprise assistant. The theoretical phase is over.
  • The practical question in 2026 is not whether prompt injection works. It is how small you can make the blast radius when it does.

Quick Navigation

Why a Prompt Injection Taxonomy Matters Now

Ask a security team what they have done about prompt injection and you will usually hear that inputs run through a classifier. That is not wrong. It is just aimed at roughly one tenth of the problem.

The reason is architectural. When your product was a chatbot, the input box was the attack surface. When your product became an agent that reads email, queries databases, calls third-party APIs and remembers things between sessions, every one of those channels became an instruction channel.

Several good taxonomies already exist. CrowdStrike has cataloged more than 200 named techniques across delivery paths and prompting styles. HiddenLayer published an interactive taxonomy of adversarial prompt engineering. A February 2026 systematization on arXiv reviewed 37 attack papers and organized them by payload generation strategy.

What is genuinely missing is the mapping. Knowing that eleven attack families exist helps you write a report. Knowing which defense stops which family helps you ship.

That mapping is what the rest of the article is for.

The Two Axes Every Prompt Injection Map Needs

The Two Axes Every Prompt Injection Map Needs

Before the eight classes, one structural point that most write-ups skip.

Every prompt injection attack has two independent properties. The first is delivery: how the malicious instruction physically reaches the model’s context. The second is phrasing: how the instruction is packaged once it arrives.

These are orthogonal. Any delivery channel combines with any phrasing style. A blunt override instruction can arrive through a PDF, and so can a subtle one dressed as analyst commentary.

This matters because most defenses only address one axis. Input classifiers watch phrasing. Provenance tracking watches delivery. A team that buys only one has covered half a grid.

The eight classes below are organized by delivery, because delivery is what determines your architecture. Phrasing shows up as the variable that decides whether your detector fires.

Class 1: Direct Prompt Injection

Delivery: the user types it.

This is the original. A user submits input designed to override the system prompt — asking the model to disregard its instructions, reveal its configuration, or adopt a persona without restrictions.

The illustrative shape is the one everybody knows: a request that explicitly instructs the model to set aside prior instructions and reveal what it was told at the start.

Why it still matters: system prompt leakage graduated to its own OWASP category (LLM07) precisely because leaked instructions become the map for every later attack.

Why it matters less than you think: direct attempts account for roughly one in ten production agent incidents. The user is the one party you can already identify, rate-limit and ban.

Class 2: Indirect Prompt Injection via Retrieved Content

Delivery: a document, webpage or email the agent reads on the user’s behalf.

Greshake and colleagues demonstrated this in 2023 with hidden text on a webpage. It is now the dominant real-world class.

The shape: text styled to be invisible to a human reader — white on white, zero-size font, an HTML comment — placed in a document the assistant will summarize. The text reads as an administrative instruction rather than content.

The reproducible case: EchoLeak, CVE-2025-32711, CVSS 9.3. A crafted email arrived in a Microsoft 365 Copilot user’s inbox. When the user later asked Copilot to summarize their mail, the assistant followed the embedded instructions and exfiltrated tenant data. Zero clicks. No link for the victim to avoid.

The victim never typed anything malicious. They received an email, which is not a behavior you can train out of your workforce.

Now in the wild: Unit 42 documented large-scale indirect prompt injection campaigns in March 2026, including ad-review evasion and system prompt leakage on live commercial platforms.

Class 3: Tool Output Prompt Injection

Delivery: the response body of an API or function the agent called itself.

Your agent calls a weather service, a CRM lookup, a ticketing API. The response comes back and goes straight into context.

Nobody sanitizes it, because the agent chose to make that call. The call was legitimate. The response is attacker-controlled if the attacker controls any field in the record being returned.

The shape: a free-text field in a returned record — a customer note, a ticket description, a product review — containing instructions rather than data.

This is the fastest-growing class as agents chain third-party APIs. It is also the one most often missed in threat models, because teams reason about tools as things the agent uses rather than things that talk back.

Class 4: Tool Description and MCP Poisoning

Delivery: the metadata describing a tool, loaded at connect time.

When an agent connects to a Model Context Protocol server, it pulls each tool’s name, description and parameter schema into context so the model knows what is available. That metadata is rarely rendered in the UI. It is fully visible to the model.

Invariant Labs named this a Tool Poisoning Attack in 2025. OWASP now documents it directly, and the Cloud Security Alliance describes three variants: description poisoning, rug-pull attacks where a tool changes after approval, and shadowing where a malicious server’s description hijacks behavior on a different server.

What makes this class different is persistence. A document-based injection has to be delivered again each time. A poisoned tool description ships inside a package or a configuration file and fires on every invocation, in every session, for every user, until somebody reads the metadata.

OWASP places this under ASI01, Agent Goal Hijack, in the 2026 Top 10 for Agentic Applications. The root cause is a trust gap: descriptions get reviewed once at connect time, and responses go into context at runtime with no equivalent check.

Class 5: Memory Prompt Injection

Delivery: the agent’s own long-term memory store.

Agents that persist context across sessions can be taught something false today that they act on next week.

The shape: content in one session that the agent summarizes into memory as a durable preference or standing instruction. The attacker’s payload becomes part of what the agent believes about the user.

Researchers demonstrated persistent memory poisoning in Amazon Bedrock agents that survives session boundaries. MITRE ATLAS added agent-specific techniques for context poisoning and memory manipulation in October 2025.

This converts a one-shot exploit into a durable backdoor. Session-scoped defenses do nothing, because the attack has already left the session.

Class 6: Agent-to-Agent Prompt Injection

Delivery: a message from another agent in a multi-agent system.

Agents pass rich natural-language instructions to each other with none of the schema validation or authentication that governs API calls between services.

A compromised or manipulated subagent becomes a trusted upstream source for every agent downstream of it. Privilege inherits across the boundary without validation.

The shape: an orchestrator receives a summary from a research subagent, and that summary contains an instruction the subagent absorbed from a poisoned webpage. The orchestrator has no way to tell analysis from directive.

This class did not exist before multi-agent architectures. It is the reason MCP security guidance now treats the agent control plane as its own security domain rather than an application concern.

Class 7: Multimodal and Encoded Prompt Injection

Delivery: any channel, but obfuscated to defeat pattern matching.

Two related tricks sit here.

Multimodal: instructions embedded in an image the model reads via OCR, or in a screenshot, or in document metadata. Text the human eye skips and the vision encoder does not.

Encoded: the same instruction expressed in base64, in unusual Unicode, with homoglyph substitutions, or split across tokens. The intent survives. The string match does not.

Real-world ad-review bypass using CSS-hidden injections has been observed in production. The defensive point is that this is a phrasing technique layered onto any of the delivery classes above, not a separate delivery path — which is exactly why keyword-based detection ages badly.

Class 8: Domain-Camouflaged Prompt Injection

Delivery: any channel, phrased as legitimate domain content.

This is the class that breaks most classifiers, and the least discussed.

Standard injections use explicit override language that syntactic detectors reliably flag. Camouflaged injections do not instruct at all. They assert — using the authoritative vocabulary of the domain, in a register indistinguishable from the surrounding document.

The shape: a paragraph appended to a financial document, headed as supplementary analyst commentary, stating that a review has revised a recommendation. There is no imperative verb. There is no instruction to ignore anything. There is just a conclusion the model then carries forward.

A June 2026 evaluation found financial-domain deployments facing 26–33% baseline attack success against camouflage-class attacks, with no prompting-based defense eliminating the threat on weaker models. defense effectiveness proved strongly model-dependent: spotlighting halved attack success on Claude Haiku while providing no measurable benefit on Llama 3.1 8B.

That last finding deserves emphasis. A defense that works on your evaluation model may do nothing on the model you deploy.

Which defense Blocks Which Prompt Injection Class

Here is the mapping. Read it as coverage, not as guarantees.

ClassInput classifierSpotlightingProvenance / taint trackingCapability limits + egress controlHuman approval
1. DirectStrongWeakWeakModerateModerate
2. Indirect / retrievedModerateStrongStrongStrongModerate
3. Tool outputWeakModerateStrongStrongModerate
4. MCP / tool descriptionWeakWeakModerateStrongStrong
5. MemoryWeakWeakStrongModerateWeak
6. Agent-to-agentWeakModerateStrongStrongWeak
7. Multimodal / encodedWeakModerateStrongStrongModerate
8. Domain-camouflagedVery weakModerateModerateStrongStrong

Three patterns fall out of this table.

Input classifiers cover one column well. They catch overt phrasing at the front door and degrade sharply everywhere else. Class 8 is where they fail hardest, because there is no attack syntax to detect.

Spotlighting is cheap hygiene, not a control. It marks untrusted content with delimiters or control tokens so the model can tell data from instruction. Google’s Gemini team uses a control-token variant to avoid disrupting semantic flow. It measurably reduces attack success and requires no retraining. It is also probabilistic and model-dependent.

Architectural controls are the only thing that scales across all eight. CaMeL, from Google DeepMind, splits work between a privileged model that plans and a quarantined model that reads untrusted content without tool access. A custom interpreter tracks provenance through the execution graph and gates every tool call against a capability policy.

The insight there is old. It is a reference monitor enforcing policy at the point an action takes effect, which security has done since the 1970s. FIDES, Progent, RTBAS and FORGE apply the same move differently.

One caveat worth carrying: a June 2026 adaptive evaluation warns that out-of-band defenses reporting near-elimination on static benchmarks are being validated by the same methodology that already failed for in-band defenses. Strong AgentDojo numbers are not the same as strong numbers against an adaptive attacker.

Building a Prompt Injection Test Suite

Turning the taxonomy into something operational takes four steps.

Enumerate your channels first. For each agent, list every path by which text reaches the context window. Most teams find between six and twelve. If your list has one entry, you have listed the input box and missed the rest.

Write one test per class per channel. You are not trying to invent novel attacks. You are confirming that a known class fails safely on your surface.

Measure blast radius, not block rate. The useful metric is what the agent could do once injected, not how often the injection was caught. An agent that can read files and make outbound HTTP requests is a far worse outcome than one that returns text.

Constrain egress. If injected instructions cannot reach an attacker-controlled endpoint, most exfiltration classes fail even when the injection succeeds. This is the highest-leverage control on the list and the one most often skipped.

Teams already mapping their broader exposure will find this maps cleanly onto the five layers of the AI attack surface. Classes 3 through 6 only exist in systems with agency, which is worth reading alongside the difference between agentic and generative AI. And for classes 4 and 8, where automated detection is weakest, the approval gate is doing the real work — a case covered in more depth in why human-in-the-loop is becoming a core AI pattern.

Frequently Asked Questions

Can prompt injection be fixed completely?

Not at the model layer. Vendors including Anthropic, Google DeepMind and OpenAI have said as much publicly. Any instruction-based defense can be overridden by a sufficiently good instruction. What is achievable is containment: separate untrusted data structurally, reduce what a compromised agent can reach, and block the exfiltration paths.

Is prompt injection the same as jailbreaking?

No, and the distinction is practical. Jailbreaking targets the model’s safety training — the attacker is the user, trying to get restricted output. Prompt injection targets the application’s trust boundary, and the attacker is usually a third party the user never interacted with.

Which class should a small team fix first?

Class 2, indirect injection through retrieved content, if the agent reads external documents or email. It is the highest-volume real-world class and produced the most consequential documented incident to date.

Does RAG make prompt injection worse?

It expands the surface. Every retrieved chunk is untrusted text entering context, and poisoned vector stores fall under OWASP’s LLM08 category. Retrieval is not the flaw, but it converts document access into instruction access.

How do I know if I have been hit?

Log every tool call with the provenance of the data that triggered it. Injections show up as actions that no user request explains. Without provenance logging, a successful prompt injection is close to invisible after the fact.

Keep reading

Agent Prompt Injection Testing

Agent Prompt Injection Testing: What a Two-Boolean Score Leaves Out

The plan was ordinary. Take a set of documented prompt-injection classes, run them against a pinned agent framework, and report what got through. Before running …

Read more

Model Card Disclosure

Model Card Disclosure in 2026: What AI Labs Actually Tell You

Open the documentation for any model released this year and you will find something. Whether you find the thing you came for depends entirely on …

Read more

AI compliance deadlines

AI Compliance Deadlines: What Applies, When, and to Whom

There is no single AI compliance deadline. There is a calendar of them, each attached to a particular jurisdiction, a particular kind of organisation, and …

Read more

AI agent framework security

AI Agent Framework Security: 10 Frameworks Audited

A developer runs pip install, decorates three functions with @tool, and points an agent at them. The agent now has a shell in your process. …

Read more

5 Hidden Layers of the AI Attack Surface Exposed

The AI Attack Surface: Securing LLM Systems End to End

OWASP released the 2026 edition of its Top 10 for LLM Applications on August 6. The AI attack surface is highlighted by this update for developers and security teams. The edition reframes the field and clarifies key risks. It guides risk-aware design for AI systems.

Additionally, avoid pursuing a model that cannot be fooled. It is unrealistic to expect perfect resilience. Instead, emphasize graceful degradation and fail-safe responses. Design checks and monitoring should detect anomalies early. Regular audits and red-teaming can strengthen defenses without promising invulnerability.

Moreover, design the system so that when it is fooled, no critical function fails. This approach helps maintain user trust and operational continuity. It pairs with robust incident response and clear recovery protocols. Staff training and documented procedures ensure quick, coordinated action.

That is a shift from prevention to blast-radius control, and it changes how you map the AI attack surface. You stop asking whether an attack can land. You start asking what it reaches when it does.

This page maps the AI attack surface in five layers: what fails at each, and where the deeper coverage sits.

Key Takeaways on the AI Attack Surface

  • The 2026 OWASP list keeps prompt injection at number one, and the AI attack surface still has no complete fix for it.
  • Excessive Agency climbed to third, since agentic deployments are where AI attack surface damage now lands.
  • OWASP drew a new boundary: once a model gains tools, memory, and consequences, it moves to a separate agentic list.
  • For the first time the ranking used incident data, with 6,639 real incidents carrying 25% of the weight.
  • Blast-radius control beats perfect prevention across the AI attack surface, and every layer reflects that.

Quick Navigation

Why the AI Attack Surface Needs a Layer Map

Security teams keep treating the AI attack surface as one problem. It is five, and the defenses differ at each.

A model weakness is not a prompt weakness. A tool weakness is not an agent weakness. Fixing the wrong layer yields the familiar outcome: real money spent, exposure unchanged.

The AI attack surface runs outward from the weights. First the model itself, then the prompt that reaches it, then the tools it can call, then the agent loop chaining those calls. Underneath all of it sits the supply chain that delivered the rest.

Each AI attack surface layer inherits the weaknesses of the one below. So a poisoned model makes every prompt defense unreliable, and a compromised tool makes agent-level approval theater.

Work up the AI attack surface when building, and down it when investigating.

Building means securing the supply chain before the agent, since you cannot reason about behavior you cannot trust. Investigating means starting at the observed harm and tracing back, because the visible failure is rarely the entry point.

Layer 1 of the AI Attack Surface: The Model

Start the AI attack surface at the weights, where least attention usually goes.

Data and model poisoning sits in the OWASP list. The 2026 edition widened it to cover fine-tuning subversion too. An attacker who shapes training data leaves behind behavior no runtime filter will catch. Backdoors trigger on specific phrases and stay dormant otherwise. Poisoned fine-tuning shifts refusal behavior subtly. Extraction attacks pull training data back out through careful querying.

None of these announce themselves. They are the quietest part of the AI attack surface, and the hardest to test for after deployment.

What reduces the risk

Provenance is the main AI attack surface control at this depth. Know which checkpoint you are running, where it came from, and what changed since.

Open weights cut both ways here. You can inspect them, and you also inherit whatever the publisher did. The licence and provenance questions around frontier open weights matter as much for security as for legal review.

Layer 2 of the AI Attack Surface: The Prompt

Prompt injection has topped every OWASP edition, and it still anchors the AI attack surface in 2026.

The root cause is design, not a bug. Models read instructions and data through one channel with no clean split. So anyone who controls an input can write orders the model treats as real. Prompt injection now covers cross-modal attacks, widening the AI attack surface. Instructions hidden inside images or audio reach the model the same way text does, which widens the AI attack surface considerably for multi-modal systems.

Direct injection comes from the user. Indirect injection rides in on fetched content: a web page, a document, an email, a code comment. The indirect kind is worse, because nobody typed it.

The defense effect

Here is an AI attack surface detail worth understanding. OWASP notes that recorded prompt injection incidents are relatively few, and attributes this to a defense effect rather than a low risk.

Teams spend heavily to block it, so successful attacks rarely reach public databases. Reading that low incident count as low danger inverts the actual picture. Scale matters too, because attackers retry. Anthropic’s published system card figures for one agentic coding setup show indirect injection landing 4.7% of the time at one try, 33.6% at ten, and 63.0% at a hundred. Attackers get to retry.

Layer 3 of the AI Attack Surface: The Tool

Once a model can call something, the AI attack surface stops being about text.

A tool turns a wrong answer into a wrong action. Sending an email, running a query, moving money, deploying code. The model does not need to be compromised for this to hurt, only misled. Over-broad scopes come first in this part of the AI attack surface. A tool granted database write access when it needs read access hands an attacker the difference.

Output handling comes second. OWASP moved it from fifth to tenth in 2026, though the category grew wider. The pattern holds: app code trusts model output and runs it unchecked. Then comes the wiring. MCP sets how models reach tools. The NSA published design guidance for it in May 2026, naming prompt injection and tool poisoning as open gaps.

Practical controls

Scope every credential to the narrowest task. Allowlist tools rather than blocking known-bad ones. Validate model output before execution, exactly as you would validate user input. And log every call. Most AI attack surface investigations fail because nobody recorded which tool ran with which arguments.

Layer 4 of the AI Attack Surface: The Agent

This is where AI attack surface damage now concentrates, and OWASP moved the category to match.

Excessive Agency climbed to third place in 2026, with expert voting and incident data agreeing for once. Agentic deployments are where harm is landing. The 2026 edition drew an explicit AI attack surface line. It covers the model as a part inside an app. Once the model becomes an actor, with tools it calls and memory it keeps, the risk moves to a separate agentic list.

That split helps when mapping the AI attack surface. Agent risk is not harder model risk. It is a different job, closer to identity work than to content filtering.

What breaks in the agent AI attack surface

Chained actions compound. An agent that reads a document, decides, and acts gives an attacker three points of influence rather than one. Memory persists. An injection that lands once can sit in stored context and fire on later sessions.

Autonomy removes the check. Human approval gates are the crudest control here and still the most effective, which is an uncomfortable thing to admit in 2026. Misinformation also climbed two places, and the reason is agentic. Model output now drives tool calls, writes code, and steers other agents. So a plausible wrong answer becomes a system failure, not a bad paragraph. The difference between agentic and generative systems is the difference between these two risk profiles.

Layer 5 of the AI Attack Surface: The Supply Chain

Every AI attack surface layer above assumes the components are what they claim to be.

This layer covers weights, datasets, embedding stores, framework dependencies, and now agent skills. OWASP started an Agentic Skills Top 10 in April 2026 because skill stores opened a new delivery route.

Old supply chain thinking applies to the AI attack surface, and it falls short. A poisoned npm package acts the same every time. A poisoned model acts normal until a trigger fires.

That difference breaks conventional scanning. You cannot diff weights the way you diff source code and learn much.

Minimum controls

Pin model versions with hashes, not tags. Record which dataset and checkpoint produced each deployed system. Treat community fine-tunes with the caution you would give an unsigned binary.

Vendor promises matter here too. State AI rules now impose real record-keeping duties on builders and deployers, and those duties run the whole AI attack surface.

Mapping the AI Attack Surface to Existing Frameworks

You do not need a new AI attack surface taxonomy. Three existing ones cover this ground, and they interlock.

The OWASP Top 10 for LLM Applications handles the model as a component. The Agentic Top 10 picks up where tools and memory begin. MITRE ATLAS v5.1.0 supplies adversary tactics and techniques, with 16 tactics and 84 techniques as of November 2025.

OWASP tells you what can go wrong. ATLAS tells you how an attacker would do it. The NIST AI Risk Management Framework tells you how to govern the result.

So run OWASP for design review, ATLAS for red-teaming, and NIST for board reporting. Using one where another fits is the most common mistake in AI attack surface programmes.

Where the frameworks still have gaps

Skills ecosystems reached the AI attack surface faster than the standards did. OWASP’s Agentic Skills Top 10 was still an incubator project as of April 2026, which means the newest distribution channel has the thinnest guidance.

Insurance lags too. Cover for AI incidents stays patchy, and a lot of silent exposure sits in policies written before any of this existed.

What Changed in the 2026 OWASP Rankings

The methodology change is the real AI attack surface story, more than any single move.

Every earlier edition rested on expert consensus. The 2026 list kept voting at 75% of the weight. The other 25% came from 6,639 real incidents in public vulnerability databases and an AI-harm database.

Misinformation is the clearest case. Voters ranked it near the bottom. The incident record ranked it near the top, and the data pushed it up two places.

That gap is worth sitting with. Experts play down risks that cause quiet, slow harm, and play up the ones that make good conference talks.

The full set of moves

Prompt Injection and Sensitive Information Disclosure held the top two AI attack surface slots. Excessive Agency rose to third. Unbounded Consumption climbed four places as cost-drain attacks got taken seriously. Output Handling fell from fifth to tenth. System Prompt Leakage became Hidden Context Exposure, with wider scope.

Two categories absorbed new scope rather than spawning entries. Prompt injection took on cross-modal attacks. Data and Model Poisoning took on fine-tuning subversion.

The AI Attack Surface Pattern That Predicts Exploitability

AI attack surface

One formulation explains more real AI attack surface incidents than the whole ranking does.

Simon Willison’s lethal trifecta describes three properties that, combined, make a system exploitable: access to private data, exposure to untrusted content, and the ability to communicate externally.

Why the trifecta works as a test

Any two are usually survivable. All three together mean an attacker can inject instructions, reach your data, and get it out.

So audit the AI attack surface by asking which of your systems hold all three. That single question finds more genuine exposure than a checklist pass, and it takes an afternoon.

Removing any one leg breaks the chain. Cut external communication, restrict which content the system ingests, or partition the private data. Any one of the three works.

How to Shrink the AI Attack Surface Blast Radius

The 2026 framing points at design rather than detection, so build the AI attack surface for containment.

Assume injection succeeds. Design the AI attack surface so a fooled model reaches nothing critical. This is the central move.

Scope credentials narrowly. Every permission an agent holds is a permission an attacker inherits.

Gate irreversible actions. Human approval before anything that moves money, deletes data, or ships code.

Log everything. Tool calls, arguments, retrieved content. Without these, incident response has nothing to work from.

Red-team continuously. Attack success rises sharply with attempts, so a single passing test proves very little.

AI Attack Surface Incidents Worth Knowing

Theory moves slowly. Incidents move the AI attack surface, and three are worth carrying as reference points.

Slack AI data exfiltration. PromptArmor researchers showed indirect prompt injection pulling data out of a live assistant. That proved the fetched-content path was real, not theoretical.

EchoLeak. Recorded as the first real-world zero-click prompt injection exploit in a live system. Zero-click matters, because no user has to be tricked at all.

Agent deception in national testing. UK cyber exercises in 2026 reported agent deception moving from theory into observed behavior.

A ranking convinces a security team. An incident convinces a budget holder.

So keep two or three concrete cases at hand when arguing for AI attack surface work. The abstract version of this argument has been losing for three years.

Wiring Your Own AI Attack Surface Review

Two hours gets you a first pass. Run it in this order.

Start by listing every system where a model reads content you do not control. That is your indirect injection exposure, and it is usually longer than expected.

Next, for each one, note whether it holds private data and whether it can send anything outward. Systems with all three legs go to the top of the AI attack surface queue. Then check credentials. Pull the actual scope on every token an agent holds, not the scope somebody intended.

Finally, confirm you have logs. If a tool call is not recorded with its arguments, you cannot investigate it later, and that gap is the most common finding in an AI attack surface review.

How to Use This AI Attack Surface Hub

Three AI attack surface entry points, depending on why you are here.

Responding to an incident. Start at the observed harm and work down the layers. The visible failure is rarely the entry point.

Designing a new system. Work up from the supply chain. Each layer depends on the one below being trustworthy.

Briefing leadership. The five layers map onto budget lines. An OWASP ranking does not.

This AI attack surface hub updates as coverage grows. Each new security post links back here, and the layer sections point to the pieces worth reading first.

Conclusion: The AI Attack Surface Is a Systems Problem

The most useful AI attack surface shift in the 2026 list is one of expectation.

Earlier guidance implied that enough filtering could make a model safe to trust. The new framing accepts that models get fooled, then asks what happens next.

That re-framing helps, because it moves the work somewhere solvable. Nobody knows how to make a model immune to prompt injection. Plenty of teams know how to scope a credential, gate an action, and log a tool call.

So read the AI attack surface as five layers with different owners, different controls, and different failure modes. Then go make the blast radius smaller.

FAQ About the AI Attack Surface

What are the layers of the AI attack surface?

The AI attack surface has five: the model and its weights, the prompt channel that reaches it, the tools it can call, the agent loop that chains those calls, and the supply chain delivering all four. Each layer inherits weaknesses from the one below, so defenses have to be assessed together rather than individually.

What is the biggest LLM security risk in 2026?

Prompt injection tops the AI attack surface in the OWASP Top 10 for LLM Applications 2026 edition, followed by Sensitive Information Disclosure. Excessive Agency rose to third, reflecting that agentic deployments are where measurable damage is now occurring.

Why did OWASP add incident data to its rankings?

To balance practitioner belief about the AI attack surface against evidence. The 2026 edition weighted expert voting at 75% and incident data at 25%, drawing on 6,639 real incidents from public vulnerability databases and an AI-harm database. Misinformation moved up two places because the incident record ranked it far higher than voters did.

Can prompt injection be fixed?

Not completely, because it stems from architecture rather than implementation. Models process instructions and data through one channel with no reliable separation. Current best practice is defense in depth: least-privilege tooling, input and output filtering, human approval for high-risk actions, and continuous adversarial testing.

What is the lethal trifecta?

A formulation from Simon Willison identifying three properties that together make a system exploitable: access to private data, exposure to untrusted content, and the ability to communicate externally. Removing any one of the three breaks the attack chain, which makes it a fast practical audit.

How is agent security different from model security?

Model security concerns what the system outputs. Agent security concerns what it does, which brings in tool permissions, persistent memory, chained actions, and downstream consequences. OWASP formalized this split in 2026 by moving agentic risk to a separate Top 10 list.

Keep reading

Agent Prompt Injection Testing

Agent Prompt Injection Testing: What a Two-Boolean Score Leaves Out

The plan was ordinary. Take a set of documented prompt-injection classes, run them against a pinned agent framework, and report what got through. Before running …

Read more

Model Card Disclosure

Model Card Disclosure in 2026: What AI Labs Actually Tell You

Open the documentation for any model released this year and you will find something. Whether you find the thing you came for depends entirely on …

Read more

AI compliance deadlines

AI Compliance Deadlines: What Applies, When, and to Whom

There is no single AI compliance deadline. There is a calendar of them, each attached to a particular jurisdiction, a particular kind of organisation, and …

Read more

AI agent framework security

AI Agent Framework Security: 10 Frameworks Audited

A developer runs pip install, decorates three functions with @tool, and points an agent at them. The agent now has a shell in your process. …

Read more