Agent security
Agents read text written by strangers and then act with your credentials. This page sets out the core rule, both 2026 OWASP lists, the incidents so far and the defences that would have stopped them.
The lethal trifecta
- Access to private dataEmail, repositories, databases, files, or anything else that sits behind your credentials.
- Exposure to untrusted contentWeb pages, issues, support tickets, inbound email and tool descriptions: any text an attacker can write.
- Ability to communicate externallySending mail, opening pull requests, fetching URLs or rendering remote images. Simon Willison named this three-part combination in June 2025.
If one agent holds all three, assume a prompt injection can make it hand your data to an attacker. Don't count on the model refusing; remove at least one leg from every agent and session.
Security checklist
22 checks across inputs, tools, runtime, MCP, CI and operations. Tick them off as you go and export the result as a Markdown file.
Incident log
For each public agent security incident since April 2025: what happened, why the attack worked and what to change in your own setup.
Recent incidents
- MCP Python SDK session hijack and cross-session task access
- Claude Code WebFetch auto-approval let data out via huggingface.co
- Comment and Control: PR text steals secrets from CI agents
- MCP TypeScript SDK leaked responses between clients
- ClawHavoc: hundreds of malicious ClawHub skills spread AMOS
OWASP Top 10 for Agentic Applications 2026
The OWASP GenAI Security Project published this list on December 9, 2025. It covers the risks that appear once a model plans, keeps memory, calls tools and works alongside other agents.
ASI01Agent Goal HijackAn attacker changes what the agent is trying to achieve, usually through instructions hidden in content it reads. The agent then pursues the attacker's goal with the user's tools and permissions.
ASI02Tool Misuse and ExploitationThe agent uses legitimate tools in harmful ways, such as deleting records, sending messages or chaining calls, because it was manipulated or holds tools broader than the task requires.
ASI03Identity and Privilege AbuseAgents act through credentials, delegated tokens and inherited permissions. Attackers abuse those identities, or the gaps between them, to escalate privileges or act as someone else.
ASI04Agentic Supply Chain VulnerabilitiesTools, MCP servers, skills, plugins, models and prompts loaded at build time or run time can be malicious or compromised. Because agents load many of them dynamically, one bad component reaches every session that uses it.
ASI05Unexpected Code Execution (RCE)Agents that write and run code, or hand model output to shells and interpreters, can be steered into running commands the attacker chose on the host.
ASI06Memory & Context PoisoningAttackers plant false facts or instructions in an agent's memory, retrieved documents or saved context. The poison persists and shapes later sessions long after the original input is gone.
ASI07Insecure Inter-Agent CommunicationMessages between agents travel without proper authentication, integrity checks or validation, so they can be spoofed, replayed or altered to mislead the agent that receives them.
ASI08Cascading FailuresOne fault, such as a poisoned input, a bad tool result or a compromised agent, spreads through connected agents and automated steps faster than people can catch it.
ASI09Human-Agent Trust ExploitationAgents sound confident and helpful, so people tend to approve what they propose. Attackers use that trust to get a human to confirm a harmful action or reveal information.
ASI10Rogue AgentsAn agent that has been compromised or has drifted from its intended behaviour keeps acting on its own, outside the scope and oversight it was given.
OWASP Top 10 for LLM Applications 2026
This edition was published on August 4, 2026 and replaces the 2025 list. Excessive Agency rose from sixth to third place, and System Prompt Leakage was renamed Hidden Context Exposure.
LLM01Prompt InjectionInput alters the model's behaviour in ways the developer did not intend. It can come straight from the user or indirectly from documents, web pages and tool results the model reads.
LLM02Sensitive Information DisclosureThe model or application reveals personal data, credentials, business secrets or other confidential material in its output, drawn from training data, context or connected systems.
LLM03Excessive AgencyThe application gives the model more functions, permissions or autonomy than the task needs, so a manipulated or mistaken output causes real damage. It moved from sixth place in 2025 to third in 2026.
LLM04Supply ChainThird-party models, datasets, adapters, packages and plugins can be tampered with or vulnerable, and they carry that risk into your application.
LLM05Data and Model PoisoningAttackers manipulate pre-training, fine-tuning or embedding data to plant backdoors, biases or faulty behaviour that only surfaces later in production.
LLM06Unbounded ConsumptionWithout limits on requests, input size or compute, attackers can run up your bill, exhaust resources or copy a model through high-volume queries.
LLM07MisinformationThe model produces false or misleading output that looks credible, and users or downstream systems act on it without checking.
LLM08Hidden Context ExposureFormerly System Prompt Leakage. System prompts, hidden instructions and other context the user is not meant to see can be extracted, exposing the rules, logic or secrets placed there.
LLM09Vector and Embedding WeaknessesFlaws in how embeddings are generated, stored and retrieved let attackers inject content, leak data across tenants or recover source text, which hits RAG systems hardest.
LLM10Improper Output HandlingModel output reaches browsers, shells, databases or other components without validation or encoding, which opens the way to XSS, SQL injection, code execution and data exfiltration.
Defences
Fifteen documented defences, each linked to its source. Layer several of them, because none stops every attack on its own.
Break the lethal trifecta
Simon Willison's rule: an agent that can read private data, sees untrusted content and can send data out can be turned against you by any text it reads. Remove at least one of the three from each agent or session. For example, the agent that triages public issues gets no secrets, and the one that holds secrets gets no outbound channel.
Constrain the agent after untrusted input
Research on design patterns for agent security sets one rule: once an agent has ingested untrusted input, that input must not be able to trigger consequential actions. The patterns include action-selector, plan-then-execute, dual LLM and context minimisation. Pick one per workflow, for instance fix the plan before the agent reads any untrusted data.
Track data flow with CaMeL
CaMeL splits the agent in two: a privileged planner writes code from the user's request, and a quarantined model handles untrusted data. Values from the quarantined side carry capability tags, and policies check those tags before any tool runs. In the paper it solved 77% of AgentDojo tasks with provable security, against 84% for an undefended agent.
Use least-privilege credentials
Give agents read-only, project-scoped access by default and never an admin or service_role key, which in Supabase bypasses row-level security. Scope CI and repository tokens to the one job they do. The Amazon Q Developer incident traced back to an over-scoped GitHub token in CodeBuild.
Require approval for consequential actions
Make a person confirm tool calls that write, delete, send or spend, and fail closed when nobody answers. Supabase recommends manual approval of MCP tool calls, and in OpenClaw you set tools.exec.ask to always with askFallback left at deny. Show the full arguments so the reviewer sees what will actually run.
Sandbox code and tool execution
Run shell commands and generated code in a container or VM that holds no credentials and sees only the workspace. OpenClaw ships with sandboxing off and tools.exec.security at full on gateway hosts, so switch sandboxing on, set exec security to deny or allowlist, set fs.workspaceOnly to true and keep elevated mode disabled. Confirm the result with openclaw sandbox explain.
Restrict outbound network access
Deny outbound traffic from agents and MCP servers by default, then allow only the hosts each one needs. Never auto-approve fetches to multi-tenant hosts where anyone can publish, which is how CVE-2026-54316 turned huggingface.co into an exfiltration channel. An email server should reach its mail API and nothing else, as postmark-mcp showed.
Keep control planes off the internet
Bind agent gateways, dashboards and debug proxies to loopback and require a token of at least 24 characters, for example from openssl rand -hex 32. Reach them remotely through an SSH tunnel or Tailscale Serve, and use Tailscale Funnel only with password auth. Run openclaw security audit --deep on a schedule.
Validate MCP token audience
The MCP authorisation spec requires a server to reject access tokens that were not issued for it, and forbids passing a client's token through to an upstream API. Clients send RFC 8707 resource indicators so each token is bound to one server. A proxy server needs consent from each client, or it becomes a confused deputy.
Isolate MCP sessions and tenants
Create a separate server and transport instance for each session instead of sharing one across clients. Bind every session and task to the authenticated principal that created it, and check that binding on each request. Both 2026 MCP SDK advisories came from shared or unbound state.
Vet skills, plugins and MCP servers
Pin exact versions, read the diff before each update and check scanner verdicts such as VirusTotal and the ClawHub security audit status. Hash tool descriptions when you approve a server and alert when they change, which catches rug pulls. OpenClaw does no built-in blocking at install time, so set security.installPolicy yourself.
Control who can message the agent
Keep DM access on pairing or an allowlist, require a mention before the agent acts in group chats, and set session.dmScope to per-channel-peer so senders never share context. Anyone who can message the agent can try to instruct it, so the sender list is part of your attack surface.
Keep secrets out of untrusted CI runs
Don't run an agent with repository secrets on workflows that outsiders can trigger through a pull request, issue or comment. Treat titles, descriptions and comments from those events as hostile. If a step really needs secrets, run it only after a maintainer has approved the run.
Treat model output as untrusted
Encode or sanitise model output before rendering it, and never pass it unchecked to a shell, SQL query or browser. Block automatic loading of markdown images and links to external domains, and set a strict Content Security Policy. EchoLeak moved data out through URLs that loaded on their own.
Verify agent signatures, then authorise
To identify an agent calling your site or API, verify its Web Bot Auth signature against the keys the operator publishes in /.well-known/http-message-signatures-directory. ChatGPT agent signs as Signature-Agent https://chatgpt.com. A valid signature tells you who operates the agent, not which user sent it or what that user may do, so authorise each request separately.
Questions about agent security
What is prompt injection in AI agents?
Prompt injection is text the model treats as instructions even though it arrived as data, for example in a web page, an email or a tool description. In an agent, those instructions can trigger tool calls that run with your credentials. OWASP lists it as LLM01:2026, and Agent Goal Hijack (ASI01) covers the agentic form.
What is the lethal trifecta for AI agents?
It is Simon Willison's name for an agent that combines access to private data, exposure to untrusted content and a way to communicate externally. With all three in place, a prompt injection can read your data and send it out. The 2025 Supabase MCP token leak is a textbook case.
How do I secure an MCP server?
Reject access tokens that were not issued for your server, and never pass a client's token through to an upstream API. Create one server and transport instance per session, and bind sessions and tasks to the authenticated user. Run the TypeScript SDK at 1.26.0 or later and the Python SDK at 1.27.2 or later, which fix a cross-client response leak and session hijacking.
Can a better system prompt stop prompt injection?
Don't rely on it. Research on design patterns for agents argues that once an agent has ingested untrusted input, it must be constrained so that the input cannot trigger consequential actions. CaMeL, which enforces this with capability tracking, solved 77% of AgentDojo tasks with provable security, against 84% for an undefended agent.
What is the difference between the OWASP LLM Top 10 and the Agentic Top 10?
The OWASP Top 10 for LLM Applications, whose 2026 edition came out on 4 August 2026, covers risks in any application built on a language model. The OWASP Top 10 for Agentic Applications, released on 9 December 2025, covers what changes when the model plans, uses tools, keeps memory and talks to other agents. Most agent builders need both.
Is it safe to expose an OpenClaw gateway to the internet?
No. Leave gateway.bind at its default of loopback, and use an SSH tunnel or Tailscale Serve when you need remote access. OpenA2A counted 192,492 exposed gateways on 1 September 2026, and CVE-2026-25253 showed that even loopback-only installs need prompt patching.