---
title: "AI agent security: OWASP top 10s, incidents, defences · DotsAgent"
description: "Agent security reference: both 2026 OWASP top 10 lists, 17 real incidents from MCP to OpenClaw, 15 defences and a checklist for teams shipping AI agents."
url: https://dotsagent.io/security
---

Secure · reference

# Agent security

Agents read text written by strangers and then act with your credentials. This page sets out the core rule, both 2026 OWASP lists, the incidents so far and the defences that would have stopped them.

## The lethal trifecta

1. **Access to private data** Email, repositories, databases, files, or anything else that sits behind your credentials.
2. **Exposure to untrusted content** Web pages, issues, support tickets, inbound email and tool descriptions: any text an attacker can write.
3. **Ability to communicate externally** Sending mail, opening pull requests, fetching URLs or rendering remote images. Simon Willison named this three-part combination in June 2025.

If one agent holds all three, assume a prompt injection can make it hand your data to an attacker. Don't count on the model refusing; remove at least one leg from every agent and session.

- [Security checklist](https://dotsagent.io/security/checklist): 22 checks across inputs, tools, runtime, MCP, CI and operations. Tick them off as you go and export the result as a Markdown file.

- [Incident log](https://dotsagent.io/security/incidents): For each public agent security incident since April 2025: what happened, why the attack worked and what to change in your own setup.

## Recent incidents

- Jun 5, 2026 [MCP Python SDK session hijack and cross-session task access](https://dotsagent.io/security/incidents/2026-06-mcp-python-sdk-session-hijack)
- Jun 2026 [Claude Code WebFetch auto-approval let data out via huggingface.co](https://dotsagent.io/security/incidents/2026-06-claude-code-webfetch-exfiltration)
- Apr 15, 2026 [Comment and Control: PR text steals secrets from CI agents](https://dotsagent.io/security/incidents/2026-04-comment-and-control-ci-secrets)
- Feb 4, 2026 [MCP TypeScript SDK leaked responses between clients](https://dotsagent.io/security/incidents/2026-02-mcp-typescript-sdk-response-leak)
- Feb 1, 2026 [ClawHavoc: hundreds of malicious ClawHub skills spread AMOS](https://dotsagent.io/security/incidents/2026-02-clawhavoc-malicious-skills)

## OWASP Top 10 for Agentic Applications 2026

The OWASP GenAI Security Project published this list on December 9, 2025. It covers the risks that appear once a model plans, keeps memory, calls tools and works alongside other agents.

1. `ASI01` **Agent Goal Hijack** An attacker changes what the agent is trying to achieve, usually through instructions hidden in content it reads. The agent then pursues the attacker's goal with the user's tools and permissions.
2. `ASI02` **Tool Misuse and Exploitation** The agent uses legitimate tools in harmful ways, such as deleting records, sending messages or chaining calls, because it was manipulated or holds tools broader than the task requires.
3. `ASI03` **Identity and Privilege Abuse** Agents act through credentials, delegated tokens and inherited permissions. Attackers abuse those identities, or the gaps between them, to escalate privileges or act as someone else.
4. `ASI04` **Agentic Supply Chain Vulnerabilities** Tools, MCP servers, skills, plugins, models and prompts loaded at build time or run time can be malicious or compromised. Because agents load many of them dynamically, one bad component reaches every session that uses it.
5. `ASI05` **Unexpected Code Execution (RCE)** Agents that write and run code, or hand model output to shells and interpreters, can be steered into running commands the attacker chose on the host.
6. `ASI06` **Memory & Context Poisoning** Attackers plant false facts or instructions in an agent's memory, retrieved documents or saved context. The poison persists and shapes later sessions long after the original input is gone.
7. `ASI07` **Insecure Inter-Agent Communication** Messages between agents travel without proper authentication, integrity checks or validation, so they can be spoofed, replayed or altered to mislead the agent that receives them.
8. `ASI08` **Cascading Failures** One fault, such as a poisoned input, a bad tool result or a compromised agent, spreads through connected agents and automated steps faster than people can catch it.
9. `ASI09` **Human-Agent Trust Exploitation** Agents sound confident and helpful, so people tend to approve what they propose. Attackers use that trust to get a human to confirm a harmful action or reveal information.
10. `ASI10` **Rogue Agents** An agent that has been compromised or has drifted from its intended behaviour keeps acting on its own, outside the scope and oversight it was given.

## OWASP Top 10 for LLM Applications 2026

This edition was published on August 4, 2026 and replaces the 2025 list. Excessive Agency rose from sixth to third place, and System Prompt Leakage was renamed Hidden Context Exposure.

1. `LLM01` **Prompt Injection** Input alters the model's behaviour in ways the developer did not intend. It can come straight from the user or indirectly from documents, web pages and tool results the model reads.
2. `LLM02` **Sensitive Information Disclosure** The model or application reveals personal data, credentials, business secrets or other confidential material in its output, drawn from training data, context or connected systems.
3. `LLM03` **Excessive Agency** The application gives the model more functions, permissions or autonomy than the task needs, so a manipulated or mistaken output causes real damage. It moved from sixth place in 2025 to third in 2026.
4. `LLM04` **Supply Chain** Third-party models, datasets, adapters, packages and plugins can be tampered with or vulnerable, and they carry that risk into your application.
5. `LLM05` **Data and Model Poisoning** Attackers manipulate pre-training, fine-tuning or embedding data to plant backdoors, biases or faulty behaviour that only surfaces later in production.
6. `LLM06` **Unbounded Consumption** Without limits on requests, input size or compute, attackers can run up your bill, exhaust resources or copy a model through high-volume queries.
7. `LLM07` **Misinformation** The model produces false or misleading output that looks credible, and users or downstream systems act on it without checking.
8. `LLM08` **Hidden Context Exposure** Formerly System Prompt Leakage. System prompts, hidden instructions and other context the user is not meant to see can be extracted, exposing the rules, logic or secrets placed there.
9. `LLM09` **Vector and Embedding Weaknesses** Flaws in how embeddings are generated, stored and retrieved let attackers inject content, leak data across tenants or recover source text, which hits RAG systems hardest.
10. `LLM10` **Improper Output Handling** Model output reaches browsers, shells, databases or other components without validation or encoding, which opens the way to XSS, SQL injection, code execution and data exfiltration.

## Defences

Fifteen documented defences, each linked to its source. Layer several of them, because none stops every attack on its own.

### Break the lethal trifecta

Simon Willison's rule: an agent that can read private data, sees untrusted content and can send data out can be turned against you by any text it reads. Remove at least one of the three from each agent or session. For example, the agent that triages public issues gets no secrets, and the one that holds secrets gets no outbound channel.

[simonwillison.net](https://simonwillison.net/2025/jun/16/the-lethal-trifecta/)

### Constrain the agent after untrusted input

Research on design patterns for agent security sets one rule: once an agent has ingested untrusted input, that input must not be able to trigger consequential actions. The patterns include action-selector, plan-then-execute, dual LLM and context minimisation. Pick one per workflow, for instance fix the plan before the agent reads any untrusted data.

[arxiv.org](https://arxiv.org/abs/2506.08837)

### Track data flow with CaMeL

CaMeL splits the agent in two: a privileged planner writes code from the user's request, and a quarantined model handles untrusted data. Values from the quarantined side carry capability tags, and policies check those tags before any tool runs. In the paper it solved 77% of AgentDojo tasks with provable security, against 84% for an undefended agent.

[arxiv.org](https://arxiv.org/abs/2503.18813)

### Use least-privilege credentials

Give agents read-only, project-scoped access by default and never an admin or service_role key, which in Supabase bypasses row-level security. Scope CI and repository tokens to the one job they do. The Amazon Q Developer incident traced back to an over-scoped GitHub token in CodeBuild.

[supabase.com](https://supabase.com/blog/defense-in-depth-mcp)[aws.amazon.com](https://aws.amazon.com/security/security-bulletins/AWS-2025-015/)

### Require approval for consequential actions

Make a person confirm tool calls that write, delete, send or spend, and fail closed when nobody answers. Supabase recommends manual approval of MCP tool calls, and in OpenClaw you set tools.exec.ask to always with askFallback left at deny. Show the full arguments so the reviewer sees what will actually run.

[supabase.com](https://supabase.com/blog/defense-in-depth-mcp)[docs.openclaw.ai](https://docs.openclaw.ai/tools/exec-approvals)

### Sandbox code and tool execution

Run shell commands and generated code in a container or VM that holds no credentials and sees only the workspace. OpenClaw ships with sandboxing off and tools.exec.security at full on gateway hosts, so switch sandboxing on, set exec security to deny or allowlist, set fs.workspaceOnly to true and keep elevated mode disabled. Confirm the result with openclaw sandbox explain.

[docs.openclaw.ai](https://docs.openclaw.ai/gateway/sandboxing)[docs.openclaw.ai](https://docs.openclaw.ai/gateway/security/hardened-baseline)

### Restrict outbound network access

Deny outbound traffic from agents and MCP servers by default, then allow only the hosts each one needs. Never auto-approve fetches to multi-tenant hosts where anyone can publish, which is how CVE-2026-54316 turned huggingface.co into an exfiltration channel. An email server should reach its mail API and nothing else, as postmark-mcp showed.

[nvd.nist.gov](https://nvd.nist.gov/vuln/detail/CVE-2026-54316)[koi.ai](https://www.koi.ai/blog/postmark-mcp-npm-malicious-backdoor-email-theft)

### Keep control planes off the internet

Bind agent gateways, dashboards and debug proxies to loopback and require a token of at least 24 characters, for example from openssl rand -hex 32. Reach them remotely through an SSH tunnel or Tailscale Serve, and use Tailscale Funnel only with password auth. Run openclaw security audit --deep on a schedule.

[docs.openclaw.ai](https://docs.openclaw.ai/gateway/security/network-exposure)[docs.openclaw.ai](https://docs.openclaw.ai/gateway/tailscale)[docs.openclaw.ai](https://docs.openclaw.ai/gateway/security/running-the-audit)

### Validate MCP token audience

The MCP authorisation spec requires a server to reject access tokens that were not issued for it, and forbids passing a client's token through to an upstream API. Clients send RFC 8707 resource indicators so each token is bound to one server. A proxy server needs consent from each client, or it becomes a confused deputy.

[modelcontextprotocol.io](https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization/security-considerations)[modelcontextprotocol.io](https://modelcontextprotocol.io/docs/tutorials/security/security_best_practices)

### Isolate MCP sessions and tenants

Create a separate server and transport instance for each session instead of sharing one across clients. Bind every session and task to the authenticated principal that created it, and check that binding on each request. Both 2026 MCP SDK advisories came from shared or unbound state.

[github.com](https://github.com/advisories/GHSA-345p-7cg4-v4c7)[github.com](https://github.com/advisories/GHSA-JPW9-PFVF-9F58)

### Vet skills, plugins and MCP servers

Pin exact versions, read the diff before each update and check scanner verdicts such as VirusTotal and the ClawHub security audit status. Hash tool descriptions when you approve a server and alert when they change, which catches rug pulls. OpenClaw does no built-in blocking at install time, so set security.installPolicy yourself.

[openclaw.ai](https://openclaw.ai/blog/virustotal-partnership)[docs.openclaw.ai](https://docs.openclaw.ai/clawhub/security-audits)[invariantlabs.ai](https://invariantlabs.ai/blog/mcp-security-notification-tool-poisoning-attacks)[docs.openclaw.ai](https://docs.openclaw.ai/help/faq/security-and-access-control)

### Control who can message the agent

Keep DM access on pairing or an allowlist, require a mention before the agent acts in group chats, and set session.dmScope to per-channel-peer so senders never share context. Anyone who can message the agent can try to instruct it, so the sender list is part of your attack surface.

[docs.openclaw.ai](https://docs.openclaw.ai/gateway/security/access-control)[docs.openclaw.ai](https://docs.openclaw.ai/gateway/security/hardened-baseline)

### Keep secrets out of untrusted CI runs

Don't run an agent with repository secrets on workflows that outsiders can trigger through a pull request, issue or comment. Treat titles, descriptions and comments from those events as hostile. If a step really needs secrets, run it only after a maintainer has approved the run.

[oddguan.com](https://oddguan.com/blog/comment-and-control-prompt-injection-credential-theft-claude-code-gemini-cli-github-copilot/)

### Treat model output as untrusted

Encode or sanitise model output before rendering it, and never pass it unchecked to a shell, SQL query or browser. Block automatic loading of markdown images and links to external domains, and set a strict Content Security Policy. EchoLeak moved data out through URLs that loaded on their own.

[genai.owasp.org](https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/)[nvd.nist.gov](https://nvd.nist.gov/vuln/detail/cve-2025-32711)

### Verify agent signatures, then authorise

To identify an agent calling your site or API, verify its Web Bot Auth signature against the keys the operator publishes in /.well-known/http-message-signatures-directory. ChatGPT agent signs as Signature-Agent https://chatgpt.com. A valid signature tells you who operates the agent, not which user sent it or what that user may do, so authorise each request separately.

[datatracker.ietf.org](https://datatracker.ietf.org/doc/draft-ietf-webbotauth-httpsig-protocol/)[help.openai.com](https://help.openai.com/en/articles/11845367-chatgpt-agent-allowlisting)

## Questions about agent security

### What is prompt injection in AI agents?

Prompt injection is text the model treats as instructions even though it arrived as data, for example in a web page, an email or a tool description. In an agent, those instructions can trigger tool calls that run with your credentials. OWASP lists it as LLM01:2026, and Agent Goal Hijack (ASI01) covers the agentic form.

### What is the lethal trifecta for AI agents?

It is Simon Willison's name for an agent that combines access to private data, exposure to untrusted content and a way to communicate externally. With all three in place, a prompt injection can read your data and send it out. The 2025 Supabase MCP token leak is a textbook case.

### How do I secure an MCP server?

Reject access tokens that were not issued for your server, and never pass a client's token through to an upstream API. Create one server and transport instance per session, and bind sessions and tasks to the authenticated user. Run the TypeScript SDK at 1.26.0 or later and the Python SDK at 1.27.2 or later, which fix a cross-client response leak and session hijacking.

### Can a better system prompt stop prompt injection?

Don't rely on it. Research on design patterns for agents argues that once an agent has ingested untrusted input, it must be constrained so that the input cannot trigger consequential actions. CaMeL, which enforces this with capability tracking, solved 77% of AgentDojo tasks with provable security, against 84% for an undefended agent.

### What is the difference between the OWASP LLM Top 10 and the Agentic Top 10?

The OWASP Top 10 for LLM Applications, whose 2026 edition came out on 4 August 2026, covers risks in any application built on a language model. The OWASP Top 10 for Agentic Applications, released on 9 December 2025, covers what changes when the model plans, uses tools, keeps memory and talks to other agents. Most agent builders need both.

### Is it safe to expose an OpenClaw gateway to the internet?

No. Leave gateway.bind at its default of loopback, and use an SSH tunnel or Tailscale Serve when you need remote access. OpenA2A counted 192,492 exposed gateways on 1 September 2026, and CVE-2026-25253 showed that even loopback-only installs need prompt patching.

## Sources

1. [genai.owasp.org](https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/)/resource/owasp-top-10-for-agentic-applications-for-2026/
2. [genai.owasp.org](https://genai.owasp.org/download/52117)/download/52117
3. [genai.owasp.org](https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/)/resource/owasp-genai-llm-top-10-2026/
4. [github.com](https://github.com/GenAI-Security-Project/GenAI-LLM-Top10)/GenAI-Security-Project/GenAI-LLM-Top10

Independent reference for people who build AI agents. Not affiliated with any vendor named here.

© 2026 DotsAgent · Facts checked October 1, 2026
