dotsagent.io
Language:English
Secure · reference

Agent security

Agents read text written by strangers and then act with your credentials. This page sets out the core rule, both 2026 OWASP lists, the incidents so far and the defences that would have stopped them.

The lethal trifecta

  1. Access to private dataEmail, repositories, databases, files, or anything else that sits behind your credentials.
  2. Exposure to untrusted contentWeb pages, issues, support tickets, inbound email and tool descriptions: any text an attacker can write.
  3. Ability to communicate externallySending mail, opening pull requests, fetching URLs or rendering remote images. Simon Willison named this three-part combination in June 2025.

If one agent holds all three, assume a prompt injection can make it hand your data to an attacker. Don't count on the model refusing; remove at least one leg from every agent and session.

Recent incidents

OWASP Top 10 for Agentic Applications 2026

The OWASP GenAI Security Project published this list on December 9, 2025. It covers the risks that appear once a model plans, keeps memory, calls tools and works alongside other agents.

  1. ASI01
    Agent Goal Hijack

    An attacker changes what the agent is trying to achieve, usually through instructions hidden in content it reads. The agent then pursues the attacker's goal with the user's tools and permissions.

  2. ASI02
    Tool Misuse and Exploitation

    The agent uses legitimate tools in harmful ways, such as deleting records, sending messages or chaining calls, because it was manipulated or holds tools broader than the task requires.

  3. ASI03
    Identity and Privilege Abuse

    Agents act through credentials, delegated tokens and inherited permissions. Attackers abuse those identities, or the gaps between them, to escalate privileges or act as someone else.

  4. ASI04
    Agentic Supply Chain Vulnerabilities

    Tools, MCP servers, skills, plugins, models and prompts loaded at build time or run time can be malicious or compromised. Because agents load many of them dynamically, one bad component reaches every session that uses it.

  5. ASI05
    Unexpected Code Execution (RCE)

    Agents that write and run code, or hand model output to shells and interpreters, can be steered into running commands the attacker chose on the host.

  6. ASI06
    Memory & Context Poisoning

    Attackers plant false facts or instructions in an agent's memory, retrieved documents or saved context. The poison persists and shapes later sessions long after the original input is gone.

  7. ASI07
    Insecure Inter-Agent Communication

    Messages between agents travel without proper authentication, integrity checks or validation, so they can be spoofed, replayed or altered to mislead the agent that receives them.

  8. ASI08
    Cascading Failures

    One fault, such as a poisoned input, a bad tool result or a compromised agent, spreads through connected agents and automated steps faster than people can catch it.

  9. ASI09
    Human-Agent Trust Exploitation

    Agents sound confident and helpful, so people tend to approve what they propose. Attackers use that trust to get a human to confirm a harmful action or reveal information.

  10. ASI10
    Rogue Agents

    An agent that has been compromised or has drifted from its intended behaviour keeps acting on its own, outside the scope and oversight it was given.

OWASP Top 10 for LLM Applications 2026

This edition was published on August 4, 2026 and replaces the 2025 list. Excessive Agency rose from sixth to third place, and System Prompt Leakage was renamed Hidden Context Exposure.

  1. LLM01
    Prompt Injection

    Input alters the model's behaviour in ways the developer did not intend. It can come straight from the user or indirectly from documents, web pages and tool results the model reads.

  2. LLM02
    Sensitive Information Disclosure

    The model or application reveals personal data, credentials, business secrets or other confidential material in its output, drawn from training data, context or connected systems.

  3. LLM03
    Excessive Agency

    The application gives the model more functions, permissions or autonomy than the task needs, so a manipulated or mistaken output causes real damage. It moved from sixth place in 2025 to third in 2026.

  4. LLM04
    Supply Chain

    Third-party models, datasets, adapters, packages and plugins can be tampered with or vulnerable, and they carry that risk into your application.

  5. LLM05
    Data and Model Poisoning

    Attackers manipulate pre-training, fine-tuning or embedding data to plant backdoors, biases or faulty behaviour that only surfaces later in production.

  6. LLM06
    Unbounded Consumption

    Without limits on requests, input size or compute, attackers can run up your bill, exhaust resources or copy a model through high-volume queries.

  7. LLM07
    Misinformation

    The model produces false or misleading output that looks credible, and users or downstream systems act on it without checking.

  8. LLM08
    Hidden Context Exposure

    Formerly System Prompt Leakage. System prompts, hidden instructions and other context the user is not meant to see can be extracted, exposing the rules, logic or secrets placed there.

  9. LLM09
    Vector and Embedding Weaknesses

    Flaws in how embeddings are generated, stored and retrieved let attackers inject content, leak data across tenants or recover source text, which hits RAG systems hardest.

  10. LLM10
    Improper Output Handling

    Model output reaches browsers, shells, databases or other components without validation or encoding, which opens the way to XSS, SQL injection, code execution and data exfiltration.

Defences

Fifteen documented defences, each linked to its source. Layer several of them, because none stops every attack on its own.

Break the lethal trifecta

Simon Willison's rule: an agent that can read private data, sees untrusted content and can send data out can be turned against you by any text it reads. Remove at least one of the three from each agent or session. For example, the agent that triages public issues gets no secrets, and the one that holds secrets gets no outbound channel.

simonwillison.net

Constrain the agent after untrusted input

Research on design patterns for agent security sets one rule: once an agent has ingested untrusted input, that input must not be able to trigger consequential actions. The patterns include action-selector, plan-then-execute, dual LLM and context minimisation. Pick one per workflow, for instance fix the plan before the agent reads any untrusted data.

arxiv.org

Track data flow with CaMeL

CaMeL splits the agent in two: a privileged planner writes code from the user's request, and a quarantined model handles untrusted data. Values from the quarantined side carry capability tags, and policies check those tags before any tool runs. In the paper it solved 77% of AgentDojo tasks with provable security, against 84% for an undefended agent.

arxiv.org

Use least-privilege credentials

Give agents read-only, project-scoped access by default and never an admin or service_role key, which in Supabase bypasses row-level security. Scope CI and repository tokens to the one job they do. The Amazon Q Developer incident traced back to an over-scoped GitHub token in CodeBuild.

supabase.comaws.amazon.com

Require approval for consequential actions

Make a person confirm tool calls that write, delete, send or spend, and fail closed when nobody answers. Supabase recommends manual approval of MCP tool calls, and in OpenClaw you set tools.exec.ask to always with askFallback left at deny. Show the full arguments so the reviewer sees what will actually run.

supabase.comdocs.openclaw.ai

Sandbox code and tool execution

Run shell commands and generated code in a container or VM that holds no credentials and sees only the workspace. OpenClaw ships with sandboxing off and tools.exec.security at full on gateway hosts, so switch sandboxing on, set exec security to deny or allowlist, set fs.workspaceOnly to true and keep elevated mode disabled. Confirm the result with openclaw sandbox explain.

docs.openclaw.aidocs.openclaw.ai

Restrict outbound network access

Deny outbound traffic from agents and MCP servers by default, then allow only the hosts each one needs. Never auto-approve fetches to multi-tenant hosts where anyone can publish, which is how CVE-2026-54316 turned huggingface.co into an exfiltration channel. An email server should reach its mail API and nothing else, as postmark-mcp showed.

nvd.nist.govkoi.ai

Keep control planes off the internet

Bind agent gateways, dashboards and debug proxies to loopback and require a token of at least 24 characters, for example from openssl rand -hex 32. Reach them remotely through an SSH tunnel or Tailscale Serve, and use Tailscale Funnel only with password auth. Run openclaw security audit --deep on a schedule.

docs.openclaw.aidocs.openclaw.aidocs.openclaw.ai

Validate MCP token audience

The MCP authorisation spec requires a server to reject access tokens that were not issued for it, and forbids passing a client's token through to an upstream API. Clients send RFC 8707 resource indicators so each token is bound to one server. A proxy server needs consent from each client, or it becomes a confused deputy.

modelcontextprotocol.iomodelcontextprotocol.io

Isolate MCP sessions and tenants

Create a separate server and transport instance for each session instead of sharing one across clients. Bind every session and task to the authenticated principal that created it, and check that binding on each request. Both 2026 MCP SDK advisories came from shared or unbound state.

github.comgithub.com

Vet skills, plugins and MCP servers

Pin exact versions, read the diff before each update and check scanner verdicts such as VirusTotal and the ClawHub security audit status. Hash tool descriptions when you approve a server and alert when they change, which catches rug pulls. OpenClaw does no built-in blocking at install time, so set security.installPolicy yourself.

openclaw.aidocs.openclaw.aiinvariantlabs.aidocs.openclaw.ai

Control who can message the agent

Keep DM access on pairing or an allowlist, require a mention before the agent acts in group chats, and set session.dmScope to per-channel-peer so senders never share context. Anyone who can message the agent can try to instruct it, so the sender list is part of your attack surface.

docs.openclaw.aidocs.openclaw.ai

Keep secrets out of untrusted CI runs

Don't run an agent with repository secrets on workflows that outsiders can trigger through a pull request, issue or comment. Treat titles, descriptions and comments from those events as hostile. If a step really needs secrets, run it only after a maintainer has approved the run.

oddguan.com

Treat model output as untrusted

Encode or sanitise model output before rendering it, and never pass it unchecked to a shell, SQL query or browser. Block automatic loading of markdown images and links to external domains, and set a strict Content Security Policy. EchoLeak moved data out through URLs that loaded on their own.

genai.owasp.orgnvd.nist.gov

Verify agent signatures, then authorise

To identify an agent calling your site or API, verify its Web Bot Auth signature against the keys the operator publishes in /.well-known/http-message-signatures-directory. ChatGPT agent signs as Signature-Agent https://chatgpt.com. A valid signature tells you who operates the agent, not which user sent it or what that user may do, so authorise each request separately.

datatracker.ietf.orghelp.openai.com

Questions about agent security

What is prompt injection in AI agents?

Prompt injection is text the model treats as instructions even though it arrived as data, for example in a web page, an email or a tool description. In an agent, those instructions can trigger tool calls that run with your credentials. OWASP lists it as LLM01:2026, and Agent Goal Hijack (ASI01) covers the agentic form.

What is the lethal trifecta for AI agents?

It is Simon Willison's name for an agent that combines access to private data, exposure to untrusted content and a way to communicate externally. With all three in place, a prompt injection can read your data and send it out. The 2025 Supabase MCP token leak is a textbook case.

How do I secure an MCP server?

Reject access tokens that were not issued for your server, and never pass a client's token through to an upstream API. Create one server and transport instance per session, and bind sessions and tasks to the authenticated user. Run the TypeScript SDK at 1.26.0 or later and the Python SDK at 1.27.2 or later, which fix a cross-client response leak and session hijacking.

Can a better system prompt stop prompt injection?

Don't rely on it. Research on design patterns for agents argues that once an agent has ingested untrusted input, it must be constrained so that the input cannot trigger consequential actions. CaMeL, which enforces this with capability tracking, solved 77% of AgentDojo tasks with provable security, against 84% for an undefended agent.

What is the difference between the OWASP LLM Top 10 and the Agentic Top 10?

The OWASP Top 10 for LLM Applications, whose 2026 edition came out on 4 August 2026, covers risks in any application built on a language model. The OWASP Top 10 for Agentic Applications, released on 9 December 2025, covers what changes when the model plans, uses tools, keeps memory and talks to other agents. Most agent builders need both.

Is it safe to expose an OpenClaw gateway to the internet?

No. Leave gateway.bind at its default of loopback, and use an SSH tunnel or Tailscale Serve when you need remote access. OpenA2A counted 192,492 exposed gateways on 1 September 2026, and CVE-2026-25253 showed that even loopback-only installs need prompt patching.

Sources

  1. genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/
  2. genai.owasp.org/download/52117
  3. genai.owasp.org/resource/owasp-genai-llm-top-10-2026/
  4. github.com/GenAI-Security-Project/GenAI-LLM-Top10

Independent reference for people who build AI agents. Not affiliated with any vendor named here.

© 2026 DotsAgent · Facts checked October 1, 2026