./resources / blog

OpenClaw and the Agent Attack Surface

In under a month, a personal AI assistant went from a weekend project to more than a hundred thousand internet-exposed instances — most of them holding shell access, API keys, and a mailbox. This is how attackers take them over, drawn out at the protocol level, and what a defensible deployment actually looks like.

OpenClaw is the clearest example yet of a category we would call the personal agent runtime: software that connects a large language model to your messaging apps on one side and your operating system on the other, then lets it act autonomously in between. Send it a WhatsApp message, and it can read your email, clone a repository, run a shell command, drive a browser, and reply — unattended.

It is genuinely useful, which is why it reached 100,000 GitHub stars faster than any comparable open-source project. It is also, from a security architecture perspective, close to a worst case: a remote-code-execution service, deliberately built, wired to untrusted input channels, holding long-lived credentials for everything it touches, and installed by people who are not systems administrators.

This article is not a takedown of one project. OpenClaw simply got popular first, so it got scrutinised first. Every finding below generalises to the wider class — self-hosted agent gateways, MCP servers with tool access, autonomous coding agents, and the "just point it at your inbox" assistants arriving weekly. If you understand how this one falls over, you can assess the next one in an afternoon.

What the architecture actually looks like

Before you can reason about the attack surface, you need an accurate mental model. Most people picture a chatbot. What is really running is a broker sitting between untrusted input and privileged execution.

Figure 1 — the broker in the middle LLM Provider API sees the entire context window OpenClaw Gateway :18789/tcp — Control UI auth · session authority tool orchestration config store + API keys device pairing registry Untrusted ingress WhatsApp / Telegram Slack / Discord Email / IMAP Web pages / documents Node hosts — execution Shell — system.run Filesystem read/write Headless browser Skills / plugins Everything on the left is attacker-reachable input. Everything on the right is code execution. The gateway is the only thing standing between them.
Figure 1. The gateway brokers untrusted input into privileged execution. Traditional appsec asks "can an attacker reach the execution path?" Here, reaching it is the product's advertised feature — the only question is who gets to drive.

Three properties make this architecture unusually unforgiving:

  • The security boundary is a single process. There is no privilege separation between "read a Telegram message" and "run a shell command". One authentication check gates both.
  • Credentials are aggregated by design. To be useful, the agent holds tokens for email, GitHub, cloud consoles and calendars simultaneously. Compromising the gateway compromises all of them at once.
  • The operator is usually not an administrator. These tools are installed by developers, analysts and enthusiasts on laptops — outside change management, outside the asset inventory, outside the patch cycle.

The exposure problem, in numbers

The default gateway port, 18789/tcp, is identical across every deployment — which makes internet-wide enumeration trivial. Multiple research teams scanned for it, and the counts vary sharply by methodology and scan window: Bitsight identified more than 30,000 distinct instances between 27 January and 8 February 2026; SecurityScorecard reported 40,214 in February; other trackers have since put the figure above 200,000.

Treat any single number with suspicion — they measure different things over different windows. The proportions are the alarming part. In SecurityScorecard's dataset, roughly 63% of observed deployments were vulnerable, 12,812 were exploitable for remote code execution, and 93.4% of publicly reachable instances carried a critical authentication bypass. Exposure clustered not only in technology but in healthcare, finance, government and insurance.

Three default-configuration decisions drive most of this:

  • Bind mode. The gateway offers loopback, LAN (0.0.0.0) and VPN binding. Users who want to reach their assistant from a phone pick LAN — the documentation's warning to "use only if you trust the network" does not convey that this means the public internet behind most home routers with UPnP.
  • Trivial credentials are accepted. The gateway requires a password or token for non-local connections but enforces no complexity — researchers confirmed a single character such as a is a valid credential.
  • An explicit safety-off switch. gateway.controlUi.allowInsecureAuth disables device identity verification. It exists to work around "localhost only" friction, and forum answers recommend it freely.

Where the trust model breaks

Exposure alone would be bad enough. What turns it into reliable compromise is a chain of authority decisions inside the gateway that each looked reasonable in isolation. Figure 2 traces a request from browser to shell and marks where each disclosed vulnerability lands.

Figure 2 — request path and where authority is lost Control UI (browser) reads gatewayUrl from query string WebSocket connect → :18789 stored token sent in connect payload Authentication password / token · loopback exempt from limits Session authority operator.admin bound to the connection RPC dispatch — node.invoke approval fields trusted from the client Node host — system.run arbitrary commands as the host user Disclosed failure points CVE-2026-25253 — 1-click RCE token exfiltrated to attacker gateway · CVSS 8.8 ClawJacked — loopback brute force no rate limit from localhost · auto device pairing CVE-2026-53821 — authority bypass operator.admin obtained on a live WebSocket CVE-2026-28466 — authz bypass client-controlled approval fields reach system.run
Figure 2. Four independent disclosures, four consecutive stages of the same pipeline. Each was patched quickly; the pattern — authority granted early and never re-verified — is the durable lesson.

CVE-2026-25253 — one click, full compromise

The highest-impact issue needs no exposed port at all. The Control UI accepts a gatewayUrl parameter from the query string, trusts it without validation, and auto-connects on page load — sending the stored gateway token in the connect payload. An attacker who gets a victim to click a crafted link receives that token at a server they control, then replays it against the victim's real gateway.

Figure 3 — CVE-2026-25253 token exfiltration chain Attacker Victim browser Real gateway 1 — crafted link ?gatewayUrl=wss://evil 2 — UI auto-connects no origin validation 3 — token sent in connect payload 4 — token captured by attacker 5 — replay system.run → RCE
Figure 3. Fixed in v2026.1.29. Note that steps 2 and 3 require no interaction beyond opening a page — the victim's own browser performs the exfiltration.

ClawJacked — the localhost exemption

The ClawJacked chain is the most instructive, because every individual decision was defensible. Browsers do not apply cross-origin restrictions to WebSocket connections to localhost, so any website you visit can open a socket to your gateway. The gateway's rate limiter exempts loopback connections — reasonable, since loopback was assumed trusted. And device pairings from loopback are auto-approved without a user prompt, for the same reason.

Chained: a malicious web page opens a WebSocket to your local gateway, brute-forces the password at hundreds of guesses per second with no lockout, silently registers itself as a trusted device, and then reads your configuration, enumerates connected nodes, and drives the agent. No plugin, no extension, no interaction beyond visiting a page. The lesson generalises past this product: "it's only bound to localhost" is not an access control when the victim runs a browser.

The authorisation pair

CVE-2026-28466 (fixed in 2026.2.14) let an authenticated attacker execute arbitrary commands on connected node hosts because approval fields inside node.invoke requests were taken from the client and forwarded to nodes unsanitised — the client could simply assert that a command had already been approved. CVE-2026-53821 (fixed in 2026.5.18) let restricted Control UI clients prematurely obtain operator.admin authority on a live WebSocket connection and issue administrative RPCs.

Both are the same bug in different clothing: authority is decided once, early, and carried forward as a property of the connection rather than re-checked against the specific action being requested.

Prompt injection: the vulnerability that has no patch

Everything above is fixed. This is not.

An agent that reads content is an agent that can be instructed by that content. Because the LLM receives retrieved data and operator instructions in the same context window, with no cryptographic or structural distinction between them, text that looks like an instruction is an instruction. Researchers extracted OpenClaw's system prompt with an 84.6% success rate by simply asking the agent to convert its rules to JSON.

Figure 4 — indirect prompt injection Attacker plants text page · email · ticket · doc Agent ingests it during a normal task Context window data = instructions Tool call fires with full authority Exfiltrate API keys Pivot to connected apps Destroy or encrypt data The model cannot separate "content to process" from "instructions to obey". This is architectural. No model update closes it — only constraining what the tools can do.
Figure 4. Phishing for agents. The victim never sees the payload, because the payload was never addressed to the human.

The practical form is mundane. An attacker messages the bot directly on a connected channel. Or files a GitHub issue containing hidden instructions, knowing a coding agent will read it. Or emails an invoice with white-on-white text. When the agent processes that content in the course of an ordinary task, it executes the attacker's instructions with the operator's full authority — and, in most deployments, silently.

Because this cannot be fixed at the model layer, it must be contained at the capability layer. That means the question stops being "can we prevent injection?" and becomes "what is the worst thing a successful injection can accomplish?" That reframing is the entire basis of a defensible design.

Supply chain: the plugin registry

Agent runtimes are extended through community plugins — "skills" — installed with a single command. Security teams have found disguised malware, credential stealers and remote-access payloads in official registries, with harmful packages reappearing under new names after removal. The install path typically pulls arbitrary code that runs in-process with the gateway's full privileges and its whole credential store. Every lesson from npm and PyPI typosquatting applies, minus a decade of ecosystem tooling.

Designing a deployment that can survive this

Banning the technology outright is a defensible policy for regulated environments, and some universities have already prohibited it on managed devices. For most organisations it is also unrealistic — the productivity gains are real, and prohibition drives usage onto unmanaged personal machines where you have no visibility at all. The workable position is a sanctioned, hardened pattern.

Figure 5 — hardened reference deployment Operator access VPN / Tailscale only 64-char token + MFA Isolated agent host bind 127.0.0.1:18789 container · non-root · seccomp read-only FS · scoped mounts human approval on system.run signed / pinned plugins only Controlled egress egress allowlist proxy short-lived secrets broker LLM provider API no lateral network reach read-only integration scopes Monitoring asset inventory · gateway auth + node.invoke logs → SIEM · egress anomaly alerts · release watch
Figure 5. The design goal is not preventing agent compromise — assume it happens. It is ensuring a compromised agent reaches nothing that matters and cannot act unobserved.

Configuration baseline

Two settings account for the overwhelming majority of real-world exposure:

# hardened gateway baseline
gateway:
  bind: 127.0.0.1              # never 0.0.0.0 — reach it over VPN instead
  port: 18789
  controlUi:
    allowInsecureAuth: false   # never true, under any circumstances
  auth:
    token: "<64 chars from a CSPRNG>"   # not "a", not a passphrase

Finding what is already running

Most organisations we work with discover their first agent runtime during an assessment, not through the asset register. Start with the fixed port — but pair it with endpoint telemetry, because the laptop deployments are the ones that will not answer a network scan:

# network sweep of your own estate
nmap -p 18789 --open -sV 10.0.0.0/8

# endpoint inventory — process and listener check
# (Windows) Get-NetTCPConnection -LocalPort 18789
# (Linux)   ss -lntp | grep 18789

Then close the loop with egress telemetry: an internal host establishing sustained outbound sessions to an LLM provider API is a reliable indicator of an agent runtime, whether or not you ever find the listener.

The controls that matter, in order

  • Never bind to 0.0.0.0. Loopback plus VPN. This single control eliminates the internet-exposure population entirely.
  • Patch aggressively and watch releases. Four of the issues above were fixed within days of disclosure — one within 24 hours. The exposure window is created by operators who never update, not by slow maintainers.
  • Constrain the blast radius before you constrain the model. Read-only scopes, per-integration credentials, short-lived tokens from a broker rather than API keys in a config file. Assume injection succeeds; make it worthless.
  • Require human approval for irreversible actions. Shell execution, sending mail, deleting files, moving money. Approval fatigue is real, so scope the gate narrowly enough that it stays meaningful.
  • Segment the host. Container, non-root, no lateral reach to production, databases or corporate mail. Treat the agent host as a DMZ asset, because functionally that is what it is.
  • Pin and review plugins. No install-on-demand from public registries into a process holding your credential store.
  • Log to the SIEM. Authentication attempts, device pairings, node.invoke calls and egress destinations. Without these, a compromise is simply invisible.

How Preemptive helps

Agent runtimes fall between the cracks of a conventional testing programme. They are not quite applications, not quite infrastructure, and they are rarely in the asset register — so a standard scope misses them entirely. Our work in this area covers four things:

  • Agent discovery and exposure assessment. We enumerate agent runtimes across your internal estate and external perimeter — fixed-port sweeps, endpoint telemetry, egress analysis and attack-surface monitoring — and tell you what is reachable, from where, and holding which credentials.
  • Configuration review against a hardened baseline. Bind addresses, credential strength, insecure-auth flags, plugin provenance, integration scopes and version currency against the disclosed CVE set.
  • Adversarial testing of the agent itself. Indirect prompt injection through every ingested channel, tool-abuse chains, authority-escalation attempts against the gateway, and credential-exfiltration paths — testing the deployment as an attacker would, not as a scanner would.
  • Detection engineering and policy. SIEM content for gateway authentication, device pairing and tool invocation; egress baselines; and a sanctioned-use policy that gives your teams a supported path rather than driving the tools underground.

If your organisation has adopted AI agents — or suspects it has, informally — the first step is knowing what is deployed and what it can reach. That is usually a short, high-yield engagement, and it almost always finds something.

Key takeaways

  • A personal agent runtime is a remote code execution service wired to untrusted input — assess it as infrastructure, not as a chatbot.
  • A fixed default port and a LAN bind option produced tens of thousands of internet-exposed instances, most with critical authentication weaknesses.
  • The disclosed CVEs share one root cause: authority granted early and never re-verified against the action being requested.
  • "Bound to localhost" is not an access control — browsers reach loopback WebSockets from any page you visit.
  • Prompt injection has no patch. Contain it at the capability layer: least privilege, short-lived credentials, approval gates on irreversible actions.
  • Patch velocity is not the problem — operator patch adoption is. Watch releases and update.
#ai-security #openclaw #prompt-injection #agent-security #misconfiguration
All articles Assess your agent deployment