./resources / blog

OpenClaw 1.0 → 2.0: What Changes, and What to Watch for Tomorrow

OpenClaw 1.0 spent its first months as a case study in how autonomous agents get owned — a single-process broker wiring untrusted input straight to shell access. 2.0 is a genuine re-architecture that closes the root cause of that CVE chain. It also moves the attack surface somewhere new. This is what actually changes, what doesn't, and what security teams should be preparing for now.

We pulled OpenClaw 1.0 apart at the protocol level in OpenClaw and the Agent Attack Surface. The short version: a personal agent runtime that brokers messaging apps into operating-system execution, shipped with a fixed gateway port, a LAN-bind footgun, and a chain of disclosed vulnerabilities — one-click RCE, a loopback brute-force, and two authorisation bypasses — all sharing a single root cause: authority granted early and never re-verified against the action being requested. Prompt injection sat underneath all of it, unpatchable at the model layer.

2.0 is the maintainers' answer to that scrutiny. It is a bigger change than a version bump usually implies, and much of it is genuinely good security engineering. But a redesign that fixes yesterday's bugs does not automatically fix tomorrow's — it changes where you have to look. A note on framing: 2.0 is still stabilising as it ships, so treat the specifics below as the announced direction. The durable analysis is not any single feature — it is how the attack surface shifts.

The core change: from one process to separated planes

1.0's defining property was that the security boundary was a single process. Reading a Telegram message and running a shell command lived behind one authentication check, in one address space, holding every credential at once. 2.0 breaks that apart into distinct planes with authority re-checked at each hop.

1.0 monolith vs 2.0 separated planes OpenClaw 1.0 Single-process broker Untrusted ingress Auth (once, at connect) Tool orchestration Credential store (all) system.run — execution one boundary gates everything 2.0 OpenClaw 2.0 Control plane — identity, sessions (often hosted) Policy & guardrail engine — per-action authz Capability broker — short-lived, scoped tokens Sandboxed tool executors — isolated per tool Signed / pinned plugin registry Secure defaults — loopback, strong tokens authority re-checked at every hop Preemptive Cyber Security
Figure 1. 2.0's headline change is privilege separation: the policy engine re-authorises each action, and executors run sandboxed with short-lived capability tokens instead of a shared credential vault.

What genuinely improves

Give credit where it is due — several of these changes retire whole vulnerability classes from the 1.0 era.

  • Per-action authorisation. The policy engine re-checks authority against the specific action, not the connection. That directly kills the "decide once, carry forward" root cause behind CVE-2026-53821 and CVE-2026-28466.
  • Capability tokens instead of a shared vault. Tools receive short-lived, narrowly scoped grants from a broker rather than reaching into one process holding every credential. A compromised executor no longer means every integration falls at once.
  • Sandboxed executors by default. Shell, browser and filesystem run isolated rather than in-process, so a single tool compromise is contained.
  • Secure defaults. Loopback binding and strong generated tokens out of the box — the two configuration decisions that produced the tens of thousands of internet-exposed 1.0 instances.
  • A signed plugin path. Provenance and pinning on the skills registry raise the bar on the supply-chain problem that plagued 1.0's install-on-demand model.
2.0 is what you get when a project takes its incident history seriously. The single-process broker — the thing that made 1.0 uniquely unforgiving — is gone.

What does not change

Architecture cannot fix a problem that lives in the model. Prompt injection is still unpatched, because the LLM still receives retrieved data and operator instructions in one undifferentiated context window. 2.0's policy engine and sandboxing change the consequences of a successful injection — a very good thing — but they do not stop the model from being persuaded by content it reads. The defensive question remains the one from our 1.0 write-up: not "can we prevent injection?" but "what is the worst a successful injection can accomplish?"

Two other constants: the operator is still usually not an administrator, and the runtime is still rarely in the asset register. Better defaults help, but shadow deployment is a governance problem, not an engineering one.

Where the attack surface moves

This is the part defenders should internalise. 2.0 does not remove attack surface so much as relocate it — from a single exposed port on a laptop to a distributed system with new trust relationships. Some of these are strictly better; all of them need fresh eyes.

The surface moves — it does not vanish 1.0 attack surface • Internet-exposed :18789 • Trivial / single-char token • "Localhost is trusted" • Single-process RCE • Shared credential vault • Unsigned plugin registry • Prompt injection mostly patched / defaulted away 2.0 attack surface • Hosted control plane — tenant isolation, account takeover • Capability-token theft & replay • Agent-to-agent (A2A) trust & delegation abuse • Policy-engine bypass & rule mis-config • More autonomy, less human-in-the-loop • Expanded tool / MCP ecosystem • Long-term memory poisoning prompt injection persists — now with more reach Preemptive Cyber Security
Figure 2. The 1.0 problems were largely about a laptop on the internet. The 2.0 problems are about distributed trust: hosted control planes, delegated capability tokens, and agents talking to agents.

The five things to watch in 2.0

  • The hosted control plane. If identity and orchestration move to a vendor cloud, you inherit a SaaS threat model — tenant isolation, account takeover, and a supply-chain dependency on the provider. A control-plane compromise is now a fleet-wide event, not a single laptop.
  • Capability tokens. Short-lived scoped tokens are a big improvement over a shared vault — if they cannot be stolen and replayed. Watch token lifetime, audience binding, and whether an injected agent can mint or broaden its own grants.
  • Agent-to-agent (A2A) trust. 2.0-class runtimes lean into multi-agent delegation. Every delegation edge is a new authority-transfer point, and "authority granted early and never re-verified" is exactly the bug that will resurface between agents if the delegation model is loose.
  • Policy-engine correctness. A guardrail engine is only as good as its rules. Mis-configuration, over-broad allow rules, and bypasses become the new high-value target — the single point that decides every action.
  • More autonomy, fewer prompts. As agents get more capable, the temptation is to remove the human approval gate for "routine" actions. That is precisely the mitigation that makes prompt injection survivable — do not trade it away for convenience.

1.0 vs 2.0 at a glance

DimensionOpenClaw 1.0OpenClaw 2.0Net for defenders
BoundarySingle processSeparated planesBetter — contained blast radius
AuthorisationOnce, per connectionPer action (policy engine)Better — kills the 1.0 root cause
CredentialsShared vaultShort-lived capability tokensBetter — but token hygiene is now critical
ExecutionIn-processSandboxed executorsBetter — isolation by default
DeploymentSelf-hosted laptopHosted control plane optionNew — SaaS/tenant threat model
Multi-agentMinimalA2A delegationNew — delegation abuse surface
Prompt injectionUnpatchedUnpatched (better contained)Unchanged — still architectural

What to do now — before you roll 2.0 out

The migration from 1.0 to 2.0 is a golden window: you are touching the deployment anyway, so bake the controls in rather than bolting them on.

2.0 readiness pipeline 01Inventoryfind agents 02Re-baseline2.0 config 03Token hygienescope · TTL 04Test policyengine rules 05A2A trustdelegation 06Monitorplane + logs Preemptive Cyber Security
Figure 3. Treat the 1.0→2.0 migration as a security project, not an upgrade. Every stage is a control you can bake in while you are already touching the deployment.

Defender checklist for tomorrow

  • Re-inventory. The fixed 1.0 port is fading as a signal — lean on endpoint telemetry and egress-to-LLM-provider patterns to find 2.0 deployments, including the hosted ones.
  • Treat the control plane as production. If it is hosted, apply your SaaS due diligence: tenant isolation, SSO, admin MFA, and a data-processing review.
  • Audit capability tokens. Short TTLs, tight scopes, audience binding, and no self-minting — verify an injected agent cannot broaden its own grants.
  • Test the policy engine like the target it is. Fuzz the rules, hunt allow-rule over-breadth, and confirm every irreversible action re-authorises.
  • Map agent-to-agent trust. Enumerate delegation edges and make sure authority is re-verified on transfer, not inherited.
  • Keep the human gate. Approval on shell, mail, deletion and money stays — autonomy is not a reason to remove the one control that makes injection survivable.
  • Extend the SIEM. Control-plane auth, capability-token issuance, delegation events and executor egress — not just the old gateway logs.

How Preemptive helps

We assess agent runtimes as what they are — infrastructure that executes code on untrusted input. For teams moving to a 2.0-class deployment, that means a migration-time review of the new planes: control-plane and tenant-isolation testing, capability-token and policy-engine assessment, agent-to-agent trust modelling, indirect prompt-injection testing through every ingested channel, and detection engineering for the delegation and control-plane events your current SIEM content will miss. If you adopted 1.0 informally and are about to adopt 2.0 the same way, the highest-yield first step is still the simplest: find what you have, and prove what it can reach.

Key takeaways

  • 2.0 fixes the 1.0 root cause — the single-process broker and connection-level authority are gone, replaced by separated planes and per-action authorisation.
  • It relocates the attack surface rather than removing it: hosted control planes, capability tokens, and agent-to-agent trust are the new frontier.
  • Prompt injection still has no patch — 2.0 contains its impact better, but the containment controls (least privilege, approval gates) must stay in place.
  • Watch capability-token hygiene, policy-engine correctness, and A2A delegation — that is where "authority granted early and never re-verified" will try to resurface.
  • Use the migration as a security project: inventory, re-baseline, and extend monitoring to the control plane before you scale 2.0 out.
#openclaw #ai-security #agent-security #prompt-injection #capability-tokens #a2a
Read the 1.0 breakdown Assess your agent deployment