OpenClaw 1.0 → 2.0: What Changes, and What to Watch for Tomorrow
OpenClaw 1.0 spent its first months as a case study in how autonomous agents get owned — a single-process broker wiring untrusted input straight to shell access. 2.0 is a genuine re-architecture that closes the root cause of that CVE chain. It also moves the attack surface somewhere new. This is what actually changes, what doesn't, and what security teams should be preparing for now.
We pulled OpenClaw 1.0 apart at the protocol level in OpenClaw and the Agent Attack Surface. The short version: a personal agent runtime that brokers messaging apps into operating-system execution, shipped with a fixed gateway port, a LAN-bind footgun, and a chain of disclosed vulnerabilities — one-click RCE, a loopback brute-force, and two authorisation bypasses — all sharing a single root cause: authority granted early and never re-verified against the action being requested. Prompt injection sat underneath all of it, unpatchable at the model layer.
2.0 is the maintainers' answer to that scrutiny. It is a bigger change than a version bump usually implies, and much of it is genuinely good security engineering. But a redesign that fixes yesterday's bugs does not automatically fix tomorrow's — it changes where you have to look. A note on framing: 2.0 is still stabilising as it ships, so treat the specifics below as the announced direction. The durable analysis is not any single feature — it is how the attack surface shifts.
The core change: from one process to separated planes
1.0's defining property was that the security boundary was a single process. Reading a Telegram message and running a shell command lived behind one authentication check, in one address space, holding every credential at once. 2.0 breaks that apart into distinct planes with authority re-checked at each hop.
What genuinely improves
Give credit where it is due — several of these changes retire whole vulnerability classes from the 1.0 era.
- Per-action authorisation. The policy engine re-checks authority against the specific action, not the connection. That directly kills the "decide once, carry forward" root cause behind CVE-2026-53821 and CVE-2026-28466.
- Capability tokens instead of a shared vault. Tools receive short-lived, narrowly scoped grants from a broker rather than reaching into one process holding every credential. A compromised executor no longer means every integration falls at once.
- Sandboxed executors by default. Shell, browser and filesystem run isolated rather than in-process, so a single tool compromise is contained.
- Secure defaults. Loopback binding and strong generated tokens out of the box — the two configuration decisions that produced the tens of thousands of internet-exposed 1.0 instances.
- A signed plugin path. Provenance and pinning on the skills registry raise the bar on the supply-chain problem that plagued 1.0's install-on-demand model.
2.0 is what you get when a project takes its incident history seriously. The single-process broker — the thing that made 1.0 uniquely unforgiving — is gone.
What does not change
Architecture cannot fix a problem that lives in the model. Prompt injection is still unpatched, because the LLM still receives retrieved data and operator instructions in one undifferentiated context window. 2.0's policy engine and sandboxing change the consequences of a successful injection — a very good thing — but they do not stop the model from being persuaded by content it reads. The defensive question remains the one from our 1.0 write-up: not "can we prevent injection?" but "what is the worst a successful injection can accomplish?"
Two other constants: the operator is still usually not an administrator, and the runtime is still rarely in the asset register. Better defaults help, but shadow deployment is a governance problem, not an engineering one.
Where the attack surface moves
This is the part defenders should internalise. 2.0 does not remove attack surface so much as relocate it — from a single exposed port on a laptop to a distributed system with new trust relationships. Some of these are strictly better; all of them need fresh eyes.
The five things to watch in 2.0
- The hosted control plane. If identity and orchestration move to a vendor cloud, you inherit a SaaS threat model — tenant isolation, account takeover, and a supply-chain dependency on the provider. A control-plane compromise is now a fleet-wide event, not a single laptop.
- Capability tokens. Short-lived scoped tokens are a big improvement over a shared vault — if they cannot be stolen and replayed. Watch token lifetime, audience binding, and whether an injected agent can mint or broaden its own grants.
- Agent-to-agent (A2A) trust. 2.0-class runtimes lean into multi-agent delegation. Every delegation edge is a new authority-transfer point, and "authority granted early and never re-verified" is exactly the bug that will resurface between agents if the delegation model is loose.
- Policy-engine correctness. A guardrail engine is only as good as its rules. Mis-configuration, over-broad allow rules, and bypasses become the new high-value target — the single point that decides every action.
- More autonomy, fewer prompts. As agents get more capable, the temptation is to remove the human approval gate for "routine" actions. That is precisely the mitigation that makes prompt injection survivable — do not trade it away for convenience.
1.0 vs 2.0 at a glance
| Dimension | OpenClaw 1.0 | OpenClaw 2.0 | Net for defenders |
|---|---|---|---|
| Boundary | Single process | Separated planes | Better — contained blast radius |
| Authorisation | Once, per connection | Per action (policy engine) | Better — kills the 1.0 root cause |
| Credentials | Shared vault | Short-lived capability tokens | Better — but token hygiene is now critical |
| Execution | In-process | Sandboxed executors | Better — isolation by default |
| Deployment | Self-hosted laptop | Hosted control plane option | New — SaaS/tenant threat model |
| Multi-agent | Minimal | A2A delegation | New — delegation abuse surface |
| Prompt injection | Unpatched | Unpatched (better contained) | Unchanged — still architectural |
What to do now — before you roll 2.0 out
The migration from 1.0 to 2.0 is a golden window: you are touching the deployment anyway, so bake the controls in rather than bolting them on.
Defender checklist for tomorrow
- Re-inventory. The fixed 1.0 port is fading as a signal — lean on endpoint telemetry and egress-to-LLM-provider patterns to find 2.0 deployments, including the hosted ones.
- Treat the control plane as production. If it is hosted, apply your SaaS due diligence: tenant isolation, SSO, admin MFA, and a data-processing review.
- Audit capability tokens. Short TTLs, tight scopes, audience binding, and no self-minting — verify an injected agent cannot broaden its own grants.
- Test the policy engine like the target it is. Fuzz the rules, hunt allow-rule over-breadth, and confirm every irreversible action re-authorises.
- Map agent-to-agent trust. Enumerate delegation edges and make sure authority is re-verified on transfer, not inherited.
- Keep the human gate. Approval on shell, mail, deletion and money stays — autonomy is not a reason to remove the one control that makes injection survivable.
- Extend the SIEM. Control-plane auth, capability-token issuance, delegation events and executor egress — not just the old gateway logs.
How Preemptive helps
We assess agent runtimes as what they are — infrastructure that executes code on untrusted input. For teams moving to a 2.0-class deployment, that means a migration-time review of the new planes: control-plane and tenant-isolation testing, capability-token and policy-engine assessment, agent-to-agent trust modelling, indirect prompt-injection testing through every ingested channel, and detection engineering for the delegation and control-plane events your current SIEM content will miss. If you adopted 1.0 informally and are about to adopt 2.0 the same way, the highest-yield first step is still the simplest: find what you have, and prove what it can reach.
Key takeaways
- 2.0 fixes the 1.0 root cause — the single-process broker and connection-level authority are gone, replaced by separated planes and per-action authorisation.
- It relocates the attack surface rather than removing it: hosted control planes, capability tokens, and agent-to-agent trust are the new frontier.
- Prompt injection still has no patch — 2.0 contains its impact better, but the containment controls (least privilege, approval gates) must stay in place.
- Watch capability-token hygiene, policy-engine correctness, and A2A delegation — that is where "authority granted early and never re-verified" will try to resurface.
- Use the migration as a security project: inventory, re-baseline, and extend monitoring to the control plane before you scale 2.0 out.