Frontier Models for Zero-Day & Vulnerability Research
Vulnerability research used to be a craft measured in months. Frontier models are collapsing that timeline — reading code, reasoning across call graphs, and drafting proof-of-concepts at a pace no human team can match. This is a full field guide: which models to use, how to wire them into a multi-agent harness with a persona per bug class, how to persist findings into an Obsidian knowledge base, and what the ROI, true-positive, and false-positive numbers actually look like when you run it in a lab.
At Preemptive Cyber Security we have spent the last year rebuilding our research workflow around large language models — not as a novelty, but as a force multiplier for the slow, expensive parts of finding bugs. The results have been striking, and so have the failure modes. This post distils what we have learned into something you can actually stand up in your own lab.
The changing world of vulnerability research
Two clocks govern security. MTTD — mean time to detect — is how long a weakness lives before someone finds it. MTTR — mean time to respond or remediate — is how long it takes to fix once known. For decades attackers and defenders fought over the gap between them. That gap is now measured in hours, not weeks: n-day exploitation routinely begins within a day of a public advisory, and the same models that help defenders trip-wire a codebase help adversaries weaponise a diff.
The strategic shift is simple to state and hard to live with: whoever compresses MTTD fastest sets the tempo. If a frontier model can surface an exploitable flaw in your dependency before the vendor's advisory drops, you patch pre-emptively. If the adversary gets there first, you are already in an incident. Research is no longer a back-office function — it is the front line of the detect-to-remediate race.
How frontier models actually find bugs
A modern model does not "scan" in the classic sense. It reads. Given source, a decompilation, or a diff, it builds a mental model of data flow, spots where untrusted input reaches a dangerous sink, and reasons about the conditions to reach it. In practice its strengths and weaknesses are complementary to traditional tooling:
- Where it excels: understanding unfamiliar code fast, explaining a suspicious pattern, triaging thousands of static-analysis hits, correlating a patch diff to root cause, drafting a PoC, and writing detection logic or a Semgrep/YARA rule from a described flaw.
- Where it struggles: deep multi-step exploit primitives (heap grooming, precise ROP), long-range reasoning across huge codebases without help, and — critically — confidently hallucinating vulnerabilities that do not exist. Every finding needs verification.
Treat the model as a brilliant, tireless junior researcher with no fear of being wrong. Its output is a lead, never a conclusion.
Comparing models for cybersecurity work
No single model wins on every axis. The right answer is usually a portfolio: a strong closed model for hard reasoning, an open-weight model for bulk triage you can run privately, and a security-tuned small model for cheap first-pass classification. The table below reflects how we weigh them for research use — capability tiers are directional, not benchmark gospel.
| Model class | Examples | Licence | Best for | Watch-outs |
|---|---|---|---|---|
| Frontier closed | Claude Opus / Sonnet 5, GPT-series, Gemini | Paid API | Hardest reasoning, exploit logic, patch-diff root cause, agentic tool use | Data leaves your boundary; per-token cost; rate limits |
| Open-weight general | Llama, Qwen, DeepSeek, Mistral | Open | Private bulk triage, self-hosting, fine-tuning on your corpus | Needs GPUs; weaker at long-range reasoning than frontier |
| Security-tuned | Foundation-Sec, WhiteRabbitNeo-style tunes | Open | Cheap first-pass classification, CWE tagging, log/alert triage | Narrower; still needs a stronger model to confirm |
| Code-specialist | Code-tuned open + closed variants | Mixed | Taint tracing, repo-scale comprehension, refactoring PoCs | Context limits on very large repos; needs retrieval |
Open source vs paid — the honest trade
Paid frontier APIs give you the strongest reasoning with zero infrastructure, but every prompt — including proprietary or client source — leaves your control, and cost scales with volume. Open-weight models you host cost more up front (GPUs, ops) but keep sensitive code on-premises, run at a flat rate once provisioned, and can be fine-tuned on your own vulnerability corpus. For consultancy work under NDA, that data-residency property is often the deciding factor, which is why our research rigs lean on self-hosted models for anything touching client code and reserve frontier APIs for sanitised, high-difficulty problems.
The harness: many agents, one per bug class
A single "find the bugs" prompt underperforms badly. Vulnerability classes demand different mental models — a memory-corruption hunter thinks about lifetimes and bounds; a web hunter thinks about trust boundaries and encoding; an auth hunter thinks about state machines. So we build a harness: an orchestrator that fans work out to specialised agents, each with a tailored persona, system prompt, toolset, and model choice for its class.
Designing a persona
Each agent is defined by four things: a system prompt that frames its worldview ("You are a memory-safety auditor; you care only about object lifetimes, bounds, and integer arithmetic…"), a toolset (grep/AST queries, a sandbox to run candidate PoCs, a symbolic helper), a model matched to the class, and an output contract — structured JSON with a CWE, a location, a confidence, and the reproduction steps. The output contract is what makes the pipeline verifiable and lets the orchestrator dedupe overlapping hits.
Standing the harness up — step by step
- Pick your runtime. An agent framework or a thin custom loop; either way you need tool-calling and per-agent system prompts.
- Define the personas as config: name, class, system prompt, model, tools, and the JSON output schema.
- Give agents safe tools — read-only code access, an AST/grep query tool, and an isolated sandbox for running candidate PoCs. Never let an agent execute untrusted code outside the sandbox.
- Add the orchestrator to chunk the target (retrieval for big repos), fan out to the relevant personas, set a token/time budget, and collect results.
- Insert a verification stage that re-runs each candidate, attempts reproduction in the sandbox, and drops anything unproven below a confidence threshold.
- Persist confirmed findings to your knowledge base (next section) with full provenance — model, prompt, and evidence.
Wiring the harness into an Obsidian Vault
Findings are worthless if they evaporate at the end of a run. We route every confirmed lead into an Obsidian vault — a folder of Markdown notes — so research compounds: notes link to related CVEs, tag by CWE and target, and become queryable with Dataview. The agents read and write the vault through the Local REST API plugin, wrapped by an Obsidian MCP server so any MCP-capable client can treat the vault as a tool.
Step-by-step: connect agents to your vault
- Create the vault. Install Obsidian, create a vault such as
vuln-research, and add folders like/findings,/targets, and/cwe. - Install the Local REST API plugin. Settings → Community plugins → browse → Local REST API → install and enable. It listens on
https://127.0.0.1:27124. - Copy the API key. In the plugin settings, copy the generated key and note the port. Keep the server bound to localhost.
- Add an Obsidian MCP server. Install a community obsidian-mcp server and point it at the REST API URL and key, so it exposes tools like
read_note,create_note,search, andappend. - Register it with your MCP client / harness. Add the server to your client config (the same pattern as any MCP server: command + args, or an SSE/HTTP URL) and restart so the vault tools appear.
- Define a finding template. Give the writer agent a Markdown template with YAML frontmatter —
target,cwe,severity,confidence,status,model— plus body sections for evidence and repro. Consistent frontmatter is what makes Dataview dashboards work. - Verify round-trip. Ask the agent to write a test finding, confirm the note appears in the vault, then query it with a Dataview block. You now have a self-updating research knowledge base.
Example finding frontmatter
target: acme-router-fw-2.3cwe: CWE-416 (Use-After-Free)severity: high·confidence: 0.7·status: unverifiedmodel: memory-safety-agent / code-model- Body: root-cause, tainted-path, candidate PoC, links to
[[CWE-416]]and prior[[acme-router]]notes.
ROI, true positives, and false positives
The honest question every lead asks is: does this pay for itself, and can I trust it? Measure three things continuously — true-positive rate (confirmed real bugs ÷ flagged), false-positive rate (noise ÷ flagged), and yield per dollar (confirmed findings ÷ compute + analyst cost). The figures below are directional from our own lab runs, not industry benchmarks — your mileage depends heavily on target, prompts, and verification rigour.
| Stage | What it produces | Typical TP signal | FP behaviour |
|---|---|---|---|
| Raw agent pass | Many candidate leads, fast | High recall — catches most real issues | High FP; confident hallucinations common |
| + Verification stage | Reproduced / discarded | Real bugs survive reproduction | FP drops sharply once repro is required |
| + Human review | Report-ready findings | Precision approaches manual review | Residual FP mostly edge-case logic |
The pattern is consistent: models give you recall for cheap and precision only after verification. A raw pass might flag ten "criticals" of which two are real — useless if you ship it, gold if you gate it behind automated reproduction and a human. The ROI comes not from replacing the researcher but from letting one researcher cover an order of magnitude more code, spending their time confirming and exploiting rather than reading boilerplate. Where it pays best: large dependency trees, diff-triage after upstream releases, and first-pass audits of unfamiliar code.
Local AI labs vs cloud-hosted
Where you run the models is a security decision as much as a cost one. A self-hosted lab keeps client code on-premises and runs at a flat rate; a cloud/API setup gives you frontier capability instantly but sends prompts off-box.
A starter local lab is more approachable than it sounds: a workstation with a capable GPU (or a small multi-GPU box), an inference server for open-weight models, your harness, and the Obsidian vault — all on an isolated research network with no path to client production. Scale the GPUs as your throughput needs grow. Keep the frontier-API path for problems where the reasoning gap justifies sending sanitised inputs off-box, and log every external call for your own audit trail.
Doing this responsibly
Everything here is for authorised research — your own code, code you are contracted to test, or software with a clear vulnerability-disclosure programme. The same harness that helps you patch pre-emptively is dual-use, so the guardrails are not optional.
Research guardrails
- Authorisation and scope before any target touches the harness — no unsanctioned code, no live third-party systems.
- Sandboxed execution only for candidate PoCs; never run model-generated code on your host or network.
- Coordinated disclosure for anything real — work with the vendor, respect embargoes, and never weaponise beyond proof.
- Data residency — keep client and proprietary code on self-hosted models; log and sanitise anything sent to external APIs.
- Human sign-off on every finding before it leaves the lab. The model proposes; a researcher disposes.
Key takeaways
- The research edge now belongs to whoever compresses MTTD — frontier models pull discovery ahead of disclosure.
- Use a portfolio of models: frontier for hard reasoning, open-weight for private bulk triage, security-tuned for cheap first passes.
- Build a harness with a persona per bug class and a strict output contract — one prompt for everything underperforms.
- Persist findings to an Obsidian vault via the Local REST API + an MCP server so research compounds and stays queryable.
- Models buy recall cheaply and precision only after verification — gate every lead behind reproduction and a human.
- Go hybrid: self-host for client code, reach for frontier APIs on sanitised hard problems, and disclose responsibly.