- 1 GET /ghost/api/content/posts/?key=<content-key>
- 2 200 OK — a members-only post returns its full body
- 3 no member session, no paywall — the gate never applied
The AI hacker that proves it
Warden points a focused agent at one target, lets it propose vulnerabilities, then re-runs every claim. A claim is not a finding until it survives being run again.
No model sits in the judging seat. The agent proposes an exploit; a separate, deterministic engine proves it — and records everything that did not reproduce as a false positive, against the agent's own name.
A real finding — the Ghost restricted-content bypass, re-run to a verdict.
Every security tool has the same disease: false positives
Scanners and “ask an LLM to find bugs” tools bury teams under confident-sounding reports, most of which are wrong. The real bugs get lost in the noise.
An engineer burns a week chasing forty “critical” findings, only to learn thirty-five never existed. That costs the team both tokens and time — and while they dig through the noise, the bugs that matter sit untouched.
There is a second problem underneath it. A single AI told to “find everything at once” spreads itself thin: it skims broadly, reports shallow issues, and misses the deep ones that actually matter.
The model proposes. A deterministic engine proves.
No AI sits in the judging seat. The model never gets to grade its own homework.
It maps the target and proposes an exploit — with the exact request sequence required to trigger it.
provisionalA separate, deterministic engine re-runs that exploit against a control. Same inputs, same verdict, every time. It replays the evidence itself.
provenClaims that do not reproduce are recorded as false positives — in the open, against the agent's own name. Only the true positives are reported.
Find it, prove it, patch it, report it.
The whole loop runs on a single confirmed vulnerability — from the first request to the pull request. You point Warden at a target; it does the rest.
-
Find
Deep source review and black-box pentest against a live target — the way an outside adversary would.
-
Prove
Re-run the exploit, then an equivalent harmless sequence. It is only a finding if the attack fires and the control stays quiet.
-
Patch
Write a minimal fix and demonstrate the exploit no longer reproduces against the patched code.
-
Report
A clean, human-readable write-up for every confirmed finding — the exact steps that trigger it, and nothing that did not survive replay.
-
Notify
Open the fix as a pull request, file a GitHub issue, or message the developers on Slack — automatically.
It works in short, independent rounds
Map the target, attack it, prove or discard each claim, learn from what failed — then throw the round away and go again. When the budget runs out or the target is exhausted, Warden writes the report.
- Map
- Attack
- Prove
- Learn
- Map
- Attack
- Prove
- Learn
- Map
- Attack
- Prove
- Learn
Each round starts fresh
So the agent stays sharp and cannot quietly carry forward a belief the engine already rejected. Only what survived replay is handed to the next round.
Forced to keep exploring
Warden keeps re-running until enough time has passed, so it is pushed to explore more branches instead of giving up early — and you watch the whole thing live in a console.
The standard every finding is held to
This is what directly attacks the false-positive problem: proof, not model confidence, is the basis for a finding.
Every finding is a reproducible exploit
Warden never reports a bug from static reasoning alone. It produces the exact request sequence required to trigger the vulnerability, and hands it to you.
accepted only if no “potential” findings
Every exploit is validated against a control
It replays the attack, then runs an equivalent harmless sequence. A finding is accepted only when the attack succeeds and the control stays quiet — so proof, not model confidence, is the basis for a finding.
accepted only if attack fires + control quiet, or nothing
Even invisible vulnerabilities leave evidence
When success produces no visible output — SSRF, RCE, exfiltration — Warden plants a secret marker the agent never sees and waits for the target to call back. The callback is objective evidence the attack really ran.
accepted only if no callback, no finding
Source-only bugs are demonstrated, not inferred
With no live system to hit, Warden writes a test that fails against the vulnerable code, applies a minimal fix, and shows the test now passes. The vulnerability is demonstrated empirically.
accepted only if a failing test, or it did not happen
What powers Warden
These primitives are what let it work like a real researcher — continuously, autonomously, and with the context of one who has been on the target for a while.
A live exploit database
Current, real-world vulnerabilities and the techniques used against them — kept up to date, not frozen at training time.
A technique & template library
Proven attack playbooks and reusable starting points, so agents spend their time reasoning about a target instead of rebuilding basic tactics.
State-of-the-art attack tools
Warden runs inside a Kali Linux sandbox with the tools a professional reaches for on a real engagement.
A credential vault
It authenticates and tests as a real user without ever holding the password — the access without the secret.
An agent-native inbox
It handles sign-ups, email verification, password resets — the workflows that need a real identity to get past the front door.
A shared wiki
Agents working the same target share discoveries, hypotheses and context — parallel exploration becomes cumulative intelligence, not duplicated effort.
Proven in the wild
Warden has claimed real vulnerabilities in popular open-source software — running on the cheapest model on the internet.
- Ghost Security advisory Affected · >=v4.0.0, <v6.63.0
A high-severity restricted-content bypass: the content API let unauthenticated visitors read gated post content. Disclosed as a GitHub security advisory and fixed upstream.
Confirmed - Joplin Security advisory Affected · 3.7.10
A moderate-severity flaw: an unvalidated x-success / x-error callback target reached Electron's shell.openExternal(), so a crafted link could dispatch arbitrary URI schemes with attacker-controlled data. Disclosed as a GitHub security advisory and fixed upstream.
Confirmed - Ghost Fixes merged upstream Affected · >=v4.0.0, <v6.63.0
The issue Warden reported in Ghost — editor-tier roles could post system-wide notifications rendered as unsafe HTML in Ghost Admin — fixed and merged upstream. One pull request restricts notification creation to administrators; the other sanitises the notification markup.
Point it at one target.
A GitHub repo, a running app, or both. Warden finds it, proves it, patches it and reports it — working in short rounds until the budget runs out or the target is exhausted, then handing you only what survived replay.