
The skill that grew into Cloudflare's 7,245 bug findings. What will it find in your repository?
Can you trust Claude Code with a security audit? Cloudflare open-sourced a skill where a different agent checks every finding. Its phases, setup and gaps.
In this article6
security-audit-skill is an open-source Cloudflare skill under the MIT license that turns a coding agent such as Claude Code into a security auditor. The audit runs in six phases: reconnaissance, a hunt by attack class, a check of every candidate by a fresh agent that tries to refute it, a JSON findings ledger, a second independent pass against the source, and a report. Cloudflare's internal harness grew out of this skill: 20,799 raw candidates, of which 7,245 findings reached engineering teams after validation and deduplication. But those numbers come from the harness, and only the skill has been released. It installs with a single npx skills add command and needs a model that can run parallel subagents. According to the authors, one run finds roughly half of what several runs find together.
Cloudflare has released the tool that found 7,245 bugs in its own code. More precisely, it released what that tool started as. On 18 June 2026 Cloudflare's security team published a write-up of its vulnerability harness and on the same day opened the cloudflare/security-audit-skill repository. By 29 September the repository had 22,845 stars, 3,962 of them in the past week, which put it at the top of GitHub Trending.
A security audit run by an agent is something you want to try on your own project today, and here is a ready recipe from a team that has already run the approach across 128 of its repositories. Let's look at what the recipe is, how it works in Claude Code and what not to expect from it.
A skill for coding agents: a six-phase security audit with an independent check of every finding, a JSON ledger and zero-dependency Node.js validators. MIT license.
What security-audit-skill is and where it came from
A skill is a set of markdown instructions that an agent loads when it sees a matching request. If you haven't installed skills before, we have a guide to how skills work and get connected in Claude Code. Here the main SKILL.md file sets the rules and order of work, and next to it sit prompts for each phase, a shared list of attack classes and ten companion files for different kinds of code: native code and binaries, LLM applications, HTTP and authorization, the browser client, the supply chain, cloud, RPC and queues, resource exhaustion, data isolation, desktop and mobile apps.
According to the blog, Cloudflare started with a skill of about 450 lines, ran it against a single repository and tuned the prompts until real bugs started coming out. Then each phase became a separate agent, with a database and an orchestrator behind them. Going from the first slash-command run to a scanner across 128 repositories took about six weeks. The authors write that most of the value lives in the prompts, and those barely changed inside the harness.
Six audit phases: some agents hunt, others refute
The core idea of the skill fits into one line of the README: the agent that validates a finding is never the one that found it. Everything else is built around it.
- 1
Reconnaissance
Parallel agents read the code and write architecture.md: the stack, entry points, trust boundaries. At the same time a coverage-ledger.json is created, a ledger of what needs to be checked and against which attack class.
- 2
Coverage hunt
Isolated hunters take their areas from the ledger and look for violations of specific invariants. After each wave a separate critic looks for missed entry points.
- 3
Candidate validation
Each candidate goes to a fresh agent with the instruction «You did not write this candidate. Try to refute it». It never sees the output of other validators.
- 4
Structured output
Findings are written to findings.json with a verdict of confirmed, needs_validation or rejected. The file is checked by a JSON schema and a Node.js validator.
- 5
Second pass
New agents check every final record against the source once more. Any substantial edit is checked by yet another independent agent.
- 6
Report
REPORT.md, FINDINGS-DETAIL.md and NEEDS-VALIDATION.md are assembled from the checked records.
What we like most is how the skill handles doubt. The needs_validation verdict does not mean «a vulnerability, but we're not sure». It is a specific hypothesis with one unknown fact: say, a proxy setting that isn't in the repository. Such a record has no severity, but it does have a plan for checking that fact. Severity on confirmed findings is strictly tied to what was proven: critical only if an unauthenticated attacker gets code execution, full storage access or other people's accounts. A missing recommended practice does not count as a vulnerability.
7,245 bugs at Cloudflare: which numbers belong to the skill
Retellings pass around the line «12,000 confirmed vulnerabilities across 128 repositories». In Cloudflare's blog the numbers mean something different. 20,799 is the harness's raw candidates over its whole lifetime. About 12,057 of them passed the first validation. Findings from another harness were then added, the combined pool grew to 13,841 across 145 repositories, a deduplication agent folded 5,442 duplicates, and another 1,154 went into the «wrong repository or not a risk» bucket.
findings reached Cloudflare engineering teams after validation and deduplication, out of 20,799 raw candidates
The main discrepancy isn't in the arithmetic. These numbers came from the harness, not the skill. In the harness one model does the hunting and an entirely different one does the validation, so that a finding is judged by different weights and different training data. In the skill a different agent refutes the finding, but usually of the same model, just with a fresh context. On top of that the harness has things the skill cannot have by definition: a database that lets work resume after a failure, vulnerability tracing across repositories, and separate deduplication and fix agents. Cloudflare promises to release the harness itself later; as of 29 September the repository contains only the skill.
And the skill is no longer the one described in the blog. The blog describes seven phases and about 450 lines. The repository has six phases, and on 10 September a large PR reworked the workflow, the findings contract and the validators: now SKILL.md, HUNTING.md and VALIDATION-AND-REPORTING.md alone run past 600 lines, and the two validators come to about 1,650. A coverage ledger, run profiles and strict sandbox requirements appeared. The skill became more careful and heavier.
How to run a security audit in Claude Code
The skill isn't tied to one agent: it needs a model with tool calling and parallel subagents. Claude Code fits. It installs through the skills.sh CLI:
npx skills add https://github.com/cloudflare/security-audit-skill \--skill security-audit# add --global to install it for every project on the machine# then in an agent session, at the root of the repository:security audit this codebasedo a security review, output to ~/audits/my-project
By default the skill works in guidance mode: it will answer a question about a specific vulnerability, but it won't create files or launch all six phases. A full audit only starts on an explicit request to audit or pentest the codebase. Results are written outside the repository, to ~/security-audit-skill/<repo-name>/run-<N>. Inside the project the skill only writes to a folder you named yourself and that git ignores.
The size of a run is set by a profile: quick for a first look, standard by default, deep for large and critical systems. You can also set a budget as a number of agent calls. If the budget doesn't cover reconnaissance and validation, the skill will honestly refuse to start rather than skip the check of findings. Repeat runs complement each other: the skill reads previous ledgers and aims at the gaps.
Where the skill hits its limits
Three limits before your first run
You need an OS-level sandbox: no network, an empty environment and resource limits. Without one the skill won't execute your code and leaves findings as needs_validation. One run, according to the authors, finds roughly half of what several runs find together, and more often the simple bugs. And this is a periodic audit, not a check for every PR: Cloudflare's worst full scan took more than 14 hours.
The sandbox isn't a formality: during validation the agent builds and runs the target's code, and that code can do anything. We covered how agents find their way out of a poorly closed container in OpenAI's report on an agent escaping through DNS.
What we take for ourselves
Our opinion, which you are free to dispute: the most valuable thing in the repository isn't the attack classes but the rule «whoever found it doesn't approve it». It works far beyond security. Code review, fact-checking a text, a data migration: wherever an agent grades its own work, it tends to approve it. Cloudflare writes in the blog that an agent may edit the source so its exploit works and then triumphantly report a bug it created itself. A separate agent with no access to the first one's reasoning catches this more cheaply than a human does. We'll be wrong if, on your tasks, a second agent of the same model simply repeats the first one's mistakes; then the check needs a different model, which is exactly what Cloudflare's harness does. We worked out what the scaffolding around a model costs and when it pays off in our comparison of Claude Code, Codex and pi.
How to start without extra risk
Start with quick on one service that has external input: an API, webhooks, file uploads. Read NEEDS-VALIDATION.md: it lists the facts only you know. Run a second pass a week later, it will pick up what was missed. The skill only describes fixes, it doesn't change your code.
Versions and numbers as of 29 September 2026: the last commit in the repository was on 14 September, 22,845 stars. The skill has changed noticeably since June, so check the current README.
Sources4expand
- Cloudflare, «cloudflare/security-audit-skill», README and SKILL.md, checked 29 September 2026 — https://github.com/cloudflare/security-audit-skill
- Dan Jones, Alexandra Godoi, Grant Bourzikas, «Build your own vulnerability harness», Cloudflare Blog, 18 June 2026 — https://blog.cloudflare.com/build-your-own-vulnerability-harness/
- Cloudflare, «security-audit-skill commits», change history, checked 29 September 2026 — https://github.com/cloudflare/security-audit-skill/commits/main
- GitHub, «Trending repositories this week», 29 September 2026 — https://github.com/trending?since=weekly
Read next
The alarm went off after 12 minutes. The agent was stopped two and a half hours later. What actually broke in OpenAI's sandbox?September 27, 2026
Same model, same success rate, twice the money. What are you actually paying Claude Code for?September 17, 2026
Agents now get a boss, a task queue and a budget. Do you need one if you already work in Claude Code?September 26, 2026
Comments