Cloudflare's Open-Source Audit Skill Is at 25,900 Stars, but Its Own Post Shows 20,799 Candidates Became 7,245 Actionable Findings
Security / analysis
Cloudflare's Open-Source Audit Skill Is at 25,900 Stars, but Its Own Post Shows 20,799 Candidates Became 7,245 Actionable Findings
The repository trending on GitHub is the 450-line starting point. The funnel numbers in Cloudflare's June write-up describe a different system that has not been released.

A coding agent given Cloudflare's security-audit skill and a repository will map the code, hunt for bugs, try to disprove each one with a fresh agent and write only the survivors into a schema-checked findings.json. The skill sits at roughly 25.9k stars and 1.6k forks on GitHub under an MIT licence, and it gained 617 stars on October 7 alone. Stars measure curiosity, not adoption. The more useful numbers are in the Cloudflare blog post the repository points to, and they describe something else.
What the repository contains
The repo is the single-repository starting point of a harness Cloudflare built internally. Its README lays out six phases:
- Reconnaissance writes
architecture.mdandcoverage-ledger.json, mapping trust boundaries and input surfaces. - Coverage-led hunting assigns isolated hunters to ledger units, with critics looking for gaps.
- Candidate validation has a fresh verifier try to disprove each unique candidate.
- Structured output sorts records into
confirmed,needs_validationandrejected. - Record verification has fresh agents re-check the source claims.
- Reporting generates
REPORT.md,FINDINGS-DETAIL.mdandNEEDS-VALIDATION.md.

A confirmed finding needs a complete source trace and a bounded observed result. A needs_validation item has one unresolved fact and no severity. The README's rule is blunt: "Only confirm established boundary failures." Defence-in-depth gaps become hardening notes, not vulnerabilities.
The skill needs an agent with tool use and parallel sub-agents, plus Node.js for the validators. Target code should run only inside an OS-enforced sandbox with no external network. Without one, leads stay as needs_validation and nothing is executed.
The funnel in Cloudflare's post
The post, written by Grant Bourzikas and published June 18, 2026, says the skill began as a roughly 450-line, single-session, seven-phase audit of one repository. Its limits were context exhaustion, no persistence after a crash and no view of cross-repo relationships. A single run found only about half the bugs that repeated runs found, which is why the README says multiple runs are additive.
Turning it into a pipeline took about six weeks. The result is a fleet scanner across 128 repositories and a shared validation system covering 145. These are Cloudflare's own figures, not independently audited.
- Raw candidates from discovery21K findings
- Findings in validation pool14K findings
- Survived validation (about)12K findings
- Actionable for teams7245 findings
Source: Cloudflare, Build your own vulnerability harness, June 18, 2026 (company-supplied)
The validation pool arithmetic is clean: 13,841 findings, minus 5,442 duplicates folded together and 1,154 routed as wrong-repo or low-risk, leaves 7,245. The post does not reconcile the 20,799 discovery figure with the 13,841 pool, so the chart is a set of reported counts, not a single chain.
What defenders can take from it
The design choices carry more information than the headline totals.
- Separation of roles: In the post's words, "A Hunter has to state the threat model before it's allowed to file anything," and a separate validator cannot file findings. The reasoning: "If a Hunter is allowed to grade its own homework, it will confidently validate everything it outputs."
- Proof or nothing: Each confirmed finding ships with a test that runs against untouched code and a git diff. The initial validation rejection rate fell from 40% to 11%, and the high-integrity share rose from 35% to 58%.
- Humans merge: "The Fixer never merges code on its own; a human must review the branch." The post calls the Fixer the youngest and slowest part of the system.
One measured surprise: hunters invoked Semgrep zero times in a month, while a "wishlist" tool, with which agents ask for missing tools or environments, drew 25,472 entries across the 128 repositories.
What the numbers do not show
Cloudflare states it has no false-negative rate, since no labelled ground truth exists, and measures success by re-run discovery and coverage growth. It also says the findings came from an isolated research environment and are not active unpatched production vulnerabilities. That is a fair caveat, and it means the 7,245 figure cannot be read as a count of exploitable bugs shipped to customers.
The post gives one end-to-end benchmark: a single repository of about 30,000 lines produced about 100 initial findings in three to four hours, refined to 80 distinct bugs, with roughly 14 hours to open pull requests. About 10 critical or high, exploitable bugs were fast-tracked for production in about five days.
The skill in the repository does none of the fleet-level work. Deduplication, the cross-repo tracer and the fixer belong to the harness, which the post says will follow "shortly" and which was not in the repository when we read it. Anyone installing the skill gets the first six phases, one repository at a time.
Why it matters to patching
Volume is the pressure point. DIVD reported an AI agent chaining two Zammad zero-days to reach root. Maintainers who receive 7,245 candidates need the dedup and validation layers more than they need another hunter. Defenders on the receiving end are already working to three-day federal deadlines for flaws that are exploited, not merely found.
The practical test is the one Cloudflare's own design implies: run the skill several times, treat needs_validation as unconfirmed, and read the sandbox requirement as non-negotiable. Whether the released harness reproduces the 12,057-survivor funnel outside Cloudflare is the open question, and nothing public answers it yet.
Sources
More in Security
- 01Cling Botnet Hides Commands in the STUN Transaction ID and Spreads Through Realtek Flaw CVE-2021-35394Nozomi Networks says the malware sends traffic that resembles ordinary Google STUN replies, so defenders have to hunt for all-zero transaction IDs instead of blocking an address.
- 02Apple Fixes CoreGraphics Flaw CVE-2026-86950 in iOS 26.7.1 After Meta Reports Targeted AttacksApple's entry says a crafted file can run code and that exploitation may have hit specific people on iOS versions before iOS 27, but it names no victims and no attacker.
- 03NetScaler SAML Zero-Day CVE-2026-88779 Was Exploited Days After Two Others, and CISA's Deadline Is TodayCitrix rates the flaw 8.7 and calls it a denial of service, but a researcher's honeypot ran a downloaded binary, and appliances patched for last week's bugs need a second upgrade.
- 04FortiMail CVE-2026-104286: Unauthenticated File Write Exploited at Disclosure, With No Fixed Build YetFortinet's advisory FG-IR-26-175 lists fixes for 7.4, 7.6 and 8.0 as pending; the interim steps are disabling IBE and closing the management interface.