
Cloudflare Security Audit
FreePerform authorized security audits on codebases.
Free · Opens the source repo
What Cloudflare Security Audit does
The Cloudflare Security Audit skill is designed for security professionals and developers looking to conduct thorough assessments of their codebases for vulnerabilities. This skill provides a structured methodology to audit authorized codebases, focusing on identifying exploitable vulnerabilities through scoped reconnaissance, adversarial review, and validation. It emphasizes the importance of obtaining explicit permission from the system owner before any security testing is conducted, ensuring ethical practices in security assessments.
This skill operates through a series of defined phases, beginning with the establishment of the target codebase and its identity. Users are guided through a mandatory confirmation process to ensure that all actions are authorized and within the agreed scope. The skill facilitates the generation of detailed reports, including findings and recommendations based on the audit, which can be crucial for developers looking to improve the security posture of their applications.
The output of the audit is organized into a dedicated directory, ensuring that all artifacts are systematically stored for easy reference. This includes human-readable reports, detailed findings, and machine-readable outputs that can be utilized for further analysis. The skill is particularly useful for web applications, APIs, and other software components where security vulnerabilities can have significant consequences.
Overall, the Cloudflare Security Audit skill is tailored for professionals who need a reliable and structured approach to perform security audits, ensuring that they can identify and address vulnerabilities effectively while adhering to ethical guidelines.
When to use it
Use this skill when you need to perform a security audit or penetration test on a codebase you own or have explicit permission to assess.
When not to use it
Do not use this skill for unauthorized testing or on systems you do not own, as it is designed strictly for authorized assessments.
What you can build with it
Auditing a Web Application
Use this skill to perform a security audit on your web application, identifying potential vulnerabilities before deployment.
Conducting a Penetration Test
Activate this skill to execute a penetration test on a codebase you have permission to assess, ensuring compliance with security standards.
Reviewing API Security
Utilize this skill to audit APIs for security flaws, helping to secure data exchanges and protect against common vulnerabilities.
How to install Cloudflare Security Audit
View source1. Install with the skills CLI
npx skills add sickn33/agentic-awesome-skills/cloudflare-security-audit --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by sickn33⚠️ AUTHORIZED USE ONLY This skill is for educational purposes or authorized security assessments only. You must have explicit, written permission from the system owner before using this tool. Misuse of this tool is illegal and strictly prohibited.
Mandatory confirmation gate Before running any command that probes, exploits, changes, persists on, extracts data from, or attempts credential access against a target:
- Ask the user to state the exact target URL, IP, account, or resource.
- Ask the user to confirm written authorization and the permitted scope.
- Show the exact command(s) and explain their expected effect.
- Wait for explicit confirmation in the current conversation.
Without that confirmation, remain read-only and provide defensive guidance only. Prefer a sandbox, disposable VM, or controlled lab.
Security Audit
[!WARNING] Authorized Use Only. Audit only code and systems the user owns or is explicitly authorized to assess. Keep testing inside the approved scope and avoid destructive exploitation.
Example
User: Audit this repository for authorization bypasses and injection paths. Keep testing local and non-destructive.
Agent: I will confirm the repository scope, map trust boundaries, validate each candidate, and report only reproducible findings.
You are a security auditor. Your job is to find exploitable vulnerabilities with real impact.
When to Use
Use this skill when asked to perform a security audit, find security bugs, do a security review, audit for vulnerabilities, or pen-test a codebase. Activate it for web apps, APIs, services, CLI tools, libraries, daemons, and more.
Platform terminology
This skill is agent-neutral. In the methodology:
- Task tool means the coding agent's delegation or sub-agent mechanism.
researchagent means a delegated agent optimized for focused codebase exploration and factual verification.generalagent means a delegated agent that can investigate broadly and spawn focused research agents.subagent_typemeans the equivalent delegated-agent role supported by the current platform.
Use the platform's equivalent capabilities while preserving the specified roles, parallelism, prompts, and independence boundaries.
Setup
Before starting, establish two paths and one target identity:
- Target: the codebase to audit (from the user's request or the current working directory)
- Target identity: the canonical physical repository path plus its normalized
originowner/repository URL. Hash both values to create a stable target ID; do not key history by repository basename alone. - Output directory: where all audit artifacts go. Ask the user if not specified, or default to
~/security-audit-skill/<target-id>/run-<N>where<N>is the next unused integer. Create it if it doesn't exist. This ensures same-named repositories cannot share audit history.
All files written during the audit go in the output directory:
architecture.md— Phase 1 output, fed into Phase 2 agent promptsREPORT.md— human-readable report (Phase 4)FINDINGS-DETAIL.md— detailed data flows for MEDIUM+ findings (Phase 4)findings.json— machine-readable structured output (Phase 5)target.json— canonical path, normalized origin, and target ID used to bind this run
Subagents (Phases 1, 2, 3, 6) do NOT write files — they return results to you via the Task tool. You are responsible for writing all files to the output directory.
Coverage and prior runs
Each audit run explores different code paths depending on which agents find what and where they dig. No single run finds everything. Testing shows the best single run finds roughly half the total vulnerabilities across multiple runs.
If prior runs exist for the exact target ID, first require their target.json canonical path and normalized origin to match the current target byte-for-byte. Treat missing or mismatched manifests as unrelated and never read or summarize their findings. Do not search or reuse prior runs from a basename-only directory. After that identity check, read matching findings.json files before starting Phase 2. Use them to:
- Skip known findings — don't waste agents re-discovering the same status bypass. Mention prior findings in the report but focus hunting effort on new ground.
- Target gaps — if prior runs focused heavily on injection and auth, weight this run toward business logic, creative attacks, and the wildcard agent. If prior runs missed public endpoints, focus there.
- Resolve disagreements — if prior runs gave conflicting verdicts on the same finding, validate it definitively.
Include a brief summary of prior runs in the architecture summary so Phase 2 agents know what's already been found.
If no prior runs exist, note in the report that coverage improves with additional runs and recommend the user run the audit again to catch findings this run may have missed.
Core Principles
Only report what you can exploit
Every finding must have a concrete attack scenario: who is the attacker, what do they do, and what do they get? "An attacker could theoretically..." is not a finding. "Send this request, get this result" is.
Confirm dynamically when you can
This is a source-first audit, but a claim you can execute beats one you can only argue. Where the target is locally buildable — a parser, a library, a CLI, a native component — build and run it: reproduce the crash, run the payload, diff the two parsers on the same bytes. Better still, extract the suspect code into a minimal standalone harness and test the hypothesis in isolation — fuzz the one function, feed it the crafted input, watch what it does. Where confirmation needs infrastructure you don't have — a proxy chain, a live cache, production auth — you cannot confirm from source alone: mark it "requires deployment testing" and do not report it as confirmed. Dynamic evidence is what resolves the memory-safety and request-framing classes that static reading leaves ambiguous.
Determine the baseline dynamically
In Phase 1, identify what this application is and what comparable applications exist. Use those comparables to calibrate -- not to dismiss findings, but to focus effort. If the comparable has the same pattern and it's been exploited there, that's a STRONGER finding, not a weaker one. If the comparable has the same pattern and nobody's ever exploited it in 20 years, you should understand why before reporting it.
Do NOT hardcode a specific comparable. A CMS gets compared to other CMSes. An API gateway gets compared to other API gateways. A novel application may have no meaningful comparable.
Defense-in-depth gaps are not vulnerabilities
If Layer A prevents the attack, the absence of Layer B is a hardening note, not a finding. Report it separately if you want, but do not inflate its severity.
Severity requires impact
Severity is the combination of likelihood (how easy to exploit, what access is needed) and impact (what damage is achieved). Use both axes:
- CRITICAL: Unauthenticated RCE, full database dump, admin account takeover without credentials
- HIGH: Authenticated RCE, SQL injection with data exfiltration, stored XSS that fires for all users, auth bypass. Also: any finding where the RBAC/permission model is completely defeated for an action — e.g., a user can perform an action that the system explicitly gates behind a higher role, and the action has real consequences (publishing content, deleting resources, modifying other users' data).
- MEDIUM: Targeted XSS requiring specific conditions, CSRF with meaningful state change, information disclosure of secrets/credentials. Also: business logic bypasses with real but limited consequences — e.g., the action is possible but requires authentication, or the impact is confined to the attacker's own data, or the bypass requires uncommon conditions.
- LOW: Information disclosure of non-secret data, DoS requiring sustained effort
- INFORMATIONAL: A confirmed but minimal-impact observation with no standalone exploit — useful mainly as a building block for another finding. Pure defense-in-depth gaps belong in hardening notes, not here.
The key distinction between HIGH and MEDIUM for business logic findings: does the finding defeat an explicit security boundary? Defeating one — acting past a role the system explicitly enforces — is HIGH; a data inconsistency, a finding that requires privileged access to exploit, or one with limited blast radius is MEDIUM.
If you cannot describe the concrete damage an attacker achieves, the severity is probably lower than you think.
These principles are enforced operationally by the validation rules in HUNTING.md — the canonical bar every hunter applies before reporting a finding, and that Phase 3 re-applies adversarially. The domain companion files add domain-specific checks on top of that bar; they do not replace it.
Workflow overview
Follow all six phases in order:
- Recon — Run Phase 1 from RECONNAISSANCE.md to map the application's architecture, trust boundaries, and input surfaces.
- Hunt — Use HUNTING.md for Phase 2 orchestration, methodology, and validation rules; select scopes from ATTACK-CLASSES.md, which routes native, AI/LLM, HTTP-protocol/auth, and client-side targets to specialized companion files (MEMORY-SAFETY-AND-BINARY.md, AI-AND-LLM.md, WEB-PROTOCOL-AND-AUTH.md, CLIENT-SIDE.md).
- Validate — Use Phase 3 in VALIDATION-AND-REPORTING.md to consolidate duplicates and independently try to disprove every finding.
- Report — Use Phase 4 in VALIDATION-AND-REPORTING.md to write
REPORT.mdandFINDINGS-DETAIL.md. - Structured output — Use Phase 5 in VALIDATION-AND-REPORTING.md and
resources/report-schema.jsonto writefindings.json, then validate it with a trusted JSON Schema validator already available in the user's environment. - Independent verification — Use Phase 6 in VALIDATION-AND-REPORTING.md to verify every factual claim and reconcile all outputs.
Limitations
- Requires a coding agent with a model that supports tool use and parallel sub-agents
- A trusted JSON Schema validator is required for structural validation in Phase 5
- Multiple runs are needed for full coverage — a single run typically finds roughly half of the total vulnerabilities
- The skill does not replace manual penetration testing or automated SAST/DAST tools
Anti-Patterns to Avoid
These are the mistakes that make security audits useless:
- Listing everything that deviates from OWASP as a finding. OWASP is a checklist, not a bug list. Every real application makes tradeoffs.
- Rating defense-in-depth gaps as HIGH/CRITICAL. "Missing validateIdentifier where the query builder already quotes identifiers" is not HIGH severity.
- Ignoring the deployment model. Rate limiting at the CDN layer is a valid architecture. Not every app needs application-level rate limiting.
- Treating designed behavior as a bug. Understand the trust model before auditing. If the design says admins are fully trusted, admin-does-admin-things is not a finding.
- Padding the report with LOW findings to look thorough. Ten LOWs don't make a useful report. Three MEDIUMs do.
- "Potential" findings without proof. Either you can exploit it or you can't. If you need the word "potentially" or "theoretically", you haven't done enough research.
- Ignoring what the codebase does well. If auth is solid, say so. It builds trust in the findings you DO report and helps the team prioritize.
- Constructing exploits from incorrect parser/runtime assumptions. The most convincing false positives come from reasoning "the parser/runtime will interpret this as..." without verifying. If your exploit depends on parser or runtime behavior, cite the spec or test it. Don't assume.
- Skipping business logic and creative attacks. The standard vulnerability classes (SQLi, XSS, SSRF) are what every scanner checks. The value of a manual audit is finding the things scanners can't: logic errors, state machine violations, chained attacks, implicit trust assumptions.
- Giving up too easily. "The codebase uses parameterized queries so there's no SQL injection" is a lazy conclusion. Check EVERY use of sql.raw(). Check dynamic identifiers. Check search/FTS. Check if there's a code path that bypasses the query builder. Push.
Frequently asked questions about Cloudflare Security Audit
Similar skills
Authenticated Scan with OpenVAS
Perform deep vulnerability scans using OpenVAS with credentials.
Active Directory Penetration Test
Conduct focused AD penetration tests with ease.
Active Directory BloodHound Analysis
Visualize Active Directory attack paths and risks.
Orchestrating LLM Attacks with PyRIT
Automate multi-turn adversarial attacks against LLMs.
Operating Sliver C2
Deploy and manage Sliver C2 for red-team engagements.
Operating Havoc C2
Deploy and manage a modern command-and-control framework.
