The security-guidance Plugin, Explained
Anthropic ships an official, free plugin that turns Claude Code into its own security reviewer: security-guidance watches every edit for known-dangerous patterns, has a fast LLM re-read the diff when Claude finishes a turn, and runs an agentic reviewer across the codebase at commit time. How the three layers work, what they actually catch, how to tune them for your project, and where they honestly fall short.
Explain it like I’m five
Imagine a construction site where a worker builds scaffolding all day. Now imagine a building inspector who does three things: every time the worker adds a board, the inspector glances at it and says “careful, that one wobbles”; at the end of the day, the inspector walks the whole floor looking for mistakes the worker was too tired to see; and before the building opens to the public, the inspector does a deep check — reading the blueprints, tracing the wiring behind the walls — to make sure nothing hidden will hurt anyone.
That’s the security-guidance plugin. The worker is Claude Code, writing and editing your code. The inspector checks the scaffolding three times over — at every edit, at the end of every turn, and before every commit — so security problems get caught when they’re still cheap to fix, not after the building opens.
The core idea: review the agent’s work while it’s still cheap to fix
Coding agents are fast, and fast is exactly the problem: a vulnerability Claude writes at 2pm can be in production by 2:05. The usual safety nets — a security scan in CI, a human reviewer’s comment three days later — arrive after the bad code already exists.
The plugin’s bet is that the cheapest place to catch a security bug is inside the agent session itself, seconds after it’s written, while the full context is still warm and the fix is one more edit away. Rather than one gate at the end, it layers three reviewers at three different moments: a free instant check on every edit, a smart re-read of the diff when Claude pauses, and a deep cross-file trace when you commit.
This also explains what the plugin is not: it’s not a scanner that says yes or no to your pull request. It doesn’t block writes or commits. Findings come back to Claude as instructions — here’s a suspected SQL injection, go fix it — and the agent resolves them inside the same session. Anthropic positions it as one layer of defense-in-depth: it complements static analysis, dynamic scanning, dependency checks, and human review; it replaces none of them.
Don’t gate the agent’s output at the end — review the agent’s work at three points where the fix is still cheap: on edit, on turn end, on commit.
Layer 1: instant pattern warnings on every edit
The cheapest check of the three runs the moment Claude edits or writes a file — and it costs nothing at all, because there’s no model involved.
It’s a regex-based check for roughly 25 known-dangerous patterns. When Claude writes something like eval(userInput), drops in yaml.load on untrusted data, deserializes with pickle.load, uses torch.load(weights_only=False), assigns raw user input to innerHTML, or pastes a hardcoded sk_live_ secret into code — the plugin fires an instant reminder that this pattern is risky and why.
Think of it as a spell-checker for security: dumb, fast, and free. It can’t understand logic, so it can’t judge whether that eval is actually exploitable — but it can stop the most common accidents the instant they happen, before the agent moves on and loses the thread. The patterns cover the classics: injection, cross-site scripting, server-side request forgery, hardcoded secrets, unsafe deserialization, and path traversal.
Zero model cost means this layer is always on with no latency and no token spend. It’s the tripwire that makes the two smarter layers affordable — the obvious stuff never reaches them.
Layer 2: the Opus 4.7 diff review
Regex can spot a dangerous function call. It can’t spot a logic bug — code that uses perfectly safe functions in a perfectly unsafe way. That’s what the second layer is for.
When Claude finishes a turn, the plugin sends the diff to a fast LLM call — Opus 4.7 by default — for a security review, then feeds high-severity findings back to Claude as instructions so it can fix them before you ever see the response. The diff is small (just what changed this turn), so the review call is cheap and focused.
The textbook example the plugin is designed for: an insecure direct object reference (IDOR). Suppose Claude changes an API call from /api/users/123 to /api/users/124 so a page can fetch another user’s data — but forgets to add an authorization check on the server side. No regex can flag that; the code looks completely normal. But a model reading the diff can reason: this changed which user’s data is requested, and I don’t see an auth check — that’s a suspected auth bypass. The finding goes back to Claude, which adds the check in the same session.
This layer is where the plugin earns its keep: logic-level vulnerabilities — auth bypass, IDOR, business-logic flaws — that pattern matching physically cannot see. The tradeoff is latency: every turn waits on one extra model call. In practice it’s a fast model on a small diff, so the pause is short — but it’s not zero, and that’s the honest price of the review.
Patterns catch the dangerous ingredients; the diff review catches the dangerous recipes.
Layer 3: agentic review that traces data flow
Some vulnerabilities don’t live in any single diff. They live in the relationship between files — a request enters through one endpoint, passes through two helpers, and gets executed by a third. No turn-by-turn review can see that whole path. So the third layer waits for the commit.
On git commit, an SDK-driven reviewer agent opens the diff and then actively reads the surrounding code — using Read, Grep, and Glob — to trace how data flows through the codebase. It can follow a user-supplied value from the API route into the helper that (maybe) sanitizes it, into the database call that uses it. If the sanitizer turns out to be missing or bypassable three files away, it flags the finding with the full trace attached.
This is the layer that catches the multi-file classes regex and single-diff review both miss: cross-file SSRF (a URL parameter flowing untouched into an outbound request), auth checks that exist on one route but not the sibling route, deserialization of user data assembled across modules. It costs more than the other two layers — a real agent session doing real reads — but it only runs at commit time, when you’re already pausing.
Layer 1 is always on and free. Layer 2 is per-turn and cheap. Layer 3 is per-commit and thorough. Each catches what the cheaper ones can’t see — patterns, then logic, then data flow.
Make it yours: project-specific rules
Generic security advice only goes so far — your project has its own rules, and the plugin lets you teach them to all three reviewers.
.claude/claude-security-guidance.md is where you write threat-model rules in plain language. The model reviewers read this file before judging your diffs, so you can encode house policy: how you hash passwords, which auth library is canonical, what counts as a secret. An example, from the plugin’s own documentation shape:
# Our project's security rules
- We hash passwords with bcrypt. Never use MD5 or SHA-1 for passwords.
- All user-facing API routes require the requireAuth middleware.
- Secrets live in environment variables, never in source files.
- The payments/ directory is the highest-risk area: be extra strict there..claude/security-patterns.yaml is the counterpart for Layer 1: custom regex or substring patterns that extend the per-edit tripwire. If your codebase has a legacy internal function that’s dangerous with untrusted input, add a pattern for it — now Claude gets a warning every time it reaches for that function, in every session.
# project-specific tripwires, checked on every edit
patterns:
- name: "legacy-exec-wrapper"
regex: "internalExec\\("
message: "internalExec() does not sanitize input — use safeExec() instead"Both files live in version control with the project, so the rules travel with the repo — every teammate’s Claude session enforces the same policy. For organizations, the plugin can be declared in .claude/settings.json under enabledPlugins, and admins can push it through managed settings so nobody can quietly disable it.
Source: the official security-guidance README ↗Install it
Per the official README, the marketplace ships enabled by default in Claude Code — so for most people this is one command, with no configuration beyond having the CLI.
# from inside a Claude Code session:
/plugin install security-guidance@claude-plugins-officialPrerequisites before you run it:
- Claude Code CLI v2.1.144 or newer. Check yours with claude --version and upgrade if you’re behind.
- Python 3.8+ on your PATH. The plugin’s hooks and tooling run on Python.
- A working API path — a Claude subscription, an API key, or a third-party provider configuration that Claude Code already uses.
Choosing the review model — the LLM layers are configurable through environment variables:
# which model reviews the diff each turn (default shown)
export SECURITY_REVIEW_MODEL=claude-opus-4-7
# Bedrock: us.anthropic.claude-opus-4-7
# Vertex: claude-opus-4-7@20260218
# which model runs the agentic commit reviewer (defaults to the same)
export SG_AGENTIC_MODEL=claude-opus-4-7The plugin is free on all Claude Code plans — Anthropic doesn’t charge for it separately. (The review calls themselves ride on whatever API path Claude Code already uses.) It was announced around May 26, 2026.
Source: install command, prerequisites, and env vars from the official README ↗Honest limits
The plugin is good at what it does. Here’s what it doesn’t do — and why you should know before you trust it.
It never blocks anything. Findings are instructions Claude is asked to resolve, not gates that stop a write or a commit. If the agent misunderstands the finding — or if you tell it to push ahead anyway — the vulnerable code still ships. This is guidance, not enforcement.
The 30–40% number is Anthropic’s, not independent. Press coverage reports that Anthropic’s internal testing showed security-related PR comments dropping 30–40% after the plugin was introduced. That’s the company measuring its own tool — encouraging, but not an independent benchmark, and it measures review comments, not vulnerabilities actually prevented. Treat it as a signal, not a guarantee.
It complements scanners; it doesn’t replace them. The plugin sees your working tree in-session; it doesn’t run your full dependency graph, doesn’t execute your code, and doesn’t know about CVEs in your libraries. You still want SAST, dependency scanning, and DAST in your pipeline — and a human reviewer who can think adversarially about the whole design.
LLM reviewers can be wrong in both directions. A diff review can miss a subtle flaw (false negative) and can nag about a pattern that’s actually safe in context (false positive). If your team gets alert fatigue from the nagging, the customization files from the previous section — teaching it your actual rules — are the fix, not the off switch.
A security-conscious pair programmer that never sleeps: excellent at catching the mistakes agents make at speed, but still one layer in a defense-in-depth stack — not the stack itself.
FAQ
Layer 1 is instant — it’s regex, no model call. Layer 2 adds one fast LLM call on a small diff at the end of each turn; noticeable but short. Layer 3 runs an agent session, but only at commit time, when you’re already pausing. The per-turn cost is the real one: if you want zero added latency, the turn-end review is the knob to question.
Yes — the plugin itself is free on all Claude Code plans. The diff and commit reviews use LLM calls, which ride on the API path Claude Code already uses (subscription, API key, or third-party provider), so those calls are billed whatever your normal usage costs.
No, and Anthropic doesn’t claim it does. It’s positioned as one layer of defense-in-depth: it catches agent-written bugs in-session, while SAST, dependency scanning, and DAST still cover the full pipeline, CVEs, and runtime behavior. Use both.
No — it reviews what changes, not what exists. Layer 1 watches new edits, Layer 2 reviews the turn’s diff, and Layer 3 traces the commit. Your untouched legacy code never passes through it. If you want a full-repo review, check out Anthropic’s separate claude-code-security-review repo, which demonstrates agents autonomously hunting and patching vulnerabilities across a codebase.
Then say so — findings come back as instructions to Claude, not as hard blocks. You (or Claude) can dismiss a false positive and move on. If the same finding keeps nagging, add a rule to .claude/claude-security-guidance.md explaining why the pattern is safe in your project, and the reviewers will learn it.
By default, Opus 4.7 reviews the turn-end diff. Override it with SECURITY_REVIEW_MODEL (e.g. us.anthropic.claude-opus-4-7 on Bedrock, claude-opus-4-7@20260218 on Vertex), and set SG_AGENTIC_MODEL separately for the agentic commit reviewer — it defaults to the same model.
Key takeaways
- security-guidance is Anthropic’s official, free plugin for Claude Code: three security reviewers at three moments — on edit, on turn end, on commit.
- Layer 1 is instant regex tripwires for ~25 known-dangerous patterns (eval, innerHTML, hardcoded secrets, unsafe deserialization) — zero model cost, zero latency.
- Layer 2 sends the turn’s diff to an Opus 4.7 review that catches logic bugs regex can’t — IDOR, auth bypass — and feeds fixes back to Claude before you see the response.
- Layer 3 is an agentic reviewer at commit time that reads across files to trace data flow, catching multi-file vulnerabilities like cross-file SSRF and missing auth checks on sibling routes.
- Make it project-specific: .claude/claude-security-guidance.md for plain-language threat-model rules, .claude/security-patterns.yaml for custom per-edit patterns; enforce org-wide via .claude/settings.json enabledPlugins and managed settings.
- Honest limits: it guides, never blocks; the 30–40% PR-comment reduction is Anthropic’s internal claim, not independent measurement; it complements SAST, dependency scanning, and human review — it replaces none of them.
Sources
Every claim, benchmark, and number in this guide comes from the sources below. The explanations are entirely our own — the measurements and figures are the original authors’.
security-guidance README — anthropics/claude-code on GitHub ↗
The source of the three-layer architecture, the pattern list, the install command, the prerequisites, the environment variables, and the customization files in this guide.
Cyber Security News — “Free security plugin for Claude Code” ↗
Coverage of the plugin’s release and feature set.
The Tech Outlook — Claude Code’s security-guidance plugin ↗
Coverage of how the plugin identifies and fixes vulnerabilities in-session.
GBHackers — Anthropic launches free Claude Code terminal plugin ↗
The source of the reported 30–40% reduction in security-related PR comments, framed in this guide as Anthropic’s internal claim.
anthropics/claude-code-security-review on GitHub ↗
Anthropic’s reference repo demonstrating agents autonomously hunting and patching SQL injection, XSS, RCE, IDOR, and hardcoded credentials.
The 30–40% PR-comment reduction figure is reported in press coverage as Anthropic’s internal testing result. We have not independently verified it — it’s their measurement, presented here as their claim. The install steps were verified against the official README; READMEs move faster than blog posts, so re-check before running.