The Code Modernization Plugin, Explained
Anthropic ships an official, free plugin that points Claude Code at your oldest systems — COBOL, legacy Java, .NET Framework, PHP monoliths — and walks them through a strict sequence: understand first, get a human to approve the plan, then build, then prove the new code behaves like the old. How the five specialist agents work, the three ways to build, and where it honestly falls short.
Explain it like I’m five
Imagine your family has lived in the same big old house for fifty years. The wiring is ancient, the plumbing groans, and nobody remembers why the third-floor radiator only works if you tap it twice. You want to renovate — but people still live there, and if the builders misunderstand one weird quirk, the whole house stops working.
So you hire a careful crew. A surveyor measures every room and writes down what’s actually there. An architect draws a floor plan showing how the rooms connect. A historian interviews the house and writes down every rule — tap the radiator twice — with the exact page of the logbook it came from. A foreman makes you sign the renovation plan before a single wall comes down. The builders do the work. Then an independent inspector checks the new wiring against the old, wire by wire, and a safety officer does one last sweep.
That’s the code-modernization plugin. The house is your legacy codebase. The crew is a team of AI specialists, and the foreman never lets them skip a step — because modernization fails when teams renovate before they understand the house.
The core idea: modernization fails when teams skip steps
Ask any engineer who has lived through a rewrite: the failure mode is almost never the new technology. It’s the skipped homework — transforming code before understanding it, or shipping without a harness that catches behavior drift.
The plugin’s answer is to enforce a sequence. Per the official README, it runs: preflight → assess → map → extract-rules → brief → (reimagine | transform | uplift) → verify → harden. Each step stands alone, so you can stop and review after any of them — and the build steps refuse to run until a human approves the brief.
This also explains the plugin’s shape: it’s not one smart agent but seven slash commands plus five specialist agents that argue with each other — a critic whose job is to call out over-engineering, a test engineer who checks whether your tests actually pin behavior, a security auditor tuned for the mistakes cross-stack rewrites introduce. The sequence is the product; the agents are the crew that walks it.
Don’t point an AI at legacy code and say “modernize it” — make it prove it understands the system, write down the plan, get your signature, and then prove the new code behaves like the old.
The enforced path: understand, approve, build, prove
Every command ends with a “next step” line telling you exactly what to run next. Here’s the full path, from the README’s own table.
Step 0 — say what you want. /code-modernization:modernize asks a couple of short questions, writes your answers to INTENT.md once, and every later command reads it — nothing is asked twice. Not sure what you want yet? Choose understand it first: you get the assessment, the map, the business rules, and a plan, and no code is rebuilt.
Step 1 — preflight. modernize-preflight checks the environment is ready: it asks five questions only a person can answer (is this the whole system or a slice, can it build and test here, is there custom build tooling, has anyone tried before, is anything off limits), proves a build works on this code, and finds missing source. Your legacy code is linked, never copied, and nothing edits your legacy source — commands write only to analysis/<name>/ and modernized/.
Steps 2–4 — assess, map, extract-rules. modernize-assess inventories the codebase: complexity, debt, security posture, and a recommended pattern. modernize-map draws the structure — dependencies, data flow, entry points, business flows — as an interactive map. modernize-extract-rules mines the business rules out of the code as testable Given/When/Then cards, each with a file:line citation, and each card is re-checked by a second agent. A companion modernize-review step lets a person confirm or correct the rules that look wrong — your corrections flow into the plan and the build.
Step 5 — the brief, the approval gate. modernize-brief synthesizes everything into a phased plan a steering committee can approve. Nothing is built until you approve it. The README is explicit: a person decides at six points — the five preflight answers, the rules that look wrong, approving the brief, accepting any difference the proof finds, signing the proof, and applying the security patch — and the plugin never decides any of them for you.
Steps 6–8 — build, verify, harden. One of three build methods (next section), then modernize-verify — the proof — then modernize-harden, a security scan of the legacy system with a reviewed patch you apply yourself. A modernize-status command tells you where you are and the exact command to paste, any time.
Each artifact feeds the next: the map needs the inventory, the rules need the map, the brief needs the rules, the build needs the brief, and the proof needs the rules to know what “behaves the same” even means. Skip a step and you’re modernizing on vibes.
The five agents: a team that argues with itself
The plugin’s own addition notes describe five specialist agents. Each does one job — and the adversarial ones exist so the team checks its own work.
Does deep reads on unfamiliar dialects — COBOL, legacy Java, and friends. This is the agent that can actually sit with a 40-year-old program and tell you what it does, instead of guessing from the file names.
Pulls business rules out of procedural code with citations — the Given/When/Then cards from the previous section. The “legacy code is the oracle” principle lives here: the old system’s behavior is the specification, and every rule points at the exact lines it came from.
The adversarial reviewer. Its job is to flag over-engineering in the proposed target design — the teammate who asks “do we really need a microservice for that?” before you build one.
An OWASP-flavored review tuned for cross-stack translation issues — the specific class of bugs a rewrite introduces, like an auth check that existed in the old framework’s middleware and silently disappears in the new one.
Audits test suites for behavior-pinning versus coverage theater — the difference between tests that prove the system does what it does, and tests that merely execute lines. Only the former count as proof.
The heavy steps run many agents at once: the README says extract-rules started 50 to 200 agents in its trial runs, and every fan-out step announces how many it will start before it does. That’s also why the README warns to expect real usage on a large system — and to start with one module as a pilot, not the whole estate.
Source: the agent roster from the plugin’s addition notes ↗Three ways to build: transform, reimagine, uplift
The brief recommends one of three methods — they’re genuinely different operations, not three names for “rewrite it.”
A cross-stack rewrite from extracted intent, one module at a time, while the old system keeps running. Example from the README: COBOL → Java. The plan is approved first, pinning tests capture the old behavior, then an idiomatic rewrite — and the proof compares old and new on the same inputs.
A greenfield rebuild on a new architecture. A spec is mined from the code, the architecture is reviewed and approved, then services are scaffolded with executable acceptance tests. For when the old design itself is the problem, not just the language.
A same-stack version bump that preserves the code and fixes only the version deltas — .NET Framework → .NET 8, Java 8 → 17, Spring Boot 2 → 3 — driven by a catalog of the breaking changes this code actually hits. A version move can skip extract-rules and review entirely: preflight, assess, uplift, verify is a complete path.
Two things the brief won’t recommend a build for: rehost (move the system as-is) and replace (buy a product instead) — neither changes code, so no build command applies, though the analysis is still useful for both. And if the delta catalog shows an “uplift” would rewrite most of the code anyway, the command says so and points you at transform instead.
Source: the three build methods from the official README ↗The proof step: trust, but byte-compare
This is the part that separates the plugin from “AI rewrote my code, hope it works.” Modernization fails quietly: the new code passes the tests someone wrote, and differs on the inputs nobody thought of.
So modernize-verify builds its proof from evidence a script can check, not from the model’s opinion — and it re-does the work instead of trusting the build’s notes:
- The tests, run again from clean. Build output is deleted, the full suite runs, and counts come from the runner’s own result files. A run that executed no tests is a failure; a count the command merely typed in does not count.
- Old and new on the same inputs. Where the legacy code can run, both versions run on identical inputs and a script compares every byte — only fields that must vary (timestamps, generated IDs) are masked, and every masked field is listed.
- Inputs nobody used. It invents at least ten new inputs — boundaries, empty, huge, unusual order, malformed records — and compares again, so a suite that only passes on cases the model chose gets caught.
- Tests that can fail. A deliberate one-line break must turn the tests red — the canary — and the comparison itself is checked by changing one byte of an output.
- Every critical rule is backed by a test that ran. Each P0 business rule must be named by a test that executed and passed; a rule named only by a skipped test is listed as “named, not run.”
Each module gets one verdict, recomputed from current evidence every time: PROVEN, PARTLY PROVEN (with what’s missing), or NOT PROVEN. If the old system can’t run where you work — a mainframe program, a system that only runs in production — the best possible verdict is PARTLY PROVEN, and the report says so plainly. Anything a person must decide is listed as waiting for a person and never ticked for you.
Most “AI modernization” demos end at “look, it runs.” The proof step is the plugin admitting that running isn’t the bar — behaving identically on inputs nobody wrote tests for is the bar, and it built machinery to check exactly that.
Real-world use cases
Anthropic ran the plugin for real, headlessly, on public codebases while building it — and published what happened, including the runs that didn’t pass. Every number below is the README’s own reported figure from those runs.
AWS CardDemo — COBOL with CICS and JCL, rewritten in Java one job at a time. The plugin mined 35 rules (8 critical) and wrote a five-phase plan, then rewrote the monthly interest job in Java 21 with 213 tests. Five independent checks compared it against the real COBOL program on inputs written after the fact; each found something the tests had missed — a blank field that halted the job, a non-numeric account key, a negative zero — and the final check ended NOT PROVEN. The plugin reported that plainly instead of rounding up.
Eclipse Jetty — Java 8 to 17, 2,600 files. A pilot on the jetty-util module: the same 946 tests gave the same results on both versions. Building and running both versions surfaced six changes that reading the code had not predicted — including a build that would have shipped Java 8-labeled classes that crash on Java 8. This is the uplift story: the delta catalog earns its keep.
osCommerce — PHP 5, ~44,000 lines, moved to Python and FastAPI. 97 rules (26 critical), six-phase plan; the first slice (the product page) finished with 17,957 tests passing and 17,727 comparison cases against the real PHP files identical. Two deliberate canary breaks each turned tests red. Independent verification: PROVEN.
eShop — .NET Framework 4.7.2 to .NET 10. The delta catalog listed 12 silent behavior changes and insisted the first phase build a test harness, because the solution had none. The classic enterprise trap — “it’s just a version bump” — caught at the catalog stage instead of in production.
Understand-it-first mode. Join a team, inherit a system nobody can explain, choose understand it first at the front door: you get the assessment, the interactive map, the cited business rules, and a plan — and no code is rebuilt. For many teams, that artifact set alone is worth the install.
The pattern across all of them: the value isn’t the rewrite — it’s the cited understanding and the proof. The rewrite is just what you do once you have both.
Source: trial runs documented in the official README ↗Install it
You need Claude Code. Then it’s one command, per the README:
# from inside a Claude Code session:
/plugin install code-modernization@claude-plugins-officialThen open a folder for the work and type /code-modernization:modernize — it asks what you want to do, finds the code, and gives you the exact first command. Or go step by step with modernize-preflight <name> --source <path> and follow each command’s “next step” line.
Prerequisites (helpful, checked by preflight):
- Python 3.8+ on your PATH — the map, the proof, and the report run on Python.
- A build toolchain for your stack — this enables the strongest proof (old and new running side by side). Without one, the plugin falls back to recorded-output tests and says so.
- scc or cloc for size metrics; the whole system in the tree (deployment descriptors, copybooks, DDL) so entry points and data lineage resolve.
Lock down the workspace — the commands never edit your code by convention, and the README recommends backing that up with a deny rule in .claude/settings.json (preflight checks for it):
{
"permissions": {
"allow": ["Read(**)", "Edit(analysis/**)", "Edit(modernized/**)"],
"deny": ["Edit(/legacy/**)"]
}
}Honest limits
The plugin is unusually candid about its own limits — the README documents them at length. Here’s the short version.
Heavy steps cost real usage. Rule extraction fanned out 50 to 200 agents in the trial runs. On a large system, expect meaningful model spend — which is why the README says to start with one module as a pilot, not the whole estate. This is not a fire-and-forget tool; it’s a project.
Rule extraction isn’t deterministic. A model does the mining, so two runs can find different rules. The README says to treat BUSINESS_RULES.md as reviewed output, not a deterministic build artifact — which is exactly why the human review step exists.
Some systems cap at PARTLY PROVEN. If the legacy code can’t run where you work — a mainframe program, a production-only system — old-versus-new comparison is impossible, and the report says PARTLY PROVEN plainly. That’s a property of your environment, not a bug, but it bounds what the proof can promise.
Analyzed code is untrusted input. A hostile codebase can plant “ignore previous instructions” comments or credential-shaped traps. The plugin treats file content as data, lists instruction-shaped text it found without following it — Anthropic even tested this with a booby-trapped codebase, and reports that nothing planted was obeyed — but the README still says to treat discovery artifacts with the same skepticism as the code.
It counts what you do. The plugin sends usage telemetry — whole numbers only, never code, paths, or prompts — through Claude Code’s own telemetry. You can turn it off with CODE_MODERNIZATION_TELEMETRY=0, Claude Code’s DISABLE_TELEMETRY=1, or the plugin’s own Usage-counts option.
A rigorous, expensive, human-gated way to modernize — not a magic “rewrite my monolith” button. Its best feature may be the discipline it enforces: understand, approve, build, prove, in that order, with the receipts to show for each step.
FAQ
No. The commands never edit legacy source by convention — they write only to analysis/<name>/ and modernized/ — and the README recommends a .claude/settings.json deny rule on the source tree that preflight verifies. Your original code stays byte-identical; the proof step even checks that via file times.
The README’s rough times for systems of tens of thousands of lines: preflight 3–4 minutes, assess 5–8, map 5–15, extract-rules 5–15, brief about 5. Building one module takes 15–30 minutes; an uplift pilot about 15. On million-line systems, work one module at a time.
Then old-versus-new comparison is impossible and the best verdict is PARTLY PROVEN — checked against recorded outputs instead. The report states this plainly; it’s a bound on the proof, not a failure of the tool.
Not necessarily — a model does the extraction. Treat BUSINESS_RULES.md as reviewed output: read it, correct it in the review step, and don’t expect byte-identical results across runs.
No. Telemetry sends whole numbers only — counts of commands, rules, agents — never code, file names, paths, or prompts, and counts of 100+ are rounded. It rides Claude Code’s own telemetry, so nothing is sent when that’s off, and the plugin itself makes no network calls.
The plugin is free. The agent and model usage rides whatever API path your Claude Code already uses — but the heavy steps are genuinely heavy (up to ~200 agents on extract-rules), so budget real usage on a large system and pilot one module first.
Key takeaways
- code-modernization is Anthropic’s official, free plugin for Claude Code that enforces a strict modernization sequence — understand first (preflight, assess, map, extract-rules), get human approval (brief), then build (transform, reimagine, or uplift), then prove (verify) and secure (harden).
- Five specialist agents do the deep work: legacy-analyst reads unfamiliar dialects like COBOL, business-rules-extractor mines Given/When/Then rules with file-and-line citations, architecture-critic argues against over-engineering, security-auditor reviews cross-stack translation risks, and test-engineer checks that tests pin behavior rather than chase coverage.
- Three genuinely different build methods: transform rewrites cross-stack (COBOL to Java), reimagine rebuilds greenfield on a new architecture, and uplift bumps the version (Java 8 to 17, .NET Framework to .NET 8) while preserving your code — the brief recommends which one fits.
- The proof step is the differentiator: old and new run on the same inputs with byte-level comparison, invented edge-case inputs, and a canary break that must turn tests red. Verdicts are PROVEN, PARTLY PROVEN, or NOT PROVEN — recomputed from evidence on every run, never carried over.
- Safety by design: legacy source is never edited (backed by a permission deny rule), analyzed code is treated as untrusted input, discovered secrets are masked, and a person decides at six gates — including brief approval and proof sign-off.
- Honest limits: heavy steps fan out dozens to hundreds of agents (pilot one module first), rule extraction isn’t deterministic across runs, systems that can’t execute cap at PARTLY PROVEN, and usage telemetry sends counts only — with an opt-out.
Sources
Every claim, number, and trial-run result in this guide comes from the sources below. The explanations are entirely our own — the measurements and figures are the original authors’.
code-modernization README — anthropics/claude-plugins-official on GitHub ↗
The source of the enforced sequence, the step table, the three build methods, the proof design, the trial-run results, the safety rules, and the telemetry details in this guide.
“Add code-modernization plugin” — anthropics/claude-plugins-official ↗
The source of the seven-slash-command inventory and the five-agent roster (legacy-analyst, business-rules-extractor, architecture-critic, security-auditor, test-engineer).
The trial-run numbers (CardDemo’s 213 tests, Jetty’s 946 tests, osCommerce’s 17,957 tests) are the README’s own reported figures from Anthropic’s test runs on public codebases — presented here as their measurements, not independent benchmarks. The install steps were verified against the official README; READMEs move faster than blog posts, so re-check before running.