Spotify’s Shunt Plugin, Explained
A Spotify engineer open-sourced the setup behind a reported 90% cut in Claude Code token usage: stop paying a frontier model to do input/output. A plugin called shunt intercepts bulk file reads and boilerplate writing, hands them to a cheap worker model running on Spotify’s Portal, and keeps Claude on the hard reasoning. Here’s how the three layers work, what the author’s own benchmarks actually measured, where the idea breaks, and how to install it.
Explain it like I’m five
Imagine you hire a world-famous chef for a dinner party. You would never pay them by the hour to wash lettuce and chop onions — you’d hand the prep work to a capable assistant and save the chef for the dishes only they can cook.
That’s the whole idea. When you ask Claude Code to do something, it spends a lot of its time on prep work: reading files, summarizing code, writing boilerplate — all billed at your most expensive model’s rates. Spotify’s shunt plugin is the kitchen manager: it watches for prep work, reroutes it to a cheaper assistant (a small model running on Spotify’s Portal), and hands only the finished answer back to Claude. The author reports this cut their Claude Code token usage by about 90% on bulk-reading tasks — their measurement, on their codebase, which we’ll unpack in full below.
The core idea: pay for reasoning, not I/O
The author’s starting observation is disarmingly simple: most of what a coding agent does inside a session isn’t reasoning. It’s input/output — reading files, summarizing what it found, generating scaffolding. A frontier model is overkill for that.
The author frames this as a cost problem that’s getting worse fast. They cite a Gartner projection that by 2028, AI coding costs could exceed the average developer’s salary, and report that a quarter of engineering leaders already spend $200–$500 per developer per month on tokens. Those are the author’s cited figures, not ours — but the direction is hard to argue with. Every file your agent reads twice, you pay for twice.
The fix the author proposes is model routing as configuration: decide per task which model is good enough, instead of running everything through the most expensive one. Keep the frontier model for debugging, architecture, and judgment calls; send the mechanical work to a worker model that costs a fraction as much. The insight isn’t that cheap models are smart — it’s that most agent work doesn’t need smart.
Route the boring work to the cheapest model that can do it; reserve the frontier model for the work that actually needs reasoning.
Where it actually helps
Not every task is worth rerouting. The pattern that pays off: the input is large, the output is small, and the work is mechanical. Four everyday examples:
You join a team with a sprawling service and ask “where is retry logic handled?” Instead of Claude reading forty files at frontier-model prices, the worker reads them and returns one short answer. This is the exact bulk-read shape the author benchmarked.
A new module needs unit tests written in the repo’s house style. Generating them from existing test patterns is pattern-matching, not reasoning — ideal worker-model work.
Spinning up a new service means type stubs and configuration files that follow existing conventions. Mechanical, convention-bound, delegable.
Before reviewing a big pull request, get a worker-generated summary of what changed across the touched files — then spend your expensive attention on the diff itself, not on discovering what’s in it.
If the task needs judgment — debugging a race condition, designing an API, reviewing security-sensitive code — it stays with Claude. The author is explicit that reasoning never gets delegated.
Portal AiKA Modes: the workers
Portal is Spotify’s system for running small, single-purpose agents. The author describes them as declarative agents on an ephemeral runtime: you declare what the agent is, Portal runs it on demand, and there’s no server to keep alive.
Each agent is called an AiKA Mode. You configure four things: the instructions (what the mode does), the model (which model runs it), the temperature, and which MCP tools it can use. Modes can be public or private, and they’re invoked through the Portal CLI or API. The author’s key point: no infrastructure to manage, no API keys to juggle, no servers — sending a task to a different model becomes a configuration choice, not an infrastructure project.
The two example modes
bulk-reader reads many files and returns a single answer about them. The author configures it to respond in short bullet points with no extra prose, so the answer stays cheap to consume as well as cheap to produce. code-writer generates tests, configuration scaffolding, and type stubs that match the patterns in your existing code — and the author stresses that constraining it to output only the code matters, because any chatter around the code is tokens you pay the worker to produce and Claude to re-read.
The worked examples run on Gemini 2.5 Flash, but the author notes any cheap model can be slotted in — the model is a config field, not a hard dependency.
# a mode declares WHAT the worker is; Portal handles the rest
name: "bulk-reader"
instructions: "Read the given files, answer the question."
instructions+: "Reply in short bullets. No prose."
model: "gemini-2.5-flash" # any cheap model works
temperature: 0.1
tools: ["filesystem"] # MCP tools the mode may use
visibility: "private"This shape is illustrative — the four knobs (instructions, model, temperature, tools) are the author’s description of what a mode configures. Check the Portal docs for the exact schema.
Source: spotify/portal-ai-plugins on GitHub (Apache-2.0) ↗The three shunt layers
Portal provides the workers. The shunt plugin is what wires them into Claude Code automatically — three layers that decide when to hand work off, and how.
Layer 1 · Hooks — the tripwires
A PreToolUse hook inspects every Claude Code Read before it runs. If the file is above the size threshold, the read is blocked and rerouted to the worker instead. The default threshold is 350 lines, adjustable through SHUNT_MIN_LINES — and the author is explicit about why a threshold exists at all: below it, the overhead of delegating costs more than it saves. A second hook, check-bash-read, closes the sneaky path: it catches cat, head, tail, less, and more aimed at large files. Piped commands pass through untouched, and targeted Read calls with offset/limit stay allowed — the hook only stops full reads of big files.
Layer 2 · Bash wrappers — the translators
Two small scripts wrap the Portal CLI so Claude can call the workers without knowing Portal’s API: bulk-read takes a question plus file paths; code-write takes a spec plus reference and target files. Thin glue, deliberately boring.
# bulk-read: one question, many files, one cheap answer
bulk-read --question "where is retry logic handled?" --paths "src/**/*.ts"
# code-write: a spec plus files showing the house style
code-write --spec "unit tests for the new cache module" \
--reference "src/cache/__tests__/" --target "src/cache/new-module.ts"Layer 3 · Skills — the playbooks
Markdown skills teach Claude when to reach for those scripts — for example, “about to read a huge file? delegate it to the worker instead.” And if Portal isn’t set up, the system degrades gracefully: Claude just reads files the normal way. Nothing breaks; you simply don’t get the savings.
The plugin decides when to hand work off; the mode decides how the work gets done. Routing is configuration — not infrastructure.
Benchmarks, honestly
The headline number — about 90% — comes from the author’s own measurements, and it deserves its full context.
The author benchmarked four scenarios on a Java monorepo and reports mean token savings of roughly 90% on bulk-read tasks. That is their codebase, their scenarios, their measurement — not an independent benchmark, and not a promise about your repository. Treat it as a strong signal for one specific shape of work (large input, small output, mechanical) rather than a discount on your entire Claude Code bill.
Two overhead numbers the author is upfront about: each delegation adds roughly 10–30 seconds, and every worker invocation is capped at 30 seconds. The worker is slower per task than Claude reading directly — you are trading latency for tokens. If a task can’t tolerate that pause, don’t route it.
Source: the author’s benchmark write-up on Spotify Engineering ↗Where it breaks
The author is unusually candid about what shunt can’t do — and these limits are the most useful part of the write-up.
Edits are never delegated. The worker’s line numbers come back unreliable, so generating edits through it is asking for corrupted files. Reading is delegable; writing changes is not.
Reasoning is never delegated. In the author’s testing, the worker once missed a subtle thread-safety bug. Debugging, architecture decisions, and safety-critical work stay with the frontier model — full stop.
Latency is real. Ten to thirty seconds per delegation, with a 30-second cap per invocation. That’s the price of the savings; interactive back-and-forth doesn’t belong on the worker.
The threshold is load-bearing. Below 350 lines (the default), the author found that delegation overhead exceeds the savings — so small reads rightly stay local. Tune SHUNT_MIN_LINES to your own usage, but don’t set it to zero.
Shunt is a bulk-I/O discount, not a general intelligence discount. It makes the mechanical parts of your session dramatically cheaper and leaves the thinking parts exactly where they were.
Install it
The plugin is open source (Apache-2.0) in Spotify’s portal-ai-plugins repo, which includes the shunt plugin directory. The commands below are from the repo’s README — verify them there before running, since READMEs move faster than blog posts.
claude plugin marketplace add spotify/portal-ai-plugins
claude plugin install portal@portal
claude plugin install shunt@portal
# then start a new Claude Code session and run:
/portal:setupOne thing to know before you expect savings: delegation only works once Portal itself is set up — the worker modes need somewhere to run. Without it, the plugin degrades gracefully and Claude reads files the normal way. Run /portal:setup in a fresh session and confirm the modes respond before you trust the routing.
Source: install commands verified against the repo README ↗FAQ
That’s the author’s reported number on their Java monorepo, for bulk-read tasks — not a guarantee. Your savings scale with how much of your session is bulk I/O. If you mostly do small targeted reads and heavy reasoning, expect far less.
Yes — every file it reads goes to whichever worker model you configure (Gemini 2.5 Flash in the author’s examples). That’s our observation, not the article’s claim: if your codebase holds secrets or customer data, think about where those bytes go before you delegate.
Yes — set SHUNT_MIN_LINES. The author’s guidance: keep a threshold, because below it delegation overhead costs more than it saves.
Nothing breaks. Per the author, the system degrades gracefully — Claude falls back to reading files directly until Portal is configured.
The author tried the idea and dropped it: the worker’s line numbers come back unreliable. Reading is delegable; writing changes is not.
The author’s examples use Gemini 2.5 Flash, but any cheap model can be configured. The author’s rule of thumb, in our words: pick the cheapest model that clears your quality bar for the task.
Key takeaways
- Most of what a coding agent does is I/O, not reasoning — and paying frontier-model rates for I/O is the waste Spotify’s shunt plugin attacks.
- Portal AiKA Modes are small declarative agents on an ephemeral runtime: you configure instructions, model, temperature, and MCP tools — no servers, no API keys.
- Two example modes: bulk-reader (many files in, one short answer out) and code-writer (tests, scaffolding, and type stubs in your repo’s patterns).
- Shunt wires this into Claude Code in three layers: PreToolUse hooks that block big reads (default 350 lines, via SHUNT_MIN_LINES), bash wrappers around the Portal CLI, and skills that teach Claude when to delegate.
- The author reports ~90% mean token savings on bulk reads across four scenarios on a Java monorepo — their measurement, for mechanical I/O work, not a general discount.
- Honest limits: no delegated edits (unreliable line numbers), no delegated reasoning (it once missed a thread-safety bug), 10–30s added latency per delegation, and a 30-second cap per invocation.
Sources
Every claim, benchmark, and number in this guide comes from the sources below. The explanations are our own — the measurements and figures are the original authors’.
“Portal by Spotify cut my Claude Code token usage by 90%” — Spotify Engineering, September 2026 ↗
The source of the core thesis, the Portal/AiKA architecture, the three shunt layers, the benchmark results, the limitations, and the cost figures cited in this guide.
spotify/portal-ai-plugins on GitHub (Apache-2.0) ↗
The open-source repository containing the shunt plugin; the install commands in this guide were verified against its README.
The Gartner 2028 cost projection is cited by the article’s author; we have not independently verified the underlying Gartner report. Treat all benchmark figures as the author’s reported results on their own codebase.