Build a 3AM debugging agent with the Claude Agent SDK
Your Playwright suite fails at 3AM. This agent reads the trace, searches git history through MCP servers, finds the guilty pull request, and posts the diagnosis — built with the Claude Agent SDK in about fifty lines of code.
Watch the companion reel ↗ Build a 3AM debugging agent with the Claude Agent SDK
Explain it like I’m five
Imagine your school has a night watchman. Every morning he checks which windows broke overnight, looks up who visited that hallway in the logbook, and leaves you a note: “Room 4’s window — the basketball team, here’s the proof.” That is the 3AM debugging agent: a night watchman for your test suite.
The big idea is simple. Tests fail at night, and humans debug in the morning — groggy, contextless, clicking through traces. An agent flips that around: it does the detective work while you sleep, using the same tools you would use — the test report, the git history, the pull request — and leaves you a solved case instead of a mystery.
What it is
A small program built with the Claude Agent SDK that triages failed Playwright end-to-end tests automatically. Your CI runs the suite at night; when specs fail, CI wakes the agent with a failure bundle — the failing spec names, error messages, trace files, and screenshots.
The detective story, step by step:
- CI wakes the agent with the failure bundle: three red specs, their error text, traces, screenshots.
- The agent reads the trace. Timeout waiting for [data-testid="checkout-btn"], plus a screenshot showing a restyled button → hypothesis: the button changed.
- It searches git history through a Git MCP server: commits from the last 24 hours touching the checkout component and the spec → finds PR #482, “redesign checkout button,” merged at 6PM.
- It diffs the PR. data-testid="checkout-btn" became data-testid="pay-now-btn" — the test was never updated → root cause, with evidence.
- It reports back: a comment on the PR with the diagnosis and a suggested fix, plus a Slack ping to the author.
You define a goal and hand the agent tools; the SDK runs perceive → reason → act until the agent writes its report. The intelligence lives in the loop and the tools — not in custom code.
The tools
The agent’s hands are MCP servers — the Model Context Protocol, Anthropic’s open standard for plugging tools into language models. Three servers do the heavy lifting in this build.
- Git MCP server (@modelcontextprotocol/server-git). Gives the agent git_log, git_diff, git_blame and git_show, scoped to one repository with --repository. This is how it answers “what changed in the last 24 hours?”
- GitHub MCP server (@modelcontextprotocol/server-github). Reads pull-request metadata and posts the diagnosis as a PR comment. Exact tool names vary by server version — check the README of the version you install.
- Playwright MCP server (@playwright/mcp, from Microsoft). Browser tools like browser_navigate, browser_snapshot, browser_take_screenshot and browser_console_messages — the agent can re-run or inspect the failing flow itself.
- File tools for reading the trace files and the CI failure bundle that starts the whole run.
One config block wires them all into the SDK:
from claude_agent_sdk import ClaudeAgentOptions
options = ClaudeAgentOptions(
mcp_servers={
"git": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-git",
"--repository", "."],
},
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"cm"> # token with repo scope; name varies by server version,
"cm"> # check the server's README for the current variable
"env": {"GITHUB_PERSONAL_ACCESS_TOKEN": "ghp_..."},
},
"playwright": {
"command": "npx",
"args": ["-y", "@playwright/mcp@latest"],
},
},
allowed_tools=[
"mcp__git__*",
"mcp__github__*",
"mcp__playwright__*",
],
system_prompt=(
"You are a build detective. Diagnose the failing "
"Playwright specs, correlate them with recent commits, "
"and report with evidence. "
"Read-only: never modify code."
),
permission_mode="acceptEdits",
)Source: microsoft/playwright-mcp — Playwright MCP server ↗
Code walkthrough
Five steps from an empty file to a working detective. Every SDK API name below is real, verified against Anthropic’s Agent SDK documentation; anything version-dependent is marked.
Step 1 — Install and set your key
pip install claude-agent-sdk
export ANTHROPIC_API_KEY="sk-ant-..." # your key from the Anthropic ConsoleStep 2 — Declare the MCP servers
The mcp_servers dict in ClaudeAgentOptions tells the SDK which tool servers to launch. Each entry is a command the SDK spawns — here, npx fetching the published MCP packages. allowed_tools uses the mcp__<server>__* pattern to grant the agent every tool on those servers, and system_prompt sets its role and the read-only rule.
(See the full options block in The tools above.)
Step 3 — Write the brief
CI writes one markdown file per failed run — spec names, error text, a trace excerpt, screenshot paths. That file is the prompt. A good brief also states the mission: “Diagnose these failures. Correlate with commits from the last 24 hours. Report with evidence: trace lines and commit hashes.”
Step 4 — Run the loop
import asyncio
from claude_agent_sdk import query, ResultMessage
"cm"># CI writes this file: failing spec names, error text,
"cm"># trace excerpt, screenshot paths
brief = open("nightly-failure.md").read()
async def main():
async for message in query(prompt=brief, options=options):
if isinstance(message, ResultMessage) \
and message.subtype == "success":
print(message.result) # the diagnosis
asyncio.run(main())query() streams the SDK’s message stream: tool calls, tool results, and the agent’s reasoning, ending in a ResultMessage whose result field holds the final report. You don’t write the perceive-reason-act loop — the SDK runs it.
Step 5 — Read the report
What comes back is the diagnosis text. In the full build, the agent itself posts it as a PR comment through the GitHub MCP server and pings the author on Slack before finishing — so “reading the report” means opening the PR and finding the detective’s comment already there.
The exact tool names the agent invokes inside the MCP servers depend on the server version you install. Your code never names them — the SDK passes the servers’ tool schemas to Claude, and Claude picks the right tool for each step. Upgrade a server, and the agent just gets smarter tools.
Guardrails
An agent with git history and a PR-comment button needs boundaries. Three rules keep this one safe.
- Read-only by default. It diagnoses; it never reverts code, never pushes, never merges. The enforcement is mechanical: allowed_tools grants it no write tools, and permission_mode governs anything the SDK would otherwise ask about. Give it no write tools and it can’t write.
- Every claim links evidence. “PR #482 renamed the test id” must arrive with the commit hash and the diff hunk attached. If it can’t cite it, it doesn’t say it — write that into the system prompt, not just the brief.
- The human owns the fix. The agent suggests; the PR author decides. The Slack ping is a notification, not an assignment — and the suggested fix ships only when a human applies it.
Confident misdiagnosis. An agent that sounds certain about the wrong commit sends a human chasing ghosts at 9AM. The evidence-linking rule is the antidote: a wrong claim with a visible citation gets caught in seconds; a wrong claim without one wastes a morning.
Real-world use cases
The same pattern — failure bundle in, evidence-linked diagnosis out — fits anywhere machines fail at night and humans pay in mornings.
The nightly e2e suite. Three hundred Playwright specs run at 3AM; four fail. By standup, each failure already has a comment on the guilty PR with the trace line, the suspect commit, and a suggested fix. Triage goes from an hour to a glance.
Flaky-test triage. Point the same agent at your flake tracker instead of CI. It correlates each flaky failure with recent commits and separates “real regression” from “timing-sensitive test that needs a better wait” — with the diff to prove it.
The deploy gate. Run it as a pre-deploy check: if the canary fails, the agent diagnoses before the rollback decision. The on-call engineer gets a root cause with the page, not just the page.
The pattern: anywhere the debugging loop is “read logs, check git, correlate,” an agent can run that loop faster than a human — as long as a human still approves the fix.
FAQ
The SDK ships with built-in tools (file reading, bash, and more), which is enough for a first prototype — the agent can already read traces and run git commands. MCP servers matter when you want maintained, scoped toolsets: the git server restricts the agent to one repository, the Playwright server brings a real browser. Start built-in, graduate to MCP.
Each run bills the tokens the agent consumes: the failure bundle in, plus tool results and reasoning, times your model’s price. A focused triage run is typically cents, not dollars — but scope the git history window (24 hours, not 24 months) or a chatty agent will happily read your entire repo history.
Then the agent’s correlation finds no guilty commit — and that is the diagnosis: “no recent change explains this; the trace points at the payment service timing out.” Ruling out the usual suspect is valuable triage, and the evidence trail shows its work.
Technically yes — technically you shouldn’t let it, yet. Auto-applying fixes turns a misdiagnosis into a bad commit at 3AM with nobody watching. Keep the agent read-only until its diagnoses earn trust over weeks of correct calls; then graduate one repo at a time, with CI re-running the suite on every agent-proposed fix.
Takeaways
- An agent is a goal plus tools plus a loop. The Claude Agent SDK runs perceive → reason → act; you supply the failure bundle and the MCP servers.
- MCP servers are the hands. Git history, PR comments, and a real browser arrive as tool schemas — the agent picks the right tool per step, and your code never names them.
- The build is about fifty lines. Install, declare servers, write the brief, stream through query(), read the ResultMessage.
- Read-only by default. Diagnose, never revert; every claim cites a trace line or commit hash; the human applies the fix.
- Watch the failure mode. A confident misdiagnosis with visible evidence gets caught in seconds — one without evidence wastes a morning.
Sources
Every API name in this guide — query, ClaudeAgentOptions, mcp_servers, allowed_tools, system_prompt, permission_mode, ResultMessage — and the MCP server packages and their tools were verified against the documentation and repositories linked below, read on October 9, 2026. The detective scenario, analogies, use cases, and explanations are our own.
- Anthropic — Claude Agent SDK overview ↗
- Anthropic — Agent SDK Python reference ↗
- Model Context Protocol — official docs ↗
- modelcontextprotocol/servers — git & GitHub MCP servers ↗
- microsoft/playwright-mcp — Playwright MCP server ↗
Companion reel: this guide will be linked from @theclaudecraft’s “Build a 3AM debugging agent” reel once it posts.