What Happens When You Hit Enter in Claude: Streaming, Explained
You press enter. Silence. Then words flow in, one by one. Underneath, your prompt became a single HTTPS request with one flag — stream: true — and the reply arrived as a stream of server-sent events. Here’s the full journey: time to first token, why it’s SSE and not NDJSON or WebSockets, the event sequence your client stitches into words, and why Claude Code’s ‘Cogitating…’ spinner is honest theater.
Explain it like I’m five
Imagine you order dinner at a restaurant. The old way: the kitchen cooks your whole meal in silence and the waiter brings everything at once — you stare at an empty table wondering if anyone heard you. The new way: the kitchen sends out each dish the moment it’s ready, and the waiter keeps you posted between courses.
That’s streaming. When you hit enter in Claude, your words travel to Anthropic’s computers as one request with a small flag that says “don’t wait — send me each piece the moment it’s ready.” The reply arrives word by word because the kitchen is plating dishes as they finish, not because someone is typing fast. And that little Cogitating… message with the timer? That’s the waiter making conversation while the kitchen works — the app being polite during the silence, not the chef talking.
Your prompt becomes one HTTPS request
Everything starts the same way: your app — claude.ai, Claude Code, or your own code — sends one HTTPS POST to Anthropic’s Messages API carrying your prompt, the model name, and a single boolean: stream, set to true.
Without that flag, the API waits until the model has finished the entire response, then returns it as one JSON object. With it, the server keeps the HTTP connection open and starts pushing chunks the moment the model produces them. The request you never see looks roughly like this:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model": "claude-sonnet-4-5",
"max_tokens": 64,
"stream": true,
"messages": [{"role": "user", "content": "Say hi"}]}' \
--no-bufferOne connection, one response — but instead of one big answer at the end, a series of small events while it’s being written. That single flag is the entire difference between staring at a spinner and watching words appear.
The pause has a name: time to first token
The most mysterious part is the silence before the first word. It has a name engineers use everywhere: time to first token, TTFT.
The model can’t emit its first word until it has read and understood your entire prompt — every word you wrote, plus the system instructions and conversation history riding along. That first token is the most expensive one: all the comprehension happens up front, and each following token gets cheaper because the model is just continuing a thought it already started.
So when words start flowing fast after a slow start, that’s not the model “warming up” — it’s the shape of the work. Comprehension first, continuation after. A long pause usually means a long prompt, not a broken connection.
TTFT in one sentence: the silence is the model reading; the stream is the model talking.
SSE, not NDJSON, not WebSockets
A correction worth making precisely: the stream is not newline-delimited JSON. Anthropic uses Server-Sent Events — SSE — a small but meaningful difference.
With NDJSON, the server sends one JSON object per line and the client has to infer what each line means from its shape. With SSE, every chunk carries a named event — event: content_block_delta — followed by its data. The name tells the client exactly what arrived, no guessing. Same idea, one upgrade: labeled envelopes instead of blank ones.
And why not WebSockets? A WebSocket is a two-way phone call — great when both sides need to talk freely. But this flow is one-way: you already said everything in the request. SSE is a one-way firehose over plain HTTPS: simpler to run, automatically reconnectable, and readable natively by every browser through the EventSource API. The right tool is the simplest one that fits the shape of the conversation.
NDJSON — bare JSON lines, the client guesses what each chunk is. WebSockets — a two-way phone call for a one-way flow. SSE — named events over plain HTTP, the client always knows what arrived.
Anatomy of the stream: a tiny state machine
The stream isn’t random chunks — it’s a tiny state machine with a fixed vocabulary. Learn the six events and you can read any Claude stream like sheet music.
1. message_start — a new reply is beginning, with metadata (message id, model) and empty content. 2. content_block_start — block zero begins. A block is one unit of the reply: a run of text, a thinking block, or a tool call. 3. content_block_delta — the flood. Each carries a few characters (text_delta) or a fragment of tool-call JSON (input_json_delta); your client appends each fragment to the running text. 4. content_block_stop — block complete. 5. message_delta — final bookkeeping: why the model stopped (end_turn, max_tokens, tool_use) and the final token counts. 6. message_stop — the stream is done; close the connection.
# every event: a name, then its data
event: message_start
data: {"type":"message_start","message":{"id":"msg_01...",...}}
event: content_block_start
data: {"type":"content_block_start","index":0,...}
event: content_block_delta
data: {"type":"content_block_delta","index":0,
"delta":{"type":"text_delta","text":"Hello"}}
# ...more deltas, one per fragment...
event: content_block_stop
data: {"type":"content_block_stop","index":0}
event: message_delta
data: {"type":"message_delta",
"delta":{"stop_reason":"end_turn"},"usage":{...}}
event: message_stop
data: {"type":"message_stop"}Sprinkled through it all: ping events — keep-alives the server sends so intermediaries don’t get bored and close an idle connection. And the vocabulary can grow over time, so Anthropic’s docs advise clients to ignore event types they don’t recognize rather than crash on them.
Source: Anthropic API reference — Messages (streaming event types) ↗“Cogitating…”: honest theater
Now the spinner. In Claude Code you’ll see ✱ Cogitating… with a live timer and a token count while the model works. None of that comes from the API.
The API sends tokens, not status messages. Everything around the words — the whimsical verbs (Cogitating, Percolating, Pondering), the elapsed-time counter, the tool-call cards that fill in as JSON streams — is the client app painting UI over an open stream while it waits. The timer counts wall-clock time since your request; the token counter tallies deltas as they arrive.
That’s why it deserves the name honest theater: it doesn’t fake progress, it narrates the wait. The alternative — a dead, silent terminal for thirty seconds — would have you mashing Ctrl+C. Good streaming UX is mostly about respecting the silence between events.
One subtlety worth knowing: tool calls stream as partial JSON, and the client only executes the tool once the JSON block closes. That’s why you see the tool card assemble itself before anything runs — the app refuses to act on half a thought.
Where this matters in real work
This isn’t trivia — the streaming protocol shapes how you build with Claude. Four places it shows up:
Render each text_delta fragment as it arrives instead of waiting for the full reply. Your app feels instant even when the model is slow — perceived latency drops to TTFT, not total generation time.
message_delta carries cumulative token usage, so tools can show a running cost ticker mid-reply instead of surprising you at the end of the turn.
input_json_delta lets you preview a tool call — its name, its partial arguments — before it executes. Human-in-the-loop approval UI hooks in right here, between the deltas and the execution.
Long TTFT with fast deltas after? Your prompt is big or the model is loaded — not your network. Deltas stalling mid-stream? That’s generation or the network. The events tell you which half is slow.
If you only remember one debugging rule: the events before the first text_delta measure comprehension; everything after measures generation.
Read a raw stream yourself
The fastest way to make this concrete: open a terminal and watch the events with curl.
Run the request from the earlier section with --no-buffer so curl doesn’t hide the streaming from you, and watch the terminal: message_start, the run of content_block_delta events spelling out the reply fragment by fragment, then message_stop. Once you’ve seen the envelopes, every “AI is typing…” indicator you’ve ever used becomes legible — it was always just a client, stitching deltas, narrating the wait.
Takeaways
- One flag changes everything. stream: true turns a single silent wait into a live event flow over one HTTPS connection.
- The silence is the model reading. Time to first token is comprehension before continuation — a long pause usually means a long prompt.
- SSE, not NDJSON, not WebSockets. Named events mean no guessing; a one-way flow means the simplest fitting tool.
- Six events, one state machine. Your client stitches content_block_delta fragments into the words you see appearing.
- Spinners are honest theater. Timers, verbs, and tool cards are the app narrating the wait over an open stream — not data from the API.
Sources
All technical claims about the streaming protocol come from Anthropic’s public API documentation; the UI observations describe Claude Code’s terminal interface, and the analogies are our own.
- Anthropic API reference — Messages: streaming, event types, and delta types ↗
- MDN — EventSource: how browsers consume server-sent events natively ↗
Companion reel: What happens when you hit enter in Claude — @theclaudecraft.