Streaming
Runs stream Server-Sent Events (SSE) by default. The SDK wraps these in a RunStream iterator so you can process events as they arrive.
Quick Patterns
Just the text, live:
Text + post-run data (run_id, full text). Use the with context manager so the connection closes cleanly:
Just the final output (no streaming):
Follow-up replies (runs.reply()) stream the same way and inherit the run's settings. See Follow up on a run.
Event Types
These are the event.type values the Python SDK yields (it normalizes the raw
Claude-native SSE frames into these friendlier names; see SSE Frame Format
for the raw wire if you consume the HTTP stream directly).
Filter events by class
Detecting Failures
A run can fail mid-stream (expired credential, model rate limit, quota). The default iter_text() / stream.text path drops error events, so a failed run can otherwise look like an empty success. Opt into raising, or check after iterating:
Reconnecting to a dropped stream
If the connection drops mid-run (proxy idle-timeout, network blip), rejoin with the run_id from the metadata event. runs.stream(run_id) replays the run's full history then live deltas, so reset any local accumulation on reconnect; it returns 409 once the run is no longer executing (use runs.get(run_id) for the result):
The server emits a 15s keepalive on the streaming path so a long-silent tool call doesn't trip the read timeout; raise it for very long runs with M8tes(timeout=...).
SSE Frame Format
If you consume the HTTP stream directly (your own SSE client), the wire carries
flat, Claude-native frames: text rides inside content_block_delta and tools inside
content_block_start / tool_use / tool_result. The Python SDK translates these
into the friendlier event.type names in the Event Types table above;
if you're using an SDK you never see the raw frames. The @m8tes/react
client does the same normalization in the browser.
Raw wire example
data: {"type": "metadata", "run_id": 123, "mode": "chat", "sandbox_enabled": true}
data: {"type": "content_block_start", "id": "block_1", "block_type": "text"}
data: {"type": "content_block_delta", "id": "block_1", "delta": {"type": "text_delta", "text": "Here are "}}
data: {"type": "content_block_delta", "id": "block_1", "delta": {"type": "text_delta", "text": "the open tickets..."}}
data: {"type": "content_block_stop", "id": "block_1"}
data: {"type": "tool_use", "id": "tc_1", "name": "gmail_search", "input": {"query": "is:open"}}
data: {"type": "tool_result", "tool_use_id": "tc_1", "content": "...", "result": "..."}
data: {"type": "run_metrics", "execution_time_ms": 4200, "input_tokens_used": 1200, "completion_state": "complete"}
data: {"type": "done", "completion_state": "complete", "stop_reason": "end_turn", "message_count": 4}
Delegation (subagents) in the stream
An agent may delegate part of a job to a subagent: a worker that runs in its own context and reports a result back. You will see this as a tool call named Agent:
data: {"type": "tool_use", "id": "tc_2", "name": "Agent", "input": {"description": "Pull last week's spend", "subagent_type": "general-purpose", "prompt": "..."}}
data: {"type": "tool_result", "tool_use_id": "tc_2", "content": "...the subagent's report..."}
- Treat it like any other tool call. The subagent's own steps are not streamed; you get the delegation and its result. Nothing about your event handling has to change.
- Its tokens are part of the run. Delegated work is included in the run's cost and counts against the run's spending limits.
run_metricscarriessubagent_countandsubagent_tokensso you can see how much of a run was delegated. - Permissions still apply. A subagent's tool calls go through the same approval rules as the agent's own. If a tool needs a human, it still asks, and the approval says which subagent asked for it.
The Python SDK normalizes these into the events (text-delta, tool-call-start, tool-result-end) shown in the table above.
Next: Runs · Human-in-the-Loop · Webhook Events
