# Streaming > One WebSocket for live events and replies that render as they are written — the agent stream, its frames, limits, and how it relates to long-poll and webhooks. > > Markdown mirror of https://nmbr.ai/developers/docs/streaming/ — part of the nmbr Agent API docs (https://nmbr.ai/developers/docs/). Spec: https://nmbr.ai/developers/openapi.yaml · Site index: https://nmbr.ai/llms.txt ## What the stream is `wss://nmbr.ai/ws/agent/v1/stream` is a third transport next to [long-poll and webhooks](/developers/docs/events/): the same events, over one socket, with no polling latency — and one thing the others can't do: your agent can **stream a reply**, so the person watches the answer appear in their chat as your model produces it, the way they would with an assistant in a browser. Authenticate at the upgrade with the agent token, either as `?token=agent:…` or an `Authorization: Bearer agent:…` header. Every check that guards an HTTP call runs here too (a paused agent or a revoked token gets the same `403` / `401` and the usual `{ "error": { "code", "message" } }` body instead of a socket). Optional `?afterSeq=` replays events after that cursor; without it you start from now. ## Frames Both directions are JSON text frames of at most 64 KB. **Server → agent** | Frame | Meaning | |---|---| | `hello { seq }` | Connected; `seq` is your cursor. | | `event { event }` | An event, in the same envelope as `GET /updates` (`id`, `seq`, `type`, `ts`, `agentId`, `payload`). | | `pong` | Answer to your `ping`. | | `reply.started { streamId, messageId, conversationId }` | Your reply's bubble is on the person's screen. | | `reply.ended { streamId, messageId, message }` | Your reply is final; `message` is the stored message. `aborted: true` when withdrawn, `idle: true` when nmbr ended it for you. | | `error { streamId?, error: { code, message } }` | A frame was refused. With `streamId`, it concerns that reply. | **Agent → server** | Frame | Meaning | |---|---| | `ping` | Keep-alive; nmbr also pings you every 30 s. | | `typing { conversationId, typing }` | Show or clear the typing indicator (needs their `messages:write`). | | `reply.start { streamId, conversationId \| to }` | Open a reply. `streamId` is yours (≤ 64 chars, unique on this socket). | | `reply.delta { streamId, text }` | Append text. The person sees it at once. | | `reply.end { streamId }` | Finish the reply. | | `reply.abort { streamId }` | Withdraw it — the bubble disappears. | ## A streamed reply, step by step 1. `reply.start` passes the same rules as `POST /messages`: the person added your agent, granted `messages:write`, no block, and their replies-only window is open. A refusal is an `error` with the usual code (`not_a_contact`, `scope_not_granted`, `reply_window_closed`, `blocked`) and is audited like any refused send. It counts as **one message** for the per-minute limit and the daily ceiling, however long it streams. 2. nmbr posts a placeholder bubble to the person's devices (no push yet) and answers `reply.started`. 3. Each `reply.delta` is pushed to their screen immediately and written to the database at most once a second, so a reload mid-stream shows the text so far. 4. `reply.end` writes the final text, sends the push notification once, and settles the bubble. `reply.abort` (or ending with nothing written) removes it. ```ts const stream = agent.stream(); await stream.connect(); for await (const event of stream.events()) { if (event.type !== "message.received" || !event.payload.message) continue; const reply = stream.reply({ conversationId: event.payload.conversationId }); for await (const token of llm.stream(event.payload.message.content)) reply.write(token); await reply.end(); } ``` The SDK's `agent.stream()` uses the platform `WebSocket` (Node 22+, Deno, Bun, browsers); on older Node pass `{ webSocket: WebSocket }` from the `ws` package. Batch tokens if your model emits more than 20 per second — nmbr refuses deltas above that rate (`rate_limited`) rather than flooding the person's device. ## Limits 3 connections per agent · 120 inbound frames per minute per connection · 20 deltas per second per reply · 3 open replies per connection · 10,000 characters per reply (the reply is ended for you at the cap) · a reply with no delta for 60 seconds is ended with the text so far · a connection that drops ends its open replies the same way. Over the cap the socket closes with code `1008` (rate) or `1009` (frame size). ## Reconnecting Nothing is lost: events live in the same queue the long-poll reads. Keep `lastSeq` (the SDK does) and reconnect with `?afterSeq=` to resume; replies in flight when a socket dropped were already ended by nmbr, so start a new one.