Skip to content

Streaming ​

Group: Inference & Interface | Previous group exit: can write a model-calling loop with error-family classification, retry semantics, and usage observability | This topic exit: can consume one model stream — accumulate chunks, cancel at any time, recognize interruption, and read the finish reason Prerequisites: Model API Contract | Next: Session and State, Generative UI

1. Overview ​

Streaming solves the problem of perceived latency. Generating an answer often takes seconds; waiting for the complete payload shows the user a spinner over blank space. Streaming brings Time To First Token (TTFT) under a second with the rest arriving progressively — the user sees work happening, which changes the waiting experience completely.

One streaming interaction consists of four things: transport (SSE, Server-Sent Events), accumulation (chunks joined into full content), cancellation (AbortController propagating so the server stops generating), and finish semantics (finish_reason tells you whether the ending is natural, truncated, or a refusal).

When to use / when not to ​

  • Use: human-facing generation (chat, writing, code); unpredictable generation length with a waiting user.
  • Do not use: background batch jobs (whole-payload handling is simpler); structured output that must be validated before display (see Structured Output; the two can combine — stream the transport, buffer, then validate); moderation-sensitive production traffic — partial completions are harder to moderate (per OpenAI's official note, retrievedAt 2026-09-01).

Decision table: SSE vs WebSocket vs polling ​

ApproachDirectionControlStateTrust domainMinimum complexity
SSE (HTTP stream)one-way: server → clientstandard HTTP; browser auto-reconnects (EventSource)connectionless semanticsplain HTTP authlowest: response header + data lines
WebSocketbidirectionalyou own protocol, heartbeats, reconnectoptionally statefulupgraded protocol, extra gateway configmedium: worth it only for two-way needs
Pollingclient pullsfully manualserver stores resultsplain HTTPlow, but poor latency and wasted requests

Default to SSE: model streams are pure one-way pushes over ordinary HTTP infrastructure; move to WebSocket only when you must send input mid-generation (e.g., real-time barge-in). OpenAI Responses now ships an official WebSocket mode for exactly that incremental input (retrievedAt 2026-09-01) — it solves "send input mid-generation", not resuming an interrupted generation (see Principles); SSE remains the default HTTP streaming path.

Historical milestones ​

SSE is an old WHATWG HTML-spec technology (predating LLM apps); vendors define their own streaming event vocabularies — OpenAI Chat Completions uses data: JSON chunks plus a [DONE] sentinel, the Responses API switched to typed semantic events (response.output_text.delta and friends), Anthropic uses lifecycle message events (message_start → content_block_delta → message_stop). Event lists are governed by the vendor docs of the day (retrievedAt 2026-09-01).

2. Usage ​

Minimal hands-on: zero-key stream consumption and cancellation (≤15 minutes) ​

A local mock serves an OpenAI-style SSE stream; the client demonstrates full consumption and mid-stream cancellation. Environment: Node ≥ 23.6 (add --experimental-strip-types on 22.6–23.5). Save as streaming-mock.mts:

ts
// fixture: zero-key, zero-dependency SSE consumption and cancellation. Deterministic output.
import * as http from 'node:http';

// ---- mock server: OpenAI-style SSE stream, deterministic chunks ----
const TOKENS = ['流式', '传输', '把', '等待', '变成', '逐步', '呈现', '。'];
const encoder = new TextEncoder();
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));

const server = http.createServer(async (req, res) => {
  if (req.url !== '/v1/chat/stream') { res.writeHead(404).end(); return; }
  res.writeHead(200, {
    'Content-Type': 'text/event-stream',   // the SSE MIME type
    'Cache-Control': 'no-cache',
  });
  let finished = false;
  req.on('close', () => {
    if (!finished) console.log('[server] client disconnected -> generation stopped');
  });

  for (let i = 0; i < TOKENS.length; i++) {
    if (res.destroyed) return;             // stop generating right after cancellation
    const chunk = { choices: [{ delta: { content: TOKENS[i] }, finish_reason: null }] };
    res.write(encoder.encode(`data: ${JSON.stringify(chunk)}\n\n`));
    await sleep(80);
  }
  finished = true;
  const done = { choices: [{ delta: {}, finish_reason: 'stop' }], usage: { prompt_tokens: 10, completion_tokens: 8, total_tokens: 18 } };
  res.write(encoder.encode(`data: ${JSON.stringify(done)}\n\n`));
  res.write(encoder.encode('data: [DONE]\n\n'));
  res.end();
});

await new Promise<void>((resolve) => server.listen(0, '127.0.0.1', resolve));
const base = `http://127.0.0.1:${(server.address() as { port: number }).port}`;

// ---- client: SSE parsing + chunk accumulation + AbortController cancellation ----
interface StreamResult { text: string; finishReason: string | null; aborted: boolean }

async function consumeStream(url: string, signal?: AbortSignal): Promise<StreamResult> {
  const res = await fetch(url, { signal });
  if (!res.ok || !res.body) throw new Error(`HTTP ${res.status}`);
  const reader = res.body.getReader();
  const decoder = new TextDecoder();
  let buffer = '';
  let text = '';
  let finishReason: string | null = null;

  try {
    while (true) {
      const { value, done } = await reader.read();
      if (done) break;
      buffer += decoder.decode(value, { stream: true });   // a chunk may split a line; buffer first
      let sep: number;
      while ((sep = buffer.indexOf('\n\n')) >= 0) {        // events end with a blank line
        const rawEvent = buffer.slice(0, sep);
        buffer = buffer.slice(sep + 2);
        for (const line of rawEvent.split('\n')) {
          if (!line.startsWith('data:')) continue;          // handle only data: lines
          const data = line.slice(5).trim();
          if (data === '[DONE]') return { text, finishReason, aborted: false };
          const chunk = JSON.parse(data) as {
            choices: { delta: { content?: string }; finish_reason: string | null }[];
          };
          text += chunk.choices[0]?.delta.content ?? '';   // accumulate
          if (chunk.choices[0]?.finish_reason) finishReason = chunk.choices[0].finish_reason;
        }
      }
    }
    return { text, finishReason, aborted: false };
  } catch (e) {
    if (signal?.aborted) return { text, finishReason, aborted: true };  // cancelled: keep partial output
    throw e;
  }
}

console.log('--- full consumption ---');
const full = await consumeStream(`${base}/v1/chat/stream`);
console.log('text:', full.text);
console.log('finish_reason:', full.finishReason);

console.log('--- cancel after 320ms ---');
const controller = new AbortController();
setTimeout(() => controller.abort(), 320);
const partial = await consumeStream(`${base}/v1/chat/stream`, controller.signal);
console.log('text:', partial.text);
console.log('aborted:', partial.aborted, '| finish_reason:', partial.finishReason);

await sleep(50);
server.close();

Run and normal output:

text
$ node streaming-mock.mts
--- full consumption ---
text: 流式传输把等待变成逐步呈现。
finish_reason: stop
--- cancel after 320ms ---
text: 流式传输把等待
aborted: true | finish_reason: null
[server] client disconnected -> generation stopped

The negative output is the second section: after cancellation aborted: true and finish_reason is empty — partial output has no finish reason, so the UI must mark it "interrupted" rather than treat it as complete. The server log proves generation actually stopped (that is the evidence of "cancellable"). Acceptance: both paths match the output above.

Cleanup: delete the file; the mock listens on a random 127.0.0.1 port, released on process exit.

Scenario matrix ​

ScenarioInputActionOutputFitsDoes not fit
Basic: progressive displayan SSE streamaccumulate deltas + renderfull textchat interfacesJSON needing whole-payload validation
Common: stop buttonuser clickAbortController.abort()partial text + interrupted markerany generation UI—
Combined: streaming + structurea stream of component JSONbuffer per line + validate, then renderwhitelist componentsGenerative UIstrict low-frequency validation (plain non-streaming)

3. Principles ​

SSE wire format (WHATWG-spec behavior, retrievedAt 2026-09-01, MDN) ​

  • Response headers: Content-Type: text/event-stream, paired with Cache-Control: no-cache.
  • The body is a UTF-8 text stream; each event ends with a blank line (\n\n).
  • Each line is field: value; four fields are recognized: data (payload), event (event name), id (resume marker), retry (reconnect milliseconds).
  • Multiple consecutive data: lines are joined with newlines; a leading : makes a comment line, usable as a keep-alive.
  • The EventSource API fires onmessage for unnamed events and addEventListener for event:-named ones; it reconnects automatically (tunable via retry), and close() terminates.
  • Connection limits: over HTTP/1.x a browser allows at most 6 concurrent SSE connections per origin (marked Won't fix in Chrome/Firefox); HTTP/2 negotiates its stream limit (default 100).

Model-API streams almost universally use data: lines carrying JSON; vendors differ in the event vocabulary:

ConceptOpenAI Chat CompletionsOpenAI ResponsesAnthropic Messages
Text deltachoices[0].delta.contentdelta on response.output_text.deltatext_delta on content_block_delta
Finish reasonfinish_reason on the last chunkresponse.completedstop_reason carried by message_delta
End sentineldata: [DONE]the event type itselfmessage_stop
First chunkdelta containing only the roleresponse.createdmessage_start

Finish reason semantics ​

finish_reason/stop_reason answers "why did this stop", which decides the next action:

  • Natural end (stop / end_turn): content is complete; safe to write into session history.
  • Length cap (length / max_tokens): content truncated; continue or inform the user.
  • Stop sequence (stop_sequence): hit a stop string you configured.
  • Tool call (tool_calls / tool_use): hand off to tool execution (see Tool Calling Contract).
  • Safety refusal (Anthropic refusal): content may be a refusal notice; check before displaying.
  • Cancellation (abort): no finish reason — it is a client action, not a server semantic; the UI must mark it.

Backpressure and render throttling ​

Network arrival speed and rendering speed are decoupled. Calling setState per chunk at high stream rates overwhelms the render pipeline. The fix is to separate accumulation from rendering — flush the buffer once per frame (requestAnimationFrame) or at a fixed interval. The Node-side equivalent is the ReadableStream reader loop, which naturally consumes chunk by chunk without piling up.

Interruption and recovery ​

  • A broken transport is not a generation semantic. A fetch stream does not auto-reconnect (only EventSource does); keep the partial content received so far and mark it "incomplete".
  • Generation cannot be resumed. There is no universal "continue from token N" interface; recovery means a new request built around the partial content (or an explicit user retry). EventSource's Last-Event-ID reconnection is meaningful only for replayable event sources, not one-shot generation.
  • Proxy buffering is the number-one enemy. Reverse proxies such as Nginx buffer responses by default, collapsing the stream into one block. Servers need X-Accel-Buffering: no (or the equivalent) and should flush response headers early.

Spec vs local test ​

ClaimSpec/official docsLocal mock test (fixture above)
Events end with a blank line; data: carries payloadWHATWG/MDN (L0)verified: parser splits on \n\n
A chunk may split a line; buffer and joinMDN streaming semantics (L0)verified: decode with { stream: true } + buffer
The server can sense connection close after abortHTTP connection semanticsverified: req.on('close') fires and generation stops
A cancelled stream has no finish_reasonvendor streaming docs, inferredverified: aborted: true with null finishReason
HTTP/1.x 6-connections-per-origin capMDN (L0)not tested (single-connection scenario), see open questions

4. Development ​

Integration and compatibility ​

  • Prefer fetch + AbortController on the client (POST-friendly); EventSource is GET-only, suited to simple read-only streams.
  • Flush headers early on the server; when a CDN/gateway sits in the path, verify "chunks arrive progressively", not "one block arrives".
  • Memoize historical messages in the render layer and update only the streaming one, avoiding whole-list re-renders.

Debug runbooks ​

Symptom → Evidence → Action → Done when ​

Symptom: the frontend displays everything only after generation completes. Evidence: curl -N <endpoint> shows chunked vs one-block arrival; response headers hint at an intermediate buffer. Action: add X-Accel-Buffering: no / disable gzip / flush early on the server; confirm the intermediate layers pass streams through. Done when: curl -N prints tokens progressively; browser TTFT lands under a second.

Symptom → Evidence → Action → Done when ​

Symptom: the user clicked stop, but billing shows generation continued. Evidence: server logs show generation-loop output after the abort. Action: listen for req.on('close') and forward the cancellation upstream (pass the AbortSignal to the vendor API), stopping on receipt. Done when: server generation logs stop immediately after abort; reconciliation shows no over-billing.

Symptom → Evidence → Action → Done when ​

Symptom: an interrupted answer is stored as complete, and the model continues from the half sentence next turn. Evidence: assistant messages without a finish_reason exist in session history. Action: check aborted/finishReason before writing back; mark truncated messages or do not persist them. Done when: every assistant message in history carries an explicit end state.

Symptom → Evidence → Action → Done when ​

Symptom: opening a few more tabs makes all new connections fail. Evidence: browser console connection errors; HTTP/1.x with 6 open same-origin connections. Action: upgrade to HTTP/2, or multiplex over a single event channel. Done when: multiple tabs stream concurrently.

Anti-pattern list ​

  • Calling JSON.parse on each chunk of incomplete JSON — buffer per event first.
  • Aborting only in the frontend UI layer without server propagation — generation and cost continue.
  • Silently swallowing a broken stream and rendering it as a normal ending — success-looking but missing the finish reason.
  • Calling setState on every network chunk — render jitter at high stream rates.
  • Ignoring finish_reason and persisting length-truncated content as complete.

5. Resource Library ​

Four-level reading route ​

  • Beginner (2): MDN Using server-sent events (wire format and EventSource); run this page's fixture to exercise cancellation.
  • Builder (2): the OpenAI streaming guide (delta chunk structure); the Anthropic streaming event reference.
  • Operator (2): proxy/CDN configuration checks for the streaming path (X-Accel-Buffering and friends); add TTFT metrics per Observability.
  • Researcher (2): the WHATWG HTML spec section on event-stream syntax; HTTP/2 stream multiplexing material.

Resource table ​

NameLevelCanonical URLUseSupported claimNext
MDN Using SSEL0https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_eventswire-format authorityfields/blank-line separation/reconnect/6-connection capwrite a parser
OpenAI Streaming guideL0https://developers.openai.com/api/docs/guides/streaming-responsesvendor streaming overviewResponses semantic events, moderation noteconnect a real vendor
Anthropic Streaming referenceL0https://docs.anthropic.comevent lifecyclemessage_start/delta/stop event familyconnect a real vendor
WHATWG HTML specL0https://html.spec.whatwg.org/multipage/server-sent-events.htmlsyntax normevent-stream grammar and parsing rulesconsult for exact definitions
This page's fixtureEstreaming-mock.mts (inline)zero-key verificationaccumulation/cancellation/finish-reason behaviorswap in a real vendor endpoint

retrievedAt: all web resources 2026-09-01.

Active falsification and open questions ​

  • "Sub-second TTFT meaningfully improves the experience" is an engineering consensus; your product's threshold needs its own measurement.
  • The 6-connection cap was not reproduced in the fixture (needs a multi-tab browser environment); the claim is from MDN, marked L0.
  • Vendor event vocabularies follow their official docs; this page's table lists only verified fields.

Where learn-ai stops / where to go next ​

This page owns "progressive arrival". Sending and trimming multi-turn history → Session and State; rendering components inside the stream → Generative UI; TTFT/throughput measurement → Cost and Performance; framework streaming wrappers → the Products framework pages.

Built for frontend engineers · Powered by VitePress