AI + DevelopmentBuild
Building a Streaming API with Node.js and TypeScript
What Netflix-scale video streaming teaches us about streaming data, from LLM tokens to live events.
Nobody waits for the whole film
Two versions of the same AI feature. The first sends a question, shows a spinner for twenty seconds, then drops a full answer on the screen. The second shows the first words in under a second and keeps writing while you read.
Both do the same work and take the same total time. Users will tell you the second one is faster, and in the way that matters, they're right. People judge speed by the first byte, not the last.
Video platforms learned this long ago. Nobody downloads a whole film before pressing play. The video starts almost immediately and keeps arriving in pieces while you watch. That idea, plus the engineering that makes it reliable, is exactly what you need when you build a streaming API in Node.js, whether you are streaming LLM tokens, logs or live events.
This article takes the principles behind large-scale video streaming and turns them into a working Node.js and TypeScript streaming API, then covers what breaks when it reaches production.
What "streaming" actually means in an API
A streaming API sends a response in pieces as they become ready, instead of building the whole thing first. On the web there are four common ways to do it:
| Technology | Direction | Best for | Trade-offs |
|---|---|---|---|
| Chunked HTTP response | Server to client | Large downloads, NDJSON, simple progressive output | No standard event format; the client must parse the stream itself |
| Server-Sent Events (SSE) | Server to client | LLM token streaming, notifications, live feeds | One-way only; text-based; the browser EventSource API supports GET requests only |
| WebSockets | Both directions | Chat, collaboration, games | Separate protocol; harder to proxy, cache and scale |
| HTTP range requests | Client pulls byte ranges | Video and audio playback, resumable downloads | Designed for files, not for live data |
For streaming LLM responses, SSE is usually the right default. It runs over plain HTTP, works through most proxies once configured, has a simple event format, and supports event IDs that let a client tell the server where it left off. You only need WebSockets when the client must keep sending data on the same connection while the server streams back.
What Netflix-scale streaming teaches a Node.js developer
You are not going to build Netflix. But the problems video platforms solved, at enormous scale, are the same ones every streaming API runs into: unreliable networks, slow clients, expensive bandwidth and connections that drop halfway.
Video is sent in small segments, not as one file. Adaptive bitrate streaming splits video into short segments, each encoded at several quality levels. The player fetches one segment at a time and picks the quality of the next one based on how quickly the last ones arrived. Formats such as HLS and MPEG-DASH standardise this.
Netflix moves the content close to the viewer. Its video is delivered through Open Connect, Netflix's own content delivery network, started in 2011. Netflix places its Open Connect Appliances in internet exchange points and, for qualifying providers, directly inside internet service providers' networks, so most video travels a short distance to the viewer.
Netflix sends only the bits a title actually needs. In 2015 Netflix described per-title encoding: instead of one fixed set of bitrates for every title, each title gets its own bitrate ladder based on how complex it is. A simple animated film doesn't need the bits an action film with fast motion does.
Those ideas translate directly to a data streaming API:
| Video streaming principle | What it means for your streaming API |
|---|---|
| Small segments | Send data as small, self-contained events instead of one large response |
| Adaptive delivery | Adapt to the client: slow consumers should slow the producer down, not overflow it |
| Content close to the user | Serve static and repeatable content from a CDN or cache; stream only what is unique to this request |
| Send only what's needed | Send compact events; don't stream data the client will throw away |
| Buffering and backpressure | Never produce faster than the network and client can consume |
| Resumable playback | Give events IDs so a client can reconnect and continue where it stopped |
| Graceful degradation | A bad network should degrade the experience, not break the session |
The rest of this article builds a streaming API that follows these principles, using LLM token streaming as the example, because that is where most Node.js developers meet streaming today.
The architecture
Two views of the same system: where the parts sit, and what happens over time on one connection.
mermaid
flowchart LR
C[Client<br/>browser or app] --> E[CDN / Edge]
E -->|static assets,<br/>cached answers| C
E --> G[API Gateway<br/>auth, rate limits]
G --> S[Node.js<br/>streaming service]
S -->|stream enabled| L[LLM provider]
L -.->|token chunks| S
S -.->|SSE events| CThe edge serves everything that is the same for every user. Only the unique, per-request part travels the full path, which is the Open Connect idea applied to an API.
mermaid
sequenceDiagram
participant C as Client
participant S as Node.js service
participant L as LLM provider
C->>S: POST /api/chat/stream
S-->>C: 200 text/event-stream (headers flushed)
S->>L: request with streaming enabled
loop each chunk
L-->>S: token chunk
S-->>C: event: token (id: n)
end
Note over S,C: heartbeat comment every 15s while idle
alt client disconnects
C--xS: connection closed
S->>L: abort request
else generation completes
S-->>C: event: done
endNotice the abort branch. When the user closes the tab, the server stops the upstream request, so you stop paying for tokens nobody will read.
Building the streaming API in Node.js and TypeScript
The server below streams an LLM response over SSE with Express. The model call sits behind a small provider-neutral type, so you can plug in whichever SDK you use, as long as it exposes the response as an async iterable of text chunks and accepts an AbortSignal. Most major LLM SDKs offer both, but check yours.
import express, { Request, Response } from "express";
import { once } from "node:events";
// Wrap your provider's SDK so it yields text chunks and honours the abort signal.
type ModelStream = (prompt: string, signal: AbortSignal) => AsyncIterable<string>;
declare const streamFromModel: ModelStream;
const app = express();
app.use(express.json());
async function send(res: Response, signal: AbortSignal, event: string, data: unknown, id?: number) {
const frame = (id !== undefined ? `id: ${id}\n` : "")
+ `event: ${event}\ndata: ${JSON.stringify(data)}\n\n`;
// Backpressure: if the socket buffer is full, wait for 'drain' (or give up on abort).
if (!res.write(frame)) await once(res, "drain", { signal });
}
app.post("/api/chat/stream", async (req: Request, res: Response) => {
res.writeHead(200, {
"Content-Type": "text/event-stream",
"Cache-Control": "no-cache, no-transform", // no-transform: proxies must not compress or alter it
Connection: "keep-alive",
"X-Accel-Buffering": "no", // tell nginx not to buffer this response
});
res.flushHeaders(); // the client knows the stream is open before the first token
const upstream = new AbortController();
res.on("close", () => upstream.abort()); // client left: stop paying for tokens
const heartbeat = setInterval(() => res.write(": ping\n\n"), 15_000); // keeps idle proxies from closing
let id = 0;
try {
for await (const chunk of streamFromModel(req.body.prompt, upstream.signal)) {
await send(res, upstream.signal, "token", { text: chunk }, ++id);
}
await send(res, upstream.signal, "done", { chunks: id });
} catch (err) {
if (!upstream.signal.aborted) {
// Headers are already sent, so errors travel as events, not status codes.
await send(res, upstream.signal, "error", { message: "Generation failed" }).catch(() => {});
}
} finally {
clearInterval(heartbeat);
res.end();
}
});
app.listen(3000);
Each line maps back to a streaming principle:
- Small segments: every chunk becomes its own SSE event with an
id. - Backpressure:
res.write()returnsfalsewhen the socket buffer is full. Waiting fordrainmeans a slow client slows the producer instead of filling the server's memory. - Graceful degradation: errors after the first byte are sent as an
errorevent, because the HTTP status has already gone out. - Don't send what nobody needs: the
closehandler aborts the upstream call, and the abort signal ononcestops the server waiting for adrainthat will never come.
On the client, the browser's EventSource only supports GET requests without a body, so for a POST with a prompt, read the stream with fetch:
const controller = new AbortController(); // wire this to a Stop button
const res = await fetch("/api/chat/stream", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ prompt: "Explain backpressure in one paragraph" }),
signal: controller.signal,
});
const reader = res.body!.pipeThrough(new TextDecoderStream()).getReader();
let buffer = "";
for (;;) {
const { value, done } = await reader.read();
if (done) break;
buffer += value;
const events = buffer.split("\n\n");
buffer = events.pop() ?? ""; // keep the incomplete event for the next read
for (const e of events) handleEvent(e); // parse 'event:', 'id:' and 'data:' lines
}
This parser is deliberately minimal. In production, use a tested SSE parsing library, since events can arrive split across network reads in more ways than this handles.
What breaks in production
Streaming works on localhost on the first try. Then it meets real infrastructure.
Proxies buffer the response. A reverse proxy may collect the whole response before forwarding it, which turns your stream back into a twenty-second wait. In nginx, disable proxy_buffering for the streaming route, or send the X-Accel-Buffering: no header as the code above does. Check your load balancer and API gateway for the equivalent setting.
Compression delays chunks. Compression middleware can hold data until it has enough to compress efficiently. Exclude text/event-stream responses from compression.
Idle timeouts close long streams. Load balancers and gateways often close connections that look idle. The heartbeat comment keeps traffic flowing; make sure its interval is shorter than the shortest idle timeout in the path.
Reconnects need a replay plan. EventSource reconnects automatically and sends the last event ID it received in a Last-Event-ID header; with fetch you send it yourself. But you can't ask an LLM to resume a half-finished generation. If resuming matters, write each event to a short-lived store (for example a Redis stream keyed by response ID) as you send it, and replay from the requested ID. That is the resumable-playback idea, and it costs storage, so do it only where it's worth it.
Browsers limit connections. Over HTTP/1.1, browsers allow only a few concurrent connections per domain, and MDN notes that this limit applies to EventSource too, so several open tabs can block each other. Serve streams over HTTP/2, which multiplexes many streams over one connection.
Long-lived connections change scaling. Each open stream holds a connection on one instance. Plan capacity by concurrent streams, not requests per second. If any instance needs to push events to any client, put a pub/sub layer between producers and the instances holding the connections.
Measure the right things. Track time to first token (the number users feel), chunks or tokens per second, total stream duration, client disconnects and aborted generations. A spike in aborts often means the answer is too slow or too long, not that users are impatient.
When not to stream
Streaming adds connection management, proxy configuration and harder error handling. Skip it when:
- The response is short. A two-line answer gains nothing from streaming.
- The output must be validated as a whole first. Strict JSON for another system, or an answer that must pass a policy check before anyone sees it, can't be shown half-finished.
- Nobody is watching. Batch jobs and background pipelines want the complete result, not a progress feed.
Streaming is a promise, not a transport trick
Netflix doesn't make films shorter; it makes sure you never wait for parts of the film it has already prepared. A good streaming API makes the same promise: the user never waits for work the system has already done.
Everything in this article (small events, backpressure, aborting work nobody needs, resumable IDs, an edge that serves what's already known) is how you keep that promise once real networks, proxies and impatient users get involved. Build your streaming API in Node.js around those principles, and the transport becomes the easy part.