LLMsBuild

Understanding Context Windows: How LLMs Actually Use Information

A context window is the model's whole workspace for one request. Here's what fills it, why models miss what's buried inside, and how to manage it in real applications.

By Faizan Syed7 min read98 views

What is an LLM context window?

A context window is the maximum number of tokens a model can process in one request, counting both what you send and what it generates. It is per request. The model keeps nothing between calls, so anything it should know has to be inside the window every time.

In a typical chat or agent application, the window holds:

text
  System instructions
+ Tool definitions
+ Conversation history
+ Retrieved documents
+ Tool results
+ The user's message
+ The model's output (including reasoning tokens on some models)
= Must fit inside the context window

The exact composition depends on the API. Providers wrap messages in their own formatting tokens, and some store conversation state for you. The principle holds everywhere: if something influences the answer, it took up space in the window.

These terms get used interchangeably, but they mean different things:

TermWhat it is
Context windowThe model's token limit for one request
PromptThe instructions you write: one part of the context
Conversation historyEarlier turns you choose to send again
Model memoryInformation stored outside the model and loaded into the context when relevant
RAGA way to choose external documents to put into the context
Context engineeringDeciding what goes into the window on every call

Tokens: the unit behind the limit

Context windows are measured in tokens, not characters or words. A tokenizer splits text into units from a fixed vocabulary. Common words often become one token. Rarer words, identifiers and long numbers are split into pieces.

text
"Understanding context windows"  ->  a few tokens, not one string and not three neat words
"ctx_window_cfg_v2"              ->  more tokens than its length suggests

(Illustrative. The exact split depends on the tokenizer.)

OpenAI's docs give a rough English rule of about four characters per token, but treat that as napkin maths. Code, JSON and many non-English languages usually need more tokens, and each model family has its own tokenizer, so the same text can produce different counts on different models. For budgets and cost, count with the model's own tokenizer or a provider counting endpoint, and trust the usage figures the API returns.

Input and output share one budget

The classic example: send 195k tokens to a model with a 200k window, and only about 5k remain for the answer. That's true as far as it goes, but two more constraints apply.

  • Models usually have a separate output cap, often much smaller than the window. If that cap is lower than the space you left, the cap wins.
  • APIs handle overflow differently. Some reject a request when input plus the requested maximum output exceeds the window. Others accept it and stop the answer early. On reasoning models, hidden reasoning tokens can use up output budget before the visible answer begins.

Check your provider's docs for the exact behaviour, and check the finish reason on every response: a "length" stop means the answer was cut off. Then reserve output space on purpose:

ts
interface ModelLimits {
  contextWindow: number;   // from the provider's model docs, kept in config
  maxOutputTokens: number; // the model's own output cap
}

function inputBudget(limits: ModelLimits, desiredOutput: number, margin = 0.05) {
  const output = Math.min(desiredOutput, limits.maxOutputTokens);
  const safety = Math.ceil(limits.contextWindow * margin); // formatting tokens, estimate error
  return limits.contextWindow - output - safety;
}

Models don't use long context evenly

A model that accepts 200k tokens doesn't necessarily use 200k tokens well. In Lost in the Middle (Liu et al., TACL 2024), models answered most accurately when the relevant document sat at the start or end of the input, and worse when it sat in the middle. Chroma's 2025 Context Rot report found performance becoming less reliable as input length grew, even on simple tasks. Newer models have improved, so treat these as tendencies to test, not fixed rules.

What that means in practice:

  • Keep stable instructions at the start, and put the user's actual question near the end, after long documents.
  • Order retrieved material by relevance, and drop weak matches instead of padding.
  • Never assume "it's in the context" means "the model used it". Test with real questions.

That's the answer to the opening puzzle. The document was present, but buried among less relevant tokens, in a position the model used less reliably.

Bigger windows cost more, but they aren't always worse

Input tokens are billed, and the model has to process the whole input before it produces the first output token. Long contexts therefore raise both cost and time to first token. Prompt caching, offered by several providers, can cut both for a repeated, stable prefix, but it doesn't make irrelevant tokens useful.

More context isn't automatically wrong, though. When a task needs the whole artifact (a full contract, a long log, several related files), long context can beat retrieved fragments that miss the connections between them. Research comparing the two approaches (Xu et al., 2023) found that retrieval improved results even for long-context models, so they work best together, not as rivals.

The useful question isn't "does it fit?" It's "does every part earn its tokens?"

Managing context in long conversations

History is the part of the window that grows on its own. Four common strategies, often combined:

StrategyHow it worksGood forTrade-off
Sliding windowKeep the last N turns verbatimShort tasks where recent detail mattersOlder decisions are forgotten entirely
Running summaryReplace older turns with a summaryLong sessionsSummaries lose detail and can drift
Retrieval over historyStore turns, fetch only relevant onesAssistants used over days or weeksAdds an embedding and search step
Structured stateKeep facts (plan, preferences, IDs) in an object you controlAgents and workflowsYou must design and maintain the schema

Here's a minimal version that combines structured state, a summary and the most recent turns within a token budget:

ts
type Turn = { role: "user" | "assistant"; content: string };

function buildHistory(
  turns: Turn[],
  summary: string,
  state: Record<string, unknown>,
  budget: number,
  count: (text: string) => number, // the model's tokenizer
): string[] {
  const fixed = [`Known facts: ${JSON.stringify(state)}`];
  if (summary) fixed.push(`Earlier in this conversation: ${summary}`);
  let remaining = budget - fixed.reduce((n, p) => n + count(p), 0);

  const recent: string[] = [];
  for (const t of [...turns].reverse()) {   // newest first
    const line = `${t.role}: ${t.content}`;
    if (count(line) > remaining) break;     // older turns live in the summary
    recent.unshift(line);
    remaining -= count(line);
  }
  return [...fixed, ...recent];
}

Choose by what the feature must remember. A support bot needs the current issue and the customer's plan, not last month's chat. A coding assistant needs the files being edited and the error message, not the whole repository.

Where RAG and context engineering fit

RAG and history management answer the same constraint from two sides. RAG chooses which external knowledge enters the window. History strategies choose which past conversation stays in it. Context engineering is the discipline of making both choices together, on every call, within one token budget. For the full picture, read Tokens Explained: What Actually Happens When an LLM Reads Your Prompt?

The one idea to keep: a context window is not storage, it's a workspace. Its size tells you the most the model can look at. Your application decides what's worth looking at.