AI PerformanceBuild

Caching LLM Responses Without Serving Stale Answers

Caching can cut AI latency and cost dramatically, or serve wrong and leaked answers. What to cache, how to key it and when to invalidate.

By The AI Build Room2 min read90 views

Three different caches

"Caching LLM calls" can mean three different things, with different risks:

  1. Provider prompt caching: the provider reuses computation for a repeated prompt prefix. Output is still generated fresh. Low risk, often a large saving on long system prompts.
  2. Exact response caching: identical input returns a stored output. Safe when the key truly captures every input.
  3. Semantic caching: similar questions return a stored answer. Highest savings and highest risk.

Build the key from everything that matters

An exact cache key must include the model and version, the full prompt (including system prompt and retrieved context), and generation parameters. If you leave out retrieved context, a cached answer will survive a documentation update and keep serving the old answer.

Never share personalized answers

If the prompt contains user-specific data, the cache must be scoped to that user, or not used at all. A shared cache that returns one user's account details to another user is a data breach, not a performance bug.

Semantic caching needs a high bar

"How do I reset my password?" and "How do I reset my 2FA?" are close in embedding space and need different answers. Use a strict similarity threshold, restrict semantic caching to curated FAQ-style questions, and measure the false-hit rate before rolling it out.

Invalidation

Tie cache entries to the versions of what produced them: prompt template version, model version and the ids or hashes of the source documents. When any of these changes, the key changes and stale entries simply stop being used. Add a time-to-live as a safety net, not as the primary invalidation strategy.