LLMsBuild

Why LLM Responses Become Inconsistent

Same prompt, different answers. Sampling, prompt drift, model updates and hidden context changes all contribute. Here is how to make output predictable enough to ship.

By The AI Build Room2 min read52 views

Sampling is randomness by design

Models generate text by sampling from a probability distribution over the next token. Temperature controls how adventurous that sampling is. Higher temperature means more variety; lower temperature means more repeatable output. Even at low temperature, outputs are not guaranteed to be identical across calls.

The prompt is rarely as stable as you think

Inconsistency often comes from inputs that change without anyone noticing:

  • Retrieved documents differ between runs because the index changed.
  • Conversation history is truncated at a different point.
  • A timestamp, user name or locale is injected into the system prompt.
  • Tool results vary because the underlying data changed.

Log the full final prompt for a sample of requests. Comparing two prompts that produced different answers frequently reveals the cause in minutes.

Model versions change

If you call a model alias rather than a pinned version, the provider may update the model behind it. Behavior shifts with no change in your code. Pin model versions in production and upgrade deliberately, with an evaluation run before and after.

Make the output format non-negotiable

Free-form text drifts. Ask for structured output with an explicit schema, validate it in code and retry with the validation error when it fails. Structured output turns "the answer looks different" into a type error you can detect and handle.

Measure consistency

Run each evaluation question several times and measure agreement, not just accuracy on a single run. A prompt that is right 9 times out of 10 has a 10% failure rate in production, no matter how good the demo looked.