LLMsBuild
Why LLM Responses Become Inconsistent
Same prompt, different answers. Sampling, prompt drift, model updates and hidden context changes all contribute. Here is how to make output predictable enough to ship.
Sampling is randomness by design
Models generate text by sampling from a probability distribution over the next token. Temperature controls how adventurous that sampling is. Higher temperature means more variety; lower temperature means more repeatable output. Even at low temperature, outputs are not guaranteed to be identical across calls.
The prompt is rarely as stable as you think
Inconsistency often comes from inputs that change without anyone noticing:
- Retrieved documents differ between runs because the index changed.
- Conversation history is truncated at a different point.
- A timestamp, user name or locale is injected into the system prompt.
- Tool results vary because the underlying data changed.
Log the full final prompt for a sample of requests. Comparing two prompts that produced different answers frequently reveals the cause in minutes.
Model versions change
If you call a model alias rather than a pinned version, the provider may update the model behind it. Behavior shifts with no change in your code. Pin model versions in production and upgrade deliberately, with an evaluation run before and after.
Make the output format non-negotiable
Free-form text drifts. Ask for structured output with an explicit schema, validate it in code and retry with the validation error when it fails. Structured output turns "the answer looks different" into a type error you can detect and handle.
Measure consistency
Run each evaluation question several times and measure agreement, not just accuracy on a single run. A prompt that is right 9 times out of 10 has a 10% failure rate in production, no matter how good the demo looked.