Non-determinism
Updated 2026-09-15
Non-determinism means the same prompt can produce different answers on repeated runs, even with identical settings. It is a property of how these systems sample text, amplified by retrieval that returns a different source set minute to minute.
What it forces on measurement
Three consequences follow, and they are not optional. Measurement must be sampled rather than checked once. Comparisons must hold the prompt set and sample size constant, or the difference measured is the instrument, not the brand. And any promise of a specific future percentage is unfounded, because the system producing the number is not controlled by the vendor promising it.
Illustrative run (example numbers)
One prompt, ten runs, same engine, same day: the brand appears in four. Run the same ten again and it appears in six. Neither result is wrong; both are samples of a rate somewhere near 0.5, which is exactly why a single check cannot be reported as a finding.
Common mistakes
- Screenshotting one good answer as proof of visibility, or one bad answer as proof of a problem.
- Explaining every cycle-to-cycle wobble with a content change, when variance alone accounts for most small moves.
- Buying a tool that promises a guaranteed percentage outcome — that promise cannot be kept.
Frequently asked questions
- Can temperature settings remove non-determinism?
- They reduce sampling variance but not retrieval variance: the sources returned for a query change on their own, so repeated measurement stays necessary even with conservative settings.
- How large a change is real?
- Large enough to exceed your own cycle-to-cycle noise at your chosen N. Measuring two consecutive cycles with no changes made is the cheapest way to learn what that threshold is for your prompt set.
