You set temperature to zero, sent the same prompt twice, and got two different answers. Your snapshot tests, your response cache, and last week's eval baseline all assumed that couldn't happen. It did anyway. A customer refreshed a page on Monday and got a different assistant reply than on Sunday. A CI snapshot test flaked overnight and passed on the retry. Your Wednesday eval showed a 1.4-point drop against Tuesday's baseline, and now the whole team is arguing about whether the new prompt template regressed or whether you're inside the noise floor.
Backend engineers grew up on databases and pure functions. Same input, same output, always. LLM endpoints don't promise that, even at temperature zero, and the reasons live below the layer you have control over. Server load changes the batch size the model gets served in. Batch size changes the order of floating-point reductions in the transformer's attention kernels. Different reductions change the softmax probabilities by a fraction of a percent. On the token where two candidates were nearly tied, that fraction picks a different token. From that point on, the sequence diverges.
Issue 18 named sampling as one of the four training-and-inference layers that shape production behaviour. This issue is the inference-layer detail that catches the most teams off guard: how variance leaks in even when you've set temperature to zero, and the six design changes that let your tests, caches, and evals survive it. Horace He's Thinking Machines paper from September 2025 (see Further reading) is the clearest public write-up of the actual mechanism to date.
Why temperature zero doesn't mean deterministic
Temperature zero is greedy sampling. At each step, the model computes a probability distribution over the next token and picks the argmax. If the top token is meaningfully ahead of the runner-up, temperature zero picks the same token every time. If the top two are close, a tiny numerical perturbation flips the choice.
The perturbation is real. Transformer inference on a GPU does a very large number of floating-point multiplications and additions, and floating-point addition isn't associative. (a + b) + c isn't guaranteed to equal a + (b + c) for arbitrary floats. When the server batches your request with three others, the reductions happen in one order. When it batches with seven, the reductions happen in a different order. The final logits differ by one part in ten thousand. Most of the time, that doesn't matter. On the token where the top two are close, it changes the output. Everything downstream of that token now differs.
Batch invariance is the fix, and none of the major vendors ship it by default in mid-2026. Horace He's paper walked through how to write attention and reduction kernels that produce bit-identical outputs regardless of batch size, and the answer requires custom kernels and a performance hit. Vendors optimise for latency and throughput, so they trade determinism for speed. You inherit that trade-off whether you want it or not.
There's a second source of variance that's easier to name. Vendors ship model updates behind aliases. gpt-4o isn't a model, it's a pointer to whatever snapshot is current. When OpenAI rolls a new snapshot, gpt-4o starts returning outputs from the new model, and nothing in your code changed. OpenAI ships a system_fingerprint in the response that tells you which backend served the request; Anthropic pins model versions with a date suffix. Reading either one is how you find out the model has moved, and if you're not logging them, you'll debug the wrong problem the first time behaviour shifts.
What backend engineers usually reach for, and why it doesn't hold
The default backend playbook for flakiness is retries plus idempotency keys plus a good cache. That playbook was designed for network flakiness and transient database contention, and it doesn't cover the variance you're actually facing here. Retries don't help because the endpoint is answering correctly; it just answers slightly differently on two consecutive correct answers. Idempotency keys don't help because the vendor doesn't check them for LLM calls, and even if you cached at your own layer, you'd still be recomputing on the first call. The cache helps but only after the first sample lands, and it introduces staleness problems of its own.
The other default is "just set the seed". OpenAI's API accepts a seed parameter and returns a system_fingerprint. In principle, the same seed and the same fingerprint should give bit-identical output. In practice, the docs are careful to call it a best-effort guarantee, and the reason is exactly the batch-invariance issue above. Set the seed, still expect drift, log the fingerprint so you know when the model rolled underneath you. Seed helps; it isn't the fix.
The pattern that does hold is different. Design your code and your tests to assume the LLM is a stochastic function, not a pure one. Same input, distribution of outputs, with statistical guarantees on properties of that distribution. Every recommendation in this issue lives inside that reframing.
The design-for-variance checklist
Six components, each of which fixes one class of failure. Adopt them incrementally in the order below; the first three cost you almost nothing, and the last three are the ones that let you tell a real regression from noise.
Pin dated model snapshots. Never target an alias in production. Anthropic gives you dated versions like claude-3-5-sonnet-20241022; OpenAI gives you dated snapshots like gpt-4o-2024-08-06. Pin the one you tested against. Migration to a new snapshot is a deliberate, planned event (Issue 14's playbook), not something that happens to you overnight because a vendor rolled an alias.
Log the model version and any fingerprint the provider returns. OpenAI's system_fingerprint, Anthropic's response headers, the request ID. Put them next to the trace from Issue 4. When someone six weeks from now says "the model started behaving differently around this date", the logs are what tell you whether the vendor changed something under you.
Store outputs rather than regenerating them. If you generated a response for a request and shipped it, keep it. Every fresh call re-samples from the distribution. A cache keyed on the exact prompt (plus system prompt, plus tool set, plus model version) means the second call returns the first call's output, which is what your users experienced and what your tests were written against. This is separate from the semantic prompt-caching layer from Issue 9; that one reuses the KV cache on the vendor's side to save cost. Output caching on your side is what pins user-visible behaviour.
Swap exact-match snapshot tests for property assertions. A test that says "the output equals this exact string" is going to flake, and when it does, you'll add a retry and move on, and now your test suite doesn't catch real regressions either. Rewrite the test to assert properties instead: "the output is valid JSON matching this schema", "the output mentions the customer's account number", "the output is under 500 tokens". Properties are what you actually care about; exact strings are an implementation detail of the sampled trajectory.
Run evals over several samples with confidence intervals. A single-shot eval on 200 questions gives you a point estimate with no error bars. Two runs a week apart will show a delta, and you can't tell whether the delta is signal or noise. Run each eval prompt three to five times, take the mean and a bootstrap 95 percent confidence interval, and compare intervals rather than points. This roughly triples your eval cost. It also stops the "is that a regression or is it drift" argument permanently.
Agree a minimum detectable change (MDC) before anyone calls a regression. How large a shift in your customer-facing metric would a rational team ship on? A one-point drop on a hundred-example set is inside the noise of most LLM evals. A five-point drop probably isn't. Pick the MDC by looking at run-to-run variance on the current model with the current prompt, and set the MDC to at least two standard deviations of that noise. Nothing below the MDC is a regression, no matter how the graph looks in the retro meeting.

The causal chain that produces the variance you see in production, and the layer where you actually fix it. Reads left to right: same prompt goes in, the vendor's server-side conditions perturb the logits by a tiny amount, and when two candidate tokens are close the choice flips. The fix isn't on the vendor's side of the box; it's in how you design the tests, caches, and evals downstream.
A worked example: turning a snapshot test into a property test
You have a snapshot test that asserts the assistant's response to a specific customer-support query equals a specific 400-word string. It's been flaking on CI once every twenty runs since you shipped. Here's what you're actually trying to test, phrased as properties:
The response acknowledges the customer's issue by name. Assert with a regex on the customer's account identifier.
The response cites the correct refund policy by ID. Assert with a lookup against the policy registry.
The response ends with a call to action from the approved set. Assert with membership in a small list.
The response is between 200 and 500 tokens. Assert with the tokeniser count.
The response contains no PII from other customers. Assert with a regex against the anonymised training corpus.
Five specific, checkable properties replace one brittle string match. Any real regression will show up in at least one of them. The one-in-twenty flake goes away because the exact word choice of the assistant's polite acknowledgement was never the thing you cared about.
Common mistakes
Treating temperature zero as a determinism guarantee. This is the failure that this issue exists to warn against. Temperature zero is greedy sampling, and greedy sampling is deterministic in exact arithmetic. Your vendor doesn't do exact arithmetic. Set temperature zero for tasks where you want the mode of the distribution, but design everything downstream as if it were stochastic.
Snapshot testing LLM outputs by string equality. Every LLM app I've seen with this pattern eventually adds a retry to the test, then a bigger retry, then an "allowed diff" tolerance, and finally the test stops finding regressions altogether. Skip the whole arc. Write property tests from day one.
Reading eval deltas without confidence intervals. A 1.2-point drop looks like a regression. It's often noise. Run the eval three to five times, plot the confidence interval, and ship when the new interval sits above the old one by more than the minimum detectable change. Confidence intervals turn eval graphs from a source of arguments into a source of decisions.
Regenerating cached outputs "in case the model got smarter". The model didn't get smarter; the vendor rolled a snapshot behind the alias you never should have been using. The user-visible behaviour that people wrote docs and workflows around just changed. Cache your outputs, pin your snapshots, and treat model migration as the deliberate event Issue 14 laid out.
The takeaway
LLM inference isn't a pure function even at temperature zero, and the reason is a stack of engineering decisions on the vendor's side that trade determinism for throughput. Same prompt, same settings, sometimes a different output, and that's not going to change soon. The six-item design-for-variance checklist (pin dated snapshots, log fingerprints, cache outputs, replace snapshot tests with property assertions, run evals with confidence intervals, and agree a minimum detectable change before anyone calls a regression) is what lets your tests, caches, and evals stop pretending the endpoint is idempotent and still catch real regressions. The vendor won't fix this for you. You design around it.
Production checklist
Replace every model alias in production and CI with a dated snapshot string. Grep for the alias name and delete it.
Log the model identifier, the provider's fingerprint (OpenAI
system_fingerprint, Anthropic response headers), and the request ID on every LLM call, next to the Issue 4 trace fields.Cache LLM outputs keyed on
(model_snapshot, system_prompt_hash, user_prompt_hash, tool_set_hash, temperature, seed). Re-serve the cached output rather than re-sampling.Audit your test suite for exact-string snapshot assertions against LLM output. Convert each to a set of property assertions on the response's structure, content, and length.
Add a "seed" column to your eval runs. Where the vendor supports seed, pin it and log the returned fingerprint alongside.
Run every eval prompt at least three times, compute the bootstrap 95 percent confidence interval, and gate ship decisions on intervals rather than point estimates.
Measure the current-model run-to-run variance on your golden set, and set the team's minimum detectable change to at least two standard deviations of that noise. Document it in the eval README.
When a vendor announces a new model snapshot, run the Issue 14 shadow-eval gate against the pinned old snapshot before switching. The variance floor from this issue is what tells you whether the new number is a real win.
Add a "vendor fingerprint changed unexpectedly" alert. If a request comes back with a fingerprint you haven't seen before, page the on-call engineer; it means an alias or a routing rule shifted under you.
Further reading
Horace He, Thinking Machines Lab, "Defeating Nondeterminism in LLM Inference" (September 2025) - thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference
Anthropic API reference, "Versions" - docs.anthropic.com/en/api/versioning
David Goldberg, "What Every Computer Scientist Should Know About Floating-Point Arithmetic" - docs.oracle.com/cd/E19957-01/806-3568/ncg_goldberg.html
Hamel Husain, "Your AI product needs evals" - hamel.dev/blog/posts/evals