Stop using greedy decoding for LLM inference
publication date: August 14, 2026, reading time: 10 minutes
Language models are by definition probabilistic: they estimate the probability distribution of a language. Yet, even though this property makes them incredibly powerful and useful, we like to pretend they are fully deterministic machines when generating outputs from them, essentially hiding their true nature behind a greedy-decoding blanket, typically under the guise of reproducibility.
The goal of this blogpost is to show how this is a terrible
default decoding choice.To my own surprise. I've been guilty of using greedy decoding quite frequently in the past. In simple terms, greedy decoding (setting temperature=0) ignores the full distribution of next-token probabilities and selects the one token
that happens to be the most likely one. As the name suggests, it's a greedy approach to finding the most probable output. In the next three sections, we will see that 1) this decoding method is not deterministic in practice, 2) it is heavily biased, and 3) the goal of that bias, maximizing the output likelihood, is wrong and degenerate.
Most of the things covered here are not new but already well established; the first papers that come to mind are The Curious Case of Neural Text DegenerationAri Holtzman et al. (2020). The Curious Case of Neural Text Degeneration. ICLR 2020. and The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism,Yifan Song et al. (2025). The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism. NAACL 2025. with many others following them more recently. Despite that, greedy decoding is as popular as ever. It is the default in Hugging Face's Transformers, it is still the de facto standard in evaluation harnesses, and just last year, it was used in 1 166 out of 7 493 published ACL papers.All papers in the main proceedings and Findings of the three major ACL conferences of 2025: ACL, NAACL, and EMNLP. This count is very approximate, based on a simple regex.
(The interactive map of the mentions needs JavaScript to render.)
1. Greedy decoding is not deterministic
We can immediately see that something is not working as advertised when generating
multiple continuations of a single prompt with the supposedly deterministic
greedy decoding – here using Qwen3-30B-A3B-Base and the otherwise
default vLLM settings:
(The divergence tree needs JavaScript to render.)
The divergence gets worse with longer outputs. At 1 000 output tokens, we get 9 560 unique continuations of “Determinism is” out of 10 000 – that's very far from deterministic!
The non-determinism is caused by the unintuitive quirks of floating-point computations, as explained very nicely in Defeating Nondeterminism in LLM Inference.Horace He et al. (2025). Defeating Nondeterminism in LLM Inference. Thinking Machines Lab: Connectionism. The root of the problem is that floating-point arithmetic is not associative: (a + b) + c is not always equal to a + (b + c). A single forward pass sums enormous numbers of floats in parallel, so the exact result depends on the exact order in which they happen to be summed.
What actually changed the order between the runs of Figure 2 then? First, the batch-dependence of matrix multiplication: the shape of a batch and the position inside it affect how the GPU kernels split their work and in which order they reduce it.Thus, any request to an LLM server should be automatically treated as non-deterministic, because each response is affected by the simultaneous requests of other users. See also Berk Atil et al. (2025). Non-Determinism of “Deterministic” LLM System Settings in Hosted Environments. Eval4NLP 2025. The second cause is atomic addition, a low-level GPU operation that sums elements in whatever order they happen to arrive – which the default mixture-of-experts kernel uses when combining the outputs of the top-k experts.Atomic additions are rarely used in deep-learning kernels precisely because of this inherent non-determinism – but the default MoE kernel in vLLM v0.20.1 for bfloat16 tensors, the FlashInfer CUTLASS backend, uses them to aggregate the expert outputs.
In the end, as these small imprecisions accumulate layer by layer, token by token, the difference in the final logits can be large enough to flip one of the generated tokens. And since every following token is then conditioned on it, the whole generation snowballs into a completely different output.
But are these deviations large enough to affect performance on benchmarks? Unfortunately, they are, especially when using reasoning models. Their long chains of thought give the divergence thousands
of tokens to build up, so two inference runs often end up arriving at entirely different
answers. Here are some popular benchmarks run with
Qwen3-30B-A3B-Thinking ten times. The run-to-run differences
are clear evidence of the probabilistic monster peeking out of the
false safety of the deterministic blanket:
(The benchmark chart needs JavaScript to render.)
That being said, it is absolutely possible to make model inference run-to-run deterministic. We just have to make sure the software stack itself is fully deterministic. In the case above, that means forcing vLLM to use a different MoE backend,Specifically the Triton backend, as it is actually deterministic. and to use a batch size of 1 to avoid any batch-dependence.Apparently, the vLLM team is now developing deterministic batch-invariant inference, currently in beta. Then all 10 000 completions result in a single output:
(The divergence tree needs JavaScript to render.)
However, my larger point is that this kind of determinism is mostly useless! It can only be reproduced when exactly matching the GPU type and the full software stack; even a small change can slightly alter the order of the floating-point operations, completely changing the output. Figure 5 shows how a little difference is enough to flip the outcome. Such determinism is as useful as claiming that a dice roll is reproducible just because you get the same outcome when you exactly match the hand movement each time. It is technically true, but practically useless. Instead, a much more useful description would treat the dice roll as a random event and provide the chances of all six possible outcomes. I believe we should adopt a similar mindset when evaluating language models.
(The divergence tree needs JavaScript to render.)
enforce_eager). Every run is still
perfectly repeatable, yet half of the completions quickly diverge from the outputs in Figure 4.2. Greedy decoding is biased
The point demonstrated in this section might be obvious – of course greedy decoding is biased! But I think it’s still useful to show this explicitly, because it highlights how greedy decoding is a non-trivial choice that greatly affects what the outputs look like, and as such, it should always be properly justified. Even ignoring the problems with non-determinism, there is still an infinite number of other decoding methods that are theoretically deterministic, so why choose greedy decoding specifically?
But first, what does “biased” mean here, exactly? Random sampling with temperature 1 is unbiased by construction, as it draws every token exactly according to the probability the model assigns to it. Repeating this process gives us a reference distribution of outputs that can be used to judge the bias of greedy decoding. Plotting the distribution of perplexities of the 10 000 greedy continuations from the previous section against the unbiased continuations of the same prompts, we see two very different behaviors. Greedy decoding produces text that the model itself considers atypically likely:
(The distribution chart needs JavaScript to render.)
Yet even if you do insist on a deterministic decoding method – despite the points in the previous section – greedy decoding is still not the natural choice. For example, sampling with temperature 1 and a set random seed is deterministic in exactly the same way, on exactly the same terms (a frozen deterministic software stack), but its output remains an unbiased draw from the model's distribution:
(The distribution chart needs JavaScript to render.)
In other words, determinism does not require the bias, and the reproducibility argument for greedy decoding falls apart. What remains is the bias itself. To be fair, one could defend it as fully intentional – greedy decoding is biased by design, its very purpose being to approximate the most-probable output.
3. Greedy decoding is degenerate
So is generating the most-probable output – the very target of greedy decoding – worth aiming at? Even though it seems like a good idea at first sight, it is flawed at a closer look.
We can see that when thinking about what the most-probable output actually looks like. Let's start by considering the problem backwards: given a word, can you find a left context that would maximize its likelihood? A simple and effective solution is to repeat that same word multiple times in the left context, let's say 100 times – after seeing so many repetitions in a row, any reasonable language model should be almost certain that the 101st repetition follows. And then crucially, the same logic applies to all further repetitions: once a repetition loop is established, every further repetition is nearly free in terms of likelihood. So a very effective recipe for a highly likely output is to simply repeat the same text over and over. Putting this to the test, let's see how the perplexity changes with these three generation strategies: 1) random sampling, 2) greedy decoding, and 3) repeating the prompt over and over:
(The trajectory chart needs JavaScript to render.)
It is clear that we have found our counter-example: mindless repetition consistently achieves the lowest perplexity of the three strategies, beating greedy decoding at its own objective, while being obviously unusable.It is likely that an even better strategy exists (in terms of maximizing likelihood), but that would also mean approaching the asymptote of perplexity 1 faster, so being even further away from natural language. Biasing a sampling method towards such outputs cannot be the right goal. Paradoxically, the only thing that saves greedy decoding in practice is its greediness, which prevents it from looking ahead and finding such outputs consistently.The same paradox is well documented for beam search in machine translation: increasing the beam size finds translations that are more likely but measurably worse. Philipp Koehn and Rebecca Knowles (2017). Six Challenges for Neural Machine Translation. WNMT 2017. But whenever repetition happens to be locally optimal, nothing stops it from falling into an endless repetition loop. This is especially problematic for long unconstrained outputs, such as reasoning traces of modern language models – sooner or later, some repetition loop becomes cheap enough to enter. It’s not surprising that the Qwen3 model card even explicitly warns users against this decoding method:
“DO NOT use greedy decoding, as it can lead to performance degradation and endless repetitions.”The model card of Qwen3-30B-A3B on Hugging Face.
This performance degradation can partly be observed when repeating the five benchmarks from Figure 3 with random sampling. Random sampling performs on par with greedy decoding or better: the one statistically significant difference, GSM8K, is in favor of random sampling, while all the remaining differences fall well within the run-to-run variations. Furthermore, the two methods show a comparable amount of that variation, so there is no additional stability coming from greedy decoding.
(The comparison chart needs JavaScript to render.)
Conclusion
I hope that it's now clear that greedy decoding is far from an obvious choice for LLM inference. On top of that, ensuring reproducible LLM inference is anything but simple – and, just like insisting on the reproducibility of a dice roll, it is ultimately pointless and misleading about the actual behavior of a language model. Instead of being scared by the probabilistic monster and hiding it behind the greedy-decoding blanket, we should embrace it and gather the benefits it gives us. With unbiased random sampling, modern language models score just as high on benchmarks, and we can easily estimate not only the expected performance on a target task, but also its variance! Reporting these statistics is much more useful and informative than a single biased pseudo-reproducible score.I really recommend reading the following paper for a practical statistical toolkit: Evan Miller (2024). Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations. arXiv:2411.00640.
Citation
@misc{samuel2026greedy,
author = {Samuel, David},
title = {Stop Using Greedy Decoding for {LLM} Inference},
year = {2026},
month = {August},
url = {https://davidsamuel.no/blog/greedy-decoding/}
}