Core Concepts Entry

Temperature, explained

Reviewed September 2026

TL;DR: Temperature controls how a model picks its next token from the probabilities it computed. Low temperature always takes the likeliest choice, producing consistent, focused output; high temperature spreads the choice across less likely tokens, producing variety and surprise. It changes selection, not intelligence - a wrong answer at temperature zero is still wrong every time.

How it works

Every generation step ends with the model holding a probability for each token in its vocabulary. Something still has to pick one, and that step is called sampling. Temperature is a number that reshapes the distribution before the pick: near zero it sharpens the distribution until the top token wins essentially every time; around one it samples proportionally to the model's actual probabilities; higher than that it flattens the distribution so unlikely tokens get real chances. Ask "The sky is" at low temperature and you get "blue" on every run; turn it up and "turning orange over the harbor" becomes possible.

Choosing a value is about matching the task. Extraction, classification, code generation, and anything downstream code must parse - especially structured output - want low temperature: you need the same input to yield the same answer. Brainstorming, marketing copy, and dialogue want moderate settings, where variety is a feature. Temperature usually travels with a companion control, top-p, which limits sampling to the smallest set of tokens covering a probability threshold; the standard advice is to adjust one and leave the other alone.

Two misconceptions are worth killing. First, temperature is not an accuracy dial: lowering it makes output more repeatable, not more correct, and raising it does not add knowledge - though pushing it very high does increase the odds of incoherent or fabricated continuations, which is where it brushes against hallucination. Second, temperature zero does not guarantee bit-identical runs in practice - batching and floating-point nondeterminism on inference servers can still produce small variations, a real nuisance when you want reproducible evals.

Note that reasoning-focused models increasingly fix, ignore, or outright reject sampling controls - on some current models any value but the default returns an error, and an effort or thinking level is the knob you turn instead. Check what your target model actually honors before building logic around a setting.

Where it sits in the AI stack

Temperature acts at the last instant of inference - between the model's probabilities and the token that ships:

Key tools and implementations

  • Temperature parameter

    The classic API knob, still offered on most chat and completion endpoints - but no longer universal across current models.

  • Top-p (nucleus) sampling

    The companion control that caps sampling to the most probable slice of the vocabulary.

  • Greedy decoding

    Always taking the single likeliest token - what temperature zero approximates.

  • Seed parameters

    Provider options that improve run-to-run reproducibility where determinism matters.