Skip to content

temperature

Temperature is a decoding parameter that rescales the model’s logits before softmax, thereby controlling how deterministic or diverse the output is.

Lowering temperature sharpens the distribution, pushing probability mass toward the top tokens and approaching greedy decoding in the limit. Raising temperature flattens the distribution, allowing more exploration and novelty at the cost of coherence or correctness. Even the low end doesn’t guarantee reproducibility, though. A temperature of 0.0 still won’t make results fully deterministic.

Temperature is implemented by dividing logits by the temperature constant before softmax, and it is often used together with top-k or nucleus (top-p) sampling. In practice, one chooses temperature per task: low for factual or precise generation, and higher for creative or open-ended generation.

Not every model exposes the parameter. Some current reasoning models reject non-default temperature values with an error, and some vendors have deprecated the knob outright. Anthropic’s Messages API, for example, marks temperature as deprecated and rejects any value other than 1.0 on models released after Claude Opus 4.6. Where the parameter is unavailable, prompting or a reasoning-effort setting steers the output instead.

Prompt Engineering: A Practical Example

Tutorial

Prompt Engineering: A Practical Example

Learn prompt engineering techniques with a practical, real-world project to get better results from large language models. This tutorial covers zero-shot and few-shot prompting, delimiters, numbered steps, role prompts, chain-of-thought prompting, and more. Improve your LLM-assisted projects today.

intermediate ai data-science

For additional information on related topics, take a look at the following resources:

Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.


By Leodanis Pozo Ramos • Updated Sept. 21, 2026