From chat to API: the system slot
Temperature — and why it was taken away
Every prompting tutorial written in the last few years has a temperature section. It says the same thing: a dial from 0 to 1, low for facts, high for creativity, 0.3 for production. You will have read it a dozen times.
On current Claude models, sending that parameter returns a 400 error. Anthropic lists temperature, top_p and top_k under API parameter deprecations and tells you to remove them from your request payloads. The recommended replacement is one word: prompting.
This lesson is here rather than deleted because the reason the dial was removed is worth more than the dial ever was.
What temperature actually did
When the model picks the next token, it is choosing from a probability distribution. Temperature scaled that distribution before sampling. Low temperature flattened the long tail and made the top choice near-certain. High temperature flattened the peak and let unlikely tokens surface.
That mechanism is real, and it still applies wherever the parameter is still accepted — older Claude snapshots, and other vendors, each with their own range and default. If you are working against one of those, read that vendor's parameter reference rather than a blog post; the ranges are not the same across providers, which is itself a good reason not to memorise a number.
The part everyone got wrong
The standard advice was "use low temperature when you need repeatability." Anthropic's own migration guidance retires that idea directly:
If you were using
temperature = 0for determinism, note that it never guaranteed identical outputs on prior models.
It never worked. Not "works less well now" — it did not do the job it was famous for, on any model, ever. Temperature narrowed the distribution; it did not pin the output. People ran a prompt twice at 0.0, got two similar answers, and concluded the dial was holding the line, when what was actually holding the line was a prompt with only one sensible answer.
That is the whole lesson, and it generalises past this one parameter: a setting that correlates with the outcome you want is not the same as a setting that controls it.
What you do instead
Everything the dial was reached for, the prompt does better and more explicitly:
| What you wanted the dial for | What actually gets you there |
|---|---|
| Consistent, repeatable replies | An output spec tight enough that only one shape of answer fits — length, format, required and banned elements |
| Varied phrasing, several options | Ask for it: "give me five alternatives, each with a different opening verb" |
| Less rambling | A length constraint and a "no preamble, no summary" rule |
| Less invention | Grounding: paste the source and require answers from it only (Module 8) |
The right-hand column is stateable, reviewable, and survives a model change. A number in a config file is none of those things.
You reached for the temperature dial. Now what?
What were you trying to fix?
Why this happened at all
Read the replacement advice once more — use prompting to guide model behaviour. The vendor removed the knob because the knob was the weaker instrument. On models that plan before they answer, nudging token probabilities is a blunt way to ask for something you could simply state.
This course has been making that argument since Module 2, in the lesson on the output spec. The platform has now made it non-optional. If you have been tightening prompts instead of tuning settings, nothing about your practice changes — which is the nicest thing an upstream deprecation can say about a skill.
The Bayt Coffee assistant Hagar is about to build gets its consistency from a tight system prompt and a locked output format, not from a number. That was always the better build; now it is the only one.
Next module: scaling the 5-slot prompt skeleton up into a real production system prompt. :::
Sign in to rate