Sampling knobs are how you turn a probability distribution into a product behavior. The same prompt can yield a crisp JSON object or a meandering poem depending on decoding. Engineers who treat these as afterthoughts spend their weeks chasing “flaky” models that were never configured for the job.
At each step the model scores the whole vocabulary. Decoding is the policy that picks one token:
Temperature changes the shape of the distribution; top-k / top-p change which slice you sample from; max tokens and stop sequences decide when the loop ends. Latency and cost scale with how long you let that loop run.
| Parameter | Controls | Notes |
|---|---|---|
| temperature | Sharpness of softmax | Primary diversity dial |
| top_k | Keep only k highest-prob tokens | Hard shortlist |
| top_p | Keep nucleus with mass ≥ p | Adaptive shortlist |
| frequency / presence penalty | Discourage reused tokens | Helps long prose; can hurt code |
| max_tokens / max_output | Hard length cap | Cost + latency ceiling |
| stop sequences | End when a string appears | Great for delimiters and turn ends |
z.z' = z / T.Using both is fine if you understand the order your API applies them; setting both extremely tight (tiny k and tiny p) can leave an empty or brittle candidate set.
Overdoing penalties on code or legal text makes identifiers and defined terms drift. Prefer clearer prompts and lower temperature before cranking penalties.
Stop sequences are surgical: end at \n\nUser:, ` , or </json>. Max tokens is a blunt instrument: it prevents runaway cost but can cut mid-JSON. Pair them: stop on a closer when possible; set max as a safety net.
Toy decode step with temperature + top-k:
import math
import random
def softmax(xs):
m = max(xs)
exps = [math.exp(x - m) for x in xs]
z = sum(exps)
return [e / z for e in exps]
def apply_temperature(logits, T: float):
return [x / T for x in logits]
def top_k_mask(probs, k: int):
order = sorted(range(len(probs)), key=lambda i: probs[i], reverse=True)[:k]
allowed = set(order)
masked = [probs[i] if i in allowed else 0.0 for i in range(len(probs))]
s = sum(masked) or 1.0
return [p / s for p in masked]
def sample(logits, T=0.8, k=3):
probs = softmax(apply_temperature(logits, T))
probs = top_k_mask(probs, k)
return random.choices(range(len(probs)), weights=probs, k=1)[0]
random.seed(1)
vocab = ["yes", "no", "maybe", "unknown"]
logits = [2.2, 1.1, 0.4, -0.5]
print(vocab[sample(logits, T=0.2, k=2)])
Stop-sequence trimming for structured replies:
def apply_stops(text: str, stops: list[str]) -> str:
cut = len(text)
for s in stops:
i = text.find(s)
if i != -1:
cut = min(cut, i)
return text[:cut]
raw = '{"ok": true}\n\nUser: ignore'
print(apply_stops(raw, ["\n\nUser:", "\n\n"]))
Frequency penalty sketch:
def penalize(logits, token_counts: list[int], penalty: float):
# subtract penalty * count from each logit
return [z - penalty * c for z, c in zip(logits, token_counts)]
Config object you can version in experiments:
PRESETS = {
"extract": {"temperature": 0.1, "top_p": 0.3, "max_tokens": 256},
"ideate": {"temperature": 0.9, "top_p": 0.95, "max_tokens": 800},
}
Decoding parameters are the runtime policy over next-token probabilities — tune temperature, truncation, penalties, and stops to match stability, cost, and format needs.