The learning rate is how big each training step is. One fixed value for the whole run is rarely ideal. Early training often needs larger steps to make progress; late training needs smaller steps so the model can settle safely. A schedule changes the learning rate over time.
The dilemma in one sentence:
You want both stories in one run: bigger steps early, smaller steps later.
Training also has phases:
| Knob | Plain-English idea |
|---|---|
| Warmup length | How long the learning rate ramps from ~0 to the peak (often about 3% of steps; raise to 5–10% if the run is short or unstable) |
| Peak learning rate | The maximum. Rough guide: about 1×10⁻⁵–2×10⁻⁵ for full fine-tuning; about 1×10⁻⁴–2×10⁻⁴ for LoRA |
| Decay shape | How you go down from the peak (linear, cosine, steps, restarts, …) |
| Floor (min learning rate) | Where decay ends — zero, or a small floor if you want learning to stay slightly alive |
| Schedule | Plain-English idea | When it fits |
|---|---|---|
| Constant (+ warmup) | After warmup, keep one learning rate forever | Short, controlled runs (for example short LoRA); more spike risk on long runs |
| Linear decay | After warmup, fall in a straight line toward zero | Clean, simple baseline |
| Cosine decay | Stay productive near the peak longer, then glide gently down | Modern default for many LLM fine-tunes |
| Step decay | Stay flat, then drop by a factor at milestones (or when validation stalls) | When you want sharp, reactive drops |
| Cosine with restarts | Decay, jump back up, decay again | Exploring more than one “basin” in one budget |
| Warmup–Stable–Decay (WSD) | Warm up, hold a plateau, then decay at the end | Longer or open-ended budgets |
Schedules give you large steps when learning is easy and small steps when the model needs to settle — warmup plus a decaying shape (often cosine) is the usual safe pattern.