QLoRA is not only about 4-bit weights. It works because several memory-saving methods are stacked. Two big ones are gradient checkpointing and a paged optimizer.
The previous lessons packed the frozen textbook and kept the sticky note small. This lesson is about two other ways GPU memory still dies.
GPU memory fails for different reasons:
Think of training as cooking with a small counter:
QLoRA already shrinks the books. Checkpointing puts some bowls back in the fridge and recooks them when needed. Paging lets overflow notes sit on a side table instead of knocking the whole kitchen over.
During normal backpropagation, the model walks forward through the layers and stores many activations (intermediate outputs) so the backward pass can compute gradients later. That storage can be huge — sometimes bigger than the weights.
Gradient checkpointing discards some of those activations on purpose and recomputes them in the backward pass.
| Benefit | Cost |
|---|---|
| Less activation memory | Extra compute during backprop |
It is like keeping bookmarks instead of leaving every page open:
A common placement intuition: put checkpoints roughly every √n layers (for n layers). Then the backward pass can restart from a nearby checkpoint instead of replaying the whole network from layer 1 every time.
Tiny example:
√36 = 6If backward needs layer 20, it restarts from the bookmark at 18 — not from layer 1. Too few bookmarks → long recomputation. Too many → you are almost back to storing everything.
You do not need to hand-pick this every time. Libraries often place checkpoints for you once you turn the feature on. The √n picture is why “a few bookmarks, not all pages” is the usual design.
The optimizer is not free. Adam-style methods keep extra numbers per trainable parameter (running averages). That optimizer state can cause sudden out-of-memory spikes — especially with long sequences or uneven batch sizes.
A paged optimizer keeps that state more flexibly, often by paging data between GPU and CPU memory, so a spike is less likely to crash the run.
Same idea as a computer swapping RAM to disk:
It does not remove memory cost entirely. It smooths the spikes that cause failures.
| Piece | Memory pain it targets |
|---|---|
| 4-bit / NF4 weights | Base weight storage |
| LoRA adapters | Trainable parameter count |
| Double quantization | Scale / metadata overhead |
| Gradient checkpointing | Activation memory |
| Paged optimizer | Optimizer spikes |
Tiny sketch (concept only):
from transformers import TrainingArguments
model.gradient_checkpointing_enable() # bookmarks, not every activation
args = TrainingArguments(
optim="paged_adamw_8bit", # smoother optimizer memory
per_device_train_batch_size=1,
gradient_accumulation_steps=16, # keep the step small; accumulate
)
Read the names as the ideas above:
gradient_checkpointing_enable() → save activation memory, pay extra computepaged_adamw_8bit → Adam-style update that can page optimizer stateThe next lesson puts these flags next to NF4 and LoRA in one stack.
Checkpointing saves activation memory by recomputing; a paged optimizer softens optimizer spikes — both help QLoRA fit on fewer GPUs.