Catastrophic forgetting happens when you fine-tune a pretrained model on a narrow task and the updates overwrite the weights that held broader skills. The new task gets better. Everyday abilities — following instructions, coding, other languages, simple reasoning — can quietly get worse.
Picture a generalist who studies only medical Q&A for weeks and somehow “forgets” how to write Python or speak French. The model did not delete a folder named “French.” The same shared weights store many skills. Push those weights hard toward one narrow dataset, and other skills can fade.
Why it is scary:
| Cause | Plain-English idea |
|---|---|
| Shared weights | Many skills live tangled in the same parameters. |
| Narrow data | Fine-tune examples cover a tiny slice of what pretraining saw. |
| Weight drift | The farther you wander from the starting model, the more old skills suffer. |
| No memory of old tasks | A normal optimizer does not “remember” earlier abilities unless you re-show old data or add a penalty for drifting. |
Imagine a 7B model fully fine-tuned for a few epochs on legal contracts: contract accuracy jumps, while general benchmarks drop. That pattern is the warning light. Keep one fixed general eval set and compare base vs fine-tune every time.
| Strategy | Plain-English idea | When to use it |
|---|---|---|
| Conservative hyperparameters | Smaller learning rate, fewer epochs, early stop on canary evals | Always — your default baseline (almost free) |
| Rehearsal / data mixing | Mix a little general instruction data (often about 1–10%) into training so old skills keep getting practice | When you have (or can proxy) general data |
| Regularize toward the start | Penalize big drift from the pretrained weights (L2-SP or EWC-style ideas) | Full fine-tune required and little general replay data |
| Weight averaging | After training, blend fine-tuned weights with the original model to recover generality | Model already fine-tuned and general quality dropped |
| Freeze layers | Lock lower layers; tune only the top part closest to the task | Tight compute or very small datasets |
Start with conservative hyperparameters. Add rehearsal if you can. Use freezing when data or compute is tight. Reach for regularization or averaging when you must fully fine-tune and general skills are slipping.
Narrow fine-tuning can overwrite broad skills; measure general canaries, then reduce forgetting with smaller updates, rehearsal, freezing, regularization, or blending back toward the base model.