Clever schedules cannot fully rescue bad data. This chapter is about the data problems that quietly break fine-tunes, and the habits that keep training honest.
Two things matter most: important characteristics of data, then how to prepare it. Five quality issues show up again and again.
| Problem | What goes wrong | Typical fix |
|---|---|---|
| Noisy data | Menus, ads, OCR junk, spam, or AI-generated slop pollute the corpus | Filter hard; remove obvious garbage |
| Incorrect labels / “truth” | Wrong answers, buggy code, mis-tagged language taught as fact | Audit labels; drop or fix bad rows |
| Duplicates | Memorization rises; evals get contaminated | Remove exact and near duplicates |
| Inconsistent formatting | HTML, Markdown, PDF text, and LaTeX mixed badly | Normalize structure before training |
| Outdated information | Old APIs, old policies, old officeholders stated as current | Refresh corpus; track cutoffs |
| Risk | Plain-English idea | What to do |
|---|---|---|
| Small datasets | Too few examples → easy overfitting, high variance | Prefer transfer / light tuning; augment carefully; do not train forever |
| Missing edge cases | Rare but important failures never appear in train | Actively collect those cases; mine production misses |
| Lack of diversity | Model works for one group and fails for others | Set coverage targets; measure per subgroup |
| Imbalanced classes | Always predicting the majority looks “accurate” | Use precision/recall/F1; rebalance or weight classes |
| Bad splits | Test scores look great because of leakage | Split first; never fit transforms on the full dataset |
| Domain / distribution shift | Train on web docs, deploy on live chats and tools | Align training mix to deployment; use retrieval for fresh facts |
A common honest split is roughly 70% train / 15% validation / 15% test (adjust to your size).
Pitfalls (leakage):
Better habits:
| Shift | Example |
|---|---|
| Usage shift | Trained on documents; deployed on chat and tools |
| Temporal shift | World and APIs moved on after the training cutoff |
| Domain shift | Web-trained model dropped into clinical notes or contracts |
Mitigations in practice: instruction-style post-training, more deployment-like data late in training, domain refresh with replay, retrieval for current facts, and monitoring live quality.
Clean, deduplicated, well-split, deployment-like data beats clever training tricks; treat noise, labels, duplicates, format, and staleness as first-class bugs.