A language model is a model of text. Given the words so far, it estimates what comes next. A large language model (LLM) is that idea scaled up — many parameters, huge training data, usually built on transformers.
Core job in one line: predict the next token.
Example: “The cat sat on the ___”
The model assigns probabilities to possible next words (mat, sofa, roof, …). That next-token engine powers autocomplete, chat, translation-style systems, and more.
Scale (parameters + data + compute) is why modern LLMs can:
Well-known families include GPT, Claude, Gemini, and Llama-style open-weight models.
| Type | Plain-English idea | Typical uses |
|---|---|---|
| Encoder-only | Looks both left and right (bidirectional). Great at understanding. | Classification, NER, similarity, search embeddings (BERT family) |
| Decoder-only | Reads left-to-right; generates the next token. | Chat and free-form generation (GPT-style) |
| Encoder–decoder | Encoder reads input; decoder writes output. | Translation-style and many sequence-to-sequence tasks |
| Stage | What happens |
|---|---|
| Pretraining | Learn general language from huge unlabeled text |
| Fine-tuning | Adapt to a task, domain, or instruction style with more focused data |
| Safety / alignment | Extra training so the model follows policies and behaves more helpfully/safely |
A general pretrained model is a strong starting point, but many products need:
Fine-tuning updates the model on task-specific data so those habits become part of the weights.
LLMs learn next-token prediction at huge scale; fine-tuning is how you turn that general skill into stable, task-specific behavior.