Soft prompting moves adaptation out of the architecture and into token space. Instead of only writing a hand-crafted text prompt, we learn virtual prompt tokens whose embeddings are trained — while the backbone model can stay frozen.
| Approach | Where adaptation happens | Plain-English idea |
|---|---|---|
| Adapters | Architecture space | Add small modules into the network |
| Soft prompts | Token space | Add and learn task-specific virtual tokens / context |
The big idea: good context can steer the language model without changing its weights.
| Kind | Meaning | Catch |
|---|---|---|
| Discrete prompt | A normal string of real tokens written by a person | Small wording changes can swing quality a lot |
| Continuous / soft prompt | A sequence of trainable virtual token embeddings | Not real vocabulary words — learned vectors that act like extra context |
Why continuous prefix-style tuning exists: manual prompts are brittle. Learning a soft prompt can be more stable and compact than rewriting English by hand.
Soft prompts can be sensitive to:
Longer soft prompts also mean more tokens to process — so cost can grow with prompt length even when the backbone is frozen.
Soft prompting steers a frozen model by learning virtual prompt tokens in embedding space, instead of rewriting the network.