A model that “knows” your domain still fails if the prompt is mush. Prompting is not poetry; it is an interface contract: tell the model the job, the constraints, the data, and the shape of a successful answer — then keep that contract stable across versions.
Chat APIs are not a single blob of text. They are a role-tagged transcript: system (or developer) messages set policy, user messages state the task, assistant messages carry prior replies. The model continues the transcript under those roles. Confusing “who said what” is how you get leaked instructions, ignored policies, and brittle few-shots.
Good prompts look boring on purpose: short instruction, explicit format, one example if needed, clear failure behavior when uncertain.
Minimal pattern:
Instruction: Classify sentiment as Positive, Neutral, or Negative.
Input: "The food was okay."
Output: Sentiment:
| Role | Typical contents | Lifetime |
|---|---|---|
| system / developer | Persona, safety, tools policy, house style | Stable across turns |
| user | Task + payload | Per request / turn |
| assistant | Prior model replies (and sometimes tool results, depending on API) | History |
Multi-turn apps append assistant and user turns. That history is context — it burns tokens and can contradict a new system policy if you never refresh it.
""" or XML-ish tags) so instructions and payloads do not blend.Providers differ on how strongly system messages outweigh user text. Defense in depth still matters: put non-negotiables in system and repeat critical format rules near the user task for long contexts. Never assume a buried system line survives a 50-turn thread packed with RAG.
Prefer clarifying the contract over stacking adjectives (“be extremely careful and very precise…”). Models respond better to checkable rules than to emotional intensity.
Version prompts like code (support_v3). Keep a golden set of inputs with expected properties (label, JSON keys, refusal). On model upgrades, run the suite before you celebrate the new default. For multi-tenant apps, never put another customer’s history in the few-shot block.
Represent messages as structured objects — never concatenate roles into one ambiguous string in production:
from typing import Literal, TypedDict
class Message(TypedDict):
role: Literal["system", "user", "assistant"]
content: str
def build_sentiment_messages(text: str) -> list[Message]:
return [
{
"role": "system",
"content": (
"You label sentiment. Reply with exactly one of: "
"Positive, Neutral, Negative. If unclear, Neutral."
),
},
{
"role": "user",
"content": f'Text:\n"""{text}"""\nSentiment:',
},
]
print(build_sentiment_messages("I think the food was okay.")[1]["content"])
Few-shot as prior turns (keeps the final user turn clean):
def with_few_shots(text: str) -> list[Message]:
shots = [
("Loved the quick refund.", "Positive"),
("Package arrived damaged and late.", "Negative"),
]
msgs: list[Message] = [
{
"role": "system",
"content": "Classify sentiment. One word: Positive, Neutral, or Negative.",
}
]
for example, label in shots:
msgs.append({"role": "user", "content": f'Text: """{example}"""'})
msgs.append({"role": "assistant", "content": label})
msgs.append({"role": "user", "content": f'Text: """{text}"""'})
return msgs
Illustrative request body (no live API):
payload = {
"model": "chat-mid",
"messages": build_sentiment_messages("Service was fine."),
"temperature": 0.2,
}
# requests.post(url, json=payload, headers={"Authorization": "Bearer ..."})
Prompt lint for empty instruction or missing delimiters:
def lint_user_prompt(prompt: str) -> list[str]:
issues = []
if '"""' not in prompt and "<<<" not in prompt:
issues.append("Consider fencing untrusted input.")
if len(prompt) < 20:
issues.append("Prompt looks underspecified.")
return issues
Prompting is a structured contract — instruction, context, input, and output shape — delivered through stable chat roles so the model can do the job you meant.
`)