Free-form prose is for humans. Downstream code wants objects: enums, IDs, date ranges, confidence scores. Structured outputs turn the model from a storyteller into a component you can wire into queues, UI forms, and agents — but only if you validate every response like it came from an untrusted client.
Ask for “JSON” in a prompt and you will get JSON… until you do not. Trailing commas, markdown fences, renamed keys, and partially truncated objects appear under load. Reliability is a stack:
Always assume vendor features differ. Your app-side validator is the portable guarantee.
Libraries in the Pydantic family (or dataclasses + hand checks) give you:
Example conceptual schema:
SearchQuery
rewritten_query: string
published_daterange: { start: date, end: date }
domains_allow_list: list[string]
On validation failure, do not only “try again colder.” Send the error:
Your previous output failed validation: published_daterange.end is before start. Reply with corrected JSON only.
Cap retries (e.g. 2). Log failure rates per prompt version — a spike means schema or model drift.
For UX, you may accept a partial object with defaults for optional fields. For actions (refunds, deletes, emails), fail closed: no valid object → no side effect.
Schemas fail in production when they are written for lawyers instead of generators:
"Positive" | "Neutral" | "Negative") instead of free strings.status: "ok" | "unknown" beats forcing a fake value.When the model must emit dates, demand ISO-8601 in the schema description and reject everything else in the validator. When it must emit IDs, prefer copying from provided context over inventing new ones — and check membership against your database.
Agents often chain: plan JSON → tool calls → final answer JSON. Validate each stage. A pretty final paragraph that skipped a failed plan object is how silent wrong workflows ship. Persist the validated objects, not only the chat text, so support can replay what the system believed.
Stdlib validation without external deps:
from dataclasses import dataclass
from datetime import date
import json
@dataclass
class DateRange:
start: date
end: date
def __post_init__(self):
if self.end < self.start:
raise ValueError("end before start")
@dataclass
class SearchQuery:
rewritten_query: str
published_daterange: DateRange
domains_allow_list: list[str]
def parse_search_query(raw: str) -> SearchQuery:
data = json.loads(raw)
dr = data["published_daterange"]
return SearchQuery(
rewritten_query=str(data["rewritten_query"]).strip(),
published_daterange=DateRange(
start=date.fromisoformat(dr["start"]),
end=date.fromisoformat(dr["end"]),
),
domains_allow_list=[str(d) for d in data.get("domains_allow_list", [])],
)
good = '''
{"rewritten_query": "llm evaluation",
"published_daterange": {"start": "2024-01-01", "end": "2024-12-31"},
"domains_allow_list": ["arxiv.org"]}
'''
print(parse_search_query(good))
Strip accidental markdown fences before parse:
def strip_fences(text: str) -> str:
t = text.strip()
if t.startswith("```"):
lines = t.splitlines()
# drop first and last fence lines
if lines and lines[0].startswith("```"):
lines = lines[1:]
if lines and lines[-1].startswith("```"):
lines = lines[:-1]
t = "\n".join(lines).strip()
if t.lower().startswith("json"):
t = t[4:].lstrip()
return t
Retry loop sketch:
def complete_structured(call_model, schema_errors_max=2):
feedback = None
for _ in range(schema_errors_max + 1):
raw = call_model(feedback)
try:
return parse_search_query(strip_fences(raw))
except (json.JSONDecodeError, KeyError, ValueError) as e:
feedback = f"Validation error: {e}. Return corrected JSON only."
raise RuntimeError("structured_output_failed")
Illustrative API flag (vendor-specific):
payload = {
"model": "chat-mid",
"messages": [...],
"response_format": {"type": "json_object"}, # syntax, not full schema
"temperature": 0.1,
}
json.loads.user_id belongs to another tenant — validators must include authz invariants.Structured GenAI outputs need a schema, constrained generation when available, strict validation, and bounded repair — never trust free text as a typed API.