Some questions die in a single search. "Which teams own services that still call the deprecated Auth v1 API, and what is their on-call?" needs a hop: find services using Auth v1, then look up ownership, then on-call. Multi-hop retrieval chains those lookups. Agentic RAG lets an LLM decide the next query, tool, or stop condition instead of a fixed one-shot pack.
| Pattern | Plain-English idea |
|---|---|
| Single-hop RAG | One library visit—retrieve once, answer |
| Multi-hop RAG | Research session—read, realize you need more, search again |
| Agentic RAG | Model plans what to search, which tool to use, whether to search again |
Not every question deserves an agent. Hopping increases latency, cost, and failure modes. Use it when the answer structurally depends on intermediate entities.
Two bad assumptions:
Example task: find the error in the payment container, check if it relates to yesterday's auth commit, draft a Slack update for QA. One embedding pass gets confused; the system needs to search, inspect, decide what is missing, search again, then generate.
| Pattern | Plain-English idea |
|---|---|
| Decomposition | Split into sub-questions |
| Seed-and-expand | Retrieve once; extract entity IDs; targeted follow-up queries |
| CoRAG (chain-of-RAG) | Chain of sub-questions and sub-answers; each step guides the next |
| Pattern | Plain-English idea |
|---|---|
| ReAct | Interleave reasoning traces with search/tool actions |
| Router | Classify single-hop vs multi-hop vs SQL tool vs refuse |
| Planner–executor | Planner emits a JSON plan; executor runs steps |
Minimal two-hop toy: find services on Auth v1, then look up owners.
SERVICES = [
{"id": "svc_billing", "text": "billing-api still calls auth-v1 verifyToken"},
{"id": "svc_edge", "text": "edge-gateway migrated to auth-v2"},
{"id": "svc_jobs", "text": "jobs-worker uses auth-v1 for batch auth"},
]
OWNERS = {
"svc_billing": "Payments Team — oncall +payments",
"svc_jobs": "Data Platform — oncall +data-plat",
}
def retrieve(corpus, query: str, k: int = 5):
q = set(query.lower().split())
scored = [(len(q & set(r["text"].lower().split())), r) for r in corpus]
scored.sort(reverse=True)
return [r for s, r in scored[:k] if s > 0]
hop1 = retrieve(SERVICES, "auth-v1")
hop2 = [{"id": h["id"], "owner": OWNERS[h["id"]]}
for h in hop1 if h["id"] in OWNERS]
print("hop1:", [h["id"] for h in hop1])
print("hop2:", hop2)
Multi-hop and agentic retrieval chain planned searches when one lookup cannot gather all entities, under strict budgets and citation discipline.