A vector index is a powerful hammer. Hitting it on every raw chat message is how you leak tenants, burn embed+ANN budget on duplicate questions, and retrieve from the wrong corpus. A retrieval gateway is the thin control plane in front of search: authenticate, authorize, shape the query, maybe answer from cache, then route to the right index with the right filters.
Treat retrieval like a privileged API, not a library call sprinkled through prompt templates. The gateway’s job is four verbs: who, what, whether, where.
Order matters:
authenticate → authorize (tenant, roles, doc ACL) → embed / cache / search → generate
Never embed first and “filter later” in the prompt. Metadata filters in the vector DB are necessary; they are not a substitute for rejecting unauthenticated calls. Log subject_id, tenant_id, and policy_version on every retrieval span so incidents are reconstructible.
Multi-tenant rule of thumb: tenant_id is a mandatory filter, enforced in gateway code that builds the search request — not an optional argument the model can omit.
Cache keys are often the query embedding (or a normalized string + embedding). On a new query, if cosine(q, cached_q) >= threshold (and same tenant, locale, product surface), return the cached answer or cached chunks.
Tune on a labeled pair set: same-intent vs different-intent. Optimize for low false-hit rate first; accept a lower hit rate. Invalidate on corpus or policy change (TTL alone is not enough when docs update). For regulated answers, cache chunk ids + generation inputs, not only final prose, so you can rebuild with the new model.
Before search:
Shaping is a small model or rules + templates. Cap rewrite length; log both raw and shaped queries for eval.
Not every query should hit the same top-k dense search:
| Signal | Route |
|---|---|
| Exact FAQ / policy id | Keyword or id lookup |
| Navigational “section 4.2” | Structured doc store |
| Open semantic question | Dense / hybrid ANN |
| Multi-entity comparison | Multi-hop planner |
| Toolable (“order status 123”) | API tool, not RAG |
A lightweight classifier or rules router sits in the gateway. Wrong route is a quality bug: dense search over a tracking-number question invents prose instead of calling the order API.
A sketch of gateway decisions: ACL gate, cache with a conservative threshold, and a tiny router.
from dataclasses import dataclass
import numpy as np
@dataclass
class Principal:
user_id: str
tenant_id: str
roles: set[str]
@dataclass
class CacheEntry:
query_vec: np.ndarray
tenant_id: str
answer: str
CACHE: list[CacheEntry] = []
THRESHOLD = 0.92 # tune up if false hits appear
def cosine(a, b):
return float(np.dot(a, b) / ((np.linalg.norm(a) * np.linalg.norm(b)) + 1e-9))
def authorize(p: Principal, resource_tenant: str) -> bool:
return p.tenant_id == resource_tenant and "reader" in p.roles
def shape(raw: str) -> str:
q = raw.strip().lower()
for noise in ("please ", "hey ", "can you "):
if q.startswith(noise):
q = q[len(noise):]
return q
def cache_lookup(tenant_id: str, qvec: np.ndarray) -> str | None:
best, best_s = None, -1.0
for e in CACHE:
if e.tenant_id != tenant_id:
continue
s = cosine(qvec, e.query_vec)
if s > best_s:
best, best_s = e, s
if best is not None and best_s >= THRESHOLD:
return best.answer
return None # miss — safer than a weak hit
def route(shaped: str) -> str:
if shaped.startswith("order status") or shaped[:3].isdigit():
return "orders_api"
if "section" in shaped and any(c.isdigit() for c in shaped):
return "structured_docs"
return "hybrid_rag"
def gateway(p: Principal, raw_query: str, embed_fn, search_fn, generate_fn):
if not authorize(p, p.tenant_id):
raise PermissionError("forbidden")
shaped = shape(raw_query)
qvec = embed_fn(shaped)
hit = cache_lookup(p.tenant_id, qvec)
if hit is not None:
return {"source": "cache", "answer": hit}
path = route(shaped)
if path != "hybrid_rag":
return {"source": path, "answer": f"delegate:{path}:{shaped}"}
chunks = search_fn(qvec, tenant_id=p.tenant_id) # filter mandatory
answer = generate_fn(shaped, chunks)
CACHE.append(CacheEntry(qvec, p.tenant_id, answer))
return {"source": "rag", "answer": answer}
Wire real embed/search clients; keep authorize and tenant filters in this layer, not in the prompt.
Track auth denies, cache hit rate, sampled false-hit rate, route mix, and empty-result rate. Start with tenant filter + FAQ string cache; add semantic cache and multi-route when traces show duplicate traffic and intent diversity.
A retrieval gateway authenticates first, uses a conservative semantic cache (false hits hurt more than misses), shapes queries, and routes to the right index or tool with tenant filters enforced in code.