A vector database (or vector index inside a general store) saves embeddings and answers "which stored vectors are closest to my query?" fast enough for interactive RAG. At ten documents, a simple loop in numpy is fine. At ten million chunks, you need indexing, filtering, and operational discipline.
Think of a library shelved by "topic coordinates." When a question arrives, the librarian walks to the right region instead of reading every spine.
| Approach | Plain-English idea | Trade-off |
|---|---|---|
| Brute-force KNN (k-nearest neighbors) | Compare query to every vector | Exact but slow at scale |
| ANN (approximate nearest neighbor) | Skip most vectors using smart indexes | Much faster; may miss a true neighbor |
Common ANN methods include HNSW (hierarchical navigable small world graph), IVF (inverted file index), and PQ (product quantization). You trade a little recall for huge speed and memory savings.
id → vector plus metadata (source, tenant, updated_at).top_k + optional metadata filter (tenant = acme AND lang = en).| Index | Plain-English idea | Good when |
|---|---|---|
| Flat / brute force | Compare to every vector | Small corpora, recall baselines |
| HNSW | Layered graph of near neighbors | Strong recall/latency on CPU; RAM-heavy |
| IVF | Cluster vectors; search only nearby buckets | Millions of vectors with tuning |
| PQ | Compress vectors into short codes | Tight memory budgets |
Essential for multi-tenant apps and access control. Always enforce authorization in your app layer too—the index filter is necessary but not sufficient if misconfigured.
Pinecone, Weaviate, Milvus, Qdrant, Chroma (dev/light), pgvector. Choose based on ops model, filter strength, hybrid search, and cost at your scale.
A model change requires full re-embed. Track index build time, recall@k vs a flat baseline, p95 query latency, and how stale upserts are. Store canonical text elsewhere; the vector store is an index, not your CMS.
A tiny in-memory vector store with cosine search and metadata filters.
import numpy as np
from dataclasses import dataclass
@dataclass
class Row:
id: str
vector: np.ndarray
meta: dict
class TinyVectorDB:
def __init__(self):
self.rows: list[Row] = []
def upsert(self, id: str, vector: np.ndarray, meta: dict):
v = vector / (np.linalg.norm(vector) + 1e-9)
self.rows = [r for r in self.rows if r.id != id]
self.rows.append(Row(id, v, meta))
def query(self, vector: np.ndarray, k: int = 3, where: dict | None = None):
q = vector / (np.linalg.norm(vector) + 1e-9)
scored = []
for r in self.rows:
if where and any(r.meta.get(key) != val for key, val in where.items()):
continue
scored.append((float(np.dot(q, r.vector)), r))
scored.sort(key=lambda x: x[0], reverse=True)
return scored[:k]
db = TinyVectorDB()
rng = np.random.default_rng(2)
db.upsert("a", rng.normal(size=4), {"tenant": "acme", "topic": "hr"})
db.upsert("b", rng.normal(size=4), {"tenant": "acme", "topic": "eng"})
db.upsert("c", rng.normal(size=4), {"tenant": "other", "topic": "hr"})
hits = db.query(rng.normal(size=4), k=2, where={"tenant": "acme"})
print([(score, r.id, r.meta["topic"]) for score, r in hits])
Production APIs mirror upsert / query; the difference is ANN structures, durability, and distributed filters.
tenant in the query leaks another customer's chunks into the prompt.top_k blows token cost; the DB is happy, your bill is not.Vector databases index embeddings for fast approximate similarity search with metadata filters so RAG can retrieve the right chunks at interactive latency.