Blog/vector-search/Do you need a vector database? Brute-force search benchmarked from 10k to 1M vectors
Do you need a vector database? Brute-force search benchmarked from 10k to 1M vectors
Exact numpy search over 100k embeddings takes 3.5 ms on 2 vCPUs. Measured latency, memory and HNSW recall to show when an index actually pays off.
What this post covers
How this was tested
2 vCPU Intel Xeon @ 2.1 GHz, 7 GB RAM, Python 3.11, numpy 2.4.4, faiss-cpu 1.15.1, hnswlib. Synthetic normalized float32 vectors. Median of 50 (exact) or 200 (HNSW) single queries.
Most RAG tutorials start by installing a vector database. For a lot of projects, that is a dependency you don’t need yet. I measured how long plain exact search takes as the corpus grows, and what an HNSW index buys you in return.
Short answer: below about 100,000 chunks, a numpy matrix multiply is fast enough for almost any RAG app. Between 100k and 1M it depends on your traffic. Past that, use an index.
The setup
Every embedding is normalized, so cosine similarity is a dot product. Exact search is one matrix-vector product and a partial sort:
import numpy as np
def top_k(X: np.ndarray, q: np.ndarray, k: int = 10) -> np.ndarray:
"""X: (n, d) normalized float32 matrix. q: (d,) normalized query."""
scores = X @ q
idx = np.argpartition(-scores, k)[:k]
return idx[np.argsort(-scores[idx])]
I timed this against faiss.IndexFlatIP (also exact) for three common embedding sizes: 384 (MiniLM-class models), 768 (BERT-base-class) and 1536 (OpenAI text-embedding-3-small default). The machine is deliberately small: 2 vCPUs, the size of a cheap cloud instance.
Exact search latency
Single-query latency, median of 50 queries:
| Dim | Vectors | RAM for vectors | numpy | faiss Flat |
|---|---|---|---|---|
| 384 | 10,000 | 15 MB | 0.35 ms | 0.78 ms |
| 384 | 100,000 | 154 MB | 3.5 ms | 8.3 ms |
| 384 | 500,000 | 768 MB | 31 ms | 66 ms |
| 384 | 1,000,000 | 1.5 GB | 62 ms | 122 ms |
| 768 | 100,000 | 307 MB | 10.6 ms | 22.6 ms |
| 768 | 500,000 | 1.5 GB | 57 ms | 120 ms |
| 768 | 1,000,000 | 3.1 GB | 114 ms | 240 ms |
| 1536 | 10,000 | 61 MB | 1.2 ms | 3.1 ms |
| 1536 | 100,000 | 614 MB | 20 ms | 51 ms |
| 1536 | 500,000 | 3.1 GB | 107 ms | 248 ms |
Three things stand out:
- Latency scales linearly with
n × d. It is a memory-bandwidth problem. Doubling either the corpus or the dimension doubles the time. - Memory is the real limit, not speed. 1M vectors at 1536 dims is 6 GB of float32 before any metadata. I skipped that row: it won’t fit comfortably on a 7 GB machine.
- Plain numpy beat faiss Flat for single queries here, by about 2×. Faiss is built for batches; one query at a time pays overhead that numpy’s direct matrix-vector call doesn’t. If you batch queries, test again before assuming this holds.
What an HNSW index buys you
Next, I built hnswlib indexes (M=16, ef_construction=200) and measured query time and recall@10: the fraction of the true top 10 the index actually returns.
I used two synthetic datasets, because recall depends heavily on how your data is shaped:
- Uniform: random Gaussian directions. This is the worst case for any approximate index: no structure to exploit.
- Clustered: 2,000 tight clusters. Real embeddings sit between these two, usually much closer to clustered, since documents about the same topic land near each other.
| Data | Dim | Vectors | Build time | ef=32 | ef=64 | ef=128 |
|---|---|---|---|---|---|---|
| clustered | 384 | 100k | 23 s | 0.10 ms, recall 0.954 | 0.17 ms, 0.995 | 0.34 ms, 1.000 |
| clustered | 768 | 100k | 48 s | 0.22 ms, 0.878 | 0.36 ms, 0.985 | 0.68 ms, 1.000 |
| uniform | 384 | 100k | 37 s | 0.16 ms, 0.129 | 0.31 ms, 0.206 | 0.55 ms, 0.281 |
| uniform | 768 | 100k | 69 s | 0.34 ms, 0.058 | 0.63 ms, 0.105 | 1.07 ms, 0.172 |
| uniform | 384 | 500k | 279 s | 0.19 ms, 0.034 | 0.35 ms, 0.057 | 0.67 ms, 0.099 |
HNSW queries are 10 to 100 times faster than exact search at these sizes. The cost:
- Build time. 500k vectors took over 4.5 minutes on 2 cores. You pay this again on every full re-index, such as after switching embedding models.
- Recall is not guaranteed. On structured data,
ef=64gave near-perfect recall. On structureless data, the same settings returned less than a quarter of the true neighbors. Your data will be somewhere in between, so measure recall on your own embeddings before trusting the defaults.
Where the time actually goes in a RAG request
A typical RAG request spends 1 to 5 seconds waiting for the LLM to generate. Next to that:
- 3.5 ms (100k × 384, exact) is invisible.
- 60 to 110 ms (1M exact) is noticeable only if you care about p99, or serve many queries per second on the same box.
- Exact search gives perfect recall for free, which removes one variable when you debug bad answers.
A rough sizing rule: one PDF page is usually 2 to 3 chunks. 100k chunks is roughly 30,000 to 50,000 pages. Many internal “chat with our docs” projects never get there.
My rule of thumb
| Your corpus | What I’d use |
|---|---|
| under 100k chunks | numpy (or pgvector with no index). Store vectors in a .npy file or Postgres. |
| 100k to 1M chunks | exact search is still fine at low QPS. Add HNSW when p95 latency or QPS demands it, and measure recall. |
| over 1M chunks, or many tenants | a proper ANN index (pgvector HNSW, Qdrant, Milvus, OpenSearch), plus filtering and sharding. |
Most importantly, keep the vector store behind one interface so you can swap numpy for pgvector later without touching the rest of the pipeline. That one seam is worth more than picking the “right” database on day one.
Reproduce it
The exact-search benchmark:
import time, numpy as np, faiss
faiss.omp_set_num_threads(2)
rng = np.random.default_rng(0)
norm = lambda x: x / np.linalg.norm(x, axis=1, keepdims=True)
for d in (384, 768, 1536):
for n in (10_000, 100_000, 500_000, 1_000_000):
if n * d * 4 > 3.2e9: continue # skip what won't fit in RAM
X = norm(rng.standard_normal((n, d), dtype=np.float32))
Q = norm(rng.standard_normal((50, d), dtype=np.float32))
t = []
for q in Q:
s = time.perf_counter()
sc = X @ q
idx = np.argpartition(-sc, 10)[:10]
t.append(time.perf_counter() - s)
print(d, n, f"{np.median(t)*1e3:.2f} ms")
For HNSW, build with hnswlib.Index(space="ip", dim=d), init_index(max_elements=n, ef_construction=200, M=16), and compare knn_query results against the exact top 10 to get recall.
Caveat: these are synthetic vectors on one small machine. Latency for exact search doesn’t depend on what the vectors mean, so those numbers transfer well. HNSW recall does depend on your data, which is exactly why you should measure it yourself.
Want the production version of this?
A production RAG + MCP starter kit for FastAPI: hybrid search, validated citations, evals in CI, Docker, 32 tests. It runs offline in about 60 seconds.
See ShipRAG or get the free RAG checklist first.