MongoDB Vector Search · benchmark · August 2026
Does MongoDB Vector Search’s binary quantization hold at 75 million vectors?
MongoDB’s public Vector Search benchmark reaches 15.3 million vectors. I wanted to know what happens much further out—when MongoDB Vector Search runs approximate search over 75 million real-world embeddings but can rescore at most 10,000 candidates.
On this page
01 · The benchmark gap
Fifteen million vectors does not answer the 75-million-vector question
MongoDB’s large public Vector Search benchmark uses 15.3 million, 2048-dimensional embeddings and reports 90–95% recall with quantization. That is useful evidence, but it is a different regime from 75 million vectors at 512 dimensions.
The part that bothered me was the fixed search budget. MongoDB Vector Search caps numCandidates at 10,000. The corpus can keep growing; the pool eligible for exact rescoring cannot. If true neighbours stop reaching that pool, no reranking stage can recover them.
At what scale does a fixed 10,000-candidate budget stop being enough—and had 75 million vectors already crossed that line?
02 · Mechanism
MongoDB Vector Search’s binary and scalar indexes fail in different places
Both MongoDB Vector Search indexes use compressed values during approximate nearest-neighbour search. The difference comes after retrieval.
Approximate retrieval, exact final ordering
Approximate retrieval, approximate final ordering
MongoDB’s binary path only has to retrieve a true neighbour into the candidate pool. MongoDB Vector Search then uses the retained full-precision vectors to repair its ordering. Its scalar path ranks on the approximation, so increasing the pool cannot necessarily buy back information already lost in that comparison.
03 · Getting the data
Seventy-five million real embeddings from public road-camera feeds
The corpus came from public traffic-camera streams available online. It combined open HLS feeds from state transport agencies with a public catalogue of HD road cameras. Across the run, 2,126 cameras on 23 hosts contributed footage from at least nine countries.
Each stream was decoded at one frame per second. YOLO detected the vehicles in each frame, the detections were cropped, and an embedding model converted every crop into a 512-dimensional vector. Capture, detection, and embedding ran as separate stages across two hosts. Sharding camera groups over four capture processes lifted the pipeline from 5.5 to a peak of 1,104 documents per second.
| network | documents | share | cameras |
|---|---|---|---|
| ARDOT · Arkansas | 30,876,071 | 41.2% | 522 |
| Caltrans · California | 16,834,275 | 22.4% | 318 |
| DelDOT · Delaware | 11,675,979 | 15.6% | 377 |
| Bison Futé · France | 3,992,793 | 5.3% | 150 |
| Weacom · Russia | 3,223,126 | 4.3% | 81 |
| Smartech LATAM · Peru | 2,634,624 | 3.5% | 31 |
| ITIC · Thailand | 2,215,975 | 3.0% | 44 |
| KazToll · Kazakhstan | 1,056,016 | 1.4% | 190 |
| SHA · Maryland | 948,796 | 1.3% | 28 |
| KT · Kyrgyzstan | 939,486 | 1.3% | 29 |
| Remaining hosts | 603,508 | 0.8% | — |
| total | 75,000,649 | 100% | 2,126 |
The collection ran for about 46 hours and produced roughly 260 GB of documents. Each record contained a normalized vehicle bounding box and a 512-dimensional embedding stored as BSON binary data.
04 · Method
Exact ground truth without 300 exhaustive MongoDB searches
The benchmark measures MongoDB Vector Search over 75 million real-world visual embeddings, anonymized here to keep the experiment independent from any product or organisation. I selected 300 fixed query vectors and stripped each query’s self-match from both the exact and approximate result sets.
recall@k = | top-k(approximate) ∩ top-k(exact) | / kRunning an exact database search once per query was impractical at this scale. Instead, the corpus was streamed once in chunks. Every chunk was scored against all queries with exact float32 cosine similarity, while each query maintained a running top-200 pool.
- Arms: MongoDB Vector Search indexes using binary and scalar quantization.
- Candidate sweep: 360, 720, 1080, 1440, 4000, and 10000.
- Primary metric: document-ID overlap at k=72—not the score reported by the index.
- Tail metrics: median, p5, p95, and per-query perfect recall.
The exact pool was validated independently by rereading stored vectors and recomputing their cosine values. Agreement was within float32 precision.
05 · Results
More candidates rescue MongoDB binary search. They barely affect scalar.
The controls below are linked across every chart. Select one candidate setting and the mean recall and query-tail readouts update together.
At 1,440 candidates, binary is 13.0 points ahead. Its p5 recall is 59.7%, compared with 19.4% for scalar.
06 · Reading the result
The mean is only half the finding
At 360 candidates, scalar begins 1.4 recall points ahead. By 720, binary has crossed it. At 10,000 candidates, binary reaches 95.8% mean recall while scalar remains at 73.1%.
The tail is more revealing. Binary’s fifth percentile rises from 40.3% to 81.9%. Scalar’s fifth percentile is exactly 19.4% at every setting. The least successful scalar queries recover the same small fraction whether the search receives 360 candidates or 10,000.
A floor that more search cannot lift points to information lost during ranking, not neighbours that merely need more exploration.
07 · Latency benchmark
Latency Benchmark
The recall experiment was designed to compare correctness, not to produce a clean latency number. Binary, scalar-quantized, and unquantized index builds shared the environment while evaluation scripts generated additional CPU, memory, and I/O work. Timing those queries would have measured the research harness as much as MongoDB Vector Search.
I therefore ran latency as a separate experiment over the same 75-million-document collection. I isolated one 512-dimensional binary index, cleared the host page cache, and then accessed only that index so unrelated vector indexes would not occupy memory. This removed the cross-index cache eviction present in the earlier test setup.
The sweep covered 13 numCandidates settings from 100 to 10,000. At each setting, I ran 100 warm queries using vectors from previously unseen documents—1,300 measured queries in total. Every query passed its stored 512-dimensional BinData vector directly as queryVector, with no decoding, no pre-filter, and limit: 72.
The latency host had 62.7 GB of RAM, 16 cores, and no swap. Its working configuration capped the mongod container at 32 GB, reduced the WiredTiger cache to 8 GB, and left mongot uncapped.
| numCandidates | p50 | p95 | p99 | max |
|---|---|---|---|---|
| 100 | 46 ms | 63 ms | 82 ms | 82 ms |
| 360 | 176 ms | 210 ms | 268 ms | 268 ms |
| 720 | 198 ms | 265 ms | 730 ms | 730 ms |
| 1,000 | 201 ms | 229 ms | 375 ms | 375 ms |
| 2,000 | 255 ms | 278 ms | 387 ms | 387 ms |
| 3,000 | 299 ms | 322 ms | 400 ms | 400 ms |
| 4,000 | 346 ms | 373 ms | 407 ms | 407 ms |
| 5,000 | 397 ms | 428 ms | 981 ms | 981 ms |
| 6,000 | 446 ms | 477 ms | 1,082 ms | 1,082 ms |
| 7,000 | 494 ms | 530 ms | 723 ms | 723 ms |
| 8,000 | 537 ms | 575 ms | 608 ms | 608 ms |
| 9,000 | 591 ms | 637 ms | 984 ms | 984 ms |
| 10,000 | 635 ms | 668 ms | 945 ms | 945 ms |
None of the 1,300 measured queries exceeded two seconds. At MongoDB’s 10,000-candidate ceiling, median latency was 635 ms—7.9× under that budget—and p95 was 668 ms. The worst observation in the full sweep was 1,082 ms at 6,000 candidates.
From 1,000 to 10,000 candidates, the median increased by 434 ms. A linear fit over that range adds 0.0481 ms per candidate, with no high-candidate inflection point.
This is a warm steady-state result, not a cold-start promise. It applies when the index is isolated, the memory allocation matches this setup, and the working set has been warmed.
The recall curve says a wider candidate pool buys accuracy. The isolated latency run says that, once warm, the same lever remained affordable on this host.
08 · The 10,000-candidate ceiling
So? Can MongoDB binary quantization still be trusted?
It depends. For this corpus and this query set, yes—at a measured operating point. The recall curve is still rising at 10,000 candidates, and its tail improves alongside its mean. There is no sign of a catastrophic ceiling at 75 million vectors.
The more general answer depends less on document count than on the geometry of the embedding space. By “dense,” I mean a neighbourhood with many vectors separated by very small similarity margins. In that setting, the distortion introduced by a 1-bit comparison is more likely to change which edges the ANN traversal follows. A true neighbour that never enters the candidate pool cannot be recovered by exact rescoring.
Distortion stays small relative to the similarity margins.
True neighbours are likely to reach the candidate pool.
The 1-bit comparison can erase fine distinctions.
True neighbours can be lost before exact rescoring.
I have tested other, denser vector sets and observed lower binary recall. That does not contradict this result; it shows why trust has to be earned on the corpus being searched. When neighbourhoods are crowded, test a more accurate representation for ANN traversal as well as a wider candidate pool.
09 · Resource tradeoff
MongoDB binary quantization buys search memory, not equivalent disk savings
MongoDB Vector Search’s compact 1-bit representation reduces the vector working set used by approximate search. But its automated binary quantization retains full-precision vectors for rescoring. Storage planning must account for both representations and the graph rather than assuming a 32× smaller index on disk.
10 · Limitations
What this benchmark does not prove
- One corpus is not a scaling law. The result does not isolate corpus size from embedding distribution.
- No query-time filters. Selective filters can materially change graph exploration.
- Geometric recall is not relevance. This measures nearest-neighbour fidelity, not whether a human prefers the result.
References