Evaluate and tune hybrid search
Compare MiniLM, SPLADE, and hybrid retrieval on SQuAD, select weights on tuning questions, and validate on held-out articles.
Use this workflow to select a dense-to-sparse ratio from measured retrieval results. It separates weight selection from evaluation and keeps both single-model scoring baselines eligible to win.
Start with the hybrid-search example, which installs the maintained SDK, runs MiniLM and SPLADE, and verifies retrieval on a small corpus. This benchmark uses a separate space and takes longer because it embeds 10,250 sentences and runs thousands of queries.
Release pending: The results below were measured on a local build with PR #1844, scheduled for v1.0.325. Run this benchmark against a corrected build. The released v1.0.324 can silently omit scores during fusion and is unsuitable for reproducing it.
Prepare an isolated benchmark
From the example's hybrid-search directory, with its virtual environment and
GoodMem credentials loaded:
DENSE_ID=$(python -c 'import json; print(json.load(open("run/resources.json"))["dense"])')
SPARSE_ID=$(python -c 'import json; print(json.load(open("run/resources.json"))["sparse"])')
python setup.py --output-dir run/benchmark --name 'SQuAD hybrid benchmark' \
--dense-id "$DENSE_ID" --sparse-id "$SPARSE_ID"
SPACE_ID=$(python -c 'import json; print(json.load(open("run/benchmark/resources.json"))["space_id"])')
mkdir -p squad_data
curl -fL https://raw.githubusercontent.com/rajpurkar/SQuAD-explorer/master/dataset/dev-v1.1.json \
-o squad_data/dev-v1.1.json
python insert_squad_sentences_goodmem.py --space-id "$SPACE_ID" --dry-run
python insert_squad_sentences_goodmem.py --space-id "$SPACE_ID" \
--output run/benchmark/manifest.json --sample-size 1000 --seed 42The loader:
- Splits SQuAD paragraphs into sentences using the bundled sentence breaker.
- Maps each annotated answer span to a sentence that contains the complete answer. It accepts any mapped annotator answer for a question and counts unmappable questions instead of inventing ground truth.
- Stores each sentence as one memory with further chunking disabled.
- Shuffles article titles deterministically, assigns half to tuning and half to testing, and samples 1,000 questions from each group.
- Checks every page of memory processing status and waits for the entire corpus to complete. Failed processing stops the command.
The manifest records the dataset hash, corpus, question IDs, split, seed, registered models, and server version. Rerunning the loader resumes by stable memory IDs; it rejects unexpected memories or a conflicting manifest.
Both tuning and test articles are in the searchable corpus. This is required for answer retrieval. Only the tuning questions are used to select weights; test questions come from different articles and remain unused until selection is complete.
Search the ratio space
python optimize_embedder_weights.py --manifest run/benchmark/manifest.json \
--output-dir run/benchmark/search --threads 4The search fixes the dense coefficient at 1 and varies the sparse coefficient. The initial grid includes dense-only, sparse-only, 1:0.005, 1:1, and logarithmically spaced sparse coefficients from 0.00001 through 100. One nine-point logarithmic refinement explores the neighborhood of the best tuning result. These are finite searches: the output is the best tested ratio, not a global optimum.
selection.json freezes the winning ratio before test queries are issued. The
test phase compares that choice with dense-only, sparse-only, 1:0.005, and 1:1.
The script does not change its choice based on those test results. If scores tie
during tuning, it prefers a single-model score over a hybrid score.
Both embedders still generate candidates when one coefficient is zero. Thus “dense-only” and “sparse-only” here mean scoring weights of 1:0 and 0:1 within the same two-embedder space. This controls candidate discovery while comparing weights; it does not measure an installation with only one model.
To change the search grid for a new experiment:
python optimize_embedder_weights.py --manifest run/benchmark/manifest.json \
--output-dir run/another-search --ratios '0,0.001,0.003,0.005,0.01,0.03,0.1,1,sparse' \
--fine-points 9Once you inspect held-out results, do not repeatedly adjust the grid against the same test set and continue calling it held out. Use new test data for subsequent model or search-design choices.
Interpret the output
- MRR@10: average reciprocal rank of the first accepted answer sentence, with zero for questions whose answer is absent from the first ten results.
- Hit@k: fraction of questions with at least one accepted answer sentence in
the first
kresults. Because some questions have multiple accepted sentences, this is an answer hit rate rather than document recall. - Coverage: fraction of queries returning any results. A query that returns ten wrong sentences has full coverage and zero MRR.
The evaluator fails on API errors, partial/warning statuses, incomplete streams, or invalid weight IDs. Those are operational failures, not retrieval misses. It never reports Hit@10 from a run that requested only five results.
The held-out comparison joins results by question ID. It reports a paired article-cluster bootstrap interval for the MRR difference: whole articles are resampled so questions from the same article remain together. The seed and number of resamples are recorded. An interval spanning zero does not establish an improvement. It is not evidence that the configurations are equivalent.
These intervals describe sampling uncertainty in this experiment. They do not prove a production improvement, account for all sources of dataset bias, or replace a benchmark drawn from your own queries.
Measured results
The best tested ratio was 1:0.168702397557, or roughly 1:0.17. It was selected on the tuning split; all numbers below are from the separate test split.
| Scoring configuration | MRR@10 | Hit@1 | Hit@5 | Hit@10 |
|---|---|---|---|---|
| Dense-only scoring (1:0) | 0.7069 | 60.7% | 83.9% | 90.2% |
| Sparse-only scoring (0:1) | 0.8117 | 74.0% | 90.3% | 92.6% |
| Previous ratio (1:0.005) | 0.7553 | 66.8% | 87.2% | 91.9% |
| Equal weights (1:1) | 0.8122 | 74.0% | 90.3% | 92.7% |
| Selected ratio (about 1:0.17) | 0.8147 | 74.2% | 90.7% | 92.9% |
All five configurations returned results for every test query (100% coverage). The searchable corpus contained 10,250 sentences from 48 SQuAD articles. The experiment used 1,000 tuning questions from 24 articles and 1,000 test questions from the other 24. One question had no complete answer span in a sentence and was excluded before sampling.
The tuning curve is fairly flat around its peak. The exact grid winner is recorded for reproduction, not because twelve decimal places are meaningful for another application.
Compared with 1:0.005, the selected ratio improved test MRR@10 by 0.0593 (95% paired article-cluster bootstrap interval 0.0434 to 0.0714). Compared with dense-only scoring, the difference was 0.1078 (0.0911 to 0.1209).
The comparison with sparse-only scoring is less decisive: 0.0030 MRR, with an interval of −0.0019 to 0.0091. This experiment therefore supports replacing 1:0.005 with about 1:0.17 for this hybrid setup, but does not establish an improvement over SPLADE-only scoring. If your own evaluation shows the same pattern, evaluate whether operating both models is worthwhile; measure a true single-embedder deployment separately before drawing latency or cost conclusions.
The measurements used MiniLM and SPLADE with pinned model revisions, TEI 1.9.4
on CPU, one sentence per memory, no postprocessor, requested_size=10, and the
server's default HNSW settings. Ratios were selected on a local build containing
the fusion fix, which is
scheduled for GoodMem v1.0.325; that release is still pending. Do not compare these results with an uncorrected
v1.0.324 run: that version could silently omit per-embedder scores during fusion.
The run summary records all tuning points, held-out metrics, confidence intervals, model revisions, dataset hashes, and server-build provenance. The paired question ranks allow the reported comparisons to be checked without redistributing the corpus.
Inspect misses
Each successful evaluation writes its weights, metrics, question IDs, ranks,
and retrieved memory IDs to a JSON file under search/evaluations/. The filenames
are fingerprints of the full run configuration, not abbreviated embedder IDs.
Choose the exact evaluation to inspect:
python analyze_missing_gt_similarity.py --manifest run/benchmark/manifest.json \
--evaluation PATH_TO_EVALUATION_JSON --limit 10 --search-size 100 \
--output run/benchmark/misses.jsonFor each miss, the analyzer searches with the selected weights and each model alone. It shows returned text, scores, accepted memory IDs, and the first accepted answer's rank if found. It uses the same GoodMem API and embedders as the experiment; no direct database access or separate query model is needed.
A larger requested size can change candidate discovery and ranking. The analysis explains what these searches returned; it does not compute an exact full-corpus rank or compare raw scores from different models as though they shared a scale.
Reproduce and adapt the experiment
The example source
includes pinned dependencies, the Compose environment, all scripts, and regression
tests. Keep the manifest, selection.json, summary.json, and evaluation files
with your results. Completed evaluations can resume from matching caches; changed
questions, weights, retrieval depth, or manifests use different cache keys.
Create a new run when changing the corpus, model weights, chunking, or server configuration. Replace SQuAD with representative questions and accepted answers from your application before adopting the measured ratio as a default.