Combine keyword and vector search
Fuse keyword and semantic rankings, and filter by metadata and permissions the right way.
- Explain when keyword and semantic search each win
- Merge rankings with reciprocal rank fusion
- Apply metadata and permission filters before top-k, not after
Keyword search and semantic search fail in opposite ways: keywords miss paraphrases (“money back” vs “refund”), embeddings blur exact strings (“E4012” vs “E4021”). Hybrid search runs both and merges the results - and it usually beats either alone.
Their scores aren’t comparable (a BM25 score of 7.3 vs a cosine of 0.82), so the simplest robust merge ignores scores and uses ranks: reciprocal rank fusion (RRF, Cormack et al. 2009) gives each document 1 / (k + rank) from every list it appears in, with k = 60 by convention, and sorts by the sum.
Try it
Which search wins?
For each query, which approach is more likely to find the right passage?
““error E4012””
““how do I get my money back””
““SKU 88-1045-B dimensions””
““my parcel never showed up””
““Kubernetes 1.31 release notes””
““can I bring my dog to work?” (policy says “pets”)”
Filters: before, not after
Real queries come with constraints: this customer’s tenant, documents the user is allowed to see, the right language or product version, recent dates. Store them as chunk metadata and filter during search.
Filter before picking the top k. If you take the top 10 overall and then drop chunks the user can’t see, you may be left with one result - or none - even though plenty of allowed, relevant chunks existed further down.
1keyword = ["d3", "d1", "d2"]
2semantic = ["d1", "d4", "d3"]
3scores = {}
4for ranking in (keyword, semantic):
5 for rank, doc in enumerate(ranking, start=1):
6 scores[doc] = scores.get(doc, 0) + 1 / (60 + rank)
7print(max(scores, key=scores.get))d1
Key takeaways
Keywords catch exact identifiers; embeddings catch paraphrases - hybrid gets both.
Reciprocal rank fusion merges rankings without comparing incompatible scores.
Apply metadata and permission filters before top-k selection.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Reciprocal rank fusion
Each input line is name: id id id ..., a ranked list. Fuse them with RRF (k = 60, ranks from 1) and print id score (4 decimals) for every document, highest first, ties by id.
- Keyword and vector lists
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Filter before top-k
The first line is tenant k. Each following line is id tenant score. Print post-filter: IDS (top k overall, then keep the tenant’s) and pre-filter: IDS (keep the tenant’s, then top k), both in score order, or none.
- Acme’s chunks
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…