Benchmark four RAG retrieval strategies head-to-head using Cohere's full API stack — Embed v3, Rerank v3, and Command R+ — with automated LLM-as-judge evaluation.
Strategy Faithfulness Relevance Context Mean ↑
──────────────────────────────────────────────────────────
multi-query 0.91 0.88 0.85 0.88 🥇
hyde 0.87 0.85 0.82 0.85
reranked 0.83 0.80 0.79 0.81
naive 0.71 0.68 0.65 0.68
Most RAG tutorials implement exactly one retrieval strategy and call it done. In practice, the right strategy depends on your query distribution, corpus size, and latency budget. This tool makes it trivial to run a controlled comparison so you can make the choice with data.
| Strategy | How it works | When to use |
|---|---|---|
| Naive | Embed query → cosine ANN search → generate | Baseline; simplest latency profile |
| Reranked | Naive retrieval → Cohere Rerank v3 cross-encoder | Most queries; best precision/latency tradeoff |
| HyDE | Generate hypothetical answer → embed that → search | Complex questions with lexical gap between query and docs |
| Multi-Query | Decompose → N parallel retrievals → union → rerank | Multi-faceted questions; improves recall |
Cohere's embed-english-v3.0 requires an input_type parameter (search_document vs search_query). These map to different representation heads trained with a contrastive objective — mixing them degrades retrieval quality. This library enforces the distinction throughout.
Reranking is a cross-encoder: it attends jointly over the query + each candidate, giving far richer interaction signals than bi-encoder dot products. The two-stage pattern (bi-encoder first pass, cross-encoder second pass) is the standard approach for production retrieval systems.
Gao et al. 2022 showed that embedding a hypothetical answer to the query bridges the lexical gap between short queries and long passages. The hypothetical answer lives in the same vector subspace as real documents, giving better nearest-neighbour geometry.
Decomposing a complex question into sub-queries expands the coverage of the embedding-space neighbourhood, improving recall for multi-faceted questions. The union of retrieved chunks is reranked against the original query.
git clone https://github.com/YOUR_USERNAME/cohere-ragbench
cd cohere-ragbench
pip install -e .
export COHERE_API_KEY=your_key_here# Index your documents (one chunk per line)
ragbench index --file my_corpus.txt --db bench.db
# Run all strategies and compare
ragbench run --db bench.db "How does grouped query attention reduce memory?"
# Run specific strategies only
ragbench run --db bench.db --strategies naive,reranked "..."
# Export results to JSON
ragbench run --db bench.db --json results.json "..."Or run the built-in example (transformer architecture QA):
python examples/transformers_qa.pyragbench/
├── embedder.py # Cohere Embed v3 with input_type + batching
├── reranker.py # Cohere Rerank v3 cross-encoder
├── store.py # SQLite vector store (no external vector DB needed)
├── evaluator.py # LLM-as-judge scoring (faithfulness / relevance / context)
├── strategies/
│ ├── base.py # Abstract RAGStrategy
│ ├── naive.py
│ ├── reranked.py
│ ├── hyde.py
│ └── multi_query.py
└── cli.py # Typer CLI (ragbench index / ragbench run)
The vector store uses SQLite with embeddings serialised as float32 BLOBs and cosine search done in numpy. This keeps the dependency list minimal (no Pinecone, Weaviate, etc.) and makes the project portable.
- Self-evaluation bias: LLM-as-judge uses Command R, the same model family as generation. An independent judge model would be more rigorous.
- Scale: The numpy cosine search is O(N·D). For corpora >100k chunks, swap
store.pyfor a proper ANN index (FAISS, hnswlib). - Latency tracking: Adding wall-clock timing per strategy would make the precision/latency tradeoff quantitative.
- Streaming: Command R+ supports streaming; a streaming mode would make the CLI feel faster.
MIT