Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

cohere-ragbench

Benchmark four RAG retrieval strategies head-to-head using Cohere's full API stack — Embed v3, Rerank v3, and Command R+ — with automated LLM-as-judge evaluation.

Strategy     Faithfulness   Relevance   Context   Mean ↑
──────────────────────────────────────────────────────────
multi-query  0.91           0.88        0.85      0.88 🥇
hyde         0.87           0.85        0.82      0.85
reranked     0.83           0.80        0.79      0.81
naive        0.71           0.68        0.65      0.68

Why this exists

Most RAG tutorials implement exactly one retrieval strategy and call it done. In practice, the right strategy depends on your query distribution, corpus size, and latency budget. This tool makes it trivial to run a controlled comparison so you can make the choice with data.

Strategies

Strategy How it works When to use
Naive Embed query → cosine ANN search → generate Baseline; simplest latency profile
Reranked Naive retrieval → Cohere Rerank v3 cross-encoder Most queries; best precision/latency tradeoff
HyDE Generate hypothetical answer → embed that → search Complex questions with lexical gap between query and docs
Multi-Query Decompose → N parallel retrievals → union → rerank Multi-faceted questions; improves recall

Why input_type matters for Embed v3

Cohere's embed-english-v3.0 requires an input_type parameter (search_document vs search_query). These map to different representation heads trained with a contrastive objective — mixing them degrades retrieval quality. This library enforces the distinction throughout.

Reranking as a second stage

Reranking is a cross-encoder: it attends jointly over the query + each candidate, giving far richer interaction signals than bi-encoder dot products. The two-stage pattern (bi-encoder first pass, cross-encoder second pass) is the standard approach for production retrieval systems.

HyDE

Gao et al. 2022 showed that embedding a hypothetical answer to the query bridges the lexical gap between short queries and long passages. The hypothetical answer lives in the same vector subspace as real documents, giving better nearest-neighbour geometry.

Multi-Query

Decomposing a complex question into sub-queries expands the coverage of the embedding-space neighbourhood, improving recall for multi-faceted questions. The union of retrieved chunks is reranked against the original query.

Installation

git clone https://github.com/YOUR_USERNAME/cohere-ragbench
cd cohere-ragbench
pip install -e .
export COHERE_API_KEY=your_key_here

Quick start

# Index your documents (one chunk per line)
ragbench index --file my_corpus.txt --db bench.db

# Run all strategies and compare
ragbench run --db bench.db "How does grouped query attention reduce memory?"

# Run specific strategies only
ragbench run --db bench.db --strategies naive,reranked "..."

# Export results to JSON
ragbench run --db bench.db --json results.json "..."

Or run the built-in example (transformer architecture QA):

python examples/transformers_qa.py

Architecture

ragbench/
├── embedder.py         # Cohere Embed v3 with input_type + batching
├── reranker.py         # Cohere Rerank v3 cross-encoder
├── store.py            # SQLite vector store (no external vector DB needed)
├── evaluator.py        # LLM-as-judge scoring (faithfulness / relevance / context)
├── strategies/
│   ├── base.py         # Abstract RAGStrategy
│   ├── naive.py
│   ├── reranked.py
│   ├── hyde.py
│   └── multi_query.py
└── cli.py              # Typer CLI (ragbench index / ragbench run)

The vector store uses SQLite with embeddings serialised as float32 BLOBs and cosine search done in numpy. This keeps the dependency list minimal (no Pinecone, Weaviate, etc.) and makes the project portable.

Limitations & next steps

  • Self-evaluation bias: LLM-as-judge uses Command R, the same model family as generation. An independent judge model would be more rigorous.
  • Scale: The numpy cosine search is O(N·D). For corpora >100k chunks, swap store.py for a proper ANN index (FAISS, hnswlib).
  • Latency tracking: Adding wall-clock timing per strategy would make the precision/latency tradeoff quantitative.
  • Streaming: Command R+ supports streaming; a streaming mode would make the CLI feel faster.

License

MIT

About

Benchmark RAG strategies head-to-head using Cohere Embed v3, Rerank v3, and Command R+

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages