Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
105 commits
Select commit Hold shift + click to select a range
9b59d2e
[Automated Commit] Format Codebase
mlcommons-bot Dec 20, 2024
fe51c12
Merge branch 'mlcommons:master' into master
v-shobhit Sep 18, 2025
1f2666c
[Automated Commit] Format Codebase
github-actions[bot] Sep 18, 2025
ec84225
Add download_pdf.py
v-shobhit Sep 19, 2025
a77eb4d
Add artefacts downloading scripts
v-shobhit Sep 22, 2025
9f0db17
Fix downloading script
v-shobhit Sep 22, 2025
b4dcd57
[Automated Commit] Format Codebase
github-actions[bot] Sep 22, 2025
e0f28b2
cleanup read_pdf.py
v-shobhit Sep 22, 2025
d301bff
add env setup files
v-shobhit Sep 22, 2025
daa07f8
add single-shot retrieval
v-shobhit Sep 22, 2025
23aaea6
renmae setup file
v-shobhit Sep 22, 2025
a15bc6d
Add README
v-shobhit Sep 22, 2025
fb5f157
name fix
v-shobhit Sep 22, 2025
94b11d1
Add reranker
v-shobhit Sep 22, 2025
9275aaa
Add README.md
v-shobhit Sep 22, 2025
c5c7213
Save vector db
hans-intel Sep 26, 2025
5142a06
Support Intel XPU for embedding and reranker model
hans-intel Sep 26, 2025
9f5d8c5
Separate tok_k for retrieval and reranking
hans-intel Sep 26, 2025
41cd5c4
save url mapping in passage for evaluation
hans-intel Sep 26, 2025
7b9390f
Implement evaluation
hans-intel Sep 26, 2025
142d41d
Support bm25. RagDB is a superclass of vectordb and bm25db
hans-intel Sep 29, 2025
d1ed96d
Fix a bug in scoring; code cleanup
hans-intel Sep 29, 2025
73d1350
Add BM25 params
hans-intel Sep 29, 2025
9ab82f9
Change url_mapping to exclude file extension to support both pdf and txt
hans-intel Sep 29, 2025
beeb74f
Support txt file to ingest: --passages renamed to --ingest
hans-intel Sep 30, 2025
1f70b6c
Add stemmer to bm25
hans-intel Oct 1, 2025
94a46cb
Add missing ingest function for vectordb
hans-intel Oct 1, 2025
1140a8a
Add metrics recall, precision, F1, MAP (Mean Average Precision)
hans-intel Oct 1, 2025
4fd651c
implment retrieval strategies (top_p, relative, elbow, ...)
hans-intel Oct 2, 2025
140add3
rename evaluate() to evaluate_retrieval()
hans-intel Oct 2, 2025
4a8deb1
support html parsing (bs4)
hans-intel Oct 3, 2025
9710776
code cleanup
hans-intel Oct 3, 2025
1e84b9c
Fix vector db ingestion
hans-intel Oct 3, 2025
e72e2c7
fix r2r variation (torch seed)
hans-intel Oct 3, 2025
35f5bc2
improve parsing (remove wiki metadata, default sentence boundary)
hans-intel Oct 3, 2025
6c6d160
Add perf monitoring feature (--benchmark)
hans-intel Oct 6, 2025
fba871e
Add indexing trend measurement
hans-intel Oct 7, 2025
9396387
Add vector indexing option
hans-intel Oct 8, 2025
bb3737e
fix read_docs measurement
hans-intel Oct 10, 2025
f6b4396
Merge branch 'e2e-rag-eval-adv' into e2e-rag
hans-intel Oct 10, 2025
a385e67
Merge branch 'e2e-rag-r2r-fix' into e2e-rag
hans-intel Oct 10, 2025
9b28f7b
Merge branch 'e2e-rag-bsoup' into e2e-rag
hans-intel Oct 10, 2025
b4b0e02
Merge branch 'e2e-rag-latency-measurement' into e2e-rag
hans-intel Oct 10, 2025
7aee971
Merge branch 'e2e-rag-vector-indexing-option' into e2e-rag
hans-intel Oct 10, 2025
3c8a004
Implment IVF nprobe option
hans-intel Oct 14, 2025
5a9120c
Update README
hans-intel Oct 15, 2025
ac61fc0
Add feature save/load embeddings (--load-embeddings)
hans-intel Oct 15, 2025
b39e8ce
Embedding is done in multiple devices
hans-intel Oct 17, 2025
bfcb570
Centralized parameter management
hans-intel Oct 17, 2025
f5775df
Support top_p in vector db
hans-intel Oct 17, 2025
5ded358
Support parallelization in read_docs (--processes)
hans-intel Oct 17, 2025
1ae5f56
Change default max retrieval to 20 from 100
hans-intel Oct 17, 2025
f47c4f5
fix result display -- deduplicated urls
hans-intel Oct 17, 2025
a004b93
Fix correctly passing retrieval strategy params to filter
hans-intel Oct 17, 2025
ccef3a0
Merge branch 'e2e-rag-fix-vector-top_p' into e2e-rag
hans-intel Oct 18, 2025
464f0d8
Add new metrics (@N where N=#retrieved docs)
hans-intel Oct 18, 2025
9156fc6
Add detailed analysis after eval based on reasoning types and number …
hans-intel Oct 18, 2025
0fedf4f
More detailed analyses for FRAMES prompts
hans-intel Oct 26, 2025
c2e78df
basic multi-shot implementation
hans-intel Oct 27, 2025
2be6029
Implement --difficulty option (target difficult prompts only)
hans-intel Oct 27, 2025
731e101
implement doc grader
hans-intel Oct 27, 2025
0e9e44c
Combined docgrader into query_rewriter
hans-intel Oct 27, 2025
1a58fab
improve query writer prompt (WIP)
hans-intel Oct 29, 2025
a4171c8
Support hpu devices
hans-intel Oct 29, 2025
9d9254f
Query writer prompt WIP
hans-intel Oct 30, 2025
34f5501
Query writer prompt WIP
hans-intel Oct 30, 2025
6c7b862
Query rewriter prompt WIP
hans-intel Oct 30, 2025
9b26053
HPU reranking+embedding workaround (run on CPU)
hans-intel Nov 6, 2025
d966e8b
HPU reranking+embedding workaround (run on CPU)
hans-intel Nov 6, 2025
936d5ab
HPU embedding fix
hans-intel Nov 8, 2025
2a00054
fix read_docs to include .infobox
hans-intel Nov 9, 2025
248e23f
HPU reranking+embedding workaround (run on CPU)
hans-intel Nov 6, 2025
0eedaa6
HPU embedding fix
hans-intel Nov 8, 2025
ea50009
Support hpu devices
hans-intel Oct 29, 2025
3994990
generate LLM answer single_shot_retrieval
hans-intel Nov 6, 2025
eb57581
LLM answer for single shot retrieval and evaluation script
hans-intel Nov 7, 2025
b4da009
change default params to match FRAMES paper (k=5, n=5)
hans-intel Nov 9, 2025
3c35831
HPU reranking+embedding workaround (run on CPU)
hans-intel Nov 6, 2025
f2d4e55
copy url_mapping to output folder
hans-intel Nov 9, 2025
8bfd4ad
Support hpu devices
hans-intel Oct 29, 2025
ab9f902
Saving result of multi-shot retrieval for evaluate.py
hans-intel Nov 9, 2025
1139d0d
Simplified query rewriter prompt as baseline
hans-intel Nov 9, 2025
068d9b3
improve scoring prompt
hans-intel Nov 9, 2025
8b3b155
Fix bm25 for single shot retrieval
hans-intel Nov 9, 2025
8d22e1f
add scripts
hans-intel Nov 10, 2025
1bfebe5
multi hop accuracy improvement with oracle test into the same branch
mkankana Apr 29, 2026
9a038e4
hybrid model usage 120B+20B. removed dead code
mkankana May 6, 2026
b9a73fc
LLM calls to openrouter. logging all LLM calls and results
mkankana May 14, 2026
bbc0e66
added log sample result
mkankana May 17, 2026
6670aa9
added log sample result
mkankana May 17, 2026
3cad55f
Add parallel multi-shot retrieval with per-query threading
hans-intel May 18, 2026
3e8a0b4
Add OUTPUT_DIR and TEMPERATURE as parameters
hans-intel May 18, 2026
b6eacfa
Add retry with backoff for LLM calls and fix silent failures
hans-intel May 19, 2026
ce33ce8
Fix colbert to be used in late interaction way (query, doc embedding …
hans-intel May 21, 2026
46a9d95
Refactor device detection for cross-vendor support
hans-intel May 22, 2026
777c5bf
- OMP_NUM_THREADS derived from sched_getaffinity (respects upstream n…
hans-intel May 22, 2026
25b9b24
Remove hard-coded NUMA config, now in python script (membind cannot b…
hans-intel May 22, 2026
0b5fa9c
Add per-system config.sh + entry-point wrappers
hans-intel May 22, 2026
b701630
Add per-process GPU index allocator
hans-intel May 22, 2026
4c22295
Per-process reranker; per-worker NUMA pinning
hans-intel May 22, 2026
88df0be
add cross-system DB verification scripts
hans-intel May 22, 2026
ff80999
Decouple endpoint routing from OpenRouter
hans-intel May 23, 2026
9527fdd
Update readme
hans-intel May 23, 2026
99835eb
add a script that calculates prefix cache hit rate
hans-intel Jun 2, 2026
32041c5
Fixes to download docs (with delay and proper URL formatting), WARN a…
rpoornac Jun 2, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions e2e/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
*.db
output_*
passages/
data/
run_container.sh
.claude/
config.sh
313 changes: 313 additions & 0 deletions e2e/CLAUDE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,313 @@
# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

## Overview

This is a RAG (Retrieval-Augmented Generation) benchmark system for evaluating multi-hop question answering using Wikipedia documents from the FRAMES dataset. The system supports multiple retrieval methods (BM25, vector search), reranking, and multi-shot iterative retrieval with query decomposition.

## Architecture

The codebase follows a modular pipeline architecture:

1. **Document Ingestion** → **Passage Chunking** → **Vector/BM25 Indexing**
2. **Query** → **Retrieval** → **Optional Reranking** → **LLM Answer Generation** → **Evaluation**

### Core Components

- **`retrieve/` module**: Defines abstract `RagDB` base class with two implementations:
- `VectorDB`: Dense vector search using FAISS (flat, HNSW, or IVF indexing)
- `BM25DB`: Sparse lexical search using bm25s library
- Both support reranking with ColBERTv2 or similar models

- **Retrieval Scripts**:
- `single_shot_retrieval.py`: Single-step retrieval and evaluation
- `multi_shot_retrieval.py`: Multi-hop retrieval with LLM-based query decomposition (iterative retrieval with query rewriting)
- `oracle_single_shot.py`: Upper-bound evaluation using ground truth documents

- **Parameter Management**: `params.py` centralizes all CLI parameters with Optuna optimization metadata

- **Utilities**:
- `download_docs.py`: Downloads Wikipedia pages from FRAMES dataset URLs
- `read_docs.py`: Extracts text and chunks documents into passages (uses `text_splitter.py`)
- `evaluate.py`: LLM judge-based evaluation of generated answers
- `utils.py`: Common helpers (device config, LLM setup, seeding)

## Common Development Workflows

### Initial Setup

```bash
# Install dependencies (requires Ubuntu-based environment)
./setup.sh

# Download Wikipedia documents from FRAMES dataset
python3 download_docs.py --output_dir doc_html --format html --processes 30

# Build vector database (run once)
bash scripts/run_ingestion.sh
```

The setup script will:
1. Extract passages from HTML documents → `passages/doc_html_len2048_ov32_word.json`
2. Build FAISS HNSW index → `vector_html_hnsw_len2048_ov32_word.db`

### Running Experiments

**Single-shot retrieval:**
```bash
# Run evaluation on existing database
python3 single_shot_retrieval.py \
--db vector_html_hnsw_len2048_ov32_word.db \
--retrieval_method vector \
--eval 100 \
--generate-answer

# Compare BM25 vs Vector
python3 single_shot_retrieval.py --db bm25.db --retrieval_method bm25 --eval
```

**Multi-shot retrieval with query decomposition:**
```bash
# Run multi-shot experiment (requires LLM server on port 8123)
bash scripts/run_multi_shot.sh 50 # Evaluate 50 queries
bash scripts/run_multi_shot.sh all # Full evaluation
```

**Oracle upper bound (using ground truth docs):**
```bash
python3 oracle_single_shot.py \
--dataset data/frames_dataset.tsv \
--wiki-articles-dir wiki_articles \
--batch-size 16
```

### LLM Server Setup

The system expects a vLLM-compatible OpenAI API server:

```bash
# Start vLLM server (example from scripts/start_vllm_server.sh)
python3 -m vllm.entrypoints.openai.api_server \
--model /model/gpt-oss-20b-mxfp4 \
--dtype bfloat16 \
--host 0.0.0.0 \
--port 8123 \
--gpu-memory-util=0.95 \
--enable-prefix-caching \
--max-model-len=131072
```

Default service URL: `http://127.0.0.1:8123/v1/chat/completions`

### Evaluation

Evaluation uses an LLM judge to score answers:

```bash
# Score results from any experiment
python3 evaluate.py result_single_shot.json
python3 evaluate.py result_multi_shot.json
python3 evaluate.py oracle_checkpoint.pkl # For oracle results
```

## Key Parameters (via params.py)

All parameters are centralized in `params.py` with CLI definitions. Key categories:

**Retrieval Method:**
- `--retrieval_method {bm25,vector}`: Choose retrieval backend
- `--vector_index_method {flat,hnsw,ivf}`: FAISS index type (default: hnsw)
- `--bm25_k1`, `--bm25_b`, `--bm25_stemmer`: BM25 tuning parameters

**Retrieval Strategy:**
- `--retrieval_strategy {fixed_k,top_p,relative}`: How many docs to retrieve
- `--top_k_retriever N`: Number of docs to retrieve (default: 10)
- `--top_k_reranking N`: Number of docs after reranking (default: 10)

**Device & Performance:**
- `--device {auto,xpu,cuda,hpu,cpu}`: Hardware accelerator
- `--num_embedding_devices N`: Parallel embedding generation across devices
- `--benchmark`: Enable performance monitoring

**Multi-shot specific (multi_shot_retrieval.py):**
- `--max-iterations N`: Max retrieval rounds (default: 5)
- `--max-sub-queries N`: Sub-queries per iteration (default: 3)

**Oracle specific (oracle_single_shot.py):**
- `--batch-size N`: Batch requests to LLM (default: 1)
- `--timeout N`: Request timeout in seconds (default: 2400)
- `--enable-thinking`: Use chain-of-thought reasoning

## Hardware Support

The system supports multiple accelerators via `--device`:
- **XPU** (Intel GPUs): Primary development target
- **CUDA** (NVIDIA GPUs)
- **HPU** (Habana Gaudi): Embeddings/reranking fall back to CPU due to compatibility
- **CPU**: Fallback option

Device selection is abstracted in `utils.py:get_device_config()` and `ragdb.py:_determine_device()`.

## Database Persistence

- Vector databases: `.db` file (FAISS index) + `_data/` directory (docstore, metadata)
- BM25 databases: Pickled retriever object in `.db` file
- Embeddings cache: `.emb.pkl` files (use `--load-embeddings` to reuse)
- Checkpoints: `oracle_checkpoint.pkl` for oracle runs (resumable via pandas DataFrame)

## Multi-shot Retrieval Logic

The multi-shot system (multi_shot_retrieval.py) implements iterative retrieval:
1. LLM evaluates retrieved docs and decides if sufficient to answer
2. If insufficient, generates up to k focused sub-queries
3. Each sub-query retrieves additional docs
4. Process repeats for max N iterations
5. Final docs are reranked and passed to answer generator

The query rewriter prompt includes failure analysis to escalate search strategies when stuck (e.g., switching from specific queries to broader entity searches).

## Evaluation Methodology

- **Retrieval Accuracy**: Checks if correct Wikipedia URLs are in top-K results
- **Answer Quality**: LLM judge scores generated answers against ground truth
- **Difficulty Filtering**: `--difficulty N` filters queries by number of required source documents

Results are saved to JSON files with schema:
```json
{
"query": "...",
"retrieved_urls": [...],
"correct_urls": [...],
"llm_answer": "...",
"ground_truth_answer": "..."
}
```

## Testing & Debugging

- Use `--eval N` to test on first N queries (faster iteration)
- Use `--no-rerank` to compare retrieval methods fairly
- Use `--no-save` to skip writing database during optimization
- Use `--benchmark` to track component performance
- Check logs: Scripts redirect output to `log_*.txt` files

## Important Notes

- **LLM timeout**: Oracle and multi-shot runs may need `--timeout` adjustment for reasoning models
- **Checkpointing**: Oracle script saves progress after each batch and can resume from checkpoint
- **Determinism**: Use `--seed` for reproducible results (affects sampling, not LLM generation)

---

## Experiments and Results

### Accuracy Benchmarks

| Type | Queries | Precision@N | Recall@N | F1@N | LLM Judge Accuracy |
|---|---|---|---|---|---|
| Oracle | 824 | 100% | 100% | 100% | 68% |
| Single-shot retrieval | 827 | 39% | 70% | 42% | 20% |
| Multi-shot baseline | 50 | 12% | 36% | 16% | 20% |
| Multi-shot + fixes below | 50 | 69% | 64% | 61% | 42% |
| Multi-shot + fixes below | 400 | 73% | 67% | 66% | 34% |
| **Multi-shot + fixes below** | **824** | **72%** | **67%** | **66%** | **36%** |

The retrieval recall is the primary bottleneck: theoretical max accuracy ≈ Recall × Oracle_accuracy.

---

### Fix 1: Split Monolithic Query Rewriter (HIGH IMPACT)

**Root Cause:** The original `query_rewriter()` was a single LLM call with 3 simultaneous tasks:
1. Evaluate relevance of new documents
2. Decide if accumulated docs are sufficient to answer
3. Generate new search queries

This cognitive overload caused the LLM to mark **all documents as irrelevant** (relevance: [0,0,0,...]), leading to 0 kept docs and 60% "Unknown" answers.

**Fix:** Split into two focused LLM calls:
- `evaluate_document_relevance()` — binary relevance classification only (temp=0.0, short prompt)
- `generate_search_queries()` — query generation or final answer (temp=0.1, full context)

**Additional fixes bundled with Fix 1:**
- Added best-effort final answer after max iterations (instead of always returning "Unknown")
- Added "return Unknown if insufficient" guard in prompt to reduce hallucination
- Added fallback to original query when no sub-queries are generated
- Guarded `sufficient=True` to require non-empty `kept_docs`
- Added `reasoning_content` fallback + `max_tokens=10240` for thinking-model compatibility
- Fixed `UnboundLocalError` from `import re` inside `try` blocks (4 locations)

**Result:** Accuracy 20% → 30%, Recall 36.8% → 51.8% (n=50)

---

### Fix 2: Context Length Reduction (REVERTED — made things worse)

**Motivation:** Fix 1 still produced empty LLM responses due to long prompts when many docs accumulated (11+ docs × 1200 chars ≈ 20KB+ prompts hitting token limits).

**Change:** Limit kept-doc context shown to LLM — query generation: 10 most recent docs; relevance check: 5 most recent docs.

**Result:** Fewer empty responses, but accuracy dropped 30% → 28%, Recall 51.8% → 46.4%.

**Why it failed:** Multi-hop reasoning needs to connect facts across all retrieved documents, not just the most recent ones. Truncating context broke cross-document reasoning chains.

**Decision:** Reverted Fix 2. The right fix is instead: increase `max_tokens` for the judge/LLM calls so they don't hit length limits.

---

### Chunking Strategy Experiments

**Baseline:** 2048-char chunks, 32-char overlap (1.5% overlap) — too large, dilutes embeddings.

**Hypothesis:** Smaller chunks produce more focused embeddings → better retrieval precision for multi-hop facts.

#### Results Across Chunk Sizes (multi-shot, n=50)

| Chunk Size | Overlap | Recall@N | Precision@N | LLM Accuracy |
|---|---|---|---|---|
| 2048 chars | 32 (1.5%) | 51.8% | 62.5% | 30% |
| 512 chars | 100 (20%) | — | — | Tested, no improvement |
| **768 chars** | **32 (4%)** | **67.7%** | **73.0%** | **34–37%** |

**Winner: 768-char chunks with 32-char overlap** — significant retrieval improvement over 2048.

**Why 768 works better than 512:**
- 512-char chunks split related facts across too many chunks; retrieval becomes noisy
- 768-char chunks fit 2–3 complete sentences; good balance of focus vs. context
- More manageable passage count than 512

**Overlap finding:** The 1.5% overlap (32/2048) in the baseline was too low. With 768-char chunks, even 32-char overlap (4%) provides measurably better boundary coverage than 32/2048 did.

**Strategies considered but not implemented:**
- **Hierarchical chunking** (512 retrieval / 2048 context): Promising but complex; worth trying if further gains needed
- **Semantic sentence grouping**: More implementation complexity for marginal benefit over fixed-length with word boundary
- **Token-based chunking**: Aligns better with embedding model limits (e5-base-v2 = 512 tokens); worth trying

**Key insight for future experiments:** The chunk size primarily affects retrieval recall. Every ~10% recall improvement translates to ~7% accuracy gain (based on Recall × Oracle_accuracy formula).

---

### Iteration Distribution (full 824-query run, max_iterations=5)

| Iterations Used | Questions | % |
|---|---|---|
| 2 | 297 | 36.0% |
| 3 | 113 | 13.7% |
| 4 | 37 | 4.5% |
| 5 | 377 | 45.8% |

45.8% of queries hit the max-iterations limit — suggesting accuracy gains are available by increasing `--max-iterations` to 7 or 10.

---

### Future Experiment Candidates

In priority order based on findings above:

1. **Increase `--max-iterations`** (7 or 10) — 45.8% of queries are cut off at 5 iterations
2. **Hierarchical chunking** (retrieve 512-char children, answer with 2048-char parents)
3. **Token-based chunking** at 256 tokens to align with e5-base-v2 limits
4. **Hybrid BM25 + vector retrieval** — lexical search catches exact name matches that dense search misses
5. **Larger top_k_retriever** (15 → 25) given high iteration cutoff rate
6. **Better judge model** — current judge (same gpt-oss-20b) hits token limits on complex questions; a stronger judge would give more accurate accuracy measurements
Loading