Repository navigation
Conversation
- Each query is scored 0 to 1 depending on the number of correct links - Final score is averaged
Support no-save feature Add more bm25 params Refactor single shot script to consolidate vector and bm25 db
Change default bm25 backend to numba (faster)
add no-rerank option to compare vector and bm25 method
Simplified evaluation logic
Refactored max_token and context_char_limit (now use context_token_limit to calculate it)
* --num-workers N * sharing embedding/reranking model across threads * LLM requests expect continuous batching from the serving side
- Retry on 429 (rate limit), 502/503/504, and timeouts with exponential backoff - Raise on errors instead of returning empty string silently - Warn when LLM returns empty content with raw response for debugging - Fix NoneType error in logger when API returns content: null
- Centralize detection in utils.detect_device(); RagDB._determine_device now delegates instead of duplicating the logic. - Reorder priority to CUDA/ROCm -> XPU -> HPU -> CPU. - Distinguish AMD ROCm from NVIDIA CUDA via torch.version.hip (logging only; both return "cuda" since PyTorch ROCm uses the cuda namespace). - Add "rocm" as a user-facing alias for "cuda" in --device choices.
…umactl); honors E2E_OMP_NUM_THREADS override. - New env vars: E2E_OMP_NUM_THREADS
…e done this way) Added parameters: embedding_device, reranker_device New env vars: DEVICE, EMBEDDING_DEVICE, RERANKER_DEVICE
(use config.template.sh to system-specific config.sh) - INGESTION_* (run_ingestion.sh — chunk size, embedding device, doc/passage/db paths) - INFERENCE_* (run_multi_shot/single_shot) - INFERENCE_ORACLE_* (run_oracle) - CPU_* (Python NUMA/OMP, shared). Script renames for naming consistency (all entry points are run_*.sh): - scripts/setup_db.sh -> scripts/run_ingestion.sh - scripts/run_oracle_eval.sh deleted (replaced by scripts/run_oracle.sh) - scripts/run_single_shot.sh, scripts/run_oracle.sh added.
- Looking for empty GPUs to load embedding and reranker - INFERENCE_EMBEDDING_GPU_DEVICES, INFERENCE_RERANKER_GPU_DEVICES to override - Fixed a bug in embedding index and GPU indices were the same (GPU indices could start from non-zero)
- INFERENCE_RERANKER_NUMA_NODE pin reranker child to NUMA node N - INFERENCE_RERANKER_OMP_NUM_THREADS override reranker OMP threads - INFERENCE_EMBEDDING_NUMA_NODES CSV (one per --num_embedding_devices) - INFERENCE_EMBEDDING_OMP_NUM_THREADS cap per worker (default = even split)
Removed --llm_service_url / --llm_model and added url endpoint and model for each component
…bout OPENROUTER_API_KEY Signed-off-by: Rajesh Poornachandran <rajesh.poornachandran@amd.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes in e2e/download_docs.py to download appropriate URLs
ERROR -> WARN for lack of OPENROUTER_API_KEY