Skip to content

Multi vendor support: Fixes for download_docs & WARN for lack of OPENROUTER_API_KEY - #5

Open
rpoornac wants to merge 105 commits into
v-shobhit:masterfrom
hans-intel:multi-vendor-support
Open

rpoornac wants to merge 105 commits into
v-shobhit:masterfrom
hans-intel:multi-vendor-support

Conversation

@rpoornac

@rpoornac rpoornac commented Jun 2, 2026

Copy link
Copy Markdown

Fixes in e2e/download_docs.py to download appropriate URLs
ERROR -> WARN for lack of OPENROUTER_API_KEY

mlcommons-bot and others added 30 commits December 20, 2024 22:46
- Each query is scored 0 to 1 depending on the number of correct links
- Final score is averaged
Support no-save feature
Add more bm25 params
Refactor single shot script to consolidate vector and bm25 db
Change default bm25 backend to numba (faster)
add no-rerank option to compare vector and bm25 method
hans-intel and others added 30 commits November 9, 2025 04:50
Refactored max_token and context_char_limit (now use context_token_limit to calculate it)
* --num-workers N
* sharing embedding/reranking model across threads
* LLM requests expect continuous batching from the serving side
- Retry on 429 (rate limit), 502/503/504, and timeouts with exponential backoff
- Raise on errors instead of returning empty string silently
- Warn when LLM returns empty content with raw response for debugging
- Fix NoneType error in logger when API returns content: null
- Centralize detection in utils.detect_device(); RagDB._determine_device
  now delegates instead of duplicating the logic.
- Reorder priority to CUDA/ROCm -> XPU -> HPU -> CPU.
- Distinguish AMD ROCm from NVIDIA CUDA via torch.version.hip (logging
  only; both return "cuda" since PyTorch ROCm uses the cuda namespace).
- Add "rocm" as a user-facing alias for "cuda" in --device choices.
…umactl); honors E2E_OMP_NUM_THREADS override.

- New env vars: E2E_OMP_NUM_THREADS
…e done this way)

Added parameters: embedding_device, reranker_device
New env vars: DEVICE, EMBEDDING_DEVICE, RERANKER_DEVICE
(use config.template.sh to system-specific config.sh)
- INGESTION_* (run_ingestion.sh — chunk size, embedding device, doc/passage/db paths)
- INFERENCE_* (run_multi_shot/single_shot)
- INFERENCE_ORACLE_* (run_oracle)
- CPU_* (Python NUMA/OMP, shared).

Script renames for naming consistency (all entry points are run_*.sh):
- scripts/setup_db.sh        -> scripts/run_ingestion.sh
- scripts/run_oracle_eval.sh deleted (replaced by scripts/run_oracle.sh)
- scripts/run_single_shot.sh, scripts/run_oracle.sh added.
- Looking for empty GPUs to load embedding and reranker
- INFERENCE_EMBEDDING_GPU_DEVICES, INFERENCE_RERANKER_GPU_DEVICES to override
- Fixed a bug in embedding index and GPU indices were the same (GPU indices could start from non-zero)
- INFERENCE_RERANKER_NUMA_NODE       pin reranker child to NUMA node N
- INFERENCE_RERANKER_OMP_NUM_THREADS  override reranker OMP threads
- INFERENCE_EMBEDDING_NUMA_NODES     CSV (one per --num_embedding_devices)
- INFERENCE_EMBEDDING_OMP_NUM_THREADS  cap per worker (default = even split)
Removed --llm_service_url / --llm_model and added url endpoint and model for each component
…bout OPENROUTER_API_KEY

Signed-off-by: Rajesh Poornachandran <rajesh.poornachandran@amd.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants