Problem
The ML reranker is always in-process ONNX (getRerankProvider() returns TransformersRerankProvider; search/rerank.ts only branches on config.reranker === 'onnx' with no other provider path). There is no way to point reranking at a remote endpoint, unlike the LLM layer which has a full custom/openai-compatible backend (custom-backend.ts, run.ts).
For local-first / low-RAM deployments this is a real gap: hosting a second CPU-heavy cross-encoder (bge-reranker-v2-m3) in-process adds RAM+latency on already-loaded boxes, and some users already run a dedicated embedding/rerank service remotely (e.g. an external model server) — embedding has an openai_compatible path (WIGOLO_EMBEDDING_PROVIDER/WIGOLO_EMBEDDING_API_BASE), but reranking has no equivalent.
Request
- Add a configurable rerank backend, parallel to the embedding layer's
openai_compatible option: e.g. WIGOLO_RERANK_PROVIDER=onnx|remote|none (keep onnx default) plus WIGOLO_RERANK_API_BASE / WIGOLO_RERANK_API_KEY / WIGOLO_RERANK_MODEL for the remote path.
- Ideally accept an OpenAI-compatible rerank endpoint (several local servers expose a rerank/embedding adapter), so it composes with the existing
openai_compatible embedding setup on the same host.
- Keep the
RerankProvider interface (rerank(query, candidates, topK) → {id, score}[]) and register a remote adapter behind getRerankProvider() so the rest of the pipeline is unchanged.
This mirrors the gap that was already solved on the LLM side (#276 / openai-compatible provider). Happy to help test against a local rerank server if useful.
Env sketch
WIGOLO_RERANK_PROVIDER=remote
WIGOLO_RERANK_API_BASE=http://127.0.0.1:8082/v1
WIGOLO_RERANK_API_KEY=... # optional, if the endpoint requires auth
WIGOLO_RERANK_MODEL=bge-reranker-v2-m3
Problem
The ML reranker is always in-process ONNX (
getRerankProvider()returnsTransformersRerankProvider;search/rerank.tsonly branches onconfig.reranker === 'onnx'with no other provider path). There is no way to point reranking at a remote endpoint, unlike the LLM layer which has a full custom/openai-compatiblebackend (custom-backend.ts, run.ts).For local-first / low-RAM deployments this is a real gap: hosting a second CPU-heavy cross-encoder (bge-reranker-v2-m3) in-process adds RAM+latency on already-loaded boxes, and some users already run a dedicated embedding/rerank service remotely (e.g. an external model server) — embedding has an
openai_compatiblepath (WIGOLO_EMBEDDING_PROVIDER/WIGOLO_EMBEDDING_API_BASE), but reranking has no equivalent.Request
openai_compatibleoption: e.g.WIGOLO_RERANK_PROVIDER=onnx|remote|none(keeponnxdefault) plusWIGOLO_RERANK_API_BASE/WIGOLO_RERANK_API_KEY/WIGOLO_RERANK_MODELfor the remote path.openai_compatibleembedding setup on the same host.RerankProviderinterface (rerank(query, candidates, topK) → {id, score}[]) and register a remote adapter behindgetRerankProvider()so the rest of the pipeline is unchanged.This mirrors the gap that was already solved on the LLM side (#276 /
openai-compatibleprovider). Happy to help test against a local rerank server if useful.Env sketch