Skip to content

[Feat][Router]: Add automatic retry with exponential backoff and jitter - #939

Open
ikaadil wants to merge 4 commits into
vllm-project:mainfrom
ikaadil:request-retry
Open

ikaadil wants to merge 4 commits into
vllm-project:mainfrom
ikaadil:request-retry

Conversation

@ikaadil

@ikaadil ikaadil commented May 1, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Adds opt-in automatic retry for transient backend failures using exponential backoff with jitter, and folds the existing instance-failover reroute into the same mechanism.

Motivation: Mitigating the thundering herd problem

A forwarded request fails in one of two ways, which call for opposite responses:

Failure Example Response
Transport failure connection refused, timeout Engine excluded, request rerouted to another engine immediately
Retryable status 408 429 500 502 503 504 Engine stays eligible, request re-issued after a backoff

Both draw on one budget: --max-retries. Retrying is disabled by default, so a request is attempted exactly once unless --enable-retries is passed.

CLI

Flag Default Description
--enable-retries off Enable retrying
--max-retries 5 Total attempts including the initial one; must be >= 2
--initial-backoff-ms 50 Backoff before the first retry
--max-backoff-ms 30000 Upper bound on the backoff
--backoff-multiplier 1.5 Growth factor between retries
--jitter-factor 0.2 Randomisation applied to each delay (0.0-1.0)

Without --enable-retries the other flags are ignored, so a shared config template carrying them cannot change behaviour or block startup.

vllm-router --port 8000 \
    --service-discovery static \
    --static-backends "http://localhost:9001,http://localhost:9002" \
    --static-models "facebook/opt-125m,facebook/opt-125m" \
    --routing-logic roundrobin \
    --enable-retries \
    --max-retries 5 \
    --initial-backoff-ms 100 \
    --max-backoff-ms 60000 \
    --backoff-multiplier 2.0 \
    --jitter-factor 0.1

Backoff

delay  = min(initial_backoff_ms * multiplier ^ retry, max_backoff_ms)
delay' = delay * (1 + U[-jitter_factor, +jitter_factor])

With the defaults, retries land at 50ms, 75ms, 112.5ms, 168.8ms, each spread across a +/-20% window. Without the jitter, routers that backed off from the same incident retry in lockstep and re-create the overload they backed off from.

Changes

File Change
services/request_service/retry.py New. RetryConfig (frozen, self-validating), RetryState (attempt sequencing), is_retryable_status
services/request_service/request.py route_general_request drives RetryState.attempts(); shared router dispatch extracted into _select_backend()
parsers/parser.py Retry flags; validation delegated to RetryConfig
app.py app.state.retry_config = RetryConfig.from_args(args)
routers/routing_logic.py Retry policy moved out; failover wiring removed

Validation lives in RetryConfig.__post_init__, so an invalid config cannot be constructed.

Behaviour

  • A retryable status does not blacklist the engine, so it stays a candidate after the backoff. This is what makes retrying useful on a single-engine deployment.
  • Failover between healthy engines is never delayed; the backoff applies only once every engine has been excluded and the pool is given another chance.
  • Once the budget is spent the backend response passes through, so the client sees the engine's own status and body rather than a synthesised router error.
  • Non-retryable statuses return on the first attempt.
  • Retries apply only before streaming begins; once the first chunk is forwarded the headers are already with the client, so a mid-stream error passes through untouched.

Breaking change

--max-instance-failover-reroute-attempts is removed. It rerouted a failed request to another engine and is now subsumed by --max-retries, which covers the same case and adds backoff.

Its default was 0 (a single attempt), which is also the default here, so deployments that never set it are unaffected. Deployments that did set it should migrate:

- --max-instance-failover-reroute-attempts 2
+ --enable-retries --max-retries 3

Note for reviewers

process_request reports a backend error status by yielding it (yield backend_response.headers, backend_response.status), not by raising, and there is no raise_for_status on that call. A retry implementation that only catches HTTPException never fires for an engine 5xx. This PR inspects the status after the first anext, and aclose()s a discarded response so the upstream connection is released.

Tests

56 new tests, 281 passing suite-wide.

  • test_retry_config.py — status classification, backoff growth, cap, jitter bounds, validation, from_args mapping.
  • test_request_retry.py — end to end through route_general_request: default single attempt, transport reroute, no-backoff-on-failover, single-engine retry, budget exhaustion passthrough, non-retryable passthrough, generator close on discard, exact sleep sequence.

Every backend-status stand-in yields the status, matching process_request; one that raises HTTPException would pass whether or not retrying works.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a retry mechanism with exponential backoff and jitter for transient failures, aligning the router's behavior with the sglang model gateway. The changes include a new RetryConfig dataclass, CLI arguments for configuration, updated documentation, and logic in the request service to handle retryable HTTP status codes (408, 429, 500, 502, 503, 504). Feedback identifies a logic error where retries are effectively disabled by default due to the max_attempts calculation, the inclusion of an unused last_response variable, and a concern that blacklisting URLs for transient errors prevents retrying the same backend in single-node environments.

Comment thread src/vllm_router/services/request_service/request.py Outdated
Comment thread src/vllm_router/services/request_service/request.py Outdated
Comment thread src/vllm_router/services/request_service/request.py Outdated
@ikaadil
ikaadil force-pushed the request-retry branch 3 times, most recently from a5bac11 to 1c7980e Compare May 2, 2026 09:54
@ikaadil

ikaadil commented May 4, 2026

Copy link
Copy Markdown
Contributor Author

@ruizhang0101 could you please review the MR? Thanks!

@ruizhang0101

Copy link
Copy Markdown
Collaborator

@aeon-x Could you take a look at this?

@aeon-x

aeon-x commented May 6, 2026

Copy link
Copy Markdown
Collaborator

Hey @ikaadil, i think retrying should be an optional, as most users would need fast fail over.

Can you make sure that this retry mechanism is turned off unless it is explictly turned on by a flag?

@ikaadil

ikaadil commented May 6, 2026

Copy link
Copy Markdown
Contributor Author

Can you make sure that this retry mechanism is turned off unless it is explictly turned on by a flag?

Done

@waelrabah11 waelrabah11 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall looks good to me. You can consider my comments as nitpicks but I just disagree with using the retry prefix. It's a bit redundant and could be avoided with documentation

Comment thread src/vllm_router/README.md Outdated
Comment thread src/vllm_router/README.md Outdated
Comment thread src/vllm_router/app.py Outdated
Comment thread src/vllm_router/parsers/parser.py Outdated
Comment thread src/vllm_router/parsers/parser.py Outdated
Comment thread src/vllm_router/parsers/parser.py Outdated
Comment thread src/vllm_router/parsers/parser.py Outdated
Comment thread src/vllm_router/parsers/parser.py Outdated
…transient failures

Signed-off-by: Ifta khairul Alam Adil <ikaadil007@gmail.com>
Signed-off-by: Ifta khairul Alam Adil <ikaadil007@gmail.com>
Signed-off-by: Ifta khairul Alam Adil <ikaadil007@gmail.com>
Signed-off-by: Ifta khairul Alam Adil <ikaadil007@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants