Skip to content

[Bugfix][Router] loadaware: score bursts against live load, not the arrival snapshot - #1108

Open
ibrahimnd2000 wants to merge 1 commit into
vllm-project:mainfrom
ibrahimnd2000:loadaware-live-load-for-bursts
Open

ibrahimnd2000 wants to merge 1 commit into
vllm-project:mainfrom
ibrahimnd2000:loadaware-live-load-for-bursts

Conversation

@ibrahimnd2000

@ibrahimnd2000 ibrahimnd2000 commented Sep 28, 2026 •

Copy link
Copy Markdown

Problem

Under a burst, loadaware places almost every request on one endpoint.

route_general_request snapshots request_stats before it awaits route_request. LoadAwareRouter.route_request then awaits the controller lookup (and the instance-map refresh). A request only counts as in flight once process_request calls on_new_request. So when many requests arrive together, every one of them is scored against the same pre-burst load. The load term is equal for all endpoints, and whichever endpoint holds even a short shared prefix (for example a common system prompt) wins every placement.

Impact: a GLM-5.x deployment with 3 vLLM replicas ran a benchmark cell of 128 concurrent ~32k-token prompts.

  • The opening burst was placed 95 / 30 / 22.
  • The overloaded replica needed ~130 s to prefill its queue, so a large share of the burst waited past the client's first-token timeout.

Fix

  • After the last await in route_request, re-read request_stats from request.app.state.request_stats_monitor (new helper LoadAwareRouter.live_request_stats). This applies to the select path and to the no-cache fallback path.
  • Why this is enough: nothing awaits between the placement decision and on_new_request. After route_request returns, route_general_request runs only synchronous code until await anext(stream_generator), which runs process_request synchronously up to on_new_request. So each decision sees every request placed before it.
  • Without a reachable monitor (e.g. unit tests without an app), the caller's snapshot is used, as before.

Tests

New tests in src/tests/test_loadaware_router.py route 128 concurrent requests through route_request with a real RequestStatsMonitor. The controller round-trip is mocked with 1–20 ms of latency, and a 512-of-32,000-token prefix is cached on one of three endpoints. As in the real request path, each placement is followed by on_new_request.

Placement over 3 endpoints
Stale snapshot (old behaviour) 128 / 0 / 0
With this fix 43 / 43 / 42

Also added: unit tests for live_request_stats, covering the live read and the fallback to the snapshot without a monitor.

pytest src/tests/test_loadaware_router.py src/tests/test_kvaware_router.py src/tests/test_request_stats.py
47 passed

Four of the new tests fail on main. pre-commit (black, isort, ruff, codespell) passes.

Validation

In the deployment above, with the fix, a live 30-request burst with a shared cached system prompt was placed 11 / 10 / 9. The router's per-endpoint in-flight counts matched the engines' running requests.

Notes

  • This complements [Bugfix][Router] Fix leaked in-flight counters in RequestStatsMonitor #1072, which fixed in-flight counters leaking on errors and client disconnects. That fix keeps the counts correct over time; this one keeps them current within a burst.
  • "loadaware: KV lookup for chat completions" (separate PR) adds a lookup-failure fallback in route_request. Once both are merged, that fallback should also use live_request_stats. Both PRs add a test section at the same place in test_loadaware_router.py, so whichever lands second needs a trivial rebase.

…rrival snapshot

route_general_request snapshots request_stats before awaiting
route_request, and LoadAwareRouter.route_request awaits the controller
lookup (and the instance-map refresh). A request only counts as in flight
once process_request calls on_new_request. So every request of a burst was
scored against the same pre-burst load, and a small shared-prefix match
(e.g. a common system prompt cached on one endpoint) won every tie: a
128-request burst over three endpoints was placed 128/0/0.

Re-read request_stats from the RequestStatsMonitor after the awaits, on the
select path and the no-cache fallback path. Nothing awaits between the
placement decision and on_new_request, so each decision now sees every
earlier one; the same burst is placed 43/43/42. Without a reachable monitor
the caller's snapshot is used as before.

Signed-off-by: Muhammad Ibrahim <ibrahimnd2000@gmail.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a live_request_stats method to fetch up-to-date request statistics from the monitor during routing, rather than relying on stale snapshots taken at the request's arrival. This prevents concurrent request bursts from being routed to the same endpoint due to outdated load information. Corresponding unit tests have been added to verify burst routing behavior and live stats retrieval. There are no review comments, so no further feedback is provided.

@ibrahimnd2000 ibrahimnd2000 changed the title [Bugfix][Router] loadaware: score bursts against live load, not the a… [Bugfix][Router] loadaware: score bursts against live load, not the arrival snapshot Sep 28, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant