Skip to content

[Bugfix][Router] Reconcile the engine list when the K8s watch reconnects - #1092

Open
jStrider wants to merge 3 commits into
vllm-project:mainfrom
jStrider:fix/service-discovery-resync-on-reconnect
Open

jStrider wants to merge 3 commits into
vllm-project:mainfrom
jStrider:fix/service-discovery-resync-on-reconnect

Conversation

@jStrider

Copy link
Copy Markdown

FIX #1091

K8sPodIPServiceDiscovery._watch_engines() treats the watch stream as its only source of engine removals. When the stream drops, the loop reconnects but never re-lists, so a DELETED event delivered while the connection is down is lost permanently: the engine stays in available_engines and keeps receiving traffic.

Requests then round-robin onto a pod IP that no longer exists and hang until ClientConnectorError, while /v1/models and /health stay green — both are served from the router's own state, so health checks never notice. vllm:num_requests_running for that server climbs and never drains, which also skews least-loaded routing.

This is distinct from #1008 / #1012, which fixed the eviction case (MODIFIED with podIP=None). Here the pod is deleted normally; the event is simply never delivered.

The change

_reconcile_engines() lists the pods and drops the engines whose pod is gone. It is called at the top of the while self.running loop, so it runs before each stream is opened — that is, exactly when events may have been missed, and only once at startup in steady state. No new thread, no new flag, no periodic cost.

Removal happens under available_engines_lock rather than through _delete_engine(), which takes the same lock and would del without a guard — avoiding both a deadlock and a race with the watch thread.

Testing

Two unit tests in the existing test_k8s_pod_ip_service_discovery.py: an engine whose pod is gone is dropped, and live engines are never dropped. Both fail if the detection is neutralised (verified by mutating stale to an empty set).

Validated end to end on a scale-to-zero GPU cluster (KEDA, workers 0↔1), same protocol both times — block the router's egress to the API server, wait for the watch to drop on its own, delete the worker while blind, restore the network:

without the patch with the patch
ghost present on restore yes yes
after 30 s still served no longer exists but was still registered: dropping it
after 3 min /v1/models advertises the model, real request times out at 45 s engine list empty

Note for anyone reproducing this: the NetworkPolicy must exclude the real control-plane IP, not the kubernetes.default ClusterIP (kube-proxy DNATs before the policy applies), and an already-established watch survives the policy — the router is only truly blind once its next reconnect fails.

🤖 Generated with Claude Code

The watch stream is the only source of engine removals in
K8sPodIPServiceDiscovery, so a DELETED event that lands while the
connection is down is lost for good: the engine stays in
available_engines and keeps taking traffic. Requests hang until they
fail with ClientConnectorError, while /v1/models and /health stay green
because both are served from the router's own state.

This is distinct from vllm-project#1012, which handles the eviction case (MODIFIED
with podIP=None): here the pod is deleted normally and the event is
simply never delivered.

List the pods before opening each stream and drop the engines that no
longer exist. The loop already restarts on every disconnect, so the list
runs exactly when events may have been missed, and only once at startup
in steady state.

Reproduced on a scale-to-zero GPU cluster by blocking the router's
egress to the API server, deleting the worker while blind, then
restoring: without this change the router served the dead pod
indefinitely; with it the engine is dropped on reconnect.

Fixes vllm-project#1091

Signed-off-by: Julien Renaud <julien.renaud@sancare.fr>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a reconciliation mechanism (_reconcile_engines) in K8sPodIPServiceDiscovery to clean up stale engines whose Kubernetes pods were deleted while the watch stream was disconnected, along with corresponding unit tests. The review feedback highlights two important points: first, a similar watch-reconnect vulnerability exists in K8sServiceNameServiceDiscovery which should also implement a reconciliation mechanism; second, a potential race condition exists between the reconciliation list call and the start of the watch stream, which can be resolved by using the resource_version from the list call to initialize the watch stream.

Comment thread src/vllm_router/service_discovery.py Outdated
def _watch_engines(self):
while self.running:
try:
self._reconcile_engines()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The K8sServiceNameServiceDiscovery class implements a very similar watch loop in its _watch_engines method but does not have a reconciliation mechanism. It is vulnerable to the exact same watch-reconnect bug where a service deleted while the watch is down becomes a permanent ghost service.\n\nPlease implement a similar reconciliation method (e.g., _reconcile_services) for K8sServiceNameServiceDiscovery to ensure consistency and prevent the same bug there.

Comment thread src/vllm_router/service_discovery.py Outdated
return None
return pod.metadata.labels.get("model")

def _reconcile_engines(self):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

There is a subtle race condition between _reconcile_engines() and self.k8s_watcher.stream(). Because they are two separate API calls (list_namespaced_pod in _reconcile_engines and another list_namespaced_pod internally inside self.k8s_watcher.stream), if a pod is deleted in the short window between these two calls:\n\n1. _reconcile_engines will see the pod as alive (since it was alive during the first list call) and won't remove it.\n2. self.k8s_watcher.stream will start its watch after the pod is already deleted, so it won't see the pod in its initial list and won't receive a DELETED event for it.\n\nAs a result, the deleted pod will remain in self.available_engines as a ghost pod indefinitely until the next reconnect.\n\nTo completely eliminate this race condition and also avoid redundant HTTP requests to all pods on every reconnect (since the watch stream without resource_version lists all pods and triggers ADDED events for all of them), you can:\n1. Perform the list call once in _reconcile_engines and return both the live pods and the resource_version (pods.metadata.resource_version).\n2. Populate/reconcile self.available_engines using that list.\n3. Start the watch stream using that resource_version so it only streams subsequent events.

…rsion

Addresses the review on the first revision.

The list now also populates available_engines and returns its
resourceVersion, which the watch starts from. Without it there was still a
window: a pod deleted between the list and the stream was seen alive by the
list and absent from the stream's own initial list, so no DELETED event ever
arrived. Starting the watch at the list's resourceVersion closes it, and
drops the redundant /v1/models request to every pod on each reconnect.

Since the stream no longer replays ADDED events, the reconciliation is also
the startup discovery path. A ready pod already registered at the same URL,
model label and sleep label is skipped, so the common case costs one list.

K8sServiceNameServiceDiscovery gets the same treatment: it had the identical
watch-reconnect hole.

Each object is reconciled in isolation. One unreachable pod, or a Service
with no Endpoints object, would otherwise abort the whole pass and leave the
router with no engines at all.

Tests: startup discovery, skip of an unchanged pod, readiness and cleared-IP
transitions missed during a disconnect, sleep-label change, resourceVersion
propagation, per-object error isolation, and the service-name equivalents.

Signed-off-by: Julien Renaud <julien.renaud@sancare.fr>
@jStrider

Copy link
Copy Markdown
Author

Thanks — both points were right, and acting on them found a third problem. Pushed a second commit.

Race between the list and the watch. Fixed as suggested: _reconcile_engines() now returns pods.metadata.resource_version and the watch starts from it. This also removes the redundant /v1/models request to every pod on each reconnect, which was real: timeout_seconds is always passed, so disable_retries is True in watch.py and every stream() call restarted without a resource_version — a full list plus an ADDED per pod, every time.

One consequence worth stating explicitly: with a resource_version, the stream no longer replays ADDED, so the reconciliation is now the only discovery path, not just a cleanup. It therefore populates as well as removes, and a ready pod already registered at the same URL, model label and sleep label is skipped so the steady state still costs a single list. I verified on a live cluster that discovery still works — a cold-started worker is picked up and served normally.

K8sServiceNameServiceDiscovery. Same treatment, it had the identical hole.

Third problem, found while doing the above. Because the reconciliation is now the discovery path, a single failing object was enough to break everything: a Service matched by the selector but with no Endpoints object (headless, ExternalName, or just created) makes _check_service_ready raise, the exception escapes to _watch_engines, and the loop retries at 2 Hz with no engine ever discovered. Each object is now reconciled in isolation.

Known limitation, not addressed here: the reconciliation is sequential, and an unknown pod that is k8s-ready but HTTP-unreachable costs health_check_timeout_seconds. With enough such pods the pass could outlive the API server's resource_version retention, yielding a 410 and another list — self-healing, but it would stay list-only without ever streaming. Worth a follow-up if anyone runs into it.

Testing. 13 unit tests covering startup discovery, the skip, readiness and cleared-IP transitions missed during a disconnect, sleep-label changes, resourceVersion propagation, per-object isolation, and the service-name equivalents. Each one was checked against a mutation that neutralises the behaviour it asserts.

End to end on a scale-to-zero GPU cluster, same protocol as before — block the router's egress to the API server, wait for the watch to drop, delete the worker while blind, restore:

WARNING: Serving engine ...-b47s6 no longer exists but was still registered: dropping it
engines still listed: NONE

Against the unpatched image, the same sequence leaves /v1/models advertising the model with no pod behind it and a real request timing out at 45 s.

Comment on lines +730 to +734
return (
known.url == f"http://{pod.status.pod_ip}:{self.port}"
and known.model_label == self._get_model_label(pod)
and known.sleep == (labels.get("sleeping") == "true")
)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we also compare the pod's lora-modified annotation (store it on EndpointInfo)?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes — good catch, it was a real hole. Done in 7adf99a.

lora-modified is the annotation the operator stamps on the pod (triggerPodEvent, loraadapter_controller.go) for the sole purpose of forcing the router to re-read /v1/models. It is the only pod-side signal that the adapter list moved: when an adapter is loaded or unloaded, the URL, the readiness, the model label and the sleeping label are all identical before and after. So _is_unchanged returned True and, on reconnect, a lora-modified event lost during the disconnect was swallowed — the router kept serving a stale adapter list until the next real event on that pod. Same class of bug as the one this PR fixes, applied to adapters.

EndpointInfo now carries lora_modified, _handle_pod reads it through _get_lora_modified() and passes it down to _add_engine, and _is_unchanged compares it. New test test_reconcile_rehandles_pod_whose_lora_annotation_changed, checked against a mutation removing the comparison.

K8sServiceNameServiceDiscovery is deliberately untouched on this point: the operator annotates pods, not Services, so there is no equivalent signal to compare there.

The same commit also applies black to the two previous ones, which is what CI was failing on.

The LoRA operator stamps lora-modified on the pod every time it loads or
unloads an adapter, for the sole purpose of forcing the router to re-read
/v1/models. Nothing else about the pod moves: URL, readiness, model label
and sleep label are all identical before and after.

_is_unchanged therefore skipped such a pod on reconnect, and an adapter
change that happened while the watch was down was never picked up: the
router kept serving a stale adapter list until the next real event. Same
class of bug as the one this series fixes, applied to adapters.

EndpointInfo now carries the annotation and the reconciliation compares it.
K8sServiceNameServiceDiscovery is untouched: the operator annotates pods,
not Services, so there is no equivalent signal there.

Also reformats the two previous commits to black, which CI rejected.

Signed-off-by: Julien Renaud <julien.renaud@sancare.fr>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug: router keeps a deleted pod as a healthy backend when a DELETED event is missed during a watch reconnect

2 participants