You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Load-balancer-aware rolling deploys: take each node out of rotation, drain, deploy, rejoin when healthy #2975
Roll a replicated deploy_component through the cluster one node at a time behind the load balancer, the way upgrades already roll:
Take the node out of rotation.
Drain its traffic. This is the default; a setting skips it for faster deploys.
Deploy and restart.
Wait until the node is healthy.
Put it back into rotation, then move on to the next node.
Every node is out of rotation the whole time it changes, so no client ever reaches a node mid-swap, and a release that fails its health gate never takes traffic on any node.
Why
restart: 'rolling' today restarts nodes one at a time, but each node stays in rotation the whole time. Its workers are replaced while it keeps taking traffic, and the release is live on its disk before any of them has loaded it.
Step 2's planning review asked for certification before the swap, so that an uncertified release can never become a node's boot target or take traffic. Taking the node out of rotation for the whole change gives that guarantee at the cluster level. It does so without making the per-node activation transaction (.deploy-staging swap, journal, recovery) any more complicated.
Upgrades already roll behind the traffic director. Deploys do not (harper-pro#899).
One node at a time: the restart: 'rolling' job (restart_service with replicated: true, bin/restart.ts). It restarts the origin, then each peer in sequence.
Build everywhere, switch later: staged deploys. activate: false builds the release on every node without loading it, and deployment_id then activates it at that node's turn without rebuilding, certifying it there in a canary when that activation restarts (Staged component deploys: land build-aside-then-swap as a sequence of small PRs #2315 steps 2, 5 and 6).
A health signal: component status from get_status and the step 2 canary verdict.
Sketch, open for design
Settings: per deploy, with a config default.
Rotation mode: drain (the default), immediate (no drain), or off, which is today's behavior.
A drain deadline.
A health-gate deadline.
What a failure does: stop the rollout, or continue and report.
Order:
Stage the release on every node first.
Then, for each node in turn: mark it Unavailable, drain (bounded), activate and restart, run the health gate, mark it Available, move on.
Quorum guard: never take the last node, or a majority, out of rotation at once. A single-node cluster deploys in place.
Failure: stop at the failing node, restore its previous release (step 5), return it to rotation, and report which nodes run which release.
Open questions
Who owns rotation in each environment: the GTM, the Symphony load balancer, or only the availability status?
What counts as drained: in-flight HTTP requests, idle keep-alive connections, WebSocket/MQTT sessions, replication traffic?
How long-lived connections (WebSocket, MQTT, SSE) are handed off, and whether drain waits for them or closes them.
How it relates to restart: true and 'rolling': a new mode, or what 'rolling' becomes?
Summary
Roll a replicated
deploy_componentthrough the cluster one node at a time behind the load balancer, the way upgrades already roll:Every node is out of rotation the whole time it changes, so no client ever reaches a node mid-swap, and a release that fails its health gate never takes traffic on any node.
Why
restart: 'rolling'today restarts nodes one at a time, but each node stays in rotation the whole time. Its workers are replaced while it keeps taking traffic, and the release is live on its disk before any of them has loaded it..deploy-stagingswap, journal, recovery) any more complicated.What already exists
availabilitystatus (set_status { id: 'availability', status: 'Unavailable' },server/status/definitions.ts). The traffic director reads it. Optionally keep a node out of rotation (or gate component load) until secondary indexes are built #1358 proposes driving it from index-build state.restart: 'rolling'job (restart_servicewithreplicated: true,bin/restart.ts). It restarts the origin, then each peer in sequence.activate: falsebuilds the release on every node without loading it, anddeployment_idthen activates it at that node's turn without rebuilding, certifying it there in a canary when that activation restarts (Staged component deploys: land build-aside-then-swap as a sequence of small PRs #2315 steps 2, 5 and 6).get_statusand the step 2 canary verdict.Sketch, open for design
Unavailable, drain (bounded), activate and restart, run the health gate, mark itAvailable, move on.Open questions
availabilitystatus?restart: trueand'rolling': a new mode, or what'rolling'becomes?Related: #2315, #641, #1358, #2294, harper-pro#899.