Skip to content

Load-balancer-aware rolling deploys: take each node out of rotation, drain, deploy, rejoin when healthy #2975

Description

@dawsontoth

Summary

Roll a replicated deploy_component through the cluster one node at a time behind the load balancer, the way upgrades already roll:

  1. Take the node out of rotation.
  2. Drain its traffic. This is the default; a setting skips it for faster deploys.
  3. Deploy and restart.
  4. Wait until the node is healthy.
  5. Put it back into rotation, then move on to the next node.

Every node is out of rotation the whole time it changes, so no client ever reaches a node mid-swap, and a release that fails its health gate never takes traffic on any node.

Why

  • restart: 'rolling' today restarts nodes one at a time, but each node stays in rotation the whole time. Its workers are replaced while it keeps taking traffic, and the release is live on its disk before any of them has loaded it.
  • Staged component deploys: land build-aside-then-swap as a sequence of small PRs #2315 step 2 (canary rollout certification, Certify a restarting deploy in a canary worker before it rolls out #2981, shipping in v5.4) narrows that, but only per worker. It holds one replacement worker out of traffic until the component loads, and restores the previous release if it does not. The node stays in rotation, though, and its other workers keep serving while the canary decides.
  • Step 2's planning review asked for certification before the swap, so that an uncertified release can never become a node's boot target or take traffic. Taking the node out of rotation for the whole change gives that guarantee at the cluster level. It does so without making the per-node activation transaction (.deploy-staging swap, journal, recovery) any more complicated.
  • Upgrades already roll behind the traffic director. Deploys do not (harper-pro#899).

What already exists

Sketch, open for design

  • Settings: per deploy, with a config default.
    • Rotation mode: drain (the default), immediate (no drain), or off, which is today's behavior.
    • A drain deadline.
    • A health-gate deadline.
    • What a failure does: stop the rollout, or continue and report.
  • Order:
    • Stage the release on every node first.
    • Then, for each node in turn: mark it Unavailable, drain (bounded), activate and restart, run the health gate, mark it Available, move on.
  • Quorum guard: never take the last node, or a majority, out of rotation at once. A single-node cluster deploys in place.
  • Failure: stop at the failing node, restore its previous release (step 5), return it to rotation, and report which nodes run which release.

Open questions

Related: #2315, #641, #1358, #2294, harper-pro#899.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Fields

    Priority

    P2

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions