From c5ce6c4bd88d583a21b8ac73821466cd91363cb6 Mon Sep 17 00:00:00 2001 From: Zeke Date: Thu, 23 Jul 2026 23:44:27 -0700 Subject: [PATCH 1/2] docs(k8s): document the rolling-upgrade failover-before-drain gap (P1-6) P1-6 investigation concluded: DOCUMENT the gap, do NOT wire a preStop failover. There is no self-failover command, and a safe controlled handoff must identify the Raft leader, poll a candidate's lag to 0 cluster-wide, and CLUSTER FAILOVER the candidate -- none of which a single terminating node can drive within its grace window. Cluster orchestration is required; the #392 `ironcache upgrade --cluster` driver already provides it. Adds a "Kubernetes (Helm) rolling upgrades" section to docs/UPGRADE.md (+ a "which procedure applies" row + a pointer in values.yaml near updateStrategy): - Case A -- default chart (no per-slot replicas): a `helm upgrade` image bump takes each node's slots briefly UNAVAILABLE during its graceful restart. RPO=0 (save-on-exit + retained PVC); the availability gap is confined to one node at a time by the readiness gate + ordinal serialization. Safe, not zero-downtime. - Case B -- replicated (HA) cluster: a bare rollout causes an UNPLANNED failover (in-sync replica self-promotes after failover_timeout_secs=5s -> a write-rejection blip; RPO bounded by replica_max_lag=256, 0 for the graceful case). The controlled CLUSTER FAILOVER fence (no blip, RPO=0) is what `ironcache upgrade --cluster` orchestrates; the native StatefulSet rollout does NOT invoke it and the preStop is a lame-duck sleep only. - The gap + why a preStop can't self-fence; drive the controlled failover out of band, and automating failover-before-drain on every pod deletion is operator-charter work. - RPO knob: cluster.minReplicasToWrite (0 by default; must stay 0 with the no-replica topology, raise to >=1 once you have runtime replicas to bound RPO). Grounded in cluster_upgrade_driver.rs (the RPO=0 fence), serve.rs raft-mode CLUSTER mutators (REPLICATE/FAILOVER), and ironcache-config defaults (failover_timeout_secs=5, replica_max_lag=256). Co-Authored-By: Claude Opus 4.8 Claude-Session: https://claude.ai/code/session_01FfFZ8gkkNhDBASuntB72HR --- deploy/helm/ironcache/values.yaml | 8 ++++ docs/UPGRADE.md | 75 +++++++++++++++++++++++++++++++ 2 files changed, 83 insertions(+) diff --git a/deploy/helm/ironcache/values.yaml b/deploy/helm/ironcache/values.yaml index ebf885c..602cd9f 100644 --- a/deploy/helm/ironcache/values.yaml +++ b/deploy/helm/ironcache/values.yaml @@ -302,6 +302,14 @@ preStop: # StatefulSet RollingUpdate: pods roll ONE AT A TIME by descending ordinal, each gated by the # readinessProbe + minReadySeconds. `partition` is a canary lever: freeze ordinals < partition, # verify, then patch it down to 0 to roll the rest. 0 = roll everything (the normal case). +# +# UPGRADE SEMANTICS: a `helm upgrade` that changes image.tag rolls the StatefulSet, but the +# native rollout has NO failover-before-drain hook and the preStop is a lame-duck sleep only. +# With the default no-replica topology, each node's slots are briefly UNAVAILABLE during its +# graceful restart (RPO=0 -- save-on-exit + retained PVC). For a zero-downtime upgrade of a +# REPLICATED cluster, drive the controlled `CLUSTER FAILOVER` fence out of band (the +# `ironcache upgrade --cluster` driver), not a bare `helm upgrade`. See docs/UPGRADE.md +# "Kubernetes (Helm) rolling upgrades". updateStrategy: partition: 0 # A "ready" pod must stay ready this long before the rollout advances to the next ordinal, so a diff --git a/docs/UPGRADE.md b/docs/UPGRADE.md index e6d73f5..19f10ae 100644 --- a/docs/UPGRADE.md +++ b/docs/UPGRADE.md @@ -26,6 +26,7 @@ peer of `docs/RUNBOOK.md` (symptom-to-action) and `DEPLOY.md` (install / config) | One node, no persistence | none available (data is in-memory) | [Single node](#single-node-rolling-upgrade); accept the cold working set, or add persistence first | | One node + `data_dir` + `ironcache.socket` | socket-activation handoff | [Single node](#single-node-rolling-upgrade) -- zero refused connections, warm restart | | A raft cluster (`cluster_mode = raft`) with replicas | replica failover | [HA cluster](#ha-cluster-rolling-upgrade) -- ownership moves off the node before you touch it | +| A Kubernetes StatefulSet (the Helm chart) | per-node availability gap (no-replica default), or out-of-band failover-freeze | [Kubernetes](#kubernetes-helm-rolling-upgrades) -- the native rollout has no failover-before-drain hook | `cluster_mode` and `data_dir` are config keys (`Config::cluster_mode` / `Config::data_dir` in `crates/ironcache-config/src/lib.rs`; TOML `cluster_mode = "raft"` / `data_dir = "..."`, @@ -533,6 +534,80 @@ real raft cluster is load-sensitive (flaky as a hard CI gate). The split --- +## Kubernetes (Helm) rolling upgrades + +A `helm upgrade` that changes `image.tag` triggers a **StatefulSet RollingUpdate**: the +controller deletes and recreates pods one at a time, highest ordinal first, gated by the +readiness probe + `minReadySeconds` (see `deploy/SCALING.md` for why the readiness gate -- +NOT the PodDisruptionBudget -- is what paces the roll). **The native rollout has no +failover-before-drain hook**, so what an image bump costs a slot-owner depends entirely on +whether that slot has an in-sync replica. + +### Case A -- the default chart (no per-slot replicas) + +Out of the box the chart's static topology assigns each of the `replicas` nodes a slot range +as its **sole owner, with no replica** (`cluster.minReplicasToWrite: 0`; replicas are a +runtime-only, opt-in thing -- Case B). So there is **nothing to fail over to**: when the +rollout deletes a primary pod, its slots are **unavailable for the duration of that pod's +graceful restart** -- SIGTERM save-on-exit, pod recreated with the new image, snapshot + +raft-log reload from the **retained** PVC, then `/readyz` passes. + +- **RPO = 0.** The save-on-exit persists the working set and the PVC is retained across the + pod recreate (`persistentVolumeClaimRetentionPolicy`), so the node reloads exactly what it + had. No data is lost. +- **Availability: a per-node gap.** Each node's slots are down for its restart window. The + readiness gate + one-at-a-time ordinal serialization confine it to one node at a time; size + `startupProbe.failureThreshold * periodSeconds` and `terminationGracePeriodSeconds` to your + worst-case reload so a healthy node is not CrashLooped or SIGKILLed mid-save. + +This is **safe** (no data loss) but **not zero-downtime** for the slots on the pod being +rolled. For many caches that is acceptable; if it is not, use Case B. + +### Case B -- a replicated (HA) cluster + +If you have assigned in-sync replicas (`CLUSTER REPLICATE ...` puts a replica +of a slot on another node -- see [HA cluster](#ha-cluster-rolling-upgrade)), a slot can fail +over instead of going dark. But note **how** it fails over under a bare `helm upgrade`: + +- **Bare rollout = UNPLANNED failover.** Deleting a primary pod is an ungraceful owner loss + from the cluster's view: an in-sync replica self-proposes promotion only after + `failover_timeout_secs` of continuous downtime (default 5 s), so those slots see a brief + write-rejection blip + the down-timeout window. RPO is bounded by the replica's lag + (`replica_max_lag`, default 256) and is 0 for the graceful save-on-exit case, but the blip + is client-visible. +- **Controlled failover = no blip, RPO = 0.** The `CLUSTER FAILOVER` fence moves ownership to + a caught-up replica *before* the old primary is touched, with no down-timeout window. That + is exactly what the `ironcache upgrade --cluster` driver orchestrates (pause writes -> drain + the candidate to lag 0 -> `CLUSTER FAILOVER` -> commit; fail-closed on drain timeout). See + [HA cluster rolling upgrade](#ha-cluster-rolling-upgrade). + +**The gap on Kubernetes:** the native StatefulSet rollout does not invoke the controlled +fence, and the chart's `preStop` is a **lame-duck sleep only** (it deprograms the Service +endpoint; it does not fail over slots). A pod's `preStop` also *cannot* safely run the fence +itself: there is no self-failover command -- a controlled handoff has to identify the Raft +leader, poll a candidate's lag to 0 cluster-wide, and `CLUSTER FAILOVER` the *candidate*, none +of which a single terminating node can drive within its grace window. So for a zero-downtime +upgrade of a replicated cluster you must run the controlled failover **out of band**: + +1. Before touching a primary, move its ownership to an upgraded, in-sync replica with the + `ironcache upgrade --cluster` driver or the [manual](#3-promote-a-replica-then-upgrade-the-old-owner) + `CLUSTER FAILOVER` sequence. On Kubernetes this means driving the roll from the driver + (e.g. an `--actuator-command` that recreates each pod in order) rather than a bare + `helm upgrade`, or performing the failover by hand and then letting the rollout proceed. +2. Automating failover-before-drain on *every* pod deletion (so a bare `helm upgrade` becomes + safe) requires a controller reconciling cluster state -- it is out of scope for a + declarative chart and is part of the planned operator (see `deploy/K8S_READINESS_PLAN.md`). + +### RPO knob + +`cluster.minReplicasToWrite` (`min_replicas_to_write`) is 0 by default, so a write is +acknowledged without waiting for any replica. With the default no-replica topology (Case A) +it MUST stay 0 -- `>= 1` would fail every write `-NOREPLICAS`. Once you have assigned runtime +replicas (Case B), raising it to `>= 1` makes an acknowledged write survive that many node +losses (bounding RPO under an ungraceful failover), at a write-latency cost. + +--- + ## Verify the new version took over After any upgrade, confirm the NEW binary is the one serving: From 60d74555142cdd83fd50d7f348a792e81698020a Mon Sep 17 00:00:00 2001 From: Zeke Date: Thu, 23 Jul 2026 23:49:40 -0700 Subject: [PATCH 2/2] docs(k8s): condition the Case A RPO=0 claim on an active save policy (review note) Fact-check was SHIP (all 6 claims confirmed with file:line evidence). Taking the reviewer's optional note: the Case A "RPO=0 on a graceful rolling restart" guarantee holds because save-on-exit fires, which is gated on a save policy being active (has_save_policy() = save_interval_secs > 0). The default chart sets saveIntervalSecs: 900 so it holds, but if an operator disables persistence or sets saveIntervalSecs: 0, save-on-exit is a no-op and the rolled pod comes back cold. Stated the dependency explicitly. Co-Authored-By: Claude Opus 4.8 Claude-Session: https://claude.ai/code/session_01FfFZ8gkkNhDBASuntB72HR --- docs/UPGRADE.md | 9 ++++++--- 1 file changed, 6 insertions(+), 3 deletions(-) diff --git a/docs/UPGRADE.md b/docs/UPGRADE.md index 19f10ae..0c3fbc5 100644 --- a/docs/UPGRADE.md +++ b/docs/UPGRADE.md @@ -552,9 +552,12 @@ rollout deletes a primary pod, its slots are **unavailable for the duration of t graceful restart** -- SIGTERM save-on-exit, pod recreated with the new image, snapshot + raft-log reload from the **retained** PVC, then `/readyz` passes. -- **RPO = 0.** The save-on-exit persists the working set and the PVC is retained across the - pod recreate (`persistentVolumeClaimRetentionPolicy`), so the node reloads exactly what it - had. No data is lost. +- **RPO = 0** -- *provided a save policy is active* (it is by default: `persistence.enabled: + true` + `saveIntervalSecs: 900`). On that policy SIGTERM runs the save-on-exit, which + persists the working set, and the PVC is retained across the pod recreate + (`persistentVolumeClaimRetentionPolicy`), so the node reloads exactly what it had. If you + disable persistence or set `saveIntervalSecs: 0`, save-on-exit does not fire and a rolled + pod comes back cold -- the restart is no longer RPO=0. - **Availability: a per-node gap.** Each node's slots are down for its restart window. The readiness gate + one-at-a-time ordinal serialization confine it to one node at a time; size `startupProbe.failureThreshold * periodSeconds` and `terminationGracePeriodSeconds` to your