From 39f7efd856332f1db7b8c168b68d61ce7ebb0547 Mon Sep 17 00:00:00 2001 From: Zeke Date: Fri, 24 Jul 2026 07:30:41 -0700 Subject: [PATCH 1/2] feat(chart): auto-provision the Grafana dashboard as a sidecar ConfigMap (P2) The starter Grafana dashboard shipped as a standalone file but was not auto-provisioned (the plan flagged this PARTIAL). Adds an opt-in, sidecar-discoverable ConfigMap so `metrics.grafanaDashboard.enabled=true` lands the dashboard in Grafana with no manual import. - Moved deploy/grafana/ironcache-dashboard.json -> deploy/helm/ironcache/dashboards/ (single source; .Files.Get can only read within the chart) and updated the README pointer. - New templates/grafana-dashboard.yaml: a gated ConfigMap labelled `grafana_dashboard: "1"` (the kube-prometheus-stack / grafana sidecar discovery label, both the label key + value configurable), embedding the JSON via .Files.Get, with an optional grafana_folder annotation. Off by default; requires the Grafana sidecar to be enabled + watching the ns. - values: metrics.grafanaDashboard.{enabled,label,labelValue,folder}; schema: typed the same. - CI: a metrics.grafanaDashboard.enabled=true value-set in the deploy-lint matrix + a guard asserting the ConfigMap renders WITH the sidecar label (so a broken gate can't silently ship no dashboard). Verified with helm 3.15.4 + kubeconform 0.6.7: default renders no dashboard; enabled renders a valid ConfigMap whose embedded JSON parses byte-identically to the source (title "IronCache overview", 18 panels); full 7-way matrix + schema all pass. Co-Authored-By: Claude Opus 4.8 Claude-Session: https://claude.ai/code/session_01FfFZ8gkkNhDBASuntB72HR --- .github/workflows/deploy-lint.yml | 11 +++++++- deploy/helm/ironcache/README.md | 8 +++--- .../dashboards}/ironcache-dashboard.json | 0 .../templates/grafana-dashboard.yaml | 25 +++++++++++++++++++ deploy/helm/ironcache/values.schema.json | 22 ++++++++++++++++ deploy/helm/ironcache/values.yaml | 12 +++++++++ 6 files changed, 74 insertions(+), 4 deletions(-) rename deploy/{grafana => helm/ironcache/dashboards}/ironcache-dashboard.json (100%) create mode 100644 deploy/helm/ironcache/templates/grafana-dashboard.yaml diff --git a/.github/workflows/deploy-lint.yml b/.github/workflows/deploy-lint.yml index b773b24e..91f8d784 100644 --- a/.github/workflows/deploy-lint.yml +++ b/.github/workflows/deploy-lint.yml @@ -74,7 +74,8 @@ jobs: # The k3s overlay is a real values file; render THROUGH it so a schema violation or # a bad key in values-k3s.yaml fails CI (helm auto-validates -f files against the schema). k3s_set="-f deploy/helm/ironcache/values-k3s.yaml" - for extra in "" "--set console.enabled=true --set console.replicas=2" "$cluster_tls_set" "--set topologySpread.enabled=true" "$np_set" "$k3s_set"; do + gd_set="--set metrics.grafanaDashboard.enabled=true" + for extra in "" "--set console.enabled=true --set console.replicas=2" "$cluster_tls_set" "--set topologySpread.enabled=true" "$np_set" "$k3s_set" "$gd_set"; do echo "::group::helm template ${extra:-(defaults)}" # $extra must WORD-SPLIT into multiple `--set` flags, so it is # intentionally unquoted (SC2086 does not apply here). @@ -117,6 +118,14 @@ jobs: printf '%s\n' "$rendered" | grep -qE "^ replicas: 1$" \ || { echo "ERROR: k3s overlay did not render replicas: 1"; exit 1; } ;; + *grafanaDashboard.enabled=true*) + # Assert the dashboard ConfigMap rendered AND carries the sidecar discovery label + # (a silently-empty gate would ship no dashboard while kubeconform stays happy). + printf '%s\n' "$rendered" | grep -q "ic-ironcache-grafana-dashboard" \ + || { echo "ERROR: grafanaDashboard.enabled rendered no dashboard ConfigMap"; exit 1; } + printf '%s\n' "$rendered" | grep -qE "grafana_dashboard: \"1\"" \ + || { echo "ERROR: dashboard ConfigMap missing the grafana_dashboard sidecar label"; exit 1; } + ;; esac echo "::endgroup::" done diff --git a/deploy/helm/ironcache/README.md b/deploy/helm/ironcache/README.md index d4b345e7..f569a709 100644 --- a/deploy/helm/ironcache/README.md +++ b/deploy/helm/ironcache/README.md @@ -61,9 +61,11 @@ See `values.yaml` for the full commented surface. The ones that matter first: ## Observability Every node serves `/metrics` + `/livez` + `/readyz` on `metrics.port`; the -starter Grafana dashboard and Prometheus alert rules ship in -`deploy/grafana/ironcache-dashboard.json` and -`deploy/prometheus/ironcache-alerts.yml` (see `docs/METRICS.md`). +starter Grafana dashboard ships in the chart at +`dashboards/ironcache-dashboard.json` (set `metrics.grafanaDashboard.enabled=true` +to auto-provision it as a sidecar-discoverable ConfigMap, or import it manually), +and the Prometheus alert rules in `deploy/prometheus/ironcache-alerts.yml` (see +`docs/METRICS.md`). The chart is CI-gated by `helm lint` + `helm template | kubeconform` for the default and console-enabled value sets (`.github/workflows/deploy-lint.yml`). diff --git a/deploy/grafana/ironcache-dashboard.json b/deploy/helm/ironcache/dashboards/ironcache-dashboard.json similarity index 100% rename from deploy/grafana/ironcache-dashboard.json rename to deploy/helm/ironcache/dashboards/ironcache-dashboard.json diff --git a/deploy/helm/ironcache/templates/grafana-dashboard.yaml b/deploy/helm/ironcache/templates/grafana-dashboard.yaml new file mode 100644 index 00000000..1e2238fc --- /dev/null +++ b/deploy/helm/ironcache/templates/grafana-dashboard.yaml @@ -0,0 +1,25 @@ +{{- /* SPDX-License-Identifier: MIT OR Apache-2.0 */ -}} +{{- if .Values.metrics.grafanaDashboard.enabled }} +# The IronCache Grafana dashboard as a sidecar-discoverable ConfigMap. The Grafana +# dashboard sidecar (kube-prometheus-stack / the grafana chart) watches for ConfigMaps +# carrying the label below and loads them into Grafana automatically -- no manual import. +# Requires that sidecar to be enabled AND watching this namespace (or all namespaces). +# Off by default. The dashboard JSON lives at dashboards/ironcache-dashboard.json in the +# chart (also usable for a manual import). +apiVersion: v1 +kind: ConfigMap +metadata: + name: {{ include "ironcache.fullname" . }}-grafana-dashboard + namespace: {{ .Release.Namespace }} + labels: + {{- include "ironcache.labels" . | nindent 4 }} + {{ .Values.metrics.grafanaDashboard.label }}: {{ .Values.metrics.grafanaDashboard.labelValue | quote }} + {{- with .Values.metrics.grafanaDashboard.folder }} + annotations: + # Place the dashboard in a named Grafana folder (sidecar `grafana_folder` annotation). + grafana_folder: {{ . | quote }} + {{- end }} +data: + ironcache-dashboard.json: |- +{{ .Files.Get "dashboards/ironcache-dashboard.json" | indent 4 }} +{{- end }} diff --git a/deploy/helm/ironcache/values.schema.json b/deploy/helm/ironcache/values.schema.json index 4de226e7..d30e5d62 100644 --- a/deploy/helm/ironcache/values.schema.json +++ b/deploy/helm/ironcache/values.schema.json @@ -236,6 +236,28 @@ "description": "Scrape interval as a Prometheus duration, e.g. 30s." } } + }, + "grafanaDashboard": { + "type": "object", + "description": "Auto-provision the Grafana dashboard as a sidecar-discoverable ConfigMap.", + "properties": { + "enabled": { + "type": "boolean", + "description": "Render the dashboard ConfigMap (off by default)." + }, + "label": { + "type": "string", + "description": "The label key the Grafana sidecar watches (default grafana_dashboard)." + }, + "labelValue": { + "type": "string", + "description": "The label value the Grafana sidecar matches (default \"1\")." + }, + "folder": { + "type": "string", + "description": "Optional Grafana folder (via the grafana_folder annotation); empty = sidecar default." + } + } } } }, diff --git a/deploy/helm/ironcache/values.yaml b/deploy/helm/ironcache/values.yaml index f7fa38c3..c584495b 100644 --- a/deploy/helm/ironcache/values.yaml +++ b/deploy/helm/ironcache/values.yaml @@ -141,6 +141,18 @@ metrics: serviceMonitor: enabled: false interval: 30s + # Auto-provision the IronCache Grafana dashboard as a sidecar-discoverable ConfigMap. + # The Grafana dashboard sidecar (kube-prometheus-stack / the grafana chart) loads any + # ConfigMap carrying `label: labelValue` into Grafana. Requires that sidecar enabled + + # watching this namespace. Off by default; the dashboard JSON also ships at + # dashboards/ironcache-dashboard.json for a manual import. + grafanaDashboard: + enabled: false + # The label + value the Grafana sidecar watches for (kube-prometheus-stack default). + label: grafana_dashboard + labelValue: "1" + # Optional Grafana folder to file the dashboard under (empty = the sidecar default). + folder: "" # --- NetworkPolicy (cache pods) ----------------------------------------------- # A NetworkPolicy for the CACHE StatefulSet (distinct from console.networkPolicy). From 1b5db9cf2a1b1e006e3dca50adc7d0173b3e7e19 Mon Sep 17 00:00:00 2001 From: Zeke Date: Fri, 24 Jul 2026 07:35:09 -0700 Subject: [PATCH 2/2] fix(ci): here-string the deploy-lint content guards + fix moved-dashboard doc refs CI caught a real bug in the new grafanaDashboard guard: `printf '%s\n' "$rendered" | grep -q` under `set -o pipefail` fails on a LARGE render (the ~55KB dashboard-embedded output) -- grep -q closes the pipe on its first match, the still-writing printf takes SIGPIPE, and pipefail propagates that non-zero status, turning a SUCCESSFUL match into a false failure ("rendered no dashboard ConfigMap" + "printf: write error: Broken pipe"). The small renders never tripped it. Converted ALL the content guards to `grep <<< "$rendered"` here-strings (no upstream writer to break) and documented why. Also fixed three doc references left dangling by the dashboard move (README.md, DEPLOY.md, docs/METRICS.md now point at deploy/helm/ironcache/dashboards/ + mention the auto-provision flag). Verified in bash+pipefail on the 55KB render: the here-string guards pass cleanly. Co-Authored-By: Claude Opus 4.8 Claude-Session: https://claude.ai/code/session_01FfFZ8gkkNhDBASuntB72HR --- .github/workflows/deploy-lint.yml | 22 ++++++++++++++-------- DEPLOY.md | 8 +++++--- README.md | 2 +- docs/METRICS.md | 8 +++++--- 4 files changed, 25 insertions(+), 15 deletions(-) diff --git a/.github/workflows/deploy-lint.yml b/.github/workflows/deploy-lint.yml index 91f8d784..b1b754e2 100644 --- a/.github/workflows/deploy-lint.yml +++ b/.github/workflows/deploy-lint.yml @@ -85,45 +85,51 @@ jobs: | kubeconform -strict -summary -verbose \ -ignore-missing-schemas \ -kubernetes-version "${KUBE_VERSION}" + # Content guards below use `grep <<< "$rendered"` here-strings, NOT + # `printf ... | grep -q`: under `set -o pipefail`, `grep -q` closes the pipe on + # its first match, and for a LARGE render (e.g. the embedded Grafana dashboard) + # the still-writing printf then takes SIGPIPE, whose non-zero status pipefail + # propagates -- turning a SUCCESSFUL match into a false failure. A here-string + # has no upstream writer to break. case "$extra" in *console.enabled=true*) # Guard against a broken `console.enabled` gate silently rendering # nothing (kubeconform is happy with zero console objects, so assert # the console Deployment/Service actually appeared). - printf '%s\n' "$rendered" | grep -q "ic-ironcache-console" \ + grep -q "ic-ironcache-console" <<< "$rendered" \ || { echo "ERROR: console.enabled=true rendered no console resources"; exit 1; } ;; *clusterTls.enabled=true*) # #660: with clusterTls on and an empty ca, the Secret must still emit cluster.ca # (defaulted to the cert), or the StatefulSet mounts a missing key and the pod # fails at start. Assert the key is present in the rendered Secret. - printf '%s\n' "$rendered" | grep -q "cluster.ca:" \ + grep -q "cluster.ca:" <<< "$rendered" \ || { echo "ERROR: clusterTls.enabled rendered no cluster.ca Secret key"; exit 1; } ;; *networkPolicy.enabled=true*) # Assert the NetworkPolicy actually rendered AND carries the peer-only bus/repl # rule (client+10000 / client+20000). A silently-empty gate would pass kubeconform # (zero NetworkPolicies is valid) while shipping no isolation at all. - printf '%s\n' "$rendered" | grep -q "kind: NetworkPolicy" \ + grep -q "kind: NetworkPolicy" <<< "$rendered" \ || { echo "ERROR: networkPolicy.enabled=true rendered no NetworkPolicy"; exit 1; } - printf '%s\n' "$rendered" | grep -q "port: 16379" \ + grep -q "port: 16379" <<< "$rendered" \ || { echo "ERROR: NetworkPolicy missing the cluster-bus (client+10000) rule"; exit 1; } ;; *values-k3s.yaml*) # The k3s overlay is the single-node standalone posture: no PDB (replicas=1 # gate) and exactly one replica. Guards the edge posture from silent regression. - if printf '%s\n' "$rendered" | grep -q "kind: PodDisruptionBudget"; then + if grep -q "kind: PodDisruptionBudget" <<< "$rendered"; then echo "ERROR: k3s overlay rendered a PDB (expected none at replicas=1)"; exit 1 fi - printf '%s\n' "$rendered" | grep -qE "^ replicas: 1$" \ + grep -qE "^ replicas: 1$" <<< "$rendered" \ || { echo "ERROR: k3s overlay did not render replicas: 1"; exit 1; } ;; *grafanaDashboard.enabled=true*) # Assert the dashboard ConfigMap rendered AND carries the sidecar discovery label # (a silently-empty gate would ship no dashboard while kubeconform stays happy). - printf '%s\n' "$rendered" | grep -q "ic-ironcache-grafana-dashboard" \ + grep -q "ic-ironcache-grafana-dashboard" <<< "$rendered" \ || { echo "ERROR: grafanaDashboard.enabled rendered no dashboard ConfigMap"; exit 1; } - printf '%s\n' "$rendered" | grep -qE "grafana_dashboard: \"1\"" \ + grep -qE "grafana_dashboard: \"1\"" <<< "$rendered" \ || { echo "ERROR: dashboard ConfigMap missing the grafana_dashboard sidecar label"; exit 1; } ;; esac diff --git a/DEPLOY.md b/DEPLOY.md index f7f42b0e..f25f93ae 100644 --- a/DEPLOY.md +++ b/DEPLOY.md @@ -559,9 +559,11 @@ The endpoint is ON by default at `127.0.0.1:9091` (override with `--metrics-addr raft gauges). Scrape it directly, or enable the chart's `metrics.serviceMonitor`. Full catalog of every `ironcache_*` series and the key `INFO` fields is in -[`docs/METRICS.md`](docs/METRICS.md); a starter Grafana dashboard and Prometheus -alert rules ship in [`deploy/grafana/`](deploy/grafana/) and -[`deploy/prometheus/`](deploy/prometheus/). +[`docs/METRICS.md`](docs/METRICS.md); a starter Grafana dashboard ships in the +chart at +[`deploy/helm/ironcache/dashboards/`](deploy/helm/ironcache/dashboards/) (set +`metrics.grafanaDashboard.enabled=true` to auto-provision it) and the Prometheus +alert rules in [`deploy/prometheus/`](deploy/prometheus/). When something is wrong at 3am, [`docs/RUNBOOK.md`](docs/RUNBOOK.md) is the symptom-to-action index: every operator-visible error string, log line, and probe diff --git a/README.md b/README.md index c86146d6..0badc7bd 100644 --- a/README.md +++ b/README.md @@ -211,7 +211,7 @@ full contract. stateless dashboard server that polls the nodes as a scoped ACL user and never sits on the data path. The HA deployment runbook is [`deploy/CONSOLE_DEPLOY.md`](deploy/CONSOLE_DEPLOY.md). -- A shipped **Grafana dashboard** (`deploy/grafana/ironcache-dashboard.json`) and +- A shipped **Grafana dashboard** (`deploy/helm/ironcache/dashboards/ironcache-dashboard.json`) and **Prometheus alert rules** (`deploy/prometheus/ironcache-alerts.yml`) over the `/metrics` series cataloged in [`docs/METRICS.md`](docs/METRICS.md). - **CalVer rolling releases** on every push to `main` plus formal `v*` releases: diff --git a/docs/METRICS.md b/docs/METRICS.md index 2ac1c240..389f3314 100644 --- a/docs/METRICS.md +++ b/docs/METRICS.md @@ -11,9 +11,11 @@ nothing here is aspirational. Companion artifacts: -- `deploy/grafana/ironcache-dashboard.json` -- a starter Grafana dashboard built - on these series (p99/p99.9, ops/sec, hit ratio, evictions, connections, memory, - per-shard hot-shard detection, replication, persistence). +- `deploy/helm/ironcache/dashboards/ironcache-dashboard.json` -- a starter Grafana + dashboard built on these series (p99/p99.9, ops/sec, hit ratio, evictions, + connections, memory, per-shard hot-shard detection, replication, persistence). + Set `metrics.grafanaDashboard.enabled=true` to auto-provision it as a + sidecar-discoverable ConfigMap. - `deploy/prometheus/ironcache-alerts.yml` -- starter Prometheus alerting rules. ## The ops endpoint