Skip to content

feat(chart): alert rules as a PrometheusRule + ServiceMonitor scrape knobs (P2-6) - #764

Merged
ELares merged 1 commit into
mainfrom
feat/monitoring-completeness
Jul 24, 2026
Merged

ELares merged 1 commit into
mainfrom
feat/monitoring-completeness

Conversation

@ELares

@ELares ELares commented Jul 24, 2026

Copy link
Copy Markdown
Owner

What

Closes the monitoring asymmetry #758 introduced (dashboard auto-provisioned, alert rules left as a bare file) and lands plan item P2-6.

  • templates/prometheusrule.yaml -- a gated PrometheusRule embedding the rule bodies via .Files.Get. The alerts file moved into the chart (deploy/prometheus/ -> deploy/helm/ironcache/alerts/) because .Files.Get is chart-relative -- the same single-source move the dashboard got. The file is still exactly a groups: document, so a plain Prometheus can still load it via rule_files: and promtool check rules it unchanged; the template just indents it under spec:. Render verified to parse with all 6 groups / 8 active rules intact (the 2 other alert: hits are its documented commented-out INFO-only block, not lost rules).
    Includes the real operator gotcha: the Operator only adopts rules its ruleSelector matches, so metrics.prometheusRule.additionalLabels exists for the usual release: <kps> selector.
  • ServiceMonitor (P2-6) -- sampleLimit, scrapeTimeout, relabelings, metricRelabelings, additionalLabels passthrough: the levers for staggering/undersampling scrapes and dropping high-cardinality series. All optional -- at defaults nothing extra renders (no empty keys).
  • values + values.schema.json typed for both.
  • deploy-lint: a mon_set value-set with guards asserting the PrometheusRule renders and carries real rule groups -- a moved/misnamed file would otherwise yield a valid-but-empty spec that kubeconform skips (CRDs have no core schema) -- plus the sampleLimit passthrough.

Also: doc staleness the moves/earlier PRs left behind

The 4 references to the old alerts path, and three chart-README rows still documenting pre-#755/#747/#660 behavior (image.tag "latest"; "a bare helm upgrade ROTATES the cluster secret"; clusterTls.ca "REQUIRED"). Plan's alerts row flipped to DONE.

Verified

helm 3.15.4 + kubeconform 0.6.7: lint clean; full 8-way matrix validates; defaults render byte-identically to main (both new templates gated off).

…r scrape knobs (P2-6)

Closes the monitoring asymmetry #758 introduced (the dashboard was auto-provisioned but the
alert rules stayed a bare file) and lands plan item P2-6.

  - templates/prometheusrule.yaml: a gated PrometheusRule embedding the rule bodies via
    .Files.Get. The alerts file MOVED into the chart (deploy/prometheus/ ->
    deploy/helm/ironcache/alerts/) because .Files.Get is chart-relative -- same single-source
    move the dashboard got in #758. The file is still exactly a `groups:` document, so a plain
    Prometheus can load it via `rule_files:` and `promtool check rules` it unchanged; the
    template just indents it under `spec:`. Verified the render parses with all 6 groups / 8
    active rules intact (the 2 remaining `alert:` hits in the file are its documented
    commented-out INFO-only block, not lost rules).
    Includes the operator gotcha: the Operator only adopts rules its `ruleSelector` matches, so
    metrics.prometheusRule.additionalLabels exists for the usual `release: <kps>` selector.
  - servicemonitor.yaml (P2-6): sampleLimit, scrapeTimeout, relabelings, metricRelabelings and
    additionalLabels passthrough -- the levers for staggering/undersampling scrapes and dropping
    high-cardinality series on a large cluster. All optional: at defaults NOTHING extra renders
    (no empty keys), verified.
  - values + values.schema.json typed for both.
  - deploy-lint: a `mon_set` value-set with guards asserting the PrometheusRule renders AND
    carries real rule groups (a moved/misnamed file would otherwise yield a valid-but-EMPTY spec
    that kubeconform SKIPS, since CRDs have no core schema) plus the sampleLimit passthrough.

Also refreshes docs the moves/earlier PRs left stale: the 4 references to the old alerts path
(README, DEPLOY.md, chart README, docs/METRICS.md), and three chart-README rows that still
documented pre-#755/#747/#660 behaviour (image.tag "latest"; "a bare helm upgrade ROTATES the
cluster secret"; clusterTls.ca "REQUIRED"). Plan updated: the alerts row is now DONE.

Verified with helm 3.15.4 + kubeconform 0.6.7: lint clean; the full 8-way matrix validates;
defaults render byte-identically to main (both new templates are gated off).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FfFZ8gkkNhDBASuntB72HR
@ELares
ELares merged commit c5f249c into main Jul 24, 2026
8 checks passed
@ELares
ELares deleted the feat/monitoring-completeness branch July 24, 2026 18:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant