Skip to content

Show canary certification and each node's outcome when a deploy restarts #1785

Description

@dawsontoth

Part of HarperFast/create-harper#143.

Problem

Studio deploys with restart: 'rolling' (src/integrations/api/instance/applications/deployComponent.ts). From Harper 5.4.0, a restarting deploy is certified in a canary worker on each node before that node serves it (HarperFast/harper#2981). A release that fails to load is rejected. The release it replaced is put back, or, if there was none, the component is failed closed on that node.

Studio reads none of this:

  • A rejection on another node shows as success. With 'rolling', the response covers only the node that received the deploy. The other nodes take the release one at a time, in a job (restartJobId) that runs after the response. A node that rejects the release fails that job, not the deploy, and Studio never follows the job. The deployment record doesn't help either: it says success before the job reaches the other nodes, and it still says so after a node rejects the release.
  • certification isn't read. Its values are certified, uncertified, unavailable and not-requested, and Studio doesn't type it, so a restart that went out unchecked looks like a certified one.
  • A rejection shows a generic error. A rejected deploy fails with certification as an object: status (rejected or interrupted), reason, failures, restored and failed_closed. Studio shows a generic error toast instead.
  • Warnings are dropped. streamOperation.ts drops event names it doesn't know. One of them is warning, which is how the origin says a node installed something different from it (5.3.1).
  • Nothing links a deploy to its record. deployment_id comes back on the result and on SSEOperationError, but nothing uses it.

Proposal

  1. Show certification with the deploy's outcome. For uncertified and unavailable, add a sentence saying why the release went out unchecked.
  2. On a rejection, show the decision: which node, why, and whether the previous release was restored or the component failed closed.
  3. When the result has a restartJobId, keep the deploy open and follow the job with get_job until it is COMPLETE or ERROR. A failed job's message lists each node's outcome under activated. Show each node as it lands, and treat a failed job as a failed deploy that names the node.
  4. Handle warning events in the progress stream, and keep them with the result.
  5. Link the result, and any error that carries a deployment_id, to the deployment's record (Deployments page: what is live, where it came from, and how each node took it #1786).
  6. Gate the certification parts on 5.4.0. Following the job applies to older versions too, since a rolling restart's other nodes already restart in a job.

The get_components health poll that runs after a deploy (useComponentHealthCheck.ts) can stay, as the fallback for older versions and for an inconclusive stream.

One thing to check: the job runs on the node that received the deploy. Confirm that get_job reaches that node through the central-manager proxy.

Related

Done when

On a two-node 5.4 cluster, a release whose canary fails on one node shows as failed for that node, with the reason and the release that node went back to. A release that loads everywhere shows certified for each node, and Studio doesn't report the deploy as done until the last node has taken it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Fields

    Priority

    P1

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions