You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Show canary certification and each node's outcome when a deploy restarts #1785
Studio deploys with restart: 'rolling' (src/integrations/api/instance/applications/deployComponent.ts). From Harper 5.4.0, a restarting deploy is certified in a canary worker on each node before that node serves it (HarperFast/harper#2981). A release that fails to load is rejected. The release it replaced is put back, or, if there was none, the component is failed closed on that node.
Studio reads none of this:
A rejection on another node shows as success. With 'rolling', the response covers only the node that received the deploy. The other nodes take the release one at a time, in a job (restartJobId) that runs after the response. A node that rejects the release fails that job, not the deploy, and Studio never follows the job. The deployment record doesn't help either: it says success before the job reaches the other nodes, and it still says so after a node rejects the release.
certification isn't read. Its values are certified, uncertified, unavailable and not-requested, and Studio doesn't type it, so a restart that went out unchecked looks like a certified one.
A rejection shows a generic error. A rejected deploy fails with certification as an object: status (rejected or interrupted), reason, failures, restored and failed_closed. Studio shows a generic error toast instead.
Warnings are dropped.streamOperation.ts drops event names it doesn't know. One of them is warning, which is how the origin says a node installed something different from it (5.3.1).
Nothing links a deploy to its record.deployment_id comes back on the result and on SSEOperationError, but nothing uses it.
Proposal
Show certification with the deploy's outcome. For uncertified and unavailable, add a sentence saying why the release went out unchecked.
On a rejection, show the decision: which node, why, and whether the previous release was restored or the component failed closed.
When the result has a restartJobId, keep the deploy open and follow the job with get_job until it is COMPLETE or ERROR. A failed job's message lists each node's outcome under activated. Show each node as it lands, and treat a failed job as a failed deploy that names the node.
Handle warning events in the progress stream, and keep them with the result.
Gate the certification parts on 5.4.0. Following the job applies to older versions too, since a rolling restart's other nodes already restart in a job.
The get_components health poll that runs after a deploy (useComponentHealthCheck.ts) can stay, as the fallback for older versions and for an inconclusive stream.
One thing to check: the job runs on the node that received the deploy. Confirm that get_job reaches that node through the central-manager proxy.
On a two-node 5.4 cluster, a release whose canary fails on one node shows as failed for that node, with the reason and the release that node went back to. A release that loads everywhere shows certified for each node, and Studio doesn't report the deploy as done until the last node has taken it.
Part of HarperFast/create-harper#143.
Problem
Studio deploys with
restart: 'rolling'(src/integrations/api/instance/applications/deployComponent.ts). From Harper 5.4.0, a restarting deploy is certified in a canary worker on each node before that node serves it (HarperFast/harper#2981). A release that fails to load is rejected. The release it replaced is put back, or, if there was none, the component is failed closed on that node.Studio reads none of this:
'rolling', the response covers only the node that received the deploy. The other nodes take the release one at a time, in a job (restartJobId) that runs after the response. A node that rejects the release fails that job, not the deploy, and Studio never follows the job. The deployment record doesn't help either: it sayssuccessbefore the job reaches the other nodes, and it still says so after a node rejects the release.certificationisn't read. Its values arecertified,uncertified,unavailableandnot-requested, and Studio doesn't type it, so a restart that went out unchecked looks like a certified one.certificationas an object:status(rejectedorinterrupted),reason,failures,restoredandfailed_closed. Studio shows a generic error toast instead.streamOperation.tsdrops event names it doesn't know. One of them iswarning, which is how the origin says a node installed something different from it (5.3.1).deployment_idcomes back on the result and onSSEOperationError, but nothing uses it.Proposal
certificationwith the deploy's outcome. Foruncertifiedandunavailable, add a sentence saying why the release went out unchecked.restartJobId, keep the deploy open and follow the job withget_jobuntil it isCOMPLETEorERROR. A failed job'smessagelists each node's outcome underactivated. Show each node as it lands, and treat a failed job as a failed deploy that names the node.warningevents in the progress stream, and keep them with the result.deployment_id, to the deployment's record (Deployments page: what is live, where it came from, and how each node took it #1786).The
get_componentshealth poll that runs after a deploy (useComponentHealthCheck.ts) can stay, as the fallback for older versions and for an inconclusive stream.One thing to check: the job runs on the node that received the deploy. Confirm that
get_jobreaches that node through the central-manager proxy.Related
successeven when another node rejects the release harper#3060 records each node's outcome on the deployment record, so the history agrees with what this shows.harper deploy restart=rollingexits 0 before the other nodes have taken the release harper#3058 makesharper deploywait for the job, and weighs having the server stream the job's progress, which would replace pollingget_jobhere.Done when
On a two-node 5.4 cluster, a release whose canary fails on one node shows as failed for that node, with the reason and the release that node went back to. A release that loads everywhere shows
certifiedfor each node, and Studio doesn't report the deploy as done until the last node has taken it.