Skip to content

Deploying from CI: lead with OIDC, and correct the deploy reference where it disagrees with Harper #711

Description

@dawsontoth

Part of HarperFast/create-harper#143.

These are the gaps found while planning create-harper's deploy-on-merge workflow, checked against this repository at 2cfb813 and Harper main at be4009ca5 on 2026-10-06. Line numbers are from 2cfb813. No open issue or PR covers sections 1–3, 5 or 6. Some items were added or left by #707, #710 and #685, so they are recent.

1. The CI guide predates OIDC

learn/developers/deploying-from-ci.mdx was written in #594, before OIDC trusted publishing was documented in #637.

What to change:

  • Lead with OIDC, using one complete workflow that runs as written:
    • on, jobs and runs-on, a checkout, and npm install -g harper;
    • permissions: id-token: write on the deploy job, plus environment;
    • vars.HARPER_CLI_TARGET;
    • harper deploy project=… restart=rolling json=true > deploy.json;
    • a wait that polls harper get_job. Progress goes to stderr (bin/deployRenderer.ts), so stdout stays parseable.
  • Next to the workflow, show the trust policy and the deploy-only role. The role needs get_job (see section 2).
  • Explain how to write the policy for a tag-triggered workflow: pin workflow_path and gate on environment, since workflow_ref can't pin a tag (operations.md L777).
  • Keep the curl paths as the alternative for other CI systems. Use secret references instead of literal tokens, since a literal token needs super user (operations.md L742).
  • Mark the canary text as 5.4.0+.

2. OIDC examples and roles

  • The two OIDC examples are fragments. Both reference/cli/authentication.md (authentication.md L269-279) and operations.md (operations.md L623-632) put environment and steps at the top level, with no on, jobs or runs-on, and no CLI install. The first is even labelled .github/workflows/deploy.yml. They also disagree with each other: by_ref=true restart=true replicated=true against by_ref=true alone, which restarts nothing and so is never certified. Replace both with a link to the guide's workflow.

  • The role and policy examples omit get_job. These examples list operations without it:

    But the guide's wait step polls get_job. An operations allowlist and a token's scope both refuse operations they don't list (utility/operation_authorization.ts, utility/operation_authorization.ts), so the wait step gets a 403. The jobs table's "Role Required: any" for get_job (operations.md L1912) hides this. restart_service isn't needed, because the rolling job is created internally.

  • The CLI overview still recommends the refresh token. reference/cli/overview.md calls the refresh token "recommended for CI/CD" and doesn't mention workload identity (overview.md L161-168). That contradicts authentication.md (authentication.md L55).

3. harper deploy and the CLI reference

4. applications.md → Remote Management

This is the source of the deploying-to-harper-fabric skill, so it is what agents learn from.

5. Deployment records (operations.md)

  • deployment_id is a random UUID, not a content hash. operations.md L1208 calls it a "content hash", but each deploy gets a random UUID (components/deploymentRecorder.ts). The content hash is payload_hash.
  • list_deployments has no default limit. operations.md L1179 says it defaults to 100, but leaving limit out returns every row (components/deploymentOperations.ts).
  • The statuses are incomplete (operations.md L1211-1212):
    • loading, replicating and restarting occur while a deploy runs, so filtering on pending misses deploys in flight.
    • rolled_back and rollback_of are never written. Going back writes a new success record with activated_from, which operations.md L1221 describes only for staged builds.
    • A stage whose peers fail returns an error, but its record stays staged, with error set.
  • Some fields are undocumented, and error is misdescribed. origin_node, restart_mode (immediate, rolling or null), credentials (references only) and payload_blob_present are returned but not documented. error is described as a message (operations.md L1223), but it's an object: { message, code, phase }.
  • The record and "what is deployed" disagree for rolling deploys. With restart: 'rolling', the record says success before the other nodes take the release, and keeps saying it if one of them rejects the release. This stays true until A rolling deploy's record says success even when another node rejects the release harper#3060 ships. Reconcile it with the sentence recommending list_deployments for "what exactly is deployed right now".
  • delete_deployment_payload also accepts staged. operations.md L1261 lists success, failed and rolled_back, and operations.md L959 already says a staged deployment works.

6. Streaming progress

  • Document the SSE events. A streamed deploy_component sends phase, install, peer, warning, payload_dropped, error and done, with done carrying { result }. A replayed log can include truncated.
    • phase can also be pending, success or staged.
    • On 5.4, load brackets the canary's decision, so a deploy that doesn't restart no longer emits it (operations.md L1167, commands.md L131).
  • The live tail only follows a deploy running on the node that answers. get_deployment's live tail (operations.md L1195) depends on this. Through a load balancer, it replays the log and returns, possibly with a status that isn't terminal.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

content📝 Content specific issues and requests - text, examples, missing info, or clarity

Type

Fields

Priority

P2

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions