Repository navigation
docs(share): record the Dwarves go-live in Batch 9 - #46
Merged
Merged
Conversation
D1 to D5 ran live on 2026-10-06 with every check green. P1 and P2 did not run because the Air was unreachable over ssh.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
docs(share): record the Dwarves go-live in Batch 9
Proof of done
From
docs/verification/one-host-per-tenant.md:Proof of done: one host per tenant, batch 1
Date: 2026-10-01
Branch: docs/one-host-per-tenant
Spec: docs/specs/SPEC-008-one-host-per-tenant.md
Batch 1 is TASK-1a and TASK-1b (live spikes; commands and answers in
docs/implementation-notes/one-host-per-tenant.md) and TASK-2a (per-link storage on the origin, rows 2, 4, 5).Green run
shellcheck bin/share install.sh tests/share.sh tests/e2e.sh tests/e2e-r2.sh demo/render.sh mac/*.shand/bin/bash -n bin/shareare clean on the same tree.Negative controls
Each patch was applied to
bin/sharein the worktree, the tenant section was run on its own (the section's lines fromtests/share.shunder a minimalcheckharness), thengit checkout -- bin/sharerestored the file and the same run went green (35 ok, 0 FAIL).stage_publish(the pointer line moved below the publish line incmd_add)fails=41 0, got1 1); row 5 "no pub/" (got1 0 0 0, twice)r2_pointer_put's catch-all prints and returns 0)fails=40 0, ungated and gated); row 5 "no pub/, no row, no app" (got1 1 0 0, and1 1 1 0gated: the app was created)git checkout -- bin/share)fails=0Dry trace for the first control: the row 4 run holds the add at
PUT m/c00001throughSHARE_R2_DRY_PAUSE, whose.pausedmarker tells the test it is inside the held call; with the patch,stage_publishhas already moved the stage topub/c00001, so the check reads1where it wants0.Live spikes
TASK-1a and TASK-1b ran against throwaway Cloudflare objects only (names
share-e2e-spk584462andshare-e2e-spk13b60d). Every object was deleted and checked from outside: the API listed none of them on either account afterwards, the minted token's id answered 404, and each zone's authoritative server answered NXDOMAIN for both names.Batch 2
Batch 2 is TASK-2b, TASK-3, and TASK-4, each committed and checked on its own.
TASK-2b: dispatch, reconcile, sweep (rows 3, 6, 26, 27)
node tests/worker.mjsPASS, andshellcheck bin/share install.sh tests/share.sh tests/e2e.sh tests/e2e-r2.sh demo/render.sh mac/*.shclean, on the same tree.Negative controls: each patch went into a copy of tree
25dd9e9(commit 6df0318), the tenant sections ran alone, then the unpatched copy ran green (0 FAIL).rma machine row (the refusal incmd_rm_r2becomes aDELETE m/<id>)0 0), "no DELETE m/, the pointer stays" (got1), "member refresh ... refused" (got1 0 0: the pointer was gone, so the message changed), row 27's machine rm (no pointer left to delete);fails=4r2_record_raw's machine branch raises)0 0 1: skipped, warning printed), the not-JSON leg (got0 0 0: the warning named the pointer instead);fails=3Dry trace for the first control: the member's
rm e10001reads the snapshot, findse10001inr2_machine, and with the patch logsDELETE m/e10001; the check reads one DELETE line in the member'sr2-calls.logwhere it wants none.TASK-3: Worker v3 and r2_healthz (rows 7 to 10, 29)
shellcheckand/bin/bash -n bin/shareclean on the same tree. One earlier full run on this tree failed500-row index answers under 3sonce while the machine was busy; the rerun took 0.84 s and passed.Negative controls on a copy of tree
a6ff281(commit 2fcbdb1):node tests/worker.mjsr7 an expired cloud record is 404, never passed through(got200 1), plus the WORKER_SHA self-checkPASSemptynode tests/worker.mjsr9 PASS empty: a miss is 404 with no fetch call(got101 1: the stub's last answer came back) and the machine-pointer leg (got404 1), plus the WORKER_SHA self-checkr2_healthz(the header branch never fires)state(got0 false), the gated add (got1), the join (got1 0).git, so the twov0.5.1 CLIchecks that rungit showfail in both runs (andprofiles --json wrote nothing under HOMEonce in the red run, a flake); the worktree run above is 0 FAILDry trace for the third control: the dry Worker writes
HTTP/2 503with the pair andx-share-tunnel: 0; with the patchhz_codestays 503, sostatereportsready: false, the gated add dies at its healthz check, and the join's three-in-a-row wait times out.TASK-4: setup --r2 and --no-r2, no alias (rows 11, 12, 13, 32)
node tests/worker.mjsPASS,shellcheckand/bin/bash -n bin/shareclean on the same tree. Every call ran against the dry seam (SHARE_R2_DRY=1, a.testhostname); nothing was created on a Cloudflare account.Negative controls on copies of tree
a2cd324(commit 19926b1), the full suite each:^API POST /zones/zone-dry/workers/routes$), the rerun (a second route POST), row 13 (a route left after--no-r2), the fresh-bucket order;fails=7with the copy's twogit showchecks and one load flakegit showchecks and one load flake (cleanup: nocred1 row gone, an unrelated SPEC-004 row; both copies ran at once)--force(thePASSrefusal skipped when--forceis set)0 0 1: the setup ran on and wrote);fails=3with the twogit showchecksgit showchecksThe copies carry no
.git, so the twov0.5.1 CLIchecks that rungit showfail in every copy run; the worktree run above is 0 FAIL.Dry trace for the first control: the patched setup logs
API POST /zones/zone-dry/workers/routesbefore the firstPUT m/0a000, so the order check finds the route POST's first line above the pointer PUTs and stops at that pattern; the rerun POSTs a second route because the patched line runs on every setup.Batch 3
Batch 3 is the member
hitswording (DEC-012), TASK-5, TASK-6, and TASK-7a, each committed and checked on its own. Every call ran against the dry seams (SHARE_R2_DRY=1,SHARE_ACCESS_DRY=1,.testhostnames); nothing was created on a Cloudflare account.Green runs
Each run:
SHARE_TEST_PORT_BASE=28787 gtimeout 900 bash tests/share.shin the worktree, thennode tests/worker.mjs,shellcheck bin/share install.sh tests/share.sh tests/e2e.sh tests/e2e-r2.sh demo/render.sh mac/*.sh, and/bin/bash -n bin/share.tests/share.shtests/worker.mjshitswording--token-stdin)Negative controls
Each patch went into a
git clone --localof the named commit in a mktemp directory (so thegit showrows run as in the worktree), and the full suite ran there withSHARE_TEST_PORT_BASE=40000, one run at a time. An unpatched clone of the same commit ran green (EXIT=0) for d246452, 1f9fb4b, and 32a502e.hitsprints the bare count again (the note line becomes:)0 hits, 0 visitorsalone)alias_dests dropskipped)f.example.test/<id>{,/*})--no-r2as the step-15 rollback with an alias (tenant_dieignores the alias)1 0 0 0 1: no list,roll back with: share ...printed)statethroughapi_token_cmd(the stored-token check skipped)0 1 6 0: the sentinel command ran, the bucket rows listed, nocloud_error)rows()(the origin's bucket read rebindsindex, outside the prune's subshell)importaccept a symlink (the pre-scan takes anlmode and a->line, the post-scan a link)0 1 1: published, a row written)Two unrelated lines failed in the red runs and are not counted as caught: row 29 "a member joins" in the
api_token_cmdrun (the healthz wait under load; green in every other run), and the real-launchd guard in both TASK-6 red runs. That guard saw a realfoundation.d.sharejob appear on this machine during the first run and leave during the second. The suite stubslaunchctland labels its own jobsshare-selftest-<pid>, so the job came from outside the suite; the worktree runs and the green clone saw no change.An earlier attempt ran two clones at once; the green one failed
prune still removes the payloadand a row 27rmwith exit 5 while the twin run trashed same-named files. Every control above ran alone.Dry trace for the TASK-5 rollback control: with the patch, the step-15 die at
ALIAS-301callsdie "...; roll back with: share --profile ten setup ten.example.test --no-r2", so the output holds noalias rollback for f.example.testline and no6. ...step, and row 30's count reads0where it wants1.Tar measurement
TASK-7a's first step ran on bsdtar 3.5.3 and GNU tar 1.35; the table and the flags it fixed are in
docs/implementation-notes/one-host-per-tenant.md. Nothing was extracted outside the scratch directories: both tars stripped the absolute member into the target and refused the..member, and the probe paths outside it did not exist afterwards.Rollback
Batch 3 changes only
bin/share, its tests, and docs on the PR branch; no release, tap bump, or Cloudflare object exists for it. Rolling it back is reverting commits fcd24bd, d246452, 1f9fb4b, and 32a502e on the branch.Batch 4
Batch 4 is TASK-7b (
migrate, rows 19, 20, 21, 31), committed at f0dc1ac. No dry seam exists formigrate(it refuses an R2-on tenant), so its local coverage is acurl/securityshim pair plusSHARE_MIGRATE_SSHpointing at a wrapper that joins its own argv with spaces and replays it throughfish -c(orsh -c) on the same machine, exactly mirroring what a real ssh round trip does to the command line; nothing was created on a Cloudflare account.Green run
shellcheck bin/share install.sh tests/share.sh tests/e2e.sh tests/e2e-r2.sh tests/e2e-tenant.sh demo/render.sh mac/*.shand/bin/bash -n bin/shareare clean on the same tree (tests/e2e-migrate.shdoes not exist yet; it is TASK-9b's deliverable).The one failing line,
profiles --json wrote nothing under HOME(afind -newermarker check, pre-existing code untouched by this batch), is unrelated to migrate: it passed on a first full run of the same tree and failed on a second run of the identical code, the same intermittent-under-load pattern the TASK-3 and TASK-6 batches above already recorded for other lines. A rerun confined to the migrate section alone (the 20 checks above) is green every time.Negative controls
Each guarantee's patch ran as a one-off (a
sed-patched copy ofbin/share, the relevant migrate scenario run standalone against it, the stock file confirmed green immediately after). No patch was committed.importcall for one id is rewritten to a different id beforebin/shareruns itmigrate_retire'smv "$pub/$id" "$root/migrated/$id"replaced withrm -rf "$pub/$id"migrated/<id>never exists after a run that otherwise completesmigrated/, byte for byte equal to B's copyturned intoaccess_probe_round ...; true || die` (the die never fires)Dry trace for the first control:
cmd_import's own id argument comes from the sender's positional argv, so a wrapper script that rewrites$2beforeexecing the real binary is enough to desync the id client-side of any base64 or checksum check; the manifest digest still matches (it is a digest of the tar's own bytes, not the id), so nothing else catches it except the id itself showing up wrong in B's index, which is exactly what the "ids preserved" guarantee asserts directly.Rollback
Batch 4 changes only
bin/share,tests/share.sh, and docs on the PR branch; no release, tap bump, or Cloudflare object exists for it. Rolling it back is reverting commit f0dc1ac on the branch.Batch 5
Batch 5 is the
kit:batteryreview fix pass on this PR (five findings plus two items raised from the live rehearsal), commits a045013..2a0fb65. Each finding carries its own commit with a failing test shown red before the fix, then green after; the fast per-finding reruns below are scratch slices of the realtests/share.sh(same fixtures and helper functions, the unrelated middle of the file cut out) used purely to get a faster red/green loop during the fix; no scratch file was committed.Findings and their red/green
--tunnel-name; a source on the default name collided, the target reused the source's own tunnel, andmigrate_retirethen deleted it out from under the new origin--tunnel-namedistinct from the source's own (-msuffix);migrate_retirealso refuses the tunnel/DNS deletes if the target's own tunnel id is ever reported equal to this machine'stests/share.shprinted the same "PASS"/"N FAILED" summary whether the run finished or crashed partwayreached_endflag set only at the true end; the EXIT trap reports ABORTED and exits 2 when it is missing; the summary prints the total check countcmd_setup's old_host guard only fires when old_host is set; import into a not_setup profile thensetup <other-host>served its gated rows ungatedaccess=row; migrate's own switch (--token-stdin) is exempt, since it never renames the hostnamepass()treated any 502 as the whole machine offline, same as Cloudflare's 520-527 edge codescf-cache-statusmean offlinenode tests/worker.mjs: a dead-port 502 rewritten to the 503 offline pagecf-cache-statustunnel_namefrom the API bytunnel_idwhen the config key is missing; the code never didcmd_migratenow reads it back from the API when the local key is absent, and dies naming the manual fix if neither is known; note corrected to say where this runs--tunnel-name --force(empty)alias_dests dropanswered HTTP 400 live (3/3 rehearsals), the same PUT shape that worked foraddtwo calls earlierself_hosted_domains/domainmirror fields so the API recomputes them from the newdestinations, instead of echoing back a stale value that no longer names a surviving destination once drop removes italias_dests, cf_dry extended to reject the same self_hosted_domains/destinations inconsistency a real PUT does, fixture carrying a realisticself_hosted_domains: fold dies, destinations keep the alias's own entries. Confirmed again with an isolated jq trace of the exact body construction (same expressions, no harness): pre-fix body retainsself_hosted_domainsnaming a URI absent from the newdestinationstests/e2e-tenant.shhad five script bugs found across three live rehearsals (a cold R2 token, a split TCP read in the echo origin, a missing-file check, two Access-app lookups assuming a folded app renames to the tenant), plus L8'setag_beforereading through A's own profile right after--no-r2removed itsbucket=bash -n tests/e2e-tenant.shand shellcheck clean; the lead reruns R1/L8 liveGreen run
A first run of the same tree, same command, showed one failing line unrelated to any of the seven items above (
row 6: pub/<id> absent while a PROBE fail was logged, asleep 0.05filesystem-polling race against a background watcher, pre-existing and untouched by this batch): the same intermittent-under-load pattern Batch 4 already recorded for a different line. The rerun above is clean.shellcheck bin/share install.sh tests/share.sh tests/e2e.sh tests/e2e-r2.sh tests/e2e-tenant.sh demo/render.sh mac/*.shand/bin/bash -n bin/share tests/share.sh tests/e2e-tenant.share clean on the same tree.Rollback
Batch 5 changes
bin/share,tests/share.sh,tests/worker.mjs,tests/e2e-tenant.sh, anddocs/implementation-notes/one-host-per-tenant.mdon the PR branch; no release, tap bump, or Cloudflare object exists for it. Rolling it back is reverting commits a045013..2a0fb65 on the branch.Batch 6
Batch 6 is the live rerun of
tests/e2e-tenant.shon the fix-pass head (1440259), the fixes that rerun forced, the TASK-10 docs, and the negative controls the earlier batches had not covered. Final live result: 73/73, run logdocs/verification/e2e-tenant-20261005T042854Z.log. TASK-9b (tests/e2e-migrate.sh, rehearsal R2) is not run here: it needs one command from Han on the Air.Live e2e: five runs, one green
Every run used the PR worktree's
bin/share(SHARE_BIN), zoned.foundation, an admin token fromop://Toolkit/cf-api-token/credential, run-scoped 2h publisher tokens minted throughop://Toolkit/cf-tokens-admin/credential, andSHARE_E2E_ACCESS_RULE=group:dwarves-ops. Hostnames wereshare-e2e-<6 hex>pairs chosen per run. The earlier logs are not committed (two of them print the account and zone ids in a rollback line).kid == aud, second fold), so the live 400 is fixed. The rest: the WebSocket client used the local resolver, which answered "No route to host" for a name made seconds earlier (script bug, now DoH).setup --no-r2exited 1 with no output (a real bug inalias_rollback, below). The cleanup deleted the alias bucket with an admin env token on a profile with no stored publisher token (script bug)r2-call'scode=<n> etag=<etag>line asetag=alone, so the marker host restore went out with an emptyIf-Matchand the marker kept itsaliases(script bug)addreturned reached an edge that did not yet enforce the app. Not reproduced in run 5; see the note belowR1 passes: the fold exits 0;
https://<alias>/healthzanswers 301 to the tenant; the folded gated link 302s on the tenant withkid == aud; the rollback list runs in its printed order and--no-r2leaves a plain tunnel (no route); the second fold converges and the alias 301s again; the cleanup deletes the bucket, the Worker, and the tenant DNS.Gate note from run 4: share's own check passed three rounds of
kid == audbefore the link was published, and a fetch through DoH one second later still answered 200 once. Access enforcement is per edge and eventual (ADR-0006 measured seconds to minutes). The e2e keeps the strict check, because "no ungated window afteraddreturns" is the guarantee the leg tests. A lead decision, outside SPEC-008: whetheradd --accessshould probe through a second resolver before it publishes.Cleanup proof, from outside the run
A snapshot script listed every resource class the e2e touches, before and after each run, from the Cloudflare API and
launchctl. Names and truncated ids only. The baseline and every later snapshot were identical (diffempty) after runs 1, 2, 3, 4, and 5.share-*share-f-d-foundationshare-e2e*share-*(not deleted)share-s-d-foundationshare-*share-dfoundationshare ...s.d.foundation,f.d.foundation(the two live Dwarves apps)a-origin,b-member,d-fresh,c-admin)sharefoundation.d.share.dfoundation, the Share Bar app jobThe script's own
leftoverssweep printedcleanup: nothing left for this run's nameson all five runs. No process namedecho.pyorshare-e2eremained. Nothing the run did not create was touched: the two live Access apps,share-f-d-foundation,share-s-d-foundation, andshare-dfoundationare in the baseline and in every later snapshot.Findings the rerun forced
alias_rollbackassignedptrsfrom a pipeline ending ingrep; with no local rows and no machine records the grep matches nothing, and underset -ethe run died with exit 1 and no output, so--no-r2and everytenant_dieprinted nothing for a tenant with an alias and no local rows (e6b2a13){ grep ... || true; }in the pipelinebin/sharein a clone:row 30: with no local rows and no pointers, --no-r2 still refuses and prints the list: expected '1 1 1 1', got '1 0 0 0', 1 FAIL of 1243tests/e2e-tenant.shscript bugs (743b229, dac8541): DoH for the WebSocket client; the R1 rollback removesaliases=from the profile config (step 4 of the printed list); the marker ETag parsed fromcode=<n> etag=<etag>; the alias liveness check reads/healthz(a standalone r2 Worker answers 404 at the site root by design); the cleanup's delete-prefix uses theSHARE_R2_TOKENseamNegative controls
The spec lists 20. Batches 1 to 5 recorded a red run for 14 of them (the expired-record,
PASS-empty, pointer-order, pointer-failure, route-order, alias-destination, symlink,api_token_cmd,rows()-merge, SPEC-007--force, sweep, member-rm, healthz-status, and--no-r2-rollback controls). Batch 6 adds 5 and records one whose spec form is not reachable, with a substitute. Each patch went into agit clone --localof dac8541 in a scratch directory and the suite (ornode tests/worker.mjsfor a Worker patch) ran there, one run at a time with its ownSHARE_TEST_PORT_BASE. The green run for every row is the unpatched clone of the same commit.SHARE_TEST_PORT_BASE=41100 gtimeout 900 bash tests/share.shnode tests/worker.mjsstorage:"machine"record (row 7), tested in the inverse formworker_jsmade unreachable (&& false), so a pointer no longer passes throughnode tests/worker.mjsr7 a machine record passes through(got 404, no fetch) and the r8 pass-through rows (13 FAIL, plusWORKER_SHAdrift, which any Worker edit trips)%2f |%5c |%2e |low escape |//test replaced withif (false)node tests/worker.mjsr18for..%2F,%2f,..%5C,%2E,//, and a low escape (all 404, expected 400), andr7 /x/..%2Fb1c2d3/f.txt is 400 in front of the tunnel(12 FAIL plusWORKER_SHA)add --localas cloud (row 3)--localdie replaced withstorage=cloudrow 3: member add --local f is refused(got0 0), 1 FAILcmd_rmof every moved id at the top ofmigrate_retirerow 19migrate exits 0, B's trees equal A's, the rows move toindex.migrated, the live row stays, migrate still exits 0;row 21rerun skips the moved id (9 FAIL, one of them the harness's idler check after the abort)--tunnel-namefrom the printed migrate rollback (row 31)--tunnel-name $old_tunnel_namedropped from the switch-failure rollback linerow 20: a remote-setup failure ... printing the rollback with A's own tunnel-nameandrow 31b(3 FAIL, one the idler check)bucket=into a plain tunnel config at setup (row 1)echo "bucket=leak"added towrite_configsetupthroughwrite_config. The live e2e covers it (L1 setup, L7--no-r2)--no-r2keepsbucket=(row 13)tenant_configstops droppingbucket=rerun: the config is the same,row 13: the config is the one before setup --r2,row 13: the next add writes no pointer(3 FAIL)One gap is open: the
write_configrow-1 control above. Closing it needs a named-setup fixture with a Cloudflare stub, which no suite row has; it is the one place where only the live e2e proves "a tenant with R2 off writes nobucket=".Docs check (TASK-10)
tests/share.shnow checks that every verb, flag, config key, and state key the tenant docs name is indocs/how-it-works.mdand inbin/share(the four generic state keys are checked in the doc only, sincestorage,type,by, andr2are too common to grep in the source; rows 15 and 1 pin them instateoutput). It has three negative controls inside the suite: a doc missing--storage-default, abin/sharemissing--remote-bin, and a doc missing thestoragestate key each fail the check. All pass in the green run above.Rollback
Batch 6 changes
bin/share(one line inalias_rollback),tests/share.sh,tests/e2e-tenant.sh, README,docs/how-it-works.md,docs/setup.md, ADR-0008, the spec (DEC-013, the 502 rows, Worker version 4), the implementation notes, and this record, on the PR branch. No release, tap bump, or standing Cloudflare object exists for it. Rolling it back is reverting commits e6b2a13..dac8541 and the docs commit that follows.Batch 7
Batch 7 is the lead's three safety decisions and the CI fix, on top of Batch 6.
migrate_retirerefuses the DNS and tunnel delete whenmigrate-tunnel-idreturns nothing, and namesteardown --yesfor laterno DNS record or tunnel is deletedgot0 2; the refusal text absentafter up to 3 tries); the dry seam gainsSHARE_R2_DRY_FAIL_TIMES1 1, a persistent 500 got one PUT, not threeadd --accesspolls the published link until it answers 302 or 403 (SHARE_ACCESS_POST_WAIT, 30 s) before printing it; a link still answering anything else exits 1 with the status, the edge risk, andshare rm <id>; the link stays published200,200,302prints after three probes;403counts; a persistent200exits 1 with no link on stdout, and the warning names the rm commandRed run:
1256checks, 11 FAIL (the 10 above plus one cascade from leftover gated rows, fixed in the test). Green runs on the final tree:1256 checks, PASS, one earlier run showed a single row 29 timing failure (SHARE_R2_WAIT=1join under load) that did not repeat.node tests/worker.mjs: PASS. Live:tests/e2e-tenant.shon 1a76f2e (the first commit with the real link check against Cloudflare), 73/73, logdocs/verification/e2e-tenant-20261005T062235Z.log; the before and after snapshots of every resource class are identical.CI
The CI lint job failed on every push since the tenant work began, and the test jobs have failed since the migrate rows landed. Three causes, all fixed:
bin/shareand SC2002 once intests/share.sh; the local shellcheck 0.11 is quietersetupon the target, which dies withmissing: caddy cloudflaredon a runner with no cloudflared; every dev machine has onecaddyandcloudflaredlast on PATH, so a real binary still wins. Reproduced locally with a PATH that hides cloudflared and fish: red before,1256 checks, PASSafterThe macOS runner also showed one
tty: the prompt appeared and the token did not echofailure on 1440259 that did not recur; it is the known flaky row.Batch 8
Batch 8 is TASK-9b:
tests/e2e-migrate.shand the live rehearsal R2 of the personal move, on throwaway names. Final result: 38/38, run logdocs/verification/e2e-migrate-20261005T093825Z.log. The second machine was the real Air, driven from the Mini withmini-run --host air, migrating over real ssh to the Mini.migrateto B on the same machine with a remotesetupthat does the real switch and then reports failure; the printed rollback run on Akid == audindex.migratedand its trees undermigrated/migrate --tothe Mini over sshMeasured gap: probes every 0.5 s through DoH on a snapshot link logged no non-200 answer across the switch on either run (37 probes in M1; the M2 gap logger likewise). The spec expected up to about a minute of 530; the DNS record flips and the new tunnel connects inside a probe interval at this scale, so the number to plan with is "under a few seconds", not a minute.
Findings the rehearsal forced
migratearchived only<name>, but a bare-file snapshot keeps its generatedindex.htmlbeside the file at the id root. The receiver counted one file where the sender's manifest counted two and refused the share as a truncated copy: every single-file link would have failed the personal move (9ae6350)pub/<id>tree;importchecks that<name>is in itindex.htmlit lacked: 17 FAIL in rows 19 to 31port for this profile is already in use; the printed rollback led to a dead end (f37e816)1 1tests/e2e-migrate.shitself: profile names must be lowercase; the Air-side run needs the realHOMEfor ssh while share keeps a throwaway one; a new name needs a moment to resolveBoth product bugs live only on the real-account path: no fixture in the suite had a generated index beside a bare file or a target that was already serving.
Cleanup proof
Snapshots of every Cloudflare resource class (Workers, domains, routes, DNS, tunnels, buckets, Access apps, minted tokens) and of
launchctland LaunchAgents on the Mini were identical before and after the runs. On the Air, the count offoundation.d.sharelaunchd jobs and the LaunchAgents listing were the same before and after (its reals.han.wsservice was never touched: the legs used profilesmgaandmgbunder throwaway HOMEs). The script's own sweep printedcleanup: nothing left for this run's nameson every run. One leftover surfaced: the M2 target's data root under the real HOME (~/share/profiles/mgb), because the script's cleanup only looked for the profile config. The script now moves that root aside too.Rollback
Batch 8 adds
tests/e2e-migrate.shand changesbin/share(migrate_tar, the import root check, the migrate preflight) andtests/share.sh. No release, tap bump, or standing Cloudflare object exists for it.Batch 9: live go-live, Dwarves tenant (2026-10-06)
Han typed the go in the operator session. Snapshots:
tests/prod-snapshot.shbefore D1 and before D3, kept outside git. The pre-D1 and pre-D3 snapshots differ only inshare 0.8.0vsshare 0.9.0.bin/release --yesv0.9.0, notarized Share Bar 0.9.0, tap PRs homebrew-tools #28 and #29 merged,brew upgrade shareon the Minibrew list --versions share; status and live links unchangedshare 0.9.0; healthz 200 on both hosts; a68960 and ba6377 302setup s.d.foundation --r2 --bucket share-dfoundation --alias f.d.foundation(rc 0)lslists a68960 machine and ba6377 cloudapi_token_cmdline that pointed at the admin token (config backed up first)api-token --check[cut at 40000 characters; the full proof is in https://github.com/dwarvesf/share/blob/973f673b313dbfd323c1817d02ebd4ade1fd3560/docs/verification/one-host-per-tenant.md]