Skip to content

feat(p5): integrate durable migration jobs - #57

Merged
flyxl merged 816 commits into
mainfrom
codex/p5-main-pr
Oct 9, 2026
Merged

flyxl merged 816 commits into
mainfrom
codex/p5-main-pr

Conversation

@flyxl

@flyxl flyxl commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

范围

将 P5 JobRuntime 与 Schema Diff、Data Sync、Data Transfer 的最终集成提交到本地 main,并提议将本地 main 合入 GitHub main。

本地 main 在 P5 集成前已比 origin/main 超前 586 个提交;P5 集成带来 222 个提交,另有 1 个交付清理提交。因此 PR 共包含 809 个提交、1,205 个文件(+230,349 / -9,271 行)。该 PR 同时包含此前已存在于本地 main、尚未提交到 GitHub main 的工作。

验证

  • DATAZEN_DRIVERS=all pnpm tauri:build:webdriver:通过(EXIT=0,包含 WebDriver feature)。
  • Schema Diff WDIO:9 个 spec 全部失败,46 个失败用例;SD-006 等不到自动生成的计划,其余主要是窗口未加载或 Next 不可用。
  • Data Sync WDIO:5 个 spec 中 1 通过、4 失败;另跑 unknown-outcome spec 失败。包括同步窗口未出现/stale element、SYNC-REAL-009 执行未启动、SYNC-REAL-026 超时;fault injection 预期 unknown,实际得到 committed。
  • Data Transfer WDIO:12 个 spec 中 2 通过、10 失败,识别到 18 个失败用例。失败点包括类型映射页 Next 禁用、若干 data 路径未出现 Preview、旅程中的 Execute 禁用,以及 DT-OBJ-1 结构预览/执行失败。
  • PostgreSQL/MySQL fixture setup 均成功;WDIO 输出没有数据库错误码。只运行了 WDIO E2E,没有运行单元测试、类型检查或其他门禁。

已知待处理项

ColumnMappingEditor 勾选“新建目标表”后会清空目标表名并要求明确填写。部分旧 WDIO 步骤只点击输入框再 Tab,没有输入名字;这解释了类型映射、structure/both 模式和 DT-OBJ-1 的一部分 Next 禁用失败。其余 WDIO 失败尚未全部定位,当前测试状态不应视为通过。

probe added 30 commits October 5, 2026 23:16
接入 src/lib/session:per-pane QueryPanelSession controller + registry,
面板卸载时移除;数据库切换仅在 session 返回 status==='switched' 时落
switchDatabase,事务打开/丢失时先经 needsTransactionSwitchConfirm 确认。
门禁:tsc EXIT=0;全量 vitest 574 files / 5973 tests passed,EXIT=0。
schemaStore 关系列缓存改按 identity(connectionId + configRevision +
dbSessionId + target)寻址,新增 bindMetadataIdentity 与
metadataIdentityBinding 接缝(避免 activeConnectionStore ↔ schemaStore
模块环)。门禁:tsc EXIT=0;全量 vitest 571 files / 5965 tests passed。
完成上会话遗留 WIP:prepare 把实际用于源侧读取的过滤捕获一次
(spec.filters → source_filter),同一过滤器贯穿读取与产物冻结,
重验/预览从 Artifact 复用同一源侧过滤;TableMeta 新增
unchanged_count(review 面板基线)与 source_filter(重验/预览复用),
均带 #[serde(default)] 保持旧产物可回读。

补齐契约测试:
- CM-46 产物元数据:生效过滤与相同行基线落在
  ChangeSetArtifact.table_meta 上,且经受 serde 往返(落盘/回读不丢);
- CM-43 重新比较用例追加断言:key=2 两侧一致 ⇒
  table_meta[0].unchanged_count == 1、无过滤 ⇒ source_filter 为 None。

本 crate 整体过 rustfmt;其中 filter/recordset/session 等既有文件的
重排为纯格式漂移(import 排序/换行),无语义变化,一并收进提交以
满足 crate 级 fmt 门禁。
Replace the direct preview/execute entry points with the two-Job contract from
docs/architecture/platform/data-migration-jobs.md §2.1: a dataTransferPrepare Job
carries only the frozen plan contract, and a dataTransferApply Job carries only
planId + digest + selectionRevision + the reviewed selection, reading the plan
body back from the stored artifact.

- scope.rs: fail-closed §8 admission. The backend scope declaration is
  mandatory and only the local desktop backend is accepted, and every endpoint
  reference must have the shape of an in-memory session token, so migration
  across backends is refused instead of resolved lazily.
- admission.rs: one planId serves at most one apply Job. Consumption is checked
  before expiry/digest/revision so a replay reports the one-shot fact, and every
  refusal names the re-prepare action.
- assembly.rs: derives the frozen plan from a stored plan through the existing
  validated preview/execution context, so apply never re-renders a plan from a
  second inspection.
- runtime.rs: runs the Jobs on the shared runtime with a shared data-transfer
  service key for both endpoints, replaying an idempotent key instead of
  creating a second Job for an unknown commit.
- plans.rs: plan revision and expiry become stored facts, so a stale review is
  refused; the SQL-file destination is content addressed (CM-49).
开发期台账:四道门禁的命令、逐字结论行与首尾工作区 sha,CM-46/47/48/49
覆盖矩阵,以及前端调用方迁移、SQL 文件端点单测缺口等遗留项。合并时删除。
铁律原文「全系统只能有一个隧道引用计数器,即 TunnelLedger::refs」此前**只靠审计**:
transport.rs 登记过一个已实证的反例 —— 给 RecordingTunnelTransport 加一个
`close_tally: Mutex<usize>` 并让 `close_calls()` 改读它,编译通过且全轨测试全绿,
于是第二本账可以伪装成第一本账。本次把它落成测试期机械闸门。

采用哪个修法:两个候选**都**实现,并补上它们各自都挡不住的那一层。
- 候选 (b) 落成结构:`TunnelEntry` 的计数只经四个注册方法(established / refs /
  add_reference / take_reference)按值进出,`drain()` 不再直读 `refs` 字段。
- 候选 (a) 落成断言:闸门 R1 检查 `refs` 的点访问只出现在那四个方法里。
- 新增 R4 + R5(只有 a/b 拦不住伪装的那一半):实现 TunnelTransport 的结构体
  **不允许**持有任何整数标量字段(剥掉 Mutex/Cell/Arc/Option 包装、本地新类型与
  类型别名**递归展开**),且每个端口文件里「事件 → 账」的纯折函数必须恰好一个,
  所有返回整数的 &self 观测方法必须是它的投影。
- R3:台账的观测账改名 `teardown_calls`(对外仍 `close_calls()`),禁止出现在任何
  判断条件里;字段与访问器刻意不同名,扫描器才分得清两者。
- R6:登记表随 `tunnel/mod.rs` 的模块表自证,`TunnelEntry` 必须私有,`drain` 必须
  唯一且住在 `impl TunnelLedger`;把闸门裁小这件事本身会转红。
闸门住在 `src/tunnel/single_counter_audit.rs`(`#[cfg(test)]`),因此随 `--lib` 跑
**并进 CI**;基元拆到 `source_scan.rs` 只为守单文件规模纪律。
运行时侧另加两条不变量:端口每个读数逐步与「现场从 journal 数出来的原始账」对账;
以及一条绕过台账直接经 trait 发 close 的用例,实证两本账的相等来自构造而非巧合。

反例实验(独立 detached worktree + 独立 CARGO_TARGET_DIR,实测结论):
- 改动**前**复现原始反例:`cargo test -p datazen-runtime --lib -- tunnel`
  → `test result: ok. 23 passed; 0 failed`,EXIT=0 —— 确认缺陷真实存在。
- 改动**后**重做同一次变异:闸门转红(详见 progress.md 的逐字结论行)。

文档:transport.rs 的「已知名洞 / 本轮不裁决」与 mod.rs / ledger.rs 对应措辞
改写为已实现事实,并写明边界(测试期结构闸门,挡得住自然写出来的第二本账,
挡不住 unsafe / 跨 crate 静态 / 宏生成)。§4 的越界禁区未触碰:`mod.rs:52-60`
那条「实现缺失」登记原样保留。
变异实验(独立 detached worktree /tmp/p3fu-tun-mut,独立 target)实测发现 506f424
那版闸门有一个**真实漏洞**:R1 / R2 只盯 `refs` 这个名字,于是「`refs` 照旧留着、
在 TunnelEntry 旁边挂一份逐笔相同的 `shadow_refs: u32`、让注册窄口同步维护它」
编译通过且 `cargo test --lib -- tunnel` 全绿(39 passed / EXIT=0)—— 镜像账与真账
数值永远相同,任何**值**断言都看不出差别。这条正是 CM-32-FU1 要防的伪装本身。

补两条规则关掉它:
- R7:台账**声明**层面,与名字无关。`TunnelEntry` 里「一份落地的计数」形状的字段
  必须**恰好一个**且名为 `refs`;`TunnelLedger` 的这样的字段**至多一个**(那份观测账)。
  判据复用 R4 的剥包装 + 本地新类型递归展开,`Mutex<usize>` / `Cell<isize>` /
  `struct Mirror(Mutex<usize>)` 同样在网里。
- R2+:`drain()` 交给调用方的那个数只许来自 `take_reference()` 的返回值、注册读方法
  或字面量 0 —— 堵「字段照旧由窄口管,但释放结果里的数改从观测账取」这条路。

顺带关掉 R3 的一处静默失效:管辖名单原先只从 `struct TunnelLedger` 的字段清单推导,
植入变异的片段没有该声明时名单退化成只剩访问器名,`if self.teardown_calls == 2`
扫不出来。改为「字段清单 ∪ 已知点名集合」,已知集合的完整性由 R7 背书。

结构:闸门拆成 `single_counter_audit/`(mod.rs 881 行 + kill_tests.rs 475 行),
连同 source_scan.rs 全部落在 800 行纪律内;kill test 作为子模块仍看得见父模块的
私有规则函数,判据未变。R6 因此学会承认目录模块(`<name>.rs` 或 `<name>/mod.rs`
至少登记一个),把闸门拆成目录这件事本身不会绕过登记表检查。
文档:transport.rs / ledger.rs / mod.rs 的口径补上 R7 与两条运行时不变量的名字。

验证:cargo fmt -- --check = 0;--lib 461 passed;tunnel 集成两条 3 + 7 passed;
p4_usecase_journeys 7 passed。原始反例与全部 9 次变异在新代码上重跑,见 progress.md。
Round-1 review returned TEST_FAILED with eight defects; this commit closes
D1, D2, D3, D4, D5, D6 and D8.

- D1: scope.rs no longer forces `job.database_target()` on a SQL-file Job, so
  §6.1 jobs with no target endpoint clear the same-backend gate.
- D2: PipelineBudget::capacity() is now consulted before every stage runs, and
  an over-commit is rejected instead of silently overrunning §6.2.
- D3: a budget overrun keeps the confirmed commit boundaries and reports the
  stage as unbounded instead of returning an empty boundary list (§7).
- D4: the handler receives the runtime cancel token; the pipeline re-checks it
  after each execute and rolls the in-flight batch back without committing a
  boundary, and cancel_data_transfer now takes the request through the P5 job
  repository first (job_api/cancel.rs), falling back to the legacy registry
  only for job ids this client never accepted.
- D5: value_bytes charges the real payload size for Json/Timestamp/Bool/Float
  instead of a flat 16 bytes.
- D6: the apply view publishes the §7 recovery verdict over its own boundaries
  (recoveryVerdict / recoveryResumeThrough / recoveryReason).
- D8: the apply path claims the stored plan before it runs, so one planId can
  never drive two apply Jobs.

Runtime-side limits are recorded in progress.md for the runtime track: the
runtime only samples cancel_requested between stages, StageOutcome carries no
error message, JobResult.error is hardcoded None, and the checkpoint's
recovery_policy is hardcoded to resumeAfterVerify.

Gates (CARGO_TARGET_DIR=/tmp/p5-dt-target, worktree sha 4d13d177d2fe…):
  cargo test -p datazen-data-transfer   -> ok. 209 passed; 0 failed
  cargo test -p datazen-runtime         -> aggregate passed=834 failed=0
  cargo test -p datazen --lib           -> ok. 1732 passed; 0 failed; 6 ignored
  pnpm --config.verify-deps-before-run=false typecheck -> 0 error TS
- 新增「round-1 缺陷 → 修复 → 覆盖测试」表(D1–D6、D8)。
- 修正 round-1 的门禁计数归属(D7):08493c333 实测 1710,1726 来自
  422341c 新增的 16 条 job_api 用例,本轮再 +6 条为 1732。
- 记录 runtime 禁区五项待裁定事实:取消只在 stage 之间采样、StageOutcome
  无错误消息、JobResult.error 恒为 None、recovery_policy 硬编码、
  request_cancel 对终态 Job 返回 CasConflict。
- 更新 round-2 四道门禁逐字结论行与首尾 sha。
四个注册方法在 established() 建账并把新条目 refs 置 1;refs() 是唯一读法,
add_reference() / take_reference() 是唯一两种记账变更。R1/R2/R7 的判据
原先散在散文里,读者只能靠通读全文推断哪条规则管哪个方法——注释把它们
按方法名对齐到规则编号上,纯注释修订,不改任何闸门语义。
single_counter_audit/mod.rs 已 884 行,越过 AGENTS.md 的单文件规模线。
R6(闸门查自己)与 R1–R5(闸门查台账与端口)是两种职责,且 R6 自带两个
纯函数(declared_modules / unregistered_modules)专供 kill test 喂变异
样本,把它们按 R6 单独成文件,读闸门时不必再在 884 行里找。

闸门语义一字未改:
- 三个 R6 用例逐条随 --lib 重跑通过(gate_self_audit::* 三个 ok)
- 新文件自己进 AUDITED 登记表并加进 R6 的必备清单,扫描面比拆分前更宽
- R4/R5 两条 kill 用例(原始反例本体)仍绿,--lib 461 passed 不变

拆分后 mod.rs 675 行 / gate_self_audit.rs 230 行。
`evict_idle_at` 原来写的是 `let Ok(Some(view)) = … else { continue }`,把四类结果
压成同一支「跳过这条」,于是:

- actor 闸门拒收(`CloseRejected("replacementInProgress")`)被当成「已摘行」,白送一笔额度;
- `SessionLost`(§9.4 例程已跑完、物理资源确已关闭)反而**不**摘行,额度压住不还;
- actor 真出错(命令没送达)也当无事发生,僵尸行留在表里。

改成四臂 `match`,逐一记账:

| 结局 | 表项 | 额度 | 计入 evicted |
|---|---|---|---|
| `Ok(Some(view))` | 摘 | 还 | 是 |
| `Ok(None)`(未到期/无物理资源/无期限) | 留 | 不动 | 否 |
| `Err(CloseRejected)`(屏障期拒收,压根没派发) | 留 | 不动 | 否 |
| `Err(SessionLost(_))`(§9.4 已关闭,仅绑定失效) | 摘 | 还 | 否 |
| 其它 `Err`(命令未送达) | 留 | 不动 | 否 |

判据是「关是否已经发出」而不是「拿到了什么答复」:`release` 只在 §9.4 四步全跑完
(句柄成批终结、`state.physical` 清空、绑定作废)之后才写 `SessionLost`,所以这一支
足以证明关闭已发生 ⇒ 摘行还额度;但视图读到的是 `Lost`/`Closing`,不能再当成功驱逐发出去。

补两条用例钉住记账:`registry/tests.rs`
- `不可判定的驱逐摘表还额度_不留僵尸行`(UndecidableCloseBackend ⇒ `SessionLost`)
- `驱逐请求没送达时保留表项_不归还额度`(CrashingCloseBackend ⇒ 表项与额度都在)
- `否定答复仍算送达_已无物理资源的会话照常摘行` 改用 `open_candidate` →
  `destroy_candidate` → `publish_candidate` 造出 `Ok(false)` 那一格,不再依赖那条
  依赖驱逐缺陷的构造路径。
`release` 第 1 步用 `drain_handles()`(`std::mem::take`)取样句柄:先把账清空再发终结
请求。于是第 2 步 (b) `!ready_to_return_to_pool(handle_count())` 读的是**自己刚清空的
那本账**,恒为 `!true == false`——`ready_to_return_to_pool(n) == (n == 0)` 在整条
释放路径上只剩 `n == 0` 一支可达。§9.4 明写「driver 返回 Clean 时若宿主仍有已登记
句柄,宿主检查必须失败」这条断言,在 actor 路径上**永远不会触发**。

改为「只读取样 → 发终结 → 确认后才注销」:

- 第 1 步改用 `handles()`(只读)取样,`registered_before` 留底;
- 新增 `HandleRegistry::retire(&[SessionHandleRef])`:只有后端**逐字确认了这一批**
  (`finalized == batch.len()` 且 `remaining == 0`)才注销,少报、多报、带剩余一律留着;
- 第 2 步的 (a) 用累计确认数比登记数,(b) 用释放后宿主自己数一遍的账,两道都不是冗余;
- `CloseResource.registered_handles` 现在如实带「释放后仍挂着几个」,成为该判据在
  后端请求形状上的可观测面;
- `close_request_for` 的 `_retry_finalize` 参数在此显式丢弃并写明理由:账上剩下的是
  那批**没被确认**的句柄,补发只是把同一个不可知的命运重掷一次骰子,§9.4 的处置是
  「关闭 + 如实记 `Undecided`」,不是重试。

测试(`packages/runtime/tests/registry_release.rs`):
- `归池时driver报Clean但宿主仍有登记句柄不得算归池成功` 从三条【现状】特征断言改为
  正确形状,并断言 `closed_with[0].registered_handles == 2`、终结请求带全部句柄、
  返回的额度可再花;
- 新增 `driver逐批少报且明说还有剩余时宿主账必须留痕并判失`:`finalize_sequence`
  `(1,0)` `(0,1)` 两资源两批 ⇒ `SessionLost("handlesStillOpen")`,
  `closed_with[0].registered_handles == 1`;
- 新增 `driver两批确认数正好对平但宿主账未清空时仍判不可知`:两批 `(2,0)` `(0,0)`,
  累计确认 2 == 登记 2 让判据 (a) 通过,只有 (b) 会响 ⇒
  `SessionLost("registeredHandlesRemained")`,`registered_handles == 2`。
  这条同时补上 `release.rs` 注释里引用却一直不存在的用例名。

`actor/tests/flow.rs` 的 `actor_终止后收掉物理资源不再重建句柄` 补上 `.finalizes(1, 0)`,
使其中 `closed_with[0].registered_handles == 0` 这条断言保持诚实(账先被清空过的话,
这个 0 说明不了任何事)。

红线条目未动:`空闲到期才驱逐_未到期一律不动手`、`未设空闲期限的会话永不被驱逐` 仍绿。
`epoch_string` 有两份实现:`registry/context.rs` 里一份私有的
`format!("rte-{:08}", …)`,集成测试夹具 `tests/registry_fixtures/mod.rs` 里又抄了一份
同样的字面量。两份都没有交叉引用,改一处不会改另一处,而 `rte-{:08}` 这个口径同时
被 `registry/context.rs` 的幂等短路判据、`directory/application/convert.rs` 的反向解析、
以及测试夹具按名字造 epoch 的地方依赖——三处只要有一处漂移,幂等重放就会拿
「格式对不上」的假象去解释真正的代际不匹配。

收敛成 `registry/epoch.rs` 里的 `pub fn epoch_string(epoch: u64) -> String`:

- `RuntimeEpoch::to_directory_string()` 变成它的转发,调用点不再各写各的格式;
- `context.rs` 删掉私有副本,改 `use crate::registry::epoch::epoch_string;`;
- `tests/registry_fixtures/mod.rs` 改为 `datazen_runtime::registry::epoch::epoch_string`
  的转发,测试不再自带第二份口径。

两条单测钉住线格式(`registry/epoch.rs`):
- `epoch_string_keeps_the_wire_shape`:`rte-00000000` / `rte-00000001` / `rte-00000007` /
  `rte-12345678`,九位不被截断,并且 `strip_prefix("rte-") + parse::<u64>()` 能按
  `application/convert.rs` 的口径原样解析回来;
- `epoch_method_delegates_to_the_single_formatter`:方法与函数对同一批 epoch 给出
  逐字相同的结果,防止将来有人只改其中一边。

说明:本提交同时带进 `context.rs` 的模块文档改动(「提交替换 × 并发空闲驱逐」的
可达性论证,FU3),因为文档与 epoch 的 import 换用在同一个文件里;FU3 的变异清单
与 550 行新用例随下一个提交进库。
新文件 `packages/runtime/tests/registry_evict_replacement.rs`(5 条用例,全绿):

- `旧句柄只在其登记时的资源上终结_新资源从不收到终结请求`:CM-74 第 2 条「提交落在旧
  resource,不出现在新 resource」的可观测面——旧句柄终结请求的 `resource_id` 恒为旧值,
  关闭请求的 `resource_id` 恒为新值,交叉一次都不许有。
- `屏障期间到来的驱逐不得摘行_旧句柄仍由替换例程在原资源终结`:替换 §7.4-6 ⑤⑥ 持屏障
  期间抵达的 `Evict` 被 actor 闸门**在派发前**拒收成 `CloseRejected`;此时若把它当成
  「已经驱逐过了」,第 ⑧ 步 `close_registered` 就会因为行已摘掉而报 `UnknownSession`,
  旧物理资源连同它的活句柄当场泄漏。
- `发布后的候选没有空闲期限_驱逐对它是不动`:`open_candidate` 的 `OpenRequest` 把
  `idle_deadline_ms` 置 `None`,候选走 `evict_idle` 的 `Ok(None)` 早退分支。
- `对照_没有屏障时到期驱逐照常摘行还额度` 与 `正例对照_带期限的会话在同一时钟下确实被
  驱逐`:两条对照组,钉住上面三条钉的确实是「屏障 + 无期限」这两个变量,而不是
  「到期就不驱逐」。

可达性论证(写进 `registry/context.rs` 的模块文档,「「提交替换 × 并发空闲驱逐」这个
交错的可达性」一节)——结论是 CM-74 第 2 条里的「句柄被终结到新 resource 上」那半
在现结构下**结构性不可达**,靠三条承重事实:

1. 终结用的是每个句柄**登记时**的 `resource_id`,关闭用的是 actor 自己的
   `state.physical`;候选是另一个 actor、另一本账,而 `HandleRegistry::register` 只在
   单个 actor 内的 `exec::apply_completion` 里被调用——没有任何路径能让旧句柄的登记
   resource 变成新值。
2. 旧会话在 ⑧ `close_registered` 处就被 `forget` 摘行,⑨ `publish_candidate` 在其之后,
   而 ⑧ 是 §9.4 释放的唯一执行者;换言之换资源与关旧资源不会并发。
3. `open_candidate` 造的 `OpenRequest` 带 `idle_deadline_ms: None`,候选 actor 的
   `evict_idle` 在第一句 `physical.is_none()` 之后立刻早退。

可达的那半是**驱逐答复的记账**:屏障期的 `Evict` 被拒收(`CloseRejected`),若当成
「已经驱逐过了」就等于给一条并未关闭的资源发一张关闭回执。

文件头自带变异清单(删掉 `CloseRejected` 分支 / 去掉闸门对 `Evict` 的拒收 / 让驱逐答复
一律变成无操作 / 第 1 步取样退回 `drain_handles` / 把 `idle_deadline_ms` 改成 `Some` /
把分组基准换成当前物理资源),逐条对应上面每一条用例的哪一格会红。
ops.rs 此前承载九个操作、已贴近 AGENTS.md 的单文件规模红线,而 op 9
closeResource 是其中最长的一段(157 行)。本次整段搬到同目录 close.rs,
逐字不改;transaction.rs 已有同构先例。

- ops.rs 只留八个操作,imports 去掉随之迁走的四个类型
- mod.rs 模块表补 close 一行,并补上此前缺失的 transaction 行
- 验证:cargo test -p datazen-runtime --lib = 442 passed; 0 failed(与基线同)
任务书里的 `RuntimeError::OrphanedState` 在本仓库并不存在;线上跑的判据是
`context.rs` 里 `epoch_of(...)` 的 match 失败后返回的
`RuntimeError::InvariantBroken("committedWithoutHostEntry")`。本轮对它的语义
做出裁定并钉死:

- 判据仍然是「在册的这一行必须是目录提交的那一代」,且判据的两半各有归属:
  `Ok(epoch)` 但世代对不上 = 撞号(T8 已钉);`Err` = 连行都不在(本轮新增)。
  只覆盖其一的用例会把另一半漏掉——两者互不遮蔽。
- 报错是唯一自洽的答案。发回执会把调用方导向 `Ok` 却立刻 `UnknownSession`;
  「按大概重开一次」把幂等键变成随机数,且每重试一次白吃一格额度。
- 码落在 `InvariantBroken` 而不是 `UnknownSession`/`SessionLost`:后者都会把
  责任/排查方向指错。代价(此后拿不到回执、需调用方换号重开)写进注释。

新增用例「目录记着已提交但注册表已无该行时同键重放不得签发回执也不得凭空重开」
用生产入口构造这一格:先成功替换,再 `invalidate_worker`(CM-58 隔离路径)摘掉
候选那一行——目录的 committed 记录不随之撤销(§13.1 :799 durable 记录只存请求
摘要与回执投影)。断言错误码逐字稳定,且拒绝是纯账目的:额度不动、表项不变、
不新关任何物理资源。

变异清单新增一条:把「查不到那一行」这一格单独放过,只有新增这条红。
实测发现原先那条 `epoch_method_delegates_to_the_single_formatter` 是**同义反复**:
把 `to_directory_string` 内联成一份**逐字相同**的 `format!("rte-{:08}", …)` 之后,
`--lib` 仍然 `test result: ok. 446 passed`(变异 EXIT=0)。值的比较在原理上就抓不到
「同一份实现被复制成两份」,而 FU4 的真实缺陷恰恰是这个形状:当时 `context.rs` 与
测试夹具各有一份,两端「必须逐字一致」只靠注释担保。原测试名和断言文案却宣称
「方法必须是那一份实现的调用」——这是一条比它能证明的更强的承诺(W-04 同型)。

改法:
1. 改名为 `epoch_method_agrees_with_the_single_formatter`,文案缩到它真正证明的
   「两端今天逐字一致」,并在文档里写明它抓不到内联复制、以及实测过它当时确实抓不到。
2. 新增 `epoch_formatter_has_one_implementation_in_track`:扫描 `src/registry/**` 与
   `tests/` 下名字以 `registry` 开头的文件,只看**非注释行**,断言非注释代码里
   `"rte-` + `{:08` 这个组合只出现在 epoch.rs 一处。文档里写明了扫描范围、排除
   本文件的原因、只扫非注释行的原因,以及已知漏网面(换写法如 `{:08x}` 不被抓,
   但那种改动会先被字面量钉住)。

顺带把一份**同型副本记在代码里**:本轨道边界之外,`src/application/convert.rs` 的
`public_handle` 仍自带一份 `format!("rte-{:08}", …)`(runtime 世代 → 平台 API 线上值)。
本轨道不改那个文件(不在文件面内),已在测试文档里点名位置与收敛动作,并作为待裁定
项交给集成方。扫描按路径前缀排除它,不是白名单,所以它修好时本条无需改动。
data-transfer 的 422341c 在自己的基线 214c445 上新增 datazen-runtime,
而集成分支早已由 schema-diff 轨在第 24 行引入同一依赖。两者行位不同,
git 按行合并不产生文本冲突,于是留下重复键,整个 workspace 无法加载。
缺陷:close_resource 对顶层资源表做**三次**求值 —— resolve() 克隆一次、归池
判定 get_mut() 第二次、归还 permit 再 get_mut() 第三次。而 resources 只在
create_resource 的 insert 出现,全目录没有任何 remove / retain / clear 作用
在它身上,第二次 get_mut 恒为 Some。

因此归池判定里的 None 分支是**可证明的死代码**,而它返回的恰恰是最危险的
一组值:request.protocol_drained + resource.registered_handles() + 空注销
清单 —— 拿宿主自述当实测值、且一条句柄都不注销,还照样写出 Closed 与 permit -1。
这不是无害冗余,是「一旦可达就静默说谎」的形状。

改动:
- evaluate_close() 在**一次**锁持有里算完一次关闭要落账的全部事实(凭证校验、
  注销前句柄数、状态迁移、permit 归还计划)。None 分支收敛为唯一入口并报
  ProviderError::SessionLost,发生在任何台账写入之前。
- resolve() 克隆这一次求值一并去掉:§3.1 的凭证校验改由 handle.verify() 对
  resourceId / runtimeEpoch / owner 三要素显式完成,owner 缺失同口径报 SessionLost。
- 新增 close_cases.rs:closing_an_unknown_resource_id_is_rejected_before_any_journal_write
  覆盖上述分支,并证明拒绝后已知资源仍可正常关闭、台账零新增、permit 余额不变。
- tests.rs 的 6 个脚手架函数(target/owner/pool_key/provider/acquire/close)上移到
  mod.rs 的 #[cfg(test)] 区,close 改名 close_and_release(避开 mod close 同名),
  5 处调用点随之改写。脚手架复制两份会漂移:同一个 pool_key 派生算法写两遍,
  「同池复用」的正例与「换 key」的反例就会各自成立。
- tests.rs 无任何断言被删或放宽:删的是 63 行脚手架 + 3 行 import,共 78 删 12 增。

门禁(HEAD 与工作区 sha 首尾一致,运行期间无第三方改动):
cargo test -p datazen-runtime --lib
  test result: ok. 443 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out
cargo fmt -p datazen-runtime -- --check → exit 0
基线缺陷:`close_resource` 做了三次资源表求值(`resolve()` 快照 + 第二次 `get_mut`
+ 归还 permit 时的第三次 `get_mut`),而 §9.4(b) 的归池判据只读**实测**的句柄数,
`CloseResourceRequest::registered_handles`(宿主按 §6.5 报的账)**从未被读取** ——
「记录一个量、消费另一个量」可以悄悄分叉而门禁全绿。CM-74(连接管理 L1324)要求
「driver 返回 Clean 时若宿主仍有已登记句柄,宿主检查必须失败(§9.4)」,基线无法
执行这条。

改动:

1. `state.rs` 新增 `ClosePrecondition` / `CloseAttempt` / `PermitRelease`,
   `ClosePrecondition::returns_to_pool()` 现在**两个账都消费**:
   `!unconfirmed && protocol_drained && measured_handles == 0 &&
   declared_handles == measured_handles`。
   四个输入都由 `FakeResource::prepare_close` 在**同一把锁里、任何状态迁移之前**
   一次算出;`declared_handles` 直接来自请求字段,分叉因此不可表达。
   在既有 `measured_handles == 0` 合取项之下,本判据与 `declared_handles == 0`
   可证等价 —— 不构成误伤,只是把「宿主检查全部通过」这条 §9.4 前置真正执行。
2. `verify_presentation` / `prepare_close` 把凭证校验(resourceId + runtimeEpoch +
   owner)与全部状态迁移收敛成单一临界区;`close.rs` 改为只调 `prepare_close`,
   消掉后两次查表。`close_receipt` 把 `state` / `resource_release` 的产出口径
   统一到 `CloseAttempt::close_outcome`。
3. `journal/entry.rs` 只改文档:`ReturnedToPool.registered_handles` 记的是
   归池判据消费的那一个实测值;账本分叉时不再记该事件。
4. `close_cases.rs` 新增 5 条用例:正例(两份账一致 ⇒ 必须归池)、宿主**少报**
   (声称 0 / 实测 1)、宿主**多报**(声称 3 / 实测 0)、声明为 0 但协议未排空、
   以及凭证 epoch 不符时必须死在校验而非查表(零台账写入)。

门禁(`cargo test -p datazen-runtime --lib`,HEAD 前后均为 ae20f3f、工作区 sha
前后均为 d48dc98d636d):

    test result: ok. 448 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.22s

`cargo fmt -p datazen-runtime -- --check` 退出码 0;编译零告警。
probe added 28 commits October 8, 2026 11:32
# Conflicts:
#	src/lib/__tests__/migrationJobHydration.test.ts
The real-driver-contract scope guard fails loudly whenever a driver crate
ships a `tests/` directory without binding the shared template. The P5
work added `tests/parameterized_transfer_rejection.rs` to both `rqlite`
and `turso`, but neither crate was added to `UNGUARDED_DRIVER_CRATES`,
so `every_driver_crate_with_tests_either_binds_this_guard_or_is_declared`
failed in both the postgres and the mysql contract suite and took the
aggregated `ci` check down with them.

Neither crate reads an env file from its tests, so they belong in the
declared-gap list with the same reason the other unbound crates carry.
The unscanned-crate report section is pinned to a golden, so its count
and its list are re-pinned in the same commit.
…claim

Three of the eight mutation probes never exercised the rule they were written
for, and the suite reported the resulting vacuous proofs as passes.

M2/M3/M5 appended to the end of `packages/runtime/Cargo.toml`, whose last
table is `[dev-dependencies]`. `normalBuildClosure` walks only normal and
build dependencies, so `tauri` / `axum` / `react` were invisible to the guard
and it stayed green. The probes now name `[dependencies]` explicitly, and the
on-disk check reads the table rather than matching the line anywhere in the
file — the old regex could not tell the two placements apart, which is why the
mismatch stayed invisible.

M4 added `platform-api → datazen-runtime`, which closes the cycle
`runtime → application → platform-api`; cargo rejects a normal dependency cycle
before the guard reads a package, so the probe proved nothing about F-04. It
now stands up a throwaway `server`-layer crate and depends on that, which is an
edge cargo accepts and the layer rule must still reject. The suite's broken-build
detector also learned to recognise "cyclic package dependency", which is how M4
was previously reported as a guard that stayed green.

The baseline guard run now gets a snapshot like every mutation. `cargo
metadata` can rewrite `Cargo.lock` with no manifest change at all, and with no
snapshot that rewrite became the content M1 backed up — so the run ended on a
`REVERT FAILED` naming a lockfile that no probe had touched.
… codes

Four PR-new tests were red. All four pinned strings the surrounding
runtime had already moved past; the behaviour each one is really after
was already correct and is still asserted alongside them.

data-transfer cancel effect (§7):

- `run_data`'s aggregate `StageOutcome` reported `PartiallyApplied`
  whenever a table failed or the pipeline got cancelled, without looking
  at `all_boundaries`. The `Err(TransferError::Cancelled(_))` path one
  screen up calls `cancelled_stage`, which does implement the documented
  rule (nothing committed => `NotStarted`), so the two cancel paths in a
  single handler disagreed.
- `data-migration-jobs.md` §7 is explicit: the stage-reported failure or
  cancel effect is decided by the handler's actual evidence, and having
  no commit boundary does not by itself prove rolledBack. The runtime
  side agrees: `job_cancel_watch.rs` pins cancelled + zero boundaries =>
  `NotStarted` and cancelled + boundaries => `PartiallyApplied`, and
  `runtime.rs` forwards the stage's own `effect_outcome` verbatim.
- The aggregate now matches: cancel + zero boundaries => `NotStarted`,
  failure + zero boundaries => `RolledBack` (what `failed_stage` and
  the runtime's `StageTerminal::Failed` both already mean), either with
  boundaries => `PartiallyApplied`, otherwise `Completed`.
- `kernel_cancel.rs` asserted `RolledBack` for a cancel that confirmed no
  boundary at all - the one thing §7 says a zero-boundary run cannot
  prove. Its own sibling test asserted `PartiallyApplied` for the very
  same zero-boundary cancel, and its boundary/row-count assertions are
  what actually pin that nothing was applied. All three now read
  `NotStarted`, which is what the runtime contract says.

sync reason codes:

- `runtime.rs` rejects an overlapping endpoint budget *before* dispatch
  with its own fixed safe code `jobBudgetRejected`; `dispatchNotStarted`
  belongs to the host's `record_dispatch_failure` path for a dispatch
  that failed while the job was still Queued.
- A stage ending non-succeeded with a known effect gets
  `handlerStageTerminated`; the `stageFailed` fallback in the
  `error_code` ladder is unreachable because `failure_reason` is set
  first.
- Both tests keep their real assertion - the refusal precedes every
  driver round trip (`open_transaction_count() == 0`) - and now pin the
  code the runtime actually persists.
`pnpm test:platform-arch:mutations` is the last rust step and its first act
is to refuse when the worktree is dirty: it edits Cargo.toml files and
reverts them, so it cannot prove the revert unless it starts clean. Exit 2
there fails the job. It never ran in the run this PR fixes - step 264 was
red, so every step after it was skipped.

It would have failed. The rust job's workspace manifest is not in one state
for the whole run. `src-tauri/Cargo.toml` tracks postgres / mysql / sqlite as
ordinary path dependencies and keeps a placeholder section that
`resolve-drivers.mjs --drivers=basic` fills at build time - redis arrives
that way, optionally. The `Restore managed files` step deinjects the
manifest again before the manifest-reading guards, so the last thing to
touch the lockfile is a cargo run over a workspace where redis is not a
dependency at all.

The committed lock had redis anyway, plus four driver crates whose
dependency lists carried the dependency this PR added in first position
instead of sorted. Both are stale: the lock was written against an injected
tree and never re-converged. So every cargo invocation rewrites it, the
worktree is dirty by the time the selfcheck looks, and the selfcheck
refuses instead of running.

The lock is regenerated, not hand-patched, and it is generated against the
committed manifest state - deinjected, as the tree is when the selfcheck
runs. `cargo metadata` over that state yields exactly the five sorts and
nothing else, so the committed file is what cargo itself writes and will
keep writing.

`origin/main` has the same stale redis line. It is harmless there only
because nothing on main inspects the worktree; the selfcheck arrives with
this PR and exposes it.

`src-tauri/Cargo.toml` keeps its empty placeholder section; only the
lockfile is committed.
…t outlive jsdom

CI frontend step 12 (pnpm test:unit:driver-set) failed on PR #57 with

  ReferenceError: document is not defined
    at Timeout._onTimeout src/windows/connection/query/useQueryExecutionGate.tsx:456

vitest.config.ts does not set `globals`, so @testing-library/react never
registers its automatic afterEach cleanup, and src/test/setup.ts does not
register one either. Every renderHook in this file therefore stayed mounted
past the end of the test. The "unassigned bind parameter" case arms a 50ms
paramFocusTimerRef timer whose callback runs document.querySelector; when the
whole file finished before that 50ms elapsed, vitest tore down jsdom first and
the callback then blew up on the missing global. The hook itself already
clears that timer in its unmount cleanup — nothing ever unmounted it.

Calling cleanup() explicitly restores the guarantee the hook was written
against, and keeps the failure deterministic instead of load-dependent.

The leaking timer and the missing cleanup both predate this PR; it only
surfaced here because this run loaded the runner differently.
@flyxl
flyxl merged commit e0f6ba2 into main Oct 9, 2026
3 checks passed
@flyxl
flyxl deleted the codex/p5-main-pr branch October 10, 2026 09:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant