From a7531442fe21df2eb119283853b6f7159a2a1fdf Mon Sep 17 00:00:00 2001 From: Drew Ritter Date: Tue, 21 Jul 2026 12:36:50 -0700 Subject: [PATCH 1/6] exp(codex): surgical spinout mitigations for PRI-2672 manual App validation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Experiment branch — NOT for merge without eval evidence (writing-skills discipline; RED/GREEN campaign to follow if the manual run validates). Two targeted changes against the measured Codex 5.6-era spinout (36% of SDD runs >8h, PRI-2672): 1. codex-tools.md — SDD dispatch rules: fork_turns "none" on every spawn (the default "all" forks the whole transcript and refuses model/effort overrides — S2's terminal 13-agent final wave was all full-history forks at sol/xhigh); capability-check for 0.145+ spawn params with an explicit terra/high role table (re-review terra/medium), no sol seats, no effort escalation between fix rounds; honest inheritance warning for <=0.144 where routing is impossible (T0-probed on 0.144.4: schema is {task_name, message, fork_turns} only, role TOMLs inert). Role-table bet: census shows task seats already ran terra/high in the spinout sessions — the surgical bet is fork_turns:none + no-sol-seats + wave termination, not task-seat downgrade. T1 cross-review matrix will refine reviewer tiers (may support terra/medium or lower). 2. SKILL.md final review — wave closure is policy, not a verdict: strong reviewers find real defects indefinitely, so "review until clean" never terminates; new breakage in the final fix diff joins residual adjudication instead of opening wave two; one-off human review procedures (competing reviewers, scoring) never become standing. Two matching rationalization-table rows. Co-Authored-By: Claude Fable 5 --- skills/subagent-driven-development/SKILL.md | 11 +++++ .../references/codex-tools.md | 41 +++++++++++++++++++ 2 files changed, 52 insertions(+) diff --git a/skills/subagent-driven-development/SKILL.md b/skills/subagent-driven-development/SKILL.md index 6c0b8349d2..819849941f 100644 --- a/skills/subagent-driven-development/SKILL.md +++ b/skills/subagent-driven-development/SKILL.md @@ -413,6 +413,15 @@ rulings, or stop on load-bearing ones. There is no second fix wave — residual load-bearing findings surface to your human partner when finishing-a-development-branch presents the options. +The wave closing is policy, not a verdict. A sufficiently strong reviewer +finds real defects indefinitely, so "review until one comes back clean" +never terminates — the completed wave is the exit, not a clean report. +New Critical/Important breakage in the final fix diff joins the residuals +for adjudication; it does not start a second wave. And review procedures +your human partner sets up for one review — competing reviewers, scoring, +extra seats — apply to that review only. Never adopt them as standing +procedure for reviews they didn't ask about. + ## Finish When the final whole-branch review is clean and its fixes are merged, @@ -434,6 +443,8 @@ Use superpowers:finishing-a-development-branch. | "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. | | "Reviews slow the loop down" | The loop without reviews is just unverified churn. Reviews are the loop's brakes and steering. | | "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Controllers without one have re-dispatched entire completed task sequences. | +| "This new finding is real — one more wave" | Real findings are infinite under a strong reviewer. The completed wave is the exit; adjudicate and route. | +| "They liked competing reviewers earlier, I'll run them again" | One-off review procedures apply to the review they were given for. Re-adopting them unasked is scope creep in review clothing. | ## Example Workflow diff --git a/skills/using-superpowers/references/codex-tools.md b/skills/using-superpowers/references/codex-tools.md index b14b585821..5aa3f5f025 100644 --- a/skills/using-superpowers/references/codex-tools.md +++ b/skills/using-superpowers/references/codex-tools.md @@ -9,6 +9,47 @@ multi_agent = true This enables `spawn_agent`, `wait_agent`, and `close_agent` for skills like `dispatching-parallel-agents` and `subagent-driven-development`. When using subagent-driven-development, close reviewer subagents when their review returns. Keep each implementer subagent open until its task's review passes — the fix loop resumes the implementer — then close it. If your harness cannot send another message to a spawned agent, dispatch each fix round as a fresh implementer carrying the brief, the report file, and the findings. +## SDD dispatch: pin fork_turns and the model on every spawn + +Every `spawn_agent` call in subagent-driven-development sets +`fork_turns: "none"`. The parameter defaults to `"all"`, which forks your +entire session transcript into the child — the opposite of the fresh, +constructed context SDD requires — and a full-history fork also refuses +model and effort overrides. Never omit the parameter, and never pass +`"all"` or a turn count for an SDD dispatch. + +Before Task 1, check your `spawn_agent` tool schema for `model` and +`reasoning_effort` parameters (present on Codex 0.145+). + +**If the parameters exist**, set both explicitly on every dispatch: + +| Role | model | reasoning_effort | +|------|-------|------------------| +| Implementer (all rounds) | `gpt-5.6-terra` | `high` | +| Task reviewer | `gpt-5.6-terra` | `high` | +| Scoped re-review | `gpt-5.6-terra` | `medium` | +| Final whole-branch review | `gpt-5.6-terra` | `high` | + +On Codex this table IS the Model Selection mapping — including the final +review, which stays on terra/high rather than "most capable available." +Never give a subagent your session's model when you run a frontier config +(sol at xhigh or max): reviewer tier never exceeds implementer tier, and +a fix round never gets an effort bump. Rounds 4-5's "more capable model" +means a fresh implementer at the same tier; a task that genuinely needs +more than terra/high is a BLOCKED escalation to your human partner, not +a quiet tier climb. Frontier-tier subagents at inherited effort are the +measured top driver of Codex SDD runs spinning out to 8+ hours +(PRI-2672): review seats that inherited sol found real-but-endless +defects every round, and fix diffs ballooned instead of converging. + +**If the parameters do not exist** (Codex 0.144 and earlier), every child +inherits your session's model and effort and no override is possible — +role files in `~/.codex/agents/` do not attach to spawns either. Say so +to your human partner before starting a plan of more than a few tasks, +and offer the choice: proceed with inheritance, or restart the session +at a lower effort so the whole run — controller and children — pays the +lower rate. + ## Environment Detection Skills that create worktrees or finish branches should detect their From f41d97b51b685a4a04759f89ce38152440b2a23f Mon Sep 17 00:00:00 2001 From: Drew Ritter Date: Tue, 21 Jul 2026 14:30:12 -0700 Subject: [PATCH 2/6] =?UTF-8?q?exp(codex):=20print=20dispatch=20tuples=20f?= =?UTF-8?q?rom=20SDD=20scripts=20=E2=80=94=20v2=20chokepoint=20for=20PRI-2?= =?UTF-8?q?672?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Run-2 dispatch-4 autopsy: after compaction, the codex-tools role table fell out of context and the controller reasoned itself into an inherited sol@max reviewer — the harness's own 'inherited parent model is preferred' plus SKILL.md's 'most capable model' outvoted the compaction summary's distilled 'terra/high' line. Its introspection named the only artifact it actually consults at the moment of dispatch: the review-package output. So the role tuple now rides the script output, reprinted fresh every round, immune to compaction by construction: - task-brief prints role=implementer terra/high fork_turns=none - review-package prints per role via a new leading --role flag (task-review default | re-review | final-review), and SKILL.md call sites pass the flag at the re-review and final-review sites - SKILL.md Model Selection now defers to platform role tables and the printed dispatch hints, closing the 'most capable model' loophole - codex-tools.md tells the controller to copy the printed tuple verbatim Harness detection considered and rejected: gating the print on CODEX_CI risks a silent no-op in exactly the environment that needs it; the line is labeled '(codex spawn_agent)' instead, and CC controllers — whose templates carry their own model fields — can ignore it. Experiment branch for PRI-2672 validation; not for merge without eval evidence. Co-Authored-By: Claude Fable 5 --- skills/subagent-driven-development/SKILL.md | 16 ++++--- .../scripts/review-package | 23 +++++++++- .../scripts/task-brief | 4 ++ .../references/codex-tools.md | 3 ++ tests/claude-code/test-sdd-workspace.sh | 42 +++++++++++++++++++ 5 files changed, 81 insertions(+), 7 deletions(-) diff --git a/skills/subagent-driven-development/SKILL.md b/skills/subagent-driven-development/SKILL.md index 819849941f..468a71a222 100644 --- a/skills/subagent-driven-development/SKILL.md +++ b/skills/subagent-driven-development/SKILL.md @@ -158,6 +158,12 @@ conflicts that only emerge from implementation. Use the least powerful model that can handle each role to conserve cost and increase speed. +When your platform's reference file (using-superpowers → Platform +Adaptation) defines a dispatch role table, that table IS this section's +mapping for your harness. Follow it over the tier language below — including +for the final review and fix-loop escalation — and follow the +`dispatch:` hint lines the task-brief and review-package scripts print. + **Mechanical implementation tasks** (isolated functions, clear specs, 1-2 files): use a fast, cheap model. Most implementation tasks are mechanical when the plan is well-specified. **Integration and judgment tasks** (multi-file coordination, pattern matching, debugging): use a standard model. @@ -340,7 +346,7 @@ output; dispatch the re-review once all three are present. Name the covering test files in the fix message — a one-line fix does not need the whole suite. -**The re-review is scoped.** Run `scripts/review-package PLAN_FILE FIX_BASE HEAD` +**The re-review is scoped.** Run `scripts/review-package --role re-review PLAN_FILE FIX_BASE HEAD` where FIX_BASE is the head the previous review saw, and dispatch [re-review-prompt.md](re-review-prompt.md) with the findings list, the brief, the report file, and the printed diff path. The re-reviewer verdicts @@ -391,7 +397,7 @@ parked-with-ruling at the cap. ## Final Review The final whole-branch review gets a package too: run -`scripts/review-package PLAN_FILE MERGE_BASE HEAD` (MERGE_BASE = the commit the +`scripts/review-package --role final-review PLAN_FILE MERGE_BASE HEAD` (MERGE_BASE = the commit the branch started from, e.g. `git merge-base main HEAD`) and include the printed path in the final review dispatch, so the final reviewer reads one file instead of re-deriving the branch diff with git commands. Dispatch @@ -406,7 +412,7 @@ with the complete findings list — not one fixer per finding. Per-finding fixers each rebuild context and re-run suites; a real session's final-review fix wave cost more than all its tasks combined. Then run exactly one scoped re-review of the fix wave -(`scripts/review-package PLAN_FILE FIX_BASE HEAD` over the fix range, +(`scripts/review-package --role re-review PLAN_FILE FIX_BASE HEAD` over the fix range, [re-review-prompt.md](re-review-prompt.md)). Adjudicate any residual findings as in the task loop's breaker: park with rulings, or stop on load-bearing ones. There is no second fix wave — @@ -494,7 +500,7 @@ Task reviewer: Spec ❌: Implementer: Added progress reporting, extracted PROGRESS_INTERVAL constant. Re-ran test/recovery.test.js — 10/10 passing. Fix report appended. -[Run review-package PLAN_FILE FIX_BASE HEAD; dispatch scoped re-review] +[Run review-package --role re-review PLAN_FILE FIX_BASE HEAD; dispatch scoped re-review] Re-reviewer: Missing progress reporting — ADDRESSED (src/recovery.js:41). Magic number — ADDRESSED (src/recovery.js:7). New breakage: none. Verdict: all findings addressed. @@ -505,7 +511,7 @@ Re-reviewer: Missing progress reporting — ADDRESSED (src/recovery.js:41). ... [After all tasks] -[Run review-package PLAN_FILE MERGE_BASE HEAD; dispatch final code-reviewer, most capable model] +[Run review-package --role final-review PLAN_FILE MERGE_BASE HEAD; dispatch final code-reviewer, most capable model] Final reviewer: All requirements met. Deferred minors triaged: none block merge. [Delete this plan's workspace — the record now lives in git] diff --git a/skills/subagent-driven-development/scripts/review-package b/skills/subagent-driven-development/scripts/review-package index 31852e2abb..26c9c9855e 100755 --- a/skills/subagent-driven-development/scripts/review-package +++ b/skills/subagent-driven-development/scripts/review-package @@ -4,13 +4,31 @@ # call. Using the recorded per-task BASE (not HEAD~1) keeps multi-commit # tasks intact. # -# Usage: review-package PLAN_FILE BASE HEAD [OUTFILE] +# Usage: review-package [--role task-review|re-review|final-review] PLAN_FILE BASE HEAD [OUTFILE] # Default OUTFILE: /.superpowers/sdd//review-...diff # (named per range, so a re-review after fixes gets a distinct fresh file). +# +# The trailing dispatch hint rides this output because the controller reads it +# immediately before spawning the reviewer; skill text loaded at session start +# does not survive context compaction, but this line is reprinted every round. set -euo pipefail +role=task-review +if [ "${1:-}" = "--role" ]; then + [ $# -ge 2 ] || { echo "usage: review-package [--role task-review|re-review|final-review] PLAN_FILE BASE HEAD [OUTFILE]" >&2; exit 2; } + role=$2 + shift 2 +fi + +case "$role" in + task-review) hint_role=task-reviewer; hint_effort=high ;; + re-review) hint_role=scoped-re-reviewer; hint_effort=medium ;; + final-review) hint_role=final-reviewer; hint_effort=high ;; + *) echo "bad --role: ${role} (task-review|re-review|final-review)" >&2; exit 2 ;; +esac + if [ $# -lt 3 ] || [ $# -gt 4 ]; then - echo "usage: review-package PLAN_FILE BASE HEAD [OUTFILE]" >&2 + echo "usage: review-package [--role task-review|re-review|final-review] PLAN_FILE BASE HEAD [OUTFILE]" >&2 exit 2 fi @@ -44,3 +62,4 @@ fi commits=$(git rev-list --count "${base}..${head}") echo "wrote ${out}: ${commits} commit(s), $(wc -c < "$out" | tr -d ' ') bytes" +echo "dispatch (codex spawn_agent): role=${hint_role} fork_turns=none model=gpt-5.6-terra reasoning_effort=${hint_effort}" diff --git a/skills/subagent-driven-development/scripts/task-brief b/skills/subagent-driven-development/scripts/task-brief index 612e14a1ee..e5cd4e8827 100755 --- a/skills/subagent-driven-development/scripts/task-brief +++ b/skills/subagent-driven-development/scripts/task-brief @@ -39,3 +39,7 @@ if [ ! -s "$out" ]; then fi echo "wrote ${out}: $(wc -l < "$out" | tr -d ' ') lines" +# The dispatch hint rides this output because the controller reads it +# immediately before spawning the implementer; skill text loaded at session +# start does not survive context compaction, but this line reprints per task. +echo "dispatch (codex spawn_agent): role=implementer fork_turns=none model=gpt-5.6-terra reasoning_effort=high" diff --git a/skills/using-superpowers/references/codex-tools.md b/skills/using-superpowers/references/codex-tools.md index 5aa3f5f025..f7126e26bb 100644 --- a/skills/using-superpowers/references/codex-tools.md +++ b/skills/using-superpowers/references/codex-tools.md @@ -32,6 +32,9 @@ Before Task 1, check your `spawn_agent` tool schema for `model` and On Codex this table IS the Model Selection mapping — including the final review, which stays on terra/high rather than "most capable available." +The task-brief and review-package scripts reprint the applicable row as a +`dispatch (codex spawn_agent):` line with their output — copy those values +onto the spawn_agent call verbatim, every time, even late in a long session. Never give a subagent your session's model when you run a frontier config (sol at xhigh or max): reviewer tier never exceeds implementer tier, and a fix round never gets an effort bump. Rounds 4-5's "more capable model" diff --git a/tests/claude-code/test-sdd-workspace.sh b/tests/claude-code/test-sdd-workspace.sh index 841723016a..1178f340a1 100755 --- a/tests/claude-code/test-sdd-workspace.sh +++ b/tests/claude-code/test-sdd-workspace.sh @@ -165,6 +165,48 @@ PLAN echo " got: $rp_explicit" fi + # --- dispatch hints ride the script output --- + if [[ "$brief_out" == *"dispatch (codex spawn_agent): role=implementer fork_turns=none model=gpt-5.6-terra reasoning_effort=high"* ]]; then + pass "task-brief prints the implementer dispatch hint" + else + fail "task-brief prints the implementer dispatch hint" + echo " got: $brief_out" + fi + + if [[ "$rp_out" == *"dispatch (codex spawn_agent): role=task-reviewer fork_turns=none model=gpt-5.6-terra reasoning_effort=high"* ]]; then + pass "review-package defaults to the task-reviewer dispatch hint" + else + fail "review-package defaults to the task-reviewer dispatch hint" + echo " got: $rp_out" + fi + + local rp_rereview + rp_rereview="$(cd "$repo" && "$SDD_SCRIPTS/review-package" --role re-review plan-a.md HEAD~1 HEAD)" + if [[ "$rp_rereview" == *"role=scoped-re-reviewer fork_turns=none model=gpt-5.6-terra reasoning_effort=medium"* ]]; then + pass "review-package --role re-review prints the medium-effort hint" + else + fail "review-package --role re-review prints the medium-effort hint" + echo " got: $rp_rereview" + fi + + local rp_final + rp_final="$(cd "$repo" && "$SDD_SCRIPTS/review-package" --role final-review plan-a.md HEAD~1 HEAD)" + if [[ "$rp_final" == *"role=final-reviewer fork_turns=none model=gpt-5.6-terra reasoning_effort=high"* ]]; then + pass "review-package --role final-review prints the final-reviewer hint" + else + fail "review-package --role final-review prints the final-reviewer hint" + echo " got: $rp_final" + fi + + rc=0 + (cd "$repo" && "$SDD_SCRIPTS/review-package" --role bogus plan-a.md HEAD~1 HEAD >/dev/null 2>&1) || rc=$? + if [[ "$rc" -eq 2 ]]; then + pass "review-package rejects an unknown --role with exit 2" + else + fail "review-package rejects an unknown --role with exit 2" + echo " exit: $rc" + fi + # --- Worktree isolation: a linked worktree resolves its own workspace --- local wt="$TEST_ROOT/wt" ( cd "$repo" && git worktree add -q "$wt" -b wt-feature ) From c988b0577d6ea052e21f0c2f7d3503e09d344feb Mon Sep 17 00:00:00 2001 From: Drew Ritter Date: Tue, 21 Jul 2026 22:57:28 -0700 Subject: [PATCH 3/6] feat(codex): SDD dispatch routing, fork hygiene, and final-review wave termination MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Codex SDD runs at frontier tiers spin out: subagents inherit the session's model and effort, review seats find real-but-endless defects every round, and the final whole-branch review has no reachable termination state. Field forensics across multiple real sessions traced the mechanism: spawn_agent's fork_turns defaults to "all" (full-context forks that also refuse model overrides), omitted model/effort params inherit the frontier parent, and dispatch rules loaded at session start do not survive context compaction — a compacted controller reverts to inheritance exactly when the session is longest and most expensive. Three coordinated changes: - codex-tools.md: every SDD spawn sets fork_turns "none" plus explicit model/reasoning_effort when the schema supports them (Codex 0.145+), with an inheritance warning for older builds; reviewer tier never exceeds implementer tier and fix rounds never get effort bumps. - task-brief/review-package print a dispatch tuple with their output (review-package grows --role for re-review/final-review), putting the routing values in front of the controller at the moment of dispatch — reprinted every round, so they survive compaction by construction. - SKILL.md: the final-review wave closing is policy, not a verdict — one fix dispatch, one scoped re-review, residuals adjudicated to the ledger; one-off review procedures never become standing; Model Selection defers to platform role tables where one exists. Validated live: a previously spun-out 14-task run (8h, 6 tasks, died mid-loop) re-executed to completion under these changes — bounded fix loops, single-cycle final review, ~20 consecutive correctly-routed dispatches across two compactions. Reviewer-tier matrix (125 cells) showed no calibration cost: clean-diff approvals and planted-defect recall are identical across tiers. Eval scenarios to follow before any upstream submission. Co-Authored-By: Claude Fable 5 --- .../scripts/review-package | 2 + .../scripts/task-brief | 2 + .../references/codex-tools.md | 45 +++++++++---------- 3 files changed, 26 insertions(+), 23 deletions(-) diff --git a/skills/subagent-driven-development/scripts/review-package b/skills/subagent-driven-development/scripts/review-package index 26c9c9855e..a6432e2f39 100755 --- a/skills/subagent-driven-development/scripts/review-package +++ b/skills/subagent-driven-development/scripts/review-package @@ -20,6 +20,8 @@ if [ "${1:-}" = "--role" ]; then shift 2 fi +# Model/effort values are mirrored in using-superpowers/references/codex-tools.md +# (they track Codex's spawn_agent allowlist) — update both together. case "$role" in task-review) hint_role=task-reviewer; hint_effort=high ;; re-review) hint_role=scoped-re-reviewer; hint_effort=medium ;; diff --git a/skills/subagent-driven-development/scripts/task-brief b/skills/subagent-driven-development/scripts/task-brief index e5cd4e8827..814df8aa83 100755 --- a/skills/subagent-driven-development/scripts/task-brief +++ b/skills/subagent-driven-development/scripts/task-brief @@ -42,4 +42,6 @@ echo "wrote ${out}: $(wc -l < "$out" | tr -d ' ') lines" # The dispatch hint rides this output because the controller reads it # immediately before spawning the implementer; skill text loaded at session # start does not survive context compaction, but this line reprints per task. +# Model/effort values are mirrored in using-superpowers/references/codex-tools.md +# (they track Codex's spawn_agent allowlist) — update both together. echo "dispatch (codex spawn_agent): role=implementer fork_turns=none model=gpt-5.6-terra reasoning_effort=high" diff --git a/skills/using-superpowers/references/codex-tools.md b/skills/using-superpowers/references/codex-tools.md index f7126e26bb..15fca4bd42 100644 --- a/skills/using-superpowers/references/codex-tools.md +++ b/skills/using-superpowers/references/codex-tools.md @@ -21,29 +21,28 @@ model and effort overrides. Never omit the parameter, and never pass Before Task 1, check your `spawn_agent` tool schema for `model` and `reasoning_effort` parameters (present on Codex 0.145+). -**If the parameters exist**, set both explicitly on every dispatch: - -| Role | model | reasoning_effort | -|------|-------|------------------| -| Implementer (all rounds) | `gpt-5.6-terra` | `high` | -| Task reviewer | `gpt-5.6-terra` | `high` | -| Scoped re-review | `gpt-5.6-terra` | `medium` | -| Final whole-branch review | `gpt-5.6-terra` | `high` | - -On Codex this table IS the Model Selection mapping — including the final -review, which stays on terra/high rather than "most capable available." -The task-brief and review-package scripts reprint the applicable row as a -`dispatch (codex spawn_agent):` line with their output — copy those values -onto the spawn_agent call verbatim, every time, even late in a long session. -Never give a subagent your session's model when you run a frontier config -(sol at xhigh or max): reviewer tier never exceeds implementer tier, and -a fix round never gets an effort bump. Rounds 4-5's "more capable model" -means a fresh implementer at the same tier; a task that genuinely needs -more than terra/high is a BLOCKED escalation to your human partner, not -a quiet tier climb. Frontier-tier subagents at inherited effort are the -measured top driver of Codex SDD runs spinning out to 8+ hours -(PRI-2672): review seats that inherited sol found real-but-endless -defects every round, and fix diffs ballooned instead of converging. +**If the parameters exist**, set both explicitly on every dispatch. The +task-brief and review-package scripts print the exact values for each +dispatch as a `dispatch (codex spawn_agent):` line with their output — +copy that line's values onto the spawn_agent call verbatim, every time, +even late in a long session. The mapping they print: every SDD seat runs +`gpt-5.6-terra` — implementers and reviewers at `reasoning_effort: high`, +scoped re-reviews at `medium`. On Codex this mapping IS the Model +Selection section — including the final review, which stays on +terra/high rather than "most capable available." Never give a subagent +your session's model when you run a frontier config (sol at xhigh or +max): reviewer tier never exceeds implementer tier, and a fix round +never gets an effort bump. Rounds 4-5's "more capable model" means a +fresh implementer at the same tier; a task that genuinely needs more +than terra/high is a BLOCKED escalation to your human partner, not a +quiet tier climb. Inherited frontier-tier subagents are a measured cause +of SDD runs spinning out for hours: review seats that inherit a frontier +model at maximum effort find real-but-endless defects every round, and +fix diffs balloon instead of converging. + +The model names here track Codex's `spawn_agent` allowlist (currently +`gpt-5.6-sol` and `gpt-5.6-terra`). When the allowlist changes, update +this file and the hint lines in task-brief and review-package together. **If the parameters do not exist** (Codex 0.144 and earlier), every child inherits your session's model and effort and no override is possible — From d1587cabcdb2aa35efca763eeac90ea2269b2be6 Mon Sep 17 00:00:00 2001 From: Drew Ritter Date: Wed, 22 Jul 2026 15:31:50 -0700 Subject: [PATCH 4/6] =?UTF-8?q?refactor(sdd):=20platform-owned=20dispatch?= =?UTF-8?q?=20hints=20=E2=80=94=20shared=20scripts=20carry=20no=20Codex=20?= =?UTF-8?q?literals?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Responds to maintainer review of this branch: the dispatch-hint lines were hardcoded Codex strings inside shared skill tooling that every harness runs. The per-role hint lines now live in a platform-owned data file, skills/using-superpowers/references/codex-dispatch.hints, and the task-brief/review-package scripts only relay the current role's line. The relay is suppressed when CLAUDECODE is set (Claude Code's dispatch templates already carry model selection) and on any harness with no hints file; unknown harnesses fail toward printing, because a silent no-op in the environment that needs the hint is the failure mode this mechanism exists to prevent. Printing at the moment of dispatch is load-bearing: skill text loaded at session start does not survive context compaction, but script output reprints every round. codex-tools.md's SDD dispatch section shrinks to the behavioral rules (fork_turns: "none" always; copy the printed hint verbatim; reviewer tier never exceeds implementer tier; no effort bumps; the <=0.144 inheritance fallback) and points at the hints file for values. Tests cover the relay path, the per-role efforts, and the new CLAUDECODE suppression case. Co-Authored-By: Claude Fable 5 --- .../scripts/review-package | 28 ++++++-- .../scripts/task-brief | 23 +++++-- .../references/codex-dispatch.hints | 11 ++++ .../references/codex-tools.md | 66 +++++++------------ tests/claude-code/test-sdd-workspace.sh | 49 +++++++++----- 5 files changed, 104 insertions(+), 73 deletions(-) create mode 100644 skills/using-superpowers/references/codex-dispatch.hints diff --git a/skills/subagent-driven-development/scripts/review-package b/skills/subagent-driven-development/scripts/review-package index a6432e2f39..b988d4246a 100755 --- a/skills/subagent-driven-development/scripts/review-package +++ b/skills/subagent-driven-development/scripts/review-package @@ -13,6 +13,8 @@ # does not survive context compaction, but this line is reprinted every round. set -euo pipefail +script_dir=$(cd "$(dirname "$0")" && pwd) + role=task-review if [ "${1:-}" = "--role" ]; then [ $# -ge 2 ] || { echo "usage: review-package [--role task-review|re-review|final-review] PLAN_FILE BASE HEAD [OUTFILE]" >&2; exit 2; } @@ -20,12 +22,12 @@ if [ "${1:-}" = "--role" ]; then shift 2 fi -# Model/effort values are mirrored in using-superpowers/references/codex-tools.md -# (they track Codex's spawn_agent allowlist) — update both together. case "$role" in - task-review) hint_role=task-reviewer; hint_effort=high ;; - re-review) hint_role=scoped-re-reviewer; hint_effort=medium ;; - final-review) hint_role=final-reviewer; hint_effort=high ;; + task-review) hint_key=task-review ;; + # re-review maps onto the fix-review hints entry; the --role value itself + # is renamed in the next commit. + re-review) hint_key=fix-review ;; + final-review) hint_key=final-review ;; *) echo "bad --role: ${role} (task-review|re-review|final-review)" >&2; exit 2 ;; esac @@ -45,7 +47,7 @@ git rev-parse --verify --quiet "$head" >/dev/null || { echo "bad HEAD: $head" >& if [ $# -eq 4 ]; then out=$4 else - dir=$("$(cd "$(dirname "$0")" && pwd)/sdd-workspace" "$plan") + dir=$("$script_dir/sdd-workspace" "$plan") out="$dir/review-$(git rev-parse --short "$base")..$(git rev-parse --short "$head").diff" fi @@ -64,4 +66,16 @@ fi commits=$(git rev-list --count "${base}..${head}") echo "wrote ${out}: ${commits} commit(s), $(wc -c < "$out" | tr -d ' ') bytes" -echo "dispatch (codex spawn_agent): role=${hint_role} fork_turns=none model=gpt-5.6-terra reasoning_effort=${hint_effort}" + +# Platform dispatch hints ride this output because the controller reads it +# immediately before spawning; the lines themselves are owned by the platform +# reference layer (using-superpowers/references/*-dispatch.hints), not this +# script. Claude Code's dispatch templates carry model selection already, so +# the relay is suppressed there and on any harness without a hints file. +hints_file="$script_dir/../../using-superpowers/references/codex-dispatch.hints" +if [ -z "${CLAUDECODE:-}" ] && [ -f "$hints_file" ]; then + hint_line=$(grep "^${hint_key}:" "$hints_file" | head -1 | cut -d: -f2- | sed 's/^ *//') || true + if [ -n "$hint_line" ]; then + echo "$hint_line" + fi +fi diff --git a/skills/subagent-driven-development/scripts/task-brief b/skills/subagent-driven-development/scripts/task-brief index 814df8aa83..269939c60e 100755 --- a/skills/subagent-driven-development/scripts/task-brief +++ b/skills/subagent-driven-development/scripts/task-brief @@ -18,10 +18,12 @@ plan=$1 n=$2 [ -f "$plan" ] || { echo "no such plan file: $plan" >&2; exit 2; } +script_dir=$(cd "$(dirname "$0")" && pwd) + if [ $# -eq 3 ]; then out=$3 else - dir=$("$(cd "$(dirname "$0")" && pwd)/sdd-workspace" "$plan") + dir=$("$script_dir/sdd-workspace" "$plan") out="$dir/task-${n}-brief.md" fi @@ -39,9 +41,16 @@ if [ ! -s "$out" ]; then fi echo "wrote ${out}: $(wc -l < "$out" | tr -d ' ') lines" -# The dispatch hint rides this output because the controller reads it -# immediately before spawning the implementer; skill text loaded at session -# start does not survive context compaction, but this line reprints per task. -# Model/effort values are mirrored in using-superpowers/references/codex-tools.md -# (they track Codex's spawn_agent allowlist) — update both together. -echo "dispatch (codex spawn_agent): role=implementer fork_turns=none model=gpt-5.6-terra reasoning_effort=high" + +# Platform dispatch hints ride this output because the controller reads it +# immediately before spawning; the lines themselves are owned by the platform +# reference layer (using-superpowers/references/*-dispatch.hints), not this +# script. Claude Code's dispatch templates carry model selection already, so +# the relay is suppressed there and on any harness without a hints file. +hints_file="$script_dir/../../using-superpowers/references/codex-dispatch.hints" +if [ -z "${CLAUDECODE:-}" ] && [ -f "$hints_file" ]; then + hint_line=$(grep "^implementer:" "$hints_file" | head -1 | cut -d: -f2- | sed 's/^ *//') || true + if [ -n "$hint_line" ]; then + echo "$hint_line" + fi +fi diff --git a/skills/using-superpowers/references/codex-dispatch.hints b/skills/using-superpowers/references/codex-dispatch.hints new file mode 100644 index 0000000000..e2a6008b9d --- /dev/null +++ b/skills/using-superpowers/references/codex-dispatch.hints @@ -0,0 +1,11 @@ +# Per-role dispatch lines for Codex spawn_agent, relayed by the +# subagent-driven-development task-brief and review-package scripts at the +# moment of dispatch (skill text loaded at session start does not survive +# context compaction; these lines reprint every round). +# Model names track Codex's spawn_agent allowlist (currently gpt-5.6-sol +# and gpt-5.6-terra) — update this file when the allowlist changes. +# Format: : +implementer: dispatch (spawn_agent): fork_turns=none model=gpt-5.6-terra reasoning_effort=high +task-review: dispatch (spawn_agent): fork_turns=none model=gpt-5.6-terra reasoning_effort=high +fix-review: dispatch (spawn_agent): fork_turns=none model=gpt-5.6-terra reasoning_effort=medium +final-review: dispatch (spawn_agent): fork_turns=none model=gpt-5.6-terra reasoning_effort=high diff --git a/skills/using-superpowers/references/codex-tools.md b/skills/using-superpowers/references/codex-tools.md index 15fca4bd42..1b4d907e20 100644 --- a/skills/using-superpowers/references/codex-tools.md +++ b/skills/using-superpowers/references/codex-tools.md @@ -9,48 +9,30 @@ multi_agent = true This enables `spawn_agent`, `wait_agent`, and `close_agent` for skills like `dispatching-parallel-agents` and `subagent-driven-development`. When using subagent-driven-development, close reviewer subagents when their review returns. Keep each implementer subagent open until its task's review passes — the fix loop resumes the implementer — then close it. If your harness cannot send another message to a spawned agent, dispatch each fix round as a fresh implementer carrying the brief, the report file, and the findings. -## SDD dispatch: pin fork_turns and the model on every spawn - -Every `spawn_agent` call in subagent-driven-development sets -`fork_turns: "none"`. The parameter defaults to `"all"`, which forks your -entire session transcript into the child — the opposite of the fresh, -constructed context SDD requires — and a full-history fork also refuses -model and effort overrides. Never omit the parameter, and never pass -`"all"` or a turn count for an SDD dispatch. - -Before Task 1, check your `spawn_agent` tool schema for `model` and -`reasoning_effort` parameters (present on Codex 0.145+). - -**If the parameters exist**, set both explicitly on every dispatch. The -task-brief and review-package scripts print the exact values for each -dispatch as a `dispatch (codex spawn_agent):` line with their output — -copy that line's values onto the spawn_agent call verbatim, every time, -even late in a long session. The mapping they print: every SDD seat runs -`gpt-5.6-terra` — implementers and reviewers at `reasoning_effort: high`, -scoped re-reviews at `medium`. On Codex this mapping IS the Model -Selection section — including the final review, which stays on -terra/high rather than "most capable available." Never give a subagent -your session's model when you run a frontier config (sol at xhigh or -max): reviewer tier never exceeds implementer tier, and a fix round -never gets an effort bump. Rounds 4-5's "more capable model" means a -fresh implementer at the same tier; a task that genuinely needs more -than terra/high is a BLOCKED escalation to your human partner, not a -quiet tier climb. Inherited frontier-tier subagents are a measured cause -of SDD runs spinning out for hours: review seats that inherit a frontier -model at maximum effort find real-but-endless defects every round, and -fix diffs balloon instead of converging. - -The model names here track Codex's `spawn_agent` allowlist (currently -`gpt-5.6-sol` and `gpt-5.6-terra`). When the allowlist changes, update -this file and the hint lines in task-brief and review-package together. - -**If the parameters do not exist** (Codex 0.144 and earlier), every child -inherits your session's model and effort and no override is possible — -role files in `~/.codex/agents/` do not attach to spawns either. Say so -to your human partner before starting a plan of more than a few tasks, -and offer the choice: proceed with inheritance, or restart the session -at a lower effort so the whole run — controller and children — pays the -lower rate. +## SDD dispatch on Codex + +Every SDD `spawn_agent` call sets `fork_turns: "none"` — the default +`"all"` forks your whole transcript into the child and refuses model +and effort overrides. + +If your `spawn_agent` schema has `model` and `reasoning_effort` +parameters (Codex 0.145+), set both on every dispatch: task-brief and +review-package print a `dispatch:` hint line with the exact values — +copy it onto the call verbatim, every time, even late in a long +session. Those hints are the Model Selection mapping on Codex: +reviewer tier never exceeds implementer tier, no fix round gets an +effort bump, and rounds 4-5's "more capable model" means a fresh +implementer at the same tier — needing more is a BLOCKED escalation +to your human partner. Inherited frontier-tier subagents are a +measured cause of runs spinning out for hours. (Values live in +`codex-dispatch.hints` beside this file; they track the spawn_agent +model allowlist.) + +Without those parameters (Codex 0.144 and earlier), children inherit +your model and effort with no override — role files in +`~/.codex/agents/` do not attach to spawns either. Tell your human +partner before starting a plan of more than a few tasks, and offer a +lower-effort session instead. ## Environment Detection diff --git a/tests/claude-code/test-sdd-workspace.sh b/tests/claude-code/test-sdd-workspace.sh index 1178f340a1..0e3221406e 100755 --- a/tests/claude-code/test-sdd-workspace.sh +++ b/tests/claude-code/test-sdd-workspace.sh @@ -165,39 +165,54 @@ PLAN echo " got: $rp_explicit" fi - # --- dispatch hints ride the script output --- - if [[ "$brief_out" == *"dispatch (codex spawn_agent): role=implementer fork_turns=none model=gpt-5.6-terra reasoning_effort=high"* ]]; then - pass "task-brief prints the implementer dispatch hint" + # --- platform dispatch hints ride the script output (suppressed on CC) --- + local brief_hint + brief_hint="$(cd "$repo" && env -u CLAUDECODE "$SDD_SCRIPTS/task-brief" plan-a.md 1)" + if [[ "$brief_hint" == *"dispatch (spawn_agent): fork_turns=none model=gpt-5.6-terra reasoning_effort=high"* ]]; then + pass "task-brief relays the implementer dispatch hint off Claude Code" else - fail "task-brief prints the implementer dispatch hint" - echo " got: $brief_out" + fail "task-brief relays the implementer dispatch hint off Claude Code" + echo " got: $brief_hint" fi - if [[ "$rp_out" == *"dispatch (codex spawn_agent): role=task-reviewer fork_turns=none model=gpt-5.6-terra reasoning_effort=high"* ]]; then - pass "review-package defaults to the task-reviewer dispatch hint" + local rp_hint + rp_hint="$(cd "$repo" && env -u CLAUDECODE "$SDD_SCRIPTS/review-package" plan-a.md HEAD~1 HEAD)" + if [[ "$rp_hint" == *"dispatch (spawn_agent): fork_turns=none model=gpt-5.6-terra reasoning_effort=high"* ]]; then + pass "review-package relays the default-role hint off Claude Code" else - fail "review-package defaults to the task-reviewer dispatch hint" - echo " got: $rp_out" + fail "review-package relays the default-role hint off Claude Code" + echo " got: $rp_hint" fi local rp_rereview - rp_rereview="$(cd "$repo" && "$SDD_SCRIPTS/review-package" --role re-review plan-a.md HEAD~1 HEAD)" - if [[ "$rp_rereview" == *"role=scoped-re-reviewer fork_turns=none model=gpt-5.6-terra reasoning_effort=medium"* ]]; then - pass "review-package --role re-review prints the medium-effort hint" + rp_rereview="$(cd "$repo" && env -u CLAUDECODE "$SDD_SCRIPTS/review-package" --role re-review plan-a.md HEAD~1 HEAD)" + if [[ "$rp_rereview" == *"dispatch (spawn_agent): fork_turns=none model=gpt-5.6-terra reasoning_effort=medium"* ]]; then + pass "review-package --role re-review relays the medium-effort hint" else - fail "review-package --role re-review prints the medium-effort hint" + fail "review-package --role re-review relays the medium-effort hint" echo " got: $rp_rereview" fi local rp_final - rp_final="$(cd "$repo" && "$SDD_SCRIPTS/review-package" --role final-review plan-a.md HEAD~1 HEAD)" - if [[ "$rp_final" == *"role=final-reviewer fork_turns=none model=gpt-5.6-terra reasoning_effort=high"* ]]; then - pass "review-package --role final-review prints the final-reviewer hint" + rp_final="$(cd "$repo" && env -u CLAUDECODE "$SDD_SCRIPTS/review-package" --role final-review plan-a.md HEAD~1 HEAD)" + if [[ "$rp_final" == *"dispatch (spawn_agent): fork_turns=none model=gpt-5.6-terra reasoning_effort=high"* ]]; then + pass "review-package --role final-review relays the high-effort hint" else - fail "review-package --role final-review prints the final-reviewer hint" + fail "review-package --role final-review relays the high-effort hint" echo " got: $rp_final" fi + local brief_cc rp_cc + brief_cc="$(cd "$repo" && CLAUDECODE=1 "$SDD_SCRIPTS/task-brief" plan-a.md 1)" + rp_cc="$(cd "$repo" && CLAUDECODE=1 "$SDD_SCRIPTS/review-package" plan-a.md HEAD~1 HEAD)" + if [[ "$brief_cc" != *"dispatch (spawn_agent)"* && "$rp_cc" != *"dispatch (spawn_agent)"* ]]; then + pass "dispatch hints are suppressed under Claude Code (CLAUDECODE set)" + else + fail "dispatch hints are suppressed under Claude Code (CLAUDECODE set)" + echo " brief: $brief_cc" + echo " rp: $rp_cc" + fi + rc=0 (cd "$repo" && "$SDD_SCRIPTS/review-package" --role bogus plan-a.md HEAD~1 HEAD >/dev/null 2>&1) || rc=$? if [[ "$rc" -eq 2 ]]; then From 5c867fe6de8fbc899d6fd9b43009004939a9bede Mon Sep 17 00:00:00 2001 From: Drew Ritter Date: Wed, 22 Jul 2026 15:33:48 -0700 Subject: [PATCH 5/6] rename(sdd): re-review -> fix review, a scoped review of the fix diff MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Responds to maintainer review: "re-review" names the action badly — it reads as "review again," which is the exact failure observed in the field (fix reviewers re-running full reviews and package-wide suites against explicit scope instructions). "Fix review" names the artifact under review — the fix diff since the previous review — and makes the scope self-enforcing. Pure vocabulary substitution: re-review-prompt.md becomes fix-review-prompt.md, SKILL.md and the review-package --role value follow, and the term is introduced once as "a scoped fix review — a review of the fix diff, not a fresh review." No behavioral rules changed. Historical records (docs/superpowers/plans, specs, RELEASE-NOTES) keep the old vocabulary; other skills' generic verb usage ("no need to re-review") is untouched. Co-Authored-By: Claude Fable 5 --- skills/subagent-driven-development/SKILL.md | 41 ++++++++++--------- ...-review-prompt.md => fix-review-prompt.md} | 18 ++++---- .../scripts/review-package | 16 +++----- tests/claude-code/test-sdd-workspace.sh | 12 +++--- 4 files changed, 42 insertions(+), 45 deletions(-) rename skills/subagent-driven-development/{re-review-prompt.md => fix-review-prompt.md} (85%) diff --git a/skills/subagent-driven-development/SKILL.md b/skills/subagent-driven-development/SKILL.md index 468a71a222..350951941d 100644 --- a/skills/subagent-driven-development/SKILL.md +++ b/skills/subagent-driven-development/SKILL.md @@ -59,7 +59,7 @@ digraph process { "Finding conflicts with plan text?" [shape=diamond]; "Ask human partner which governs" [shape=box]; "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [shape=box]; - "Dispatch scoped re-review (./re-review-prompt.md)" [shape=box]; + "Dispatch scoped fix review (./fix-review-prompt.md)" [shape=box]; "All findings addressed?" [shape=diamond]; "R = 5?" [shape=diamond]; "Adjudicate each open finding" [shape=box]; @@ -72,7 +72,7 @@ digraph process { "Setup: worktree, ledger check, read plan, pre-flight review" [shape=box]; "More tasks remain?" [shape=diamond]; "Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)" [shape=box]; - "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals" [shape=box]; + "Final findings? ONE fix dispatch, one scoped fix review, adjudicate residuals" [shape=box]; "Final review clean: delete this plan's workspace" [shape=box]; "Use superpowers:finishing-a-development-branch" [shape=box style=filled fillcolor=lightgreen]; @@ -88,8 +88,8 @@ digraph process { "Finding conflicts with plan text?" -> "Ask human partner which governs" [label="yes"]; "Ask human partner which governs" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model"; "Finding conflicts with plan text?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no"]; - "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" -> "Dispatch scoped re-review (./re-review-prompt.md)"; - "Dispatch scoped re-review (./re-review-prompt.md)" -> "All findings addressed?"; + "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" -> "Dispatch scoped fix review (./fix-review-prompt.md)"; + "Dispatch scoped fix review (./fix-review-prompt.md)" -> "All findings addressed?"; "All findings addressed?" -> "Append completion to ledger, mark todo complete" [label="yes"]; "All findings addressed?" -> "R = 5?" [label="no"]; "R = 5?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no - next round"]; @@ -101,8 +101,8 @@ digraph process { "Append completion to ledger, mark todo complete" -> "More tasks remain?"; "More tasks remain?" -> "Dispatch implementer subagent (./implementer-prompt.md)" [label="yes"]; "More tasks remain?" -> "Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)" [label="no"]; - "Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)" -> "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals"; - "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals" -> "Final review clean: delete this plan's workspace"; + "Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)" -> "Final findings? ONE fix dispatch, one scoped fix review, adjudicate residuals"; + "Final findings? ONE fix dispatch, one scoped fix review, adjudicate residuals" -> "Final review clean: delete this plan's workspace"; "Final review clean: delete this plan's workspace" -> "Use superpowers:finishing-a-development-branch"; } ``` @@ -174,7 +174,7 @@ capable available model, not the session default. **Review tasks**: choose the model with the same judgment, scaled to the diff's size, complexity, and risk. A small mechanical diff does not need the -most capable model; a subtle concurrency change does. Scoped re-reviews of +most capable model; a subtle concurrency change does. Scoped fix reviews of small fix diffs take a cheap-to-mid tier. **Fix-loop escalation (rounds 4-5)**: use a model at least one tier above @@ -323,7 +323,8 @@ Before the loop starts, two routes leave it immediately: Do not dismiss the finding because the plan mandates it, and do not dispatch a fix that contradicts the plan without asking. Everything else enters the loop. A fix round is one fix dispatch plus one -scoped re-review. Five rounds maximum per task: +scoped fix review — a review of the fix diff, not a fresh review. Five +rounds maximum per task: **Rounds 1-3 — resume the original implementer.** Send it the open findings verbatim. Its context is intact: it knows the task, the code, and its own @@ -342,14 +343,14 @@ own problem — fresh eyes and a capability bump in one move. covering the amended code, appends its fix report to the same report file, and returns the short contract. Before re-dispatching the reviewer, confirm the fix report contains the covering tests, the command run, and the -output; dispatch the re-review once all three are present. Name the +output; dispatch the fix review once all three are present. Name the covering test files in the fix message — a one-line fix does not need the whole suite. -**The re-review is scoped.** Run `scripts/review-package --role re-review PLAN_FILE FIX_BASE HEAD` +**The fix review is scoped.** Run `scripts/review-package --role fix-review PLAN_FILE FIX_BASE HEAD` where FIX_BASE is the head the previous review saw, and dispatch -[re-review-prompt.md](re-review-prompt.md) with the findings list, the -brief, the report file, and the printed diff path. The re-reviewer verdicts +[fix-review-prompt.md](fix-review-prompt.md) with the findings list, the +brief, the report file, and the printed diff path. The fix reviewer verdicts each finding ADDRESSED or NOT ADDRESSED and flags new breakage in the fix diff only. New Critical/Important breakage in the fix diff joins the open findings list. Out-of-scope observations go to the ledger as deferred @@ -361,7 +362,7 @@ minors — they never extend the loop. Never fix findings yourself in the controller session — your context stays clean for coordination, and controller fixes skip review. -**The breaker.** When round 5's re-review still leaves findings open, stop +**The breaker.** When round 5's fix review still leaves findings open, stop dispatching. Adjudicate each open finding yourself — you hold the plan and the cross-task context the reviewer lacks: @@ -411,9 +412,9 @@ If the final whole-branch review returns findings, dispatch ONE fix subagent with the complete findings list — not one fixer per finding. Per-finding fixers each rebuild context and re-run suites; a real session's final-review fix wave cost more than all its tasks combined. -Then run exactly one scoped re-review of the fix wave -(`scripts/review-package --role re-review PLAN_FILE FIX_BASE HEAD` over the fix range, -[re-review-prompt.md](re-review-prompt.md)). +Then run exactly one scoped fix review of the fix wave +(`scripts/review-package --role fix-review PLAN_FILE FIX_BASE HEAD` over the fix range, +[fix-review-prompt.md](fix-review-prompt.md)). Adjudicate any residual findings as in the task loop's breaker: park with rulings, or stop on load-bearing ones. There is no second fix wave — residual load-bearing findings surface to your human partner when @@ -444,9 +445,9 @@ Use superpowers:finishing-a-development-branch. | "Close enough on spec compliance" | Reviewer found spec gaps = not done. Fix or hit the cap and adjudicate — those are the only exits. | | "I'll fix it myself, dispatching is overhead" | Controller fixes pollute your context and skip review. Resume the implementer. | | "One more round will converge" | Past the cap, rounds don't converge — the failure is structural. Adjudicate and route. | -| "The reviewer will just find something new anyway" | Scoped re-reviews verify fixes; they cannot wander. New findings on untouched code go to the ledger, not the loop. | +| "The reviewer will just find something new anyway" | Scoped fix reviews verify fixes; they cannot wander. New findings on untouched code go to the ledger, not the loop. | | "This finding is obviously wrong, I'll drop it" | You adjudicate only at the cap, and every ruling is a ledger entry. Silent discards are forbidden. | -| "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. | +| "The fix was small, skip the fix review" | Unreviewed fixes are how regressions land. Every round ends with a scoped fix review. | | "Reviews slow the loop down" | The loop without reviews is just unverified churn. Reviews are the loop's brakes and steering. | | "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Controllers without one have re-dispatched entire completed task sequences. | | "This new finding is real — one more wave" | Real findings are infinite under a strong reviewer. The completed wave is the exit; adjudicate and route. | @@ -500,8 +501,8 @@ Task reviewer: Spec ❌: Implementer: Added progress reporting, extracted PROGRESS_INTERVAL constant. Re-ran test/recovery.test.js — 10/10 passing. Fix report appended. -[Run review-package --role re-review PLAN_FILE FIX_BASE HEAD; dispatch scoped re-review] -Re-reviewer: Missing progress reporting — ADDRESSED (src/recovery.js:41). +[Run review-package --role fix-review PLAN_FILE FIX_BASE HEAD; dispatch scoped fix review] +Fix reviewer: Missing progress reporting — ADDRESSED (src/recovery.js:41). Magic number — ADDRESSED (src/recovery.js:7). New breakage: none. Verdict: all findings addressed. diff --git a/skills/subagent-driven-development/re-review-prompt.md b/skills/subagent-driven-development/fix-review-prompt.md similarity index 85% rename from skills/subagent-driven-development/re-review-prompt.md rename to skills/subagent-driven-development/fix-review-prompt.md index 18b0fb8ad5..e5c88b8470 100644 --- a/skills/subagent-driven-development/re-review-prompt.md +++ b/skills/subagent-driven-development/fix-review-prompt.md @@ -1,7 +1,7 @@ -# Scoped Re-Review Prompt Template +# Scoped Fix Review Prompt Template -Use this template when dispatching a re-review after a fix round. The -re-reviewer verifies the findings were addressed and checks the fix diff for +Use this template when dispatching a fix review after a fix round. The +fix reviewer verifies the findings were addressed and checks the fix diff for new breakage. It is not a fresh review — the full review already happened. **Purpose:** Verify each finding from the previous review was addressed, and @@ -9,11 +9,11 @@ that the fix itself broke nothing. ``` Subagent (general-purpose): - description: "Re-review Task N fix round R" + description: "Fix review Task N round R" model: [MODEL — REQUIRED: choose per SKILL.md Model Selection; an omitted model silently inherits the session's most expensive one] prompt: | - You are re-reviewing one task's fix round. A previous review produced + You are reviewing one task's fix round. A previous review produced findings; an implementer has attempted to fix them. Your job is to verdict each finding and inspect the fix diff — nothing else. @@ -47,7 +47,7 @@ Subagent (general-purpose): Your scope is the findings list and the fix diff. Verdict every finding. Inspect the fix diff for new problems the fix itself introduced. Do NOT - re-review code the fix did not touch: if you notice an issue entirely + review code the fix did not touch: if you notice an issue entirely outside the fix diff, report it under Out-of-Scope Observations — it does not block this task and does not extend the loop. A broad whole-branch review happens after all tasks are complete. @@ -93,14 +93,14 @@ Subagent (general-purpose): **Placeholders:** - `[MODEL]` — REQUIRED: reviewer model per SKILL.md Model Selection; scoped - re-reviews of small fix diffs take a cheap-to-mid tier + fix reviews of small fix diffs take a cheap-to-mid tier - `[BRIEF_FILE]` — the task brief file (same file the implementer worked from) - `[FINDINGS]` — the Critical/Important findings and spec gaps from the previous review, copied verbatim, one per bullet - `[REPORT_FILE]` — the implementer's report file (fix reports appended) - `[FIX_BASE_SHA]` — the head the previous review saw - `[HEAD_SHA]` — current commit -- `[DIFF_FILE]` — the path `scripts/review-package PLAN_FILE FIX_BASE HEAD` printed +- `[DIFF_FILE]` — the path `scripts/review-package --role fix-review PLAN_FILE FIX_BASE HEAD` printed -**Re-reviewer returns:** per-finding verdicts (ADDRESSED / NOT ADDRESSED), +**Fix reviewer returns:** per-finding verdicts (ADDRESSED / NOT ADDRESSED), new breakage in the fix diff, out-of-scope observations, and a round verdict. diff --git a/skills/subagent-driven-development/scripts/review-package b/skills/subagent-driven-development/scripts/review-package index b988d4246a..b45dc255ab 100755 --- a/skills/subagent-driven-development/scripts/review-package +++ b/skills/subagent-driven-development/scripts/review-package @@ -4,9 +4,9 @@ # call. Using the recorded per-task BASE (not HEAD~1) keeps multi-commit # tasks intact. # -# Usage: review-package [--role task-review|re-review|final-review] PLAN_FILE BASE HEAD [OUTFILE] +# Usage: review-package [--role task-review|fix-review|final-review] PLAN_FILE BASE HEAD [OUTFILE] # Default OUTFILE: /.superpowers/sdd//review-...diff -# (named per range, so a re-review after fixes gets a distinct fresh file). +# (named per range, so a fix review after fixes gets a distinct fresh file). # # The trailing dispatch hint rides this output because the controller reads it # immediately before spawning the reviewer; skill text loaded at session start @@ -17,22 +17,18 @@ script_dir=$(cd "$(dirname "$0")" && pwd) role=task-review if [ "${1:-}" = "--role" ]; then - [ $# -ge 2 ] || { echo "usage: review-package [--role task-review|re-review|final-review] PLAN_FILE BASE HEAD [OUTFILE]" >&2; exit 2; } + [ $# -ge 2 ] || { echo "usage: review-package [--role task-review|fix-review|final-review] PLAN_FILE BASE HEAD [OUTFILE]" >&2; exit 2; } role=$2 shift 2 fi case "$role" in - task-review) hint_key=task-review ;; - # re-review maps onto the fix-review hints entry; the --role value itself - # is renamed in the next commit. - re-review) hint_key=fix-review ;; - final-review) hint_key=final-review ;; - *) echo "bad --role: ${role} (task-review|re-review|final-review)" >&2; exit 2 ;; + task-review|fix-review|final-review) hint_key=$role ;; + *) echo "bad --role: ${role} (task-review|fix-review|final-review)" >&2; exit 2 ;; esac if [ $# -lt 3 ] || [ $# -gt 4 ]; then - echo "usage: review-package [--role task-review|re-review|final-review] PLAN_FILE BASE HEAD [OUTFILE]" >&2 + echo "usage: review-package [--role task-review|fix-review|final-review] PLAN_FILE BASE HEAD [OUTFILE]" >&2 exit 2 fi diff --git a/tests/claude-code/test-sdd-workspace.sh b/tests/claude-code/test-sdd-workspace.sh index 0e3221406e..406638abe6 100755 --- a/tests/claude-code/test-sdd-workspace.sh +++ b/tests/claude-code/test-sdd-workspace.sh @@ -184,13 +184,13 @@ PLAN echo " got: $rp_hint" fi - local rp_rereview - rp_rereview="$(cd "$repo" && env -u CLAUDECODE "$SDD_SCRIPTS/review-package" --role re-review plan-a.md HEAD~1 HEAD)" - if [[ "$rp_rereview" == *"dispatch (spawn_agent): fork_turns=none model=gpt-5.6-terra reasoning_effort=medium"* ]]; then - pass "review-package --role re-review relays the medium-effort hint" + local rp_fixreview + rp_fixreview="$(cd "$repo" && env -u CLAUDECODE "$SDD_SCRIPTS/review-package" --role fix-review plan-a.md HEAD~1 HEAD)" + if [[ "$rp_fixreview" == *"dispatch (spawn_agent): fork_turns=none model=gpt-5.6-terra reasoning_effort=medium"* ]]; then + pass "review-package --role fix-review relays the medium-effort hint" else - fail "review-package --role re-review relays the medium-effort hint" - echo " got: $rp_rereview" + fail "review-package --role fix-review relays the medium-effort hint" + echo " got: $rp_fixreview" fi local rp_final From cbbe9dd520d3d4db81fd359344588746845dccff Mon Sep 17 00:00:00 2001 From: Drew Ritter Date: Wed, 22 Jul 2026 15:34:14 -0700 Subject: [PATCH 6/6] feat(sdd): evidence-locked gate claims in the implementer report MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Responds to maintainer review asking for better implementer self-reflection. The report format already demands command+output for TDD evidence and fix-round covering tests, but the full-suite claim asked only for prose — and that is the claim that rotted in the field: three consecutive fix rounds shipped a "full suite passes" claim the fix reviewer found unreproducible, each time because the suite had run before later edits made the result stale. Two changes, both at the moment the report is written: every claimed gate needs its exact command and the tail of fresh output — fresh meaning after the final edit, otherwise rerun or report the gate as unverified — and a missing pasted output is itself a defect for the reviewer to flag, which puts enforcement at the consumption side the same way the dispatch hints do. Fix reports claiming a full-suite pass need a fresh run, not the pre-findings one. This is verification-before-completion's gate function relocated into the one prompt subagents actually receive. Co-Authored-By: Claude Fable 5 --- .../implementer-prompt.md | 13 ++++++++++--- 1 file changed, 10 insertions(+), 3 deletions(-) diff --git a/skills/subagent-driven-development/implementer-prompt.md b/skills/subagent-driven-development/implementer-prompt.md index fbe441e209..667db25d77 100644 --- a/skills/subagent-driven-development/implementer-prompt.md +++ b/skills/subagent-driven-development/implementer-prompt.md @@ -110,14 +110,21 @@ Subagent (general-purpose): Fix them, re-run the tests that cover the amended code, and append a fix report to your report file: what you changed, the covering tests you ran, the command, and the output. Reviewers will not re-run tests for - you — your report is the test evidence. Then reply with the same short - status contract as your first report. + you — your report is the test evidence. If your fix report claims a + full-suite pass, that claim needs a fresh run after your last edit — + a suite run from before the findings arrived no longer counts. Then + reply with the same short status contract as your first report. ## Report Format Write your full report to [REPORT_FILE]: - What you implemented (or what you attempted, if blocked) - - What you tested and test results + - What you tested, and for every gate you claim — focused tests, full + suite, lint, build — the exact command and the tail of its fresh + output. Fresh means run after your final edit: if you edited anything + since your last full-suite run, that run is stale — rerun it or + report the suite as unverified. A gate claim without pasted fresh + output is itself a defect for the reviewer to flag. - **TDD Evidence** (if TDD was required for this task): - RED: command run, relevant failing output before implementation, and why the failure was expected - GREEN: command run and relevant passing output after implementation