Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 34 additions & 0 deletions .agents/skills/execution-playbooks/references/plan-execute-v1.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
# Execution playbook: `plan-execute@1`

## Accepted brief form

This contract accepts `strike` only, and it applies when the resolved worker lane is an execution-class model: a lower-context lane that carries out a plan instead of designing one.
The brief must carry a complete `## Execution plan` section, materialized from the planning artifact - the TLC `data/<plan>/tasks/T0X.md` contract when one exists - and `bin/sq-plan-validate.sh` structurally validates it before dispatch.
The planner owns the design; the executing worker receives the plan, not the sources it was derived from.

## Required sequence

1. Accept a brief whose `## Execution plan` section is complete. The section must carry these five labels, each written as a label line followed by its list entries: `Files to touch` (at least one path-like entry, i.e. real file paths), `Ordered steps` (steps numbered, e.g. `1.`), `Acceptance criteria`, `Verification commands` (a command to run), and `Out of scope` (an explicit out-of-scope statement).
2. Execute the plan as written; do not redesign, re-scope, or re-plan the change.
3. Follow `tlc-implement` for the implementation method and write the checklist under `data/<id>/artifacts/`.
4. Stop with `blocked:` when the plan is incomplete, internally contradictory, or contradicted by the code, naming the exact gap rather than filling it by inference.
5. Report acceptance evidence and the verification command result; leave review, fixes, delivery, and merge to the existing owners.

## Required evidence

- The brief's materialized `## Execution plan` section, with every required field present and non-empty.
- A checklist under `data/<id>/artifacts/` recording the plan fields as realized: files touched as planned, steps executed, acceptance met, and the verification command run.
- A `blocked:` report naming the missing or contradictory plan element when execution cannot proceed safely.

## Exit predicate

Every planned file and step is realized or explicitly reported, acceptance criteria are met, and the planned verification command has been run with its result recorded.

## Stop conditions

Stop when the plan is incomplete or internally contradictory, when a step contradicts the current code, when an unplanned design decision would be required, or when the change would materially exceed the materialized plan.

## Ownership and anti-patterns

The planner (with the commander) owns design, and `AGENTS.md` section 7 plus `execution-playbooks` own selection; this contract only carries the materialized plan into execution.
Do not redesign the change, silently widen scope, edit or defer the plan instead of executing it, or create playbook-owned review, delivery, approval, or terminal authority.
1 change: 1 addition & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -358,6 +358,7 @@ Once ownership is settled, validate exactly once against that final head so no o

An ask-user finding returns as `needs-decision`; Squad decides only when the configured authority permits, otherwise escalates to the commander.
For a strike brief with an explicit execution playbook, run its structural validator before sending the implementation to drill or the selected delivery path; the playbook validator never replaces drill.
When the resolved worker lane is an execution-class model that must not design, select `plan-execute@1` with the plan materialized from the planning artifact and run `bin/sq-plan-validate.sh <id>` before dispatch; `bin/sq-spawn.sh` refuses an incomplete plan.
Versioned planning and evaluation playbooks are methods only, while delivery remains owned by mode and drill and merge remains commander-authorized.
`orchestrate` is rejected because Commander, XO, backlog, and supervision already own programme coordination; `autopilot-full`, `autopilot-stack`, and `autonomous-run` remain deferred until their documented measurable triggers are proven, and none adds merge, discard, or destructive authority.
`arena` means competing approaches to one problem with one selected base, `swarm` means parallel coverage of distinct slices, and `interrogate` means independent attacks on one artifact with deduplication and judgment; these are auxiliary topologies, not playbook identities, dispatch owners, or inferences from free-form `sq-tasks` `kind`.
Expand Down
1 change: 1 addition & 0 deletions bin/sq-brief.sh
Original file line number Diff line number Diff line change
Expand Up @@ -187,6 +187,7 @@ if [ "$PLAYBOOK_SET" -eq 1 ]; then
visual-parity@1) PLAYBOOK_ID=visual-parity; PLAYBOOK_VERSION=1; PLAYBOOK_SECTION=$(cat "$SQUAD_ROOT/.agents/skills/execution-playbooks/references/visual-parity-v1.md"); EXPECTED_KIND='recon|strike' ;;
multi-phase-plan@1) PLAYBOOK_ID=multi-phase-plan; PLAYBOOK_VERSION=1; PLAYBOOK_SECTION=$(cat "$SQUAD_ROOT/.agents/skills/execution-playbooks/references/multi-phase-plan-v1.md"); EXPECTED_KIND=recon ;;
"eval@1") PLAYBOOK_ID='eval'; PLAYBOOK_VERSION=1; PLAYBOOK_SECTION=$(cat "$SQUAD_ROOT/.agents/skills/execution-playbooks/references/eval-v1.md"); EXPECTED_KIND=recon ;;
plan-execute@1) PLAYBOOK_ID=plan-execute; PLAYBOOK_VERSION=1; PLAYBOOK_SECTION=$(cat "$SQUAD_ROOT/.agents/skills/execution-playbooks/references/plan-execute-v1.md"); EXPECTED_KIND=strike ;;
orchestrate@1) echo "error: orchestrate@1 disposition=reject; Commander, XO, backlog, and supervision already own programme coordination" >&2; exit 1 ;;
autopilot-full@1|autopilot-stack@1) echo "error: $PLAYBOOK disposition=defer; re-evaluate only for a real multi-PR programme where commander merge approval is the measured bottleneck and yolo cannot cover it; no auto-merge or drill bypass" >&2; exit 1 ;;
autonomous-run@1) echo "error: autonomous-run disposition=defer; re-evaluate only with a real task whose cycle the current supervision cannot conduct, plus a verifiable terminal predicate, budget, stop conditions, and duplicate prevention" >&2; exit 1 ;;
Expand Down
106 changes: 106 additions & 0 deletions bin/sq-plan-validate.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,106 @@
#!/usr/bin/env bash
# Deterministic pre-dispatch structural gate for a brief's execution plan.
# Usage: sq-plan-validate.sh <task-id>
#
# Reads data/<id>/brief.md, bounds its `## Execution plan` section, and refuses
# when a required field is absent or structurally empty. Required fields and
# shapes (label line, then its entries as a list; short fields may carry inline
# content after the label):
# - Files to touch, with at least one path-like entry
# - Ordered steps, with at least one numbered step
# - Acceptance criteria
# - Verification command(s)
# - Out of scope
# The check is structural only: it never judges plan quality, correctness, or
# feasibility. It is the dispatch-time companion to bin/sq-playbook-validate.sh,
# which owns the post-work checklist evidence for the plan-execute@1 playbook.
set -eu
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
BASE="${SQUAD_BASE:-${SQUAD_HOME:-$ROOT}}"
DATA="${SQUAD_DATA_OVERRIDE:-$BASE/data}"
ID=${1:-}
[ -n "$ID" ] || { echo "error: task id is required" >&2; exit 2; }
BRIEF="$DATA/$ID/brief.md"
[ -f "$BRIEF" ] || { echo "missing: $BRIEF"; exit 1; }

awk -v id="$ID" '
function label_text(l) {
sub(/^[[:space:]]+/, "", l)
sub(/^[*_]+/, "", l)
sub(/[*_]+[[:space:]]*$/, "", l)
return l
}
function field(l) {
if (l ~ /^[Ff]iles[[:space:]]+to[[:space:]]+touch/) return "files"
if (l ~ /^([Oo]rdered[[:space:]]+)?[Ss]teps/) return "steps"
if (l ~ /^[Aa]cceptance[[:space:]]+criteria/) return "acceptance"
if (l ~ /^[Vv]erification[[:space:]]+commands?/) return "verification"
if (l ~ /^[Oo]ut[[:space:]]+of[[:space:]]+scope/) return "out-of-scope"
return ""
}
function flush() {
if (cur != "") { seen[cur] = 1; body[cur] = body[cur] "\n" text }
}
BEGIN { inplan = 0; heading = 0; cur = ""; text = "" }
/^##[[:space:]]+[Ee]xecution[[:space:]]+[Pp]lan[[:space:]]*$/ { inplan = 1; heading = 1; next }
inplan && /^#[[:space:]]/ { inplan = 0; next }
inplan && /^##[[:space:]]/ { inplan = 0; next }
!inplan { next }
{
l = label_text($0)
f = field(l)
if (f != "") {
flush()
cur = f; text = ""; inline[cur] = ""
if (l ~ /:/) {
sub(/^[^:]*:/, "", l)
gsub(/^[[:space:]]+|[[:space:]]+$/, "", l)
inline[cur] = l
}
next
}
if (cur != "") {
text = text "\n" $0
item = $0
sub(/^[[:space:]]+/, "", item)
if (item ~ /^[-*+][[:space:]]/ || item ~ /^[0-9]+[.)]/) has_item[cur] = 1
if (item ~ /^[0-9]+[.)]/) has_step[cur] = 1
}
}
END {
flush()
if (!heading) { print "missing: execution plan section (## Execution plan)"; exit 1 }
labels["files"] = "files to touch (exact paths)"
labels["steps"] = "ordered steps"
labels["acceptance"] = "acceptance criteria"
labels["verification"] = "verification command"
labels["out-of-scope"] = "out of scope"
split("files steps acceptance verification out-of-scope", keys, " ")
bad = 0
for (i = 1; i <= 5; i++) {
k = keys[i]
if (!(k in seen)) {
printf "missing: execution plan field \"%s\"\n", labels[k]
bad = 1
continue
}
combined = inline[k] "\n" body[k]
if (inline[k] == "" && has_item[k] != 1) {
printf "empty: execution plan field \"%s\"\n", labels[k]
bad = 1
continue
}
if (k == "files" && combined !~ /\/|[A-Za-z0-9_-]+\.[A-Za-z0-9]+/) {
printf "invalid: execution plan field \"%s\" has no path-like entry\n", labels[k]
bad = 1
}
if (k == "steps" && has_step[k] != 1 && inline[k] !~ /^[0-9]+[.)]/) {
printf "invalid: execution plan field \"%s\" has no numbered step\n", labels[k]
bad = 1
}
}
if (bad) exit 1
printf "execution plan structurally valid: %s\n", id
}
' "$BRIEF"
12 changes: 11 additions & 1 deletion bin/sq-playbook-validate.sh
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ ARTIFACTS="$DATA/$ID/artifacts"
PLAYBOOK=$(sed -n 's/^playbook=//p' "$META" | head -n 1)
VERSION=$(sed -n 's/^playbook_version=//p' "$META" | head -n 1)
case "$PLAYBOOK@$VERSION" in
bug-fix@1|investigation@1|feature@1|refactoring@1|prototype@1|perf@1|hillclimb@1|runtime-forensics@1|trace-forensics@1|visual-parity@1|multi-phase-plan@1|eval@1) ;;
bug-fix@1|investigation@1|feature@1|refactoring@1|prototype@1|perf@1|hillclimb@1|runtime-forensics@1|trace-forensics@1|visual-parity@1|multi-phase-plan@1|eval@1|plan-execute@1) ;;
*) echo "error: task $ID has unsupported execution playbook identity"; exit 1 ;;
esac
CHECKLIST=
Expand Down Expand Up @@ -172,6 +172,16 @@ case "$PLAYBOOK@$VERSION" in
"recommendation;promote|reject;production|promotion"
)
;;
plan-execute@1)
CRITERIA_LABELS=("files touched as planned" "steps executed" "acceptance met" "verification run")
CRITERIA_PATTERNS=("files" "steps" "acceptance" "verification")
CRITERIA_CHECKS=(
"files|paths;planned|materialized"
"steps;executed|followed"
"acceptance;met|verified|criteria"
"verification;run|executed|result"
)
;;
esac
COUNT=${#CRITERIA_LABELS[@]}
# Collect all proofs first for duplicate detection.
Expand Down
10 changes: 10 additions & 0 deletions bin/sq-spawn.sh
Original file line number Diff line number Diff line change
Expand Up @@ -1452,11 +1452,21 @@ if [ "$PLAYBOOK_LINES" -eq 1 ]; then
'Execution playbook: id=feature version=1') PLAYBOOK_META=feature; PLAYBOOK_VERSION_META=1; EXPECTED_KIND=strike ;;
'Execution playbook: id=refactoring version=1') PLAYBOOK_META=refactoring; PLAYBOOK_VERSION_META=1; EXPECTED_KIND=strike ;;
'Execution playbook: id=prototype version=1') PLAYBOOK_META=prototype; PLAYBOOK_VERSION_META=1; EXPECTED_KIND=recon ;;
'Execution playbook: id=plan-execute version=1') PLAYBOOK_META=plan-execute; PLAYBOOK_VERSION_META=1; EXPECTED_KIND=strike ;;
*) echo "error: malformed or unsupported execution playbook identity in $BRIEF" >&2; exit 1 ;;
esac
[ "$KIND" = "$EXPECTED_KIND" ] || { echo "error: $PLAYBOOK_META@1 is compatible only with kind: $EXPECTED_KIND" >&2; exit 1; }
# shellcheck disable=SC2016 # Backticks are literal brief syntax.
grep -q "^# Execution playbook: \`$PLAYBOOK_META@$PLAYBOOK_VERSION_META\`$" "$BRIEF" || { echo "error: $PLAYBOOK_META@$PLAYBOOK_VERSION_META identity has no materialized contract" >&2; exit 1; }
# A plan-execute brief carries the design as a materialized plan: the structural
# gate refuses dispatch before any endpoint exists when that plan is incomplete.
if [ "$PLAYBOOK_META" = plan-execute ]; then
if ! plan_validate_out=$(SQUAD_BASE="$SQUAD_BASE" "$SCRIPT_DIR/sq-plan-validate.sh" "$ID" 2>&1); then
printf '%s\n' "$plan_validate_out" >&2
echo "error: $ID brief execution plan failed structural validation; materialize a complete plan before dispatch" >&2
exit 1
fi
fi
fi

delivery_rigor_rank() { # <mode> -> 3 (most rigor) .. 1 (least); 0 = not a task mode
Expand Down
4 changes: 4 additions & 0 deletions docs/documentation-audiences.json
Original file line number Diff line number Diff line change
Expand Up @@ -1564,6 +1564,10 @@
"path": ".agents/skills/execution-playbooks/references/eval-v1.md",
"audience": "agent-runtime"
},
{
"path": ".agents/skills/execution-playbooks/references/plan-execute-v1.md",
"audience": "agent-runtime"
},
{
"path": "skills/interview-me/SKILL.md",
"audience": "public-product"
Expand Down
2 changes: 1 addition & 1 deletion docs/pt-BR/scripts.md
Original file line number Diff line number Diff line change
Expand Up @@ -127,7 +127,7 @@ A recusa compartilhada do gate drill para entrypoints do ciclo de vida da unidad
| `sq-pr-check-migrate.sh` | Quarentenair polls de tarefa antigos sem execução e reconstruir apenas polls canônicos |
| `sq-pr-check.sh` | Registrar valores validados de `pr=` e `pr_head=`, então armar atomicamente um poll estático de merge |
| `sq-pr-merge.sh` | Registrar metadados do PR, então mesclar a URL completa canônica GitHub do PR da tarefa |
| `sq-playbook-validate.sh` | Validar evidência estrutural de um playbook de execução materializado (bug-fix@1, investigation@1, feature@1, refactoring@1, prototype@1, perf@1, hillclimb@1, runtime-forensics@1, trace-forensics@1, visual-parity@1, multi-phase-plan@1, eval@1) |
| `sq-playbook-validate.sh` | Validar evidência estrutural de um playbook de execução materializado (bug-fix@1, investigation@1, feature@1, refactoring@1, prototype@1, perf@1, hillclimb@1, runtime-forensics@1, trace-forensics@1, visual-parity@1, multi-phase-plan@1, eval@1, plan-execute@1) |
| `sq-promote.sh` | Promover uma tarefa recon in-place a uma tarefa strike protegida com modo de entrega explícito |
| `sq-teardown.sh` | Teardown fail-closed: devolver worktrees ship landadas, exigir entregáveis completos de recon, aposentar bases XO |
| `sq-harness.sh` | Detectar o harness em execução e resolver crew ou XO harness, modelo e esforço |
Expand Down
6 changes: 3 additions & 3 deletions docs/verification/playbook-absorptions.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ On 2026-09-19, the focused `sq-brief` test exercised `--playbook <name>@1` for e
sq-brief.sh: absorbed playbook names are refused and leave no selectable brief
```

The test is [`test_absorbed_playbook_names_are_not_selectable`](../../tests/sq-brief.test.sh), and the registry contains wave-1 contracts `bug-fix@1`, `feature@1`, `investigation@1`, `prototype@1`, `refactoring@1` plus wave-2 contracts `perf@1`, `hillclimb@1`, `runtime-forensics@1`, `trace-forensics@1`, and `visual-parity@1` plus planning/evaluation methods `multi-phase-plan@1` and `eval@1` under `.agents/skills/execution-playbooks/references/`.
The test is [`test_absorbed_playbook_names_are_not_selectable`](../../tests/sq-brief.test.sh), and the registry contains wave-1 contracts `bug-fix@1`, `feature@1`, `investigation@1`, `prototype@1`, `refactoring@1` plus wave-2 contracts `perf@1`, `hillclimb@1`, `runtime-forensics@1`, `trace-forensics@1`, and `visual-parity@1` plus planning/evaluation methods `multi-phase-plan@1` and `eval@1` plus the materialized-plan execution contract `plan-execute@1` under `.agents/skills/execution-playbooks/references/`.

The focused test command is:

Expand All @@ -34,7 +34,7 @@ The registry inventory command is:
find .agents/skills/execution-playbooks -maxdepth 2 -type f -print | sort
```

Its expected output contains `SKILL.md` and the reference files for all twelve selectable execution playbook contracts.
Its expected output contains `SKILL.md` and the reference files for all thirteen selectable execution playbook contracts.

## Deferred and auxiliary policy

Expand All @@ -49,4 +49,4 @@ The three auxiliary topologies are not execution playbook identities or dispatch
`arena` competes on one problem and selects one base, `swarm` covers distinct slices, and `interrogate` independently attacks one artifact with deduplication and judgment.
For PR or diff surfaces, `interrogate` points to the read-only Drill surface documented in [`docs/pr-review.md`](../pr-review.md), and separate reviews are limited to requested or knowledge-only review deliverables.

The catalog remains 22 playbooks because these topology names and the four refused or deferred names are not selectable identities.
The catalog now counts 23 playbook identities: 13 selectable contracts, 5 absorbed upstream names, and 5 refused or deferred names.
15 changes: 14 additions & 1 deletion tests/sq-brief.test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -405,14 +405,27 @@ test_remaining_lifecycle_playbooks() {
assert_grep "not a layer-by-layer" "$brief" "multi-phase-plan contract missing layer boundary"
assert_grep "existing backlog" "$brief" "multi-phase-plan contract missing backlog handoff"

SQUAD_BASE="$home" "$ROOT/bin/sq-brief.sh" lifecycle-plan-execute repo --mode drill --playbook plan-execute@1 >/dev/null 2>&1 || fail "plan-execute@1 should materialize as strike"
brief="$home/data/lifecycle-plan-execute/brief.md"
assert_grep "Execution playbook: id=plan-execute version=1" "$brief" "plan-execute identity missing"
assert_grep "execution-class" "$brief" "plan-execute contract missing execution-class scope"
assert_grep "materialized from the planning artifact" "$brief" "plan-execute contract missing the plan materialization rule"
assert_grep "tlc-implement" "$brief" "plan-execute contract missing tlc-implement"
assert_grep "data/<id>/artifacts/" "$brief" "plan-execute contract missing the checklist location"
assert_grep "blocked:" "$brief" "plan-execute contract missing the stop-with-blocked rule"
out=$(SQUAD_BASE="$home" "$ROOT/bin/sq-brief.sh" lifecycle-plan-execute-recon repo --recon --playbook plan-execute@1 2>&1); status=$?
[ "$status" -ne 0 ] || fail "plan-execute@1 must refuse recon"
assert_contains "$out" "accepts only kind: strike" "plan-execute refusal should name strike compatibility"
assert_absent "$home/data/lifecycle-plan-execute-recon/brief.md" "plan-execute recon refusal must not leave a partial brief"

SQUAD_BASE="$home" "$ROOT/bin/sq-brief.sh" lifecycle-eval repo --recon --playbook eval@1 >/dev/null 2>&1 || fail "eval@1 should materialize"
brief="$home/data/lifecycle-eval/brief.md"
assert_grep "Execution playbook: id=eval version=1" "$brief" "eval identity missing"
assert_grep "candidate-visible" "$brief" "eval contract missing candidate blinding"
assert_grep "chain-elicitation" "$brief" "eval contract missing chain-elicitation prevention"
assert_grep "does not enter production" "$brief" "eval contract missing promotion boundary"

for playbook in shipping@2 multi-phase-plan@2 eval@2; do
for playbook in shipping@2 multi-phase-plan@2 eval@2 plan-execute@2; do
out=$(SQUAD_BASE="$home" "$ROOT/bin/sq-brief.sh" "lifecycle-invalid-$RANDOM" repo --mode drill --playbook "$playbook" 2>&1); status=$?
[ "$status" -ne 0 ] || fail "$playbook should be refused"
assert_contains "$out" "unknown execution playbook" "$playbook should identify invalid identity"
Expand Down
Loading
Loading