Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .github/workflows/test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -5,8 +5,8 @@ name: test
# hosted minutes are free, so master pushes and ready PRs run the suites again.
#
# A DRAFT pull request skips the suites (see the job-level `if` below): iterate
# locally with `bash tests/run-all.sh`, which is the same entry CI runs, and mark
# the PR ready when it is done. `ready_for_review` is listed because it is not one
# locally with `bash tests/run-all.sh --changed` (only the suites the diff touches; the
# full glob is the entry CI runs), and mark the PR ready when it is done. `ready_for_review` is listed because it is not one
# of the default pull_request types, so without it a draft marked ready would never
# trigger a first run. Master push still runs unconditionally, so the 2026-08-10
# silent-drift failure mode stays closed.
Expand Down
1 change: 1 addition & 0 deletions docs/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@ All notable changes to dwarves-kit are documented here.
- Config surface (new section, additive, MINOR): `[intake]` with `url_ledger`, `verdicts`, `boards`, `notes`, all defaulting `""`. Every key resolves root-only, so a project `.kit.toml` cannot set them. An install that sets none of them keeps `intake gate` answering from this kit's own inventory and open pull requests alone.

### Added
- `tests/run-all.sh --changed [<base>]` runs only the suites the diff against `<base>` (default: the merge-base with `origin/master`) touches: suites whose code lines name a changed file's basename, changed suites themselves, `tests/test-<mod>*.sh` for `lib/<mod>/`, plus every suite with an `# always:` header (the tree-wide lints: kit-contract, config-registry, no-personal-paths, no-scattered-ids, boundary-lint, meta). The local pre-push check; CI keeps the full glob. `RUN_ALL_JOBS` now defaults to `auto` on macOS, where every parallel run has been green, and stays `1` on Linux until the ubuntu flake is understood.
- `session observe entry-fee [--days N] [--project SLUG-OR-NAME] [--top N] [--trend] [--json]`
sizes the fixed preamble every agent turn re-reads before any work. The per-session
total is measured (the first main-chain assistant turn's input + cache-creation +
Expand Down
2 changes: 1 addition & 1 deletion docs/FEATURES.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ GENERATED , do not hand-edit. Regenerate: `bash lib/registry/feature-registry.sh
| `/kit:design` | `[H/I]` | Opt-in interactive solution-design beat between /think and /spec. Explores 2-3 approaches one question at a time, holds for your approval p… | SPEC-003, SPEC-004, SPEC-005 +105 | test-command-emit-sweep.sh, test-command-triggers.sh, test-design-record.sh +21 |
| `/kit:devs-team` | `[H/I]` | Parallel multi-lens critique of a solution design (the active spec if present, else the decision brief). Dispatches 5 engineering lenses, m… | SPEC-016, SPEC-018, SPEC-019 +11 | test-gate-vocab-recording.sh, test-meta.sh, test-outcome-emit-sweep.sh |
| `/kit:dispatch` | `[H/I]` | Fire several disjoint VALIDATED specs concurrently, each in its own worktree, then converge. Cross-goal fan-out behind a disjointness gate … | SPEC-002, SPEC-016, SPEC-017 +80 | test-advisor-ledger-emit.sh, test-agent-effectiveness.sh, test-attempt-state.sh +31 |
| `/kit:docs` | `[H/I]` | Update all project documentation to match the current codebase. Cross-references the diff against every doc file and fixes drift. | SPEC-001, SPEC-002, SPEC-003 +194 | proof-loop-09-scenario-b.sh, run-all.sh, run-workflow.sh +70 |
| `/kit:docs` | `[H/I]` | Update all project documentation to match the current codebase. Cross-references the diff against every doc file and fixes drift. | SPEC-001, SPEC-002, SPEC-003 +194 | proof-loop-09-scenario-b.sh, run-all.sh, run-workflow.sh +71 |
| `/kit:draft-agent` | `[H/I]` | Meta-agent agent-builder. From a one-line description, generates a new subagent definition OR a mega-goal sub-goal file and (by default) in… | SPEC-089, SPEC-108, SPEC-139 +1 | test-agent-effectiveness.sh, test-command-emit-sweep.sh, test-meta-agent.sh +1 |
| `/kit:execute` | `[H/I]` | Autonomous spec execution with verification. Dispatches worker subagents per task, verifies each with task-verifier, retries fixable failur… | SPEC-001, SPEC-003, SPEC-004 +60 | test-break-it.sh, test-gate-vocab-recording.sh, test-hooks.sh +9 |
| `/kit:explain` | `[H/I]` | Turn a merged change into a literate-diff explainer a human READS to understand: background -> goal + intuition -> a prose-ordered diff -> … | SPEC-050, SPEC-060, SPEC-094 +18 | proof-loop-09-scenario-b.sh, test-boundary-lint.sh, test-command-emit-sweep.sh +11 |
Expand Down
48 changes: 48 additions & 0 deletions docs/verification/run-all-changed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
# Verification -- run-all-changed

`tests/run-all.sh --changed [<base>]` runs only the suites the diff touches, and `RUN_ALL_JOBS` defaults to `auto` on macOS.

The full glob costs 13-15 minutes sequential on a Mac, and every in-session worker was paying that as its pre-push check (two sessions were sitting on it side by side when this was filed). Selection: a suite whose code lines name a changed file's basename, a changed suite itself, `tests/test-<mod>*.sh` for a changed `lib/<mod>/` file, plus every suite carrying an `# always:` header. Six suites carry it: kit-contract, config-registry, no-personal-paths, no-scattered-ids, boundary-lint, meta. They are the tree-wide lints that fail on a file you ADDED while naming no file you touched, so no diff-derived pick can reach them.

## Green run

```
Command: bash tests/test-run-all-changed.sh
Exit: 0
Verdict: PASS (7/7: named pick, always-on pick with the unnamed file listed, changed suite picks itself, lib/<mod> pick, empty diff runs everything, explicit base, a red always-on lint fails the run)
```

The real primary flow, end to end, on this branch at c0d2918 on the Air (M4, 10 cores, `RUN_ALL_JOBS` unset so the macOS default applied):

```
Command: bash tests/run-all.sh --changed
Output: run-all: --changed against 2ca7966: 11 changed files -> 11 suites (11 named, the rest always-on)
run-all: 11 suites, 4 at a time, 0 serial
run-all: all 11 suites passed, 0 skipped for missing tooling
Exit: 0
Wall clock: 124s
Verdict: PASS
```

The always-on floor, each lint alone on the same machine: kit-contract 4s, config-registry 47s, no-personal-paths 11s, no-scattered-ids 4s, boundary-lint 6s; test-meta about 140s and the bound of any `--changed` run.

## Negative control

```
Command: git checkout origin/master -- tests/run-all.sh && bash tests/test-run-all-changed.sh; git checkout HEAD -- tests/run-all.sh
Exit: 1
Output: test-run-all-changed: 1 passed, 6 FAILED
Verdict: RED as expected, then restored (git status clean)
```

With the master runner, `--changed` is an unknown argument, so every fixture run executes the whole glob: the six cases that assert a suite was NOT run go red, and only the empty-diff case (which expects the whole glob) stays green.

## The first dogfood run, and what it changed

The first cut matched comments too and ran 7 suites in 420s: `test-orchestrate-wavefront` was picked because its header comment mentions `run-all.sh`, and it hit the 300s ceiling under the load of two other sessions' full runs. Matching now reads code lines only; the second run above did not pick it. The same first run also went red on `test-meta`'s registry freshness pin, because the new test file changed what the registry scans, which is the exact failure class the `# always:` pin exists to surface locally rather than in CI.

## Not proven

- The ubuntu-only parallel flake (#647) is untouched. Linux keeps `RUN_ALL_JOBS=1` by default; macOS gets `auto` on the evidence that every parallel run there has been green, and this branch's two parallel runs add to it.
- CI still runs the full glob; `--changed` is a local check and nothing in the workflow calls it.
- A suite that reaches a changed file only through a helper it sources (no basename in its own code lines) is not picked. None found by hand; the always-on lints do not depend on it.
71 changes: 66 additions & 5 deletions tests/run-all.sh
Original file line number Diff line number Diff line change
Expand Up @@ -10,8 +10,10 @@
# is sub-second, and the wall clock was the sum of 146 of them. Output stays deterministic
# because results are collated in glob order after the run, not as each suite finishes.
#
# Usage: bash tests/run-all.sh [--only <pattern>]
# Env: RUN_ALL_JOBS=<n> parallel suites (default: core count; 1 = sequential)
# Usage: bash tests/run-all.sh [--only <pattern> | --changed [<base>]]
# --changed runs only the suites the diff against <base> touches (default base: the
# merge-base with origin/master). The local pre-push check; CI keeps the full glob.
# Env: RUN_ALL_JOBS=<n> parallel suites (default: auto on macOS, 1 elsewhere)
# RUN_ALL_TIMEOUT_SECS=<n> per-suite ceiling (default: 300)
# Exit: 0 all green, 1 one or more failed (every failure is listed at the end).

Expand Down Expand Up @@ -54,8 +56,10 @@ _cores() {
else echo 2
fi
}
# DEFAULT IS 1: parallel execution is opt-in via RUN_ALL_JOBS until the ubuntu
# flake below is understood.
# DEFAULT IS 1 ON LINUX: parallel execution is opt-in via RUN_ALL_JOBS there until
# the ubuntu flake below is understood. macOS defaults to auto: every macOS run of
# the parallel batch, CI and local, has been green, and the flake has never shown
# on a Mac, so the platform where the evidence holds gets the 2.3x.
#
# Parallel runs are 2.3x faster and were green on a full dispatched matrix, on
# every macOS run, and on the PR runs. They then failed twice on master, on
Expand All @@ -78,7 +82,10 @@ _cores() {
# cap is 4, not the core count: the heaviest suites spawn 100-800 subprocesses
# each, so one worker per core oversubscribes the box several times over. At -P
# 10 on a 10-core M4, three suites blew past the 300s ceiling.
JOBS="${RUN_ALL_JOBS:-1}"
JOBS="${RUN_ALL_JOBS:-}"
if [ -z "$JOBS" ]; then
case "$(uname -s)" in Darwin) JOBS=auto ;; *) JOBS=1 ;; esac
fi
if [ "$JOBS" = "auto" ]; then
JOBS="$(_cores)"
[ "$JOBS" -gt 4 ] && JOBS=4
Expand All @@ -87,6 +94,59 @@ fi
OUTDIR="$(mktemp -d)"
trap 'rm -rf "$OUTDIR"' EXIT

# --- --changed: pick suites by the diff ---------------------------------------
# The full glob is 13-15 minutes sequential on a Mac, and a branch that touches one lib
# file needs a handful of those suites. Selection is a text match, on purpose: a suite is
# picked when a CODE line of it names the basename of a changed file (a comment naming a
# file is documentation, not a dependency: the wavefront suite's header mentions this
# runner and cost a 300s timeout on the first dogfood run), when it is itself changed,
# or when it is tests/test-<mod>*.sh for a changed lib/<mod>/ file.
#
# A suite may declare that it must run on every diff:
# # always: lints every KIT_* env read in the tree against the module registry
# Those are the tree-wide lints (naming contract, config registry, personal paths,
# scattered ids, engine boundary, the registry pin). They fail on a file you ADDED while
# naming no file you touched, so no diff-derived pick can reach them.
# Over-picking is fine; a suite this misses is one the changed file never appears in.
# ponytail: basename grep, no dependency graph; add one if over-picking starts to cost.
PICKED=""
if [ "${1:-}" = "--changed" ]; then
base="${2:-}"
if [ -z "$base" ]; then
base="$(git merge-base HEAD origin/master 2>/dev/null \
|| git merge-base HEAD master 2>/dev/null \
|| echo HEAD)"
fi
changedlist="$OUTDIR/changed"
{ git diff --name-only "$base" -- . 2>/dev/null; git ls-files --others --exclude-standard; } \
| grep -v '^$' | sort -u >"$changedlist"
if [ ! -s "$changedlist" ]; then
echo "run-all: --changed found no diff against $(git rev-parse --short "$base"); running everything"
else
PICKED="$OUTDIR/picked"; : >"$PICKED"
while IFS= read -r f; do
case "$f" in
tests/test-*.sh) [ -f "$f" ] && printf '%s\n' "$f" >>"$PICKED" ;;
lib/*/*) mod="${f#lib/}"; mod="${mod%%/*}"; ls tests/test-"$mod"*.sh >>"$PICKED" 2>/dev/null ;;
esac
b="$(basename "$f")"
for s in tests/test-*.sh; do
grep -v '^[[:space:]]*#' "$s" | grep -qF -- "$b" && printf '%s\n' "$s" >>"$PICKED"
done
done <"$changedlist"
named=$(sort -u "$PICKED" | wc -l | tr -d ' ')
grep -l '^# always:' tests/test-*.sh >>"$PICKED" 2>/dev/null
sort -u -o "$PICKED" "$PICKED"
echo "run-all: --changed against $(git rev-parse --short "$base"): $(wc -l <"$changedlist" | tr -d ' ') changed files -> $(wc -l <"$PICKED" | tr -d ' ') suites ($named named, the rest always-on)"
sed 's/^/ /' "$PICKED"
if [ "$named" -eq 0 ]; then
echo "run-all: no suite names any of the changed files; only the always-on lints run"
sed 's/^/ /' "$changedlist"
fi
[ -s "$PICKED" ] || exit 0
fi
fi

# --- phase 1: decide what runs, sequentially and cheaply ---------------------
# A suite may declare external tooling it cannot run without:
# # requires: claude codex
Expand All @@ -103,6 +163,7 @@ for t in tests/test-*.sh; do
[ -f "$t" ] || continue
name="$(basename "$t" .sh)"
[ -n "$ONLY" ] && case "$name" in *"$ONLY"*) : ;; *) continue ;; esac
[ -n "$PICKED" ] && ! grep -qxF -- "$t" "$PICKED" && continue
reqs="$(sed -n 's/^# requires:[[:space:]]*//p' "$t" | head -1)"
missing=""
for r in $reqs; do command -v "$r" >/dev/null 2>&1 || missing="$missing $r"; done
Expand Down
1 change: 1 addition & 0 deletions tests/test-boundary-lint.sh
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
#!/usr/bin/env bash
# always: the engine-names-no-consumer lint scans the whole tree
# test-boundary-lint.sh -- SG-01 (learning-boundary, SPEC-285): the engine names no
# consumer, by path or by skill.
#
Expand Down
1 change: 1 addition & 0 deletions tests/test-config-registry.sh
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
#!/usr/bin/env bash
# always: lints every KIT_* env read in the tree against the module registry
# test-config-registry.sh -- SPEC-198, harness-loop sub-goal 08.
#
# Two standing lints plus a functional smoke of bin/config:
Expand Down
1 change: 1 addition & 0 deletions tests/test-kit-contract.sh
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
#!/usr/bin/env bash
# always: the naming contract scans every module in the tree
# test-kit-contract.sh -- SPEC-200: the standing contract EVERY kit module must satisfy.
#
# The rules in docs/kit-contract.md, executable. Not style policing: each rule below exists
Expand Down
1 change: 1 addition & 0 deletions tests/test-meta.sh
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
#!/bin/bash
# always: the registry pin and structural lints cover every kit artifact
# test-meta.sh -- Structural integrity tests for kit artifacts.
# Catches drift the unit tests can't see: version mismatches, missing
# frontmatter, stale references between files, schema violations.
Expand Down
1 change: 1 addition & 0 deletions tests/test-no-personal-paths.sh
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
#!/usr/bin/env bash
# always: scans the whole tree for operator paths and hostnames
# test-no-personal-paths.sh -- the kit ships no operator-specific path or hostname,
# and adopt renders none into a consumer repo.
#
Expand Down
1 change: 1 addition & 0 deletions tests/test-no-scattered-ids.sh
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
#!/usr/bin/env bash
# always: scans the whole tree for scattered spec and task ids
# test-no-scattered-ids.sh -- the provenance rule, enforced where the repo is already clean.
#
# THE RULE (CONTRIBUTING.md "Where an ID may appear"): a spec, task, ADR or ticket id belongs
Expand Down
92 changes: 92 additions & 0 deletions tests/test-run-all-changed.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,92 @@
#!/usr/bin/env bash
# run-all.sh --changed must run only the suites the diff touches.
#
# The full glob is 13-15 minutes sequential on a Mac. A branch that touches one lib file
# needs the handful of suites that name it, and the pre-push check was paying for all of
# them. Selection: a suite whose CODE lines name a changed file's basename, a changed suite
# itself, tests/test-<mod>*.sh for lib/<mod>/, plus every suite with an `# always:` header
# (the tree-wide lints, which a diff-derived pick can never reach).
#
# Each case builds a throwaway kit-shaped git repo (tests/run-all.sh plus fixture suites)
# and runs the REAL script against it, so nothing here touches the repo's own tests/.
set -uo pipefail
DIR="$(cd "$(dirname "$0")/.." && pwd)"
RA="$DIR/tests/run-all.sh"
pass=0; fail=0
ok(){ echo " ok: $*"; pass=$((pass+1)); }
no(){ echo " FAIL: $*" >&2; fail=$((fail+1)); }

TMP="$(mktemp -d)"
trap 'rm -rf "$TMP"' EXIT
export GIT_CONFIG_GLOBAL=/dev/null GIT_CONFIG_SYSTEM=/dev/null
g() { git -c user.name=t -c user.email=t@t "$@"; }

mkkit() { # $1 = dir ; a committed kit-shaped repo holding the real run-all.sh
mkdir -p "$1/tests" "$1/lib/foo" "$1/lib/baz" "$1/docs"
cp "$RA" "$1/tests/run-all.sh"
printf '#!/usr/bin/env bash\nf=lib/foo/foo.sh\nexit 0\n' > "$1/tests/test-foo.sh"
printf '#!/usr/bin/env bash\n# a comment naming foo.sh is not a dependency\nexit 0\n' > "$1/tests/test-bar.sh"
printf '#!/usr/bin/env bash\nexit 0\n' > "$1/tests/test-baz.sh"
printf '#!/usr/bin/env bash\n# always: fixture tree-wide lint\nexit 0\n' > "$1/tests/test-lint.sh"
echo 'x=1' > "$1/lib/foo/foo.sh"
echo 'y=1' > "$1/lib/baz/inner.sh"
echo '# doc' > "$1/docs/x.md"
( cd "$1" && g init -q -b master && g add -A && g commit -qm init )
}
ran() { grep -E "^$1 +ok" <<<"$2" >/dev/null; } # did suite $1 run (and pass) in output $2

echo "[1] an uncommitted lib change picks the suite whose code names the file, plus the always-on lint"
K="$TMP/k1"; mkkit "$K"; echo 'x=2' > "$K/lib/foo/foo.sh"
OUT="$(bash "$K/tests/run-all.sh" --changed 2>&1)"; RC=$?
if [ "$RC" -eq 0 ] && ran test-foo "$OUT" && ran test-lint "$OUT" && ! ran test-bar "$OUT" && ! ran test-baz "$OUT" \
&& grep -q '1 changed files -> 2 suites (1 named' <<<"$OUT"; then
ok "test-foo + test-lint; the comment-only mention in test-bar does not pick it"
else no "rc=$RC out=$OUT"; fi

echo "[2] a docs change no suite names runs only the always-on lint and says so"
K="$TMP/k2"; mkkit "$K"; echo '# more' >> "$K/docs/x.md"
OUT="$(bash "$K/tests/run-all.sh" --changed 2>&1)"; RC=$?
if [ "$RC" -eq 0 ] && ran test-lint "$OUT" && ! ran test-foo "$OUT" && ! ran test-bar "$OUT" \
&& grep -q 'no suite names any of the changed files' <<<"$OUT" && grep -q 'docs/x.md' <<<"$OUT"; then
ok "test-lint alone, the unnamed file listed"
else no "rc=$RC out=$OUT"; fi

echo "[3] a changed suite picks itself"
K="$TMP/k3"; mkkit "$K"; echo '# touched' >> "$K/tests/test-bar.sh"
OUT="$(bash "$K/tests/run-all.sh" --changed 2>&1)"; RC=$?
if [ "$RC" -eq 0 ] && ran test-bar "$OUT" && ! ran test-foo "$OUT"; then
ok "test-bar picked"
else no "rc=$RC out=$OUT"; fi

echo "[4] lib/<mod>/ picks tests/test-<mod>*.sh even when no suite names the file"
K="$TMP/k4"; mkkit "$K"; echo 'y=2' > "$K/lib/baz/inner.sh"
OUT="$(bash "$K/tests/run-all.sh" --changed 2>&1)"; RC=$?
if [ "$RC" -eq 0 ] && ran test-baz "$OUT" && ! ran test-foo "$OUT"; then
ok "test-baz via the module name"
else no "rc=$RC out=$OUT"; fi

echo "[5] no diff runs everything and says so"
K="$TMP/k5"; mkkit "$K"
OUT="$(bash "$K/tests/run-all.sh" --changed 2>&1)"; RC=$?
if [ "$RC" -eq 0 ] && grep -q 'found no diff.*running everything' <<<"$OUT" \
&& grep -q '^run-all: all 4 suites passed' <<<"$OUT"; then
ok "full glob on an empty diff"
else no "rc=$RC out=$OUT"; fi

echo "[6] an explicit base scopes the diff to the commits after it"
K="$TMP/k6"; mkkit "$K"; echo 'x=3' > "$K/lib/foo/foo.sh"; ( cd "$K" && g commit -qam foo )
OUT="$(bash "$K/tests/run-all.sh" --changed HEAD~1 2>&1)"; RC=$?
if [ "$RC" -eq 0 ] && ran test-foo "$OUT" && ! ran test-bar "$OUT"; then
ok "test-foo from the committed change"
else no "rc=$RC out=$OUT"; fi

echo "[7] a red always-on lint fails the run even when the diff names nothing"
K="$TMP/k7"; mkkit "$K"; printf '#!/usr/bin/env bash\n# always: fixture tree-wide lint\necho "FAIL: orphan"\nexit 1\n' > "$K/tests/test-lint.sh"
( cd "$K" && g commit -qam red-lint ); echo 'z=1' > "$K/lib/foo/orphan.sh"
OUT="$(bash "$K/tests/run-all.sh" --changed 2>&1)"; RC=$?
if [ "$RC" -eq 1 ] && grep -q '^run-all: FAILED ->.*test-lint' <<<"$OUT"; then
ok "the pinned lint is not skippable"
else no "rc=$RC out=$OUT"; fi

if [ "$fail" -gt 0 ]; then echo "test-run-all-changed: $pass passed, $fail FAILED" >&2; exit 1; fi
echo "test-run-all-changed: all $pass passed"
Loading