A closed-loop Claude Code workflow: you set the goal and the gates, agents loop until the verifier passes. Builder → verifier → fix-agent retry, by default.
/plugin marketplace add dwarvesf/dwarves-kit
/plugin install kit@dwarves-marketplace
Open your repo and run /kit:onboard. It previews every write and adopts the repo with sane
defaults. Then run /kit:start at the top of every session: it reads where the repo stands and
hands you the next command. Onboarding never finishes; /kit:start is the onboarding.
Condensed walkthrough: docs/QUICKSTART.md. Full command reference:
MANUAL.md.
Agent workflows are shifting from prompt -> output to goal -> loop -> evaluate -> improve -> result. dwarves-kit is the closed kind of that loop: you set the goal and the gates up front, and agents iterate inside them until a read-only verifier passes, never grading their own homework.
It ships as a toolbox, not an appliance. Every subsystem is a standalone shell command that already works on its own, bash lib/board/board.sh --help, and the same for gate/stats/classify/spec/goal/session/precedent/intake/wrap (e.g. bin/precedent find "<query>"), so there is no kit uber-binary and no install step to just poke at it (the same bash reads under pi, opencode, Claude Code, or a bare terminal). Wiring it into Claude Code (bash install.sh) adds an always-on safety spine and lets you --with exactly the modules you want and nothing else; Install has the layers.
The loop itself is one spec-driven lifecycle, think → spec → execute → review → ship → wrap → retro, with a gate at every phase boundary:
flowchart LR
goal([goal + gates<br/>set up front]) --> T
T[think<br/>forcing questions] -->|advisory| S[spec<br/>the contract]
S -->|spec-drift guard| X[execute<br/>builder → end verifiers → fix-agent]
X -->|BLOCKING: verification pipeline| R[review<br/>verdict recorded]
R -->|advisory| SH[ship<br/>version + PR]
SH -->|BLOCKING: ship-gate + push-to-main| RE[retro]
RE -.->|feeds the next cycle| T
That lifecycle is the middle of a longer arc. A board row becomes a landed PR like this:
_meta/BACKLOG.md /kit:assign lane-classify.sh
+----------------+ pull +----------------+ size +--------------------+
| ID-NNN, queued | -----> | goal draft, | ------> | tiny | normal | |
| | | scope fence | | full | bug | |
+----------------+ +----------------+ | backfill |
+--------------------+
|
+----------------------------------------------------------+
v
think --> spec --> execute --> review --> docs --> ship
| | | | | |
advisory spec- verification advisory advisory ship-gate +
drift pipeline push-to-main
guard |
v
PR open --> /kit:wrap
board flip, merge own PR,
deploy check, tidy, activity line
|
v
ID-NNN row --> shipped
Two gate classes sit on those boundaries: blocking (the verification pipeline, the ship-gate, the push-to-main blocker, mechanical, they stop a bad outcome) and advisory (think, review, they surface findings, never block). The autonomous-loop hardening adds a fresh-context re-audit of every done-claim, a kit-default cross-cutting advisor lens on top of the specialized reviewers, and a deployable-done proof gate, so a closed loop can run long without drifting into self-graded slop.
One builder implements the whole spec, then every task runs through one end verification pipeline (verifiers → fix-agent retry), and hooks enforce safety automatically (rm -rf, push-to-main, force-push, and secret-file reads are blocked). The builder is a domain specialist when one exists (a migration or data-pipeline worker, picked by a deterministic lookup), and /kit:draft-agent installs a reusable named agent (SPEC-089).
You drive it by intent, not by memorizing commands. Say what you want; the kit reads your intent, runs the right step, and stops only at the real decisions:
You: "add a --version flag to the CLI"
Kit: scopes it -> writes the spec -> builds + verifies -> ships,
pausing only where it needs your call.
The /kit:* commands below are those same actions named explicitly, for when you prefer to type them. You rarely need to.
It is bash-first (every hook readable in 30 seconds), and every component traces to a proven pattern, no novel inventions. The point of the kit is the handoff: a solo technical lead writes the spec, a contractor runs /kit:execute against the same spec.
New here? Install below, then run your first cycle. The full operator reference (every command, hook, and agent, plus troubleshooting) lives in MANUAL.md.
An open loop (the agent roams free and judges its own output) is a fast slop machine unless your standard is airtight and your budget is unlimited. The kit takes the closed shape instead: a human designs the path once, agents iterate inside it. What makes the loop trustworthy is that the gates that block are mechanical (bash hooks, tests, read-only verifiers), never the agent grading its own homework. The remaining phase gates advise and route rather than block: detect, don't dictate.
| Open loop | dwarves-kit (closed) |
|---|---|
| agent plans its own route | the spec is the contract, written before any build (validated on the full lane) |
| agent grades its own work | hard gates are mechanical: bash hooks, tests, read-only verifiers |
| loops until the budget dies | bounded: fix-agent retries max 2, then escalates to a human |
| one loop size fits all | risk lanes: tiny work skips the ceremony entirely |
The lifecycle above is what one run looks like. Across runs, the kit is organized as five stages (formerly called "legs", ADR-0034; the fifth renamed Learn -> Reflect by ADR-0036): Shape turns work into contracts, Build builds inside them, Watch records what happened, Check gates every boundary, and Reflect distills the record into proposals for the next cycle.
flowchart LR
SH([Shape]) --> BD([Build]) --> WA([Watch]) --> RF([Reflect])
RF -->|cited proposals,<br/>human promotes| SH
CH([Check]) -. gates every<br/>phase boundary .- BD
This diagram is the stages. For how DATA actually moves through them (telemetry -> proposal -> board -> ship, the ledger write/read paths, and the module map of who calls whom), see docs/data-flow.md.
Stages are metadata, not directories: each module keeps its name and install unit, and declares a primary stage. The authoritative assignment (machine copy in lib/config/module-registry.md, rendered by config list):
| Stage | Modules / subsystems |
|---|---|
| Shape (Specify) | spec, classify, precedent, intake, goal, board (input side), sync (spoke intake; outward mirror is its Watch side) |
| Build (Execute) | queue, mega, worktree, quiz_gate |
| Watch (Observe) | stats, session (capture side), telemetry, sync (outward mirror side; absorbed the bridge cockpit mirror 2026-07-16) |
| Check (Govern) | gate, money_gate, advisor, gauntlet, wrap |
| Reflect | reflect (was learn), weekend_batch, session (harvest), board (staging/promote), skill-curator, prose_rag (registry assignment, pending ADR-0034 amendment) |
| (no stage) | cosmetic (statusline; orthogonal to the loop) |
Two modules honestly span stages: board (Shape's intake on one side, Reflect's staging/promote on the other) and session (Watch's capture, Reflect's harvest).
What happens to a run's data after it ships: every gate decision and run outcome appends to the ledgers (append-only, never rewritten). stats projects them read-only; session intel writes the weekly digest, harness scorecard included; reflect propose distills cross-run evidence into cited proposals in a staging file; reflect drain renders that staging for review; board promote is the human gate that turns a proposal into a backlog row feeding the next Shape stage. Every automated stage ends at a staging file or a rendered surface, never a direct write to a board or ledger: propose, never dispose.
Layered by design: the SPINE installs unconditionally (six hooks guarding push, merge, secrets, and commit format, ADR-0024's irreversible boundary); everything else is an opt-in MODULE via --with <a,b,c>, recorded in your own project's kit.toml [modules] (a per-consumer install record, re-runnable; never a runtime registry a hook reads back). Run the installer with no --with at all to get the spine and nothing else.
| Module | What it wires | Kind |
|---|---|---|
board |
backlog-stage (SessionEnd: stage session work-items to a staging file, opt-in via BACKLOG_STAGE_AUTO=1, default off); its --surface pass also runs intake-sweep (consumer-declared deferred-link sources, config-gated); board-row-gate (PreToolUse Bash: blocks a commit that adds a new board row without a board-row-ok: <reason> line) |
2 hooks |
session |
context-readiness, output-offload, pre-compact-backup, post-compact-reinject, session-state-save, harvest, citation-guard, context-budget (warns once at 65% of the model's context window, again at 70%, KIT_CTX_WARN_PCT/KIT_CTX_STRONG_PCT/KIT_CTX_WINDOW); plus a PATH shim for the session CLI (session <intel|observe|recall|report|semantic>, ADR-0034: the five prefixed CLIs collapsed into one entry) |
8 hooks + 1 CLI |
advisor |
context-hints (session-elapsed + keyword skill hints) + tool-policy-guard (PreToolUse allow/ask/deny per tool domain; inert until a tool-policy.json exists) |
2 hooks |
cosmetic |
auto-format, notification, slop-cleaner, statusline, codebase-index, permission-auto-approve |
6 hooks |
queue |
/kit:mega + /kit:dispatch machinery (lib/queue/orchestrate.sh), the overnight queue launcher (lib/queue/queue.sh) |
hookless (lib) |
stats |
the stats CLI, a read-only projection over the run/gate ledgers |
hookless (uv CLI) |
quiz_gate |
/kit:quiz-gate (ADR-0031 understanding-gate nudge) |
hookless (command) |
weekend_batch |
the debt-paydown reader/closer (bin/reflect debt; engine lib/reflect/weekend-batch.sh, relocated per ADR-0034, renamed from learn by ADR-0036), invoked by a consumer's own skill or directly |
hookless (lib) |
bridge |
FOLDED INTO sync 2026-07-16. The git↔Hermes cockpit mirror/writeback verbs (board.sh mirror/status/writeback, bridge=on rows in boards.txt) remain runnable as the legacy engine until the SPEC-002 P2 port (kit ID-290) re-lands them as a sync edge |
absorbed |
worktree |
worktree-provision on PATH (manual worktree env-symlink + install provisioner, lib/worktree-provision/) |
hookless (CLI) |
money_gate |
money-gate (PreToolUse Edit/Write guard for money-touching edits; inert until you set MONEY_GATE_REPOS) |
1 hook |
prose_rag |
prose-rag recall inject (UserPromptSubmit, dormant until PROSE_RAG_INJECT=1) + the prose-rag CLI alias over context-kit's ctx engine verbs (lib/prose-rag/) |
1 hook + 1 CLI |
sync |
board sync two-way spoke mirror (BACKLOG.md ⇄ Apple Reminders / Notion / Hermes kanban; engine lib/sync/, per-repo .kit.toml [sync] config), inert without [sync] sources |
hookless (lib) |
team_mode is a reserved, not-yet-installable slot (parked, see docs/PHILOSOPHY.md "Team mode: parked, not absent"); naming it in --with errors on purpose.
Add a module later (module #13 never re-onboards you): edit .kit.toml's [modules] section
by hand, then re-run /kit:adopt --refresh (or bash lib/adopt.sh --refresh <repo>) to re-wire
settings.json to match. Full detail: commands/adopt.md.
bash install.sh # spine only
bash install.sh --with board,stats # spine + those two
bash install.sh --prune --with board # explicit trim: re-install down to spine + boardIn any Claude Code session:
/plugin marketplace add dwarvesf/dwarves-kit
/plugin install kit@dwarves-marketplace
That's it. Hooks, commands, agents, and the skill all install automatically. No bash, no jq, no symlinks. Updates via /plugin update kit.
To get the kit listed on Anthropic's official marketplace (claude-plugins-official), submit it via claude.ai/settings/plugins/submit. One-time manual step; not blocking the self-hosted install above.
codex plugin marketplace add dwarvesf/dwarves-kit
codex plugin add kit@dwarves-marketplaceCodex loads the five hard controls through hooks/codex-hooks.json: destructive-command safety, secret-file protection, ship completeness, commit format, and premature-completion protection. Open /hooks once after installation and trust the exact hook definitions. Each command pins the adapter and policy content hashes, so changed code fails closed before execution and also changes the trust definition when the manifest refreshes.
This first Codex package is a narrow enforcement adapter. It does not inspect prompts, assistant output, hosted tools, or every specialized tool path. The commands, agents, and skills remain Claude Code authoring surfaces until Codex-native loaders ship.
Demoted from the default doc path (Moment 1 of the onboarding design is plugin-only). Reach for it only in environments without Claude Code's plugin system (CI, project templates, older Claude Code versions), or as a kit maintainer:
git clone https://github.com/dwarvesf/dwarves-kit.git ~/.claude/dwarves-kit
cd ~/.claude/dwarves-kit && bash install.shRequires jq (for settings merge) and git; the installer refuses to start without them. No root / no package manager (CI images, locked-down containers): drop a static jq binary onto PATH instead, e.g. mkdir -p ~/bin && curl -fsSL -o ~/bin/jq https://github.com/jqlang/jq/releases/latest/download/jq-linux-arm64 && chmod +x ~/bin/jq && export PATH="$HOME/bin:$PATH" (pick the asset for your arch). Do not hand-copy the kit's files around a missing jq: the settings/hooks merge is the step that arms the guardrails, and a partial copy silently ships without them.
Hooks arm only in a NEW Claude Code session. The installer registers hooks in settings.json, but a session that was already running (and any single-shot headless run: claude -p, CI agents) never re-reads them, so every gate is advisory there: gate-ledger calls still work and record, but nothing blocks. Restart the session after installing; treat headless runs as unguarded by design.
Cloning in place is simplest, but install.sh also runs from a checkout anywhere. To uninstall: bash ~/.claude/dwarves-kit/install.sh --uninstall.
The two paths no longer collide. If the plugin is already installed, install.sh detects it and does a compat-only install: it symlinks the legacy ~/.claude/dwarves-kit/{lib,WORKFLOW.md,AGENTS.md} paths, so docs that still call bash ~/.claude/dwarves-kit/lib/<x>.sh (plain bash, where ${CLAUDE_PLUGIN_ROOT} is unset) keep resolving, and skips the hook + command registration the plugin already owns. No double-registered hooks. Force the full bash install with KIT_FORCE_FULL=1 bash install.sh (e.g. for the statusLine HUD, which the v1 plugin schema does not configure).
Invocation differs by path: installed as the plugin, commands are namespaced /kit:<name> (e.g. /kit:spec); via the bash installer they resolve bare /<name> (e.g. /spec). This README uses the plugin form.
After install, open a Claude Code session in your project and run one full lap. A tiny change is the best first try. Told as a story: an interview turns your ask into a blueprint, a crew builds it, and an inspector won't let it leave without a stamp; the full workshop tour is one page, /kit:onboard or docs/glossary.md.
/kit:startorients you and suggests the next step./kit:thinkand describe the change (e.g. "add a--versionflag to the CLI"). It throws 6 forcing questions at the idea; answer them./kit:specwrites the contract todocs/specs/SPEC-NNN-<slug>.md./kit:executeruns the autonomous build: one builder implements the whole spec from a brief (goal, acceptance, routes, territory), then one end pass checks every task against its acceptance criteria, with integration and acceptance verifiers behind it and a fix-agent retrying fixable failures (max 2)./kit:reviewthen/kit:ship: review gate, then version bump, changelog, conventional commit, PR.
That is the whole loop. The spec is the unit of handoff: a contractor running /kit:execute reads the same docs/specs/SPEC-NNN-<slug>.md you wrote. To see the artifact set without running anything, browse examples/hello-spec/.
/kit:start Detect state, suggest next command (entry point)
/kit:onboard Guided first-run: install mode, adopt, module picker, two-idea tour
/kit:think Challenge the idea (5 min)
/kit:design Opt-in: shape the solution with you before /spec
/kit:spec Generate the spec + 4 parallel researchers (15-30 min)
/kit:spec-validate Stress-test the spec (10 min)
[hand off to contractor]
/kit:execute Autonomous: builder > end verifiers > fix-agent retry loop
or
/kit:next Manual: pick next task, load context, you drive
[hooks enforce during build]
[statusline shows context budget]
[session-state-save persists progress on every stop]
[slop-cleaner flags bloat at stop points]
/kit:review Single-pass review (10 min)
/kit:review-team Parallel 3-lens review, confidence-gated + validated findings
/kit:verify Re-run tests read-only; PASS / FAIL / INCONCLUSIVE
/kit:docs Update all docs to match code (5 min)
/kit:explain Literate-diff explainer: understand a change, not click-to-merge it
/kit:quiz-gate ★-tap nudge: a 5-question quiz from the diff before merging a significant PR
/kit:pitch Outward buy-in doc from the spec + proof + impl-notes + grill record; ends in an ask, never fabricates
/kit:ship Review gate, version bump, changelog, commit, PR
/kit:retro Retrospective (10 min, after shipping)
Work is sized by risk lane before it starts (tiny / normal / full / bug, plus a backfill lane for reviewing an existing codebase and writing the operating-layer docs without changing behavior). The default is normal, and the words of a task never pick full: the classifier only suggests it, and the diff floor at push gives any hard-path change the full lane's gates. Each lane's phases are data in kit.toml ([lane.<name>]), which a repo can override in its own .kit.toml. The lanes, the gate at each phase boundary, and the operate-contract the agent follows live in AGENTS.md and WORKFLOW.md.
The whole build goes through: builder > end verifiers (task-verifier over every task, integration, acceptance) > fix-agent (if needed). The builder never grades its own work; the verifiers are separate read-only agents.
/kit:execute (orchestrator)
owns the spec's task list,
dispatches one builder for the spec
|
v
+----------------+
| builder | implements the spec
+----------------+
|
v
+----------------------+
| task-verifier | read-only gate: acceptance
| (cannot edit code) | criteria + tests
+----------------------+
| | |
PASS FAIL:fixable FAIL:escalate
| | |
v v v
mark done +-----------+ stop,
tasks | fix-agent | ask the human
+-----------+
|
+--> back to task-verifier
(max 2 retries, then escalate)
Within one spec, tasks run sequentially. Across specs, /kit:dispatch fans out disjoint VALIDATED specs into parallel git worktrees behind a disjointness gate; across sessions, a passive goal registry (lib/goal/goal-registry.sh) keeps concurrent same-machine sessions from colliding. The kit deliberately stops short of a DAG scheduler, a coordinating daemon, or cross-machine orchestration. For those, run GSD v2 or Nimbalyst alongside it.
Hooks (automatic, event-triggered)
| Hook | Event | What it does |
|---|---|---|
| codex-hook-adapter | Codex PreToolUse, Stop | Normalizes Codex payloads and dispatches the shared hard policies; contains no policy rules |
| anchor-root | Every dispatched event (except secrets-guard) | Wraps every hooks.json/settings.json entry: cds to the repo or worktree root, then execs the real hook, so no hook reads or writes relative to a subdirectory; contains no policy rules |
| safety-gate | PreToolUse(Bash) | Blocks rm -rf (build-artifact allowlist), push to main, force push, DROP TABLE, git reset --hard, kubectl delete |
| secrets-guard | PreToolUse(Read|Edit|Bash) | Blocks reads of secret files (.env, ~/.ssh, ~/.aws, .pem); canonicalizes the path first |
| commit-format | PreToolUse(Bash) | Blocks non-conventional / >72-char / spec-ID commit subjects |
| ship-gate | PreToolUse(Bash) | Blocks push/PR without a proof-of-done record + recorded lane gates (ADR-0024 boundary); a diff that touches a hard path (migration, auth, secrets, CI, kit config, data loss) owes the full lane's gates whatever lane the spec names |
| context-readiness | SessionStart | Detects project + board state (board:Nq), suggests the next step intent-first |
| context-hints | UserPromptSubmit | Injects session elapsed/idle time + keyword-matched skill hints (empty map by default; wire your own via CONTEXT_HINTS_SKILLMAP) |
| anti-rationalization | Stop | Catches Claude declaring work done prematurely |
| slop-cleaner | Stop | Flags bloated code in recently modified files |
| session-state-save | Stop, SubagentStop | Persists session state, rotates last 10 archives |
| citation-guard | Stop | Flags (or blocks, CITATION_GUARD_STRICT=1) hallucinated file:line citations in the final message |
| money-gate | PreToolUse(Edit|Write|MultiEdit) | Asks before a money-touching edit lands in a repo named in MONEY_GATE_REPOS (inert unset) |
| prose-rag | UserPromptSubmit | Injects relevant prior notes on recall-shaped prompts (dormant unless PROSE_RAG_INJECT=1) |
| board-row-gate | PreToolUse(Bash) | Blocks a git commit that adds a new board row (a first-cell ID absent from HEAD's _meta/BACKLOG.md or BACKLOG.md, any prefix) unless the message carries a board-row-ok: <reason> line. On by default; a repo opts out with [gate] board_row_gate = false in its committed .kit.toml; session kill switch DWARVES_KIT_SKIP_BOARD_ROW_GATE=1 |
| batch-debt-warn | PreToolUse(Bash) | Warns once when a session merges a 2nd PR with no lane START in the gate ledger since the first merge |
| context-budget | UserPromptSubmit | Warns once at 65% of the model's context window (advisory), again at 70% (directive) (KIT_CTX_WARN_PCT/KIT_CTX_STRONG_PCT/KIT_CTX_WINDOW); clears on a drop below the warn threshold (e.g. after /compact or /clear) |
| auto-format | PostToolUse(Write|Edit) | Runs formatter on every file change |
| output-offload | PostToolUse(*) | Offloads a >2k-token tool output to a file + leaves a terse pointer |
| spec-drift-guard | PreToolUse(Write) | Warns when creating files not in the spec |
| pre-compact-backup | PreCompact | Saves structured session snapshot before compaction |
| harvest | PreCompact, SessionEnd | Stages durable session learnings to a repo-relative ledger (PreCompact); drafts a LAB_LOG entry (SessionEnd --lab-log). Never writes a durable home; a human flushes. Stands down on a host where the scheduled harvest sweep is installed (see below) |
| backlog-stage | SessionEnd | Opt-in (BACKLOG_STAGE_AUTO=1, default off): stages forward-looking work-items from the session to a repo-relative staging file. Never writes the board directly |
| intake-sweep | SessionStart (via backlog-stage --surface, same opt-in) | Sweeps consumer-declared deferred-link sources (_meta/intake-sources.json: jsonl / command adapters) into the same staging file. Config-gated no-op; never writes the board directly |
| post-compact-reinject | SessionStart(compact) | Re-injects critical rules after compaction |
| notification | Notification | Desktop alert when Claude needs input |
| permission-auto-approve | PermissionRequest | Auto-approves allowlisted read-only commands; everything else gets the normal prompt |
| tool-policy-guard | PreToolUse | Enforces the tool-choice policy file (allow/ask/deny per tool domain; the enforcement half of the dashboard's tool-policy page) |
| statusline | StatusLine | Shows model, branch, context %, cost, thinking mode |
| codebase-index | SessionStart (opt-in) | Background-indexes the repo into codebase-memory-mcp |
Harvest sweep (scheduled capture). The per-session hook only sees a session that ends or compacts cleanly. The sweep reads transcripts on a schedule instead, so a killed or long-running session is still harvested. python3 hooks/harvest.py --sweep reads new claude sessions (plus devin when harvest.sources lists it) behind a per-source cursor, extracts learnings and pattern sightings, stages them into sweep ledgers, and writes a wrap-shaped report per run. It builds, pushes, and merges nothing.
--sweep --dry-runprints the manifest of what a run would do and changes no cursor, ledger, or report state (only the raw extract cache is written).--since <iso|epoch>widens the window.--statusprints the newest report path, its candidate count, and the queued learning count.--flush-listprints every queued learning as one JSON array.--mark-flushed <row-id> <ref>marks one routed once a human or the learning-ledger flush has written it to a durable home.bash deploy/macos/harvest-sweep/install --applyinstalls the macOS LaunchAgent and writes the per-hostinstalledmarker. It refuses unlessharvest.enableis true.--uninstallremoves both files and keeps all state. Seedeploy/macos/harvest-sweep/README.md.- The extractor is
claude -p --model <harvest.extractor_model>(defaultsonnet) with every tool, MCP server, and session write off. When that call fails for any reason but auth, a usage limit included, the same prompt runs once throughcodex exec --sandbox read-onlyon the host's ChatGPT login (harvest.extractor_fallback = "codex", the default;nonedisables it), and the report carries aSTATE ... extractor fallback used: codexrow. Codex cannot turn its shell off, so a hostile transcript can make it read local files; its reply passes the same redaction before storage. - A run where Claude hits a 5-hour or weekly usage limit and the fallback also fails holds: it stops, keeps the cursor, counts no failure, and exits 0. The next scheduled run resumes.
- With the sweep active,
wrap.distill = "harvest"lets/kit:wrapland only and leave distillation to the sweep. Config lives in[harvest]inkit.toml.
Which hooks BLOCK vs warn vs neither is a declared contract: docs/architecture.md "Hook fallback layer" (hard / advisory / convenience, parity-pinned).
Commands (manual, human-triggered)
| Command | Phase | What it does |
|---|---|---|
| /kit:start | Entry | Detect project state, suggest next command |
| /kit:grill | Intake | Universal intake interview: type-shaped questions, one at a time, answers written as they resolve |
| /kit:wayfind | Intake | User-invoked: chart a too-foggy-for-one-session effort as a decision map (_meta/megagoals/<slug>/map.md + typed tickets), resolve one per session, hand off to /kit:spec or a ROADMAP |
| /kit:think | Think | 6 forcing questions to stress-test an idea |
| /kit:design | Design | Opt-in: interactive solution-design beat (one question at a time) before /spec |
| /kit:devs-team | Design | Opt-in: 5-lens parallel critique of the solution (brief or spec), report-only |
| /kit:visual-team | Design | Opt-in: 5-lens parallel critique of a visual/UI design (downstream-facing) |
| /kit:ui-design | Design | Opt-in, downstream: UI brief -> generate (frontend-design) -> critique -> revise loop |
| /kit:prototype | Design | Opt-in: throwaway spike answering ONE design question (logic TUI or 3-5 UI variants); decision folds into the brief/spec, code survives on a prototype/<name> branch |
| /kit:assign | Orchestrate | Turn a backlog item (ID-NNN) into a scoped goal draft + route it into the lane |
| /kit:dispatch | Orchestrate | Fire N disjoint VALIDATED specs concurrently, each in its own worktree, behind a disjointness gate; lead-owned merge |
| /kit:mega | Orchestrate | Mirrors the plan-for-mega-goal skill: decompose 3-8 dependent sub-goals, front-load every clarification once, set the per-run merge config, hand off to the bounded loop; ship-layer auto-merge rides the ship-gate via lib/goal/mega-merge.sh, never bypasses it |
| /kit:spec | Spec | Generate docs/specs/SPEC-NNN-.md; a Depth: line under Lane: decides whether research agents run (none at standard) |
| /kit:spec-validate | Spec | 7 adversarial reviewers attack the spec (incl. solution-design, design record, sustainability) |
| /kit:test-plan | Spec | Opt-in: coverage matrix from acceptance criteria into the spec's ## Test plan section |
| /kit:feature-map | Spec | Source-cited, agent-checkable feature inventory for ANY target project: per-module spec + a top-level checklist. Standalone (what does this codebase do) or migration source of truth (what needs porting) when a port target is named |
| /kit:execute | Build | Autonomous: builder > end verifiers > fix-agent retry loop |
| /kit:next | Build | Lightweight: picks next undone task, loads context, you drive |
| /kit:verify | Verify | Read-only re-run of task-verifier + integration-verifier, no rebuild; verdict PASS / FAIL / INCONCLUSIVE with the claim restated falsifiably |
| /kit:battery | Verify | The full independent right arm for a finished branch: fresh-context acceptance verifier + multi-lens review + advisor in parallel at prescribed model tiers, baseline-aware, findings merged, fixes applied by the lead |
| /kit:greenlight | Ship | Post-push CI-green lane: snapshots an open PR's checks, fixes real failures via the fix-agent shape, retries flaky ones within a bounded budget, reports one terminal state; opt-in, never hard-gates merge |
| /kit:debug | Bug (off-cycle) | Systematic debug loop: root cause before any fix, evidence ledger, 3-fix wall |
| /kit:review | Review | Paranoid single-pass code review |
| /kit:review-team | Review | Parallel 3-lens review (security + architecture + test-coverage); findings confidence-gated, deduped by fingerprint, verdict-driving ones adversarially validated per finding |
| /kit:test-plan-review-team | Verify | Adversarial critique of the spec's ## Test plan, report-only: a two-lens --light pass by default, the full 6-lens team with a bounded revise loop when the spec's Depth: names a blind-spot |
| /kit:test-write | Build | Turns a SOLID-verdict ## Test plan critique into real, executing test code via test-writer, one case per matrix row |
| /kit:onboard | Entry | Guided first-run: detect install mode (plugin/bash/both/none), offer /kit:adopt, pick modules, capture consumer knobs, disclose plugin-path gaps, a two-idea tour (lane, proof of done) plus a menu; previews + confirms every write, decline = no-op |
| /kit:adopt | Entry | Retrofit the operate-contract onto an existing repo (a small AGENTS.md pointer that never overwrites a repo's own, loader, proof marker, classifiers), idempotently |
| /kit:docs | Docs | Cross-reference diff against all doc files, fix drift |
| /kit:explain | Understand | Literate-diff explainer (background -> intuition -> prose-ordered diff -> diagram); composes narrate-log + svg-knowledge-diagram, grounded in the diff not the agent's narrative |
| /kit:quiz-gate | Understand | ★-tap nudge before merging a significant+worthy gate PR: 5 diff-grounded quiz questions routed through deep-understand, three logged responses (engage/defer/wave), advisory never must-pass |
| /kit:pitch | Understand (outward) | Assembles an outward buy-in doc from the spec, proof-of-done, implementation-notes, and grill/DEBT ledger records; the outward twin of /kit:explain, ends in an ask not a quiz; never fabricates a missing source |
| /kit:ship | Ship | Review gate, version bump, changelog, commit, PR |
| /kit:wrap | Land | Session-scoped landing step after ship: board rows, the operator's own PR merges, deploy check, branch and worktree tidy, the activity line, calls /kit:retro when a shipped PR merged |
| /kit:retro | Reflect | What worked, what hurt, action items for next cycle |
| /kit:kit-health | Meta | Self-assessment against kit philosophy |
| /kit:absorb | Meta | Maintainer-only: audit upstream sources (Credits drift + seed-rescan) + draft a dated absorption proposal |
| /kit:draft-agent | Meta | Meta-agent agent-builder: generates a subagent (or sub-goal file) from a description and installs it by default (--draft to stop at a review draft) |
| /kit:gauntlet | Meta | Probe-convergence engine: converges an artifact (docs, a runbook, a spec, an API surface) toward a fixed outcome by having a fresh clean-room probe agent attempt the outcome contract unaided each round; failures revise the artifact, every round persists a full run record. Onboarding ships as the reference preset |
Agents (dispatched by commands) and Skills (Claude-triggered)
| Agent | Dispatched by | What it does |
|---|---|---|
| task-verifier | /execute, /verify | Read-only verification against spec + tests |
| integration-verifier | /execute, /verify | Read-only cross-task wiring + global acceptance check (multi-task specs) |
| acceptance-verifier | /execute, /verify | Executes the spec's ## Verification section against the build (read-only) |
| system-verifier | /verify | Runs the whole project's test suite end to end (read-only) |
| recheck-verifier | /execute | Fresh-context re-audit of a verifier PASS (sampled, plus every self-attested row): re-executes the recorded command |
| claim-verifier | any command | Adversarial N-skeptic panel over a load-bearing free-text claim |
| fix-agent | /execute | Targeted fixes on FAIL:fixable (max 2 retries) |
| data-etl-worker | /execute | Domain implementer: pipelines/transforms (DuckDB SQL first) |
| db-migration-worker | /execute | Domain implementer: schema migrations + rollback + backfill |
| code-reviewer | /review-team | Focused review with configurable lens |
| security-reviewer | /review-team | Deep OWASP-style security audit |
| api-reviewer | /review-team | API-contract lens (breaking changes, versioning, idempotency) |
| frontend-reviewer | /review-team | Frontend lens (a11y, semantic HTML, focus, responsive) |
| infra-reviewer | /review-team | Infra lens (deploy/rollback safety, CI/CD, least-privilege) |
| performance-reviewer | /review-team | Performance lens (hot paths, N+1, allocations, caching) |
| advisor | final boundary | Cross-cutting kit-default lens: critique + over-suggest modes |
| break-it | /kit:battery (escalation) | Adversarial prober: hunts one concrete input the suite does not constrain; rung 2 of coverage -> probe -> mutation |
| brief-reviewer | /think | Static review of a brief/requirement before it hardens into a spec |
| responding-to-review | /review-team | Verifies review findings, pushes back when wrong, proposes fixes (no performative agreement) |
| slop-stripper | /review-team | Behavior-preserving AI-slop strip pass: surgical edits only, never behavior changes unless fixing a real bug |
| agent-effectiveness | agent authoring | Validates a new/changed agent definition's effectiveness (4 lenses) |
| doc-verifier | /docs | Read-only check that docs match the live codebase |
| devops-triage | on-demand | Read-only production-alert triage: Workers Logs history + git around the deploy sha into a bounded root-cause verdict |
| research-stack | /spec | Maps technology stack (brownfield) |
| research-context | /spec, /kit:test-plan | Quick brownfield orientation (endpoints, models, UI, tests, recent history), capped at 80 lines |
| research-architecture | /spec | Maps architecture patterns and conventions |
| research-pitfalls | /spec | Finds landmines before implementation |
| research-features | /kit:feature-map | Deep, uncapped, source-cited feature inventory for any project: MIGRATE table + parity contract when porting, else a behavior contract |
| meta-agent | /draft-agent | Drafts a new subagent (or sub-goal file) from a one-line description |
| test-writer | /kit:test-write | Turns a reviewed test-plan coverage matrix into runnable test code, one case per matrix row |
| audit-scanner | every audit-loop skill | Shared read-only Tier-2 evidence scanner for audit-loop instances; roster physically cannot write, and it judges saved evidence rather than gathering it (a network-side Tier 1 saves output first) |
| Skill | What it does |
|---|---|
| backlog-reconcile | Audits a repo's _meta/BACKLOG.md Active queue (audit-loop instance, general-purpose): row Status verdicted against its Target artifact spec's own Status: header or a git-log match, fixes behind a PR gate |
| doc-drift | Whole-estate doc audit (audit-loop instance): enumerates every living doc, verdicts each claim against the live repo, fixes drift behind a PR gate |
| ci-drift | Whole-estate CI audit (audit-loop instance): enumerates every workflow + GitHub-side state (enabled, secrets/vars, runners, releases, environment policy), verdicts each against the live repo, fixes drift behind a PR gate |
| gauntlet-proof-audit | Audits committed gauntlet run records (audit-loop instance, general-purpose): markers well-formed, recorded verdict vs committed checker-output.txt, findings-count reconciliation, scrub clean, run-dir grammar; REMOVE disallowed, report-first |
| repo-hygiene | Whole-repo decay audit (audit-loop instance, general-purpose): five detectors over one checkout (unreferenced doc, stale staging drop, record parked in a control surface, log past its documented budget, large cold gitignored dir), each finding carrying its evidence inline, git mv the only fix it ever applies, PR-gated. Repo-scoped: the machine surface belongs to ops-toolkit tools/disk-reclaim |
| topology-drift | Maintainer-only (dwarves-kit repo dev only): audits THIS KIT's own feature estate (audit-loop instance), cross-checks the generated docs/FEATURES.md registry against the docs/workflow-paths.md path index both directions, re-places only delta features on the topology, PR-gated. To inventory a project the kit is pointed at, use /kit:feature-map instead |
| web-drift | Live-website agent-readiness audit (audit-loop instance, general-purpose): probes every site in WEB_DRIFT_SITES with lib/webcheck over read-only HTTP (groundwork, page, API tiers), verdicts each (site, check) pair, and files fixes as board rows in the repo that owns the site's source. Inert until the consumer declares its sites |
| get-api-docs | Fetches curated API docs via Context Hub before coding |
| loop-engineering | Designs a new bounded loop for the kit's orchestration: the gate (should this be a loop), then the anatomy (artifact / scanner / reviser / stop condition) on the generic bounded-revise engine |
| memory-tidy | Audits a repo's .claude/memory store: evidence-gated verdicts, PR-gated merges/deletions, index rebuild (judgment half of stats memory-sweep) |
| observe | Queries/renders the control plane (runs, gate verdicts, conformance, spend/cache economics, replay, dashboard) via the lib/bench CLIs |
| skill-review | Reviews + promotes skill drafts staged by skill-curator |
| stats | Queries/renders the ledger read plane (relocated from lib/stats/skill/ per ADR-0034 so it actually installs) |
A solo technical lead handing off implementation to contractors. The kit covers the full lifecycle with one shared spec format: the contractor running /kit:execute reads the same docs/specs/SPEC-NNN-<slug>.md you wrote with /kit:spec.
Also for a builder using Claude Code 6-8 hours a day who wants a context-budget HUD, automatic safety guards, session-state persistence across compaction, and slop detection at stop points.
- Teams of 10+ with a dedicated DevOps pipeline. The kit targets one engineer (or one lead + delegated contractors); one lead can still fan out parallel workers (
/kit:dispatch) and run concurrent same-machine sessions safely. Cross-machine orchestration, 3+ live operators, or goal-ordering chains are out of scope, pair the kit with Nimbalyst or Conductor for that. - Anyone who wants a UI. The kit is bash hooks + markdown commands. Open any file in a text editor; it's all readable.
- Projects already happy with GSD, gstack, or Trail of Bits' configs as standalone tools. The kit's value is integration; if format-translation overhead between standalone tools isn't hurting you, don't switch.
Directory layout
dwarves-kit/
tool.toml Kit metadata (name, version, language=bash, deps)
AGENTS.md Tool-agnostic operate-contract front door (any runtime reads it first)
WORKFLOW.md The cycle, the risk-tier lanes, the gates, and the flow/loop reference (ASCII diagrams)
MANUAL.md Operator reference: commands, hooks, agents, natural-language scenarios, troubleshooting
README.md / CONTRIBUTING.md / CHANGELOG.md / VERSION / LICENSE
CLAUDE.md Project template; the Claude-Code layer on top of AGENTS.md
install.sh / settings.json Bash install path
.claude-plugin/ Plugin install path (plugin.json, marketplace.json)
.github/workflows/test.yml CI: macOS + Ubuntu test matrix, on workflow_dispatch and v* tags only
bin/ STABLE consumer entrypoints (SPEC-184, one `<subsystem> <verb>` grammar per ADR-0034): `audit`/`board`/`classify`/`gate`/`goal`/`reflect`/`mega`/`precedent`/`intake`/`queue`/`session`/`spec`/`stats`/`config`/`plugin-check` thin forwarders to `lib/<subsystem>/`, plus the module CLIs (`prose-rag`, `worktree-provision`, `skill-improve`, `skill-review`) that keep their own names, and three standalone maintainer tools outside the forwarder pattern (`activate`, `release`, licensing and release cutting; `test-affected`, diff-scoped test runner with a pass cache). `learn` stays for one release as a deprecation forwarder to `reflect` (ADR-0036). A consumer (an adopted repo's board shim, the adopt-injected CLAUDE.md block) references `$DWARVES_KIT/bin/<name>`, NEVER a deep lib path, so an internal lib reorg cannot silently break it (the board-shim class of bug). Deployed by install.sh next to lib/.
agents/ Subagents dispatched by commands
commands/ Markdown command prompts
hooks/ Hook scripts + hooks.json plugin manifest
lib/gate/dispatch-gate.sh Disjointness gate + drift guard for /kit:dispatch (pure-bash concurrency moat)
lib/classify/lane-classify.sh Deterministic task-type -> risk-lane classifier + advisory floor check (used by /kit:assign + /kit:dispatch); optional `--files "<paths>"` on classify/explain/check escalates the kit-machinery gate on an actual EDIT to lib/ or hooks/, not a mere textual mention (SPEC-105, edit-vs-mention)
lib/goal/goal-registry.sh Cross-session running-goal registry: claim/list/log/release (multi-session moat + monitor)
lib/goal/goal-drafts.sh Goal-draft lifecycle: archive shipped drafts to .claude/goals/done/
lib/telemetry/lane-telemetry.sh Read-side lane-effectiveness aggregator over the run ledgers: report + misfires (reviewed at /kit:retro)
lib/queue/orchestrate.sh Non-LLM mega-goal driver: one fresh `claude -p` session per sub-goal so no session marathons; session-per-sub-goal, NOT the GSD-v2 engine (no priority/cross-machine/state-store). `run <dir>` flags: `--dry-run` (plan only), `--step` (pause for the operator between sub-goals), `--stream` (live stream-json tee'd to `.orchestrate/<id>.stream.jsonl`), `--board=roadmap|kanban|both` (event-sourced per-mega-goal kanban derived to `<dir>/BOARD.md` via `lib/board/backlog.sh`; default detects, ROADMAP stays canonical). `flip <dir> <id>` box-flip subcommand (mkdir-lock guarded, atomic). Opt-in trial `--backend orca` (or `MEGA_BACKEND=orca`, the flag wins) runs the ROADMAP through Orca Tasks and supervised Claude workers instead of one `claude -p` per sub-goal (`lib/queue/orca-backend.sh`, sourced only then); `status <dir>` prints each sub-goal's derived state and `orca-reset <dir>` rolls a run back; the ROADMAP box stays the only proof of done and the default path is unchanged. DAG-wavefront ON by default: at the default `WAVE_CAP=2` it runs dep-independent, `## Touches`-disjoint sub-goals concurrently (one worktree per session); a mega-goal whose sub-goals declare no `## Touches` still serializes (no-op), and `WAVE_CAP=1` forces the always-serial loop. `commands/mega.md` emits a `## Touches` per generated sub-goal so new mega-goals are wave-eligible. A `gate` sub-goal holds only its dependent chain; `gate!` halts the whole loop for a human. Robustness env (advisory): `WATCHDOG_STALL_SECS>0` backgrounds each session + flags it `stalled` after that long with no output (never kills); a dead/incomplete session never advances its box. Emits a `gate-ledger start` per dispatched sub-goal (rid from the goal file's `**Branch:**`) so mega-dispatched runs are tracked in `lane-telemetry`, not `?`. Multiplexer panes (SPEC-119, opt-in `MULTIPLEXER=1`): each wave session runs in its own tmux window (`TMUX_CMD`/`TMUX_SESSION` seams) for capture-pane/send-keys watch + intervene. Viewer push (SPEC-120, `PANE_VIEWER=auto` DEFAULT): on wave spawn one viewer tab auto-opens attached to the wave's tmux session (cmux/kitty/wezterm/ghostty/iterm/terminal auto-detected); `none` = pull only; headless degrades silently; unknown values rejected at pre-flight. Subagent panes: `panes <megadir> <target>...` grows one READ-ONLY tmux window per resolved transcript (a jsonl path, a directory of them, or `--latest` to derive the conductor's own `~/.claude/projects/<slug>/*/subagents` dir) for the DEFAULT background-subagent run mode, which has no wave to attach a pane to otherwise; each window runs `tail -F | jq` via the hidden `_pane-tail` re-entry (no shell in the pane -- steering stays with the conductor). `PANE_TAIL_JQ` overrides the formatter (default `lib/queue/pane-tail.jq`, a control-byte-stripping, length-capped line formatter for the subagent transcript schema). Always rc 0; skips warn on stderr and land in a `[panes] spawned N, skipped M` summary. Gitignore `.orchestrate/` + `BOARD.md` + `HANDOFF*.md` (derived/runtime)
lib/queue/queue.sh Overnight queue LAUNCHER: drives REAL interactive Claude Code `/goal` sessions via `tmux` send-keys (NOT headless `claude -p`, to sidestep the AUTH/KILL-CLASS risk; cmux was tried and dropped, no CLI-verified argv-safe launch primitive per SPEC-119/121). `run <src.tsv>` (rows `slug<TAB>repo<TAB>pointer`) opens a fresh window per queued mega, types `/goal <pointer>` + Enter, polls `capture-pane` for the completion marker (`RUNNER_DONE`/`RUNNER_GATED`, line-anchored AND blank-line-guarded so a soft-wrapped echo of the typed prompt cannot false-trigger), journals each verdict to `queue-journal.tsv`, and stops the night after two consecutive `error`-or-`stalled` megas. `--from-boards` rows get a `realpath`-resolved allow-list confinement (`QUEUE_ALLOWED_POINTER_GLOB`, defense-in-depth on top of sub-goal 04's own; a hand-authored tsv is exempt). Flags: `--dry-run`, `--max-megas N`, `--from-boards`. CONSUMER config: `TERMINAL_MUX`(tmux only)/`MUX_CMD`/`QUEUE_CLAUDE_*`/`QUEUE_JOURNAL`/`QUEUE_*_SECS`/`QUEUE_ALLOWED_POINTER_GLOB`. Aliased as `orchestrate.sh queue <src>`. Idempotent (a `done` slug is skipped on re-run)
lib/goal/mega-merge.sh Ship-layer auto-merge ENFORCEMENT for /kit:mega (ADR-0028 P2/P3): `gate <rid> <lane>` (decision, reuses `lib/gate/gate-ledger.sh check`) + `merge <pr> <rid> <lane> [--execute] [--posture=<val>]` (action; refuses unconditionally on a failing/missing gate, dry-run by default, `MEGA_MERGE_POSTURE` team-review opt-out) + `mark <pr> [repo]` (the SPEC-100 mark half, ID-089: opens a gate/gated-final PR as draft + `do-not-merge` so the `_merge_exclusion` guard always has a mark to catch; idempotent, gh via `MEGA_MERGE_GH`)
lib/board/board.sh The cockpit board command (SPEC-146; `mirror`/`status` added by SPEC-147; `writeback` added by SPEC-149): `board|next|set|states|priority [mode]|work` on one repo's BACKLOG.md (`--backlog-file`; `work` is the read-only who-is-on-what table, see `lib/board/work.sh`), `all <cmd>` for a cross-repo registry render (`--repo-root`/`REPO_ROOT`, `boards.txt`), `queue [--dry-run]` -- walks the registry, parses every repo's board via `lib/board/parse-board.sh`, and emits an allow-listed `slug<TAB>repo-path<TAB>pointer-path` feed for an overnight runner -- `mirror`/`status`: a one-way git -> Hermes kanban bridge over opt-in (`bridge=on` in `boards.txt`) repos + active mega-goals, `hermes kanban` CLI only (ADR-0001 native-first, no SQLite ATTACH), idempotent (a second run on an unchanged board is a zero-op no-op), incremental snapshot persistence -- and `writeback [--dry-run]`: the reverse leg, a Hermes-side card status move flows back into a repo's BACKLOG.md as a reviewable, HELD `chore/board-sync` PR (never auto-merged), gated by the mirror snapshot's `row_hash` conflict rule (git wins, always; a missing/corrupt snapshot refuses ALL edits rather than silently applying everything). Substantial mirror logic lives in `lib/board/board-mirror.sh`; substantial writeback logic lives in `lib/board/board-writeback.sh`; `board.sh` stays the thin dispatcher. Base kanban render still delegates to `lib/board/backlog.sh`; the `priority` quadrant awk + cross-repo `priority matrix` pivot are migrated in verbatim (byte-identical output is a pinned non-regression, see `docs/verification/board-tool/`). `run <ID>` is the single-row dispatch path: scaffolds a minimal mega-goal dir for one board row and prints the `orchestrate.sh run <dir>` command (`--exec` composes the launch; logic in `lib/board/board-run.sh`). No personal data in the kit: the registry + repo paths + Hermes target are consumer config read at runtime.
lib/board/parse-board.sh The one structured BACKLOG.md parser other tools reuse: `rows <file>` (id/status/full-line) and `queue-rows <file> <repo-name> <repo-root>` (allow-listed `#queue{repo=...,pointer=...}` token extraction -- charset gate, repo self-consistency, `../` traversal hardening, existence check; every failure is a skip with a stderr reason, never a hard error)
lib/board/board-mirror.sh The git<->Hermes kanban bridge engine, backing `board.sh mirror`/`status`: extracts opted-in BACKLOG.md rows (reusing `lib/board/parse-board.sh`) + active mega-goal roadmaps into normalized rows, diffs them (bash + `jq`/awk keyed comparison, no DuckDB) against an incremental NDJSON snapshot, and loads via `hermes kanban` CLI verbs only. Reachable native states are `{triage, ready, blocked, done}` only -- `todo`/`running` were probed live and found to have no durable CLI-only path (auto-promote back to `ready` within seconds); `_target_native`/`_create_flags_for`/`_followup_for` document the full state-mapping + Hermes-CLI-reality findings. `apply-plan` decodes each op's argv over NUL-delimited jq output (never a templated shell string), so a multi-line card body is preserved as one opaque argv element end to end -- the fix for a real bug this build's own live dev-home E2E caught (a newline-delimited decode silently split a multi-line body into extra positional args, rejected by the real CLI, masked by a naive stub). `HERMES_BIN` overrides the binary for tests (a stub logs argv; no real Hermes calls in the automated suite).
lib/board/board-writeback.sh The git<->Hermes kanban bridge WRITEBACK leg, backing `board.sh writeback`: sources `lib/board/board-mirror.sh` (not re-forked) for extract/hash/native-state machinery. `diff` reads each opted-in repo's live board (`hermes kanban --board <b> list --json`, one batched call per board) + the SPEC-147 mirror snapshot, and builds a validated changeset of rows whose Hermes status moved -- validating opted-in repo (defense in depth on top of the mirror's own filter), the CREATE-STATE RULE (a card still sitting in the column the mirror created it in never moved, so it is skipped: the snapshot records the column the mirror INTENDED, not one a human chose), the BLOCKED RULE (a card a human moved to `blocked` has NO git counterpart and is skipped, never written back as `parked`: `blocked` means someone cannot proceed right now, `parked` means deferred until a trigger and drops the row out of `board next` with no record of the blocker), the MIRROR-TERMINAL RULE (a `done` card whose completion result carries the mirror's own `board-mirror:` marker is the disappeared-row path reporting back, not a human finishing work, and is skipped rather than written back as `shipped` onto a row still live in git), a legal `backlog.sh` reverse-mapped target status, and the row_hash CONFLICT RULE (a Hermes-side edit applies only if the row's current git-side hash still equals the snapshot's recorded value; git wins, always). A missing or corrupt snapshot REFUSES ALL edits (explicit error, nonzero exit) rather than degrading to "no conflicts, apply everything". `apply` builds the `chore/board-sync` branch in an ISOLATED `git worktree` off the CURRENT HEAD (never the caller's own checkout, never a stale ref -- this is what keeps a concurrent append-only writer's row safe), edits ONLY the Status column of matched rows (reusing `lib/board/backlog.sh`'s own `set`), commits with `actor=hermes` in the body, pushes, and opens a HELD PR via `gh pr create` (argv-only, never a templated shell string; never auto-merged). Snapshot refresh updates ONLY `hermes_status` (to stop re-diffing the same not-yet-merged move); `row_hash` passes through UNCHANGED (the git SoT hasn't actually changed until the PR merges) -- `mirror`'s own idempotence self-heals the row once it does. `GH_BIN` (mirrors `HERMES_BIN`) overrides the `gh` binary for tests; no real Hermes write call or `gh` API call happens in the automated suite.
lib/board/work.sh `board work`: one read-only table of who is on what, how far along, and who is stuck. Joins the board rows, mega sub-goals, git branches and worktrees, `orca worktree ps` and the run ledger at call time and stores nothing; agent state is working, idle or unknown (file activity stands in for a worktree Orca does not list, marked `(files)`; with Orca itself missing it reads unknown); `--json` is the schema-1 form, `--idle-min N` sets the PARKED threshold
lib/board/board-run.sh The single-board-row dispatch path, backing `board.sh run <ID>`: reads the row's Item + Notes through `backlog.sh row` (the board's own parser; same no-row / duplicate-id refusals as `get`), scaffolds the MINIMAL mega-goal dir `lib/queue/orchestrate.sh` expects (one `- [ ] SG-01 <item> , auto` ROADMAP line, POINTER_PROMPT.md seeded from the row, a `goals/01-*.md` contract with `Model:`/`**Branch:**` headers, HANDOFF.md + DECISIONS.md stubs) under the repo's megagoals convention (`--dir` > `megagoal_root:` CLAUDE.md hint > first in-use root in canonical order > `_meta/megagoals/<id-slug>/` default), then prints the exact `orchestrate.sh run <dir>` command. It never launches a session itself; `--exec` composes the launch and forwards args after `--` (e.g. `-- --dry-run`). Idempotent: an existing file is reported `kept`, never overwritten (the `board init` created/kept convention).
skills/ Claude-triggered skills (one folder per skill; roster in the Skill table above)
rules/ Path-scoped coding-standard templates
examples/hello-spec/ Demo: small CLAUDE.md + SPEC.md walkthrough
tests/test-hooks.sh Hook behavior assertions
tests/test-meta.sh Structural integrity (manifests, frontmatter, cross-links)
docs/specs/SPEC-NNN-<slug>.md Specs, tracked in place via Status header (DRAFT/VALIDATED/SHIPPED); hooks pick the active one by git branch
docs/ The kit's design record (not needed to USE the kit; see docs/README.md)
README.md Map of docs/: what each file is, and that you can skip it to use the kit
PHILOSOPHY.md Design principles, target user, rejection list
architecture.md Components, data flow, the SDLC state machine, Collaborative Design Protocol, deps
execution-planes.md The four ways the kit runs agents, how they hand off, their trust models
decisions/ One ADR per file (NNNN-<slug>.md); supersession recorded in the Status line
specs/ Specs (SPEC-NNN-<slug>.md); also the live spec store the hooks detect
retro/ Per-cycle retrospectives (output of /kit:retro)
research/ Dated deep-scans that fed specific specs
_meta/BACKLOG.md Phased task backlog
For the full file listing including individual agent/hook/command names, run git ls-files or browse the repo on GitHub.
Debug mode. Set DWARVES_KIT_DEBUG=1 and every hook logs its decisions to stderr. Useful when a hook misbehaves or you want to understand why something was blocked or approved.
Hook logs. Hooks that make enforcement decisions append to ~/.claude/dwarves-kit/logs/ (anti-rationalization.log, safety-gate.log, spec-drift-guard.log, slop-cleaner.log). These build the eval corpus for future optimization.
Weekly scheduler. The kit ships ONE weekly LaunchAgent: a dispatcher over a declarative jobs list (session-intel digest, reflect propose, print-only unless BACKLOG_STAGE_AUTO=1 is set in ~/.config/kit-weekly/env; adding a job = one line, never a new plist). Consumer instantiates it: bash deploy/macos/install; runbook at deploy/macos/README.md.
Testing. bash tests/run-all.sh with no argument runs only the suites the diff touches, plus five always-on tree-wide lints (kit-contract, config-registry, no-personal-paths, no-scattered-ids, boundary-lint), about 1 to 2 minutes on a Mac. test-meta (about 200 seconds) joins only when the diff touches a path it reads (meta_input in bin/test-affected). --all is the full glob, 13 to 15 minutes, for CI and the nightly job only: it refuses (exit 64) unless CI is set or KIT_RUN_ALL=1. RUN_ALL_JOBS defaults to auto on macOS and 1 on Linux. Single suites still run on their own: bash tests/test-hooks.sh covers hook behavior, bash tests/test-meta.sh covers structural integrity (manifests, frontmatter, cross-links), bash tests/run-workflow.sh walks the CI workflow's steps locally and prints only the red ones.
CI. .github/workflows/test.yml runs on workflow_dispatch and on a v* tag push, nothing else. A push or a pull request starts no run, and merging a PR waits on no check. Run gh workflow run test before cutting a release tag. The local check is what catches a regression:
edit on a branch
|
v
bash tests/run-all.sh ....... diff-scoped suites + the five always-on lints
| (about 1-2 min on a Mac)
v
commit --> push --> PR --> merge no CI run fires anywhere on this line
|
v
cutting a release
|
+--> gh workflow run test ........ the macOS + Ubuntu matrix, on demand
+--> git push origin v<x.y.z> .... the same workflow, on the tag
|
v
KIT_RUN_ALL=1 bash tests/run-all.sh --all ... the full glob (13-15 min);
before the tag, never on a commit; without
CI or KIT_RUN_ALL=1 it refuses (exit 64)
External dependencies (install alongside, not bundled):
- Context Hub -
npm install -g @aisuite/chub - Context7 - MCP server for library docs
- codebase-memory-mcp - AST-level codebase indexing
Codebase index (opt-in). If codebase-memory-mcp is installed, the
codebase-index.sh SessionStart hook keeps the current repo's structural index fresh
in the background (built on the first session in a repo, incremental refresh after),
so /kit:spec and /kit:execute query the index (search_code, search_graph,
get_architecture, trace_path) instead of grepping, cutting orientation cost. With
the tool absent the hook no-ops and the kit greps exactly as before, nothing to
configure and nothing breaks. Enable: put the binary on PATH, run
claude mcp add --scope user codebase-memory -- codebase-memory-mcp, then re-run
install.sh.
- Prompt-type anti-rationalization hook (Haiku evaluation instead of grep patterns)
- /qa command with headless browser testing (requires Playwright)
- Intra-spec parallel task dispatch in /execute (cross-goal fan-out across specs already ships as /kit:dispatch; this is the deferred intra-spec case)
- Multi-harness packaging (Codex / Cursor / Gemini / OpenCode), deferred until real demand
See docs/CHANGELOG.md (root CHANGELOG.md is a thin pointer stub, SPEC-185). It's the source of truth; the README does not duplicate it.
Patterns extracted from:
- GSD v1 / get-shit-done - spec generation, the original planning-dir convention (since unified onto docs/specs/), 4 parallel researchers. Distinct from GSD v2 (gsd-build/gsd-2, npm
gsd-pi), a separate standalone agent on the Pi SDK referenced as an external execution runtime, not a pattern source - gstack - /office-hours, /review, /ship patterns; the /kit:ui-design loop shapes (brief schema, injection-wrap, accumulated-feedback)
- frontend-design - the external UI generator /kit:ui-design delegates to; its aesthetic-direction brief shape
- ui-ux-pro-max-skill - /kit:ui-design brief sub-shapes (token ladder, states matrix, a11y bars, voice); generator + tooling rejected per bash-over-binaries
- Trail of Bits - hook implementations, code quality rules, statusline pattern
- ClaudeKit - validation gate, adversarial review, session-state pattern
- Context Hub - API docs skill
- oh-my-claudecode - HUD/statusline, slop-cleaner pattern
- Claude-Code-Game-Studios - /start router, path-scoped rules, Collaborative Design Protocol
- Smart Ralph - fix-agent retry pattern (fail-fix-re-verify loop)
- mattpocock/skills grill-with-docs - /kit:grill mechanics: one-question-at-a-time with recommended answers, glossary/ADR write-as-you-go, the 3-criteria ADR bar, contradiction-first interviewing
- repository-harness - the FEATURE_INTAKE flag-count lane-classification model (A3: hard-gate + soft-count + auditable
explain), the@AGENTS.mdCLAUDE.md import shim (A1), adopt--dry-run/--refreshmodes (A2), and the decision-capture-at-reflect flow (A4-lite, advisory). The enforcing dual of our enforcing/CC-only kit; absorbed 2026-06-10, see docs/absorption/2026-06-10-repository-harness.md
MIT