Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion _meta/BACKLOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,7 +39,7 @@ Lane = the WORKFLOW.md risk tier (`tiny` / `normal` / `full`).
| ID-881 | bin/wrap merge retries transient GitHub failures #wrap #resilience | bin/wrap merge has no retry for a transient GitHub failure (502/503/GraphQL executing-query errors); a real refusal (not mergeable, not green, conflict, draft) must still fail immediately. Operator hand-rolled 3 shell retry loops for ~25 PR merges during a ~90min GitHub outage. Same failure class as memory note hand-rolled-merge-loop-instead-of-wrap-merge.md, a second occurrence, missing retry is the root cause. cmd_merge in lib/wrap/wrap.sh (~line 812, gh pr merge call ~line 917). lane=full. Source: board sweep session 2026-09-13 and 14. Goal draft: .claude/goals/wrap-merge-retry.md in the kit clone. | queued |
| ID-879 | wrap step 6 and the ledger flush must not commit on the checked-out default branch #wrap #ledger | Sessions commit LAB_LOG lines, ledger flushes and board rows straight onto local main because those writes land in the main checkout; main cannot be pushed, so 29 such commits plus 61 merge commits piled up on one Air by 2026-09-14 (ops-toolkit #2779 drained them). Wrap step 6, the learning-ledger flush and board capture should write on a housekeeping branch or leave the file uncommitted for the next feature PR. Companion guards: ops-toolkit pre-commit refuses commits on main; dotfiles sets pull.ff=only. Source: ops-toolkit session 2026-09-14. | queued |
| ID-882 | Dispatch keeps an Attempt state apart from the Task state: a dead or disconnected subagent is an unknown outcome, a grace window precedes LOST, and a result arriving inside the window commits with no second dispatch #kit #dispatch #u-mid #f-mid | Formalises the incident lesson resume-a-dead-subagent-never-respawn-on-its-branch. Shape from VoiceStudio backend/worker/lifecycle.py: TaskState and AttemptState enums with a legal-transition table, mark_disconnected starts grace, lose_attempt on expiry frees the task and excludes that worker, commit_result idempotent on task id so a resumed agent and its replacement cannot both land. Applies to kit:execute retries and megagoal-agent-drive. Contract tests named in ops-toolkit research/2026-09-14-voicestudio-long-running-orchestration.md | shipped [SPEC-290, proof docs/verification/dispatch-attempt-state.md] |
| ID-878 | session observe: a view that sizes the fixed context entry fee per turn #harness #kit #observability | Every agent turn re-reads a fixed preamble (skills, CLAUDE.md stack, repo memory index, agent roster) before any work. Measured by hand 2026-09-13 at ~96k tokens, about 38% of a 400-turn builder's whole bill. session observe already owns cost and burn over the same transcripts, so the entry-fee view belongs beside them rather than as a sibling script. Wants: per-component breakdown, a trend so ID-902..906 progress is visible, and a per-repo figure since the memory index and CLAUDE.md differ per checkout. lane=full. Source: ops-toolkit research/2026-09-13-token-burn-optimization.md | queued |
| ID-878 | session observe: a view that sizes the fixed context entry fee per turn #harness #kit #observability | Every agent turn re-reads a fixed preamble (skills, CLAUDE.md stack, repo memory index, agent roster) before any work. Measured by hand 2026-09-13 at ~96k tokens, about 38% of a 400-turn builder's whole bill. session observe already owns cost and burn over the same transcripts, so the entry-fee view belongs beside them rather than as a sibling script. Wants: per-component breakdown, a trend so ID-902..906 progress is visible, and a per-repo figure since the memory index and CLAUDE.md differ per checkout. lane=full. Source: ops-toolkit research/2026-09-13-token-burn-optimization.md | shipped [SPEC-289, proof lib/session/observe/docs/proof-of-done.md] |
| ID-877 | gate-ledger: record a whole lane plan in one call #kit #gate | Intent: one verb that takes a rid, a lane, and per-phase dispositions (ran, skipped, override, each with its reason) and writes every GATE line the ship-gate wants, instead of nine hand-typed record/override calls per prose-only PR. Precedent: nothing matched (bin/precedent, 2026-09-13). lane=full (touches lib/gate). Source: the skill-trigger-routing wrap. Goal draft: .claude/goals/gate-ledger-plan-record.md in the kit clone. | shipped [SPEC-287, proof docs/verification/gate-ledger-plan-record.md] |
| ID-874 | wrap apply: pull a checkout past a sibling session's dirty tracked files #wrap #git #u-mid #f-mid | Lane full (lane-classify: touches the pull path of a shared checkout). Twice this session an ff pull of the kit primary checkout was blocked by a sibling's dirty tracked files (staging log, run-all sample drift); the fix was done by hand both times: a NAMED stash of only the files git names as blocking, git pull --ff-only, pop BY REF, union-resolve the append-only log, checkout --ours for docs/verification/*/sample-*.md drift, drop only the named stash. Precedent: memory bare-stash-pop-takes-a-sibling-stash (prose, no mechanism). Home lib/wrap/wrap.sh apply pull path, behind a knob, default off. Goal draft: .claude/goals/wrap-pull-past-dirty.md. Filed by /kit:wrap 7b 2026-09-13. Rationale: recurring pull-blocked pain this session, and the manual fix is already sketched behind a knob. | shipped [#630, SPEC-286, proof docs/verification/wrap-pull-past-dirty.md] |
| ID-875 | wrap merge verifies the default branch after a squash merge #wrap #safety | After a squash merge bin/wrap merge does not confirm the default branch holds the PR head's tree, so a dropped commit goes unnoticed. Add the postcondition. Home lib/wrap/wrap.sh. Related memory note verify-main-after-squash-merge. lane=full. Source: ops-toolkit handoff 2026-09-13-wrap-build-lanes-and-leftovers. | shipped [#628, proof docs/verification/wrap-merge-tree-check.md] |
Expand Down
9 changes: 9 additions & 0 deletions docs/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,15 @@ All notable changes to dwarves-kit are documented here.
- Config surface (new section, additive, MINOR): `[intake]` with `url_ledger`, `verdicts`, `boards`, `notes`, all defaulting `""`. Every key resolves root-only, so a project `.kit.toml` cannot set them. An install that sets none of them keeps `intake gate` answering from this kit's own inventory and open pull requests alone.

### Added
- `session observe entry-fee [--days N] [--project SLUG-OR-NAME] [--top N] [--trend] [--json]`
sizes the fixed preamble every agent turn re-reads before any work. The per-session
total is measured (the first main-chain assistant turn's input + cache-creation +
cache-read); the per-component split is estimated at four characters per token over
the rendered preamble text and is labelled an estimate everywhere, with the remainder
shown as one `(unattributed)` row. Adds a per-repo median and, with `--trend`, ISO-week
medians so a preamble cleanup shows as a drop. Not part of `report`. `--project` now
also accepts a bare repo name, which matches every slug containing it, for `cost` and
`burn` too. Spec: `docs/specs/SPEC-289-observe-entry-fee.md`.
- Added a Codex plugin manifest and a Codex lifecycle adapter for the five hard guardrails: destructive-command safety, secret-file protection, ship completeness, commit format, and premature-completion protection. Claude Code keeps its existing manifest and settings. The shared secret policy now also denies Codex and Cloudflare credential files, and hard-hook logs omit command or subject text that could contain tokens. Codex requires explicit trust for each exact hook hash. This is not prompt/output DLP, and tool-hook coverage is limited to verified paths.
- `gate-ledger.sh plan-record <rid> <lane> [--ran|--skipped|--override <phase>[:<reason>]]...`
disposes every phase of a lane's plan in one call, instead of one `record`/`override` call per
Expand Down
4 changes: 2 additions & 2 deletions docs/FEATURES.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ GENERATED , do not hand-edit. Regenerate: `bash lib/registry/feature-registry.sh
| `/kit:design` | `[H/I]` | Opt-in interactive solution-design beat between /think and /spec. Explores 2-3 approaches one question at a time, holds for your approval p… | SPEC-003, SPEC-004, SPEC-005 +105 | test-command-emit-sweep.sh, test-command-triggers.sh, test-design-record.sh +21 |
| `/kit:devs-team` | `[H/I]` | Parallel multi-lens critique of a solution design (the active spec if present, else the decision brief). Dispatches 5 engineering lenses, m… | SPEC-016, SPEC-018, SPEC-019 +11 | test-gate-vocab-recording.sh, test-meta.sh, test-outcome-emit-sweep.sh |
| `/kit:dispatch` | `[H/I]` | Fire several disjoint VALIDATED specs concurrently, each in its own worktree, then converge. Cross-goal fan-out behind a disjointness gate … | SPEC-002, SPEC-016, SPEC-017 +79 | test-advisor-ledger-emit.sh, test-agent-effectiveness.sh, test-attempt-state.sh +31 |
| `/kit:docs` | `[H/I]` | Update all project documentation to match the current codebase. Cross-references the diff against every doc file and fixes drift. | SPEC-001, SPEC-002, SPEC-003 +191 | proof-loop-09-scenario-b.sh, run-all.sh, run-workflow.sh +69 |
| `/kit:docs` | `[H/I]` | Update all project documentation to match the current codebase. Cross-references the diff against every doc file and fixes drift. | SPEC-001, SPEC-002, SPEC-003 +192 | proof-loop-09-scenario-b.sh, run-all.sh, run-workflow.sh +69 |
| `/kit:draft-agent` | `[H/I]` | Meta-agent agent-builder. From a one-line description, generates a new subagent definition OR a mega-goal sub-goal file and (by default) in… | SPEC-089, SPEC-108, SPEC-139 +1 | test-agent-effectiveness.sh, test-command-emit-sweep.sh, test-meta-agent.sh +1 |
| `/kit:execute` | `[H/I]` | Autonomous spec execution with verification. Dispatches worker subagents per task, verifies each with task-verifier, retries fixable failur… | SPEC-001, SPEC-003, SPEC-004 +60 | test-break-it.sh, test-gate-vocab-recording.sh, test-hooks.sh +9 |
| `/kit:explain` | `[H/I]` | Turn a merged change into a literate-diff explainer a human READS to understand: background -> goal + intuition -> a prose-ordered diff -> … | SPEC-050, SPEC-060, SPEC-094 +18 | proof-loop-09-scenario-b.sh, test-boundary-lint.sh, test-command-emit-sweep.sh +11 |
Expand Down Expand Up @@ -99,7 +99,7 @@ GENERATED , do not hand-edit. Regenerate: `bash lib/registry/feature-registry.sh
| `get-api-docs` | `[I]` | Fetch curated API documentation using Context Hub (chub) before coding against any external API. Use when the task involves calling a third… | SPEC-285 | test-kit-foldin-hooks.sh |
| `loop-engineering` | `[I]` | Use when the user wants to design or add a new bounded loop to the kit's own SDLC orchestration ("let's build a loop", "design a new loop f… | SPEC-209, SPEC-222, SPEC-225 +4 | test-loop-engineering-contract.sh, test-research-arch-contract.sh |
| `memory-tidy` | `[I]` | Use when auditing, consolidating, or cleaning a repo's .claude/memory store, "dọn memory", "memory tidy", the biweekly memory audit, dupl… | SPEC-208, SPEC-225, SPEC-249 | test-memory-tidy-contract.sh, test-research-arch-contract.sh |
| `observe` | `[I]` | Query and render the kit's control plane from an agent session. Use when asked about the fleet's runs, gate verdicts, conformance, spend/to… | SPEC-033, SPEC-052, SPEC-128 +11 | test-bin-forwarders.sh, test-kit-contract.sh, test-orchestrate-wavefront.sh +1 |
| `observe` | `[I]` | Query and render the kit's control plane from an agent session. Use when asked about the fleet's runs, gate verdicts, conformance, spend/to… | SPEC-033, SPEC-052, SPEC-128 +12 | test-bin-forwarders.sh, test-kit-contract.sh, test-orchestrate-wavefront.sh +1 |
| `repo-hygiene` | `[I]` | Use for the whole-repo decay audit, "run the repo-hygiene loop", "what has rotted in this repo", "audit repo hygiene", "find the stale file… | SPEC-256 | test-repohygiene.sh |
| `skill-review` | `[I]` | Review and promote skill drafts that skill-curator staged from past sessions. Use when the user runs /skill-review, says "review my skill d… | SPEC-218, SPEC-285 | test-bin-forwarders.sh, test-kit-contract.sh |
| `stats` | `[I]` | Query or render the state of the scattered kit/tide/tg-cleanup/learned ledgers (the dwarves-kit gate/proof/telemetry corpus, tide file-move… | SPEC-001, SPEC-182, SPEC-192 +17 | proof-loop-09-scenario-b.sh, test-bin-forwarders.sh, test-config-stamp.sh +9 |
Expand Down
149 changes: 149 additions & 0 deletions docs/specs/SPEC-289-observe-entry-fee.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,149 @@
# SPEC-289: session observe entry-fee

**Status**: ready
Lane: full
**Owner module**: `lib/session/observe`
**Source**: backlog row ID-878; measurement `ops-toolkit/research/2026-09-13-token-burn-optimization.md`

## Problem

Every agent turn re-reads a fixed preamble before any work: the skill listing, the
CLAUDE.md stack, the repo memory index, the agent roster, the output style, the tool
schemas. A hand measurement on 2026-09-13 put that preamble near 96,000 tokens, about
38 percent of a 400-turn builder's whole bill.

That measurement is a snapshot taken by hand with a bytes-over-four rule of thumb. It
cannot say whether the number moved, which repo pays the most, or which component grew.
`session observe` already parses the same transcripts for `cost` and `burn`, so the
entry-fee view belongs beside them.

## Scope

One new view in the existing CLI. No new script, no new data source, no writes.

```
session observe entry-fee [--days N] [--project SLUG-OR-NAME] [--root DIR] [--top N] [--trend] [--json]
```

| Arg | Default | Meaning |
|---|---|---|
| `--days N` | 0 (all) | coarse file-mtime window, as every other view |
| `--project` | none | a project slug, or a bare repo name matched against every slug |
| `--top N` | 0 (all) | limit the per-repo and weekly tables |
| `--trend` | off | add a weekly median table, newest week first |
| `--json` | off | machine-readable output |

## Behaviour

### The measured total

One row per main session. The entry fee is the FIRST main-chain assistant turn's
`input_tokens + cache_creation_input_tokens + cache_read_input_tokens`. That sum is
exactly what the model read before it acted, so it is a measurement, not an estimate.

A transcript with no main-chain assistant turn contributes no row. That covers a
subagent's own transcript and a session that never got a reply. Both are excluded
rather than counted as a zero fee, which would drag every median down. A transcript
under a `subagents/` directory is excluded by path for the same reason.

### The estimated split

The transcript records each preamble block as an `attachment` entry whose `rendered`
field holds the text the model received. The view sizes each block at
`ENTRY_FEE_CHARS_PER_TOKEN` (4) characters per token and keys it by
`attachment.type` (`skill_listing`, `instructions`, `agent_listing_delta`,
`hook_additional_context`, ...). The remainder, `measured fee - sum of the sized
blocks`, is one `(unattributed)` row: the system prompt and the built-in tool schemas
reach the model but never appear in the transcript.

The split is labelled an estimate in the table header and in the JSON
(`components_estimated`). The same rule of thumb the hand measurement used is stated
alongside it.

Four characters per token overshoots on dense markdown and tables, so the sized blocks
can exceed the measured fee. That prints as an `(estimate over measured)` row carrying
the magnitude, never as a negative token count, and the component shares then read
above 100 percent, which is the estimate saying it broke here.

The measured total also covers the first user prompt and anything attached to it, so it
is the preamble plus turn one. The `(unattributed)` row absorbs that alongside the
system prompt and the tool schemas.

The split is reported for ONE real session, not as a per-component median. Medians of
separate components do not add up to the median fee, so a median-of-medians table
would not reconcile against a measured total. That session is the median of the
sessions that record a rendered preamble; older transcripts record none, and the
median across all sessions would often be one of those, reading as a 100 percent
unattributed fee.

### The per-repo and weekly figures

Per repo: one row per repo with the session count and the median measured fee, largest
median first, ties broken by session count. The memory index and the CLAUDE.md stack
differ per checkout, so this is the actionable cut.

A worktree gets its own project slug. The worktree convention is
`<repo>/.claude/worktrees/<name>`, so the slug carries a `--claude-worktrees-` marker
and the repo name sits before it. Rows key on the part before that marker, which folds
a repo's worktrees into one row. Without the fold, a one-session worktree slug ranks
above the many-session checkout of the same repo, which reads as a comparison when it
is one codebase twice.

`--trend`: ISO-week buckets of the median measured fee, newest week first, so a
reduction from a cleanup lands as a visible drop.

### Project resolution

`--project` takes the exact slug directory when it exists. Otherwise it matches every
slug containing the given string, so a bare repo name resolves without the full
cwd-derived slug. One repo's worktrees each get their own slug and a per-repo figure
wants them together, so all matches are walked. This resolution is shared with the
other views through `project_roots()`.

A substring matching more than one slug prints the resolved list to stderr. Before the
fallback existed, a wrong `--project` walked nothing and the empty output said so; a
silent multi-match would instead merge unrelated repos into one plausible figure.

## Non-goals

- Not part of `report`. `report` is the weekly digest of behaviour; the entry fee is a
standing cost measurement, and its per-repo table would double the digest's length.
- No per-component token measurement. The API reports one usage block per turn, so a
per-component count does not exist in the data.
- No subagent entry fee. A subagent pays the same preamble, but its transcript has no
main-chain turn, and counting it would mix two populations in one median.
- No writes. Read-only, like every other view.

## Verification (acceptance criteria)

Exercised by `tests/smoke.sh` against `tests/fixtures/entryfee/` (proj-alpha with fees
1000, 2000, 3000; proj-beta with fee 500; a `subagents/` transcript at 99999; a
sidechain-only transcript):

62. header reports 4 sessions, 2 projects, median 2000.
63. the split sizes `instructions` (800 chars) at 200 and `skill_listing` (400) at 100.
64. `(unattributed)` is the measured fee minus the sized blocks (2000 - 350 = 1650).
65. negative control: the `subagents/` transcript is excluded (its 99999 fee absent).
66. negative control: the sidechain-only transcript contributes no session.
67. negative control: a later, larger turn does not replace the first-turn fee.
68. per repo: proj-beta is its own row at 1 session, median 500.
69. `--trend`: 2026-W37 (3, 2000) prints above 2026-W36 (1, 1000).
70. negative control: without `--trend` the weekly table is absent.
71. `--project alpha` resolves the bare name to proj-alpha only.
72. `--json` is valid, carries median_fee 2000 and the estimate flag.
73. negative control: `report` prints no entry-fee section.

Plus, against `tests/entryfee-edge/` (an overshooting session, an untrusted-field
session, a repo with one worktree slug):

74. an estimate above the measured fee prints as `(estimate over measured)`, never a
negative token count.
75. the overshooting component's share reads above 100 percent.
76. untrusted fields (numeric timestamp, dict attachment type, numeric rendered content,
non-dict message) do not crash the scan; the valid turn's fee is still measured.
77. negative control: the dict-typed attachment contributes no component row.
78. a worktree slug folds into its repo row (one row, two sessions).
79. a multi-slug `--project` match is announced on stderr.

Plus a real run over the live transcripts, recorded in
`lib/session/observe/docs/verification/entry-fee.md`.
Loading
Loading