From d637dcd7f4705c194d05f8ba9cc5fe53bc877cb1 Mon Sep 17 00:00:00 2001 From: Han Ngo Date: Wed, 16 Sep 2026 18:06:33 +0700 Subject: [PATCH 1/5] feat(session-observe): size the fixed context entry fee per turn Every agent turn re-reads a fixed preamble before any work: the skill listing, the CLAUDE.md stack, the repo memory index, the agent roster, the output style, the tool schemas. A hand measurement put it near 96k tokens, about 38 percent of a 400-turn builder's whole bill, but a hand snapshot cannot say whether the number moved or which repo pays most. The new `entry-fee` view measures the total and estimates the split. The total is the first main-chain assistant turn's input plus cache creation plus cache read, which is exactly what the model read before it acted. The split sizes each preamble attachment from its rendered text at four characters per token and labels itself an estimate; the gap against the measured total prints as one (unattributed) row for the system prompt and the built-in tool schemas. Per-repo rows carry the median measured fee, since the memory index and the CLAUDE.md stack differ per checkout. --trend adds ISO-week medians, newest first, so a preamble cleanup lands as a visible drop. The split comes from one real session, not a per-component median: component medians do not sum to the median fee, so the table would not reconcile against a measured total. --project now also resolves a bare repo name against every slug that contains it, shared through project_roots() with cost and burn, because one repo's worktrees each own a separate slug. --- docs/CHANGELOG.md | 9 + docs/specs/SPEC-289-observe-entry-fee.md | 117 ++++++++ lib/session/observe/README.md | 3 + lib/session/observe/SPEC.md | 10 +- lib/session/observe/bin/session-observe | 276 ++++++++++++++++-- .../docs/implementation-notes/entry-fee.md | 60 ++++ .../entryfee/proj-alpha/aaaa1111.jsonl | 3 + .../entryfee/proj-alpha/aaaa2222.jsonl | 5 + .../aaaa2222/subagents/worker.jsonl | 2 + .../entryfee/proj-alpha/aaaa3333.jsonl | 2 + .../entryfee/proj-alpha/side9999.jsonl | 2 + .../entryfee/proj-beta/bbbb1111.jsonl | 2 + lib/session/observe/tests/smoke.sh | 57 ++++ 13 files changed, 515 insertions(+), 33 deletions(-) create mode 100644 docs/specs/SPEC-289-observe-entry-fee.md create mode 100644 lib/session/observe/docs/implementation-notes/entry-fee.md create mode 100644 lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa1111.jsonl create mode 100644 lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa2222.jsonl create mode 100644 lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa2222/subagents/worker.jsonl create mode 100644 lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa3333.jsonl create mode 100644 lib/session/observe/tests/fixtures/entryfee/proj-alpha/side9999.jsonl create mode 100644 lib/session/observe/tests/fixtures/entryfee/proj-beta/bbbb1111.jsonl diff --git a/docs/CHANGELOG.md b/docs/CHANGELOG.md index f656287b..faa7dfa4 100644 --- a/docs/CHANGELOG.md +++ b/docs/CHANGELOG.md @@ -12,6 +12,15 @@ All notable changes to dwarves-kit are documented here. - Config surface (new section, additive, MINOR): `[intake]` with `url_ledger`, `verdicts`, `boards`, `notes`, all defaulting `""`. Every key resolves root-only, so a project `.kit.toml` cannot set them. An install that sets none of them keeps `intake gate` answering from this kit's own inventory and open pull requests alone. ### Added +- `session observe entry-fee [--days N] [--project SLUG-OR-NAME] [--top N] [--trend] [--json]` + sizes the fixed preamble every agent turn re-reads before any work. The per-session + total is measured (the first main-chain assistant turn's input + cache-creation + + cache-read); the per-component split is estimated at four characters per token over + the rendered preamble text and is labelled an estimate everywhere, with the remainder + shown as one `(unattributed)` row. Adds a per-repo median and, with `--trend`, ISO-week + medians so a preamble cleanup shows as a drop. Not part of `report`. `--project` now + also accepts a bare repo name, which matches every slug containing it, for `cost` and + `burn` too. Spec: `docs/specs/SPEC-289-observe-entry-fee.md`. - Added a Codex plugin manifest and a Codex lifecycle adapter for the five hard guardrails: destructive-command safety, secret-file protection, ship completeness, commit format, and premature-completion protection. Claude Code keeps its existing manifest and settings. The shared secret policy now also denies Codex and Cloudflare credential files, and hard-hook logs omit command or subject text that could contain tokens. Codex requires explicit trust for each exact hook hash. This is not prompt/output DLP, and tool-hook coverage is limited to verified paths. - `gate-ledger.sh plan-record [--ran|--skipped|--override [:]]...` disposes every phase of a lane's plan in one call, instead of one `record`/`override` call per diff --git a/docs/specs/SPEC-289-observe-entry-fee.md b/docs/specs/SPEC-289-observe-entry-fee.md new file mode 100644 index 00000000..bd4e9e0f --- /dev/null +++ b/docs/specs/SPEC-289-observe-entry-fee.md @@ -0,0 +1,117 @@ +# SPEC-289: session observe entry-fee + +**Status**: ready +**Lane**: full +**Owner module**: `lib/session/observe` +**Source**: backlog row ID-878; measurement `ops-toolkit/research/2026-09-13-token-burn-optimization.md` + +## Problem + +Every agent turn re-reads a fixed preamble before any work: the skill listing, the +CLAUDE.md stack, the repo memory index, the agent roster, the output style, the tool +schemas. A hand measurement on 2026-09-13 put that preamble near 96,000 tokens, about +38 percent of a 400-turn builder's whole bill. + +That measurement is a snapshot taken by hand with a bytes-over-four rule of thumb. It +cannot say whether the number moved, which repo pays the most, or which component grew. +`session observe` already parses the same transcripts for `cost` and `burn`, so the +entry-fee view belongs beside them. + +## Scope + +One new view in the existing CLI. No new script, no new data source, no writes. + +``` +session observe entry-fee [--days N] [--project SLUG-OR-NAME] [--root DIR] [--top N] [--trend] [--json] +``` + +| Arg | Default | Meaning | +|---|---|---| +| `--days N` | 0 (all) | coarse file-mtime window, as every other view | +| `--project` | none | a project slug, or a bare repo name matched against every slug | +| `--top N` | 0 (all) | limit the per-repo and weekly tables | +| `--trend` | off | add a weekly median table, newest week first | +| `--json` | off | machine-readable output | + +## Behaviour + +### The measured total + +One row per main session. The entry fee is the FIRST main-chain assistant turn's +`input_tokens + cache_creation_input_tokens + cache_read_input_tokens`. That sum is +exactly what the model read before it acted, so it is a measurement, not an estimate. + +A transcript with no main-chain assistant turn contributes no row. That covers a +subagent's own transcript and a session that never got a reply. Both are excluded +rather than counted as a zero fee, which would drag every median down. A transcript +under a `subagents/` directory is excluded by path for the same reason. + +### The estimated split + +The transcript records each preamble block as an `attachment` entry whose `rendered` +field holds the text the model received. The view sizes each block at +`ENTRY_FEE_CHARS_PER_TOKEN` (4) characters per token and keys it by +`attachment.type` (`skill_listing`, `instructions`, `agent_listing_delta`, +`hook_additional_context`, ...). The remainder, `measured fee - sum of the sized +blocks`, is one `(unattributed)` row: the system prompt and the built-in tool schemas +reach the model but never appear in the transcript. + +The split is labelled an estimate in the table header and in the JSON +(`components_estimated`). The same rule of thumb the hand measurement used is stated +alongside it. + +The split is reported for ONE real session, not as a per-component median. Medians of +separate components do not add up to the median fee, so a median-of-medians table +would not reconcile against a measured total. That session is the median of the +sessions that record a rendered preamble; older transcripts record none, and the +median across all sessions would often be one of those, reading as a 100 percent +unattributed fee. + +### The per-repo and weekly figures + +Per repo: one row per project slug with the session count and the median measured fee, +largest median first. The memory index and the CLAUDE.md stack differ per checkout, so +this is the actionable cut. + +`--trend`: ISO-week buckets of the median measured fee, newest week first, so a +reduction from a cleanup lands as a visible drop. + +### Project resolution + +`--project` takes the exact slug directory when it exists. Otherwise it matches every +slug containing the given string, so a bare repo name resolves without the full +cwd-derived slug. One repo's worktrees each get their own slug and a per-repo figure +wants them together, so all matches are walked. This resolution is shared with the +other views through `project_roots()`. + +## Non-goals + +- Not part of `report`. `report` is the weekly digest of behaviour; the entry fee is a + standing cost measurement, and its per-repo table would double the digest's length. +- No per-component token measurement. The API reports one usage block per turn, so a + per-component count does not exist in the data. +- No subagent entry fee. A subagent pays the same preamble, but its transcript has no + main-chain turn, and counting it would mix two populations in one median. +- No writes. Read-only, like every other view. + +## Verification (acceptance criteria) + +Exercised by `tests/smoke.sh` against `tests/fixtures/entryfee/` (proj-alpha with fees +1000, 2000, 3000; proj-beta with fee 500; a `subagents/` transcript at 99999; a +sidechain-only transcript): + +62. header reports 4 sessions, 2 projects, median 2000. +63. the split sizes `instructions` (800 chars) at 200 and `skill_listing` (400) at 100. +64. `(unattributed)` is the measured fee minus the sized blocks (2000 - 350 = 1650). +65. negative control: the `subagents/` transcript is excluded (its 99999 fee absent). +66. negative control: the sidechain-only transcript contributes no session. +67. negative control: a later, larger turn does not replace the first-turn fee. +68. per repo: proj-beta is its own row at 1 session, median 500. +69. `--trend`: 2026-W37 (3, 2000) prints above 2026-W36 (1, 1000). +70. negative control: without `--trend` the weekly table is absent. +71. `--project alpha` resolves the bare name to proj-alpha only. +72. `--json` is valid, carries median_fee 2000 and the estimate flag. +73. negative control: `report` prints no entry-fee section. + +Plus a real run over the live transcripts, recorded in +`lib/session/observe/docs/verification/entry-fee.md`. diff --git a/lib/session/observe/README.md b/lib/session/observe/README.md index 470190d8..2d59f004 100644 --- a/lib/session/observe/README.md +++ b/lib/session/observe/README.md @@ -26,6 +26,7 @@ session-observe friction --days 7 # thrash / permission-friction / contex session-observe sessions --days 7 # archetype mix / circadian (by hour) / interruption rate session-observe cost --days 7 # tokens by model + cache economics + $ estimate session-observe burn --since 60 # live per-session token burn, ranked (not part of report) +session-observe entry-fee --days 14 --trend # the fixed preamble every turn re-reads (not part of report) session-observe report --days 7 --json # machine-readable, for vps-mon ingest ``` @@ -56,6 +57,8 @@ Each transcript entry already carries `hookInfos: [{command, durationMs}]`, `hoo - **cost**: from the assistant `message.usage` block (input / output / cache-read / cache-create tokens) + `message.model`. **tokens-by-model** + a **$ estimate** from an embedded dated `PRICING` table (Max plan is flat-rate, so this is *attribution*, not a bill; unknown model families like `fable` count tokens but show `?`). **cache economics** = cache-read / (read + create) hit ratio, the biggest cost lever. (cost-per-merged-PR is deferred, see proof/impl-notes: transcripts are cross-repo and this repo squash-merges, so there is no clean per-PR attribution.) - **burn**: which session is burning tokens *right now*, not over the week. One row per top-level session inside the last `--since` minutes (default 60); a subagent transcript (`/subagents/*.jsonl`) rolls its tokens into its parent row and counts toward `subs`. Usage is deduplicated by `(message.id, requestId)` (a streamed chunk repeats the same usage block). `ctx` is the live context size: input + cache-creation + cache-read of the last main-chain (non-sidechain) assistant usage, regardless of the window. Ranked by `input + cache_create + cache_read/10 + output*5` (attribution, not a bill). The `pid` column maps through the live `~/.claude/sessions/.json` files. `model share` splits the row's burn by model, biggest first, so an expensive fan-out is legible without reading the transcripts: a lead that dispatched its workers on Opus reads `opus 91%`. The split is priced through `model_cost()`, not counted, because equal Opus and Sonnet token counts are not equal spend (Opus lists at five times Sonnet on every axis, so equal tokens read `opus 83% sonnet 17%`). One model with an unknown family (fable) leaves the row unpriceable, so the whole row falls back to the same list-price token weight the ranking uses and a trailing `~` marks it. Not part of `report`. +- **entry-fee**: the fixed preamble every turn re-reads before any work (skill listing, CLAUDE.md stack, repo memory index, agent roster, output style, tool schemas). The per-session TOTAL is **measured**: the first main-chain assistant turn's input + cache-creation + cache-read, which is exactly what the model read before it acted. The per-component SPLIT is **estimated**: each preamble `attachment` entry is sized from its `rendered` text at 4 characters per token and keyed by `attachment.type`, because the API reports one usage block per turn and no per-component count exists in the data. The remainder against the measured total is one `(unattributed)` row (system prompt + built-in tool schemas, which reach the model but never the transcript). The split is reported for ONE real session rather than as a per-component median, since medians of separate components would not add up to the measured total; that session is the median of the sessions whose transcript records a rendered preamble. Per-repo rows carry the median measured fee (the memory index and CLAUDE.md stack differ per checkout), and `--trend` adds ISO-week medians newest-first so a cleanup shows as a drop. A subagent's own transcript has no main-chain turn and is excluded. Not part of `report`. + ## Output (real run, abbreviated) ``` diff --git a/lib/session/observe/SPEC.md b/lib/session/observe/SPEC.md index c737037a..28dea8b1 100644 --- a/lib/session/observe/SPEC.md +++ b/lib/session/observe/SPEC.md @@ -24,14 +24,15 @@ This is sub-goal 01 of the `cc-elevation` mega-goal (self-observability axis). S ## CLI contract ``` -cc-observe [--file F | --project=SLUG] [--root DIR] [--days N] [--top N] [--json] +cc-observe [--file F | --project=SLUG] [--root DIR] [--days N] [--top N] [--trend] [--json] ``` | Arg | Default | Meaning | |---|---|---| -| `` | required | `skills`, `tools`, `hooks`, `subagents`, `friction`, `sessions`, `cost`, or `report` (all of them) | +| `` | required | `skills`, `tools`, `hooks`, `subagents`, `friction`, `sessions`, `cost`, `burn`, `entry-fee`, or `report` (every view except `burn` and `entry-fee`) | | `--file F` | none | parse a single transcript (used by tests) | -| `--project=SLUG` | none | one project dir under the root. Slugs start with `-` (cwd-derived), so the **equals form is required**. | +| `--project=SLUG` | none | one project dir under the root. Slugs start with `-` (cwd-derived), so the **equals form is required**. A bare repo name that is not itself a slug matches every slug containing it, so a repo's worktrees aggregate. | +| `--trend` | false | (`entry-fee`) add a weekly median table, newest week first | | `--root DIR` | `~/.claude/projects` | transcripts root (all projects) | | `--days N` | 0 (all) | only files modified within N days (coarse mtime window) | | `--top N` | 0 (all) | limit to top N rows | @@ -52,6 +53,7 @@ The intended recurring use is `cc-observe report --days 7 --json` (the weekly di - friction: `Edit`/`Write`/`MultiEdit` -> per-session edit count by `file_path`, folded into `thrash` at end of each transcript (`>= THRASH_MIN`); `tool_result` content matching a `PERM_MARKERS` string -> a permission-friction event attributed to the tool via `tool_use_id` (Bash labeled by command); `isCompactSummary` entry -> a compaction for that day; a Skill's errored `tool_result` -> a skill mis-fire (skill-precision). - sessions: per transcript, track first/last `timestamp` (wall-clock), tool-use count, prompt-turn count -> `classify_session` buckets it (quick/standard/deep/marathon/automation, thresholds in `ARCH`); sidechain transcripts are skipped (not Han's sessions). Prompt-turns + tool-uses are bucketed by UTC hour (`circadian`). A prompt turn whose text holds `INTERRUPT_MARK` is an interruption. - cost: `message.usage` (input/output/cache-read/cache-create) summed per `message.model`; `model_cost` applies the dated `PRICING` table by family substring (unknown family -> tokens counted, `$` shown as `?`). Cache-hit = cache-read / (read + create). cost-per-merged-PR is out of scope (see Non-goals). + - entry-fee: per transcript, the FIRST main-chain assistant turn's `input + cache_creation + cache_read` is the session's MEASURED entry fee (the fixed preamble every turn re-reads). Preamble `attachment` entries are sized from their `rendered` text at 4 chars per token and keyed by `attachment.type`, an ESTIMATE labelled as one; the remainder against the measured fee is `(unattributed)` (system prompt + tool schemas, absent from the transcript). Transcripts with no main-chain assistant turn (a subagent's own run, a session with no reply) and transcripts under a `subagents/` dir contribute no row. Per-repo and weekly (`--trend`) tables carry the median measured fee. Full contract: `docs/specs/SPEC-289-observe-entry-fee.md`. 3. Emit the requested view(s) as aligned tables, or one JSON object with `--json`. **Hook labels**: script hooks collapse to their basename (`slop-cleaner.sh`); inline `echo` guard hooks key on a short command hash (`inline-echo:ab12cd34`); other commands use their first token. Known limitation: long-text inline/condition hooks (e.g. the `/goal` Stop-hook, whose command is the goal text) fragment by first word. Script hooks, the ones with actionable latency, label cleanly. @@ -61,6 +63,8 @@ The intended recurring use is `cc-observe report --days 7 --json` (the weekly di - No instrumentation / wrapper / daemon. Read-only over existing transcripts. - The `cost` view's `$` is an **attribution estimate** from a static dated `PRICING` table, NOT a bill (Max plan is flat-rate). No live pricing fetch. - **No cost-per-merged-PR.** It was in SG-03's outcome but has no clean data path: transcripts span all repos while merges are per-repo, and ops-toolkit squash-merges (so `git log --merges` finds ~0). Deferred to NOTES proposed-additions; needs PR data (gh), not transcript data. +- The `entry-fee` view's component split is an **estimate** (4 chars per token over the rendered preamble text). The API reports one usage block per turn, so a per-component token count does not exist in the data. Only the per-session total is measured. +- Neither `burn` nor `entry-fee` is part of `report`. `report` is the weekly digest of behaviour; `burn` is a live check and `entry-fee` is a standing cost measurement. - No live dashboard. `--json` feeds vps-mon; rendering lives there. - No writes anywhere. This tool only reads. diff --git a/lib/session/observe/bin/session-observe b/lib/session/observe/bin/session-observe index 36ec2993..bf442ad8 100755 --- a/lib/session/observe/bin/session-observe +++ b/lib/session/observe/bin/session-observe @@ -11,6 +11,11 @@ Three views, all derived from the JSONL transcripts under ~/.claude/projects/: pressure (compactions), skill mis-fires sessions - session shape: archetype mix, circadian (by hour), interruption rate cost - tokens by model + cache economics + $ estimate (attribution, not a bill) + entry-fee - the fixed preamble every turn re-reads before any work: a measured + per-session total, an estimated per-component split, a per-repo + figure, and with --trend a weekly median so a reduction is visible; + not part of `report` (report is the weekly digest of behaviour, + entry-fee is a standing cost measurement) burn - live per-session token burn in the last --since minutes (default 60), ranked by burn rate; subagent transcripts roll into their parent session; `model share` splits each row by model (priced, so an Opus fan-out is @@ -39,7 +44,7 @@ import re import sys import time from collections import Counter, defaultdict -from datetime import datetime +from datetime import date, datetime def _repo_root(): @@ -115,28 +120,54 @@ def classify_session(dur_min, tools, turns): return "standard" +def project_roots(root, project): + """Directories to walk for --project. + + The exact slug dir when it exists (the historical contract). Otherwise every + slug whose name contains the given string, so a bare repo name resolves + without the full cwd-derived slug. One repo's worktrees each get their own + slug, and a per-repo figure wants them together, so this returns all matches + rather than picking one. No match -> the literal join, which walks nothing. + """ + exact = os.path.join(root, project) + if os.path.isdir(exact): + return [exact] + try: + hits = sorted( + os.path.join(root, n) for n in os.listdir(root) + if project in n and os.path.isdir(os.path.join(root, n)) + ) + except OSError: + hits = [] + return hits or [exact] + + +def _walk_transcripts(args): + """Every *.jsonl under the resolved roots, in a stable order.""" + root = args.root or DEFAULT_ROOT + roots = project_roots(root, args.project) if args.project else [root] + for r in roots: + for dirpath, _dirs, files in os.walk(r): + for fn in sorted(files): + if fn.endswith(".jsonl"): + yield os.path.join(dirpath, fn) + + def iter_files(args): """Resolve the set of transcript files to read from the CLI args.""" if args.file: yield args.file return - root = args.root or DEFAULT_ROOT - if args.project: - root = os.path.join(root, args.project) cutoff = time.time() - args.days * 86400 if args.days else None - for dirpath, _dirs, files in os.walk(root): - for fn in files: - if not fn.endswith(".jsonl"): - continue - path = os.path.join(dirpath, fn) - # Coarse window: skip files whose last write predates the cutoff. - if cutoff is not None: - try: - if os.path.getmtime(path) < cutoff: - continue - except OSError: + for path in _walk_transcripts(args): + # Coarse window: skip files whose last write predates the cutoff. + if cutoff is not None: + try: + if os.path.getmtime(path) < cutoff: continue - yield path + except OSError: + continue + yield path def iter_entries(path): @@ -546,20 +577,13 @@ def _burn_files(args, cutoff_epoch): if args.file: yield args.file return - root = args.root or DEFAULT_ROOT - if args.project: - root = os.path.join(root, args.project) - for dirpath, _dirs, files in os.walk(root): - for fn in files: - if not fn.endswith(".jsonl"): - continue - path = os.path.join(dirpath, fn) - try: - if os.path.getmtime(path) < cutoff_epoch: - continue - except OSError: + for path in _walk_transcripts(args): + try: + if os.path.getmtime(path) < cutoff_epoch: continue - yield path + except OSError: + continue + yield path # type -> field name, for the three known title-carrying entry shapes. @@ -799,6 +823,193 @@ def run_burn(args): ) +# --- entry fee ------------------------------------------------------------- +# The fixed preamble every turn re-reads before the agent does any work: the +# skill listing, the CLAUDE.md stack, the repo memory index, the agent roster, +# the output style, the tool schemas. +# +# The TOTAL is measured. It is the first main-chain assistant turn's +# input + cache-creation + cache-read, which is exactly what the model read +# before it acted. The SPLIT is estimated: the transcript carries no per-component +# token count, so each preamble attachment is sized from its rendered text at +# ENTRY_FEE_CHARS_PER_TOKEN characters per token, the same rule of thumb the +# 2026-09-13 hand measurement used. Every output labels the split an estimate. +ENTRY_FEE_CHARS_PER_TOKEN = 4 +# Fee minus the sized components: the system prompt and the built-in tool +# schemas reach the model but never appear in the transcript. +ENTRY_FEE_UNATTRIBUTED = "(unattributed)" + + +def _rendered_chars(entry): + """Characters of rendered preamble text one attachment entry contributes.""" + r = entry.get("rendered") + if isinstance(r, str): + return len(r) + if not isinstance(r, list): + return 0 + return sum(len(b.get("content") or "") for b in r if isinstance(b, dict)) + + +def entry_fee_session(path): + """One transcript -> its entry fee + component estimate, or None. + + None means the file has no main-chain assistant turn: a subagent's own + transcript, or a session that never got a reply. Both are excluded rather + than counted as a zero fee. + """ + comps = Counter() + first_ts = None + for entry in iter_entries(path): + if not isinstance(entry, dict): + continue # malformed transcript line (untrusted input) + ts = entry.get("timestamp") + if first_ts is None and isinstance(ts, str): + first_ts = ts + if entry.get("type") == "attachment": + att = entry.get("attachment") + label = att.get("type") if isinstance(att, dict) else None + tokens = _rendered_chars(entry) // ENTRY_FEE_CHARS_PER_TOKEN + if label and tokens: + comps[label] += tokens + continue + if entry.get("type") != "assistant" or entry.get("isSidechain"): + continue + msg = entry.get("message") + u = msg.get("usage") if isinstance(msg, dict) else None + if not isinstance(u, dict): + continue + fee = ((u.get("input_tokens") or 0) + + (u.get("cache_creation_input_tokens") or 0) + + (u.get("cache_read_input_tokens") or 0)) + day = (first_ts or ts or "")[:10] + return {"fee": fee, "components": comps, "day": day, "model": msg.get("model") or "?"} + return None + + +def entry_fee_collect(args): + """One row per main session, newest-first file order irrelevant.""" + rows = [] + for path in iter_files(args): + if os.sep + "subagents" + os.sep in path: + continue # a subagent's own transcript, not a session someone started + s = entry_fee_session(path) + if s: + s["project"] = os.path.basename(os.path.dirname(path)) + rows.append(s) + return rows + + +def entry_fee_median(sessions): + """The session at the median fee, or None for an empty set.""" + if not sessions: + return None + ranked = sorted(sessions, key=lambda s: s["fee"]) + i = min(len(ranked) - 1, int(round(0.5 * (len(ranked) - 1)))) + return ranked[i] + + +def entry_fee_split_session(sessions): + """The session whose component split represents the set. + + Only transcripts that record the rendered preamble can be split at all, so + the split comes from the median of THOSE sessions rather than the median of + every session, which would often be one that records nothing and read as a + 100% unattributed fee. Returns (session, how many sessions can be split). + + The breakdown is one real session's, not a per-component median: medians of + separate components do not add up to the median fee, so a median-of-medians + table would not reconcile against the measured total. + """ + splittable = [s for s in sessions if s["components"]] + return entry_fee_median(splittable), len(splittable) + + +def entry_fee_component_rows(sess): + """Component split of one session, biggest first, with the remainder last.""" + if not sess: + return [] + fee = sess["fee"] + def share(n): + return f"{100 * n / fee:.0f}%" if fee else "-" + rows = [[name, n, share(n)] for name, n in sess["components"].most_common()] + rest = fee - sum(sess["components"].values()) + rows.append([ENTRY_FEE_UNATTRIBUTED, rest, share(rest)]) + return rows + + +def short_project(slug, n=46): + """Trim a long project slug from the LEFT: the slug is a cwd path with the + repo name at the end, so the tail is the part worth keeping.""" + return ("..." + slug[-(n - 3):]) if len(slug) > n else slug + + +def entry_fee_project_rows(sessions, top): + """Per-repo median fee, every project (the caller slices for display). + The memory index and CLAUDE.md stack differ per checkout, so the per-repo + figure is the actionable one.""" + by = defaultdict(list) + for s in sessions: + by[s["project"]].append(s["fee"]) + rows = [[short_project(p), len(v), pctl(v, 50)] for p, v in by.items()] + rows.sort(key=lambda r: -r[2]) + return rows + + +def iso_week(day): + """`2026-09-08` -> `2026-W37`. Unparseable -> `?`.""" + try: + y, w, _ = date.fromisoformat(day).isocalendar() + except ValueError: + return "?" + return f"{y}-W{w:02d}" + + +def entry_fee_week_rows(sessions): + """Weekly median fee, newest week first, so a reduction shows as a drop.""" + by = defaultdict(list) + for s in sessions: + by[iso_week(s["day"])].append(s["fee"]) + return [[w, len(by[w]), pctl(by[w], 50)] for w in sorted(by, reverse=True)] + + +def run_entry_fee(args): + """The `entry-fee` view. A separate I/O path from emit(): one row per + session rather than a model+day aggregation, and never part of `report`.""" + sessions = entry_fee_collect(args) + med_sess = entry_fee_median(sessions) + split_sess, splittable = entry_fee_split_session(sessions) + comp = entry_fee_component_rows(split_sess) + projects = entry_fee_project_rows(sessions, args.top) + weeks = entry_fee_week_rows(sessions) if args.trend else [] + med = med_sess["fee"] if med_sess else 0 + + if args.json: + print(json.dumps({ + "sessions": len(sessions), + "median_fee": med, + "chars_per_token": ENTRY_FEE_CHARS_PER_TOKEN, + "components_estimated": True, + "split_sessions": splittable, + "split_session_fee": split_sess["fee"] if split_sess else 0, + "split_components": [{"component": r[0], "est_tokens": r[1]} for r in comp], + "by_project": [{"project": r[0], "sessions": r[1], "median_fee": r[2]} for r in projects], + "by_week": [{"week": r[0], "sessions": r[1], "median_fee": r[2]} for r in weeks], + }, indent=2)) + return + + split_fee = split_sess["fee"] if split_sess else 0 + print(f"# entry-fee ({len(sessions)} sessions, {len(projects)} projects; median {med} tokens re-read per turn, measured)") + print(f" component split of one {split_fee}-token session ({splittable} sessions record the rendered preamble;") + print(f" the split is ESTIMATED at {ENTRY_FEE_CHARS_PER_TOKEN} chars/token, only the totals are measured):") + print_table(["component", "est-tokens", "share"], comp) + print(" per repo (median measured fee):") + print_table(["project", "sessions", "median-fee"], projects[: args.top] if args.top else projects) + if args.trend: + print(" weekly trend (median measured fee):") + print_table(["week", "sessions", "median-fee"], weeks[: args.top] if args.top else weeks) + print() + + def emit(args, data): if args.json: out = { @@ -892,7 +1103,7 @@ def emit(args, data): def main(argv=None): p = argparse.ArgumentParser(prog="session-observe", description="Report Claude Code skill/tool/hook usage from session transcripts.") - p.add_argument("cmd", choices=["skills", "tools", "hooks", "subagents", "friction", "sessions", "cost", "burn", "report"], help="which view") + p.add_argument("cmd", choices=["skills", "tools", "hooks", "subagents", "friction", "sessions", "cost", "burn", "entry-fee", "report"], help="which view") src = p.add_mutually_exclusive_group() src.add_argument("--file", help="parse a single transcript file (used by tests)") src.add_argument("--project", help="project slug under ~/.claude/projects/") @@ -901,6 +1112,7 @@ def main(argv=None): p.add_argument("--top", type=int, default=0, help="limit to top N rows (0 = all)") p.add_argument("--latency", action="store_true", help="(skills) add a per-skill total wall-time table, ranked by summed ms") p.add_argument("--since", type=int, default=60, help="(burn only) window in minutes (default 60)") + p.add_argument("--trend", action="store_true", help="(entry-fee) add a weekly median table, newest week first") p.add_argument("--json", action="store_true", help="machine-readable output (for vps-mon ingest)") args = p.parse_args(argv) @@ -912,6 +1124,10 @@ def main(argv=None): run_burn(args) return 0 + if args.cmd == "entry-fee": + run_entry_fee(args) + return 0 + data = collect(args) emit(args, data) return 0 diff --git a/lib/session/observe/docs/implementation-notes/entry-fee.md b/lib/session/observe/docs/implementation-notes/entry-fee.md new file mode 100644 index 00000000..562b8043 --- /dev/null +++ b/lib/session/observe/docs/implementation-notes/entry-fee.md @@ -0,0 +1,60 @@ +# Implementation notes: entry-fee + +Delta against `docs/specs/SPEC-289-observe-entry-fee.md`. The spec was written after +the transcript shape was confirmed, so these entries record the decisions the backlog +row left open, not deviations from the written spec. + +## 2026-09-16 10:00 Which number counts as the measured fee + +**Context**: the backlog row cites a hand measurement of about 96,000 tokens built from +a bytes-over-four estimate of four named components. The transcript offers no +per-component token count anywhere. + +**Decision**: the first main-chain assistant turn's +`input_tokens + cache_creation_input_tokens + cache_read_input_tokens` is the measured +whole. The per-component split stays an estimate at four characters per token over the +`rendered` text of each `attachment` entry. + +**Why**: that sum is what the API charged for the first turn, before the agent did any +work, so it is the preamble by definition. Presenting a bytes-over-four figure as the +headline would repeat the hand measurement rather than improve on it. + +**Impact**: the view carries two units side by side. Every header and the JSON label +the split an estimate, and the gap between the two appears as one `(unattributed)` +row rather than being hidden. + +**Alternatives**: size everything from bytes (loses the measurement); report only the +total (loses the per-component ask on the row). + +## 2026-09-16 10:20 The split comes from one session, not a per-component median + +**Context**: a first pass took the median of each component series independently. + +**Decision**: report one real session's split, chosen as the median of the sessions +whose transcript records a rendered preamble. + +**Why**: component medians do not sum to the median fee, so the table would not +reconcile against the measured total and the `(unattributed)` row would be an +artifact. A real session's rows always sum to its own measured fee. + +**Impact**: the header states which session's fee the split belongs to and how many +sessions could be split at all. + +**Open question**: older transcripts record no rendered preamble. On the live 14-day +window, 108 of 435 sessions carried one. If that share falls, the split loses its +base, and the view should say so louder than a count in the header. + +## 2026-09-16 10:35 --project resolution moved into a shared helper + +**Context**: `--project` took an exact cwd-derived slug. A per-repo question is asked +by repo name, and one repo's worktrees each own a separate slug. + +**Decision**: `project_roots()` returns the exact slug dir when it exists, else every +slug containing the given string. `iter_files()` and `_burn_files()` both route +through it. + +**Why**: fixing it only for the new view would leave `cost` and `burn` with the old +behaviour for the same question. One helper, both callers. + +**Impact**: a `--project` string that matches several slugs now walks all of them. +An exact slug still resolves to exactly itself, so no existing invocation changes. diff --git a/lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa1111.jsonl b/lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa1111.jsonl new file mode 100644 index 00000000..8b1c9c41 --- /dev/null +++ b/lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa1111.jsonl @@ -0,0 +1,3 @@ +{"type": "attachment", "isSidechain": false, "timestamp": "2026-09-01T08:00:00Z", "attachment": {"type": "skill_listing"}, "rendered": [{"content": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"}]} +{"type": "attachment", "isSidechain": false, "timestamp": "2026-09-01T08:00:00Z", "attachment": {"type": "instructions"}, "rendered": [{"content": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"}]} +{"type": "assistant", "isSidechain": false, "timestamp": "2026-09-01T08:00:00Z", "message": {"id": "msg_2026-09-01T08:00:00Z", "model": "claude-opus-5", "usage": {"input_tokens": 10, "cache_creation_input_tokens": 490, "cache_read_input_tokens": 500, "output_tokens": 10}, "content": []}} diff --git a/lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa2222.jsonl b/lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa2222.jsonl new file mode 100644 index 00000000..330c7635 --- /dev/null +++ b/lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa2222.jsonl @@ -0,0 +1,5 @@ +{"type": "attachment", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "attachment": {"type": "skill_listing"}, "rendered": [{"content": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"}]} +{"type": "attachment", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "attachment": {"type": "instructions"}, "rendered": [{"content": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"}]} +{"type": "attachment", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "attachment": {"type": "agent_listing_delta"}, "rendered": [{"content": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"}]} +{"type": "assistant", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "message": {"id": "msg_2026-09-08T08:00:00Z", "model": "claude-opus-5", "usage": {"input_tokens": 0, "cache_creation_input_tokens": 1000, "cache_read_input_tokens": 1000, "output_tokens": 10}, "content": []}} +{"type": "assistant", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "message": {"id": "msg_2026-09-08T08:00:00Z", "model": "claude-opus-5", "usage": {"input_tokens": 0, "cache_creation_input_tokens": 9000, "cache_read_input_tokens": 9000, "output_tokens": 10}, "content": []}} diff --git a/lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa2222/subagents/worker.jsonl b/lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa2222/subagents/worker.jsonl new file mode 100644 index 00000000..8d422dd4 --- /dev/null +++ b/lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa2222/subagents/worker.jsonl @@ -0,0 +1,2 @@ +{"type": "attachment", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "attachment": {"type": "skill_listing"}, "rendered": [{"content": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"}]} +{"type": "assistant", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "message": {"id": "msg_2026-09-08T08:00:00Z", "model": "claude-opus-5", "usage": {"input_tokens": 0, "cache_creation_input_tokens": 49999, "cache_read_input_tokens": 50000, "output_tokens": 10}, "content": []}} diff --git a/lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa3333.jsonl b/lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa3333.jsonl new file mode 100644 index 00000000..6664c03c --- /dev/null +++ b/lib/session/observe/tests/fixtures/entryfee/proj-alpha/aaaa3333.jsonl @@ -0,0 +1,2 @@ +{"type": "attachment", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "attachment": {"type": "skill_listing"}, "rendered": [{"content": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"}]} +{"type": "assistant", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "message": {"id": "msg_2026-09-08T08:00:00Z", "model": "claude-opus-5", "usage": {"input_tokens": 0, "cache_creation_input_tokens": 1500, "cache_read_input_tokens": 1500, "output_tokens": 10}, "content": []}} diff --git a/lib/session/observe/tests/fixtures/entryfee/proj-alpha/side9999.jsonl b/lib/session/observe/tests/fixtures/entryfee/proj-alpha/side9999.jsonl new file mode 100644 index 00000000..9173fb47 --- /dev/null +++ b/lib/session/observe/tests/fixtures/entryfee/proj-alpha/side9999.jsonl @@ -0,0 +1,2 @@ +{"type": "attachment", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "attachment": {"type": "skill_listing"}, "rendered": [{"content": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"}]} +{"type": "assistant", "isSidechain": true, "timestamp": "2026-09-08T08:00:00Z", "message": {"id": "msg_2026-09-08T08:00:00Z", "model": "claude-opus-5", "usage": {"input_tokens": 0, "cache_creation_input_tokens": 44444, "cache_read_input_tokens": 44444, "output_tokens": 10}, "content": []}} diff --git a/lib/session/observe/tests/fixtures/entryfee/proj-beta/bbbb1111.jsonl b/lib/session/observe/tests/fixtures/entryfee/proj-beta/bbbb1111.jsonl new file mode 100644 index 00000000..1830486e --- /dev/null +++ b/lib/session/observe/tests/fixtures/entryfee/proj-beta/bbbb1111.jsonl @@ -0,0 +1,2 @@ +{"type": "attachment", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "attachment": {"type": "skill_listing"}, "rendered": [{"content": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"}]} +{"type": "assistant", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "message": {"id": "msg_2026-09-08T08:00:00Z", "model": "claude-opus-5", "usage": {"input_tokens": 0, "cache_creation_input_tokens": 250, "cache_read_input_tokens": 250, "output_tokens": 10}, "content": []}} diff --git a/lib/session/observe/tests/smoke.sh b/lib/session/observe/tests/smoke.sh index ffe44689..66c47c95 100755 --- a/lib/session/observe/tests/smoke.sh +++ b/lib/session/observe/tests/smoke.sh @@ -320,6 +320,63 @@ echo "[61] burn F8: one unknown family (fable) drops the whole row to the token out="$(SESSION_OBSERVE_NOW="$BURNNOW" "$CC" burn --file "$BEDGE/f8-unpriced.jsonl" --since 60 2>&1)" if grep -q 'opus 50% fable 50% ~' <<<"$out"; then ok "F8 unpriced row falls back to token weight, tilde present"; else no "F8 fallback wrong: $out"; fi +EFIX="${DIR}/tests/fixtures/entryfee" # fake projects root for entry-fee: +# proj-alpha fees 1000 (2026-09-01) / 2000 / 3000, proj-beta fee 500, plus two +# negative controls (a subagents/ transcript at 99999, a sidechain-only transcript). + +echo "[62] entry-fee: 4 sessions across 2 projects, median measured fee 2000" +out="$("$CC" entry-fee --root "$EFIX")" +if grep -q '4 sessions, 2 projects' <<<"$out" && grep -q 'median 2000 tokens' <<<"$out"; then ok "4 sessions / 2 projects / median 2000"; else no "entry-fee header wrong: $out"; fi + +echo "[63] entry-fee: component split sized at 4 chars/token (instructions 800 chars -> 200, skill_listing 400 -> 100)" +if grep -Eq 'instructions[[:space:]]+200[[:space:]]+10%' <<<"$out" && grep -Eq 'skill_listing[[:space:]]+100[[:space:]]+5%' <<<"$out"; then ok "instructions 200 (10%), skill_listing 100 (5%)"; else no "component split wrong: $out"; fi + +echo "[64] entry-fee: unattributed = measured fee minus the sized components (2000 - 350 = 1650)" +if grep -Eq '\(unattributed\)[[:space:]]+1650[[:space:]]+82%' <<<"$out"; then ok "unattributed 1650 (82%)"; else no "unattributed wrong: $out"; fi + +echo "[65] entry-fee negative control: a subagents/ transcript is excluded (its 99999 fee absent, proj-alpha stays 3 sessions)" +if ! grep -q '99999' <<<"$out" && grep -Eq 'proj-alpha[[:space:]]+3[[:space:]]+2000' <<<"$out"; then ok "subagent transcript excluded"; else no "subagent transcript counted: $out"; fi + +echo "[66] entry-fee negative control: a sidechain-only transcript contributes no session (4, not 5)" +if ! grep -q '5 sessions' <<<"$out" && ! grep -q '88888' <<<"$out"; then ok "sidechain-only transcript excluded"; else no "sidechain transcript counted: $out"; fi + +echo "[67] entry-fee negative control: a later, larger turn does not replace the FIRST-turn fee (2000, not 18000)" +if ! grep -q '18000' <<<"$out"; then ok "first-turn usage wins (18000 second turn ignored)"; else no "later turn overwrote the fee: $out"; fi + +echo "[68] entry-fee per repo: proj-beta's smaller preamble is its own row (1 session, median 500)" +if grep -Eq 'proj-beta[[:space:]]+1[[:space:]]+500' <<<"$out"; then ok "proj-beta 1 session / median 500"; else no "per-repo rows wrong: $out"; fi + +echo "[69] entry-fee --trend: weekly medians, newest first (2026-W37 3/2000 above 2026-W36 1/1000)" +tout="$("$CC" entry-fee --root "$EFIX" --trend)" +w37="$(awk '/^ 2026-W37/{print NR; exit}' <<<"$tout")" +w36="$(awk '/^ 2026-W36/{print NR; exit}' <<<"$tout")" +if grep -Eq '2026-W37[[:space:]]+3[[:space:]]+2000' <<<"$tout" && grep -Eq '2026-W36[[:space:]]+1[[:space:]]+1000' <<<"$tout" && [[ -n "$w37" && -n "$w36" && "$w37" -lt "$w36" ]]; then ok "W37 3/2000 (line $w37) above W36 1/1000 (line $w36)"; else no "trend wrong: $tout"; fi + +echo "[70] entry-fee negative control: without --trend the weekly table is absent" +if ! grep -q 'weekly trend' <<<"$out"; then ok "weekly table absent without --trend"; else no "trend leaked into the default view: $out"; fi + +echo "[71] entry-fee --project resolves a bare repo name to its slug (alpha -> proj-alpha only)" +pout="$("$CC" entry-fee --root "$EFIX" --project alpha)" +if grep -q '3 sessions, 1 projects' <<<"$pout" && ! grep -q 'proj-beta' <<<"$pout"; then ok "bare name alpha resolved to proj-alpha (3 sessions, beta absent)"; else no "project resolution wrong: $pout"; fi + +echo "[72] entry-fee --json: valid JSON, median_fee 2000, split flagged as an estimate" +jout="$("$CC" entry-fee --root "$EFIX" --trend --json)" +if echo "$jout" | python3 -c ' +import json, sys +d = json.load(sys.stdin) +assert d["sessions"] == 4, d["sessions"] +assert d["median_fee"] == 2000, d["median_fee"] +assert d["components_estimated"] is True and d["chars_per_token"] == 4, d +comps = {c["component"]: c["est_tokens"] for c in d["split_components"]} +assert comps["instructions"] == 200 and comps["skill_listing"] == 100, comps +assert comps["(unattributed)"] == 1650, comps +assert [w["week"] for w in d["by_week"]] == ["2026-W37", "2026-W36"], d["by_week"] +'; then ok "json median_fee 2000, estimated split, weeks newest-first"; else no "entry-fee json wrong: $jout"; fi + +echo "[73] report negative control: no entry-fee section (existing views unchanged)" +out="$("$CC" report --file "$FIX")" +if ! grep -q '# entry-fee' <<<"$out"; then ok "report has no entry-fee section"; else no "entry-fee leaked into report: $out"; fi + echo if [[ $fail -gt 0 ]]; then echo "smoke: $pass passed, $fail FAILED" >&2; exit 1; fi echo "smoke: all $pass passed" From 1b731cfe7cb25ba7282e16e8a2ee81ba00f796b1 Mon Sep 17 00:00:00 2001 From: Han Ngo Date: Wed, 16 Sep 2026 18:15:09 +0700 Subject: [PATCH 2/5] fix(session-observe): guard untrusted entry-fee fields and fold worktree slugs The architecture and correctness review on this branch found three unguarded crash paths and two outputs that read as measurements when they were estimates. A numeric timestamp, a numeric rendered content, and a dict attachment type each aborted the whole scan rather than that one file, while every sibling collector in the module already guards the same shapes. Each now contributes nothing and the scan continues. The remainder row went negative when the four-characters-per-token estimate overshot the measured fee, printing a negative token count and a negative share. It now prints (estimate over measured) with the magnitude, and the component shares above 100 percent carry the signal. The per-repo table keyed on the raw project slug, so a repo's worktrees each ranked as a separate repo and a one-session worktree sorted above the many-session checkout of the same codebase. Rows fold on the worktree marker in the slug and break median ties on session count. A multi-slug --project match now announces its resolved list on stderr, so the widened resolution cannot merge unrelated repos in silence. pctl and entry_fee_median share one nearest-rank index, so the session shown for the median split cannot drift from the median it is shown for. --- docs/specs/SPEC-289-observe-entry-fee.md | 38 +++++++- lib/session/observe/README.md | 2 +- lib/session/observe/bin/session-observe | 89 +++++++++++++++---- .../docs/implementation-notes/entry-fee.md | 28 +++++- .../overshoot/proj-x/over1111.jsonl | 2 + .../untrusted/proj-y/bad1111.jsonl | 5 ++ .../b1111111.jsonl | 2 + .../worktrees/proj-alpha/a1111111.jsonl | 2 + lib/session/observe/tests/smoke.sh | 29 ++++++ 9 files changed, 173 insertions(+), 24 deletions(-) create mode 100644 lib/session/observe/tests/entryfee-edge/overshoot/proj-x/over1111.jsonl create mode 100644 lib/session/observe/tests/entryfee-edge/untrusted/proj-y/bad1111.jsonl create mode 100644 lib/session/observe/tests/entryfee-edge/worktrees/proj-alpha--claude-worktrees-wt-one/b1111111.jsonl create mode 100644 lib/session/observe/tests/entryfee-edge/worktrees/proj-alpha/a1111111.jsonl diff --git a/docs/specs/SPEC-289-observe-entry-fee.md b/docs/specs/SPEC-289-observe-entry-fee.md index bd4e9e0f..adf1ab5a 100644 --- a/docs/specs/SPEC-289-observe-entry-fee.md +++ b/docs/specs/SPEC-289-observe-entry-fee.md @@ -60,6 +60,15 @@ The split is labelled an estimate in the table header and in the JSON (`components_estimated`). The same rule of thumb the hand measurement used is stated alongside it. +Four characters per token overshoots on dense markdown and tables, so the sized blocks +can exceed the measured fee. That prints as an `(estimate over measured)` row carrying +the magnitude, never as a negative token count, and the component shares then read +above 100 percent, which is the estimate saying it broke here. + +The measured total also covers the first user prompt and anything attached to it, so it +is the preamble plus turn one. The `(unattributed)` row absorbs that alongside the +system prompt and the tool schemas. + The split is reported for ONE real session, not as a per-component median. Medians of separate components do not add up to the median fee, so a median-of-medians table would not reconcile against a measured total. That session is the median of the @@ -69,9 +78,16 @@ unattributed fee. ### The per-repo and weekly figures -Per repo: one row per project slug with the session count and the median measured fee, -largest median first. The memory index and the CLAUDE.md stack differ per checkout, so -this is the actionable cut. +Per repo: one row per repo with the session count and the median measured fee, largest +median first, ties broken by session count. The memory index and the CLAUDE.md stack +differ per checkout, so this is the actionable cut. + +A worktree gets its own project slug. The worktree convention is +`/.claude/worktrees/`, so the slug carries a `--claude-worktrees-` marker +and the repo name sits before it. Rows key on the part before that marker, which folds +a repo's worktrees into one row. Without the fold, a one-session worktree slug ranks +above the many-session checkout of the same repo, which reads as a comparison when it +is one codebase twice. `--trend`: ISO-week buckets of the median measured fee, newest week first, so a reduction from a cleanup lands as a visible drop. @@ -84,6 +100,10 @@ cwd-derived slug. One repo's worktrees each get their own slug and a per-repo fi wants them together, so all matches are walked. This resolution is shared with the other views through `project_roots()`. +A substring matching more than one slug prints the resolved list to stderr. Before the +fallback existed, a wrong `--project` walked nothing and the empty output said so; a +silent multi-match would instead merge unrelated repos into one plausible figure. + ## Non-goals - Not part of `report`. `report` is the weekly digest of behaviour; the entry fee is a @@ -113,5 +133,17 @@ sidechain-only transcript): 72. `--json` is valid, carries median_fee 2000 and the estimate flag. 73. negative control: `report` prints no entry-fee section. +Plus, against `tests/entryfee-edge/` (an overshooting session, an untrusted-field +session, a repo with one worktree slug): + +74. an estimate above the measured fee prints as `(estimate over measured)`, never a + negative token count. +75. the overshooting component's share reads above 100 percent. +76. untrusted fields (numeric timestamp, dict attachment type, numeric rendered content, + non-dict message) do not crash the scan; the valid turn's fee is still measured. +77. negative control: the dict-typed attachment contributes no component row. +78. a worktree slug folds into its repo row (one row, two sessions). +79. a multi-slug `--project` match is announced on stderr. + Plus a real run over the live transcripts, recorded in `lib/session/observe/docs/verification/entry-fee.md`. diff --git a/lib/session/observe/README.md b/lib/session/observe/README.md index 2d59f004..07b3a499 100644 --- a/lib/session/observe/README.md +++ b/lib/session/observe/README.md @@ -57,7 +57,7 @@ Each transcript entry already carries `hookInfos: [{command, durationMs}]`, `hoo - **cost**: from the assistant `message.usage` block (input / output / cache-read / cache-create tokens) + `message.model`. **tokens-by-model** + a **$ estimate** from an embedded dated `PRICING` table (Max plan is flat-rate, so this is *attribution*, not a bill; unknown model families like `fable` count tokens but show `?`). **cache economics** = cache-read / (read + create) hit ratio, the biggest cost lever. (cost-per-merged-PR is deferred, see proof/impl-notes: transcripts are cross-repo and this repo squash-merges, so there is no clean per-PR attribution.) - **burn**: which session is burning tokens *right now*, not over the week. One row per top-level session inside the last `--since` minutes (default 60); a subagent transcript (`/subagents/*.jsonl`) rolls its tokens into its parent row and counts toward `subs`. Usage is deduplicated by `(message.id, requestId)` (a streamed chunk repeats the same usage block). `ctx` is the live context size: input + cache-creation + cache-read of the last main-chain (non-sidechain) assistant usage, regardless of the window. Ranked by `input + cache_create + cache_read/10 + output*5` (attribution, not a bill). The `pid` column maps through the live `~/.claude/sessions/.json` files. `model share` splits the row's burn by model, biggest first, so an expensive fan-out is legible without reading the transcripts: a lead that dispatched its workers on Opus reads `opus 91%`. The split is priced through `model_cost()`, not counted, because equal Opus and Sonnet token counts are not equal spend (Opus lists at five times Sonnet on every axis, so equal tokens read `opus 83% sonnet 17%`). One model with an unknown family (fable) leaves the row unpriceable, so the whole row falls back to the same list-price token weight the ranking uses and a trailing `~` marks it. Not part of `report`. -- **entry-fee**: the fixed preamble every turn re-reads before any work (skill listing, CLAUDE.md stack, repo memory index, agent roster, output style, tool schemas). The per-session TOTAL is **measured**: the first main-chain assistant turn's input + cache-creation + cache-read, which is exactly what the model read before it acted. The per-component SPLIT is **estimated**: each preamble `attachment` entry is sized from its `rendered` text at 4 characters per token and keyed by `attachment.type`, because the API reports one usage block per turn and no per-component count exists in the data. The remainder against the measured total is one `(unattributed)` row (system prompt + built-in tool schemas, which reach the model but never the transcript). The split is reported for ONE real session rather than as a per-component median, since medians of separate components would not add up to the measured total; that session is the median of the sessions whose transcript records a rendered preamble. Per-repo rows carry the median measured fee (the memory index and CLAUDE.md stack differ per checkout), and `--trend` adds ISO-week medians newest-first so a cleanup shows as a drop. A subagent's own transcript has no main-chain turn and is excluded. Not part of `report`. +- **entry-fee**: the fixed preamble every turn re-reads before any work (skill listing, CLAUDE.md stack, repo memory index, agent roster, output style, tool schemas). The per-session TOTAL is **measured**: the first main-chain assistant turn's input + cache-creation + cache-read, which is exactly what the model read before it acted. The per-component SPLIT is **estimated**: each preamble `attachment` entry is sized from its `rendered` text at 4 characters per token and keyed by `attachment.type`, because the API reports one usage block per turn and no per-component count exists in the data. The remainder against the measured total is one `(unattributed)` row (system prompt + built-in tool schemas, which reach the model but never the transcript). The split is reported for ONE real session rather than as a per-component median, since medians of separate components would not add up to the measured total; that session is the median of the sessions whose transcript records a rendered preamble. Four characters per token overshoots on dense markdown, so a session whose sized blocks exceed its measured fee prints an `(estimate over measured)` row with the magnitude rather than a negative token count, and its component shares read above 100%. The measured total also covers the first user prompt, so it is the preamble plus turn one; the `(unattributed)` row absorbs that too. Per-repo rows carry the median measured fee (the memory index and CLAUDE.md stack differ per checkout) and fold a repo's `--claude-worktrees-` slugs into one row, since a one-session worktree ranking above its own many-session checkout reads as a comparison when it is one codebase twice. `--trend` adds ISO-week medians newest-first so a cleanup shows as a drop. A subagent's own transcript has no main-chain turn and is excluded. Not part of `report`. ## Output (real run, abbreviated) diff --git a/lib/session/observe/bin/session-observe b/lib/session/observe/bin/session-observe index bf442ad8..76133d9a 100755 --- a/lib/session/observe/bin/session-observe +++ b/lib/session/observe/bin/session-observe @@ -128,6 +128,11 @@ def project_roots(root, project): without the full cwd-derived slug. One repo's worktrees each get their own slug, and a per-repo figure wants them together, so this returns all matches rather than picking one. No match -> the literal join, which walks nothing. + + A substring that matches more than one slug prints the resolved list to + stderr. Before this fallback existed a wrong --project walked nothing and + the empty output said so; a silent multi-match would instead merge unrelated + repos into one plausible figure. """ exact = os.path.join(root, project) if os.path.isdir(exact): @@ -139,6 +144,9 @@ def project_roots(root, project): ) except OSError: hits = [] + if len(hits) > 1: + print(f"session-observe: --project {project!r} matched {len(hits)} slugs: " + + ", ".join(os.path.basename(h) for h in hits), file=sys.stderr) return hits or [exact] @@ -223,13 +231,22 @@ def hook_label(command): return toks[0].rsplit("/", 1)[-1][:40] or "?" +def _rank_index(n, p): + """Nearest-rank index into a sorted list of n values for percentile p. + + Shared so pctl() and entry_fee_median() cannot drift: the entry-fee view + reports the component split of the session AT its median fee, and a drift + here would stop that session reconciling against the median it is shown for. + """ + return min(n - 1, int(round((p / 100) * (n - 1)))) + + def pctl(vals, p): """Nearest-rank percentile (integer ms). Empty -> 0.""" if not vals: return 0 s = sorted(vals) - i = min(len(s) - 1, int(round((p / 100) * (len(s) - 1)))) - return int(s[i]) + return int(s[_rank_index(len(s), p)]) def collect(args): @@ -838,16 +855,32 @@ ENTRY_FEE_CHARS_PER_TOKEN = 4 # Fee minus the sized components: the system prompt and the built-in tool # schemas reach the model but never appear in the transcript. ENTRY_FEE_UNATTRIBUTED = "(unattributed)" +# The other direction: the estimate overshot the measured fee, so the split is +# wrong here and the row says so instead of printing a negative measurement. +ENTRY_FEE_OVERSHOT = "(estimate over measured)" +# Claude Code gives a worktree its own project slug. The worktree convention is +# /.claude/worktrees/, so the slug carries this separator and the +# repo name sits before it. Splitting there groups a repo's worktrees together. +ENTRY_FEE_WORKTREE_MARK = "--claude-worktrees-" def _rendered_chars(entry): - """Characters of rendered preamble text one attachment entry contributes.""" + """Characters of rendered preamble text one attachment entry contributes. + + Every field here is untrusted transcript data, so a non-string content or a + non-list rendered contributes zero rather than raising: one malformed line + must not abort the whole scan.""" r = entry.get("rendered") if isinstance(r, str): return len(r) if not isinstance(r, list): return 0 - return sum(len(b.get("content") or "") for b in r if isinstance(b, dict)) + total = 0 + for b in r: + c = b.get("content") if isinstance(b, dict) else None + if isinstance(c, str): + total += len(c) + return total def entry_fee_session(path): @@ -869,7 +902,7 @@ def entry_fee_session(path): att = entry.get("attachment") label = att.get("type") if isinstance(att, dict) else None tokens = _rendered_chars(entry) // ENTRY_FEE_CHARS_PER_TOKEN - if label and tokens: + if isinstance(label, str) and tokens: comps[label] += tokens continue if entry.get("type") != "assistant" or entry.get("isSidechain"): @@ -881,8 +914,10 @@ def entry_fee_session(path): fee = ((u.get("input_tokens") or 0) + (u.get("cache_creation_input_tokens") or 0) + (u.get("cache_read_input_tokens") or 0)) - day = (first_ts or ts or "")[:10] - return {"fee": fee, "components": comps, "day": day, "model": msg.get("model") or "?"} + # first_ts is the only timestamp already proven to be a string; ts is + # untrusted and slicing it directly would abort the whole scan. + return {"fee": fee, "components": comps, "day": (first_ts or "")[:10], + "model": msg.get("model") or "?"} return None @@ -904,8 +939,7 @@ def entry_fee_median(sessions): if not sessions: return None ranked = sorted(sessions, key=lambda s: s["fee"]) - i = min(len(ranked) - 1, int(round(0.5 * (len(ranked) - 1)))) - return ranked[i] + return ranked[_rank_index(len(ranked), 50)] def entry_fee_split_session(sessions): @@ -933,7 +967,14 @@ def entry_fee_component_rows(sess): return f"{100 * n / fee:.0f}%" if fee else "-" rows = [[name, n, share(n)] for name, n in sess["components"].most_common()] rest = fee - sum(sess["components"].values()) - rows.append([ENTRY_FEE_UNATTRIBUTED, rest, share(rest)]) + # A negative remainder means the 4-chars-per-token estimate overshot the + # measured fee, which happens on dense markdown and tables. Print that as + # what it is rather than as a negative measurement: the component shares + # then read above 100% and say plainly that the estimate broke here. + if rest < 0: + rows.append([ENTRY_FEE_OVERSHOT, -rest, "-"]) + else: + rows.append([ENTRY_FEE_UNATTRIBUTED, rest, share(rest)]) return rows @@ -943,16 +984,25 @@ def short_project(slug, n=46): return ("..." + slug[-(n - 3):]) if len(slug) > n else slug +def repo_of_slug(slug): + """A project slug -> the repo it belongs to, folding its worktree slugs in. + + A one-session worktree slug is not a repo, and ranking it against a + many-session checkout of the same repo reads as a comparison when it is + the same codebase twice.""" + return slug.split(ENTRY_FEE_WORKTREE_MARK, 1)[0] + + def entry_fee_project_rows(sessions, top): - """Per-repo median fee, every project (the caller slices for display). + """Per-repo median fee, largest median first, ties broken by session count. The memory index and CLAUDE.md stack differ per checkout, so the per-repo figure is the actionable one.""" by = defaultdict(list) for s in sessions: - by[s["project"]].append(s["fee"]) + by[repo_of_slug(s["project"])].append(s["fee"]) rows = [[short_project(p), len(v), pctl(v, 50)] for p, v in by.items()] - rows.sort(key=lambda r: -r[2]) - return rows + rows.sort(key=lambda r: (-r[2], -r[1])) + return rows[:top] if top else rows def iso_week(day): @@ -964,12 +1014,13 @@ def iso_week(day): return f"{y}-W{w:02d}" -def entry_fee_week_rows(sessions): +def entry_fee_week_rows(sessions, top): """Weekly median fee, newest week first, so a reduction shows as a drop.""" by = defaultdict(list) for s in sessions: by[iso_week(s["day"])].append(s["fee"]) - return [[w, len(by[w]), pctl(by[w], 50)] for w in sorted(by, reverse=True)] + rows = [[w, len(by[w]), pctl(by[w], 50)] for w in sorted(by, reverse=True)] + return rows[:top] if top else rows def run_entry_fee(args): @@ -980,7 +1031,7 @@ def run_entry_fee(args): split_sess, splittable = entry_fee_split_session(sessions) comp = entry_fee_component_rows(split_sess) projects = entry_fee_project_rows(sessions, args.top) - weeks = entry_fee_week_rows(sessions) if args.trend else [] + weeks = entry_fee_week_rows(sessions, args.top) if args.trend else [] med = med_sess["fee"] if med_sess else 0 if args.json: @@ -1003,10 +1054,10 @@ def run_entry_fee(args): print(f" the split is ESTIMATED at {ENTRY_FEE_CHARS_PER_TOKEN} chars/token, only the totals are measured):") print_table(["component", "est-tokens", "share"], comp) print(" per repo (median measured fee):") - print_table(["project", "sessions", "median-fee"], projects[: args.top] if args.top else projects) + print_table(["project", "sessions", "median-fee"], projects) if args.trend: print(" weekly trend (median measured fee):") - print_table(["week", "sessions", "median-fee"], weeks[: args.top] if args.top else weeks) + print_table(["week", "sessions", "median-fee"], weeks) print() diff --git a/lib/session/observe/docs/implementation-notes/entry-fee.md b/lib/session/observe/docs/implementation-notes/entry-fee.md index 562b8043..3a27b93b 100644 --- a/lib/session/observe/docs/implementation-notes/entry-fee.md +++ b/lib/session/observe/docs/implementation-notes/entry-fee.md @@ -57,4 +57,30 @@ through it. behaviour for the same question. One helper, both callers. **Impact**: a `--project` string that matches several slugs now walks all of them. -An exact slug still resolves to exactly itself, so no existing invocation changes. +An exact slug still resolves to exactly itself, so no invocation that already named a +valid slug changes. A typo or a partial name used to walk nothing and say so through an +empty report; it now returns a merged figure, so a multi-slug match announces its +resolved list on stderr. + +## 2026-09-16 12:10 Review fixes + +The architecture and correctness review found three unguarded crash paths and two +outputs that read as a measurement when they were not. All are fixed on this branch. + +- Three untrusted fields were sliced or hashed without a guard: a numeric `timestamp`, + a numeric `rendered[].content`, and a dict `attachment.type`. Each aborted the whole + scan, not just that file, while every sibling collector in the module guards the same + shapes. Now each contributes nothing and the scan continues. +- The remainder row went negative when the four-characters-per-token estimate overshot + the measured fee, printing `-900` and `-900%` as if they were measurements. It now + prints `(estimate over measured)` with the magnitude, and the component shares above + 100 percent carry the signal. +- The per-repo table keyed on the raw project slug, so a repo's worktrees each ranked as + a separate repo and a one-session worktree slug sorted above the many-session checkout + of the same codebase. Rows now fold on the `--claude-worktrees-` marker. + +One review point stays a documentation fix rather than a code one: the measured fee also +covers the first user prompt, not the preamble alone. The `(unattributed)` row absorbs +it, which matters at the 51 percent unattributed share the live run shows. Stated in the +spec and the README instead of subtracted, because the transcript gives no way to +separate the two. diff --git a/lib/session/observe/tests/entryfee-edge/overshoot/proj-x/over1111.jsonl b/lib/session/observe/tests/entryfee-edge/overshoot/proj-x/over1111.jsonl new file mode 100644 index 00000000..92b914fa --- /dev/null +++ b/lib/session/observe/tests/entryfee-edge/overshoot/proj-x/over1111.jsonl @@ -0,0 +1,2 @@ +{"type": "attachment", "timestamp": "2026-09-08T08:00:00Z", "attachment": {"type": "instructions"}, "rendered": [{"content": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"}]} +{"type": "assistant", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "message": {"id": "m1", "model": "claude-opus-5", "content": [], "usage": {"input_tokens": 0, "cache_creation_input_tokens": 500, "cache_read_input_tokens": 500, "output_tokens": 1}}} diff --git a/lib/session/observe/tests/entryfee-edge/untrusted/proj-y/bad1111.jsonl b/lib/session/observe/tests/entryfee-edge/untrusted/proj-y/bad1111.jsonl new file mode 100644 index 00000000..7e9be319 --- /dev/null +++ b/lib/session/observe/tests/entryfee-edge/untrusted/proj-y/bad1111.jsonl @@ -0,0 +1,5 @@ +{"type": "attachment", "timestamp": 1757318400, "attachment": {"type": "skill_listing"}, "rendered": [{"content": 12345}]} +{"type": "attachment", "timestamp": "2026-09-08T08:00:00Z", "attachment": {"type": {"nested": "dict"}}, "rendered": [{"content": "yyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyy"}]} +{"type": "attachment", "timestamp": "2026-09-08T08:00:00Z", "attachment": {"type": "instructions"}, "rendered": "not-a-list"} +{"type": "assistant", "timestamp": "2026-09-08T08:00:00Z", "message": "oops"} +{"type": "assistant", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "message": {"id": "m1", "model": "claude-opus-5", "content": [], "usage": {"input_tokens": 0, "cache_creation_input_tokens": 350, "cache_read_input_tokens": 350, "output_tokens": 1}}} diff --git a/lib/session/observe/tests/entryfee-edge/worktrees/proj-alpha--claude-worktrees-wt-one/b1111111.jsonl b/lib/session/observe/tests/entryfee-edge/worktrees/proj-alpha--claude-worktrees-wt-one/b1111111.jsonl new file mode 100644 index 00000000..ee61ccf1 --- /dev/null +++ b/lib/session/observe/tests/entryfee-edge/worktrees/proj-alpha--claude-worktrees-wt-one/b1111111.jsonl @@ -0,0 +1,2 @@ +{"type": "attachment", "timestamp": "2026-09-08T08:00:00Z", "attachment": {"type": "skill_listing"}, "rendered": [{"content": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"}]} +{"type": "assistant", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "message": {"id": "m1", "model": "claude-opus-5", "content": [], "usage": {"input_tokens": 0, "cache_creation_input_tokens": 2000, "cache_read_input_tokens": 2000, "output_tokens": 1}}} diff --git a/lib/session/observe/tests/entryfee-edge/worktrees/proj-alpha/a1111111.jsonl b/lib/session/observe/tests/entryfee-edge/worktrees/proj-alpha/a1111111.jsonl new file mode 100644 index 00000000..09629f48 --- /dev/null +++ b/lib/session/observe/tests/entryfee-edge/worktrees/proj-alpha/a1111111.jsonl @@ -0,0 +1,2 @@ +{"type": "attachment", "timestamp": "2026-09-08T08:00:00Z", "attachment": {"type": "skill_listing"}, "rendered": [{"content": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"}]} +{"type": "assistant", "isSidechain": false, "timestamp": "2026-09-08T08:00:00Z", "message": {"id": "m1", "model": "claude-opus-5", "content": [], "usage": {"input_tokens": 0, "cache_creation_input_tokens": 1000, "cache_read_input_tokens": 1000, "output_tokens": 1}}} diff --git a/lib/session/observe/tests/smoke.sh b/lib/session/observe/tests/smoke.sh index 66c47c95..16f81d28 100755 --- a/lib/session/observe/tests/smoke.sh +++ b/lib/session/observe/tests/smoke.sh @@ -377,6 +377,35 @@ echo "[73] report negative control: no entry-fee section (existing views unchang out="$("$CC" report --file "$FIX")" if ! grep -q '# entry-fee' <<<"$out"; then ok "report has no entry-fee section"; else no "entry-fee leaked into report: $out"; fi +EEDGE="${DIR}/tests/entryfee-edge" # entry-fee edge roots, deliberately OUTSIDE tests/fixtures/ +# (the untrusted root's numeric-timestamp and non-dict-message lines would otherwise +# also break session-semantic's own --root tests/fixtures walk) + +echo "[74] entry-fee: an estimate ABOVE the measured fee prints as an overshoot, never a negative measurement" +out="$("$CC" entry-fee --root "$EEDGE/overshoot")" +if grep -Eq '\(estimate over measured\)[[:space:]]+1000[[:space:]]+-' <<<"$out" && ! grep -q -- '-1000' <<<"$out"; then ok "overshoot row 1000, no negative token count"; else no "overshoot row wrong: $out"; fi + +echo "[75] entry-fee: the overshooting component's share reads above 100% (the estimate saying it broke)" +if grep -Eq 'instructions[[:space:]]+2000[[:space:]]+200%' <<<"$out"; then ok "instructions 2000 (200%)"; else no "overshoot share wrong: $out"; fi + +echo "[76] entry-fee: untrusted fields (numeric timestamp, dict attachment type, numeric content, non-dict message) do not crash" +set +e +out="$("$CC" entry-fee --root "$EEDGE/untrusted" 2>&1)" +rc=$? +set -e +if [[ $rc -eq 0 ]] && grep -q 'median 700 tokens' <<<"$out"; then ok "exit 0, the valid turn's fee 700 still measured"; else no "untrusted input crashed or row missing (rc=$rc): $out"; fi + +echo "[77] entry-fee negative control: the dict-typed attachment contributes no component row" +if ! grep -q 'nested' <<<"$out"; then ok "unhashable attachment type skipped, not keyed"; else no "dict attachment type leaked into the split: $out"; fi + +echo "[78] entry-fee: a worktree slug folds into its repo row (2 sessions under one proj-alpha, not two rows)" +out="$("$CC" entry-fee --root "$EEDGE/worktrees")" +if grep -Eq 'proj-alpha[[:space:]]+2[[:space:]]+2000' <<<"$out" && ! grep -q 'wt-one' <<<"$out"; then ok "proj-alpha 2 sessions, worktree slug folded in"; else no "worktree grouping wrong: $out"; fi + +echo "[79] entry-fee: a multi-slug --project match is announced on stderr (never a silent merge)" +err="$("$CC" entry-fee --root "$EEDGE/worktrees" --project alpha 2>&1 >/dev/null)" +if grep -q "matched 2 slugs" <<<"$err" && grep -q 'proj-alpha--claude-worktrees-wt-one' <<<"$err"; then ok "multi-match announced with the resolved slugs"; else no "multi-match not announced: $err"; fi + echo if [[ $fail -gt 0 ]]; then echo "smoke: $pass passed, $fail FAILED" >&2; exit 1; fi echo "smoke: all $pass passed" From 62e027e509427b7a428be764780630ba4f5609ef Mon Sep 17 00:00:00 2001 From: Han Ngo Date: Wed, 16 Sep 2026 18:17:10 +0700 Subject: [PATCH 3/5] docs(session-observe): proof of done for the entry-fee view Records the 18 smoke cases with their 7 negative controls, the live 14-day run over 435 sessions across 8 repos, and the negctl verdict. The negative control dropped cache_read_input_tokens from the fee sum, the term that carries most of a real preamble. The suite went red under the mutation and green again after the restore. The two largest sized components on the live run, the skill listing at 22,523 and the CLAUDE.md stack at 12,653, land in the same band as the 2026-09-13 hand measurement taken on a different repo. The split is now available on demand, per repo and per week. Flips ID-878 to shipped. --- _meta/BACKLOG.md | 2 +- lib/session/observe/docs/proof-of-done.md | 106 ++++++++++++++++++++++ 2 files changed, 107 insertions(+), 1 deletion(-) diff --git a/_meta/BACKLOG.md b/_meta/BACKLOG.md index 20d7c025..befcb8a1 100644 --- a/_meta/BACKLOG.md +++ b/_meta/BACKLOG.md @@ -38,7 +38,7 @@ Lane = the WORKFLOW.md risk tier (`tiny` / `normal` / `full`). | ID-881 | bin/wrap merge retries transient GitHub failures #wrap #resilience | bin/wrap merge has no retry for a transient GitHub failure (502/503/GraphQL executing-query errors); a real refusal (not mergeable, not green, conflict, draft) must still fail immediately. Operator hand-rolled 3 shell retry loops for ~25 PR merges during a ~90min GitHub outage. Same failure class as memory note hand-rolled-merge-loop-instead-of-wrap-merge.md, a second occurrence, missing retry is the root cause. cmd_merge in lib/wrap/wrap.sh (~line 812, gh pr merge call ~line 917). lane=full. Source: board sweep session 2026-09-13 and 14. Goal draft: .claude/goals/wrap-merge-retry.md in the kit clone. | queued | | ID-879 | wrap step 6 and the ledger flush must not commit on the checked-out default branch #wrap #ledger | Sessions commit LAB_LOG lines, ledger flushes and board rows straight onto local main because those writes land in the main checkout; main cannot be pushed, so 29 such commits plus 61 merge commits piled up on one Air by 2026-09-14 (ops-toolkit #2779 drained them). Wrap step 6, the learning-ledger flush and board capture should write on a housekeeping branch or leave the file uncommitted for the next feature PR. Companion guards: ops-toolkit pre-commit refuses commits on main; dotfiles sets pull.ff=only. Source: ops-toolkit session 2026-09-14. | queued | | ID-882 | Dispatch keeps an Attempt state apart from the Task state: a dead or disconnected subagent is an unknown outcome, a grace window precedes LOST, and a result arriving inside the window commits with no second dispatch #kit #dispatch #u-mid #f-mid | Formalises the incident lesson resume-a-dead-subagent-never-respawn-on-its-branch. Shape from VoiceStudio backend/worker/lifecycle.py: TaskState and AttemptState enums with a legal-transition table, mark_disconnected starts grace, lose_attempt on expiry frees the task and excludes that worker, commit_result idempotent on task id so a resumed agent and its replacement cannot both land. Applies to kit:execute retries and megagoal-agent-drive. Contract tests named in ops-toolkit research/2026-09-14-voicestudio-long-running-orchestration.md | queued | -| ID-878 | session observe: a view that sizes the fixed context entry fee per turn #harness #kit #observability | Every agent turn re-reads a fixed preamble (skills, CLAUDE.md stack, repo memory index, agent roster) before any work. Measured by hand 2026-09-13 at ~96k tokens, about 38% of a 400-turn builder's whole bill. session observe already owns cost and burn over the same transcripts, so the entry-fee view belongs beside them rather than as a sibling script. Wants: per-component breakdown, a trend so ID-902..906 progress is visible, and a per-repo figure since the memory index and CLAUDE.md differ per checkout. lane=full. Source: ops-toolkit research/2026-09-13-token-burn-optimization.md | queued | +| ID-878 | session observe: a view that sizes the fixed context entry fee per turn #harness #kit #observability | Every agent turn re-reads a fixed preamble (skills, CLAUDE.md stack, repo memory index, agent roster) before any work. Measured by hand 2026-09-13 at ~96k tokens, about 38% of a 400-turn builder's whole bill. session observe already owns cost and burn over the same transcripts, so the entry-fee view belongs beside them rather than as a sibling script. Wants: per-component breakdown, a trend so ID-902..906 progress is visible, and a per-repo figure since the memory index and CLAUDE.md differ per checkout. lane=full. Source: ops-toolkit research/2026-09-13-token-burn-optimization.md | shipped [SPEC-289, proof lib/session/observe/docs/proof-of-done.md] | | ID-877 | gate-ledger: record a whole lane plan in one call #kit #gate | Intent: one verb that takes a rid, a lane, and per-phase dispositions (ran, skipped, override, each with its reason) and writes every GATE line the ship-gate wants, instead of nine hand-typed record/override calls per prose-only PR. Precedent: nothing matched (bin/precedent, 2026-09-13). lane=full (touches lib/gate). Source: the skill-trigger-routing wrap. Goal draft: .claude/goals/gate-ledger-plan-record.md in the kit clone. | shipped [SPEC-287, proof docs/verification/gate-ledger-plan-record.md] | | ID-874 | wrap apply: pull a checkout past a sibling session's dirty tracked files #wrap #git #u-mid #f-mid | Lane full (lane-classify: touches the pull path of a shared checkout). Twice this session an ff pull of the kit primary checkout was blocked by a sibling's dirty tracked files (staging log, run-all sample drift); the fix was done by hand both times: a NAMED stash of only the files git names as blocking, git pull --ff-only, pop BY REF, union-resolve the append-only log, checkout --ours for docs/verification/*/sample-*.md drift, drop only the named stash. Precedent: memory bare-stash-pop-takes-a-sibling-stash (prose, no mechanism). Home lib/wrap/wrap.sh apply pull path, behind a knob, default off. Goal draft: .claude/goals/wrap-pull-past-dirty.md. Filed by /kit:wrap 7b 2026-09-13. Rationale: recurring pull-blocked pain this session, and the manual fix is already sketched behind a knob. | shipped [#630, SPEC-286, proof docs/verification/wrap-pull-past-dirty.md] | | ID-875 | wrap merge verifies the default branch after a squash merge #wrap #safety | After a squash merge bin/wrap merge does not confirm the default branch holds the PR head's tree, so a dropped commit goes unnoticed. Add the postcondition. Home lib/wrap/wrap.sh. Related memory note verify-main-after-squash-merge. lane=full. Source: ops-toolkit handoff 2026-09-13-wrap-build-lanes-and-leftovers. | shipped [#628, proof docs/verification/wrap-merge-tree-check.md] | diff --git a/lib/session/observe/docs/proof-of-done.md b/lib/session/observe/docs/proof-of-done.md index 262b5dec..320f69a8 100644 --- a/lib/session/observe/docs/proof-of-done.md +++ b/lib/session/observe/docs/proof-of-done.md @@ -617,3 +617,109 @@ bash lib/session/observe/tests/smoke.sh # -> smoke: all 56 passed bash bin/session observe burn --since 60 # live ranked table bash bin/session observe burn --json # {window_min, sessions[]} ``` + +## SPEC-289 `entry-fee`: the fixed context entry fee per turn + +**Feature:** `session observe entry-fee` sizes the preamble every agent turn re-reads before any work. The per-session total is measured (the first main-chain assistant turn's input + cache-creation + cache-read); the per-component split is estimated at four characters per token over the rendered preamble text and is labelled an estimate everywhere. Adds a per-repo median and, with `--trend`, ISO-week medians. Spec: `docs/specs/SPEC-289-observe-entry-fee.md`. Row: ID-878. +**Date:** 2026-09-16 · **Lane:** full · **Host:** dev laptop (macOS 27.0) + +### Acceptance criteria + +| # | Criterion | Source | +|---|---|---| +| E1 | Per-component breakdown of the preamble | ID-878 "per-component breakdown" | +| E2 | Per-repo figure, since the memory index and CLAUDE.md stack differ per checkout | ID-878 "a per-repo figure" | +| E3 | A trend, so a preamble cleanup is visible | ID-878 "a trend so progress is visible" | +| E4 | The measured half is never presented as an estimate, nor the estimate as a measurement | SPEC-289 Behaviour | +| E5 | Lives inside the existing CLI beside `cost` and `burn`, not as a sibling script | ID-878 "belongs beside them" | +| E6 | Read-only, stdlib only, no new dependency | module contract | + +### Run table + +| Check | Command | Expected | Result | +|---|---|---|---| +| Module suite green | `bash lib/session/observe/tests/smoke.sh \| tail -1` | all cases pass | PASS, `smoke: all 79 passed` | +| Kit suite green | `bash tests/run-all.sh \| tail -1` | every suite passes | PASS, `all 147 suites passed, 1 skipped for missing tooling` | +| Header + median (E1) | smoke 62 | 4 sessions, 2 projects, median 2000 | PASS | +| Split sized at 4 chars/token (E1) | smoke 63 | instructions 200, skill_listing 100 | PASS | +| Remainder reconciles (E4) | smoke 64 | `(unattributed)` 1650 = 2000 - 350 | PASS | +| Subagent transcript excluded | smoke 65 | its 99999 fee absent | PASS | +| Sidechain-only transcript excluded | smoke 66 | 4 sessions, not 5 | PASS | +| First turn wins, not a later one | smoke 67 | 18000 second turn ignored | PASS | +| Per repo (E2) | smoke 68 | proj-beta 1 session, median 500 | PASS | +| Trend, newest first (E3) | smoke 69 | W37 3/2000 above W36 1/1000 | PASS | +| No trend without `--trend` | smoke 70 | weekly table absent | PASS | +| Bare repo name resolves | smoke 71 | `alpha` -> proj-alpha only | PASS | +| JSON carries the estimate flag (E4) | smoke 72 | `components_estimated` true, median_fee 2000 | PASS | +| `report` unchanged (E5) | smoke 73 | no entry-fee section | PASS | +| Overshoot is not a negative measurement (E4) | smoke 74 + 75 | `(estimate over measured)` 1000, share 200% | PASS | +| Untrusted fields do not abort the scan | smoke 76 + 77 | exit 0, fee 700 still measured | PASS | +| Worktree slugs fold into their repo (E2) | smoke 78 | one proj-alpha row, 2 sessions | PASS | +| Multi-slug `--project` announced | smoke 79 | stderr names both slugs | PASS | +| Live, real data (E1 E2 E3 E6) | `bash bin/session observe entry-fee --days 14 --top 8 --trend` | tables under 5s | PASS, 1.03s wall, 435 sessions, 8 repos | + +### Live run (2026-09-16, 14-day window) + +``` +# entry-fee (435 sessions, 8 projects; median 52951 tokens re-read per turn, measured) + component split of one 94938-token session (108 sessions record the rendered preamble; + the split is ESTIMATED at 4 chars/token, only the totals are measured): + component est-tokens share + ------------------------- ---------- ----- + skill_listing 22523 24% + instructions 12653 13% + agent_listing_delta 6795 7% + hook_success 1579 2% + hook_additional_context 1114 1% + output_style_instructions 882 1% + deferred_tools_delta 411 0% + session_context 308 0% + environment 189 0% + model 39 0% + output_style 31 0% + total_tokens_reminder 21 0% + date 16 0% + (unattributed) 48377 51% + per repo (median measured fee): + project sessions median-fee + ---------------------------------------------- -------- ---------- + -Users-tieubao-workspace-tieubao-family-office 1 112659 + -Users-tieubao-workspace-tieubao-dfoundation 1 107022 + -Users-tieubao-workspace-tieubao-dotfiles 1 104169 + ...ieubao-workspace-dwarvesf-dwarves-kit-queue 1 94938 + -Users-tieubao 3 88987 + -Users-tieubao-workspace-tieubao-ops-toolkit 396 54957 + ...ce-tieubao-ops-toolkit-tools-vps-mon-worker 4 53352 + ...tieubao-ops-toolkit-experiments-webnovel-dl 6 47507 + weekly trend (median measured fee): + week sessions median-fee + -------- -------- ---------- + 2026-W38 42 54957 + 2026-W37 169 105299 + 2026-W36 221 51050 + 2026-W35 2 112171 + 2026-W34 1 113926 +``` + +The two largest sized components, `skill_listing` at 22,523 and `instructions` at 12,653, sit in the same band as the 2026-09-13 hand measurement (26,813 and 18,893) taken on a different repo. The tool now produces that split on demand, per repo and per week, instead of once by hand. + +### Negative control + +`bash lib/gate/negctl.sh` dropped `cache_read_input_tokens` from the fee sum, the one term that carries most of a real preamble. + +``` +Exit: 0 (green before mutation) +Changed: lib/session/observe/bin/session-observe +Exit: 1 (under mutation, RED expected) +Restore: git checkout HEAD -- lib/session/observe/bin/session-observe +Exit: 0 (green after restore) +Verdict: PASS +``` + +### Reproduce + +```bash +bash lib/session/observe/tests/smoke.sh # -> smoke: all 79 passed +bash bin/session observe entry-fee --days 14 --top 8 --trend # live tables +bash bin/session observe entry-fee --days 14 --json # machine-readable +``` From 97def13fd2cc173f560a712f0dbcd015f9f02810 Mon Sep 17 00:00:00 2001 From: Han Ngo Date: Wed, 16 Sep 2026 18:31:32 +0700 Subject: [PATCH 4/5] chore(registry): regenerate the feature registry The new spec raises the spec-reference counts on the /kit:docs and observe rows. Generated file, regenerated not hand-edited. --- docs/FEATURES.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/FEATURES.md b/docs/FEATURES.md index 3954836f..b508cdf2 100644 --- a/docs/FEATURES.md +++ b/docs/FEATURES.md @@ -20,7 +20,7 @@ GENERATED , do not hand-edit. Regenerate: `bash lib/registry/feature-registry.sh | `/kit:design` | `[H/I]` | Opt-in interactive solution-design beat between /think and /spec. Explores 2-3 approaches one question at a time, holds for your approval p… | SPEC-003, SPEC-004, SPEC-005 +105 | test-command-emit-sweep.sh, test-command-triggers.sh, test-design-record.sh +21 | | `/kit:devs-team` | `[H/I]` | Parallel multi-lens critique of a solution design (the active spec if present, else the decision brief). Dispatches 5 engineering lenses, m… | SPEC-016, SPEC-018, SPEC-019 +11 | test-gate-vocab-recording.sh, test-meta.sh, test-outcome-emit-sweep.sh | | `/kit:dispatch` | `[H/I]` | Fire several disjoint VALIDATED specs concurrently, each in its own worktree, then converge. Cross-goal fan-out behind a disjointness gate … | SPEC-002, SPEC-016, SPEC-017 +78 | test-advisor-ledger-emit.sh, test-agent-effectiveness.sh, test-audit-scanner-contract.sh +30 | -| `/kit:docs` | `[H/I]` | Update all project documentation to match the current codebase. Cross-references the diff against every doc file and fixes drift. | SPEC-001, SPEC-002, SPEC-003 +190 | proof-loop-09-scenario-b.sh, run-all.sh, run-workflow.sh +67 | +| `/kit:docs` | `[H/I]` | Update all project documentation to match the current codebase. Cross-references the diff against every doc file and fixes drift. | SPEC-001, SPEC-002, SPEC-003 +191 | proof-loop-09-scenario-b.sh, run-all.sh, run-workflow.sh +67 | | `/kit:draft-agent` | `[H/I]` | Meta-agent agent-builder. From a one-line description, generates a new subagent definition OR a mega-goal sub-goal file and (by default) in… | SPEC-089, SPEC-108, SPEC-139 +1 | test-agent-effectiveness.sh, test-command-emit-sweep.sh, test-meta-agent.sh +1 | | `/kit:execute` | `[H/I]` | Autonomous spec execution with verification. Dispatches worker subagents per task, verifies each with task-verifier, retries fixable failur… | SPEC-001, SPEC-003, SPEC-004 +59 | test-break-it.sh, test-gate-vocab-recording.sh, test-hooks.sh +9 | | `/kit:explain` | `[H/I]` | Turn a merged change into a literate-diff explainer a human READS to understand: background -> goal + intuition -> a prose-ordered diff -> … | SPEC-050, SPEC-060, SPEC-094 +18 | proof-loop-09-scenario-b.sh, test-boundary-lint.sh, test-command-emit-sweep.sh +11 | @@ -99,7 +99,7 @@ GENERATED , do not hand-edit. Regenerate: `bash lib/registry/feature-registry.sh | `get-api-docs` | `[I]` | Fetch curated API documentation using Context Hub (chub) before coding against any external API. Use when the task involves calling a third… | SPEC-285 | test-kit-foldin-hooks.sh | | `loop-engineering` | `[I]` | Use when the user wants to design or add a new bounded loop to the kit's own SDLC orchestration ("let's build a loop", "design a new loop f… | SPEC-209, SPEC-222, SPEC-225 +4 | test-loop-engineering-contract.sh, test-research-arch-contract.sh | | `memory-tidy` | `[I]` | Use when auditing, consolidating, or cleaning a repo's .claude/memory store, "dọn memory", "memory tidy", the biweekly memory audit, dupl… | SPEC-208, SPEC-225, SPEC-249 | test-memory-tidy-contract.sh, test-research-arch-contract.sh | -| `observe` | `[I]` | Query and render the kit's control plane from an agent session. Use when asked about the fleet's runs, gate verdicts, conformance, spend/to… | SPEC-033, SPEC-052, SPEC-128 +11 | test-bin-forwarders.sh, test-kit-contract.sh, test-orchestrate-wavefront.sh +1 | +| `observe` | `[I]` | Query and render the kit's control plane from an agent session. Use when asked about the fleet's runs, gate verdicts, conformance, spend/to… | SPEC-033, SPEC-052, SPEC-128 +12 | test-bin-forwarders.sh, test-kit-contract.sh, test-orchestrate-wavefront.sh +1 | | `repo-hygiene` | `[I]` | Use for the whole-repo decay audit, "run the repo-hygiene loop", "what has rotted in this repo", "audit repo hygiene", "find the stale file… | SPEC-256 | test-repohygiene.sh | | `skill-review` | `[I]` | Review and promote skill drafts that skill-curator staged from past sessions. Use when the user runs /skill-review, says "review my skill d… | SPEC-218, SPEC-285 | test-bin-forwarders.sh, test-kit-contract.sh | | `stats` | `[I]` | Query or render the state of the scattered kit/tide/tg-cleanup/learned ledgers (the dwarves-kit gate/proof/telemetry corpus, tide file-move… | SPEC-001, SPEC-182, SPEC-192 +17 | proof-loop-09-scenario-b.sh, test-bin-forwarders.sh, test-config-stamp.sh +9 | From 0c3503237b858594e2942a3a6b5df284867c27f0 Mon Sep 17 00:00:00 2001 From: Han Ngo Date: Wed, 16 Sep 2026 19:07:34 +0700 Subject: [PATCH 5/5] docs(spec): use the plain Lane header the ship-gate parses --- docs/specs/SPEC-289-observe-entry-fee.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/specs/SPEC-289-observe-entry-fee.md b/docs/specs/SPEC-289-observe-entry-fee.md index adf1ab5a..a3d312c7 100644 --- a/docs/specs/SPEC-289-observe-entry-fee.md +++ b/docs/specs/SPEC-289-observe-entry-fee.md @@ -1,7 +1,7 @@ # SPEC-289: session observe entry-fee **Status**: ready -**Lane**: full +Lane: full **Owner module**: `lib/session/observe` **Source**: backlog row ID-878; measurement `ops-toolkit/research/2026-09-13-token-burn-optimization.md`