Skip to content

feat(session-observe): size the fixed context entry fee per turn - #657

Merged
tieubao merged 7 commits into
masterfrom
feat/observe-entry-fee
Sep 16, 2026
Merged

tieubao merged 7 commits into
masterfrom
feat/observe-entry-fee

Conversation

@tieubao

@tieubao tieubao commented Sep 16, 2026

Copy link
Copy Markdown
Member

Closes ID-878. Spec: docs/specs/SPEC-289-observe-entry-fee.md. Proof: lib/session/observe/docs/proof-of-done.md.

What

Every agent turn re-reads a fixed preamble before any work: the skill listing, the CLAUDE.md stack, the repo memory index, the agent roster, the output style, the tool schemas. A hand measurement on 2026-09-13 put it near 96k tokens, about 38 percent of a 400-turn builder's whole bill. A hand snapshot cannot say whether the number moved, which repo pays the most, or which component grew.

session observe entry-fee [--days N] [--project SLUG-OR-NAME] [--root DIR] [--top N] [--trend] [--json]

The total is measured: the first main-chain assistant turn's input + cache_creation + cache_read, which is exactly what the model read before it acted. The split is estimated: each preamble attachment sized from its rendered text at four characters per token, the same rule of thumb the hand measurement used, labelled an estimate in the header and in the JSON. The gap between the two prints as one (unattributed) row for the system prompt and the built-in tool schemas.

Per-repo rows carry the median measured fee and fold a repo's worktree slugs into one row. --trend adds ISO-week medians, newest first, so a preamble cleanup lands as a visible drop.

Live run, 14-day window, 1.03s

# entry-fee  (435 sessions, 8 projects; median 52951 tokens re-read per turn, measured)
  component split of one 94938-token session (108 sessions record the rendered preamble;
  the split is ESTIMATED at 4 chars/token, only the totals are measured):
  component                  est-tokens  share
  -------------------------  ----------  -----
  skill_listing                   22523    24%
  instructions                    12653    13%
  agent_listing_delta              6795     7%
  hook_success                     1579     2%
  hook_additional_context          1114     1%
  output_style_instructions         882     1%
  (unattributed)                  48377    51%
  per repo (median measured fee):
  project                                         sessions  median-fee
  ----------------------------------------------  --------  ----------
  -Users-tieubao-workspace-tieubao-family-office         1      112659
  -Users-tieubao-workspace-tieubao-dfoundation           1      107022
  -Users-tieubao-workspace-tieubao-dotfiles              1      104169
  ...ieubao-workspace-dwarvesf-dwarves-kit-queue         1       94938
  -Users-tieubao                                         3       88987
  -Users-tieubao-workspace-tieubao-ops-toolkit         396       54957
  weekly trend (median measured fee):
  week      sessions  median-fee
  --------  --------  ----------
  2026-W38        42       54957
  2026-W37       169      105299
  2026-W36       221       51050

The two largest sized components land in the same band as the hand measurement taken on a different repo (26,813 and 18,893).

Design calls worth a look

  • The split is one real session's, not a per-component median. Component medians do not sum to the median fee, so a median-of-medians table would not reconcile against a measured total. That session is the median of the sessions whose transcript records a rendered preamble; older transcripts record none.
  • Four chars per token overshoots on dense markdown. An estimate above the measured fee prints (estimate over measured) with the magnitude and pushes the component shares past 100 percent, rather than printing a negative token count.
  • The measured fee also covers the first user prompt, so it is the preamble plus turn one. The (unattributed) row absorbs that. Stated in the spec and README rather than subtracted, because the transcript gives no way to separate the two.
  • --project now also takes a bare repo name, shared with cost and burn through project_roots(). A multi-slug match announces its resolved list on stderr, so the widened resolution cannot merge unrelated repos in silence.
  • Not part of report. report is the weekly digest of behaviour; this is a standing cost measurement whose per-repo table would double the digest.

Verification

Check Result
bash lib/session/observe/tests/smoke.sh smoke: all 79 passed (18 new cases, 7 of them negative controls)
bash tests/run-all.sh all 147 suites passed, 1 skipped for missing tooling
Negative control (lib/gate/negctl.sh, dropped cache_read_input_tokens from the fee) green, RED under mutation, green after restore, Verdict: PASS
Live run 435 sessions, 8 repos, 1.03s wall

Reviewed by the kit code-reviewer under the architecture and correctness lens before push. It found three unguarded crash paths on untrusted transcript fields, the negative remainder row, and the per-slug ranking; all six findings are fixed in the second commit, each with a smoke case.

Every agent turn re-reads a fixed preamble before any work: the skill
listing, the CLAUDE.md stack, the repo memory index, the agent roster,
the output style, the tool schemas. A hand measurement put it near 96k
tokens, about 38 percent of a 400-turn builder's whole bill, but a hand
snapshot cannot say whether the number moved or which repo pays most.

The new `entry-fee` view measures the total and estimates the split.
The total is the first main-chain assistant turn's input plus cache
creation plus cache read, which is exactly what the model read before
it acted. The split sizes each preamble attachment from its rendered
text at four characters per token and labels itself an estimate; the
gap against the measured total prints as one (unattributed) row for the
system prompt and the built-in tool schemas.

Per-repo rows carry the median measured fee, since the memory index and
the CLAUDE.md stack differ per checkout. --trend adds ISO-week medians,
newest first, so a preamble cleanup lands as a visible drop.

The split comes from one real session, not a per-component median:
component medians do not sum to the median fee, so the table would not
reconcile against a measured total.

--project now also resolves a bare repo name against every slug that
contains it, shared through project_roots() with cost and burn, because
one repo's worktrees each own a separate slug.
…ree slugs

The architecture and correctness review on this branch found three
unguarded crash paths and two outputs that read as measurements when
they were estimates.

A numeric timestamp, a numeric rendered content, and a dict attachment
type each aborted the whole scan rather than that one file, while every
sibling collector in the module already guards the same shapes. Each now
contributes nothing and the scan continues.

The remainder row went negative when the four-characters-per-token
estimate overshot the measured fee, printing a negative token count and
a negative share. It now prints (estimate over measured) with the
magnitude, and the component shares above 100 percent carry the signal.

The per-repo table keyed on the raw project slug, so a repo's worktrees
each ranked as a separate repo and a one-session worktree sorted above
the many-session checkout of the same codebase. Rows fold on the
worktree marker in the slug and break median ties on session count.

A multi-slug --project match now announces its resolved list on stderr,
so the widened resolution cannot merge unrelated repos in silence.

pctl and entry_fee_median share one nearest-rank index, so the session
shown for the median split cannot drift from the median it is shown for.
Records the 18 smoke cases with their 7 negative controls, the live
14-day run over 435 sessions across 8 repos, and the negctl verdict.

The negative control dropped cache_read_input_tokens from the fee sum,
the term that carries most of a real preamble. The suite went red under
the mutation and green again after the restore.

The two largest sized components on the live run, the skill listing at
22,523 and the CLAUDE.md stack at 12,653, land in the same band as the
2026-09-13 hand measurement taken on a different repo. The split is now
available on demand, per repo and per week.

Flips ID-878 to shipped.
The new spec raises the spec-reference counts on the /kit:docs and
observe rows. Generated file, regenerated not hand-edited.
# Conflicts:
#	_meta/BACKLOG.md
#	docs/FEATURES.md
@tieubao
tieubao merged commit 4c94de3 into master Sep 16, 2026
1 check passed
@tieubao
tieubao deleted the feat/observe-entry-fee branch September 23, 2026 07:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant