Skip to content

Latest commit

 

History

History
183 lines (141 loc) · 8.48 KB

File metadata and controls

183 lines (141 loc) · 8.48 KB

Design notes

The technical writeup: what lean is, what "net negative" means, how it was built, and where it differs from what already exists. Written so a future contributor (or future me) can pick it up cold.

1. What problem this solves

An AI coding agent, left alone, answers verbosely. It opens with "Sure, I'd be happy to help", restates your question, wraps the real answer in three sentences of hedging, and closes with a summary paragraph. Every one of those words is an output token you pay for and then scroll past.

Compressing that narration is a good idea. The trick is doing it without lying about the payoff and without charging you more than you save.

2. What "net negative" means

This is the core concept, so here it is plainly.

A skill like this is not free. To make the agent talk densely, you inject instructions into the model's context on every turn. That injected prompt costs input tokens. Call that the overhead.

On the other side, the skill saves you output tokens by making the reply shorter. Call that the benefit.

Net result = benefit minus overhead.

  • Net positive: the reply had lots of prose, you saved more than the prompt cost. Good.
  • Net negative: the overhead was larger than what you saved, so the skill made the turn MORE expensive, not less.

When does net negative happen? On turns where there was almost nothing to compress in the first place:

  1. Short answers. "Yes, run php artisan migrate." There are no filler paragraphs to cut, but you still paid the prompt overhead.
  2. Code heavy turns. A reply that is mostly a diff or a code block. The code must stay exact, so the compressor correctly leaves it alone. Again you paid overhead for near zero benefit.
  3. Tool call turns. The agent is editing files and running commands with little narration between. Nothing to shrink.

This matters because real coding sessions are full of exactly these turns. That is the whole reason a "65 percent" claim collapses to single digits in practice: the 65 was measured on long chat style prose, but coding output is mostly the stuff a good tool refuses to touch.

lean's answer to net negative is a gate (see section 4). The skill is told to do nothing on short, code heavy, or tool call turns, and only engage when there is genuine prose worth compressing. It cannot win those turns, so it does not try, and it does not spend anything trying.

3. How I approached it (research first)

  1. Read the existing popular skill end to end to understand its design: a prompt that tells the agent to drop filler, plus a local counter that prints a savings number, plus memory file compression.
  2. Searched for independent measurement. Found testing showing the real world saving on agentic tasks is roughly 8 percent, not the advertised 65, because agent output is dominated by verbatim code and tool calls.
  3. Searched the Laravel and PHP side. Found there are already good best practice skills for Laravel (official and community), but none of them is a token efficiency tool. So efficiency plus framework awareness is an open space, not a crowded one.
  4. Turned the three failure modes into three concrete design requirements, then built the smallest thing that satisfies all three.

4. How it is built

Four moving parts, all small.

a. The skill prompt (the always on rule)

skills/lean/SKILL.md. The rule the agent follows every turn. Kept to about 200 tokens on purpose, because this is the overhead from section 2. The less it costs, the harder it is to go net negative. Everything that is not the core rule lives in skills/lean/REFERENCE.md, which the agent only opens when asked. That keeps the per turn cost down.

b. The gate (the net negative fix)

Rule 3 in the skill. In plain terms: compress only when the reply is mostly prose and longer than about 40 words. Short, code heavy, and tool call turns are answered normally with no added scaffolding. This is a prompt instruction, not code, because the model is the thing deciding how to write each reply.

c. The Laravel layer

Rule 4 plus the conventions in REFERENCE.md. Keep framework terms exact, assume the reader knows the framework, and answer a bug in the shape Class::method, then cause, then fix. This is what makes lean useful for PHP work specifically rather than being a generic style filter.

d. The measurement tool (the honesty fix)

bin/lean-stats.mjs. This is where lean earns the word "honest".

  • It reads the real output_tokens field that the API wrote into the Claude Code session transcript. That is ground truth from the provider, not a chars divided by four estimate.
  • report <session.jsonl> prints the real output tokens for a session and, crucially, what fraction of that output is even prose. That fraction is the hard ceiling on what any compressor could possibly save. It claims no percentage of its own.
  • diff <baseline.jsonl> <lean.jsonl> is the real A/B. Run the same prompts once with lean off and once with lean on. The tool matches turns by their triggering prompt and reports the measured token difference.
  • If a turn has no usage field, the tool falls back to an approximation and labels it "approx" so nobody mistakes it for a measurement.

No lifetime counter, no invented headline number. If it cannot measure, it says so.

e. The memory file compressor (input savings)

commands/lean-compress.md plus bin/lean-compress.mjs. Everything above shrinks output. This attacks the input side. A CLAUDE.md or project notes file loads on every session, so compressing it once saves input tokens on every session after.

The rewrite itself is a model task, done by the command: keep every rule and fact, cut only the filler prose, and preserve all code, paths, URLs, and commands byte for byte. The script is the safety net that a model should not self certify. lean-compress check <original> <compressed> extracts every code block, inline code span, URL, and command from the original and confirms each still appears verbatim in the rewrite. If any is missing it exits non zero so the rewrite is rejected. It reports the exact character reduction and an honestly labeled token estimate. The command always backs up to a .bak first, so a failed check loses nothing.

f. Packaging

  • .claude-plugin/plugin.json and .claude-plugin/marketplace.json make it installable as a Claude Code plugin.
  • commands/lean.md, commands/lean-stats.md, commands/lean-compress.md are the slash commands.
  • .cursor/rules/lean.mdc is the same rule in Cursor and Windsurf format.
  • install.sh detects which agents you have and copies the right files.

5. How lean differs from what is out there

  1. Versus a prompt only compression skill: same benefit (dense replies), but a gate so it does not go net negative, a much smaller per turn prompt, and a measurement tool that reports real tokens instead of an advertised figure.
  2. Versus Laravel best practice skills (official and community): those teach the agent how to write good Laravel code. lean is orthogonal. It shapes how densely the agent talks, and it happens to preserve Laravel vocabulary so the two can run together.
  3. Versus doing nothing: lean only ever engages when there is prose to compress, so the downside is capped by design.

6. Honest limits

  1. The always on skill shrinks output tokens only. Reasoning tokens are untouched. Input savings come only from the separate memory file compressor, and only for files you choose to compress.
  2. On code heavy sessions the real saving is small, and lean will show you that rather than hide it.
  3. The gate depends on the model following the rule. It is a prompt, not a hard filter. The measurement tool exists precisely so you can check it kept its word.

7. Roadmap

Done:

  1. The output skill, the gate, and the Laravel layer.
  2. Verified measurement (lean-stats, report and diff modes).
  3. The memory file compressor and its preservation verifier (lean-compress).
  4. A reproducible 10 prompt Laravel benchmark under benchmarks/laravel-suite, with a generator, both transcripts, and the diff output recorded.

Not done yet:

  1. A session start hook for Claude Code so lean is on from the first message without typing the command.
  2. A live API measured run of the benchmark, to replace the authored estimate with a real [measured] number.
  3. Publish to the Claude Code plugin marketplace and the skills registry so install is one line.
  4. Optional: an MCP wrapper that compresses tool descriptions, which attacks the input side instead of the output side.