The technical writeup: what lean is, what "net negative" means, how it was built, and where it differs from what already exists. Written so a future contributor (or future me) can pick it up cold.
An AI coding agent, left alone, answers verbosely. It opens with "Sure, I'd be happy to help", restates your question, wraps the real answer in three sentences of hedging, and closes with a summary paragraph. Every one of those words is an output token you pay for and then scroll past.
Compressing that narration is a good idea. The trick is doing it without lying about the payoff and without charging you more than you save.
This is the core concept, so here it is plainly.
A skill like this is not free. To make the agent talk densely, you inject instructions into the model's context on every turn. That injected prompt costs input tokens. Call that the overhead.
On the other side, the skill saves you output tokens by making the reply shorter. Call that the benefit.
Net result = benefit minus overhead.
- Net positive: the reply had lots of prose, you saved more than the prompt cost. Good.
- Net negative: the overhead was larger than what you saved, so the skill made the turn MORE expensive, not less.
When does net negative happen? On turns where there was almost nothing to compress in the first place:
- Short answers. "Yes, run
php artisan migrate." There are no filler paragraphs to cut, but you still paid the prompt overhead. - Code heavy turns. A reply that is mostly a diff or a code block. The code must stay exact, so the compressor correctly leaves it alone. Again you paid overhead for near zero benefit.
- Tool call turns. The agent is editing files and running commands with little narration between. Nothing to shrink.
This matters because real coding sessions are full of exactly these turns. That is the whole reason a "65 percent" claim collapses to single digits in practice: the 65 was measured on long chat style prose, but coding output is mostly the stuff a good tool refuses to touch.
lean's answer to net negative is a gate (see section 4). The skill is told to do nothing on short, code heavy, or tool call turns, and only engage when there is genuine prose worth compressing. It cannot win those turns, so it does not try, and it does not spend anything trying.
- Read the existing popular skill end to end to understand its design: a prompt that tells the agent to drop filler, plus a local counter that prints a savings number, plus memory file compression.
- Searched for independent measurement. Found testing showing the real world saving on agentic tasks is roughly 8 percent, not the advertised 65, because agent output is dominated by verbatim code and tool calls.
- Searched the Laravel and PHP side. Found there are already good best practice skills for Laravel (official and community), but none of them is a token efficiency tool. So efficiency plus framework awareness is an open space, not a crowded one.
- Turned the three failure modes into three concrete design requirements, then built the smallest thing that satisfies all three.
Four moving parts, all small.
skills/lean/SKILL.md. The rule the agent follows every turn. Kept to about
200 tokens on purpose, because this is the overhead from section 2. The less
it costs, the harder it is to go net negative. Everything that is not the core
rule lives in skills/lean/REFERENCE.md, which the agent only opens when
asked. That keeps the per turn cost down.
Rule 3 in the skill. In plain terms: compress only when the reply is mostly prose and longer than about 40 words. Short, code heavy, and tool call turns are answered normally with no added scaffolding. This is a prompt instruction, not code, because the model is the thing deciding how to write each reply.
Rule 4 plus the conventions in REFERENCE.md. Keep framework terms exact,
assume the reader knows the framework, and answer a bug in the shape
Class::method, then cause, then fix. This is what makes lean useful for PHP
work specifically rather than being a generic style filter.
bin/lean-stats.mjs. This is where lean earns the word "honest".
- It reads the real
output_tokensfield that the API wrote into the Claude Code session transcript. That is ground truth from the provider, not a chars divided by four estimate. report <session.jsonl>prints the real output tokens for a session and, crucially, what fraction of that output is even prose. That fraction is the hard ceiling on what any compressor could possibly save. It claims no percentage of its own.diff <baseline.jsonl> <lean.jsonl>is the real A/B. Run the same prompts once with lean off and once with lean on. The tool matches turns by their triggering prompt and reports the measured token difference.- If a turn has no usage field, the tool falls back to an approximation and labels it "approx" so nobody mistakes it for a measurement.
No lifetime counter, no invented headline number. If it cannot measure, it says so.
commands/lean-compress.md plus bin/lean-compress.mjs. Everything above
shrinks output. This attacks the input side. A CLAUDE.md or project notes
file loads on every session, so compressing it once saves input tokens on
every session after.
The rewrite itself is a model task, done by the command: keep every rule and
fact, cut only the filler prose, and preserve all code, paths, URLs, and
commands byte for byte. The script is the safety net that a model should not
self certify. lean-compress check <original> <compressed> extracts every
code block, inline code span, URL, and command from the original and confirms
each still appears verbatim in the rewrite. If any is missing it exits non
zero so the rewrite is rejected. It reports the exact character reduction and
an honestly labeled token estimate. The command always backs up to a .bak
first, so a failed check loses nothing.
.claude-plugin/plugin.jsonand.claude-plugin/marketplace.jsonmake it installable as a Claude Code plugin.commands/lean.md,commands/lean-stats.md,commands/lean-compress.mdare the slash commands..cursor/rules/lean.mdcis the same rule in Cursor and Windsurf format.install.shdetects which agents you have and copies the right files.
- Versus a prompt only compression skill: same benefit (dense replies), but a gate so it does not go net negative, a much smaller per turn prompt, and a measurement tool that reports real tokens instead of an advertised figure.
- Versus Laravel best practice skills (official and community): those teach the agent how to write good Laravel code. lean is orthogonal. It shapes how densely the agent talks, and it happens to preserve Laravel vocabulary so the two can run together.
- Versus doing nothing: lean only ever engages when there is prose to compress, so the downside is capped by design.
- The always on skill shrinks output tokens only. Reasoning tokens are untouched. Input savings come only from the separate memory file compressor, and only for files you choose to compress.
- On code heavy sessions the real saving is small, and lean will show you that rather than hide it.
- The gate depends on the model following the rule. It is a prompt, not a hard filter. The measurement tool exists precisely so you can check it kept its word.
Done:
- The output skill, the gate, and the Laravel layer.
- Verified measurement (
lean-stats, report and diff modes). - The memory file compressor and its preservation verifier (
lean-compress). - A reproducible 10 prompt Laravel benchmark under
benchmarks/laravel-suite, with a generator, both transcripts, and the diff output recorded.
Not done yet:
- A session start hook for Claude Code so lean is on from the first message without typing the command.
- A live API measured run of the benchmark, to replace the authored estimate
with a real
[measured]number. - Publish to the Claude Code plugin marketplace and the skills registry so install is one line.
- Optional: an MCP wrapper that compresses tool descriptions, which attacks the input side instead of the output side.