Skip to content

Latest commit

 

History

15 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

plumbline

A plumb line is the oldest tool that proves something is truly vertical. This plugin is the same thing for agent-built features: it proves a slice runs end to end instead of taking "cut vertically" on trust.

Eight gates, eleven skills, one rule that stops the backlog from growing once per finished ticket. MIT.

once per effort:  idea -> shape -> spec -> slice ....... then close
once per ticket:  build -> review -> verdict -> the next ticket

The problem

Coding agents do not usually fail by writing bad code. They fail by building correct components that were never made to meet.

Measured on a real, non-public codebase: 292 commits, 165 tickets in a single effort. Classifying every ticket titled like "X is read by nothing", "nothing registers Y", "Z has no caller":

Ticket range Wiring defects Total Share
1–60 1 60 1.7 %
61–110 4 50 8.0 %
111–165 35 55 63.6 %

The agent made no gross mistakes. Every component was built exactly as its ticket described. The tickets each named one end of a seam, so nobody was ever forced to run producer and consumer together — until the first real run, and then 35 defects arrived at once.

The second failure mode is the mirror image: every review finding becomes a new ticket, because no rule says it shouldn't. A cross-check against a second codebase that did have such a rule: three efforts, 24 / 7 / 1 tickets, no explosion.

Full measurement with evidence: docs/DIAGNOSIS.md.

What it does

Two mechanics carry the whole plugin:

1. A ticket cannot be written without both ends of its seam. PRODUCES, CONSUMED BY, CONSUMES. An empty field is a visible error, and seam-check greps the literal the ticket named. "Slice vertically" as a sentence already exists elsewhere — it doesn't work, because nothing checks it.

2. Every finding gets exactly one verdict. Six options: fix now, add as an acceptance criterion, send back to the ticket that shipped it, already owned by an open ticket, cut a new ticket, or decline in writing. Severity decides: a blocker never becomes a deferred criterion, and a new ticket is only allowed when the work blocks something or needs a decision nobody in the effort owns.

Install

# once
/plugin marketplace add maajix/Plumbline
/plugin install plumbline@plumbline

Local checkout instead:

claude plugin marketplace add /path/to/plumbline   # this folder
claude plugin install plumbline@plumbline

After editing the plugin, bump version in .claude-plugin/plugin.json, then:

claude plugin marketplace update plumbline
claude plugin update plumbline@plumbline

Without a version bump the plugin cache keeps the old copy.

No install needed to read it: every skill is a plain Markdown file under skills/.

The flow

# Command Skill Produces
1 /shape shape-idea interview.md, shape.md, a stub shape per newly split-off sub-problem
2 /spec write-spec spec.md + cold read
3 /slice cut-slices NN-*.md tickets + cold read
4 /build build-slice code + seam test — per ticket, one session
5 /prove seam-check reads the seam report — the pass itself already ran inside /build §5
6 /review-pass review-pass findings, with severity — per ticket, right after its build
7 /verdict hold-the-line one of six verdicts — per finding
8 /close close-effort the effort ends instead of drifting

Gates 1–3 run once for the whole effort. Gates 4–7 are a loop, once per ticket: build it, review that one ticket, give every finding a verdict, then the next ticket. Do not build all the tickets and review them together — gate 6 pins its diff to one ticket's first build commit, so three tickets in the diff is two tickets of noise on every axis.

Gate 5 is the exception: it runs on its own, as step 5 of build-slice. You normally never type /prove — the command is there to re-read that report later, and re-running the full pass per ticket is exactly what the gate is written to avoid.

Every command is also reachable namespaced: /plumbline:build, /plumbline:review-pass. The review command is deliberately called /review-pass, because /review is already a built-in Claude Code command and a plugin does not shadow it.

Ticket IDs include their origin effort: billing/01. They survive moves, and build/review commits end with (ticket billing/01). Numeric command arguments still work within one unambiguous effort. Old tickets are upgraded when used, with path-scoped history checks; no old commits are rewritten.

A ticket only becomes resolved once its review cycle closes: build-slice takes it to claimed through build, bar and commit, and review-pass writes resolved once no finding still carries — verdict pending, all criteria are ticked, and no executable change followed the reviewed head. Details: docs/WORKFLOW.md.

All artefacts of an effort live under docs/issues/<effort>/: interview.md, shape.md, spec.md, live-inputs.md, and the tickets as NN-<slug>.md. A sub-problem a shaping interview splits off gets its own slug: docs/issues/<its-slug>/shape.md.

One effort, three tickets

What you actually type, start to finish.

/shape   billing needs a dunning email when a card fails
/spec
/slice                       # -> 01-*.md, 02-*.md, 03-*.md

Then the loop, once per ticket:

/build 01                    # ticket 01 is the walking skeleton: it runs end to end
/review-pass                 # no argument: takes the ticket from the newest build commit
/verdict                     # one verdict per finding, until none says "— verdict pending"

When that cycle closes clean, review-pass writes **Status:** resolved itself and makes the review commit. If a verdict requires executable work (NOW included), the review first commits its findings and an unticked criterion. Run /build 01 for the repair, then /review-pass again. Three cycles is the cap; remaining repairs at cycle three block for a decision. Prose/comment-only repairs may ride the review commit.

/build                       # no argument: 01 is resolved, so the frontier is 02
/review-pass
/verdict

/build                       # 03
/review-pass
/verdict

/close                       # settles forward references, exit criteria, statuses

That is the whole flow. Three things worth knowing while it runs:

  • The ticket stays claimed through the build. That is deliberate, not a missed step: a required finding can still land on it as a criterion, even for the last ticket of the effort. Only review-pass writes resolved.
  • A session that dies mid-ticket costs you nothing. build-slice leaves a ## Handoff block and the ticket stays claimed; the next session just runs /build, because §0 takes every claimed ticket before every open one. Handing over is not replanning — re-cutting belongs to /verdict, after the review.
  • New work found mid-flight goes through /verdict, never straight into a new ticket file. That gate is the one thing keeping the backlog from growing once per finished ticket.
  • The review runs in a fresh session. build-slice ends by saying so: the ticket file carries the state, and a builder's context should not grade its own verdicts.

The commit guard

The plugin ships one hook: a PreToolUse guard on the Bash tool that blocks a git commit when the state under docs/issues/ contradicts a rule the flow already carries. All four checks are greps over that folder — the hook knows nothing about your toolchain.

Blocked when The rule it enforces
a changed ticket file still carries a — verdict pending entry every finding gets a verdict before the commit
a ticket adds a ## Bar, heading while a criterion is unticked, a ## Handoff stands, or no ## Resolution exists the standing bar's per-ticket lines 1 and 6
**Status:** resolved is added outside a review commit only a closing review cycle writes that value
a spec gains ## Close, while a ticket in its folder is neither resolved nor declined the close walk over ticket statuses

The guard checks one staged tree, not the working tree. Stage the intended files, then invoke commit in a separate Bash call. Simple git commit -m … or -F <file>, git -C <repo>, rtk git, rtk proxy git and a plain env prefix are supported. Quoted multiline messages work. Combined commands, substitutions, alternate Git environments, file arguments and index-changing commit options such as -a are refused with a retry instruction.

This protects the documented Claude workflow, not arbitrary scripts, aliases, external commits, concurrent index writers or native Git hooks that alter the index. Fenced text is stripped from complete file snapshots. In-progress merges, rebases, cherry-picks and reverts are skipped. Expected command/Git failures block; unexpected internal errors remain fail-open with a diagnostic.

To switch it off, disable this plugin's hooks in your settings, or commit outside the session. The check that proves it works: bash hooks/guard-commit.test.sh.

The four rules

1. The first ticket built runs. The thinnest end-to-end path, real, staying in the code. If you cannot state it as "X goes in the front, Y comes out the back", you have not understood the effort yet. This rule beats rule 4.

2. Every ticket names both ends of its seam. Including the last hop, which is usually a human: operator, via <command> is a legal form and owes the same proof as any other seam. A forward reference (CONSUMED BY: ticket NN) is a debt with an address, redeemed by the session that lands NN.

3. Every finding gets exactly one verdict. Severity decides, not convenience. Without this rule the backlog grows once per finished ticket.

4. Everything is capped. Three question rounds when shaping, 25 tickets per effort, three review cycles. A cap that gets hit is a signal to split — never a reason to raise the cap.

Skills

The eight gates:

Skill What it is for
shape-idea Product questions before the spec. Technical unknowns are parked as agent work, never asked of the user.
write-spec A spec organised by paths that run, not components that exist.
cut-slices Spec to tickets that run. Walking skeleton, seam fields, caps.
build-slice One ticket, one session. Also covers the ticket turning out wrong mid-flight.
seam-check Proves what a change writes is actually read, and what it reads is actually written.
review-pass Four axes, cold read, capped at three cycles.
hold-the-line A verdict for every finding. Six options, severity decides.
close-effort Ends an effort instead of letting it drift.

Three base skills the flow calls are bundled, so the plugin runs without prerequisites. They are useful standalone:

Skill Where the flow needs it
cold-read Gates 2, 3, 5, 6. Full pass, then findings, then repair.
price-the-wall Gate 4, the moment "that's not possible" shows up.
structured-debugging Gate 4, only when the cause is unknown.

Shared identity and state transitions: references/ticket-lifecycle.md. A successor reuses a verified end-to-end path and its verify command; only a new path needs a new skeleton. Close records the handover and next command.

Shared quality bar for all of them: references/standing-bar.md.

When not to use it

Built for feature-sized efforts: several layers, an observable far end, enough tickets that a backlog could explode at all. For a two-liner the ceremony costs more than the work — there, build-slice §2/§5 (red test, seam proof) and hold-the-line on their own are enough.

Status

Run the commit-guard regressions with bash hooks/guard-commit.test.sh. The workflow has also been exercised in isolated Claude Code sessions; agent judgement is still required, not mechanically guaranteed.

A seam test proves two ends talk. It never proves they tell the truth — which is why cut-slices rule 3b makes the ticket name what checks the real thing.

Credits and prior art

This flow is built on two existing skill sets and would not exist without them.

mattpocock/skills — Matt Pocock's skill set is the foundation. The overall shape of the chain (idea → spec → tickets → review) and the two-axis review come from there, as does the habit of small, composable, hackable skills. The measurements in docs/DIAGNOSIS.md are partly a critique of specific skills in that set, and that critique is only possible because the set is public, readable and precise enough to argue with.

addyosmani/agent-skills — Addy Osmani's set contributes three mechanics the chain otherwise lacks: the stop-the-line rule when something breaks, the hard cycle limit from doubt-driven-development, and the enumeration of the ways a quality bar gets quietly lowered, which references/standing-bar.md is a direct descendant of.

What is new here is the seam rules and the verdict rule, both of which come out of the measurements in docs/DIAGNOSIS.md rather than from either set.

cold-read, price-the-wall and structured-debugging are standalone skills by the same author, bundled here so the plugin runs without prerequisites.

License

MIT — see LICENSE.

docs/DIAGNOSIS.md and docs/WORKFLOW.md are written in German. The skills, commands and this README are English.

About

Claude Code plugin that makes coding agents cut vertical slices and prove them.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages