Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@ For general loop safety guidance, see [docs/safety.md](docs/safety.md).
- [ ] No auto-merge without explicit allowlist
- [ ] MCP connectors use least privilege
- [ ] `loop-run-log.md` or equivalent observability
- [ ] Guardrails shown to fire, not just configured: `loop-drill . --record`, with `loop-drill.json` committed

## Supported versions

Expand Down
2 changes: 1 addition & 1 deletion docs/QUICKSTART.md
Original file line number Diff line number Diff line change
Expand Up @@ -282,7 +282,7 @@ Commit the scaffold + first run update so `loop-audit` sees activity on the next
|------|---------|
| End of week one | Re-run `loop-audit . --suggest` — aim for L1 (score ~40+) |
| Week two | Add a verifier skill; try one assisted fix in a worktree (L2) — see [loop-worktree](#l2-isolated-fix-attempts-loop-worktree) below |
| Before unattended (L3) | `loop-budget.md` + `loop-run-log.md` filled, human gates in `LOOP.md`, proven runs |
| Before unattended (L3) | `loop-budget.md` + `loop-run-log.md` filled, human gates in `LOOP.md`, proven runs, and `gate.yaml` proven with `npx @cobusgreyling/loop-drill . --record` (commit `loop-drill.json`) |
| Unsure which pattern | [pattern-picker.md](./pattern-picker.md) · [loop-design-checklist.md](./loop-design-checklist.md) |
| Something broke | [failure-modes.md](./failure-modes.md) · [stories/](../stories/) |

Expand Down
2 changes: 1 addition & 1 deletion docs/operating-loops.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ npx @cobusgreyling/loop-cost --pattern <id> --cadence <interval> --level L1
npx @cobusgreyling/loop-init . --pattern <id> # scaffolds loop-budget.md + loop-run-log.md + loop-budget skill
```

`loop-audit` scores cost observability and caps L3 until budget + run log + LOOP.md budget section exist.
`loop-audit` scores cost observability and caps L3 until budget + run log + LOOP.md budget section exist. It also caps L3 until the guardrails are proven: `npx @cobusgreyling/loop-drill . --record` drills `gate.yaml`, and the committed `loop-drill.json` must match the current policy ([loop-audit: present is not proven](../tools/loop-audit/README.md#present-is-not-proven)).

Rough planning factors:

Expand Down
135 changes: 135 additions & 0 deletions loop-drill.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,135 @@
{
"schema": 1,
"tool": "@cobusgreyling/loop-drill",
"guardrails": {
"breaker": {
"recordedAt": "2026-09-28T15:17:23.617Z",
"results": [
{
"id": "breaker.stagnation",
"failureMode": "Infinite Fix Loop",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "breaker.no-progress",
"failureMode": "Infinite Fix Loop",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "breaker.token-budget",
"failureMode": "Token Burn",
"direction": "sensitivity",
"outcome": "skipped",
"detail": "No tokenBudget configured — nothing caps spend mid-run. See loop-budget.md."
},
{
"id": "breaker.healthy",
"failureMode": "Infinite Fix Loop",
"direction": "specificity",
"outcome": "passed"
}
]
},
"gate": {
"recordedAt": "2026-09-28T15:17:23.617Z",
"results": [
{
"id": "gate.denylist[**/.env]",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.denylist[**/.env.*]",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.denylist[**/secrets/**]",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.denylist[**/credentials/**]",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.denylist[**/*_key*]",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.denylist[**/*_secret*]",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.denylist[**/.terraform/**]",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.denylist[**/k8s/production/**]",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.denylist[**/migrations/**]",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.denylist[**/auth/**]",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.denylist[**/payments/**]",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.denylist[**/billing/**]",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.file-count",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.auto-merge",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "sensitivity",
"outcome": "passed"
},
{
"id": "gate.benign",
"failureMode": "Over-Reach (Wrong Scope)",
"direction": "specificity",
"outcome": "passed"
}
],
"input": {
"file": "gate.yaml",
"sha256": "151e4609f03384e705735874db2ddce3b73ebd136b20888b7281d4ed43f9d53b"
}
}
}
}
16 changes: 16 additions & 0 deletions scripts/ci-audit-gates.sh
Original file line number Diff line number Diff line change
Expand Up @@ -82,6 +82,22 @@ ROOT_AUDIT_FILE="$ROOT_AUDIT_FILE" node -e '
console.error("Reference score below L2 threshold (58). Restore dogfood signals: STATE.md, skills/, AGENTS.md.");
process.exit(2);
}
// loop-drill.json is the proof loop-audit scores L3 on. ci-validate-gates.sh
// re-runs the drills, so a committed record cannot claim more than they show;
// this catches a gate.yaml edited without re-recording.
const proof = data.signals.proof || {};
const problems = [];
if (!proof.present) problems.push("loop-drill.json is missing");
if (proof.error) problems.push("loop-drill.json is unreadable: " + proof.error);
if ((proof.stale || []).length) problems.push("stale for " + proof.stale.join(", "));
if ((proof.failed || []).length) problems.push("failing: " + proof.failures.join(", "));
if (!problems.length && !(proof.proven || []).includes("gate")) problems.push("gate is not proven");
if (problems.length) {
console.error("Reference guardrails are not proven (" + problems.join("; ") + ").");
console.error("Re-record from the repo root: (cd tools/loop-drill && npm ci && npm run build) && node tools/loop-drill/dist/cli.js . --record");
process.exit(2);
}
console.log("Reference guardrails proven: " + proof.proven.join(", "));
'

if [[ -n "${LOOP_AUDIT_OUTPUT_FILE:-}" ]]; then
Expand Down
6 changes: 6 additions & 0 deletions starters/changelog-drafter/.claude/agents/verifier.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,9 @@
---
name: changelog-verifier
description: Independent checker for release note drafts produced by the changelog drafter. Accuracy and completeness focused.
model: inherit
---

You are the independent verifier for the Changelog Drafter loop.

Review the draft against the raw scan data. Flag any hallucinated items, missing breaking/security notes, or tone problems.
Expand Down
13 changes: 13 additions & 0 deletions tools/loop-audit/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,19 @@

All notable changes to `@cobusgreyling/loop-audit` are documented here.

## [Unreleased]

### Changed
- **Placeholders no longer score.** Empty or whitespace-only files, and `{}` / `[]` JSON, earn nothing and are listed under *Not counted*. A repo of 19 empty files (142 bytes) used to score 100/L3; it now scores 72/L1
- A skill counts only with a `SKILL.md` carrying `name` + `description` frontmatter (bare directories used to count); the same for Claude Code verifier agents
- `gate.yaml` must declare `version: 1` and a `denylist:`; one `loop-gate` would refuse is a failure, not a signal
- `.github/` and workflows count only when they contain a non-empty file
- **L3 requires proven guardrails**: a `loop-drill.json` (from `loop-drill --record`) showing the current `gate.yaml` passing its drills in both directions, with no recorded guardrail failing. Repos that were L3 on files alone drop to L2 until they record

### Added
- `signals.proof`: guardrails that are proven, failed, untested or stale, read from `loop-drill.json`
- A guardrail that `loop-drill` shows failing loses its points: gate → `gateYaml`, verifier → `verifier` (Verifier Theater), breaker → stall detection

## [1.9.0] - 2026-08-28

### Changed
Expand Down
32 changes: 31 additions & 1 deletion tools/loop-audit/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,14 +76,44 @@ npm publish --access public
| Human-escalation path | LOOP.md / safety docs define when to stop and hand off to a human |
| **loopActivity (v1.4)** | **Dynamic proof**: "Last run" timestamps in state, loop-related git commits, scheduled workflows, run logs |
| **Harness Runtime (v1.7)** | `.foundry/stack.yaml`, lock, sessions/traces, outerloop emit, host integrate — LE → [harness-foundry](https://github.com/cobusgreyling/harness-foundry) funnel |
| **Guardrail proof** | `loop-drill.json` from [`loop-drill --record`](../loop-drill): the gate, breaker and verifier shown firing — see below |

When score ≥ 80 and no `.foundry/stack.yaml`, audit recommends:

```bash
npx @cobusgreyling/loop-init . --with-foundry
```

L3 requires verifier + state + cost observability (budget + run log + LOOP.md budget) **and** proven loop activity (not just files on disk).
L3 requires verifier + state + cost observability (budget + run log + LOOP.md budget), proven loop activity (not just files on disk), **and proven guardrails**: a `loop-drill.json` showing the current `gate.yaml` passing its drills, with no recorded guardrail failing.

## Present is not proven

Every signal counts content, not filenames:

- **Empty files don't score.** A file that is empty or whitespace, or JSON that is just `{}` / `[]`, is a placeholder. The audit lists each one under *Not counted* instead of crediting it.
- **Skills must load.** A skill is a directory with a `SKILL.md` that has `name` and `description` frontmatter, which every host needs before it will invoke it. The same goes for a Claude Code verifier agent. A bare directory, or a file with no frontmatter, doesn't count, whatever it's called.
- **`gate.yaml` must be a policy.** It needs `version: 1` and a `denylist:`. Anything else would be refused by `loop-gate`, so it earns nothing and is reported as a failure.

Files can only show a guardrail is configured. [`loop-drill`](../loop-drill) shows whether it fires, by running it against a seeded fault and a benign case. Record the results and commit them:

```bash
npx @cobusgreyling/loop-drill . --record # gate + breaker: offline, no tokens
npx @cobusgreyling/loop-drill . --only verifier --verifier-cmd "npm test" --setup "npm ci" --record
git add loop-drill.json
```

The audit reads the record per guardrail:

| Result | Meaning | Effect |
|---|---|---|
| **proven** | A drill caught the fault **and** a drill let the benign case through, and nothing failed | Counts. The gate being proven is required for L3 |
| **failed** | A drill failed: the guardrail didn't fire, or blocked the benign case | The guardrail's points are withdrawn (gate → `gateYaml`, verifier → `verifier`, breaker → stall detection), and L3 is blocked |
| **untested** | Every drill was skipped, or only one direction passed | Not proven. For the gate this means `loop-gate` couldn't drill it, so `gateYaml` is withdrawn |
| **stale** | The gate was drilled against a different `gate.yaml`, or a canary (verifier, injection) is over 30 days old | Ignored until re-recorded |

The gate proof is tied to a sha256 of `gate.yaml` (line endings normalised), so weakening the policy after recording drops L3 until the drills are run again. Gate and breaker drills are deterministic, so they don't expire otherwise.

The record is a claim the repo makes about itself, like a `Last run:` timestamp, so a determined author can hand-write one. Re-run the drills in CI to keep it honest; this repo's own CI does (see `scripts/ci-validate-gates.sh` and `scripts/ci-audit-gates.sh`).

## Levels

Expand Down
26 changes: 26 additions & 0 deletions tools/loop-audit/dist/auditor.d.ts
Original file line number Diff line number Diff line change
@@ -1,4 +1,5 @@
import { Finding, BaseAuditResult } from '@cobusgreyling/readiness-core';
import { type ProofSignals } from './proof.js';
export interface LoopSignals {
stateFile: {
present: boolean;
Expand Down Expand Up @@ -82,10 +83,35 @@ export interface LoopSignals {
registry: boolean;
inbox: boolean;
};
/** Guardrails loop-drill showed firing (loop-drill.json). Optional so older callers still type-check. */
proof?: ProofSignals;
}
export type { Finding };
export type { ProofSignals };
export interface AuditResult extends BaseAuditResult<'L0' | 'L1' | 'L2' | 'L3', LoopSignals> {
}
/**
* A signal file counts only when it has content. Empty and whitespace-only
* files, and JSON that is just {} or [], are placeholders: `touch` is not setup.
*/
export declare function hasContent(p: string): Promise<boolean>;
/**
* Frontmatter with a name and a description: what Claude Code, Codex and Grok
* need before they will load a skill or subagent. Without it the file is
* never invoked, whatever it is called.
*/
export declare function hasSkillFrontmatter(text: string): boolean;
/**
* A gate.yaml loop-gate can load declares `version: 1` and a denylist. This is
* a shape check, not a parse; loop-drill's record proves the policy works.
*/
export declare function looksLikeGatePolicy(text: string): boolean;
/**
* L3 means unattended actions behind gates, so the gate has to be shown to
* fire, not just to exist: loop-drill must have drilled the current gate.yaml
* in both directions, and no recorded guardrail may be failing.
*/
export declare function guardrailsProven(proof: ProofSignals | undefined): boolean;
/** Activity older than this does not count toward Loop Ready. */
export declare const ACTIVITY_MAX_AGE_MS: number;
export declare function computeScore(signals: LoopSignals): {
Expand Down
Loading
Loading