From 8955d4b64865b4641946051f2bf5220777c49b61 Mon Sep 17 00:00:00 2001 From: THRISHAL12345 Date: Mon, 28 Sep 2026 20:56:35 +0530 Subject: [PATCH 1/4] fix(starters): give changelog-drafter's Claude verifier the frontmatter it needs to load Claude Code only loads a subagent file with name and description frontmatter. This one had none, so the starter's Claude verifier was never invocable. Name and description come from its Codex twin. --- starters/changelog-drafter/.claude/agents/verifier.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/starters/changelog-drafter/.claude/agents/verifier.md b/starters/changelog-drafter/.claude/agents/verifier.md index 03cd5047..9402e34b 100644 --- a/starters/changelog-drafter/.claude/agents/verifier.md +++ b/starters/changelog-drafter/.claude/agents/verifier.md @@ -1,3 +1,9 @@ +--- +name: changelog-verifier +description: Independent checker for release note drafts produced by the changelog drafter. Accuracy and completeness focused. +model: inherit +--- + You are the independent verifier for the Changelog Drafter loop. Review the draft against the raw scan data. Flag any hallucinated items, missing breaking/security notes, or tone problems. From 7f039802efd08c0df763acf9a87267fb923c6dcd Mon Sep 17 00:00:00 2001 From: THRISHAL12345 Date: Mon, 28 Sep 2026 20:56:35 +0530 Subject: [PATCH 2/4] feat(loop-drill): --record writes loop-drill.json for loop-audit Results are grouped by guardrail (gate, breaker, verifier, ...). A run replaces only the groups it drilled, so refreshing the cheap offline drills doesn't discard a slow canary. Failures and skip reasons are recorded too. The gate group carries a sha256 of the policy file, with line endings normalised so Windows and Linux checkouts agree. loop-audit uses it to tell a proof of the current gate.yaml from a proof of an older one. --- tools/loop-drill/README.md | 41 +++++-- tools/loop-drill/dist/cli.js | 31 ++++- tools/loop-drill/dist/record.d.ts | 60 ++++++++++ tools/loop-drill/dist/record.js | 83 +++++++++++++ tools/loop-drill/src/cli.ts | 31 ++++- tools/loop-drill/src/record.ts | 123 +++++++++++++++++++ tools/loop-drill/test/record.test.mjs | 165 ++++++++++++++++++++++++++ 7 files changed, 522 insertions(+), 12 deletions(-) create mode 100644 tools/loop-drill/dist/record.d.ts create mode 100644 tools/loop-drill/dist/record.js create mode 100644 tools/loop-drill/src/record.ts create mode 100644 tools/loop-drill/test/record.test.mjs diff --git a/tools/loop-drill/README.md b/tools/loop-drill/README.md index 92adc643..83648193 100644 --- a/tools/loop-drill/README.md +++ b/tools/loop-drill/README.md @@ -10,15 +10,9 @@ Nothing checked whether they fire. -The sharpest case is the verifier. `loop-audit` gives its joint-largest signal (14 points) for a verifier and gates L3 on it, by checking whether a file whose *name* contains `verifier` exists: +The sharpest case is the verifier. `loop-audit` gives its joint-largest signal (14 points) for a verifier and gates L3 on it. From files alone it can only check that a verifier *loads*: a skill or agent file with a `name` and a `description`. A verifier that approves everything loads just as well. `docs/failure-modes.md` calls the resulting failure **Verifier Theater** and rates it S2; `docs/primitives.md` calls maker/checker *"the single most important structural pattern for reliable loops"*. -```ts -if (base.includes('verifier') || base === 'loop-verifier') // auditor.ts:181 -``` - -An empty file passes. On a scratch repo, `touch .claude/agents/verifier.md` moves the score from **34 (L0)** to **55 (L1)**. `docs/failure-modes.md` calls the resulting failure **Verifier Theater** and rates it S2; `docs/primitives.md` calls maker/checker *"the single most important structural pattern for reliable loops"*. - -loop-drill closes that gap by running the real verifier against real seeded defects. +loop-drill closes that gap by running the real verifier against real seeded defects, and [`--record`](#recording-proof-for-loop-audit) hands the result back to `loop-audit`. ## Usage @@ -29,6 +23,9 @@ npx @cobusgreyling/loop-drill . # The verifier canary npx @cobusgreyling/loop-drill . --only verifier \ --verifier-cmd "npm test" --setup "npm ci" + +# Record the results for loop-audit (commit loop-drill.json) +npx @cobusgreyling/loop-drill . --record ``` ### Options @@ -44,6 +41,7 @@ npx @cobusgreyling/loop-drill . --only verifier \ | `--gate-file ` | Policy file (default: `gate.yaml`) | | `--benign-path ` | Path the specificity drill treats as ordinary | | `--token-budget ` | Token budget for the breaker drill | +| `--record` | Write the results to `loop-drill.json` for `loop-audit`. Replaces only the guardrails drilled this run | | `--json` | Machine-readable output | Exit codes match `loop-gate` and `loop-context` so control scripts chain all three: **0** all passed, **1** some skipped, **2** a guardrail failed to fire. @@ -74,6 +72,33 @@ Every guardrail is drilled twice, because only one direction is easy: A `denylist: ["**"]` catches every seeded fault and still fails, because it blocks an ordinary docs change too. +## Recording proof for loop-audit + +`loop-audit` scores what a repo has. `--record` lets it score what works. It writes `loop-drill.json` at the repo root, grouped by guardrail: + +```json +{ + "schema": 1, + "tool": "@cobusgreyling/loop-drill", + "guardrails": { + "gate": { + "recordedAt": "2026-09-28T15:08:23.379Z", + "input": { "file": "gate.yaml", "sha256": "…" }, + "results": [ + { "id": "gate.denylist[**/.env]", "failureMode": "Over-Reach (Wrong Scope)", "direction": "sensitivity", "outcome": "passed" }, + { "id": "gate.benign", "failureMode": "Over-Reach (Wrong Scope)", "direction": "specificity", "outcome": "passed" } + ] + } + } +} +``` + +- **Failures are recorded too.** `loop-audit` withdraws the points for a guardrail shown not to fire, so a verifier that approves seeded defects stops counting as a verifier. +- **Only the guardrails you ran are replaced.** `--only verifier --record` refreshes the canary and keeps the gate and breaker results, so a slow canary doesn't have to be repeated to refresh a cheap drill. +- **The gate proof is tied to the policy it drilled.** The record keeps a sha256 of `gate.yaml` (line endings normalised). Edit the file and `loop-audit` treats the old proof as stale until you record again. Canary results expire after 30 days, because the code and the agent they tested keep changing. + +`loop-audit` needs the current `gate.yaml` **proven** for L3. That means at least one drill caught its fault, at least one benign case got through, and nothing failed. It's the same [both-directions rule](#both-directions-always) the drills follow. + ## How the verifier canary works Borrowed from mutation testing. Rather than asking a model to invent a defect — non-deterministic and token-hungry — it applies a small mechanical edit to real source and checks whether the verifier notices: diff --git a/tools/loop-drill/dist/cli.js b/tools/loop-drill/dist/cli.js index 857392c8..4c6b99df 100644 --- a/tools/loop-drill/dist/cli.js +++ b/tools/loop-drill/dist/cli.js @@ -5,12 +5,14 @@ * Exit codes follow loop-gate / loop-context so control scripts can chain all * three: 0 proceed, 1 warnings, 2 escalate. */ +import { readFile } from 'node:fs/promises'; import path from 'node:path'; import { loadGateConfig } from '@cobusgreyling/loop-gate'; import { DEFAULT_BREAKER } from '@cobusgreyling/loop-context'; import { buildReport, exitCodeFor, runBreakerDrills, runGateDrills, skip, } from './drill.js'; import { runCanary } from './canary.js'; import { formatReport } from './report.js'; +import { buildGroups, fingerprint, RECORD_FILE, writeRecord } from './record.js'; const HELP = `loop-drill — fire drills for loop guardrails Injects a known fault and asserts the guardrail actually fires. Turns L3 @@ -31,6 +33,9 @@ Options: --benign-path Path the gate specificity drill treats as ordinary (default: docs/README.md) --token-budget Token budget for the breaker drill + --record Write the results to loop-drill.json, which loop-audit + reads as proof the guardrails fire. Replaces only the + guardrails drilled this run; commit the file --json Machine-readable output --help, -h Show this message @@ -40,9 +45,10 @@ Examples: loop-drill . loop-drill . --only verifier --verifier-cmd "npm test" --setup "npm ci" loop-drill . --only gate,breaker --json + loop-drill . --record `; function parseArgs(argv) { - const flags = { root: '.', json: false, help: false, mutants: 3, timeoutMs: 120_000 }; + const flags = { root: '.', json: false, help: false, mutants: 3, timeoutMs: 120_000, record: false }; const positional = []; for (let i = 0; i < argv.length; i++) { const arg = argv[i]; @@ -60,6 +66,9 @@ function parseArgs(argv) { case '--json': flags.json = true; break; + case '--record': + flags.record = true; + break; case '--only': flags.only = next(); break; @@ -121,8 +130,8 @@ async function main() { const selected = new Set((flags.only ?? 'gate,breaker').split(',').map((s) => s.trim()).filter(Boolean)); const results = []; let mutationScore = null; + const gateFile = path.resolve(root, flags.gateFile ?? 'gate.yaml'); if (selected.has('gate')) { - const gateFile = path.resolve(root, flags.gateFile ?? 'gate.yaml'); try { const config = await loadGateConfig(gateFile); results.push(...runGateDrills({ config, benignPath: flags.benignPath })); @@ -166,6 +175,24 @@ async function main() { else { console.log(formatReport(report, mutationScore)); } + if (flags.record) { + // Failures are recorded too: loop-audit withdraws credit for a guardrail + // that was shown not to fire. + const gateText = await readFile(gateFile, 'utf8').catch(() => null); + const groups = buildGroups(results, selected, { + recordedAt: report.timestamp, + gateFile: path.relative(root, gateFile).split(path.sep).join('/'), + gateSha256: gateText === null ? null : fingerprint(gateText), + verifierCommand: flags.verifierCmd, + mutationScore, + }); + await writeRecord(root, groups); + const note = `Recorded ${Object.keys(groups).join(', ')} to ${RECORD_FILE} — commit it so loop-audit can see the proof.`; + if (flags.json) + console.error(note); + else + console.log(note); + } process.exit(code); } main().catch((err) => { diff --git a/tools/loop-drill/dist/record.d.ts b/tools/loop-drill/dist/record.d.ts new file mode 100644 index 00000000..e47a9f33 --- /dev/null +++ b/tools/loop-drill/dist/record.d.ts @@ -0,0 +1,60 @@ +/** + * `loop-drill --record`: write what the drills found to loop-drill.json, so + * loop-audit can score guardrails that were shown to fire rather than files + * that merely exist. + * + * The record is grouped by guardrail (gate, breaker, verifier, ...). Recording + * a subset -- say `--only verifier` -- replaces that group and keeps the rest, + * so a slow canary run does not have to be repeated to refresh the gate. + * + * Each group carries what it was drilled against. For the gate that is a + * fingerprint of the policy file: edit gate.yaml and the old proof stops + * counting until the drills are run again. + */ +import type { DrillResult } from './drill.js'; +export declare const RECORD_FILE = "loop-drill.json"; +export declare const RECORD_SCHEMA = 1; +export interface RecordedResult { + id: string; + failureMode: string; + direction: DrillResult['direction']; + outcome: DrillResult['outcome']; + /** Only kept for drills that did not pass, to say why. */ + detail?: string; +} +export interface RecordedGuardrail { + recordedAt: string; + input?: { + file?: string; + sha256?: string | null; + command?: string; + }; + mutationScore?: number | null; + results: RecordedResult[]; +} +export interface DrillRecord { + schema: number; + tool: string; + guardrails: Record; +} +/** + * sha256 of a policy file with line endings normalised, so a Windows checkout + * with core.autocrlf fingerprints the same as the Linux runner that recorded + * it. loop-audit computes the same value; both test suites pin one vector. + */ +export declare function fingerprint(text: string): string; +/** 'gate.denylist[**\/.env]' -> 'gate'; a whole-drill skip like 'verifier' -> 'verifier'. */ +export declare function guardrailOf(id: string): string; +export interface GroupInputs { + recordedAt: string; + /** Repo-relative, forward slashes. */ + gateFile?: string; + gateSha256?: string | null; + verifierCommand?: string; + mutationScore?: number | null; +} +/** Group this run's results by the guardrails that were selected. */ +export declare function buildGroups(results: DrillResult[], selected: Iterable, inputs: GroupInputs): Record; +/** Replace the groups this run drilled; keep the others from the existing record. */ +export declare function mergeRecord(existing: unknown, groups: Record): DrillRecord; +export declare function writeRecord(root: string, groups: Record): Promise; diff --git a/tools/loop-drill/dist/record.js b/tools/loop-drill/dist/record.js new file mode 100644 index 00000000..73120330 --- /dev/null +++ b/tools/loop-drill/dist/record.js @@ -0,0 +1,83 @@ +/** + * `loop-drill --record`: write what the drills found to loop-drill.json, so + * loop-audit can score guardrails that were shown to fire rather than files + * that merely exist. + * + * The record is grouped by guardrail (gate, breaker, verifier, ...). Recording + * a subset -- say `--only verifier` -- replaces that group and keeps the rest, + * so a slow canary run does not have to be repeated to refresh the gate. + * + * Each group carries what it was drilled against. For the gate that is a + * fingerprint of the policy file: edit gate.yaml and the old proof stops + * counting until the drills are run again. + */ +import { createHash } from 'node:crypto'; +import { readFile, writeFile } from 'node:fs/promises'; +import path from 'node:path'; +export const RECORD_FILE = 'loop-drill.json'; +export const RECORD_SCHEMA = 1; +/** + * sha256 of a policy file with line endings normalised, so a Windows checkout + * with core.autocrlf fingerprints the same as the Linux runner that recorded + * it. loop-audit computes the same value; both test suites pin one vector. + */ +export function fingerprint(text) { + return createHash('sha256').update(text.replace(/\r\n/g, '\n'), 'utf8').digest('hex'); +} +/** 'gate.denylist[**\/.env]' -> 'gate'; a whole-drill skip like 'verifier' -> 'verifier'. */ +export function guardrailOf(id) { + return id.split(/[.[]/, 1)[0]; +} +/** Group this run's results by the guardrails that were selected. */ +export function buildGroups(results, selected, inputs) { + const groups = {}; + for (const name of selected) { + const mine = results.filter((r) => guardrailOf(r.id) === name); + if (mine.length === 0) + continue; + const group = { + recordedAt: inputs.recordedAt, + results: mine.map((r) => ({ + id: r.id, + failureMode: r.failureMode, + direction: r.direction, + outcome: r.outcome, + ...(r.outcome !== 'passed' && r.detail ? { detail: r.detail } : {}), + })), + }; + if (name === 'gate') + group.input = { file: inputs.gateFile, sha256: inputs.gateSha256 ?? null }; + if (name === 'verifier' && inputs.verifierCommand) { + group.input = { command: inputs.verifierCommand }; + group.mutationScore = inputs.mutationScore ?? null; + } + groups[name] = group; + } + return groups; +} +function isRecord(value) { + const r = value; + return !!r && r.schema === RECORD_SCHEMA && !!r.guardrails && typeof r.guardrails === 'object'; +} +/** Replace the groups this run drilled; keep the others from the existing record. */ +export function mergeRecord(existing, groups) { + const kept = isRecord(existing) ? existing.guardrails : {}; + const merged = { ...kept, ...groups }; + const sorted = {}; + for (const key of Object.keys(merged).sort()) + sorted[key] = merged[key]; + return { schema: RECORD_SCHEMA, tool: '@cobusgreyling/loop-drill', guardrails: sorted }; +} +export async function writeRecord(root, groups) { + const file = path.join(root, RECORD_FILE); + let existing; + try { + existing = JSON.parse(await readFile(file, 'utf8')); + } + catch { + existing = undefined; // missing or unreadable: start a fresh record + } + const record = mergeRecord(existing, groups); + await writeFile(file, `${JSON.stringify(record, null, 2)}\n`, 'utf8'); + return file; +} diff --git a/tools/loop-drill/src/cli.ts b/tools/loop-drill/src/cli.ts index 24dbc6ff..0bc41d98 100644 --- a/tools/loop-drill/src/cli.ts +++ b/tools/loop-drill/src/cli.ts @@ -6,6 +6,7 @@ * three: 0 proceed, 1 warnings, 2 escalate. */ +import { readFile } from 'node:fs/promises'; import path from 'node:path'; import { loadGateConfig } from '@cobusgreyling/loop-gate'; import { DEFAULT_BREAKER, type CircuitBreakerConfig } from '@cobusgreyling/loop-context'; @@ -19,6 +20,7 @@ import { } from './drill.js'; import { runCanary } from './canary.js'; import { formatReport } from './report.js'; +import { buildGroups, fingerprint, RECORD_FILE, writeRecord } from './record.js'; interface Flags { root: string; @@ -33,6 +35,7 @@ interface Flags { scope?: string; setup?: string; tokenBudget?: number; + record: boolean; } const HELP = `loop-drill — fire drills for loop guardrails @@ -55,6 +58,9 @@ Options: --benign-path Path the gate specificity drill treats as ordinary (default: docs/README.md) --token-budget Token budget for the breaker drill + --record Write the results to loop-drill.json, which loop-audit + reads as proof the guardrails fire. Replaces only the + guardrails drilled this run; commit the file --json Machine-readable output --help, -h Show this message @@ -64,10 +70,11 @@ Examples: loop-drill . loop-drill . --only verifier --verifier-cmd "npm test" --setup "npm ci" loop-drill . --only gate,breaker --json + loop-drill . --record `; function parseArgs(argv: string[]): Flags { - const flags: Flags = { root: '.', json: false, help: false, mutants: 3, timeoutMs: 120_000 }; + const flags: Flags = { root: '.', json: false, help: false, mutants: 3, timeoutMs: 120_000, record: false }; const positional: string[] = []; for (let i = 0; i < argv.length; i++) { @@ -85,6 +92,9 @@ function parseArgs(argv: string[]): Flags { case '--json': flags.json = true; break; + case '--record': + flags.record = true; + break; case '--only': flags.only = next(); break; @@ -148,8 +158,8 @@ async function main(): Promise { const results: DrillResult[] = []; let mutationScore: number | null = null; + const gateFile = path.resolve(root, flags.gateFile ?? 'gate.yaml'); if (selected.has('gate')) { - const gateFile = path.resolve(root, flags.gateFile ?? 'gate.yaml'); try { const config = await loadGateConfig(gateFile); results.push(...runGateDrills({ config, benignPath: flags.benignPath })); @@ -213,6 +223,23 @@ async function main(): Promise { console.log(formatReport(report, mutationScore)); } + if (flags.record) { + // Failures are recorded too: loop-audit withdraws credit for a guardrail + // that was shown not to fire. + const gateText = await readFile(gateFile, 'utf8').catch(() => null); + const groups = buildGroups(results, selected, { + recordedAt: report.timestamp, + gateFile: path.relative(root, gateFile).split(path.sep).join('/'), + gateSha256: gateText === null ? null : fingerprint(gateText), + verifierCommand: flags.verifierCmd, + mutationScore, + }); + await writeRecord(root, groups); + const note = `Recorded ${Object.keys(groups).join(', ')} to ${RECORD_FILE} — commit it so loop-audit can see the proof.`; + if (flags.json) console.error(note); + else console.log(note); + } + process.exit(code); } diff --git a/tools/loop-drill/src/record.ts b/tools/loop-drill/src/record.ts new file mode 100644 index 00000000..6c28e111 --- /dev/null +++ b/tools/loop-drill/src/record.ts @@ -0,0 +1,123 @@ +/** + * `loop-drill --record`: write what the drills found to loop-drill.json, so + * loop-audit can score guardrails that were shown to fire rather than files + * that merely exist. + * + * The record is grouped by guardrail (gate, breaker, verifier, ...). Recording + * a subset -- say `--only verifier` -- replaces that group and keeps the rest, + * so a slow canary run does not have to be repeated to refresh the gate. + * + * Each group carries what it was drilled against. For the gate that is a + * fingerprint of the policy file: edit gate.yaml and the old proof stops + * counting until the drills are run again. + */ + +import { createHash } from 'node:crypto'; +import { readFile, writeFile } from 'node:fs/promises'; +import path from 'node:path'; +import type { DrillResult } from './drill.js'; + +export const RECORD_FILE = 'loop-drill.json'; +export const RECORD_SCHEMA = 1; + +export interface RecordedResult { + id: string; + failureMode: string; + direction: DrillResult['direction']; + outcome: DrillResult['outcome']; + /** Only kept for drills that did not pass, to say why. */ + detail?: string; +} + +export interface RecordedGuardrail { + recordedAt: string; + input?: { file?: string; sha256?: string | null; command?: string }; + mutationScore?: number | null; + results: RecordedResult[]; +} + +export interface DrillRecord { + schema: number; + tool: string; + guardrails: Record; +} + +/** + * sha256 of a policy file with line endings normalised, so a Windows checkout + * with core.autocrlf fingerprints the same as the Linux runner that recorded + * it. loop-audit computes the same value; both test suites pin one vector. + */ +export function fingerprint(text: string): string { + return createHash('sha256').update(text.replace(/\r\n/g, '\n'), 'utf8').digest('hex'); +} + +/** 'gate.denylist[**\/.env]' -> 'gate'; a whole-drill skip like 'verifier' -> 'verifier'. */ +export function guardrailOf(id: string): string { + return id.split(/[.[]/, 1)[0]; +} + +export interface GroupInputs { + recordedAt: string; + /** Repo-relative, forward slashes. */ + gateFile?: string; + gateSha256?: string | null; + verifierCommand?: string; + mutationScore?: number | null; +} + +/** Group this run's results by the guardrails that were selected. */ +export function buildGroups( + results: DrillResult[], + selected: Iterable, + inputs: GroupInputs, +): Record { + const groups: Record = {}; + for (const name of selected) { + const mine = results.filter((r) => guardrailOf(r.id) === name); + if (mine.length === 0) continue; + const group: RecordedGuardrail = { + recordedAt: inputs.recordedAt, + results: mine.map((r) => ({ + id: r.id, + failureMode: r.failureMode, + direction: r.direction, + outcome: r.outcome, + ...(r.outcome !== 'passed' && r.detail ? { detail: r.detail } : {}), + })), + }; + if (name === 'gate') group.input = { file: inputs.gateFile, sha256: inputs.gateSha256 ?? null }; + if (name === 'verifier' && inputs.verifierCommand) { + group.input = { command: inputs.verifierCommand }; + group.mutationScore = inputs.mutationScore ?? null; + } + groups[name] = group; + } + return groups; +} + +function isRecord(value: unknown): value is DrillRecord { + const r = value as DrillRecord | null; + return !!r && r.schema === RECORD_SCHEMA && !!r.guardrails && typeof r.guardrails === 'object'; +} + +/** Replace the groups this run drilled; keep the others from the existing record. */ +export function mergeRecord(existing: unknown, groups: Record): DrillRecord { + const kept = isRecord(existing) ? existing.guardrails : {}; + const merged = { ...kept, ...groups }; + const sorted: Record = {}; + for (const key of Object.keys(merged).sort()) sorted[key] = merged[key]; + return { schema: RECORD_SCHEMA, tool: '@cobusgreyling/loop-drill', guardrails: sorted }; +} + +export async function writeRecord(root: string, groups: Record): Promise { + const file = path.join(root, RECORD_FILE); + let existing: unknown; + try { + existing = JSON.parse(await readFile(file, 'utf8')); + } catch { + existing = undefined; // missing or unreadable: start a fresh record + } + const record = mergeRecord(existing, groups); + await writeFile(file, `${JSON.stringify(record, null, 2)}\n`, 'utf8'); + return file; +} diff --git a/tools/loop-drill/test/record.test.mjs b/tools/loop-drill/test/record.test.mjs new file mode 100644 index 00000000..6ac63f52 --- /dev/null +++ b/tools/loop-drill/test/record.test.mjs @@ -0,0 +1,165 @@ +import { test, describe } from 'node:test'; +import assert from 'node:assert/strict'; +import path from 'node:path'; +import os from 'node:os'; +import { mkdtemp, readFile, rm, writeFile } from 'node:fs/promises'; +import { execFile } from 'node:child_process'; +import { fileURLToPath } from 'node:url'; +import { promisify } from 'node:util'; +import { buildGroups, fingerprint, guardrailOf, mergeRecord, writeRecord, RECORD_FILE } from '../dist/record.js'; + +const exec = promisify(execFile); +const CLI = fileURLToPath(new URL('../dist/cli.js', import.meta.url)); + +const GATE = 'version: 1\ndenylist:\n - "**/.env"\n - "**/secrets/**"\nmaxFiles: 10\n'; + +const result = (id, outcome = 'passed', direction = 'sensitivity', detail) => ({ + id, + name: id, + failureMode: 'Over-Reach (Wrong Scope)', + direction, + outcome, + expected: 'x', + actual: 'y', + ...(detail ? { detail } : {}), +}); + +async function withDir(fn) { + const dir = await mkdtemp(path.join(os.tmpdir(), 'loop-drill-record-')); + try { + return await fn(dir); + } finally { + await rm(dir, { recursive: true, force: true }); + } +} + +describe('record', () => { + test('fingerprint is sha256 of the LF-normalised file (vector shared with loop-audit)', () => { + const vector = 'a79338c9d44fec8fe07da24a3766aaf4e32281bd5c73628e0798d59944d1040e'; + assert.equal(fingerprint('version: 1\ndenylist:\n - "**/.env"\n'), vector); + assert.equal(fingerprint('version: 1\r\ndenylist:\r\n - "**/.env"\r\n'), vector); + }); + + test('guardrailOf maps drill ids to their guardrail', () => { + assert.equal(guardrailOf('gate.denylist[**/.env]'), 'gate'); + assert.equal(guardrailOf('gate'), 'gate'); + assert.equal(guardrailOf('breaker.token-budget'), 'breaker'); + assert.equal(guardrailOf('verifier.mutant[strict-equality-flip]'), 'verifier'); + assert.equal(guardrailOf('injection.hidden-comment'), 'injection'); + }); + + test('buildGroups groups by selected guardrail and keeps detail only for non-passing drills', () => { + const groups = buildGroups( + [ + { ...result('gate.denylist[**/.env]'), detail: 'noise' }, + result('gate.benign', 'passed', 'specificity'), + result('breaker.token-budget', 'skipped', 'sensitivity', 'No tokenBudget configured'), + ], + ['gate', 'breaker', 'verifier'], + { recordedAt: '2026-09-28T00:00:00.000Z', gateFile: 'gate.yaml', gateSha256: 'abc' }, + ); + assert.deepEqual(Object.keys(groups), ['gate', 'breaker'], 'no empty group for a guardrail with no results'); + assert.deepEqual(groups.gate.input, { file: 'gate.yaml', sha256: 'abc' }); + assert.equal(groups.gate.results[0].detail, undefined); + assert.equal(groups.breaker.results[0].detail, 'No tokenBudget configured'); + assert.equal(groups.breaker.input, undefined); + assert.deepEqual(Object.keys(groups.gate.results[0]).sort(), ['direction', 'failureMode', 'id', 'outcome']); + }); + + test('the verifier group records its command and mutation score', () => { + const groups = buildGroups( + [result('verifier.control', 'passed', 'specificity'), result('verifier.mutant[x]')], + ['verifier'], + { recordedAt: 't', verifierCommand: 'npm test', mutationScore: 1 }, + ); + assert.deepEqual(groups.verifier.input, { command: 'npm test' }); + assert.equal(groups.verifier.mutationScore, 1); + }); + + test('mergeRecord replaces the groups drilled this run and keeps the rest', () => { + const existing = { + schema: 1, + guardrails: { verifier: { recordedAt: 'old', results: [] }, gate: { recordedAt: 'old', results: [] } }, + }; + const merged = mergeRecord(existing, { gate: { recordedAt: 'new', results: [] }, breaker: { recordedAt: 'new', results: [] } }); + assert.deepEqual(Object.keys(merged.guardrails), ['breaker', 'gate', 'verifier'], 'sorted for stable diffs'); + assert.equal(merged.guardrails.gate.recordedAt, 'new'); + assert.equal(merged.guardrails.verifier.recordedAt, 'old'); + assert.equal(merged.schema, 1); + }); + + test('mergeRecord starts fresh over a missing or foreign record', () => { + for (const existing of [undefined, null, { schema: 99, guardrails: { gate: {} } }, 'text']) { + const merged = mergeRecord(existing, { gate: { recordedAt: 'new', results: [] } }); + assert.deepEqual(Object.keys(merged.guardrails), ['gate']); + } + }); + + test('writeRecord overwrites an unreadable file rather than failing', async () => { + await withDir(async (dir) => { + await writeFile(path.join(dir, RECORD_FILE), '{ corrupt'); + await writeRecord(dir, { gate: { recordedAt: 'new', results: [] } }); + const saved = JSON.parse(await readFile(path.join(dir, RECORD_FILE), 'utf8')); + assert.deepEqual(Object.keys(saved.guardrails), ['gate']); + }); + }); +}); + +describe('loop-drill --record', () => { + test('writes gate and breaker proof bound to the gate.yaml it drilled', async () => { + await withDir(async (dir) => { + await writeFile(path.join(dir, 'gate.yaml'), GATE.replace(/\n/g, '\r\n')); + const { stdout } = await exec(process.execPath, [CLI, dir, '--record']).catch((err) => err); + assert.match(stdout, /Recorded gate, breaker to loop-drill\.json/); + + const saved = JSON.parse(await readFile(path.join(dir, RECORD_FILE), 'utf8')); + assert.equal(saved.schema, 1); + assert.deepEqual(saved.guardrails.gate.input, { file: 'gate.yaml', sha256: fingerprint(GATE) }, 'CRLF on disk, LF fingerprint'); + const gate = saved.guardrails.gate.results; + assert.ok(gate.some((r) => r.direction === 'sensitivity' && r.outcome === 'passed')); + assert.ok(gate.some((r) => r.id === 'gate.benign' && r.outcome === 'passed')); + assert.ok(saved.guardrails.breaker.results.length > 0); + }); + }); + + test('records a gate loop-gate refuses as a skip with the reason', async () => { + await withDir(async (dir) => { + await writeFile(path.join(dir, 'gate.yaml'), 'version: 1\ndenylist: nope\n'); + await exec(process.execPath, [CLI, dir, '--only', 'gate', '--record']).catch((err) => err); + const saved = JSON.parse(await readFile(path.join(dir, RECORD_FILE), 'utf8')); + const [only] = saved.guardrails.gate.results; + assert.equal(only.outcome, 'skipped'); + assert.match(only.detail, /denylist/); + assert.equal(saved.guardrails.gate.input.sha256, fingerprint('version: 1\ndenylist: nope\n')); + }); + }); + + test('--only replaces one guardrail and keeps the others', async () => { + await withDir(async (dir) => { + await writeFile(path.join(dir, 'gate.yaml'), GATE); + await exec(process.execPath, [CLI, dir, '--record']).catch((err) => err); + const first = JSON.parse(await readFile(path.join(dir, RECORD_FILE), 'utf8')); + await exec(process.execPath, [CLI, dir, '--only', 'breaker', '--record']).catch((err) => err); + const second = JSON.parse(await readFile(path.join(dir, RECORD_FILE), 'utf8')); + assert.deepEqual(second.guardrails.gate, first.guardrails.gate); + assert.notEqual(second.guardrails.breaker.recordedAt, undefined); + }); + }); + + test('--json --record keeps stdout parseable', async () => { + await withDir(async (dir) => { + await writeFile(path.join(dir, 'gate.yaml'), GATE); + const { stdout, stderr } = await exec(process.execPath, [CLI, dir, '--json', '--record']).catch((err) => err); + assert.doesNotThrow(() => JSON.parse(stdout)); + assert.match(stderr, /Recorded gate, breaker/); + }); + }); + + test('without --record nothing is written', async () => { + await withDir(async (dir) => { + await writeFile(path.join(dir, 'gate.yaml'), GATE); + await exec(process.execPath, [CLI, dir]).catch((err) => err); + await assert.rejects(readFile(path.join(dir, RECORD_FILE), 'utf8')); + }); + }); +}); From ec6aa57d9d0cff4993e1c3d12b25df76748a7ba8 Mon Sep 17 00:00:00 2001 From: THRISHAL12345 Date: Mon, 28 Sep 2026 20:56:35 +0530 Subject: [PATCH 3/4] feat(loop-audit): score what is proven, not what is present A repo of 19 empty files (142 bytes) scored 100/L3, the same as the reference repo: nearly every signal was fileExists(), and a skill counted if its directory existed. Placeholders no longer score. Empty or whitespace-only files, and {} / [] JSON, earn nothing and are listed under "Not counted". A skill or Claude verifier agent needs name + description frontmatter to count. gate.yaml needs version: 1 and a denylist. .github/ needs a non-empty file. L3 now requires proven guardrails, read from loop-drill.json: the current gate.yaml must pass its drills in both directions, and no recorded guardrail may be failing. A guardrail loop-drill shows failing loses its points (gate, verifier, breaker). Proof goes stale when gate.yaml changes, or after 30 days for canaries. The fake repo now scores 72/L1. No starter's score changes. --- tools/loop-audit/CHANGELOG.md | 13 + tools/loop-audit/README.md | 32 ++- tools/loop-audit/dist/auditor.d.ts | 26 ++ tools/loop-audit/dist/auditor.js | 357 ++++++++++++++++++------ tools/loop-audit/dist/cli.js | 4 + tools/loop-audit/dist/proof.d.ts | 70 +++++ tools/loop-audit/dist/proof.js | 135 +++++++++ tools/loop-audit/src/auditor.ts | 357 ++++++++++++++++++------ tools/loop-audit/src/cli.ts | 4 + tools/loop-audit/src/proof.ts | 183 +++++++++++++ tools/loop-audit/test/auditor.test.mjs | 36 ++- tools/loop-audit/test/proof.test.mjs | 364 +++++++++++++++++++++++++ 12 files changed, 1421 insertions(+), 160 deletions(-) create mode 100644 tools/loop-audit/dist/proof.d.ts create mode 100644 tools/loop-audit/dist/proof.js create mode 100644 tools/loop-audit/src/proof.ts create mode 100644 tools/loop-audit/test/proof.test.mjs diff --git a/tools/loop-audit/CHANGELOG.md b/tools/loop-audit/CHANGELOG.md index 4521f9db..22592d8d 100644 --- a/tools/loop-audit/CHANGELOG.md +++ b/tools/loop-audit/CHANGELOG.md @@ -2,6 +2,19 @@ All notable changes to `@cobusgreyling/loop-audit` are documented here. +## [Unreleased] + +### Changed +- **Placeholders no longer score.** Empty or whitespace-only files, and `{}` / `[]` JSON, earn nothing and are listed under *Not counted*. A repo of 19 empty files (142 bytes) used to score 100/L3; it now scores 72/L1 +- A skill counts only with a `SKILL.md` carrying `name` + `description` frontmatter (bare directories used to count); the same for Claude Code verifier agents +- `gate.yaml` must declare `version: 1` and a `denylist:`; one `loop-gate` would refuse is a failure, not a signal +- `.github/` and workflows count only when they contain a non-empty file +- **L3 requires proven guardrails**: a `loop-drill.json` (from `loop-drill --record`) showing the current `gate.yaml` passing its drills in both directions, with no recorded guardrail failing. Repos that were L3 on files alone drop to L2 until they record + +### Added +- `signals.proof`: guardrails that are proven, failed, untested or stale, read from `loop-drill.json` +- A guardrail that `loop-drill` shows failing loses its points: gate → `gateYaml`, verifier → `verifier` (Verifier Theater), breaker → stall detection + ## [1.9.0] - 2026-08-28 ### Changed diff --git a/tools/loop-audit/README.md b/tools/loop-audit/README.md index eb711056..34c56215 100644 --- a/tools/loop-audit/README.md +++ b/tools/loop-audit/README.md @@ -76,6 +76,7 @@ npm publish --access public | Human-escalation path | LOOP.md / safety docs define when to stop and hand off to a human | | **loopActivity (v1.4)** | **Dynamic proof**: "Last run" timestamps in state, loop-related git commits, scheduled workflows, run logs | | **Harness Runtime (v1.7)** | `.foundry/stack.yaml`, lock, sessions/traces, outerloop emit, host integrate — LE → [harness-foundry](https://github.com/cobusgreyling/harness-foundry) funnel | +| **Guardrail proof** | `loop-drill.json` from [`loop-drill --record`](../loop-drill): the gate, breaker and verifier shown firing — see below | When score ≥ 80 and no `.foundry/stack.yaml`, audit recommends: @@ -83,7 +84,36 @@ When score ≥ 80 and no `.foundry/stack.yaml`, audit recommends: npx @cobusgreyling/loop-init . --with-foundry ``` -L3 requires verifier + state + cost observability (budget + run log + LOOP.md budget) **and** proven loop activity (not just files on disk). +L3 requires verifier + state + cost observability (budget + run log + LOOP.md budget), proven loop activity (not just files on disk), **and proven guardrails**: a `loop-drill.json` showing the current `gate.yaml` passing its drills, with no recorded guardrail failing. + +## Present is not proven + +Every signal counts content, not filenames: + +- **Empty files don't score.** A file that is empty or whitespace, or JSON that is just `{}` / `[]`, is a placeholder. The audit lists each one under *Not counted* instead of crediting it. +- **Skills must load.** A skill is a directory with a `SKILL.md` that has `name` and `description` frontmatter, which every host needs before it will invoke it. The same goes for a Claude Code verifier agent. A bare directory, or a file with no frontmatter, doesn't count, whatever it's called. +- **`gate.yaml` must be a policy.** It needs `version: 1` and a `denylist:`. Anything else would be refused by `loop-gate`, so it earns nothing and is reported as a failure. + +Files can only show a guardrail is configured. [`loop-drill`](../loop-drill) shows whether it fires, by running it against a seeded fault and a benign case. Record the results and commit them: + +```bash +npx @cobusgreyling/loop-drill . --record # gate + breaker: offline, no tokens +npx @cobusgreyling/loop-drill . --only verifier --verifier-cmd "npm test" --setup "npm ci" --record +git add loop-drill.json +``` + +The audit reads the record per guardrail: + +| Result | Meaning | Effect | +|---|---|---| +| **proven** | A drill caught the fault **and** a drill let the benign case through, and nothing failed | Counts. The gate being proven is required for L3 | +| **failed** | A drill failed: the guardrail didn't fire, or blocked the benign case | The guardrail's points are withdrawn (gate → `gateYaml`, verifier → `verifier`, breaker → stall detection), and L3 is blocked | +| **untested** | Every drill was skipped, or only one direction passed | Not proven. For the gate this means `loop-gate` couldn't drill it, so `gateYaml` is withdrawn | +| **stale** | The gate was drilled against a different `gate.yaml`, or a canary (verifier, injection) is over 30 days old | Ignored until re-recorded | + +The gate proof is tied to a sha256 of `gate.yaml` (line endings normalised), so weakening the policy after recording drops L3 until the drills are run again. Gate and breaker drills are deterministic, so they don't expire otherwise. + +The record is a claim the repo makes about itself, like a `Last run:` timestamp, so a determined author can hand-write one. Re-run the drills in CI to keep it honest; this repo's own CI does (see `scripts/ci-validate-gates.sh` and `scripts/ci-audit-gates.sh`). ## Levels diff --git a/tools/loop-audit/dist/auditor.d.ts b/tools/loop-audit/dist/auditor.d.ts index 225de07f..20cf1a1d 100644 --- a/tools/loop-audit/dist/auditor.d.ts +++ b/tools/loop-audit/dist/auditor.d.ts @@ -1,4 +1,5 @@ import { Finding, BaseAuditResult } from '@cobusgreyling/readiness-core'; +import { type ProofSignals } from './proof.js'; export interface LoopSignals { stateFile: { present: boolean; @@ -82,10 +83,35 @@ export interface LoopSignals { registry: boolean; inbox: boolean; }; + /** Guardrails loop-drill showed firing (loop-drill.json). Optional so older callers still type-check. */ + proof?: ProofSignals; } export type { Finding }; +export type { ProofSignals }; export interface AuditResult extends BaseAuditResult<'L0' | 'L1' | 'L2' | 'L3', LoopSignals> { } +/** + * A signal file counts only when it has content. Empty and whitespace-only + * files, and JSON that is just {} or [], are placeholders: `touch` is not setup. + */ +export declare function hasContent(p: string): Promise; +/** + * Frontmatter with a name and a description: what Claude Code, Codex and Grok + * need before they will load a skill or subagent. Without it the file is + * never invoked, whatever it is called. + */ +export declare function hasSkillFrontmatter(text: string): boolean; +/** + * A gate.yaml loop-gate can load declares `version: 1` and a denylist. This is + * a shape check, not a parse; loop-drill's record proves the policy works. + */ +export declare function looksLikeGatePolicy(text: string): boolean; +/** + * L3 means unattended actions behind gates, so the gate has to be shown to + * fire, not just to exist: loop-drill must have drilled the current gate.yaml + * in both directions, and no recorded guardrail may be failing. + */ +export declare function guardrailsProven(proof: ProofSignals | undefined): boolean; /** Activity older than this does not count toward Loop Ready. */ export declare const ACTIVITY_MAX_AGE_MS: number; export declare function computeScore(signals: LoopSignals): { diff --git a/tools/loop-audit/dist/auditor.js b/tools/loop-audit/dist/auditor.js index 6e6b3053..c52c70b1 100644 --- a/tools/loop-audit/dist/auditor.js +++ b/tools/loop-audit/dist/auditor.js @@ -1,7 +1,8 @@ -import { readdir, readFile } from 'node:fs/promises'; +import { readdir, readFile, stat } from 'node:fs/promises'; import path from 'node:path'; import { execSync } from 'node:child_process'; -import { fileExists, scanSkillDirectories } from '@cobusgreyling/readiness-core'; +import { fileExists } from '@cobusgreyling/readiness-core'; +import { PROOF_FILE, readProof } from './proof.js'; const STATE_FILES = [ 'STATE.md', 'pr-babysitter-state.md', @@ -105,24 +106,127 @@ const ESCALATION_HINTS = [ /exit code 2/i, /\bexit 2\b/i, ]; +const SKILL_DIRS = ['.grok/skills', '.claude/skills', '.codex/skills', 'skills']; +/** + * A signal file counts only when it has content. Empty and whitespace-only + * files, and JSON that is just {} or [], are placeholders: `touch` is not setup. + */ +export async function hasContent(p) { + let text; + try { + const info = await stat(p); + if (!info.isFile()) + return false; + if (info.size > 64 * 1024) + return true; + text = await readFile(p, 'utf8'); + } + catch { + return false; + } + const trimmed = text.replace(/^/, '').trim(); + if (!trimmed) + return false; + if (p.toLowerCase().endsWith('.json')) { + try { + const value = JSON.parse(trimmed); + if (value !== null && typeof value === 'object' && Object.keys(value).length === 0) + return false; + } + catch { + // Not valid JSON, but still content; other checks judge the format. + } + } + return true; +} +/** + * Frontmatter with a name and a description: what Claude Code, Codex and Grok + * need before they will load a skill or subagent. Without it the file is + * never invoked, whatever it is called. + */ +export function hasSkillFrontmatter(text) { + const m = text.replace(/^/, '').match(/^---\r?\n([\s\S]*?)\r?\n---[ \t]*(?:\r?\n|$)/); + if (!m) + return false; + return /^name:[ \t]*\S/m.test(m[1]) && /^description:[ \t]*\S/m.test(m[1]); +} +async function isLoadableSkill(p) { + try { + return hasSkillFrontmatter(await readFile(p, 'utf8')); + } + catch { + return false; + } +} +/** + * A gate.yaml loop-gate can load declares `version: 1` and a denylist. This is + * a shape check, not a parse; loop-drill's record proves the policy works. + */ +export function looksLikeGatePolicy(text) { + return /^version:[ \t]*1[ \t]*(?:#.*)?$/m.test(text) && /^denylist:/m.test(text); +} +async function anyFileWithContent(dir, depth, accept) { + let entries; + try { + entries = await readdir(dir, { withFileTypes: true }); + } + catch { + return false; + } + for (const e of entries) { + const p = path.join(dir, e.name); + if (e.isFile() && accept(e.name) && (await hasContent(p))) + return true; + if (e.isDirectory() && depth > 0 && (await anyFileWithContent(p, depth - 1, accept))) + return true; + } + return false; +} async function findSkills(root) { - const found = await scanSkillDirectories(root); + const found = new Set(); + const hollow = []; + // A skill is a directory with a loadable SKILL.md. A bare directory, or a + // SKILL.md with no frontmatter, used to count just for its name. + for (const dir of SKILL_DIRS) { + let entries; + try { + entries = await readdir(path.join(root, dir), { withFileTypes: true }); + } + catch { + continue; + } + for (const e of entries) { + if (!e.isDirectory()) + continue; + const rel = `${dir}/${e.name}/SKILL.md`; + if (await isLoadableSkill(path.join(root, rel))) + found.add(e.name); + else + hollow.push(rel); + } + } // Claude Code agents and Codex subagents can host the verifier role - const agentDirs = [ - path.join(root, '.claude', 'agents'), - path.join(root, '.codex', 'agents'), - ]; - for (const dir of agentDirs) { - if (!(await fileExists(dir))) + for (const dir of ['.claude/agents', '.codex/agents']) { + let entries; + try { + entries = await readdir(path.join(root, dir), { withFileTypes: true }); + } + catch { continue; - const entries = await readdir(dir, { withFileTypes: true }); + } for (const e of entries) { if (!e.isFile()) continue; const base = e.name.replace(/\.(md|toml)$/i, ''); - if (base.includes('verifier') || base === 'loop-verifier') { - found.push('loop-verifier'); - } + if (!(base.includes('verifier') || base === 'loop-verifier')) + continue; + const rel = `${dir}/${e.name}`; + const abs = path.join(root, rel); + const loadable = /\.md$/i.test(e.name) ? await isLoadableSkill(abs) : await hasContent(abs); + if (loadable) + found.add('loop-verifier'); + else + hollow.push(rel); } } // Opencode named agents in opencode.json (or the starter example before rename) @@ -137,7 +241,7 @@ async function findSkills(root) { for (const [key, def] of Object.entries(agents)) { const name = (def?.name ?? key).toLowerCase(); if (name.includes('verifier') || key.toLowerCase().includes('verifier')) { - found.push('loop-verifier'); + found.add('loop-verifier'); break; } } @@ -146,7 +250,15 @@ async function findSkills(root) { // ignore invalid JSON } } - return found; + return { names: [...found], hollow }; +} +/** + * L3 means unattended actions behind gates, so the gate has to be shown to + * fire, not just to exist: loop-drill must have drilled the current gate.yaml + * in both directions, and no recorded guardrail may be failing. + */ +export function guardrailsProven(proof) { + return !!proof && proof.proven.includes('gate') && proof.failed.length === 0; } /** Activity older than this does not count toward Loop Ready. */ export const ACTIVITY_MAX_AGE_MS = 14 * 24 * 60 * 60 * 1000; @@ -297,7 +409,8 @@ export function computeScore(signals) { signals.cost.runLog && signals.cost.loopMdBudget; const hasRealActivity = signals.loopActivity.present; - const l3Ready = costReady && hasRealActivity; + const proven = guardrailsProven(signals.proof); + const l3Ready = costReady && hasRealActivity && proven; let level = 'L0'; if (score >= LEVEL_THRESHOLDS.L3 && signals.verifier.present && signals.stateFile.present && l3Ready) level = 'L3'; @@ -313,28 +426,45 @@ export function computeScore(signals) { ? 'Strong signals but missing cost observability (loop-budget.md, loop-run-log.md, LOOP.md budget) — add before L3.' : score >= 82 && !hasRealActivity ? 'Strong structure but no proven loop runs yet — run one L1 cycle and commit state before L3.' - : score >= 62 - ? 'Good foundation — add missing verifier + safety docs for L3.' - : score >= 42 - ? 'Early loop setup — focus on L1 state + triage before enabling actions.' - : 'Not loop-ready — start with a starter from this repo (minimal-loop or pr-babysitter).'; + : score >= 82 && !proven + ? 'Strong structure but guardrails are unproven — run loop-drill --record and commit loop-drill.json before L3.' + : score >= 62 + ? 'Good foundation — add missing verifier + safety docs for L3.' + : score >= 42 + ? 'Early loop setup — focus on L1 state + triage before enabling actions.' + : 'Not loop-ready — start with a starter from this repo (minimal-loop or pr-babysitter).'; return { score, level, assessment }; } export async function auditProject(target) { const root = path.resolve(target); const findings = []; const recommendations = []; + // Every file signal goes through here, so an empty file never scores and is + // reported instead of silently counted. + const placeholders = []; + const present = async (rel) => { + const p = path.join(root, rel); + if (await hasContent(p)) + return true; + if (await fileExists(p)) + placeholders.push(rel); + return false; + }; const statePaths = []; for (const f of STATE_FILES) { - if (await fileExists(path.join(root, f))) + if (await present(f)) statePaths.push(f); } - const loopMd = await fileExists(path.join(root, 'LOOP.md')); - const agentsMd = await fileExists(path.join(root, 'AGENTS.md')) || - await fileExists(path.join(root, 'CLAUDE.md')); - const skillNames = await findSkills(root); + const loopMd = await present('LOOP.md'); + const agentsMd = (await present('AGENTS.md')) || (await present('CLAUDE.md')); + const skillScan = await findSkills(root); + const skillNames = skillScan.names; + const proof = await readProof(root); const loopSkills = skillNames.filter((s) => LOOP_SKILL_NAMES.includes(s)); - const verifier = skillNames.includes('loop-verifier'); + // A verifier loop-drill caught approving seeded defects is Verifier Theater: + // it earns nothing, whatever its file is called. + const verifierFound = skillNames.includes('loop-verifier'); + const verifier = verifierFound && !proof.failed.includes('verifier'); const triage = skillNames.includes('loop-triage') || skillNames.includes('pr-review-triage') || skillNames.includes('ci-triage') || @@ -346,19 +476,25 @@ export async function auditProject(target) { if (loopMd) { loopMdContent = await readFile(path.join(root, 'LOOP.md'), 'utf8'); } - // New expanded signals - const githubDir = await fileExists(path.join(root, '.github')); - const hasWorkflows = await fileExists(path.join(root, '.github', 'workflows')); + // New expanded signals. A .github/ of empty files is not a dogfooding setup. + const githubDir = await anyFileWithContent(path.join(root, '.github'), 3, () => true); + const hasWorkflows = await anyFileWithContent(path.join(root, '.github', 'workflows'), 0, (n) => /\.ya?ml$/i.test(n)); // Proper safety doc detection let safetyDocPresent = false; for (const f of SAFETY_FILES) { - if (await fileExists(path.join(root, f))) { + if (await present(f)) { safetyDocPresent = true; break; } } - const mcpPresent = (await Promise.all(MCP_FILES.map(f => fileExists(path.join(root, f))))).some(Boolean) || - /MCP|mcp server|plugins & connectors/i.test(loopMdContent); + let mcpConfig = false; + for (const f of MCP_FILES) { + if (await present(f)) { + mcpConfig = true; + break; + } + } + const mcpPresent = mcpConfig || /MCP|mcp server|plugins & connectors/i.test(loopMdContent); // Light evidence of worktree usage (common in patterns/starters/LOOP) let worktreeEvidence = false; const candidateMd = [ @@ -382,38 +518,15 @@ export async function auditProject(target) { } catch { } } - const registryPresent = await fileExists(path.join(root, 'patterns', 'registry.yaml')); - const budgetDoc = await fileExists(path.join(root, 'loop-budget.md')); - const runLog = await fileExists(path.join(root, 'loop-run-log.md')); + const registryPresent = await present('patterns/registry.yaml'); + const budgetDoc = await present('loop-budget.md'); + const runLog = await present('loop-run-log.md'); const loopMdBudget = BUDGET_HINTS.some((re) => re.test(loopMdContent)); - const budgetSkillDirs = [ - path.join(root, 'skills', 'loop-budget'), - path.join(root, '.grok', 'skills', 'loop-budget'), - path.join(root, '.claude', 'skills', 'loop-budget'), - path.join(root, '.codex', 'skills', 'loop-budget'), - ]; - let budgetSkill = false; - for (const dir of budgetSkillDirs) { - if (await fileExists(path.join(dir, 'SKILL.md'))) { - budgetSkill = true; - break; - } - } + // findSkills only lists skills with a loadable SKILL.md. + const budgetSkill = skillNames.includes('loop-budget'); const loopActivity = await detectLoopActivity(root); - const constraintsFile = await fileExists(path.join(root, 'loop-constraints.md')); - const constraintsSkillDirs = [ - path.join(root, 'skills', 'loop-constraints'), - path.join(root, '.grok', 'skills', 'loop-constraints'), - path.join(root, '.claude', 'skills', 'loop-constraints'), - path.join(root, '.codex', 'skills', 'loop-constraints'), - ]; - let constraintsSkill = false; - for (const dir of constraintsSkillDirs) { - if (await fileExists(path.join(dir, 'SKILL.md'))) { - constraintsSkill = true; - break; - } - } + const constraintsFile = await present('loop-constraints.md'); + const constraintsSkill = skillNames.includes('loop-constraints'); // Governance corpus: docs where scope / stall / escalation rules are written. let governanceCorpus = loopMdContent; for (const f of ['docs/safety.md', 'safety.md', 'SECURITY.md', 'loop-constraints.md']) { @@ -463,17 +576,23 @@ export async function auditProject(target) { ledgerPresent = rootEntries.some((e) => e.isFile() && /ledger/i.test(e.name)); } catch { } - const stallDetection = skillNames.includes('loop-context') || + const stallClaimed = skillNames.includes('loop-context') || skillNames.includes('loop-guard') || ledgerPresent || STALL_HINTS.some((re) => re.test(governanceCorpus)); + const stallDetection = stallClaimed && !proof.failed.includes('breaker'); const escalation = ESCALATION_HINTS.some((re) => re.test(governanceCorpus)); - const gateYaml = await fileExists(path.join(root, 'gate.yaml')); + // gate.yaml must look like a policy, and must not be one loop-drill found + // broken: failing, or impossible to drill at all (loop-gate refused it). + let gatePolicy = false; + if (await present('gate.yaml')) { + gatePolicy = looksLikeGatePolicy(await readFile(path.join(root, 'gate.yaml'), 'utf8')); + } + const gateYaml = gatePolicy && !proof.failed.includes('gate') && !proof.untested.includes('gate'); // Harness Runtime (harness-foundry) — versioned stack, sessions/traces, outerloop emit, host bridge const foundryStackPath = path.join(root, '.foundry', 'stack.yaml'); - const harnessStack = await fileExists(foundryStackPath); - const harnessLock = (await fileExists(path.join(root, '.foundry', 'stack.lock'))) || - (await fileExists(path.join(root, '.foundry', 'stack.lock.yaml'))); + const harnessStack = await present('.foundry/stack.yaml'); + const harnessLock = (await present('.foundry/stack.lock')) || (await present('.foundry/stack.lock.yaml')); let harnessSessions = false; const sessionsDir = path.join(root, '.foundry', 'sessions'); if (await fileExists(sessionsDir)) { @@ -494,13 +613,13 @@ export async function auditProject(target) { } catch { } } - if (!harnessEmit && (await fileExists(path.join(root, '.foundry', 'hooks', 'outerloop.yaml')))) { + if (!harnessEmit && (await present('.foundry/hooks/outerloop.yaml'))) { harnessEmit = true; } const harnessHost = (await fileExists(path.join(root, '.foundry', 'host', 'cursor'))) || (await fileExists(path.join(root, '.foundry', 'host', 'claude-code'))) || - (await fileExists(path.join(root, '.cursor', 'rules', 'foundry.mdc'))) || - (await fileExists(path.join(root, '.claude', 'foundry.md'))) || + (await present('.cursor/rules/foundry.mdc')) || + (await present('.claude/foundry.md')) || /foundry host integrate|harness-foundry/i.test(loopMdContent); const harness = { stack: harnessStack, @@ -510,12 +629,12 @@ export async function auditProject(target) { host: harnessHost, }; // Memory Engineering (memory-tiers.md, memory-budget.md) - const memoryTiers = await fileExists(path.join(root, 'memory-tiers.md')); - const memoryBudget = await fileExists(path.join(root, 'memory-budget.md')); + const memoryTiers = await present('memory-tiers.md'); + const memoryBudget = await present('memory-budget.md'); const memory = { tiers: memoryTiers, budget: memoryBudget }; // Fleet Engineering (fleet-registry.md, fleet-inbox.md) - const fleetRegistry = await fileExists(path.join(root, 'fleet-registry.md')); - const fleetInbox = await fileExists(path.join(root, 'fleet-inbox.md')); + const fleetRegistry = await present('fleet-registry.md'); + const fleetInbox = await present('fleet-inbox.md'); const fleet = { registry: fleetRegistry, inbox: fleetInbox }; const signals = { stateFile: { present: statePaths.length > 0, paths: statePaths }, @@ -538,8 +657,17 @@ export async function auditProject(target) { harness, memory, fleet, + proof, }; - const thinLoopWorkflow = await fileExists(path.join(root, '.github', 'workflows', 'thin-loop.yml')); + const thinLoopWorkflow = await present('.github/workflows/thin-loop.yml'); + if (placeholders.length > 0 || skillScan.hollow.length > 0) { + const hollow = [...new Set([...placeholders, ...skillScan.hollow])]; + findings.push({ + level: 'warn', + message: `Not counted — empty, or a skill/agent without name + description frontmatter: ${hollow.join(', ')}.`, + }); + recommendations.push('Fill in or delete placeholder files; a skill needs a SKILL.md with name and description frontmatter to load'); + } if (!signals.stateFile.present) { if (thinLoopWorkflow) { findings.push({ @@ -576,7 +704,14 @@ export async function auditProject(target) { else { findings.push({ level: 'ok', message: 'Triage skill present.' }); } - if (!signals.verifier.present) { + if (verifierFound && !verifier) { + findings.push({ + level: 'fail', + message: `Verifier present but loop-drill's canary caught it approving seeded defects (Verifier Theater) — not counted: ${proof.failures.filter((id) => id.startsWith('verifier')).join(', ')}.`, + }); + recommendations.push('Make the verifier reject the seeded defects, then rerun: npx @cobusgreyling/loop-drill . --only verifier --verifier-cmd "" --record'); + } + else if (!signals.verifier.present) { findings.push({ level: 'warn', message: 'No loop-verifier skill — maker/checker split incomplete.' }); recommendations.push('Add verifier: .grok/skills/loop-verifier, .claude/agents/loop-verifier.md, .codex/agents/verifier.toml, or a verifier agent in opencode.json'); } @@ -668,7 +803,30 @@ export async function auditProject(target) { else { findings.push({ level: 'ok', message: 'Tool/MCP scope constrained (least-privilege signal present).' }); } - if (!signals.governance.gateYaml) { + // An empty gate.yaml is already listed as a placeholder; only judge one with content. + const gateHasContent = (await fileExists(path.join(root, 'gate.yaml'))) && !placeholders.includes('gate.yaml'); + if (gatePolicy && proof.failed.includes('gate')) { + findings.push({ + level: 'fail', + message: `gate.yaml present but loop-drill shows it does not hold — not counted: ${proof.failures.filter((id) => id.startsWith('gate')).join(', ')}.`, + }); + recommendations.push('Fix gate.yaml so the failing drills pass, then rerun: npx @cobusgreyling/loop-drill . --record'); + } + else if (gatePolicy && proof.untested.includes('gate')) { + findings.push({ + level: 'fail', + message: `gate.yaml present but loop-drill could not drill it — not counted${proof.skipReasons.gate ? `: ${proof.skipReasons.gate}` : '.'}`, + }); + recommendations.push('Give gate.yaml a non-empty denylist that loop-gate accepts (see templates/gate.yaml.template), then rerun loop-drill --record'); + } + else if (gateHasContent && !gatePolicy) { + findings.push({ + level: 'fail', + message: 'gate.yaml is not a gate policy (needs `version: 1` and a `denylist:`) — loop-gate would refuse to load it. Not counted.', + }); + recommendations.push('Replace gate.yaml with templates/gate.yaml.template and customize the denylist'); + } + else if (!signals.governance.gateYaml) { findings.push({ level: 'warn', message: 'No gate.yaml — explicit human approval gates are not defined for loop-sync.', @@ -678,6 +836,38 @@ export async function auditProject(target) { else { findings.push({ level: 'ok', message: 'gate.yaml present (human gates defined).' }); } + // Proof that guardrails fire, from loop-drill --record. + if (proof.error) { + findings.push({ level: 'warn', message: `${PROOF_FILE} is unreadable (${proof.error}) — no guardrail counts as proven.` }); + recommendations.push(`Regenerate it: npx @cobusgreyling/loop-drill . --record`); + } + else if (!proof.present) { + findings.push({ + level: 'warn', + message: `No ${PROOF_FILE} — guardrails are present but unproven. Nothing shows the gate blocks a sensitive path or the verifier rejects a bad change.`, + }); + recommendations.push(`Prove the guardrails fire, then commit the record: npx @cobusgreyling/loop-drill . --record`); + } + else { + if (proof.proven.length > 0) { + findings.push({ level: 'ok', message: `Guardrails proven by loop-drill: ${proof.proven.join(', ')}.` }); + } + const failedElsewhere = proof.failed.filter((g) => g !== 'gate' && g !== 'verifier'); + if (failedElsewhere.length > 0) { + findings.push({ + level: 'fail', + message: `loop-drill shows guardrails failing to fire: ${failedElsewhere.join(', ')} (${proof.failures.filter((id) => failedElsewhere.some((g) => id.startsWith(g))).join(', ')}).`, + }); + recommendations.push('Fix the failing guardrails, then rerun: npx @cobusgreyling/loop-drill . --record'); + } + if (proof.stale.length > 0) { + findings.push({ + level: 'warn', + message: `${PROOF_FILE} is out of date for: ${proof.stale.join(', ')} (gate.yaml changed since it was drilled, or a canary is over 30 days old) — not counted.`, + }); + recommendations.push(`Re-record: npx @cobusgreyling/loop-drill . --record (add --only for canaries you ran before)`); + } + } if (!signals.governance.stallDetection) { findings.push({ level: 'warn', message: 'No stall / no-progress detection — a stuck loop can repeat the same failing action instead of escalating.' }); recommendations.push('Add loop-context (circuit breaker) or a max-attempts / no-progress rule in LOOP.md that escalates instead of looping'); @@ -787,6 +977,17 @@ export async function auditProject(target) { message: 'Score qualifies for L3 but no proven loop activity yet — capped at L2 until you run and commit at least one loop cycle.', }); } + if (score >= 78 && + signals.verifier.present && + signals.stateFile.present && + costReady && + signals.loopActivity.present && + !guardrailsProven(proof)) { + findings.push({ + level: 'warn', + message: `Score qualifies for L3 but the guardrails are unproven — capped at L2 until ${PROOF_FILE} shows the current gate.yaml passing its drills and no guardrail failing.`, + }); + } return { target: root, score, diff --git a/tools/loop-audit/dist/cli.js b/tools/loop-audit/dist/cli.js index 5cdd67a9..2a86dd20 100644 --- a/tools/loop-audit/dist/cli.js +++ b/tools/loop-audit/dist/cli.js @@ -109,6 +109,10 @@ try { console.log(' # IMPORTANT (v1.4): After scaffolding, actually RUN a loop (report-only) and commit the updated STATE.md.'); console.log(' # This creates the "loopActivity" evidence that pushes you toward real L2/L3 scores.'); console.log(''); + console.log(' # L3: prove the guardrails fire, not just exist, then commit the record'); + console.log(' npx @cobusgreyling/loop-drill . --record'); + console.log(' git add loop-drill.json'); + console.log(''); console.log('See docs/loop-design-checklist.md and patterns/ for full guidance.'); } if (!json && !badge && !md) diff --git a/tools/loop-audit/dist/proof.d.ts b/tools/loop-audit/dist/proof.d.ts new file mode 100644 index 00000000..ad1049f3 --- /dev/null +++ b/tools/loop-audit/dist/proof.d.ts @@ -0,0 +1,70 @@ +/** + * Proof that guardrails fire, read from the record `loop-drill --record` writes. + * + * Every other signal in the audit asks whether a file exists. That is how a + * repo of empty files used to score 100/L3. loop-drill runs each guardrail + * against a seeded fault and a benign case; this module reads what it found, + * so the audit can tell a guardrail that works from one that is only present. + * + * A guardrail's results only count while they still describe the repo: + * + * gate the record's sha256 of gate.yaml must match the current file. + * Gate drills are deterministic, so they never expire otherwise. + * breaker drills loop-context's defaults, which the repo does not edit. + * others (verifier, injection, ...) run real commands against code that + * keeps changing, so they expire after PROOF_MAX_AGE_MS. + */ +export declare const PROOF_FILE = "loop-drill.json"; +export declare const PROOF_SCHEMA = 1; +/** Canary results (verifier, injection) older than this no longer count. */ +export declare const PROOF_MAX_AGE_MS: number; +type Outcome = 'passed' | 'failed' | 'skipped'; +type Direction = 'sensitivity' | 'specificity'; +export interface RecordedResult { + id: string; + direction: Direction; + outcome: Outcome; + detail?: string; +} +export interface RecordedGuardrail { + recordedAt: string; + /** What was drilled: for the gate, the policy file and its fingerprint. */ + input?: { + file?: string; + sha256?: string | null; + command?: string; + }; + results: RecordedResult[]; +} +export interface DrillRecord { + schema: number; + guardrails: Record; +} +export interface ProofSignals { + /** loop-drill.json exists at the root. */ + present: boolean; + /** Why the record could not be used, when present but unreadable. */ + error?: string; + /** Current, no failures, and a passing drill in both directions. */ + proven: string[]; + /** Current, and at least one drill failed: the guardrail did not fire. */ + failed: string[]; + /** Current, but every drill was skipped, so nothing was tested. */ + untested: string[]; + /** Recorded against a different gate.yaml, or too long ago to count. */ + stale: string[]; + /** Ids of failed drills, for the report. */ + failures: string[]; + /** Skip reasons for untested guardrails, for the report. */ + skipReasons: Record; +} +/** + * sha256 of a policy file with line endings normalised, so a checkout with + * core.autocrlf fingerprints the same as the Linux runner that recorded it. + * loop-drill computes the same value; both test suites pin the same vector. + */ +export declare function fingerprint(text: string): string; +/** Validate the parsed record. Returns an error message, or null when usable. */ +export declare function validateRecord(value: unknown): string | null; +export declare function readProof(root: string, now?: number): Promise; +export {}; diff --git a/tools/loop-audit/dist/proof.js b/tools/loop-audit/dist/proof.js new file mode 100644 index 00000000..556a0bfa --- /dev/null +++ b/tools/loop-audit/dist/proof.js @@ -0,0 +1,135 @@ +/** + * Proof that guardrails fire, read from the record `loop-drill --record` writes. + * + * Every other signal in the audit asks whether a file exists. That is how a + * repo of empty files used to score 100/L3. loop-drill runs each guardrail + * against a seeded fault and a benign case; this module reads what it found, + * so the audit can tell a guardrail that works from one that is only present. + * + * A guardrail's results only count while they still describe the repo: + * + * gate the record's sha256 of gate.yaml must match the current file. + * Gate drills are deterministic, so they never expire otherwise. + * breaker drills loop-context's defaults, which the repo does not edit. + * others (verifier, injection, ...) run real commands against code that + * keeps changing, so they expire after PROOF_MAX_AGE_MS. + */ +import { createHash } from 'node:crypto'; +import { readFile } from 'node:fs/promises'; +import path from 'node:path'; +import { fileExists } from '@cobusgreyling/readiness-core'; +export const PROOF_FILE = 'loop-drill.json'; +export const PROOF_SCHEMA = 1; +/** Canary results (verifier, injection) older than this no longer count. */ +export const PROOF_MAX_AGE_MS = 30 * 24 * 60 * 60 * 1000; +/** Guardrails whose results depend only on their recorded input, not on time. */ +const TIMELESS = new Set(['gate', 'breaker']); +/** + * sha256 of a policy file with line endings normalised, so a checkout with + * core.autocrlf fingerprints the same as the Linux runner that recorded it. + * loop-drill computes the same value; both test suites pin the same vector. + */ +export function fingerprint(text) { + return createHash('sha256').update(text.replace(/\r\n/g, '\n'), 'utf8').digest('hex'); +} +function emptyProof(present, error) { + return { present, error, proven: [], failed: [], untested: [], stale: [], failures: [], skipReasons: {} }; +} +const OUTCOMES = new Set(['passed', 'failed', 'skipped']); +const DIRECTIONS = new Set(['sensitivity', 'specificity']); +function isResult(value) { + const r = value; + return (!!r && + typeof r.id === 'string' && + OUTCOMES.has(r.outcome) && + DIRECTIONS.has(r.direction) && + (r.detail === undefined || typeof r.detail === 'string')); +} +/** Validate the parsed record. Returns an error message, or null when usable. */ +export function validateRecord(value) { + const record = value; + if (!record || typeof record !== 'object') + return 'expected a JSON object'; + if (record.schema !== PROOF_SCHEMA) + return `unsupported schema ${String(record.schema)} (expected ${PROOF_SCHEMA})`; + if (!record.guardrails || typeof record.guardrails !== 'object') + return 'missing "guardrails"'; + for (const [name, g] of Object.entries(record.guardrails)) { + if (!g || typeof g.recordedAt !== 'string' || Number.isNaN(new Date(g.recordedAt).getTime())) { + return `guardrail "${name}" has no valid recordedAt`; + } + if (!Array.isArray(g.results) || !g.results.every(isResult)) { + return `guardrail "${name}" has malformed results`; + } + } + return null; +} +async function currentGateFingerprint(root) { + const p = path.join(root, 'gate.yaml'); + if (!(await fileExists(p))) + return null; + try { + return fingerprint(await readFile(p, 'utf8')); + } + catch { + return null; + } +} +/** Does this guardrail's record still describe the repo as it is now? */ +function isCurrent(name, g, gateSha, now) { + if (name === 'gate') { + return g.input?.file === 'gate.yaml' && !!g.input.sha256 && g.input.sha256 === gateSha; + } + if (TIMELESS.has(name)) + return true; + const age = now - new Date(g.recordedAt).getTime(); + return age <= PROOF_MAX_AGE_MS && age >= -60_000; +} +/** + * Proven means the guardrail caught a seeded fault *and* let a benign case + * through, with nothing failing. One direction is easy to fake: a denylist of + * "**" catches every fault, and a verifier that rejects everything catches + * every mutant. loop-drill always drills both; this insists on both. + */ +function classify(results) { + if (results.some((r) => r.outcome === 'failed')) + return 'failed'; + const passed = results.filter((r) => r.outcome === 'passed'); + const both = passed.some((r) => r.direction === 'sensitivity') && passed.some((r) => r.direction === 'specificity'); + return both ? 'proven' : 'untested'; +} +export async function readProof(root, now = Date.now()) { + const p = path.join(root, PROOF_FILE); + if (!(await fileExists(p))) + return emptyProof(false); + let parsed; + try { + parsed = JSON.parse(await readFile(p, 'utf8')); + } + catch (err) { + return emptyProof(true, `not valid JSON (${err.message})`); + } + const error = validateRecord(parsed); + if (error) + return emptyProof(true, error); + const record = parsed; + const gateSha = await currentGateFingerprint(root); + const proof = emptyProof(true); + for (const [name, g] of Object.entries(record.guardrails).sort(([a], [b]) => a.localeCompare(b))) { + if (!isCurrent(name, g, gateSha, now)) { + proof.stale.push(name); + continue; + } + const verdict = classify(g.results); + proof[verdict].push(name); + if (verdict === 'failed') { + proof.failures.push(...g.results.filter((r) => r.outcome === 'failed').map((r) => r.id)); + } + if (verdict === 'untested') { + const reason = g.results.find((r) => r.outcome === 'skipped' && r.detail)?.detail; + if (reason) + proof.skipReasons[name] = reason; + } + } + return proof; +} diff --git a/tools/loop-audit/src/auditor.ts b/tools/loop-audit/src/auditor.ts index d66e870e..553e39c7 100644 --- a/tools/loop-audit/src/auditor.ts +++ b/tools/loop-audit/src/auditor.ts @@ -1,7 +1,8 @@ -import { readdir, readFile } from 'node:fs/promises'; +import { readdir, readFile, stat } from 'node:fs/promises'; import path from 'node:path'; import { execSync } from 'node:child_process'; -import { Finding, BaseAuditResult, fileExists, scanSkillDirectories } from '@cobusgreyling/readiness-core'; +import { Finding, BaseAuditResult, fileExists } from '@cobusgreyling/readiness-core'; +import { PROOF_FILE, readProof, type ProofSignals } from './proof.js'; export interface LoopSignals { stateFile: { present: boolean; paths: string[] }; @@ -49,9 +50,12 @@ export interface LoopSignals { registry: boolean; inbox: boolean; }; + /** Guardrails loop-drill showed firing (loop-drill.json). Optional so older callers still type-check. */ + proof?: ProofSignals; } export type { Finding }; +export type { ProofSignals }; export interface AuditResult extends BaseAuditResult<'L0' | 'L1' | 'L2' | 'L3', LoopSignals> {} @@ -164,23 +168,121 @@ const ESCALATION_HINTS = [ /\bexit 2\b/i, ]; -async function findSkills(root: string): Promise { - const found = await scanSkillDirectories(root); +const SKILL_DIRS = ['.grok/skills', '.claude/skills', '.codex/skills', 'skills']; + +/** + * A signal file counts only when it has content. Empty and whitespace-only + * files, and JSON that is just {} or [], are placeholders: `touch` is not setup. + */ +export async function hasContent(p: string): Promise { + let text: string; + try { + const info = await stat(p); + if (!info.isFile()) return false; + if (info.size > 64 * 1024) return true; + text = await readFile(p, 'utf8'); + } catch { + return false; + } + const trimmed = text.replace(/^/, '').trim(); + if (!trimmed) return false; + if (p.toLowerCase().endsWith('.json')) { + try { + const value = JSON.parse(trimmed); + if (value !== null && typeof value === 'object' && Object.keys(value).length === 0) return false; + } catch { + // Not valid JSON, but still content; other checks judge the format. + } + } + return true; +} + +/** + * Frontmatter with a name and a description: what Claude Code, Codex and Grok + * need before they will load a skill or subagent. Without it the file is + * never invoked, whatever it is called. + */ +export function hasSkillFrontmatter(text: string): boolean { + const m = text.replace(/^/, '').match(/^---\r?\n([\s\S]*?)\r?\n---[ \t]*(?:\r?\n|$)/); + if (!m) return false; + return /^name:[ \t]*\S/m.test(m[1]) && /^description:[ \t]*\S/m.test(m[1]); +} + +async function isLoadableSkill(p: string): Promise { + try { + return hasSkillFrontmatter(await readFile(p, 'utf8')); + } catch { + return false; + } +} + +/** + * A gate.yaml loop-gate can load declares `version: 1` and a denylist. This is + * a shape check, not a parse; loop-drill's record proves the policy works. + */ +export function looksLikeGatePolicy(text: string): boolean { + return /^version:[ \t]*1[ \t]*(?:#.*)?$/m.test(text) && /^denylist:/m.test(text); +} + +async function anyFileWithContent(dir: string, depth: number, accept: (name: string) => boolean): Promise { + let entries; + try { + entries = await readdir(dir, { withFileTypes: true }); + } catch { + return false; + } + for (const e of entries) { + const p = path.join(dir, e.name); + if (e.isFile() && accept(e.name) && (await hasContent(p))) return true; + if (e.isDirectory() && depth > 0 && (await anyFileWithContent(p, depth - 1, accept))) return true; + } + return false; +} + +interface SkillScan { + names: string[]; + /** Skill and agent files that exist but could never be loaded. */ + hollow: string[]; +} + +async function findSkills(root: string): Promise { + const found = new Set(); + const hollow: string[] = []; + + // A skill is a directory with a loadable SKILL.md. A bare directory, or a + // SKILL.md with no frontmatter, used to count just for its name. + for (const dir of SKILL_DIRS) { + let entries; + try { + entries = await readdir(path.join(root, dir), { withFileTypes: true }); + } catch { + continue; + } + for (const e of entries) { + if (!e.isDirectory()) continue; + const rel = `${dir}/${e.name}/SKILL.md`; + if (await isLoadableSkill(path.join(root, rel))) found.add(e.name); + else hollow.push(rel); + } + } // Claude Code agents and Codex subagents can host the verifier role - const agentDirs = [ - path.join(root, '.claude', 'agents'), - path.join(root, '.codex', 'agents'), - ]; - for (const dir of agentDirs) { - if (!(await fileExists(dir))) continue; - const entries = await readdir(dir, { withFileTypes: true }); + for (const dir of ['.claude/agents', '.codex/agents']) { + let entries; + try { + entries = await readdir(path.join(root, dir), { withFileTypes: true }); + } catch { + continue; + } for (const e of entries) { if (!e.isFile()) continue; const base = e.name.replace(/\.(md|toml)$/i, ''); - if (base.includes('verifier') || base === 'loop-verifier') { - found.push('loop-verifier'); - } + if (!(base.includes('verifier') || base === 'loop-verifier')) continue; + const rel = `${dir}/${e.name}`; + const abs = path.join(root, rel); + const loadable = /\.md$/i.test(e.name) ? await isLoadableSkill(abs) : await hasContent(abs); + if (loadable) found.add('loop-verifier'); + else hollow.push(rel); } } @@ -195,7 +297,7 @@ async function findSkills(root: string): Promise { for (const [key, def] of Object.entries(agents)) { const name = (def?.name ?? key).toLowerCase(); if (name.includes('verifier') || key.toLowerCase().includes('verifier')) { - found.push('loop-verifier'); + found.add('loop-verifier'); break; } } @@ -204,7 +306,16 @@ async function findSkills(root: string): Promise { } } - return found; + return { names: [...found], hollow }; +} + +/** + * L3 means unattended actions behind gates, so the gate has to be shown to + * fire, not just to exist: loop-drill must have drilled the current gate.yaml + * in both directions, and no recorded guardrail may be failing. + */ +export function guardrailsProven(proof: ProofSignals | undefined): boolean { + return !!proof && proof.proven.includes('gate') && proof.failed.length === 0; } /** Activity older than this does not count toward Loop Ready. */ @@ -329,7 +440,8 @@ export function computeScore(signals: LoopSignals): { score: number; level: 'L0' signals.cost.runLog && signals.cost.loopMdBudget; const hasRealActivity = signals.loopActivity.present; - const l3Ready = costReady && hasRealActivity; + const proven = guardrailsProven(signals.proof); + const l3Ready = costReady && hasRealActivity && proven; let level: 'L0' | 'L1' | 'L2' | 'L3' = 'L0'; if (score >= LEVEL_THRESHOLDS.L3 && signals.verifier.present && signals.stateFile.present && l3Ready) level = 'L3'; @@ -344,11 +456,13 @@ export function computeScore(signals: LoopSignals): { score: number; level: 'L0' ? 'Strong signals but missing cost observability (loop-budget.md, loop-run-log.md, LOOP.md budget) — add before L3.' : score >= 82 && !hasRealActivity ? 'Strong structure but no proven loop runs yet — run one L1 cycle and commit state before L3.' - : score >= 62 - ? 'Good foundation — add missing verifier + safety docs for L3.' - : score >= 42 - ? 'Early loop setup — focus on L1 state + triage before enabling actions.' - : 'Not loop-ready — start with a starter from this repo (minimal-loop or pr-babysitter).'; + : score >= 82 && !proven + ? 'Strong structure but guardrails are unproven — run loop-drill --record and commit loop-drill.json before L3.' + : score >= 62 + ? 'Good foundation — add missing verifier + safety docs for L3.' + : score >= 42 + ? 'Early loop setup — focus on L1 state + triage before enabling actions.' + : 'Not loop-ready — start with a starter from this repo (minimal-loop or pr-babysitter).'; return { score, level, assessment }; } @@ -358,18 +472,32 @@ export async function auditProject(target: string): Promise { const findings: Finding[] = []; const recommendations: string[] = []; + // Every file signal goes through here, so an empty file never scores and is + // reported instead of silently counted. + const placeholders: string[] = []; + const present = async (rel: string): Promise => { + const p = path.join(root, rel); + if (await hasContent(p)) return true; + if (await fileExists(p)) placeholders.push(rel); + return false; + }; + const statePaths: string[] = []; for (const f of STATE_FILES) { - if (await fileExists(path.join(root, f))) statePaths.push(f); + if (await present(f)) statePaths.push(f); } - const loopMd = await fileExists(path.join(root, 'LOOP.md')); - const agentsMd = await fileExists(path.join(root, 'AGENTS.md')) || - await fileExists(path.join(root, 'CLAUDE.md')); - const skillNames = await findSkills(root); + const loopMd = await present('LOOP.md'); + const agentsMd = (await present('AGENTS.md')) || (await present('CLAUDE.md')); + const skillScan = await findSkills(root); + const skillNames = skillScan.names; + const proof = await readProof(root); const loopSkills = skillNames.filter((s) => LOOP_SKILL_NAMES.includes(s)); - const verifier = skillNames.includes('loop-verifier'); + // A verifier loop-drill caught approving seeded defects is Verifier Theater: + // it earns nothing, whatever its file is called. + const verifierFound = skillNames.includes('loop-verifier'); + const verifier = verifierFound && !proof.failed.includes('verifier'); const triage = skillNames.includes('loop-triage') || skillNames.includes('pr-review-triage') || skillNames.includes('ci-triage') || @@ -383,18 +511,23 @@ export async function auditProject(target: string): Promise { loopMdContent = await readFile(path.join(root, 'LOOP.md'), 'utf8'); } - // New expanded signals - const githubDir = await fileExists(path.join(root, '.github')); - const hasWorkflows = await fileExists(path.join(root, '.github', 'workflows')); + // New expanded signals. A .github/ of empty files is not a dogfooding setup. + const githubDir = await anyFileWithContent(path.join(root, '.github'), 3, () => true); + const hasWorkflows = await anyFileWithContent(path.join(root, '.github', 'workflows'), 0, (n) => + /\.ya?ml$/i.test(n), + ); // Proper safety doc detection let safetyDocPresent = false; for (const f of SAFETY_FILES) { - if (await fileExists(path.join(root, f))) { safetyDocPresent = true; break; } + if (await present(f)) { safetyDocPresent = true; break; } } - const mcpPresent = (await Promise.all(MCP_FILES.map(f => fileExists(path.join(root, f))))).some(Boolean) || - /MCP|mcp server|plugins & connectors/i.test(loopMdContent); + let mcpConfig = false; + for (const f of MCP_FILES) { + if (await present(f)) { mcpConfig = true; break; } + } + const mcpPresent = mcpConfig || /MCP|mcp server|plugins & connectors/i.test(loopMdContent); // Light evidence of worktree usage (common in patterns/starters/LOOP) let worktreeEvidence = false; @@ -416,42 +549,19 @@ export async function auditProject(target: string): Promise { } catch {} } - const registryPresent = await fileExists(path.join(root, 'patterns', 'registry.yaml')); + const registryPresent = await present('patterns/registry.yaml'); - const budgetDoc = await fileExists(path.join(root, 'loop-budget.md')); - const runLog = await fileExists(path.join(root, 'loop-run-log.md')); + const budgetDoc = await present('loop-budget.md'); + const runLog = await present('loop-run-log.md'); const loopMdBudget = BUDGET_HINTS.some((re) => re.test(loopMdContent)); - const budgetSkillDirs = [ - path.join(root, 'skills', 'loop-budget'), - path.join(root, '.grok', 'skills', 'loop-budget'), - path.join(root, '.claude', 'skills', 'loop-budget'), - path.join(root, '.codex', 'skills', 'loop-budget'), - ]; - let budgetSkill = false; - for (const dir of budgetSkillDirs) { - if (await fileExists(path.join(dir, 'SKILL.md'))) { - budgetSkill = true; - break; - } - } + // findSkills only lists skills with a loadable SKILL.md. + const budgetSkill = skillNames.includes('loop-budget'); const loopActivity = await detectLoopActivity(root); - const constraintsFile = await fileExists(path.join(root, 'loop-constraints.md')); - const constraintsSkillDirs = [ - path.join(root, 'skills', 'loop-constraints'), - path.join(root, '.grok', 'skills', 'loop-constraints'), - path.join(root, '.claude', 'skills', 'loop-constraints'), - path.join(root, '.codex', 'skills', 'loop-constraints'), - ]; - let constraintsSkill = false; - for (const dir of constraintsSkillDirs) { - if (await fileExists(path.join(dir, 'SKILL.md'))) { - constraintsSkill = true; - break; - } - } + const constraintsFile = await present('loop-constraints.md'); + const constraintsSkill = skillNames.includes('loop-constraints'); // Governance corpus: docs where scope / stall / escalation rules are written. let governanceCorpus = loopMdContent; @@ -499,21 +609,27 @@ export async function auditProject(target: string): Promise { const rootEntries = await readdir(root, { withFileTypes: true }); ledgerPresent = rootEntries.some((e) => e.isFile() && /ledger/i.test(e.name)); } catch {} - const stallDetection = + const stallClaimed = skillNames.includes('loop-context') || skillNames.includes('loop-guard') || ledgerPresent || STALL_HINTS.some((re) => re.test(governanceCorpus)); + const stallDetection = stallClaimed && !proof.failed.includes('breaker'); const escalation = ESCALATION_HINTS.some((re) => re.test(governanceCorpus)); - const gateYaml = await fileExists(path.join(root, 'gate.yaml')); + + // gate.yaml must look like a policy, and must not be one loop-drill found + // broken: failing, or impossible to drill at all (loop-gate refused it). + let gatePolicy = false; + if (await present('gate.yaml')) { + gatePolicy = looksLikeGatePolicy(await readFile(path.join(root, 'gate.yaml'), 'utf8')); + } + const gateYaml = gatePolicy && !proof.failed.includes('gate') && !proof.untested.includes('gate'); // Harness Runtime (harness-foundry) — versioned stack, sessions/traces, outerloop emit, host bridge const foundryStackPath = path.join(root, '.foundry', 'stack.yaml'); - const harnessStack = await fileExists(foundryStackPath); - const harnessLock = - (await fileExists(path.join(root, '.foundry', 'stack.lock'))) || - (await fileExists(path.join(root, '.foundry', 'stack.lock.yaml'))); + const harnessStack = await present('.foundry/stack.yaml'); + const harnessLock = (await present('.foundry/stack.lock')) || (await present('.foundry/stack.lock.yaml')); let harnessSessions = false; const sessionsDir = path.join(root, '.foundry', 'sessions'); @@ -535,15 +651,15 @@ export async function auditProject(target: string): Promise { if (/emit\/outerloop-evidence|outerloop/i.test(stackTxt)) harnessEmit = true; } catch {} } - if (!harnessEmit && (await fileExists(path.join(root, '.foundry', 'hooks', 'outerloop.yaml')))) { + if (!harnessEmit && (await present('.foundry/hooks/outerloop.yaml'))) { harnessEmit = true; } const harnessHost = (await fileExists(path.join(root, '.foundry', 'host', 'cursor'))) || (await fileExists(path.join(root, '.foundry', 'host', 'claude-code'))) || - (await fileExists(path.join(root, '.cursor', 'rules', 'foundry.mdc'))) || - (await fileExists(path.join(root, '.claude', 'foundry.md'))) || + (await present('.cursor/rules/foundry.mdc')) || + (await present('.claude/foundry.md')) || /foundry host integrate|harness-foundry/i.test(loopMdContent); const harness = { @@ -555,13 +671,13 @@ export async function auditProject(target: string): Promise { }; // Memory Engineering (memory-tiers.md, memory-budget.md) - const memoryTiers = await fileExists(path.join(root, 'memory-tiers.md')); - const memoryBudget = await fileExists(path.join(root, 'memory-budget.md')); + const memoryTiers = await present('memory-tiers.md'); + const memoryBudget = await present('memory-budget.md'); const memory = { tiers: memoryTiers, budget: memoryBudget }; // Fleet Engineering (fleet-registry.md, fleet-inbox.md) - const fleetRegistry = await fileExists(path.join(root, 'fleet-registry.md')); - const fleetInbox = await fileExists(path.join(root, 'fleet-inbox.md')); + const fleetRegistry = await present('fleet-registry.md'); + const fleetInbox = await present('fleet-inbox.md'); const fleet = { registry: fleetRegistry, inbox: fleetInbox }; const signals: LoopSignals = { @@ -585,9 +701,19 @@ export async function auditProject(target: string): Promise { harness, memory, fleet, + proof, }; - const thinLoopWorkflow = await fileExists(path.join(root, '.github', 'workflows', 'thin-loop.yml')); + const thinLoopWorkflow = await present('.github/workflows/thin-loop.yml'); + + if (placeholders.length > 0 || skillScan.hollow.length > 0) { + const hollow = [...new Set([...placeholders, ...skillScan.hollow])]; + findings.push({ + level: 'warn', + message: `Not counted — empty, or a skill/agent without name + description frontmatter: ${hollow.join(', ')}.`, + }); + recommendations.push('Fill in or delete placeholder files; a skill needs a SKILL.md with name and description frontmatter to load'); + } if (!signals.stateFile.present) { if (thinLoopWorkflow) { @@ -623,7 +749,13 @@ export async function auditProject(target: string): Promise { findings.push({ level: 'ok', message: 'Triage skill present.' }); } - if (!signals.verifier.present) { + if (verifierFound && !verifier) { + findings.push({ + level: 'fail', + message: `Verifier present but loop-drill's canary caught it approving seeded defects (Verifier Theater) — not counted: ${proof.failures.filter((id) => id.startsWith('verifier')).join(', ')}.`, + }); + recommendations.push('Make the verifier reject the seeded defects, then rerun: npx @cobusgreyling/loop-drill . --only verifier --verifier-cmd "" --record'); + } else if (!signals.verifier.present) { findings.push({ level: 'warn', message: 'No loop-verifier skill — maker/checker split incomplete.' }); recommendations.push('Add verifier: .grok/skills/loop-verifier, .claude/agents/loop-verifier.md, .codex/agents/verifier.toml, or a verifier agent in opencode.json'); } else { @@ -722,7 +854,27 @@ export async function auditProject(target: string): Promise { findings.push({ level: 'ok', message: 'Tool/MCP scope constrained (least-privilege signal present).' }); } - if (!signals.governance.gateYaml) { + // An empty gate.yaml is already listed as a placeholder; only judge one with content. + const gateHasContent = (await fileExists(path.join(root, 'gate.yaml'))) && !placeholders.includes('gate.yaml'); + if (gatePolicy && proof.failed.includes('gate')) { + findings.push({ + level: 'fail', + message: `gate.yaml present but loop-drill shows it does not hold — not counted: ${proof.failures.filter((id) => id.startsWith('gate')).join(', ')}.`, + }); + recommendations.push('Fix gate.yaml so the failing drills pass, then rerun: npx @cobusgreyling/loop-drill . --record'); + } else if (gatePolicy && proof.untested.includes('gate')) { + findings.push({ + level: 'fail', + message: `gate.yaml present but loop-drill could not drill it — not counted${proof.skipReasons.gate ? `: ${proof.skipReasons.gate}` : '.'}`, + }); + recommendations.push('Give gate.yaml a non-empty denylist that loop-gate accepts (see templates/gate.yaml.template), then rerun loop-drill --record'); + } else if (gateHasContent && !gatePolicy) { + findings.push({ + level: 'fail', + message: 'gate.yaml is not a gate policy (needs `version: 1` and a `denylist:`) — loop-gate would refuse to load it. Not counted.', + }); + recommendations.push('Replace gate.yaml with templates/gate.yaml.template and customize the denylist'); + } else if (!signals.governance.gateYaml) { findings.push({ level: 'warn', message: 'No gate.yaml — explicit human approval gates are not defined for loop-sync.', @@ -732,6 +884,37 @@ export async function auditProject(target: string): Promise { findings.push({ level: 'ok', message: 'gate.yaml present (human gates defined).' }); } + // Proof that guardrails fire, from loop-drill --record. + if (proof.error) { + findings.push({ level: 'warn', message: `${PROOF_FILE} is unreadable (${proof.error}) — no guardrail counts as proven.` }); + recommendations.push(`Regenerate it: npx @cobusgreyling/loop-drill . --record`); + } else if (!proof.present) { + findings.push({ + level: 'warn', + message: `No ${PROOF_FILE} — guardrails are present but unproven. Nothing shows the gate blocks a sensitive path or the verifier rejects a bad change.`, + }); + recommendations.push(`Prove the guardrails fire, then commit the record: npx @cobusgreyling/loop-drill . --record`); + } else { + if (proof.proven.length > 0) { + findings.push({ level: 'ok', message: `Guardrails proven by loop-drill: ${proof.proven.join(', ')}.` }); + } + const failedElsewhere = proof.failed.filter((g) => g !== 'gate' && g !== 'verifier'); + if (failedElsewhere.length > 0) { + findings.push({ + level: 'fail', + message: `loop-drill shows guardrails failing to fire: ${failedElsewhere.join(', ')} (${proof.failures.filter((id) => failedElsewhere.some((g) => id.startsWith(g))).join(', ')}).`, + }); + recommendations.push('Fix the failing guardrails, then rerun: npx @cobusgreyling/loop-drill . --record'); + } + if (proof.stale.length > 0) { + findings.push({ + level: 'warn', + message: `${PROOF_FILE} is out of date for: ${proof.stale.join(', ')} (gate.yaml changed since it was drilled, or a canary is over 30 days old) — not counted.`, + }); + recommendations.push(`Re-record: npx @cobusgreyling/loop-drill . --record (add --only for canaries you ran before)`); + } + } + if (!signals.governance.stallDetection) { findings.push({ level: 'warn', message: 'No stall / no-progress detection — a stuck loop can repeat the same failing action instead of escalating.' }); recommendations.push('Add loop-context (circuit breaker) or a max-attempts / no-progress rule in LOOP.md that escalates instead of looping'); @@ -847,6 +1030,20 @@ export async function auditProject(target: string): Promise { }); } + if ( + score >= 78 && + signals.verifier.present && + signals.stateFile.present && + costReady && + signals.loopActivity.present && + !guardrailsProven(proof) + ) { + findings.push({ + level: 'warn', + message: `Score qualifies for L3 but the guardrails are unproven — capped at L2 until ${PROOF_FILE} shows the current gate.yaml passing its drills and no guardrail failing.`, + }); + } + return { target: root, score, diff --git a/tools/loop-audit/src/cli.ts b/tools/loop-audit/src/cli.ts index 1408c2b9..57926361 100644 --- a/tools/loop-audit/src/cli.ts +++ b/tools/loop-audit/src/cli.ts @@ -108,6 +108,10 @@ try { console.log(' # IMPORTANT (v1.4): After scaffolding, actually RUN a loop (report-only) and commit the updated STATE.md.'); console.log(' # This creates the "loopActivity" evidence that pushes you toward real L2/L3 scores.'); console.log(''); + console.log(' # L3: prove the guardrails fire, not just exist, then commit the record'); + console.log(' npx @cobusgreyling/loop-drill . --record'); + console.log(' git add loop-drill.json'); + console.log(''); console.log('See docs/loop-design-checklist.md and patterns/ for full guidance.'); } diff --git a/tools/loop-audit/src/proof.ts b/tools/loop-audit/src/proof.ts new file mode 100644 index 00000000..ff03715d --- /dev/null +++ b/tools/loop-audit/src/proof.ts @@ -0,0 +1,183 @@ +/** + * Proof that guardrails fire, read from the record `loop-drill --record` writes. + * + * Every other signal in the audit asks whether a file exists. That is how a + * repo of empty files used to score 100/L3. loop-drill runs each guardrail + * against a seeded fault and a benign case; this module reads what it found, + * so the audit can tell a guardrail that works from one that is only present. + * + * A guardrail's results only count while they still describe the repo: + * + * gate the record's sha256 of gate.yaml must match the current file. + * Gate drills are deterministic, so they never expire otherwise. + * breaker drills loop-context's defaults, which the repo does not edit. + * others (verifier, injection, ...) run real commands against code that + * keeps changing, so they expire after PROOF_MAX_AGE_MS. + */ + +import { createHash } from 'node:crypto'; +import { readFile } from 'node:fs/promises'; +import path from 'node:path'; +import { fileExists } from '@cobusgreyling/readiness-core'; + +export const PROOF_FILE = 'loop-drill.json'; +export const PROOF_SCHEMA = 1; + +/** Canary results (verifier, injection) older than this no longer count. */ +export const PROOF_MAX_AGE_MS = 30 * 24 * 60 * 60 * 1000; + +/** Guardrails whose results depend only on their recorded input, not on time. */ +const TIMELESS = new Set(['gate', 'breaker']); + +type Outcome = 'passed' | 'failed' | 'skipped'; +type Direction = 'sensitivity' | 'specificity'; + +export interface RecordedResult { + id: string; + direction: Direction; + outcome: Outcome; + detail?: string; +} + +export interface RecordedGuardrail { + recordedAt: string; + /** What was drilled: for the gate, the policy file and its fingerprint. */ + input?: { file?: string; sha256?: string | null; command?: string }; + results: RecordedResult[]; +} + +export interface DrillRecord { + schema: number; + guardrails: Record; +} + +export interface ProofSignals { + /** loop-drill.json exists at the root. */ + present: boolean; + /** Why the record could not be used, when present but unreadable. */ + error?: string; + /** Current, no failures, and a passing drill in both directions. */ + proven: string[]; + /** Current, and at least one drill failed: the guardrail did not fire. */ + failed: string[]; + /** Current, but every drill was skipped, so nothing was tested. */ + untested: string[]; + /** Recorded against a different gate.yaml, or too long ago to count. */ + stale: string[]; + /** Ids of failed drills, for the report. */ + failures: string[]; + /** Skip reasons for untested guardrails, for the report. */ + skipReasons: Record; +} + +/** + * sha256 of a policy file with line endings normalised, so a checkout with + * core.autocrlf fingerprints the same as the Linux runner that recorded it. + * loop-drill computes the same value; both test suites pin the same vector. + */ +export function fingerprint(text: string): string { + return createHash('sha256').update(text.replace(/\r\n/g, '\n'), 'utf8').digest('hex'); +} + +function emptyProof(present: boolean, error?: string): ProofSignals { + return { present, error, proven: [], failed: [], untested: [], stale: [], failures: [], skipReasons: {} }; +} + +const OUTCOMES = new Set(['passed', 'failed', 'skipped']); +const DIRECTIONS = new Set(['sensitivity', 'specificity']); + +function isResult(value: unknown): value is RecordedResult { + const r = value as RecordedResult | null; + return ( + !!r && + typeof r.id === 'string' && + OUTCOMES.has(r.outcome) && + DIRECTIONS.has(r.direction) && + (r.detail === undefined || typeof r.detail === 'string') + ); +} + +/** Validate the parsed record. Returns an error message, or null when usable. */ +export function validateRecord(value: unknown): string | null { + const record = value as DrillRecord | null; + if (!record || typeof record !== 'object') return 'expected a JSON object'; + if (record.schema !== PROOF_SCHEMA) return `unsupported schema ${String(record.schema)} (expected ${PROOF_SCHEMA})`; + if (!record.guardrails || typeof record.guardrails !== 'object') return 'missing "guardrails"'; + for (const [name, g] of Object.entries(record.guardrails)) { + if (!g || typeof g.recordedAt !== 'string' || Number.isNaN(new Date(g.recordedAt).getTime())) { + return `guardrail "${name}" has no valid recordedAt`; + } + if (!Array.isArray(g.results) || !g.results.every(isResult)) { + return `guardrail "${name}" has malformed results`; + } + } + return null; +} + +async function currentGateFingerprint(root: string): Promise { + const p = path.join(root, 'gate.yaml'); + if (!(await fileExists(p))) return null; + try { + return fingerprint(await readFile(p, 'utf8')); + } catch { + return null; + } +} + +/** Does this guardrail's record still describe the repo as it is now? */ +function isCurrent(name: string, g: RecordedGuardrail, gateSha: string | null, now: number): boolean { + if (name === 'gate') { + return g.input?.file === 'gate.yaml' && !!g.input.sha256 && g.input.sha256 === gateSha; + } + if (TIMELESS.has(name)) return true; + const age = now - new Date(g.recordedAt).getTime(); + return age <= PROOF_MAX_AGE_MS && age >= -60_000; +} + +/** + * Proven means the guardrail caught a seeded fault *and* let a benign case + * through, with nothing failing. One direction is easy to fake: a denylist of + * "**" catches every fault, and a verifier that rejects everything catches + * every mutant. loop-drill always drills both; this insists on both. + */ +function classify(results: RecordedResult[]): 'proven' | 'failed' | 'untested' { + if (results.some((r) => r.outcome === 'failed')) return 'failed'; + const passed = results.filter((r) => r.outcome === 'passed'); + const both = passed.some((r) => r.direction === 'sensitivity') && passed.some((r) => r.direction === 'specificity'); + return both ? 'proven' : 'untested'; +} + +export async function readProof(root: string, now = Date.now()): Promise { + const p = path.join(root, PROOF_FILE); + if (!(await fileExists(p))) return emptyProof(false); + + let parsed: unknown; + try { + parsed = JSON.parse(await readFile(p, 'utf8')); + } catch (err) { + return emptyProof(true, `not valid JSON (${(err as Error).message})`); + } + const error = validateRecord(parsed); + if (error) return emptyProof(true, error); + + const record = parsed as DrillRecord; + const gateSha = await currentGateFingerprint(root); + const proof = emptyProof(true); + + for (const [name, g] of Object.entries(record.guardrails).sort(([a], [b]) => a.localeCompare(b))) { + if (!isCurrent(name, g, gateSha, now)) { + proof.stale.push(name); + continue; + } + const verdict = classify(g.results); + proof[verdict].push(name); + if (verdict === 'failed') { + proof.failures.push(...g.results.filter((r) => r.outcome === 'failed').map((r) => r.id)); + } + if (verdict === 'untested') { + const reason = g.results.find((r) => r.outcome === 'skipped' && r.detail)?.detail; + if (reason) proof.skipReasons[name] = reason; + } + } + return proof; +} diff --git a/tools/loop-audit/test/auditor.test.mjs b/tools/loop-audit/test/auditor.test.mjs index 2e48ee9c..30312986 100644 --- a/tools/loop-audit/test/auditor.test.mjs +++ b/tools/loop-audit/test/auditor.test.mjs @@ -58,7 +58,11 @@ test('computeScore: full L2 signals', () => { assert.ok(score >= 58 && score < 78); }); -test('computeScore: L3 requires verifier, high score, cost observability, and activity', () => { +function provenGate() { + return { present: true, proven: ['breaker', 'gate'], failed: [], untested: [], stale: [], failures: [], skipReasons: {} }; +} + +test('computeScore: L3 requires verifier, high score, cost observability, activity, and a proven gate', () => { const s = emptySignals(); s.stateFile = { present: true, paths: ['STATE.md'] }; s.triage = { present: true }; @@ -73,11 +77,41 @@ test('computeScore: L3 requires verifier, high score, cost observability, and ac s.registry = { present: true }; s.cost = { budgetDoc: true, runLog: true, loopMdBudget: true, budgetSkill: true }; s.loopActivity = { present: true, evidence: ['git:state update', 'state:STATE.md'] }; + s.proof = provenGate(); const { level, score } = computeScore(s); assert.equal(level, 'L3'); assert.ok(score >= 78); }); +test('computeScore: every L3 file present but guardrails unproven caps at L2', () => { + const s = emptySignals(); + s.stateFile = { present: true, paths: ['STATE.md'] }; + s.triage = { present: true }; + s.loopConfig = { present: true, path: 'LOOP.md' }; + s.agentsMd = { present: true }; + s.skills = { count: 3, loopSkills: ['loop-triage', 'minimal-fix', 'loop-verifier'] }; + s.verifier = { present: true }; + s.safety = { loopMdMentionsSafety: true, safetyDocPresent: true }; + s.github = { present: true, workflows: true }; + s.mcp = { present: true }; + s.worktreeEvidence = { present: true }; + s.registry = { present: true }; + s.cost = { budgetDoc: true, runLog: true, loopMdBudget: true, budgetSkill: true }; + s.loopActivity = { present: true, evidence: ['state:STATE.md'] }; + + assert.equal(computeScore(s).level, 'L2', 'no record at all'); + assert.match(computeScore(s).assessment, /guardrails are unproven/); + + s.proof = { ...provenGate(), proven: ['breaker'] }; + assert.equal(computeScore(s).level, 'L2', 'breaker alone does not prove the gate'); + + s.proof = { ...provenGate(), failed: ['injection'], failures: ['injection.visible'] }; + assert.equal(computeScore(s).level, 'L2', 'any failing guardrail blocks L3'); + + s.proof = provenGate(); + assert.equal(computeScore(s).level, 'L3'); +}); + test('computeScore: L3 blocked without cost observability', () => { const s = emptySignals(); s.stateFile = { present: true, paths: ['STATE.md'] }; diff --git a/tools/loop-audit/test/proof.test.mjs b/tools/loop-audit/test/proof.test.mjs new file mode 100644 index 00000000..ab8e6c34 --- /dev/null +++ b/tools/loop-audit/test/proof.test.mjs @@ -0,0 +1,364 @@ +import { test, describe } from 'node:test'; +import assert from 'node:assert/strict'; +import { mkdtemp, mkdir, writeFile, rm } from 'node:fs/promises'; +import { tmpdir } from 'node:os'; +import path from 'node:path'; +import { auditProject, hasContent, hasSkillFrontmatter, looksLikeGatePolicy } from '../dist/auditor.js'; +import { fingerprint, readProof, PROOF_MAX_AGE_MS } from '../dist/proof.js'; + +const today = () => new Date().toISOString().slice(0, 10); +const DAY = 24 * 60 * 60 * 1000; + +async function withDir(prefix, fn) { + const dir = await mkdtemp(path.join(tmpdir(), prefix)); + try { + return await fn(dir); + } finally { + await rm(dir, { recursive: true, force: true }); + } +} + +async function put(dir, rel, text) { + await mkdir(path.dirname(path.join(dir, rel)), { recursive: true }); + await writeFile(path.join(dir, rel), text); +} + +const skill = (name) => `---\nname: ${name}\ndescription: ${name} for this loop\n---\n\n# ${name}\n\nDo the ${name} job.\n`; + +const GATE = 'version: 1\ndenylist:\n - "**/.env"\n - "**/secrets/**"\nmaxFiles: 10\n'; + +/** Passing gate + breaker results in both directions, as loop-drill records them. */ +function goodGroups(gateText = GATE, recordedAt = new Date().toISOString()) { + return { + breaker: { + recordedAt, + results: [ + { id: 'breaker.stagnation', failureMode: 'Infinite Fix Loop', direction: 'sensitivity', outcome: 'passed' }, + { id: 'breaker.healthy', failureMode: 'Infinite Fix Loop', direction: 'specificity', outcome: 'passed' }, + ], + }, + gate: { + recordedAt, + input: { file: 'gate.yaml', sha256: fingerprint(gateText) }, + results: [ + { id: 'gate.denylist[**/.env]', failureMode: 'Over-Reach (Wrong Scope)', direction: 'sensitivity', outcome: 'passed' }, + { id: 'gate.benign', failureMode: 'Over-Reach (Wrong Scope)', direction: 'specificity', outcome: 'passed' }, + ], + }, + }; +} + +const record = (guardrails) => `${JSON.stringify({ schema: 1, tool: '@cobusgreyling/loop-drill', guardrails }, null, 2)}\n`; + +/** Everything L3 asks for, with real content. Only the proof is left to each test. */ +async function buildLoopRepo(dir) { + await put(dir, 'STATE.md', `# Loop State\n\nLast run: ${today()}\n\n## High Priority\n\n- none\n`); + await put( + dir, + 'LOOP.md', + '# Loop\n\nDaily triage, report-only. Gates: denylist in gate.yaml, no auto-merge.\n' + + 'Budget: 100k tokens/day, kill switch loop-pause-all. Worktree per run.\n' + + 'Escalate to a human when stuck; loop-context circuit breaker. MCP: github read-only (least privilege).\n', + ); + await put(dir, 'AGENTS.md', '# Agents\n\nRun `npm test` before proposing a change.\n'); + for (const s of ['loop-triage', 'loop-verifier', 'minimal-fix', 'loop-budget', 'loop-constraints']) { + await put(dir, `.claude/skills/${s}/SKILL.md`, skill(s)); + } + await put(dir, 'docs/safety.md', '# Safety\n\nNever touch secrets. Human review for every merge.\n'); + await put(dir, '.github/workflows/ci.yml', 'name: ci\non: [push]\njobs: {}\n'); + await put(dir, '.mcp.json', '{ "mcpServers": { "github": { "command": "gh-mcp" } } }\n'); + await put(dir, 'gate.yaml', GATE); + await put(dir, 'loop-budget.md', '# Budget\n\n100k tokens/day.\n'); + await put(dir, 'loop-constraints.md', '# Constraints\n\nDenylist: secrets/.\n'); + await put(dir, 'loop-run-log.md', `# Run log\n\n{"run_id":"${new Date().toISOString()}"}\n`); +} + +describe('placeholders do not score', () => { + test('the 142-byte repo of empty files no longer reaches L3', async () => { + // Byte-for-byte the repo that used to score 100/L3. + await withDir('loop-audit-potemkin-', async (dir) => { + await put(dir, 'STATE.md', `Last run: ${today()}\n`); + await put(dir, 'LOOP.md', 'gate denylist safety budget kill switch worktree MCP escalate stuck circuit breaker allowlist\n'); + for (const s of ['loop-triage', 'loop-verifier', 'minimal-fix', 'loop-constraints', 'loop-budget']) { + await put(dir, `skills/${s}/SKILL.md`, ''); + } + await put(dir, 'docs/safety.md', ''); + await put(dir, '.github/workflows/x.yml', ''); + for (const f of ['AGENTS.md', 'gate.yaml', 'loop-budget.md', 'loop-constraints.md', 'memory-tiers.md', 'memory-budget.md', 'fleet-registry.md', 'fleet-inbox.md']) { + await put(dir, f, ''); + } + await put(dir, '.mcp.json', '{}\n'); + await put(dir, 'loop-run-log.md', `{"run_id":"${today()}"}\n`); + + const r = await auditProject(dir); + assert.notEqual(r.level, 'L3'); + assert.notEqual(r.level, 'L2'); + assert.ok(r.score < 78, `score ${r.score}`); + const s = r.signals; + assert.equal(s.triage.present, false, 'empty SKILL.md is not a triage skill'); + assert.equal(s.verifier.present, false, 'empty SKILL.md is not a verifier'); + assert.equal(s.skills.count, 0); + assert.equal(s.agentsMd.present, false); + assert.equal(s.safety.safetyDocPresent, false); + assert.equal(s.github.present, false, '.github/ of empty files'); + assert.equal(s.github.workflows, false); + assert.equal(s.governance.gateYaml, false); + assert.equal(s.cost.budgetDoc, false); + assert.equal(s.cost.budgetSkill, false); + assert.equal(s.constraints.present, false); + assert.equal(s.memory.tiers, false); + assert.equal(s.fleet.registry, false); + const note = r.findings.find((f) => /^Not counted/.test(f.message)); + assert.ok(note, 'placeholders are reported, not silently dropped'); + for (const rel of ['AGENTS.md', 'gate.yaml', '.mcp.json', 'skills/loop-triage/SKILL.md']) { + assert.ok(note.message.includes(rel), rel); + } + assert.ok(!r.findings.some((f) => /not a gate policy/.test(f.message)), 'an empty gate.yaml is reported once, as a placeholder'); + }); + }); + + test('a skill directory without SKILL.md, or without frontmatter, does not count', async () => { + await withDir('loop-audit-hollow-skill-', async (dir) => { + await put(dir, 'STATE.md', '# State\n'); + await mkdir(path.join(dir, '.claude/skills/loop-triage'), { recursive: true }); + await put(dir, '.grok/skills/ci-triage/SKILL.md', '# CI triage\n\nNo frontmatter, so no host loads it.\n'); + await put(dir, '.claude/agents/verifier.md', 'You are the verifier.\n'); + const r = await auditProject(dir); + assert.equal(r.signals.triage.present, false); + assert.equal(r.signals.verifier.present, false); + const note = r.findings.find((f) => /^Not counted/.test(f.message)).message; + assert.ok(note.includes('.claude/skills/loop-triage/SKILL.md')); + assert.ok(note.includes('.grok/skills/ci-triage/SKILL.md')); + assert.ok(note.includes('.claude/agents/verifier.md')); + }); + }); + + test('a gate.yaml that is not a policy is a failure, not a signal', async () => { + await withDir('loop-audit-bad-gate-', async (dir) => { + await put(dir, 'gate.yaml', 'gates: yes please\n'); + const r = await auditProject(dir); + assert.equal(r.signals.governance.gateYaml, false); + assert.ok(r.findings.some((f) => f.level === 'fail' && /not a gate policy/.test(f.message))); + }); + }); + + test('hasContent rejects empty, whitespace and {} / [] JSON, accepts real content', async () => { + await withDir('loop-audit-content-', async (dir) => { + const cases = { 'a.md': '', 'b.md': ' \n\t\n', 'c.json': '{}', 'd.json': ' [ ] ', 'e.md': '\n', 'f.md': 'x', 'g.json': '{"a":1}', 'h.json': 'not json' }; + for (const [name, text] of Object.entries(cases)) await put(dir, name, text); + const got = {}; + for (const name of Object.keys(cases)) got[name] = await hasContent(path.join(dir, name)); + assert.deepEqual(got, { 'a.md': false, 'b.md': false, 'c.json': false, 'd.json': false, 'e.md': false, 'f.md': true, 'g.json': true, 'h.json': true }); + assert.equal(await hasContent(path.join(dir, 'missing.md')), false); + assert.equal(await hasContent(dir), false, 'a directory is not a file signal'); + }); + }); + + test('hasSkillFrontmatter needs a name and a description', () => { + assert.ok(hasSkillFrontmatter('---\nname: a\ndescription: b\n---\nbody')); + assert.ok(hasSkillFrontmatter('---\r\nname: a\r\ndescription: b\r\n---\r\n'), 'CRLF checkout'); + assert.ok(hasSkillFrontmatter('---\nname: a\ndescription: >\n folded\n---\n'), 'BOM, folded description'); + assert.ok(hasSkillFrontmatter('---\nname: a\ndescription: b\nallowed-tools: Read\n---'), 'frontmatter at EOF'); + assert.ok(!hasSkillFrontmatter('')); + assert.ok(!hasSkillFrontmatter('# Skill\nname: a\ndescription: b\n'), 'not frontmatter'); + assert.ok(!hasSkillFrontmatter('---\nname: a\n---\n'), 'no description'); + assert.ok(!hasSkillFrontmatter('---\nname:\ndescription: b\n---\n'), 'empty name'); + assert.ok(!hasSkillFrontmatter('---\nname: a\ndescription: b\n'), 'unterminated'); + }); + + test('looksLikeGatePolicy wants version 1 and a denylist', () => { + assert.ok(looksLikeGatePolicy(GATE)); + assert.ok(looksLikeGatePolicy('# policy\nversion: 1 # schema\ndenylist: []\n')); + assert.ok(!looksLikeGatePolicy('version: 2\ndenylist: []\n')); + assert.ok(!looksLikeGatePolicy('version: 1\n')); + assert.ok(!looksLikeGatePolicy('denylist:\n - x\n')); + }); +}); + +describe('readProof', () => { + test('fingerprint is sha256 of the LF-normalised file (vector shared with loop-drill)', () => { + const vector = 'a79338c9d44fec8fe07da24a3766aaf4e32281bd5c73628e0798d59944d1040e'; + assert.equal(fingerprint('version: 1\ndenylist:\n - "**/.env"\n'), vector); + assert.equal(fingerprint('version: 1\r\ndenylist:\r\n - "**/.env"\r\n'), vector); + }); + + test('no record, unreadable record, unsupported schema', async () => { + await withDir('loop-audit-proof-none-', async (dir) => { + assert.deepEqual((await readProof(dir)).present, false); + await put(dir, 'loop-drill.json', '{ nope'); + assert.match((await readProof(dir)).error, /not valid JSON/); + await put(dir, 'loop-drill.json', JSON.stringify({ schema: 2, guardrails: {} })); + assert.match((await readProof(dir)).error, /unsupported schema 2/); + await put(dir, 'loop-drill.json', record({ gate: { recordedAt: 'yesterday-ish', results: [] } })); + assert.match((await readProof(dir)).error, /no valid recordedAt/); + await put(dir, 'loop-drill.json', record({ gate: { recordedAt: new Date().toISOString(), results: [{ id: 'x', outcome: 'great' }] } })); + assert.match((await readProof(dir)).error, /malformed results/); + }); + }); + + test('passing in both directions against the current gate.yaml is proven', async () => { + await withDir('loop-audit-proof-ok-', async (dir) => { + await put(dir, 'gate.yaml', GATE); + await put(dir, 'loop-drill.json', record(goodGroups())); + const p = await readProof(dir); + assert.deepEqual(p.proven, ['breaker', 'gate']); + assert.deepEqual([p.failed, p.stale, p.untested], [[], [], []]); + }); + }); + + test('a CRLF checkout of gate.yaml matches a record made on LF', async () => { + await withDir('loop-audit-proof-crlf-', async (dir) => { + await put(dir, 'gate.yaml', GATE.replace(/\n/g, '\r\n')); + await put(dir, 'loop-drill.json', record(goodGroups(GATE))); + assert.ok((await readProof(dir)).proven.includes('gate')); + }); + }); + + test('editing gate.yaml after recording makes the gate proof stale', async () => { + await withDir('loop-audit-proof-edited-', async (dir) => { + await put(dir, 'gate.yaml', `${GATE} - "**/billing/**"\n`); + await put(dir, 'loop-drill.json', record(goodGroups(GATE))); + const p = await readProof(dir); + assert.deepEqual(p.stale, ['gate']); + assert.ok(!p.proven.includes('gate')); + assert.ok(p.proven.includes('breaker'), 'breaker drills do not depend on gate.yaml'); + }); + }); + + test('a gate proof recorded against another file, or with gate.yaml deleted, is stale', async () => { + await withDir('loop-audit-proof-otherfile-', async (dir) => { + const groups = goodGroups(); + groups.gate.input.file = 'policy/other.yaml'; + await put(dir, 'gate.yaml', GATE); + await put(dir, 'loop-drill.json', record(groups)); + assert.deepEqual((await readProof(dir)).stale, ['gate']); + await rm(path.join(dir, 'gate.yaml')); + await put(dir, 'loop-drill.json', record(goodGroups())); + assert.deepEqual((await readProof(dir)).stale, ['gate']); + }); + }); + + test('passing in one direction only is untested, not proven', async () => { + // A denylist of "**" catches every seeded fault; only the benign drill exposes it. + await withDir('loop-audit-proof-onedir-', async (dir) => { + await put(dir, 'gate.yaml', GATE); + const groups = goodGroups(); + groups.gate.results = groups.gate.results.filter((r) => r.direction === 'sensitivity'); + await put(dir, 'loop-drill.json', record(groups)); + const p = await readProof(dir); + assert.deepEqual(p.untested, ['gate']); + assert.ok(!p.proven.includes('gate')); + }); + }); + + test('a failed drill marks its guardrail failed and names the drill', async () => { + await withDir('loop-audit-proof-failed-', async (dir) => { + await put(dir, 'gate.yaml', GATE); + const groups = goodGroups(); + groups.gate.results.push({ id: 'gate.file-count', failureMode: 'Over-Reach (Wrong Scope)', direction: 'sensitivity', outcome: 'failed' }); + await put(dir, 'loop-drill.json', record(groups)); + const p = await readProof(dir); + assert.deepEqual(p.failed, ['gate']); + assert.deepEqual(p.failures, ['gate.file-count']); + }); + }); + + test('canaries expire; gate and breaker proofs do not', async () => { + await withDir('loop-audit-proof-age-', async (dir) => { + await put(dir, 'gate.yaml', GATE); + const old = new Date(Date.now() - PROOF_MAX_AGE_MS - DAY).toISOString(); + const groups = goodGroups(GATE, old); + groups.verifier = { + recordedAt: old, + input: { command: 'npm test' }, + results: [ + { id: 'verifier.control', failureMode: 'Verifier Theater', direction: 'specificity', outcome: 'passed' }, + { id: 'verifier.mutant[strict-equality-flip]', failureMode: 'Verifier Theater', direction: 'sensitivity', outcome: 'passed' }, + ], + }; + await put(dir, 'loop-drill.json', record(groups)); + const p = await readProof(dir); + assert.deepEqual(p.proven, ['breaker', 'gate']); + assert.deepEqual(p.stale, ['verifier']); + }); + }); +}); + +describe('auditProject scores what is proven', () => { + test('a complete loop reaches L3 only with a current record', async () => { + await withDir('loop-audit-l3-', async (dir) => { + await buildLoopRepo(dir); + const before = await auditProject(dir); + assert.equal(before.level, 'L2'); + assert.ok(before.score >= 78, `score ${before.score}`); + assert.ok(before.findings.some((f) => /No loop-drill\.json/.test(f.message))); + assert.ok(before.findings.some((f) => /guardrails are unproven — capped at L2/.test(f.message))); + + await put(dir, 'loop-drill.json', record(goodGroups())); + const after = await auditProject(dir); + assert.equal(after.level, 'L3'); + assert.ok(after.findings.some((f) => f.level === 'ok' && /Guardrails proven by loop-drill: breaker, gate/.test(f.message))); + + await put(dir, 'gate.yaml', GATE.replace('maxFiles: 10', 'maxFiles: 500')); + const edited = await auditProject(dir); + assert.equal(edited.level, 'L2', 'weakening gate.yaml after recording loses L3'); + assert.ok(edited.findings.some((f) => /out of date for: gate/.test(f.message))); + }); + }); + + test('a verifier the canary caught approving defects earns nothing', async () => { + await withDir('loop-audit-theater-', async (dir) => { + await buildLoopRepo(dir); + const groups = goodGroups(); + groups.verifier = { + recordedAt: new Date().toISOString(), + input: { command: 'true' }, + results: [ + { id: 'verifier.control', failureMode: 'Verifier Theater', direction: 'specificity', outcome: 'passed' }, + { id: 'verifier.mutant[strict-equality-flip]', failureMode: 'Verifier Theater', direction: 'sensitivity', outcome: 'failed' }, + ], + }; + await put(dir, 'loop-drill.json', record(groups)); + const r = await auditProject(dir); + assert.equal(r.signals.verifier.present, false); + assert.notEqual(r.level, 'L3'); + assert.ok(r.findings.some((f) => f.level === 'fail' && /Verifier Theater/.test(f.message) && f.message.includes('verifier.mutant[strict-equality-flip]'))); + }); + }); + + test('a gate loop-drill could not load, or saw fail, earns nothing', async () => { + await withDir('loop-audit-gate-broken-', async (dir) => { + await buildLoopRepo(dir); + const groups = goodGroups(); + groups.gate.results = [ + { id: 'gate', failureMode: 'Over-Reach (Wrong Scope)', direction: 'sensitivity', outcome: 'skipped', detail: 'Invalid gate config at gate.yaml: "denylist" must be an array of strings.' }, + ]; + await put(dir, 'loop-drill.json', record(groups)); + let r = await auditProject(dir); + assert.equal(r.signals.governance.gateYaml, false); + assert.ok(r.findings.some((f) => f.level === 'fail' && /could not drill it/.test(f.message) && /must be an array/.test(f.message))); + + groups.gate.results = [ + { id: 'gate.denylist[**/.env]', failureMode: 'Over-Reach (Wrong Scope)', direction: 'sensitivity', outcome: 'passed' }, + { id: 'gate.benign', failureMode: 'Over-Reach (Wrong Scope)', direction: 'specificity', outcome: 'failed' }, + ]; + await put(dir, 'loop-drill.json', record(groups)); + r = await auditProject(dir); + assert.equal(r.signals.governance.gateYaml, false); + assert.ok(r.findings.some((f) => f.level === 'fail' && /does not hold/.test(f.message) && f.message.includes('gate.benign'))); + assert.notEqual(r.level, 'L3'); + }); + }); + + test('a failing breaker withdraws stall detection', async () => { + await withDir('loop-audit-breaker-', async (dir) => { + await buildLoopRepo(dir); + const groups = goodGroups(); + groups.breaker.results[0].outcome = 'failed'; + await put(dir, 'loop-drill.json', record(groups)); + const r = await auditProject(dir); + assert.equal(r.signals.governance.stallDetection, false); + assert.ok(r.findings.some((f) => f.level === 'fail' && /failing to fire: breaker/.test(f.message))); + }); + }); +}); From e0fda157820c1e0e5ab426e4d39fee21853953ca Mon Sep 17 00:00:00 2001 From: THRISHAL12345 Date: Mon, 28 Sep 2026 20:56:35 +0530 Subject: [PATCH 4/4] chore: prove the reference repo's guardrails and keep the proof current Commit loop-drill.json so the reference repo stays at L3 on proof rather than files. ci-audit-gates.sh now fails when the record is missing, stale or failing, and prints the command to re-record. ci-validate-gates.sh already re-runs the drills against the live gate.yaml, so the committed record cannot claim more than they show. The L3 criteria in QUICKSTART, operating-loops and SECURITY now mention the proof. --- SECURITY.md | 1 + docs/QUICKSTART.md | 2 +- docs/operating-loops.md | 2 +- loop-drill.json | 135 ++++++++++++++++++++++++++++++++++++++ scripts/ci-audit-gates.sh | 16 +++++ 5 files changed, 154 insertions(+), 2 deletions(-) create mode 100644 loop-drill.json diff --git a/SECURITY.md b/SECURITY.md index 68e96450..db490a82 100644 --- a/SECURITY.md +++ b/SECURITY.md @@ -28,6 +28,7 @@ For general loop safety guidance, see [docs/safety.md](docs/safety.md). - [ ] No auto-merge without explicit allowlist - [ ] MCP connectors use least privilege - [ ] `loop-run-log.md` or equivalent observability +- [ ] Guardrails shown to fire, not just configured: `loop-drill . --record`, with `loop-drill.json` committed ## Supported versions diff --git a/docs/QUICKSTART.md b/docs/QUICKSTART.md index 55b670fc..bf723166 100644 --- a/docs/QUICKSTART.md +++ b/docs/QUICKSTART.md @@ -282,7 +282,7 @@ Commit the scaffold + first run update so `loop-audit` sees activity on the next |------|---------| | End of week one | Re-run `loop-audit . --suggest` — aim for L1 (score ~40+) | | Week two | Add a verifier skill; try one assisted fix in a worktree (L2) — see [loop-worktree](#l2-isolated-fix-attempts-loop-worktree) below | -| Before unattended (L3) | `loop-budget.md` + `loop-run-log.md` filled, human gates in `LOOP.md`, proven runs | +| Before unattended (L3) | `loop-budget.md` + `loop-run-log.md` filled, human gates in `LOOP.md`, proven runs, and `gate.yaml` proven with `npx @cobusgreyling/loop-drill . --record` (commit `loop-drill.json`) | | Unsure which pattern | [pattern-picker.md](./pattern-picker.md) · [loop-design-checklist.md](./loop-design-checklist.md) | | Something broke | [failure-modes.md](./failure-modes.md) · [stories/](../stories/) | diff --git a/docs/operating-loops.md b/docs/operating-loops.md index 1e06affd..731d1259 100644 --- a/docs/operating-loops.md +++ b/docs/operating-loops.md @@ -11,7 +11,7 @@ npx @cobusgreyling/loop-cost --pattern --cadence --level L1 npx @cobusgreyling/loop-init . --pattern # scaffolds loop-budget.md + loop-run-log.md + loop-budget skill ``` -`loop-audit` scores cost observability and caps L3 until budget + run log + LOOP.md budget section exist. +`loop-audit` scores cost observability and caps L3 until budget + run log + LOOP.md budget section exist. It also caps L3 until the guardrails are proven: `npx @cobusgreyling/loop-drill . --record` drills `gate.yaml`, and the committed `loop-drill.json` must match the current policy ([loop-audit: present is not proven](../tools/loop-audit/README.md#present-is-not-proven)). Rough planning factors: diff --git a/loop-drill.json b/loop-drill.json new file mode 100644 index 00000000..5be231d8 --- /dev/null +++ b/loop-drill.json @@ -0,0 +1,135 @@ +{ + "schema": 1, + "tool": "@cobusgreyling/loop-drill", + "guardrails": { + "breaker": { + "recordedAt": "2026-09-28T15:17:23.617Z", + "results": [ + { + "id": "breaker.stagnation", + "failureMode": "Infinite Fix Loop", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "breaker.no-progress", + "failureMode": "Infinite Fix Loop", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "breaker.token-budget", + "failureMode": "Token Burn", + "direction": "sensitivity", + "outcome": "skipped", + "detail": "No tokenBudget configured — nothing caps spend mid-run. See loop-budget.md." + }, + { + "id": "breaker.healthy", + "failureMode": "Infinite Fix Loop", + "direction": "specificity", + "outcome": "passed" + } + ] + }, + "gate": { + "recordedAt": "2026-09-28T15:17:23.617Z", + "results": [ + { + "id": "gate.denylist[**/.env]", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.denylist[**/.env.*]", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.denylist[**/secrets/**]", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.denylist[**/credentials/**]", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.denylist[**/*_key*]", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.denylist[**/*_secret*]", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.denylist[**/.terraform/**]", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.denylist[**/k8s/production/**]", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.denylist[**/migrations/**]", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.denylist[**/auth/**]", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.denylist[**/payments/**]", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.denylist[**/billing/**]", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.file-count", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.auto-merge", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "sensitivity", + "outcome": "passed" + }, + { + "id": "gate.benign", + "failureMode": "Over-Reach (Wrong Scope)", + "direction": "specificity", + "outcome": "passed" + } + ], + "input": { + "file": "gate.yaml", + "sha256": "151e4609f03384e705735874db2ddce3b73ebd136b20888b7281d4ed43f9d53b" + } + } + } +} diff --git a/scripts/ci-audit-gates.sh b/scripts/ci-audit-gates.sh index eb31e696..236c21d2 100755 --- a/scripts/ci-audit-gates.sh +++ b/scripts/ci-audit-gates.sh @@ -82,6 +82,22 @@ ROOT_AUDIT_FILE="$ROOT_AUDIT_FILE" node -e ' console.error("Reference score below L2 threshold (58). Restore dogfood signals: STATE.md, skills/, AGENTS.md."); process.exit(2); } + // loop-drill.json is the proof loop-audit scores L3 on. ci-validate-gates.sh + // re-runs the drills, so a committed record cannot claim more than they show; + // this catches a gate.yaml edited without re-recording. + const proof = data.signals.proof || {}; + const problems = []; + if (!proof.present) problems.push("loop-drill.json is missing"); + if (proof.error) problems.push("loop-drill.json is unreadable: " + proof.error); + if ((proof.stale || []).length) problems.push("stale for " + proof.stale.join(", ")); + if ((proof.failed || []).length) problems.push("failing: " + proof.failures.join(", ")); + if (!problems.length && !(proof.proven || []).includes("gate")) problems.push("gate is not proven"); + if (problems.length) { + console.error("Reference guardrails are not proven (" + problems.join("; ") + ")."); + console.error("Re-record from the repo root: (cd tools/loop-drill && npm ci && npm run build) && node tools/loop-drill/dist/cli.js . --record"); + process.exit(2); + } + console.log("Reference guardrails proven: " + proof.proven.join(", ")); ' if [[ -n "${LOOP_AUDIT_OUTPUT_FILE:-}" ]]; then