Skip to content

Commit 4adcf52

Browse files
Afgan0rclaude
andauthored
chore(skills): wire solidstats-process-skill-feedback; refresh project-standards (#33)
Adds the new self-improvement loop skill and refreshes solidstats-shared-project-standards (§A active-suggestion hook). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1 parent 863afdb commit 4adcf52

11 files changed

Lines changed: 747 additions & 48 deletions

File tree

Lines changed: 27 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,27 @@
1+
# Changelog — solidstats-process-skill-feedback
2+
3+
## v1.0 — 2026-06-22 — initial skill
4+
5+
- New direct-invoke (`disable-model-invocation`) process skill: the SolidStats self-improvement loop
6+
for the artefactual `solidstats-*` skills (conventions / code-review / tests / shared-*-standards).
7+
- Sibling of `estesis-process-review-feedback`, diverging on three deliberate decisions:
8+
- **Signal** — learns from *agent-discovered-during-work* divergence (skill wrong / incomplete /
9+
caused a bug), not only human edits to AI reviews. Five signals: divergence, gap, caused-bug,
10+
friction, preference.
11+
- **Threshold** — hybrid fact/preference: a **fact** promotes at one occurrence; a **preference**
12+
promotes at three (rule of three). The `class` field is load-bearing.
13+
- **Journal location** — in the canonical `solid-stats/skills` repo (`<skill>/corrections-log.md` +
14+
`regression-evals.jsonl`), committed next to the skill, no separate corrections repo and no ENV
15+
var (this repo is the source of truth the promotion edits). CAPTURE resolves that canonical
16+
checkout explicitly (§H) — the witness usually works in a *consuming* repo where the running copy
17+
is the vendored `.agents/skills/**` (wiped on sync), so the journal must never be written there.
18+
- Two modes: CAPTURE (per discovery; entry point via `capture-session-lessons` routing or the
19+
active-suggestion offer) and PROMOTE (batch, manual; applies the edit in-repo on user approval,
20+
leaves the commit to the user).
21+
- Boundary with MemPalace/memory: the one-question test — *would fixing this edit a `solidstats-*`
22+
SKILL.md?* — keeps skill-rule corrections here and product/code facts in MemPalace.
23+
- Active-suggestion protocol documented (§E); the proactive offer is wired into
24+
`solidstats-shared-project-standards` (auto-fires on every task), since this skill never
25+
auto-triggers.
26+
- Files: SKILL.md, workflows/{capture,promote}.md, references/{signal-taxonomy,journal-schema}.md,
27+
templates/correction-entry.md.

‎.agents/skills/solidstats-process-skill-feedback/SKILL.md‎

Lines changed: 260 additions & 0 deletions
Large diffs are not rendered by default.
Lines changed: 106 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,106 @@
1+
# Journal schema, cluster signature, and regression-eval shape
2+
3+
This is the contract the loop depends on. Capture writes it; promote reads it. Keep entries
4+
machine-parseable (stable field names) **and** human-readable — the log is a Markdown file the team
5+
diffs in git, next to the skill it teaches.
6+
7+
## Table of contents
8+
- [1. The normalized correction entry](#1-the-normalized-correction-entry)
9+
- [2. Field reference](#2-field-reference)
10+
- [3. The cluster signature](#3-the-cluster-signature)
11+
- [4. The regression-eval case](#4-the-regression-eval-case)
12+
13+
---
14+
15+
## 1. The normalized correction entry
16+
17+
Each correction is one block appended to `<target-skill>/corrections-log.md` (in this skills repo).
18+
A Markdown heading for scanning, a YAML fence for parsing.
19+
20+
```markdown
21+
### SC-2026-06-22-a3f9 · divergence · fact · §4
22+
23+
```yaml
24+
id: SC-2026-06-22-a3f9
25+
date: 2026-06-22
26+
target_skill: solidstats-server-ts-conventions
27+
repo: server-2 # server-2 | replays-fetcher | replay-parser-2 | web | n-a
28+
source: agent-discovered # agent-discovered | human-edit | free-form-prose
29+
signal: divergence # divergence | gap | caused-bug | friction | preference
30+
class: fact # fact (promote@1) | preference (promote@3)
31+
generalized: false # true only for a stated principle, not an instance
32+
section: "§4" # target skill's section/slug, or "unmapped"
33+
topic: di # short tag
34+
dev_change: >
35+
Conventions §4 says deps are wired in deps.py, but server-2 wires them in di/ (see
36+
src/di/container.ts). The rule names a file that does not exist in this repo.
37+
code:
38+
file: "src/di/container.ts"
39+
line: 1
40+
source: agent-snippet # agent-snippet | head-besteffort | none
41+
status: positive-example # negative-example | positive-example | needs-code-context | n-a
42+
snippet: |
43+
export const container = buildContainer({ ... })
44+
rationale: >
45+
Pure factual error — the skill points at a path that isn't there. Fixable at one occurrence.
46+
status: open # open | promoted | discarded
47+
signature: "divergence|§4|deps wired in di/ not deps.py"
48+
```
49+
```
50+
51+
> The example shows a real ```` ```yaml ```` fence nested in this doc; write it normally in the log.
52+
53+
## 2. Field reference
54+
55+
| Field | Meaning | Notes |
56+
|-------|---------|-------|
57+
| `id` | `SC-<date>-<4-hex>` | Stable; lets a promotion cite exact sources |
58+
| `target_skill` | Which skill this teaches | Resolved per SKILL.md §C; must be in-scope |
59+
| `repo` | Which consuming repo surfaced it | Provenance; `n-a` for a shared-standard-only fact |
60+
| `source` | How it arrived | `agent-discovered` is the SolidStats norm |
61+
| `signal` | One of the five | `references/signal-taxonomy.md` §2 |
62+
| `class` | `fact` or `preference` | **Drives the threshold** — §A.3 of SKILL.md |
63+
| `generalized` | A stated rule, not an instance | Relaxes the code requirement |
64+
| `section` | Target skill's section slot | `unmapped` if it fits no section (signals the taxonomy may need one) |
65+
| `dev_change` | The core observation: what the skill said vs. what is true | For `divergence` this carries the true fact |
66+
| `code.*` | Bound code | §G of SKILL.md — source priority and the HEAD caveat |
67+
| `code.status` | How to read the snippet | `negative-example` = should-flag; `positive-example` = should-not-flag/already-correct |
68+
| `rationale` | The "why" + the class reasoning | Decisive at promotion |
69+
| `status` | Lifecycle | `open` until promoted or discarded |
70+
| `signature` | Cluster key | See below |
71+
72+
## 3. The cluster signature
73+
74+
The rule of three (for preferences) counts **patterns**, not byte-identical entries. The signature
75+
is the cluster key:
76+
77+
```
78+
signature = "<signal>|<section>|<canonical-description>"
79+
```
80+
81+
- `signal` and `section` are exact-match buckets.
82+
- `canonical-description` is a short normalized phrase for *the same underlying issue*, written so
83+
three phrasings collapse to one. "deps wired in di/ not deps.py" should absorb "deps.py is wrong,
84+
it's di/" and "container lives in di/, conventions say deps.py".
85+
86+
At promote time, clustering is **semantic**: group entries whose `(signal, section)` match and whose
87+
descriptions mean the same thing, then show the cluster to the user before counting it. `fact`
88+
entries do not need clustering to reach a count (they promote at one) — but still de-dup so a single
89+
fact logged twice isn't proposed twice.
90+
91+
## 4. The regression-eval case
92+
93+
Optional, recommended for `code-review`/`conventions` targets. When a capture binds code, append one
94+
JSON object (one line) to `<target-skill>/regression-evals.jsonl`:
95+
96+
```json
97+
{"id": "SC-2026-06-22-a3f9", "section": "§4", "expect": "should-not-flag", "input_file": "src/di/container.ts", "snippet": "export const container = buildContainer({ ... })", "note": "deps wired in di/ not deps.py"}
98+
```
99+
100+
- `expect`: `should-flag` (from a `gap`/`caused-bug` negative example) or `should-not-flag` (from a
101+
positive example / corrected code).
102+
- `severity`: expected bucket when `should-flag`; omit otherwise.
103+
- `id` ties the case back to its journal entry so a promotion can graduate it into the target skill's
104+
core `evals/evals.json` when the rule lands.
105+
106+
These run on demand or right before a promotion — never on an always-run path.
Lines changed: 95 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,95 @@
1+
# Signal taxonomy and the fact/preference class
2+
3+
Capture sets two fields that drive everything downstream: `signal` (what kind of divergence) and
4+
`class` (fact or preference, which sets the promotion threshold). This file defines both and how each
5+
signal routes.
6+
7+
## Table of contents
8+
- [1. The two classes](#1-the-two-classes)
9+
- [2. The five signals](#2-the-five-signals)
10+
- [3. Assigning the class](#3-assigning-the-class)
11+
- [4. Routing each signal at promote](#4-routing-each-signal-at-promote)
12+
13+
---
14+
15+
## 1. The two classes
16+
17+
The class is the single most important field — it decides whether one occurrence is enough.
18+
19+
- **fact** — objectively verifiable as right or wrong. The skill names a path/import/method/command
20+
that does not exist or is wrong; the skill contradicts itself; following the skill produced a bug.
21+
A known falsehood does not need three repeats to be worth fixing. → **eligible to promote at one
22+
occurrence.**
23+
- **preference** — taste or style with no hard correctness. "Better to name it this way", "this rule
24+
reads ambiguously", "I'd structure this differently." One witness is a noisy sample. → **promote
25+
only after three** (rule of three).
26+
27+
When genuinely unsure, default to **preference** — the cost of waiting for a second witness is low;
28+
the cost of rewriting a shared skill from one opinion is high.
29+
30+
## 2. The five signals
31+
32+
### divergence — the skill is wrong about a fact
33+
The skill states X; the real code or the correct practice is Y. The classic case: a convention names
34+
a file/dir/import/method that the repo does not use.
35+
- Class: **fact**.
36+
- Evidence: the *true* fact (the real path/method), not necessarily a code snippet. Record it in
37+
`dev_change`; add a `file:line` in the consuming repo if one proves it.
38+
- Example: `solidstats-server-ts-conventions` says deps live in `deps.py`; server-2 wires them in
39+
`di/`.
40+
41+
### gap — a real pattern no rule covers
42+
Working in the repo surfaces a recurring pattern the skill is silent on. Not that the skill is wrong
43+
— that it is incomplete.
44+
- Class: **fact** when the pattern is canonical/agreed (it is simply missing); **preference** when
45+
it is a judgment call about whether the pattern *should* be the standard.
46+
- Evidence: the pattern, with code.
47+
- Example: every service in server-2 wraps external calls in a typed adapter, but no conventions rule
48+
states it.
49+
50+
### caused-bug — following the rule produced a defect
51+
The agent obeyed the skill and the result was a bug, or the rule actively steers toward one.
52+
- Class: **fact**. This is the highest-value signal — a rule that causes bugs is worse than a missing
53+
rule.
54+
- Evidence: the offending code **and** the bug (symptom or failing test).
55+
- Routing: qualify or fix the rule, and prefer adding a guardrail (a "don't do X because it causes
56+
Y" note) over a bare deletion.
57+
58+
### friction — the rule is ambiguous or contradictory
59+
The rule is unclear, internally contradictory, or contradicts another rule, and that cost real time.
60+
- Class: **fact** when it is a genuine internal contradiction (two rules cannot both hold);
61+
**preference** when it is merely "could be clearer."
62+
- Evidence: optional — name the two conflicting passages, or the ambiguity.
63+
- Routing: clarify the rule (or the shared standard, if the ambiguity lives there).
64+
65+
### preference — a stylistic improvement
66+
A better way to phrase, name, or structure something, with no correctness stake.
67+
- Class: **preference** (always).
68+
- Evidence: optional, example only.
69+
- Routing: only after rule of three; below that it stays as warming evidence.
70+
71+
## 3. Assigning the class
72+
73+
Decision order:
74+
1. Did following the skill cause a bug? → `caused-bug`, **fact**.
75+
2. Is the skill's statement objectively false (wrong path/method/command, or self-contradiction)? →
76+
`divergence` or `friction`, **fact**.
77+
3. Is a real, agreed pattern simply missing? → `gap`, **fact**.
78+
4. Otherwise it is a judgment call about what the standard *should* be → **preference** (`gap`,
79+
`friction`, or `preference` signal as fits).
80+
81+
Record the reasoning in `rationale` — it is decisive when PROMOTE adjudicates an edge call.
82+
83+
## 4. Routing each signal at promote
84+
85+
| Signal | Default target | Edit shape |
86+
|--------|----------------|-----------|
87+
| divergence | the target skill (or the shared standard if the wrong fact lives there) | correct the statement |
88+
| gap | the target skill's matching section, or a new section if `unmapped` | add/extend a rule |
89+
| caused-bug | the target skill | qualify the rule + add a guardrail note |
90+
| friction | the target skill, or the shared standard if the ambiguity is shared | clarify / reconcile |
91+
| preference | the target skill | adjust phrasing/structure, additively |
92+
93+
A correction whose true home is a `solidstats-shared-*-standards` skill (because the wrong/ambiguous
94+
rule is *shared*, not stack-specific) routes there, not to the per-stack skill — the same delegation
95+
the skills already use. State the routing in one line and let the user redirect.
Lines changed: 68 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,68 @@
1+
# Correction entry template
2+
3+
Append one block per correction to `<target-skill>/corrections-log.md` (in this skills repo). Heading
4+
for human scanning, YAML for machine parsing. Field meanings: `references/journal-schema.md`.
5+
6+
```markdown
7+
### <id> · <signal> · <class> · <section>
8+
9+
```yaml
10+
id: SC-<YYYY-MM-DD>-<4hex>
11+
date: <YYYY-MM-DD>
12+
target_skill: <solidstats-skill-name>
13+
repo: <server-2 | replays-fetcher | replay-parser-2 | web | n-a>
14+
source: <agent-discovered | human-edit | free-form-prose>
15+
signal: <divergence | gap | caused-bug | friction | preference>
16+
class: <fact | preference>
17+
generalized: <true | false>
18+
section: <"§X" or "unmapped">
19+
topic: <short-tag>
20+
dev_change: >
21+
<what the skill says vs. what is true; for divergence carry the true fact>
22+
code:
23+
file: <"path" or null>
24+
line: <N or null>
25+
source: <agent-snippet | head-besteffort | none>
26+
status: <negative-example | positive-example | needs-code-context | n-a>
27+
snippet: |
28+
<the few relevant lines, or omit>
29+
rationale: >
30+
<the "why" and the class reasoning (why fact vs preference)>
31+
status: open
32+
signature: "<signal>|<section>|<canonical-description>"
33+
```
34+
```
35+
36+
## A filled example
37+
38+
```markdown
39+
### SC-2026-06-22-7c1d · caused-bug · fact · §AA
40+
41+
```yaml
42+
id: SC-2026-06-22-7c1d
43+
date: 2026-06-22
44+
target_skill: solidstats-shared-backend-ts-standards
45+
repo: server-2
46+
source: agent-discovered
47+
signal: caused-bug
48+
class: fact
49+
generalized: false
50+
section: "§AA"
51+
topic: observability
52+
dev_change: >
53+
§AA's logging example logs the full request object, which includes the auth header; following it
54+
leaked a bearer token into the logs. The rule should mandate redaction of the auth header.
55+
code:
56+
file: "src/plugins/logging.ts"
57+
line: 42
58+
source: agent-snippet
59+
status: negative-example
60+
snippet: |
61+
req.log.info({ req }, 'incoming request')
62+
rationale: >
63+
Following the rule as written produces a security defect — fact, fixable at one occurrence. Add a
64+
guardrail rather than deleting the logging guidance.
65+
status: open
66+
signature: "caused-bug|§AA|request logging leaks auth header"
67+
```
68+
```
Lines changed: 67 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,67 @@
1+
# CAPTURE — turn one divergence into a journal entry
2+
3+
Run this per discovery, the moment a `solidstats-*` skill is shown wrong/incomplete/harmful — or when
4+
`capture-session-lessons` (or a proactive offer from the active skill, SKILL.md §E) routes a lesson
5+
here. Goal: classify it, bind the offending code, normalize, append to the journal in **this skills
6+
repo**, emit a regression case when code is bound, and stage it for commit. It never edits a target
7+
skill — that is PROMOTE's job.
8+
9+
## Required reading
10+
- `references/signal-taxonomy.md` — the five signals and the fact/preference class rules
11+
- `references/journal-schema.md` — the entry schema and signature
12+
- `templates/correction-entry.md` — the block to append
13+
14+
## Steps
15+
16+
0. **Resolve the canonical skills checkout** (SKILL.md §H). The journal must land in the real
17+
`solid-stats/skills` repo, NOT the vendored `.agents/skills/**` copy you may be running from in a
18+
consuming repo (it is overwritten on the next sync). Resolve: cwd/ancestor whose remote is
19+
`solid-stats/skills` → else a local clone (`../skills` or under `~/Projects/**`, remote confirmed)
20+
→ else stop and ask. All journal paths below are relative to that checkout.
21+
22+
1. **Confirm scope and resolve the target skill** (SKILL.md §C, §D). First apply the boundary test:
23+
*would fixing this mean editing a `solidstats-*` SKILL.md or its references?* If no — it is a
24+
product/code fact, not a skill divergence; stop and route it to memory/MemPalace instead. If yes,
25+
resolve which skill: infer from the cited section/skill name; if uncertain, ask the user to pick
26+
from the in-scope candidates. A misrouted correction pollutes the wrong skill's journal. Read that
27+
skill's section list so `section`/`topic` use its vocabulary.
28+
29+
2. **Classify.** Set `signal` (divergence / gap / caused-bug / friction / preference) and `class`
30+
(fact / preference) per the taxonomy decision order. The class is the load-bearing field — get it
31+
right, because it decides whether one occurrence can promote. Record the class reasoning in
32+
`rationale`.
33+
34+
3. **Bind the evidence** (SKILL.md §G):
35+
- For a **divergence**, the evidence is the *true fact* (the real path/method) — put it in
36+
`dev_change`; add a `file:line` in the consuming repo if one proves it.
37+
- For a **gap** / **caused-bug**, bind the code: prefer the exact snippet the agent was looking at
38+
(`code.source: agent-snippet`); else best-effort from local `HEAD` by `file:line`, honoring the
39+
HEAD caveat (an already-fixed snippet is a `positive-example`). For `caused-bug`, also capture the
40+
bug (symptom or failing test) in `dev_change`.
41+
- Unbindable and not a principle → `needs-code-context`; capture it but exclude from promotion
42+
until resolved.
43+
44+
4. **Author the signature** (`signal|section|canonical-description`), wording the description so a
45+
future near-duplicate collapses onto it (journal-schema §3).
46+
47+
5. **Append the entry** to `<target-skill>/corrections-log.md` using the template. Create the file on
48+
first use for that target. Never rewrite existing entries — append only; the log is an audit trail.
49+
50+
6. **Emit a regression case** (optional, recommended for conventions/code-review targets). For an
51+
entry with bound code, append one JSONL line to `<target-skill>/regression-evals.jsonl`
52+
(journal-schema §4). Skip `needs-code-context` entries.
53+
54+
7. **Stage for commit.** `git add` the journal files (`corrections-log.md`, and
55+
`regression-evals.jsonl` if written). **Do not commit or push** unless the user asks — AGENTS.md
56+
forbids autonomous commits, and this repo is shared truth. Tell the user the entry is staged.
57+
58+
8. **Print the soft nudge.** For this target, count `open` entries, `fact` entries ready to promote
59+
(each promotable at one), and preference clusters at/over three. Report, e.g.:
60+
`📊 solidstats-server-ts-conventions: 4 open · 1 fact ready · 0 preference-clusters ≥3.`
61+
Information, not an instruction — the user decides when to run PROMOTE.
62+
63+
## Output
64+
65+
A short summary: the target skill, the signal+class captured, whether code was bound (and as a
66+
positive or negative example), any `needs-code-context` flag, and the nudge line. Do not propose a
67+
skill edit here.

0 commit comments

Comments
 (0)