Skip to content
Merged
Show file tree
Hide file tree
Changes from 3 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 16 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,22 @@ All notable changes to Skill++. The format follows

## [Unreleased]

### Added

- A memory guard for the local model. Measured on an 18 GB Mac, loading
`gemma4:e4b` and the embedder took 12.8 GB of free memory, and a fold with
other apps open ran at 92 % used with macOS swapping. There was no leak; the
model is simply large. Now a fold loads the models only when they fit with
2 GB to spare, stops and unloads them at once if free memory falls below
that, and unloads them as soon as it ends instead of after Ollama's five
minutes. What the models take is measured on each computer; with
`gemma4:e4b-it-qat` a fold starts at about 9 GB free. A session that doesn't
fit is kept and folded later: when the computer is idle, at the next session
start, or with **Fold now** on the review page. A desktop notification says
when a fold waits or is stopped, and `doctor` and `stats` show the figures.
`SKILL_PLUS_PLUS_MEMORY_GUARD`, `SKILL_PLUS_PLUS_MEMORY_RESERVE_GB`,
`SKILL_PLUS_PLUS_IDLE_MINUTES` and `SKILL_PLUS_PLUS_NOTIFY` configure it.

### Changed

- `install --apply` ends by saying what to do next: start a new Claude Code
Expand Down
11 changes: 7 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -153,8 +153,10 @@ new Code session in the desktop app) and work as usual.
- `install` without `--apply` only shows what it would do, downloads included.
Your settings file is backed up before it is changed.
- `--user` instead of `--project` captures every project on the machine.
- The two models are about 10 GB on disk, and cutting a session needs about
10 GB of free memory. `--no-models` leaves Ollama alone.
- The two models are about 6.4 GB on disk, and folding a session takes about
7 GB of free memory. Skill++ only starts when that fits with 2 GB to spare,
and otherwise waits, so it never pushes your machine into swap
([Memory](docs/usage.md#memory)). `--no-models` leaves Ollama alone.
- `skill-plus-plus install --project ~/code/my-repo --remove --apply` takes the hooks
and slash commands out again.
- The latest `main`, before it is released:
Expand Down Expand Up @@ -309,8 +311,9 @@ including sessions of our own work that we cannot publish.
## 🧭 When to use · when to skip

**Good fit if you** repeat procedures in Claude Code (the CLI or the desktop
app's Code tab), on macOS or Linux, with about 10 GB of memory to spare for the
local model.
app's Code tab), on macOS or Linux, with about 9 GB of memory free now and
then: the local model takes about 7 GB while it folds a session, and Skill++
waits until that fits.

**Skip it if you** mostly do one-off work, run Windows, or use another agent:
capture is Claude Code only for now. Drafting can use any agent CLI
Expand Down
47 changes: 47 additions & 0 deletions docs/usage.md
Original file line number Diff line number Diff line change
Expand Up @@ -178,6 +178,43 @@ edits settings or spends a model call is a dry run until you add `--apply`.
| Skills you have | `lifecycle`, `tier`, `check`, `reconcile`, `bundle`, `expire`, `accuracy` |
| Internal (run by the hooks) | `hook`, `fold-session`, `fold-pending` |

## Memory

The local model is the one large thing Skill++ runs. On an 18 GB Mac, a fold
with `gemma4:e4b-it-qat` and `nomic-embed-text` takes about **7 GB** of free
memory, whatever the session: a one-word question costs as much as a whole
session. `gemma4:e4b`, the default before it, took 12.8 GB, and with other
apps open a fold ran at 92 % memory used while macOS swapped out 4.7 GB in
100 seconds. So Skill++ guards its memory use:

- **It starts only if the models fit.** A fold loads the models only when
what they take still leaves 2 GB free: with the default models, from about
9 GB free. What they take is measured on your computer: estimated from
their size at first, which asks for about 10 GB free, then the largest
amount a fold has actually used here.
- **It stops when memory runs short.** If free memory falls below 2 GB while a
fold runs, or macOS reports critical memory pressure, Skill++ stops and
unloads the models at once.
- **It lets go right away.** When a fold ends, the models are unloaded, rather
than staying in memory for Ollama's usual five minutes.
- **It never unloads what it didn't load.** A model you had loaded yourself
stays loaded.

A session that doesn't fit isn't lost. It waits, and is folded:

- **when your computer is idle,** five minutes without keyboard or mouse input,
once the models fit;
- **at the next session start,** retried at most every ten minutes;
- **when you ask:** **Fold now** on the review page's banner, or
`skill-plus-plus fold-pending --now`. The models still load only if they fit.

When a fold waits or is stopped, a desktop notification says so: Notification
Center on macOS, `notify-send` on Linux. On a Mac the first one may ask you to
allow notifications for Script Editor, which is what shows them.
`skill-plus-plus doctor` shows what the models take on your computer and from
what free memory a fold starts; `skill-plus-plus stats` counts the sessions
waiting.

## Configuration

All settings are environment variables.
Expand All @@ -198,6 +235,10 @@ All settings are environment variables.
| `SKILL_PLUS_PLUS_MATCH` | `1` | `0` banks every task without comparing it (for measuring detection) |
| `SKILL_PLUS_PLUS_DESCRIBE` | `0` | `1` asks the model to describe every tool call, inside the hook (slow) |
| `SKILL_PLUS_PLUS_MAX_STEPS` / `SKILL_PLUS_PLUS_MAX_FIELD` | `500` / `2000` | caps per session and per captured field |
| `SKILL_PLUS_PLUS_MEMORY_GUARD` | `1` | `0` turns the [memory guard](#memory) off: folds load the models whatever is free |
| `SKILL_PLUS_PLUS_MEMORY_RESERVE_GB` | `2` | free memory a fold always leaves; it starts only if the models fit with this much to spare, and stops below it |
| `SKILL_PLUS_PLUS_IDLE_MINUTES` | `5` | how long nobody must use the keyboard or mouse before a session waiting for memory is folded |
| `SKILL_PLUS_PLUS_NOTIFY` | `1` | `0` turns off the desktop notices when a fold waits for memory or is stopped |
| `SKILL_PLUS_PLUS_INTERNAL` | unset | set to anything to make the hooks do nothing, e.g. for one session |

## Troubleshooting
Expand Down Expand Up @@ -225,6 +266,12 @@ running. Stop it with Ctrl-C in its terminal, or use `skill-plus-plus web --port
**Tool calls feel slow.** Check that `SKILL_PLUS_PLUS_DESCRIBE` is not set to `1`; it asks
the local model about every tool call inside the hook.

**Sessions keep waiting for memory.** `skill-plus-plus doctor` says from what
free memory a fold starts. Close a few apps and press **Fold now** on the
review page, or leave the computer idle for a few minutes. If `doctor` says the
models take more than the machine has, sessions never fold while the guard is
on; `SKILL_PLUS_PLUS_MEMORY_GUARD=0` folds them anyway, at the risk of swapping.

**Leave one session out.** Start it as `SKILL_PLUS_PLUS_INTERNAL=1 claude`, and the hooks
do nothing for it.

Expand Down
209 changes: 170 additions & 39 deletions skill_plus_plus/capture.py
Original file line number Diff line number Diff line change
Expand Up @@ -548,22 +548,47 @@ def handle_session_end(config: Config, payload: dict) -> dict:
# keep path already attaches before judging.
_attach_reply(session, payload.get("transcript_path") or session.get("transcript"))

if config.judge_boundaries and "folded" not in session:
try:
from .boundary import judge_session
judge_session(config, session)
except Exception as exc: # noqa: BLE001 - a hook never raises at a dev
log_error(config, f"boundary judge failed: {type(exc).__name__}: {exc}")

# The last task's own completion report, said after its final tool call.
# Attached before folding so the episode carries it.
trailing = _trailing_narration(payload)
if trailing:
work = [s for s in session.get("steps", []) if not is_prompt(s)]
if work:
work[-1].setdefault("closing_note", trailing)
_save_session(config, session)
result = fold_session(config, session, persist=True)
# The memory guard (`skill_plus_plus.memory`) around everything that loads
# a model. The models load only if they fit and still leave the reserve
# free; otherwise the session is held and folded later, which costs
# nothing. A watchdog stops the fold and unloads them if memory runs short
# while it runs, and they are unloaded when it ends.
from . import memory
with memory.guarded(config) as guard:
short = guard.admit(memory.fold_models(config))
if short:
return _hold_for_memory(config, session, short)

if config.judge_boundaries and "folded" not in session:
try:
from .boundary import judge_session
judge_session(config, session)
except Exception as exc: # noqa: BLE001 - a hook never raises at a dev
log_error(config, f"boundary judge failed: {type(exc).__name__}: {exc}")
if guard.tripped:
# Stopped halfway through judging. The gaps judged so far carry
# verdicts and the rest carry `judge_session`'s up-front
# `False`, which reads as "no boundary here": banked, the
# session would merge tasks the judge never got to. Nothing is
# banked yet, so every verdict goes and the retry judges it all.
for step in session.get("steps", []):
step.pop("end", None)
return _hold_for_memory(config, session, guard.tripped, notify=False)

# The last task's own completion report, said after its final tool call.
# Attached before folding so the episode carries it.
trailing = _trailing_narration(payload)
if trailing:
work = [s for s in session.get("steps", []) if not is_prompt(s)]
if work:
work[-1].setdefault("closing_note", trailing)
_save_session(config, session)
result = fold_session(config, session, persist=True)
if guard.tripped and result.get("status") == "offline":
# Stopped between episodes: `folded` keeps the ones banked, and the
# retry resumes after them, as it does when Ollama goes away.
result["reason"] = f"memory: {guard.tripped}"
result["memory"] = True
# Offline: the session was never folded, so this file is the only copy of
# the work. Losing the step is the one thing capture exists to prevent —
# being offline costs the candidate, never the record.
Expand All @@ -576,6 +601,8 @@ def handle_session_end(config: Config, payload: dict) -> dict:
"at": datetime.now(timezone.utc).isoformat(timespec="seconds"),
"reason": result.get("reason", ""),
}
if result.get("memory"):
session["held"]["memory"] = True
_save_session(config, session)
else:
try:
Expand All @@ -585,6 +612,29 @@ def handle_session_end(config: Config, payload: dict) -> dict:
return result


def _hold_for_memory(config: Config, session: dict, reason: str, *,
notify: bool = True) -> dict:
"""Keep the session to fold later: the models do not fit in memory now.

Held like a session no model answered, so `fold_pending` retries it and
`stats` counts it, with `memory` set so the retry waits for memory rather
than for Ollama. *notify* is off when the watchdog stopped the fold, which
has already said so.
"""
session["held"] = {
"at": datetime.now(timezone.utc).isoformat(timespec="seconds"),
"reason": f"memory: {reason}",
"memory": True,
}
_save_session(config, session)
if notify:
from .notify import waiting
waiting(config, reason)
return {"status": "offline", "steps": len(session.get("steps", [])),
"episodes": [], "flagged": 0, "reason": session["held"]["reason"],
"memory": True}



def mark_ending(config: Config, session_id: str, transcript: str | None = None) -> dict:
"""Record that a session ended. The fold itself happens elsewhere.
Expand Down Expand Up @@ -648,6 +698,19 @@ def fold_session_now(config: Config, session_id: str, *,
# A ceiling on one session's fold. Bounds the damage a recycled pid can do:
# past this the lock is stale whatever `os.kill` says.
_FOLD_LOCK_SECONDS = 1800
# How often the sweep retries a session held for memory (see `_is_pending`).
MEMORY_RETRY_SECONDS = 600.0


def _age_seconds(stamp) -> float:
"""Seconds since an ISO stamp. An unreadable one counts as long ago."""
try:
at = datetime.fromisoformat(str(stamp))
except ValueError:
return float("inf")
if at.tzinfo is None:
at = at.replace(tzinfo=timezone.utc)
return (datetime.now(timezone.utc) - at).total_seconds()


def _lock_file(config: Config, session_id: str) -> Path:
Expand Down Expand Up @@ -751,7 +814,8 @@ def _find_transcript(session_id: str) -> str | None:


def fold_pending(config: Config, *, exclude: str | None = None,
idle_hours: float = PENDING_IDLE_HOURS) -> list[dict]:
idle_hours: float = PENDING_IDLE_HOURS,
now: bool = False) -> list[dict]:
"""Bank the sessions that ended without being banked.

`SessionEnd` stamps `ending` and spawns a worker, so this sweep is the
Expand All @@ -762,48 +826,115 @@ def fold_pending(config: Config, *, exclude: str | None = None,
per-session lock and the held stamp are the ones a normal end uses, and
`folded` keeps a partial retry from counting an episode twice. *exclude* is
the session that is starting, which is live by definition.

One memory guard covers the sweep, so the models load once for every
session in it and are unloaded at the end. The first session held for
memory ends the sweep: what did not fit for it will not fit for the next.
*now* skips the pause between memory retries, for the idle waiter and
"Fold now", which have just checked that the models fit.
"""
from . import memory

results = []
with _locked(config.root / "fold-pending.lock", _PENDING_LOCK_SECONDS,
"fold-pending") as got:
if not got:
return [{"status": "locked"}]
for path in sorted(config.sessions_dir.glob("*.json")):
sid = path.stem
if sid == exclude:
continue
try:
session = json.loads(path.read_text(encoding="utf-8"))
idle = (time.time() - path.stat().st_mtime) / 3600
except (OSError, json.JSONDecodeError):
continue
if not _is_pending(session, idle, idle_hours):
results.append({"session": sid, "status": "live"})
continue
# Checked before folding only so the outcome can say "folding"
# rather than "locked"; `fold_session_now` takes the lock itself,
# so nothing rests on this being race-free.
if _lock_alive(_lock_file(config, sid), _FOLD_LOCK_SECONDS):
results.append({"session": sid, "status": "folding"})
continue
results.append({"session": sid, **fold_session_now(config, sid)})
with memory.guarded(config):
for path in sorted(config.sessions_dir.glob("*.json")):
sid = path.stem
if sid == exclude:
continue
try:
session = json.loads(path.read_text(encoding="utf-8"))
idle = (time.time() - path.stat().st_mtime) / 3600
except (OSError, json.JSONDecodeError):
continue
if not _is_pending(session, idle, idle_hours, now=now):
results.append({"session": sid, "status": "live"})
continue
# Checked before folding only so the outcome can say "folding"
# rather than "locked"; `fold_session_now` takes the lock itself,
# so nothing rests on this being race-free.
if _lock_alive(_lock_file(config, sid), _FOLD_LOCK_SECONDS):
results.append({"session": sid, "status": "folding"})
continue
result = fold_session_now(config, sid)
results.append({"session": sid, **result})
if result.get("memory"):
break
_sweep_orphan_locks(config)
return results


def _is_pending(session: dict, idle: float, idle_hours: float) -> bool:
# How long a worker that held a session for memory waits for a better moment,
# and how often it looks. Past this, the next session start takes over.
MEMORY_WAIT_SECONDS = 12 * 3600
MEMORY_POLL_SECONDS = 60.0


def _held_for_memory(config: Config) -> bool:
for path in config.sessions_dir.glob("*.json"):
try:
if (json.loads(path.read_text(encoding="utf-8")).get("held") or {}).get("memory"):
return True
except (OSError, json.JSONDecodeError, AttributeError):
continue
return False


def wait_for_memory(config: Config, *, sleep=time.sleep, clock=time.time) -> list[dict]:
"""Fold the sessions held for memory once the computer is idle and they fit.

Run by the worker that held a session, after its own lock is released:
nobody waits on it, and it costs a sleeping process. One waiter at a time;
another worker that finds the lock taken leaves it to the one waiting.
Idle means nobody at the keyboard for `config.idle_minutes`, so the model
never competes with someone working. Where idle time cannot be read, the
waiter waits for memory alone.
"""
from . import memory

with _locked(config.root / "memory-wait.lock", MEMORY_WAIT_SECONDS + 600,
"memory-wait") as got:
if not got:
return []
deadline = clock() + MEMORY_WAIT_SECONDS
while clock() < deadline:
sleep(MEMORY_POLL_SECONDS)
if not _held_for_memory(config):
return []
idle = memory.idle_seconds()
if idle is not None and idle < config.idle_minutes * 60:
continue
if memory.shortfall(config):
continue
results = fold_pending(config, now=True)
if not _held_for_memory(config):
return results
return []


def _is_pending(session: dict, idle: float, idle_hours: float, *,
now: bool = False) -> bool:
"""Has this session ended without being banked?

Three rules, covering three different failures, none of them redundant:

* ``held`` — a fold that ran and could not reach a model.
* ``held`` — a fold that ran and could not reach a model, or was held for
memory. Those are retried at most every `MEMORY_RETRY_SECONDS`, unless
*now*: a retry that passes the check and trips the watchdog again would
load and unload the models at every session start.
* ``ending`` past the grace — a fold that was launched and died with it.
Without the stamp it would look like a live session and wait out the idle
rule.
* idle — the only rule that catches a session where `SessionEnd` never
fired at all: a window closed, a laptop shut down.
"""
if session.get("held"):
held = session.get("held")
if held:
if isinstance(held, dict) and held.get("memory") and not now:
return _age_seconds(held.get("at")) >= MEMORY_RETRY_SECONDS
return True
ending = session.get("ending")
if ending:
Expand Down
Loading
Loading