diff --git a/CHANGELOG.md b/CHANGELOG.md
index 28fc40f..4ff54aa 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -6,6 +6,22 @@ All notable changes to Skill++. The format follows
## [Unreleased]
+### Added
+
+- A memory guard for the local model. Measured on an 18 GB Mac, loading
+ `gemma4:e4b` and the embedder took 12.8 GB of free memory, and a fold with
+ other apps open ran at 92 % used with macOS swapping. There was no leak; the
+ model is simply large. Now a fold loads the models only when they fit with
+ 2 GB to spare, stops and unloads them at once if free memory falls below
+ that, and unloads them as soon as it ends instead of after Ollama's five
+ minutes. What the models take is measured on each computer; with
+ `gemma4:e4b-it-qat` a fold starts at about 9 GB free. A session that doesn't
+ fit is kept and folded later: when the computer is idle, at the next session
+ start, or with **Fold now** on the review page. A desktop notification says
+ when a fold waits or is stopped, and `doctor` and `stats` show the figures.
+ `SKILL_PLUS_PLUS_MEMORY_GUARD`, `SKILL_PLUS_PLUS_MEMORY_RESERVE_GB`,
+ `SKILL_PLUS_PLUS_IDLE_MINUTES` and `SKILL_PLUS_PLUS_NOTIFY` configure it.
+
### Changed
- The local model is `gemma4:e4b-it-qat`, a build of `gemma4:e4b` trained to
diff --git a/README.md b/README.md
index 689f02c..d90031f 100644
--- a/README.md
+++ b/README.md
@@ -153,8 +153,10 @@ new Code session in the desktop app) and work as usual.
- `install` without `--apply` only shows what it would do, downloads included.
Your settings file is backed up before it is changed.
- `--user` instead of `--project` captures every project on the machine.
-- The two models are about 6.4 GB on disk, and cutting a session needs about
- 7 GB of free memory. `--no-models` leaves Ollama alone.
+- The two models are about 6.4 GB on disk, and folding a session takes about
+ 7 GB of free memory. Skill++ only starts when that fits with 2 GB to spare,
+ and otherwise waits, so it never pushes your machine into swap
+ ([Memory](docs/usage.md#memory)). `--no-models` leaves Ollama alone.
- `skill-plus-plus install --project ~/code/my-repo --remove --apply` takes the hooks
and slash commands out again.
- The latest `main`, before it is released:
@@ -311,8 +313,9 @@ including sessions of our own work that we cannot publish.
## π§ When to use Β· when to skip
**Good fit if you** repeat procedures in Claude Code (the CLI or the desktop
-app's Code tab), on macOS or Linux, with about 7 GB of memory to spare for the
-local model.
+app's Code tab), on macOS or Linux, with about 9 GB of memory free now and
+then: the local model takes about 7 GB while it folds a session, and Skill++
+waits until that fits.
**Skip it if you** mostly do one-off work, run Windows, or use another agent:
capture is Claude Code only for now. Drafting can use any agent CLI
diff --git a/docs/usage.md b/docs/usage.md
index b6aedc1..f368a27 100644
--- a/docs/usage.md
+++ b/docs/usage.md
@@ -178,6 +178,43 @@ edits settings or spends a model call is a dry run until you add `--apply`.
| Skills you have | `lifecycle`, `tier`, `check`, `reconcile`, `bundle`, `expire`, `accuracy` |
| Internal (run by the hooks) | `hook`, `fold-session`, `fold-pending` |
+## Memory
+
+The local model is the one large thing Skill++ runs. On an 18 GB Mac, a fold
+with `gemma4:e4b-it-qat` and `nomic-embed-text` takes about **7 GB** of free
+memory, whatever the session: a one-word question costs as much as a whole
+session. `gemma4:e4b`, the default before it, took 12.8 GB, and with other
+apps open a fold ran at 92 % memory used while macOS swapped out 4.7 GB in
+100 seconds. So Skill++ guards its memory use:
+
+- **It starts only if the models fit.** A fold loads the models only when
+ what they take still leaves 2 GB free: with the default models, from about
+ 9 GB free. What they take is measured on your computer: estimated from
+ their size at first, which asks for about 10 GB free, then the largest
+ amount a fold has actually used here.
+- **It stops when memory runs short.** If free memory falls below 2 GB while a
+ fold runs, or macOS reports critical memory pressure, Skill++ stops and
+ unloads the models at once.
+- **It lets go right away.** When a fold ends, the models are unloaded, rather
+ than staying in memory for Ollama's usual five minutes.
+- **It never unloads what it didn't load.** A model you had loaded yourself
+ stays loaded.
+
+A session that doesn't fit isn't lost. It waits, and is folded:
+
+- **when your computer is idle,** five minutes without keyboard or mouse input,
+ once the models fit;
+- **at the next session start,** retried at most every ten minutes;
+- **when you ask:** **Fold now** on the review page's banner, or
+ `skill-plus-plus fold-pending --now`. The models still load only if they fit.
+
+When a fold waits or is stopped, a desktop notification says so: Notification
+Center on macOS, `notify-send` on Linux. On a Mac the first one may ask you to
+allow notifications for Script Editor, which is what shows them.
+`skill-plus-plus doctor` shows what the models take on your computer and from
+what free memory a fold starts; `skill-plus-plus stats` counts the sessions
+waiting.
+
## Configuration
All settings are environment variables.
@@ -198,6 +235,10 @@ All settings are environment variables.
| `SKILL_PLUS_PLUS_MATCH` | `1` | `0` banks every task without comparing it (for measuring detection) |
| `SKILL_PLUS_PLUS_DESCRIBE` | `0` | `1` asks the model to describe every tool call, inside the hook (slow) |
| `SKILL_PLUS_PLUS_MAX_STEPS` / `SKILL_PLUS_PLUS_MAX_FIELD` | `500` / `2000` | caps per session and per captured field |
+| `SKILL_PLUS_PLUS_MEMORY_GUARD` | `1` | `0` turns the [memory guard](#memory) off: folds load the models whatever is free |
+| `SKILL_PLUS_PLUS_MEMORY_RESERVE_GB` | `2` | free memory a fold always leaves; it starts only if the models fit with this much to spare, and stops below it |
+| `SKILL_PLUS_PLUS_IDLE_MINUTES` | `5` | how long nobody must use the keyboard or mouse before a session waiting for memory is folded |
+| `SKILL_PLUS_PLUS_NOTIFY` | `1` | `0` turns off the desktop notices when a fold waits for memory or is stopped |
| `SKILL_PLUS_PLUS_INTERNAL` | unset | set to anything to make the hooks do nothing, e.g. for one session |
## Troubleshooting
@@ -225,6 +266,12 @@ running. Stop it with Ctrl-C in its terminal, or use `skill-plus-plus web --port
**Tool calls feel slow.** Check that `SKILL_PLUS_PLUS_DESCRIBE` is not set to `1`; it asks
the local model about every tool call inside the hook.
+**Sessions keep waiting for memory.** `skill-plus-plus doctor` says from what
+free memory a fold starts. Close a few apps and press **Fold now** on the
+review page, or leave the computer idle for a few minutes. If `doctor` says the
+models take more than the machine has, sessions never fold while the guard is
+on; `SKILL_PLUS_PLUS_MEMORY_GUARD=0` folds them anyway, at the risk of swapping.
+
**Leave one session out.** Start it as `SKILL_PLUS_PLUS_INTERNAL=1 claude`, and the hooks
do nothing for it.
diff --git a/skill_plus_plus/capture.py b/skill_plus_plus/capture.py
index 69a26aa..6b3b531 100644
--- a/skill_plus_plus/capture.py
+++ b/skill_plus_plus/capture.py
@@ -548,22 +548,47 @@ def handle_session_end(config: Config, payload: dict) -> dict:
# keep path already attaches before judging.
_attach_reply(session, payload.get("transcript_path") or session.get("transcript"))
- if config.judge_boundaries and "folded" not in session:
- try:
- from .boundary import judge_session
- judge_session(config, session)
- except Exception as exc: # noqa: BLE001 - a hook never raises at a dev
- log_error(config, f"boundary judge failed: {type(exc).__name__}: {exc}")
-
- # The last task's own completion report, said after its final tool call.
- # Attached before folding so the episode carries it.
- trailing = _trailing_narration(payload)
- if trailing:
- work = [s for s in session.get("steps", []) if not is_prompt(s)]
- if work:
- work[-1].setdefault("closing_note", trailing)
- _save_session(config, session)
- result = fold_session(config, session, persist=True)
+ # The memory guard (`skill_plus_plus.memory`) around everything that loads
+ # a model. The models load only if they fit and still leave the reserve
+ # free; otherwise the session is held and folded later, which costs
+ # nothing. A watchdog stops the fold and unloads them if memory runs short
+ # while it runs, and they are unloaded when it ends.
+ from . import memory
+ with memory.guarded(config) as guard:
+ short = guard.admit(memory.fold_models(config))
+ if short:
+ return _hold_for_memory(config, session, short)
+
+ if config.judge_boundaries and "folded" not in session:
+ try:
+ from .boundary import judge_session
+ judge_session(config, session)
+ except Exception as exc: # noqa: BLE001 - a hook never raises at a dev
+ log_error(config, f"boundary judge failed: {type(exc).__name__}: {exc}")
+ if guard.tripped:
+ # Stopped halfway through judging. The gaps judged so far carry
+ # verdicts and the rest carry `judge_session`'s up-front
+ # `False`, which reads as "no boundary here": banked, the
+ # session would merge tasks the judge never got to. Nothing is
+ # banked yet, so every verdict goes and the retry judges it all.
+ for step in session.get("steps", []):
+ step.pop("end", None)
+ return _hold_for_memory(config, session, guard.tripped, notify=False)
+
+ # The last task's own completion report, said after its final tool call.
+ # Attached before folding so the episode carries it.
+ trailing = _trailing_narration(payload)
+ if trailing:
+ work = [s for s in session.get("steps", []) if not is_prompt(s)]
+ if work:
+ work[-1].setdefault("closing_note", trailing)
+ _save_session(config, session)
+ result = fold_session(config, session, persist=True)
+ if guard.tripped and result.get("status") == "offline":
+ # Stopped between episodes: `folded` keeps the ones banked, and the
+ # retry resumes after them, as it does when Ollama goes away.
+ result["reason"] = f"memory: {guard.tripped}"
+ result["memory"] = True
# Offline: the session was never folded, so this file is the only copy of
# the work. Losing the step is the one thing capture exists to prevent β
# being offline costs the candidate, never the record.
@@ -576,6 +601,8 @@ def handle_session_end(config: Config, payload: dict) -> dict:
"at": datetime.now(timezone.utc).isoformat(timespec="seconds"),
"reason": result.get("reason", ""),
}
+ if result.get("memory"):
+ session["held"]["memory"] = True
_save_session(config, session)
else:
try:
@@ -585,6 +612,29 @@ def handle_session_end(config: Config, payload: dict) -> dict:
return result
+def _hold_for_memory(config: Config, session: dict, reason: str, *,
+ notify: bool = True) -> dict:
+ """Keep the session to fold later: the models do not fit in memory now.
+
+ Held like a session no model answered, so `fold_pending` retries it and
+ `stats` counts it, with `memory` set so the retry waits for memory rather
+ than for Ollama. *notify* is off when the watchdog stopped the fold, which
+ has already said so.
+ """
+ session["held"] = {
+ "at": datetime.now(timezone.utc).isoformat(timespec="seconds"),
+ "reason": f"memory: {reason}",
+ "memory": True,
+ }
+ _save_session(config, session)
+ if notify:
+ from .notify import waiting
+ waiting(config, reason)
+ return {"status": "offline", "steps": len(session.get("steps", [])),
+ "episodes": [], "flagged": 0, "reason": session["held"]["reason"],
+ "memory": True}
+
+
def mark_ending(config: Config, session_id: str, transcript: str | None = None) -> dict:
"""Record that a session ended. The fold itself happens elsewhere.
@@ -648,6 +698,19 @@ def fold_session_now(config: Config, session_id: str, *,
# A ceiling on one session's fold. Bounds the damage a recycled pid can do:
# past this the lock is stale whatever `os.kill` says.
_FOLD_LOCK_SECONDS = 1800
+# How often the sweep retries a session held for memory (see `_is_pending`).
+MEMORY_RETRY_SECONDS = 600.0
+
+
+def _age_seconds(stamp) -> float:
+ """Seconds since an ISO stamp. An unreadable one counts as long ago."""
+ try:
+ at = datetime.fromisoformat(str(stamp))
+ except ValueError:
+ return float("inf")
+ if at.tzinfo is None:
+ at = at.replace(tzinfo=timezone.utc)
+ return (datetime.now(timezone.utc) - at).total_seconds()
def _lock_file(config: Config, session_id: str) -> Path:
@@ -751,7 +814,8 @@ def _find_transcript(session_id: str) -> str | None:
def fold_pending(config: Config, *, exclude: str | None = None,
- idle_hours: float = PENDING_IDLE_HOURS) -> list[dict]:
+ idle_hours: float = PENDING_IDLE_HOURS,
+ now: bool = False) -> list[dict]:
"""Bank the sessions that ended without being banked.
`SessionEnd` stamps `ending` and spawns a worker, so this sweep is the
@@ -762,48 +826,115 @@ def fold_pending(config: Config, *, exclude: str | None = None,
per-session lock and the held stamp are the ones a normal end uses, and
`folded` keeps a partial retry from counting an episode twice. *exclude* is
the session that is starting, which is live by definition.
+
+ One memory guard covers the sweep, so the models load once for every
+ session in it and are unloaded at the end. The first session held for
+ memory ends the sweep: what did not fit for it will not fit for the next.
+ *now* skips the pause between memory retries, for the idle waiter and
+ "Fold now", which have just checked that the models fit.
"""
+ from . import memory
+
results = []
with _locked(config.root / "fold-pending.lock", _PENDING_LOCK_SECONDS,
"fold-pending") as got:
if not got:
return [{"status": "locked"}]
- for path in sorted(config.sessions_dir.glob("*.json")):
- sid = path.stem
- if sid == exclude:
- continue
- try:
- session = json.loads(path.read_text(encoding="utf-8"))
- idle = (time.time() - path.stat().st_mtime) / 3600
- except (OSError, json.JSONDecodeError):
- continue
- if not _is_pending(session, idle, idle_hours):
- results.append({"session": sid, "status": "live"})
- continue
- # Checked before folding only so the outcome can say "folding"
- # rather than "locked"; `fold_session_now` takes the lock itself,
- # so nothing rests on this being race-free.
- if _lock_alive(_lock_file(config, sid), _FOLD_LOCK_SECONDS):
- results.append({"session": sid, "status": "folding"})
- continue
- results.append({"session": sid, **fold_session_now(config, sid)})
+ with memory.guarded(config):
+ for path in sorted(config.sessions_dir.glob("*.json")):
+ sid = path.stem
+ if sid == exclude:
+ continue
+ try:
+ session = json.loads(path.read_text(encoding="utf-8"))
+ idle = (time.time() - path.stat().st_mtime) / 3600
+ except (OSError, json.JSONDecodeError):
+ continue
+ if not _is_pending(session, idle, idle_hours, now=now):
+ results.append({"session": sid, "status": "live"})
+ continue
+ # Checked before folding only so the outcome can say "folding"
+ # rather than "locked"; `fold_session_now` takes the lock itself,
+ # so nothing rests on this being race-free.
+ if _lock_alive(_lock_file(config, sid), _FOLD_LOCK_SECONDS):
+ results.append({"session": sid, "status": "folding"})
+ continue
+ result = fold_session_now(config, sid)
+ results.append({"session": sid, **result})
+ if result.get("memory"):
+ break
_sweep_orphan_locks(config)
return results
-def _is_pending(session: dict, idle: float, idle_hours: float) -> bool:
+# How long a worker that held a session for memory waits for a better moment,
+# and how often it looks. Past this, the next session start takes over.
+MEMORY_WAIT_SECONDS = 12 * 3600
+MEMORY_POLL_SECONDS = 60.0
+
+
+def _held_for_memory(config: Config) -> bool:
+ for path in config.sessions_dir.glob("*.json"):
+ try:
+ if (json.loads(path.read_text(encoding="utf-8")).get("held") or {}).get("memory"):
+ return True
+ except (OSError, json.JSONDecodeError, AttributeError):
+ continue
+ return False
+
+
+def wait_for_memory(config: Config, *, sleep=time.sleep, clock=time.time) -> list[dict]:
+ """Fold the sessions held for memory once the computer is idle and they fit.
+
+ Run by the worker that held a session, after its own lock is released:
+ nobody waits on it, and it costs a sleeping process. One waiter at a time;
+ another worker that finds the lock taken leaves it to the one waiting.
+ Idle means nobody at the keyboard for `config.idle_minutes`, so the model
+ never competes with someone working. Where idle time cannot be read, the
+ waiter waits for memory alone.
+ """
+ from . import memory
+
+ with _locked(config.root / "memory-wait.lock", MEMORY_WAIT_SECONDS + 600,
+ "memory-wait") as got:
+ if not got:
+ return []
+ deadline = clock() + MEMORY_WAIT_SECONDS
+ while clock() < deadline:
+ sleep(MEMORY_POLL_SECONDS)
+ if not _held_for_memory(config):
+ return []
+ idle = memory.idle_seconds()
+ if idle is not None and idle < config.idle_minutes * 60:
+ continue
+ if memory.shortfall(config):
+ continue
+ results = fold_pending(config, now=True)
+ if not _held_for_memory(config):
+ return results
+ return []
+
+
+def _is_pending(session: dict, idle: float, idle_hours: float, *,
+ now: bool = False) -> bool:
"""Has this session ended without being banked?
Three rules, covering three different failures, none of them redundant:
- * ``held`` β a fold that ran and could not reach a model.
+ * ``held`` β a fold that ran and could not reach a model, or was held for
+ memory. Those are retried at most every `MEMORY_RETRY_SECONDS`, unless
+ *now*: a retry that passes the check and trips the watchdog again would
+ load and unload the models at every session start.
* ``ending`` past the grace β a fold that was launched and died with it.
Without the stamp it would look like a live session and wait out the idle
rule.
* idle β the only rule that catches a session where `SessionEnd` never
fired at all: a window closed, a laptop shut down.
"""
- if session.get("held"):
+ held = session.get("held")
+ if held:
+ if isinstance(held, dict) and held.get("memory") and not now:
+ return _age_seconds(held.get("at")) >= MEMORY_RETRY_SECONDS
return True
ending = session.get("ending")
if ending:
diff --git a/skill_plus_plus/cli.py b/skill_plus_plus/cli.py
index bb19439..206835b 100644
--- a/skill_plus_plus/cli.py
+++ b/skill_plus_plus/cli.py
@@ -555,6 +555,32 @@ def cmd_split(args: argparse.Namespace) -> int:
return 0
+def _memory_guarded(command):
+ """Run a command that calls the local models under the memory guard.
+
+ The same rule a fold follows (`skill_plus_plus.memory`): the models load
+ only if they fit and leave the reserve free, and are unloaded when the
+ command ends.
+ """
+ def run(args: argparse.Namespace) -> int:
+ from .memory import guarded
+ with guarded(Config(args.root)):
+ return command(args)
+ run.__doc__ = command.__doc__
+ return run
+
+
+def _model_trouble(exc: Exception) -> str:
+ from .local import MemoryShort
+ from .memory import BUSY
+ if isinstance(exc, MemoryShort):
+ if str(exc) == BUSY:
+ return f"Skill++ {exc}"
+ return f"not enough free memory: Skill++ {exc}"
+ return f"no local model: {exc}"
+
+
+@_memory_guarded
def cmd_merge(args: argparse.Namespace) -> int:
"""Fold existing candidates that are the same procedure.
@@ -590,7 +616,8 @@ def cmd_merge(args: argparse.Namespace) -> int:
try:
vectors = {e.id: vector_for(e, config, cache) for e in entries}
except LocalModelUnavailable as exc:
- print(f"No embedding model: {exc}. Nothing was compared.", file=sys.stderr)
+ why = _model_trouble(exc)
+ print(f"{why[:1].upper()}{why[1:]}. Nothing was compared.", file=sys.stderr)
return 1
finally:
save_cache(config, cache)
@@ -690,6 +717,15 @@ def cmd_fold_session(args: argparse.Namespace) -> int:
+ (f" {result['id']}" if result.get("id") else ""))
if args.verbose:
print(json.dumps(result))
+ if result.get("memory"):
+ # Held because the models did not fit. This worker is detached and
+ # nobody waits on it, so it stays to fold the held sessions once the
+ # computer is idle and they fit (`capture.wait_for_memory`).
+ from .capture import wait_for_memory
+ try:
+ wait_for_memory(config)
+ except Exception as exc: # noqa: BLE001 - detached; a traceback goes nowhere
+ log_error(config, f"memory wait failed: {type(exc).__name__}: {exc}")
return 0
@@ -716,7 +752,8 @@ def cmd_fold_pending(args: argparse.Namespace) -> int:
config = Config(args.root)
config.ensure_dirs()
- results = fold_pending(config, exclude=args.exclude, idle_hours=args.idle_hours)
+ results = fold_pending(config, exclude=args.exclude, idle_hours=args.idle_hours,
+ now=args.now)
if results == [{"status": "locked"}]:
print("Another fold-pending is running.")
return 0
@@ -726,6 +763,8 @@ def cmd_fold_pending(args: argparse.Namespace) -> int:
what = "still live, skipped"
elif r["status"] == "folding":
what = "a worker is folding it"
+ elif r.get("memory"):
+ what = f"waiting for memory: {r.get('reason', '').removeprefix('memory: ')}"
elif r["status"] == "offline":
what = f"held again: {r.get('reason', '')}"
elif episodes:
@@ -881,6 +920,7 @@ def cmd_reopen(args: argparse.Namespace) -> int:
return 1
+@_memory_guarded
def cmd_retitle(args: argparse.Namespace) -> int:
"""Name candidates the fold could not name, because no model answered.
@@ -911,7 +951,7 @@ def cmd_retitle(args: argparse.Namespace) -> int:
try:
name, sentence = name_and_sentence(config, entry)
except LocalModelUnavailable as exc:
- print(f"no local model: {exc}")
+ print(_model_trouble(exc))
return 1
if sentence:
store_summary(config, entry, sentence)
@@ -1240,6 +1280,26 @@ def cmd_doctor(args: argparse.Namespace) -> int:
here = _has_model(models, want)
print(f" {'':8} {want} {'β' if here else 'β not pulled'}")
+ # What the models take, from what free memory a fold starts, and whether
+ # this machine can ever get there (`skill_plus_plus.memory`).
+ from .memory import gb, status
+ m = status(config)
+ if m["available"] is None:
+ print(f"memory {'':8} unknown on this system: the memory guard is off")
+ elif not config.memory_guard:
+ print(f"memory {'':8} guard off (SKILL_PLUS_PLUS_MEMORY_GUARD=0); "
+ f"{gb(m['available'])} GB free of {gb(m['total'])} GB")
+ else:
+ print(f"memory {'':8} {gb(m['total'])} GB RAM, {gb(m['available'])} GB free now")
+ print(f" {'':8} the models take about {gb(m['need'])} GB here "
+ f"({'measured' if m['measured'] else 'estimated from their size'}); "
+ f"a fold starts at {gb(m['start_at'])} GB free")
+ if m["total"] is not None and m["start_at"] > m["total"]:
+ print(f" {'':8} more than this machine has: sessions wait and "
+ "never fold while the guard is on")
+ print(f" {'':8} SKILL_PLUS_PLUS_MEMORY_GUARD=0 folds them anyway, "
+ "at the risk of swapping")
+
s = _waiting_sessions(config)
pending = len(s["held"]) + len(s["waiting"])
if pending or s["folding"]:
@@ -1279,18 +1339,27 @@ def cmd_stats(args: argparse.Namespace) -> int:
# Only sessions `handle_session_end` stamped as held. A session still being
# written has a file too, and counting it would report work lost from one
# that is merely in flight.
- held = []
+ held, for_memory = [], []
for path in sorted(config.sessions_dir.glob("*.json")):
try:
- if json.loads(path.read_text(encoding="utf-8")).get("held"):
- held.append(path)
+ stamp = json.loads(path.read_text(encoding="utf-8")).get("held")
except (OSError, json.JSONDecodeError):
continue
+ if stamp:
+ (for_memory if isinstance(stamp, dict) and stamp.get("memory") else held).append(path)
if held:
print(f"held {len(held)} session(s) not banked, kept in "
f"{config.sessions_dir}")
print(f" no local model answered ({config.local_model} at "
f"{config.ollama_url}); the work is there, the candidates are not")
+ if for_memory:
+ from .memory import gb, status
+ s = status(config)
+ print(f"waiting {len(for_memory)} session(s) held for memory, kept in "
+ f"{config.sessions_dir}")
+ print(f" a fold starts when {gb(s['start_at'])} GB are free "
+ f"({gb(s['available'])} GB now); they fold once the computer is idle,")
+ print(" or run `skill-plus-plus fold-pending --now`")
return 0
@@ -1734,6 +1803,9 @@ def build_parser() -> argparse.ArgumentParser:
p.add_argument("--exclude", help="a live session to leave alone")
p.add_argument("--idle-hours", type=float, default=12.0,
help="treat a session untouched this long as ended")
+ p.add_argument("--now", action="store_true",
+ help="retry sessions held for memory now, not after the pause "
+ "(the models still load only if they fit)")
p.set_defaults(func=cmd_fold_pending)
p = sub.add_parser("web", help="browse the ledger in a local page")
diff --git a/skill_plus_plus/config.py b/skill_plus_plus/config.py
index 34c9cd6..b2b4e65 100644
--- a/skill_plus_plus/config.py
+++ b/skill_plus_plus/config.py
@@ -134,6 +134,28 @@ def __init__(self, root: str | Path | None = None) -> None:
# A boundary that would close an episode smaller than this is ignored:
# one step is not a workflow.
self.min_episode_steps = _int_env("SKILL_PLUS_PLUS_MIN_EPISODE_STEPS", 2)
+ # The memory guard (`skill_plus_plus.memory`). The local model is not
+ # small: on an 18 GB Mac, loading gemma4:e4b and the embedder took 12.8
+ # GB of available memory (gemma4:e4b-it-qat, the default after it, about
+ # 7), and a fold with other apps open ran at 92 % used with 4.7 GB
+ # swapped out in 100 s. So a fold starts only when the
+ # models fit with this much memory still free, and stops, unloading
+ # them, the moment less than this is left. An absolute figure, not a
+ # share of RAM: what keeps a machine out of swap is headroom in
+ # gigabytes. 2 GB is where that Mac turned: 2.0 GB available while
+ # loading, pressure normal and nothing swapped; 1.4-1.6 GB while
+ # folding, pressure warning and swapping. `SKILL_PLUS_PLUS_MEMORY_GUARD=0`
+ # turns the guard off.
+ self.memory_guard = _bool_env("SKILL_PLUS_PLUS_MEMORY_GUARD", True)
+ self.memory_reserve_gb = _float_env("SKILL_PLUS_PLUS_MEMORY_RESERVE_GB", 2.0)
+ # A session held for memory is folded once the computer has been idle
+ # this long and the models fit, so the model never competes with
+ # someone at the keyboard. Folding later costs nothing.
+ self.idle_minutes = _float_env("SKILL_PLUS_PLUS_IDLE_MINUTES", 5.0)
+ # Desktop notifications when the guard holds or stops a fold. The fold
+ # runs after the chat has ended, so the operating system is the one
+ # place the person will see it. `SKILL_PLUS_PLUS_NOTIFY=0` turns them off.
+ self.notify = _bool_env("SKILL_PLUS_PLUS_NOTIFY", True)
@property
def decisions_file(self) -> Path:
@@ -168,6 +190,11 @@ def archive_dir(self) -> Path:
def log_file(self) -> Path:
return self.root / "skill-plus-plus.log"
+ @property
+ def memory_file(self) -> Path:
+ """What the models took on this computer, and when a notice was last sent."""
+ return self.root / "memory.json"
+
def ensure_dirs(self) -> None:
for d in (self.ledger_dir, self.sessions_dir, self.cold_dir, self.archive_dir):
d.mkdir(parents=True, exist_ok=True)
diff --git a/skill_plus_plus/local.py b/skill_plus_plus/local.py
index 171a0a4..b7f39d8 100644
--- a/skill_plus_plus/local.py
+++ b/skill_plus_plus/local.py
@@ -43,6 +43,31 @@ class LocalModelUnavailable(RuntimeError):
"""Ollama is not reachable, or the model is not installed."""
+class MemoryShort(LocalModelUnavailable):
+ """The memory guard kept a model from loading, or stopped the fold.
+
+ A subclass, so every caller that already treats an unreachable model as
+ "no opinion" and holds the session does the same here, with no new path.
+ """
+
+
+def _guard(model: str, payload: dict) -> None:
+ """Ask the open memory guard, if any, before *model* is called.
+
+ Raises `MemoryShort` when the guard refuses. A model the guard loaded is
+ asked to leave memory soon after its last call (`memory.KEEP_ALIVE`), so a
+ fold that dies before its guard is released cannot hold it for Ollama's
+ five minutes. A model someone else loaded keeps its own `keep_alive`.
+ """
+ from .memory import KEEP_ALIVE, active
+ guard = active()
+ if guard is None:
+ return
+ guard.before_call(model)
+ if guard.owns(model):
+ payload["keep_alive"] = KEEP_ALIVE
+
+
def _num_ctx(prompt: str, reserve: int = 512) -> int:
want = (len(prompt) // _CHARS_PER_TOKEN) + reserve
return max(_CTX_FLOOR, min(_CTX_CEILING, want))
@@ -74,6 +99,7 @@ def ask(model: str, prompt: str, *, host: str = DEFAULT_HOST,
}
if think is not None:
payload["think"] = think
+ _guard(model, payload)
body = json.dumps(payload).encode("utf-8")
request = urllib.request.Request(
f"{host.rstrip('/')}/api/generate", data=body,
@@ -146,8 +172,9 @@ def embed(text: str, *, model: str = DEFAULT_EMBED_MODEL,
reached, so the caller can say so instead of matching on a fragment
silently.
"""
- body = json.dumps({"model": model, "input": text,
- "truncate": True}).encode("utf-8")
+ payload = {"model": model, "input": text, "truncate": True}
+ _guard(model, payload)
+ body = json.dumps(payload).encode("utf-8")
request = urllib.request.Request(
f"{host.rstrip('/')}/api/embed", data=body,
headers={"Content-Type": "application/json"})
diff --git a/skill_plus_plus/memory.py b/skill_plus_plus/memory.py
new file mode 100644
index 0000000..1cfcad6
--- /dev/null
+++ b/skill_plus_plus/memory.py
@@ -0,0 +1,568 @@
+"""Keep the local model from pushing the machine into swap.
+
+The local model is the one large thing Skill++ runs. Measured on an 18 GB Mac
+(docs/usage.md, "Memory"): loading gemma4:e4b and nomic-embed-text took 12.8
+GB of available memory (a fold with gemma4:e4b-it-qat, the default after it,
+about 7), and a fold with other apps open sat at 92 % used, macOS pressure at
+warning, with 4.7 GB swapped out in 100 seconds. Nothing leaked: a
+one-word question locks as much as a whole session. So the fix is not in how a
+session is read but in when the models may load, and how long they stay.
+
+Three rules, in the order they act:
+
+* **Before loading.** A model that is not loaded yet loads only if what the
+ models need still leaves `config.memory_reserve_gb` free. Otherwise nothing
+ loads and the caller holds the session. Folding later costs nothing.
+* **While folding.** A watchdog thread reads available memory twice a second.
+ Below the reserve, or at critical pressure, it trips: no further model call
+ starts, and the models this guard loaded are unloaded at once.
+* **After folding.** Whatever this guard loaded is unloaded when the guard is
+ released, not five minutes later when Ollama's own `keep_alive` would.
+
+What the models need is measured on each computer: estimated from their size
+on disk the first time, then the largest drop in available memory a fold has
+seen here. Only models this guard loaded are ever unloaded; a model someone else
+had loaded stays loaded.
+
+Every reader answers `None` when it cannot tell, and an unknown reading turns
+the guard off rather than holding every session forever.
+"""
+
+from __future__ import annotations
+
+import json
+import os
+import re
+import socket
+import subprocess
+import sys
+import threading
+import time
+import urllib.error
+import urllib.request
+from collections.abc import Iterable, Iterator
+from contextlib import contextmanager
+from pathlib import Path
+
+GB = 1024 ** 3
+
+# Size on disk times this is the first guess at what loading takes. Measured in
+# bytes on the calibration run: 9.88e9 on disk (gemma4:e4b and nomic-embed-text,
+# as Ollama lists them) took 13.7e9 of available memory, which is 12.8 GiB: the
+# weights plus the buffers and cache that loading brings along. It guesses high
+# for gemma4:e4b-it-qat: 6.42e9 on disk, 8.3 GiB guessed, 7.0 GiB measured. The
+# guess only decides the first fold of a model set on a computer; from then on
+# the measured figure does (`need_bytes`).
+NEED_FACTOR = 1.39
+# How long `unload` waits for Ollama to stop listing a model. It answers the
+# unload at once and lets go a moment later; a check made in between counted a
+# model that was leaving as loaded by someone else, and never unloaded it again.
+_UNLOAD_WAIT_SECONDS = 15.0
+# After the watchdog trips, how long it keeps watching for a model that was
+# still loading. Measured: a trip 4.3 s into a load found nothing listed to
+# unload, and the load finished anyway, 9.3 GB that stayed for `KEEP_ALIVE`.
+# A load of gemma4:e4b takes 10-14 s, so 20 s sees it land and unloads it.
+_LINGER_SECONDS = 20.0
+# A drop larger than this multiple of the estimate is something else growing at
+# the same time, a browser or a build, not the models. It is not stored.
+_OUTLIER = 2.0
+# How often the watchdog reads memory. A load moves gigabytes in seconds.
+WATCH_SECONDS = 0.5
+# macOS `kern.memorystatus_vm_pressure_level`: 1 normal, 2 warning, 4 critical.
+_CRITICAL = 4
+# Linux PSI: every task stalled on memory for this share of the last ten
+# seconds is a machine that is thrashing.
+_PSI_FULL_CRITICAL = 10.0
+# How long a model this guard loaded stays after its last call if the process
+# dies before `release()`. Ollama's own default is five minutes.
+KEEP_ALIVE = "60s"
+# How long after an unload the freed memory is read, for the notice.
+_SETTLE_SECONDS = 1.0
+
+
+# -- readers ----------------------------------------------------------------
+
+def _run(argv: list[str], timeout: float = 5.0) -> str | None:
+ try:
+ return subprocess.run(argv, capture_output=True, text=True,
+ timeout=timeout, check=False).stdout
+ except (OSError, subprocess.SubprocessError):
+ return None
+
+
+def _sysctl(name: str) -> int | None:
+ out = _run(["sysctl", "-n", name])
+ try:
+ return int(out.strip()) if out else None
+ except ValueError:
+ return None
+
+
+def _meminfo() -> dict[str, int]:
+ try:
+ text = Path("/proc/meminfo").read_text(encoding="utf-8")
+ except OSError:
+ return {}
+ out = {}
+ for line in text.splitlines():
+ m = re.match(r"(\w+):\s+(\d+)\s*kB", line)
+ if m:
+ out[m.group(1)] = int(m.group(2)) * 1024
+ return out
+
+
+def total_bytes() -> int | None:
+ if sys.platform == "darwin":
+ return _sysctl("hw.memsize")
+ if sys.platform.startswith("linux"):
+ return _meminfo().get("MemTotal")
+ return None
+
+
+def available_bytes() -> int | None:
+ """Memory the system can hand out without swapping. Not "free".
+
+ Free memory sits near zero on a Mac by design, because spare RAM holds
+ cache, and a check against it would never let a fold start: 0.3 GB free
+ while the kernel reported 81 % available. macOS says what it can hand out as
+ a percentage (`kern.memorystatus_level`); Linux as `MemAvailable`.
+ """
+ if sys.platform == "darwin":
+ level, total = _sysctl("kern.memorystatus_level"), _sysctl("hw.memsize")
+ return total * level // 100 if level is not None and total else None
+ if sys.platform.startswith("linux"):
+ return _meminfo().get("MemAvailable")
+ return None
+
+
+def pressure_critical() -> bool:
+ """Has the system itself declared a memory emergency?"""
+ if sys.platform == "darwin":
+ return _sysctl("kern.memorystatus_vm_pressure_level") == _CRITICAL
+ if sys.platform.startswith("linux"):
+ try:
+ text = Path("/proc/pressure/memory").read_text(encoding="utf-8")
+ except OSError:
+ return False
+ m = re.search(r"^full avg10=([\d.]+)", text, re.M)
+ return bool(m) and float(m.group(1)) >= _PSI_FULL_CRITICAL
+ return False
+
+
+def idle_seconds() -> float | None:
+ """How long nobody has used the keyboard or mouse, or None if unknown."""
+ if sys.platform == "darwin":
+ return mac_idle(_run(["ioreg", "-c", "IOHIDSystem", "-d", "4"]))
+ if sys.platform.startswith("linux"):
+ session = os.environ.get("XDG_SESSION_ID")
+ if session:
+ idle = loginctl_idle(_run(["loginctl", "show-session", session,
+ "-p", "IdleHint", "-p", "IdleSinceHint"]))
+ if idle is not None:
+ return idle
+ return xprintidle_idle(_run(["xprintidle"]))
+ return None
+
+
+def mac_idle(text: str | None) -> float | None:
+ """`ioreg`'s `HIDIdleTime`, in nanoseconds, as seconds."""
+ m = re.search(r'"HIDIdleTime"\s*=\s*(\d+)', text or "")
+ return int(m.group(1)) / 1e9 if m else None
+
+
+def loginctl_idle(text: str | None, now: float | None = None) -> float | None:
+ """Seconds idle from `loginctl show-session`, or None if it does not say.
+
+ `IdleSinceHint` is wall-clock microseconds. A desktop that never sets the
+ hint reports `IdleHint=no` forever; that reads as "not idle", which only
+ delays a fold, never forces one.
+ """
+ values = dict(line.split("=", 1) for line in (text or "").splitlines() if "=" in line)
+ if "IdleHint" not in values:
+ return None
+ if values["IdleHint"].strip() != "yes":
+ return 0.0
+ try:
+ since = int(values.get("IdleSinceHint", "0"))
+ except ValueError:
+ return None
+ if not since:
+ return None
+ return max(0.0, (time.time() if now is None else now) - since / 1e6)
+
+
+def xprintidle_idle(text: str | None) -> float | None:
+ """`xprintidle` prints milliseconds."""
+ try:
+ return int((text or "").strip()) / 1000
+ except ValueError:
+ return None
+
+
+# -- Ollama -----------------------------------------------------------------
+
+def _ollama(config, path: str, payload: dict | None = None,
+ timeout: float = 10.0) -> dict | None:
+ url = f"{config.ollama_url.rstrip('/')}{path}"
+ data = json.dumps(payload).encode("utf-8") if payload is not None else None
+ request = urllib.request.Request(url, data=data,
+ headers={"Content-Type": "application/json"})
+ try:
+ with urllib.request.urlopen(request, timeout=timeout) as response:
+ return json.loads(response.read().decode("utf-8"))
+ except (urllib.error.URLError, OSError, TimeoutError, ValueError):
+ return None
+
+
+def model_name(name: str) -> str:
+ """`nomic-embed-text` and `nomic-embed-text:latest` are one model.
+
+ Ollama lists the tag; the configuration usually leaves it off.
+ """
+ name = str(name or "").strip()
+ return name if ":" in name.rsplit("/", 1)[-1] else f"{name}:latest"
+
+
+def loaded_models(config) -> set[str] | None:
+ """What Ollama holds in memory now, or None if it cannot be asked."""
+ reply = _ollama(config, "/api/ps")
+ if reply is None:
+ return None
+ return {model_name(m.get("name") or m.get("model") or "")
+ for m in reply.get("models") or []}
+
+
+def model_sizes(config) -> dict[str, int]:
+ """Each installed model's size on disk, in bytes."""
+ reply = _ollama(config, "/api/tags") or {}
+ return {model_name(m.get("name") or m.get("model") or ""): int(m.get("size") or 0)
+ for m in reply.get("models") or []}
+
+
+def unload(config, models: Iterable[str], *, linger: float = 0.0) -> None:
+ """Drop *models* from memory, and return once Ollama has let them go.
+
+ The documented way: `keep_alive` 0. `/api/generate` does it for an embedding
+ model too; it answers with `done_reason: "unload"` and generates nothing.
+ It answers before the model is gone, though, so this waits until `/api/ps`
+ stops listing it (see `_UNLOAD_WAIT_SECONDS`).
+
+ *linger* keeps watching that long even once nothing is listed, and unloads
+ again whatever appears: a model still loading is not listed yet, ignores
+ the unload, and lands a few seconds later.
+ """
+ names = {model_name(m) for m in models}
+ if not names:
+ return
+
+ def ask(which):
+ for name in sorted(which):
+ _ollama(config, "/api/generate", {"model": name, "keep_alive": 0}, timeout=30.0)
+
+ ask(names)
+ asked = time.time()
+ settle = asked + linger
+ deadline = settle + _UNLOAD_WAIT_SECONDS
+ while time.time() < deadline:
+ loaded = loaded_models(config)
+ if loaded is None:
+ return
+ listed = names & loaded
+ if not listed and time.time() >= settle:
+ return
+ if listed and time.time() - asked >= 2.0:
+ ask(listed)
+ asked = time.time()
+ time.sleep(0.25)
+
+
+def fold_models(config) -> list[str]:
+ """The models a fold may load: the judge and namer, and the embedder."""
+ models = []
+ if config.judge_boundaries or config.name_candidates:
+ models.append(config.local_model)
+ if config.match_candidates:
+ models.append(config.embed_model)
+ return models
+
+
+# -- what the models take on this computer -----------------------------------
+
+def _load(config) -> dict:
+ try:
+ data = json.loads(config.memory_file.read_text(encoding="utf-8"))
+ except (OSError, json.JSONDecodeError):
+ return {}
+ return data if isinstance(data, dict) else {}
+
+
+def _store(config, data: dict) -> None:
+ try:
+ config.root.mkdir(parents=True, exist_ok=True)
+ tmp = config.memory_file.with_suffix(".json.tmp")
+ tmp.write_text(json.dumps(data, indent=2, sort_keys=True), encoding="utf-8")
+ tmp.replace(config.memory_file)
+ except OSError:
+ pass
+
+
+def _key(models: Iterable[str]) -> str:
+ return "+".join(sorted(model_name(m) for m in models))
+
+
+def need_bytes(config, models: Iterable[str],
+ sizes: dict[str, int] | None = None) -> tuple[int, bool]:
+ """What loading *models* takes on this computer, and whether it was measured."""
+ models = [model_name(m) for m in models]
+ if not models:
+ return 0, True
+ stored = (_load(config).get("need", {}).get(socket.gethostname(), {})
+ .get(_key(models)))
+ if isinstance(stored, int) and stored > 0:
+ return stored, True
+ sizes = model_sizes(config) if sizes is None else sizes
+ return int(sum(sizes.get(m, 0) for m in models) * NEED_FACTOR), False
+
+
+def record_need(config, models: Iterable[str], drop: int, estimate: int) -> None:
+ """Keep the largest drop a load of *models* caused here, within reason."""
+ models = list(models)
+ if not models or drop <= 0 or estimate <= 0 or drop > estimate * _OUTLIER:
+ return
+ data = _load(config)
+ here = data.setdefault("need", {}).setdefault(socket.gethostname(), {})
+ key = _key(models)
+ here[key] = max(int(here.get(key) or 0), int(drop))
+ _store(config, data)
+
+
+def gb(n: int | None) -> str:
+ return "?" if n is None else f"{n / GB:.1f}"
+
+
+def shortfall(config, models: Iterable[str] | None = None) -> str:
+ """Why a fold could not load its models now, or "" if it could.
+
+ A dry run of the first rule for the idle waiter and "Fold now": it loads
+ nothing and starts no watchdog.
+ """
+ return Guard(config).check(fold_models(config) if models is None else models)
+
+
+def status(config) -> dict:
+ """RAM, free memory, what a fold needs here and the free memory it starts at."""
+ need, measured = need_bytes(config, fold_models(config))
+ reserve = int(config.memory_reserve_gb * GB)
+ return {"total": total_bytes(), "available": available_bytes(), "need": need,
+ "measured": measured, "reserve": reserve, "start_at": need + reserve}
+
+
+# -- the guard --------------------------------------------------------------
+
+class Guard:
+ """One fold's, sweep's or page request's use of the local models."""
+
+ def __init__(self, config, *, unload_on_release: bool = True) -> None:
+ self.config = config
+ self.reserve = int(config.memory_reserve_gb * GB)
+ self.unload_on_release = unload_on_release
+ self.enabled = bool(config.memory_guard) and available_bytes() is not None
+ self.ours: set[str] = set() # loaded by this guard, still loaded
+ self.approved: set[str] = set() # may be called without asking again
+ self.tripped = ""
+ self._loaded: set[str] = set() # everything this guard loaded
+ self._before: int | None = None
+ self._low: int | None = None
+ self._estimate = 0
+ self._stop = threading.Event()
+ self._thread: threading.Thread | None = None
+ self._lock = threading.Lock()
+
+ def _missing(self, models: Iterable[str]) -> tuple[set[str], set[str]] | None:
+ """The requested models not yet approved, and those of them not loaded."""
+ wanted = {model_name(m) for m in models if m} - self.approved
+ if not wanted:
+ return set(), set()
+ loaded = loaded_models(self.config)
+ if loaded is None:
+ # Ollama cannot be asked, so the call will fail on its own and the
+ # session is held as offline. Nothing here can make that worse.
+ return None
+ return wanted, wanted - loaded
+
+ def check(self, models: Iterable[str]) -> str:
+ """Why *models* may not load now, or "". Changes nothing."""
+ if not self.enabled:
+ return ""
+ found = self._missing(models)
+ if not found or not found[1]:
+ return ""
+ return self._room(found[1], model_sizes(self.config))[0]
+
+ def _room(self, missing: set[str], sizes: dict[str, int]) -> tuple[str, int | None]:
+ need, _ = need_bytes(self.config, missing, sizes)
+ available = available_bytes()
+ if available is None or available - need >= self.reserve:
+ return "", available
+ return (f"needs {gb(need + self.reserve)} GB of free memory, "
+ f"{gb(available)} GB free now"), available
+
+ def admit(self, models: Iterable[str]) -> str:
+ """Approve *models* for this guard, or say why not. Starts the watchdog."""
+ if not self.enabled:
+ return ""
+ if self.tripped:
+ return self.tripped
+ found = self._missing(models)
+ if found is None:
+ return ""
+ wanted, missing = found
+ if missing:
+ sizes = model_sizes(self.config)
+ reason, available = self._room(missing, sizes)
+ if reason:
+ return reason
+ with self._lock:
+ if self._before is None and available is not None:
+ self._before = self._low = available
+ self._estimate += int(sum(sizes.get(m, 0) for m in missing) * NEED_FACTOR)
+ self.ours |= missing
+ self._loaded |= missing
+ self.approved |= wanted
+ self._watch()
+ return ""
+
+ def before_call(self, model: str) -> None:
+ """Raise `MemoryShort` unless *model* may be called now."""
+ from .local import MemoryShort
+ reason = self.tripped or self.admit([model])
+ if reason:
+ raise MemoryShort(reason)
+
+ def owns(self, model: str) -> bool:
+ return model_name(model) in self.ours
+
+ def _watch(self) -> None:
+ if self._thread is not None or not self.ours:
+ return
+ self._thread = threading.Thread(target=self._run, daemon=True,
+ name="skill-plus-plus-memory")
+ self._thread.start()
+
+ def _run(self) -> None:
+ while not self._stop.wait(WATCH_SECONDS):
+ available = available_bytes()
+ if available is None:
+ continue
+ with self._lock:
+ self._low = available if self._low is None else min(self._low, available)
+ if available < self.reserve:
+ self.trip(f"free memory fell below {gb(self.reserve)} GB")
+ return
+ if pressure_critical():
+ self.trip("memory pressure turned critical")
+ return
+
+ def trip(self, reason: str) -> None:
+ """Stop: no further model call, and unload what this guard loaded, now.
+
+ Everything it loaded, including a model still loading when this
+ tripped: that one is not listed yet, so the unload lingers until it
+ lands (`_LINGER_SECONDS`).
+ """
+ with self._lock:
+ if self.tripped:
+ return
+ self.tripped = reason
+ loaded = sorted(self._loaded)
+ before = available_bytes()
+ unload(self.config, loaded, linger=_LINGER_SECONDS if loaded else 0.0)
+ time.sleep(_SETTLE_SECONDS)
+ after = available_bytes()
+ freed = after - before if before is not None and after is not None else None
+ from .notify import stopped
+ stopped(self.config, reason, freed if freed and freed > 0 else None)
+
+ def release(self) -> None:
+ """End of the fold: remember what the load took, and unload our models.
+
+ A trip in progress finishes first, so a worker cannot exit while the
+ watchdog is still unloading.
+ """
+ self._stop.set()
+ if self._thread is not None:
+ self._thread.join()
+ with self._lock:
+ loaded, before, low = set(self._loaded), self._before, self._low
+ estimate = self._estimate
+ if loaded and before is not None and low is not None:
+ record_need(self.config, loaded, before - low, estimate)
+ if loaded and (self.unload_on_release or self.tripped):
+ unload(self.config, sorted(loaded))
+
+
+_active: Guard | None = None
+_active_lock = threading.Lock()
+
+# One Skill++ process uses the models at a time. Two folds at once, two chats
+# closed together or a session start's sweep beside a session end, each loaded
+# the models for itself, and each unloaded them from under the other. Measured
+# the same way: a fold that started while the one before was still unloading
+# counted the leaving model as someone else's and left it loaded.
+_MODELS_LOCK_SECONDS = 3600.0
+_MODELS_LOCK_POLL = 2.0
+BUSY = "is already folding a session; try again in a minute"
+
+
+def active() -> Guard | None:
+ """The guard `local.ask` and `local.embed` answer to, if one is open."""
+ return _active
+
+
+def _take_models_lock(path: Path, wait: bool) -> bool:
+ from .capture import _acquire_lock
+ deadline = time.time() + _MODELS_LOCK_SECONDS
+ while True:
+ if _acquire_lock(path, _MODELS_LOCK_SECONDS, "models"):
+ return True
+ if not wait or time.time() > deadline:
+ return False
+ time.sleep(_MODELS_LOCK_POLL)
+
+
+@contextmanager
+def guarded(config, *, wait: bool = True, **kwargs) -> Iterator[Guard]:
+ """The guard for one fold, sweep or page request.
+
+ Nested uses share the outermost: a `fold-pending` sweep keeps the models
+ loaded from one session to the next and unloads them once, at the end.
+ Across processes, one guard at a time holds the models: a fold waits for
+ the one before it, and with *wait* off (the review page, which must not
+ hang) the guard refuses every call instead.
+ """
+ global _active
+ with _active_lock:
+ outer = _active
+ if outer is None:
+ _active = guard = Guard(config, **kwargs)
+ if outer is not None:
+ yield outer
+ return
+ lock = config.root / "models.lock"
+ held = False
+ try:
+ if guard.enabled:
+ config.root.mkdir(parents=True, exist_ok=True)
+ held = _take_models_lock(lock, wait)
+ if not held:
+ guard.tripped = BUSY
+ yield guard
+ finally:
+ with _active_lock:
+ _active = None
+ guard.release()
+ if held:
+ try:
+ lock.unlink()
+ except OSError:
+ pass
diff --git a/skill_plus_plus/notify.py b/skill_plus_plus/notify.py
new file mode 100644
index 0000000..c60a5b5
--- /dev/null
+++ b/skill_plus_plus/notify.py
@@ -0,0 +1,81 @@
+"""Tell the person when the memory guard holds or stops a fold.
+
+The fold runs detached after the chat has ended, so a message in the session
+would reach nobody, and a SessionStart hook's output goes to the model, not to
+the person (Claude Code hooks reference). The operating system's own
+notifications are the one place they will see it: Notification Center through
+`osascript` on macOS, `notify-send` on Linux. Anywhere else, and whenever those
+fail, the notice reaches the log only.
+
+Nothing here raises. A missing notifier costs the notice, never the fold.
+"""
+
+from __future__ import annotations
+
+import shutil
+import subprocess
+import sys
+import time
+
+TITLE = "Skill++"
+# "Waiting" repeats while sessions wait, so it is sent at most this often.
+# "Stopped" is sent every time: it only happens when memory ran short.
+WAITING_EVERY_SECONDS = 3600
+
+
+def command(message: str, platform: str | None = None) -> list[str] | None:
+ """The command that shows *message* on this system, or None."""
+ platform = platform or sys.platform
+ if platform == "darwin":
+ return ["osascript", "-e",
+ f"display notification {_apple(message)} with title {_apple(TITLE)}"]
+ if platform.startswith("linux") and shutil.which("notify-send"):
+ return ["notify-send", "--app-name", TITLE, TITLE, message]
+ return None
+
+
+def _apple(text: str) -> str:
+ """An AppleScript string literal."""
+ return '"' + str(text).replace("\\", "\\\\").replace('"', '\\"') + '"'
+
+
+def send(config, message: str, *, kind: str, every: float = 0.0) -> bool:
+ """Show *message*, at most once per *every* seconds for this *kind*."""
+ from .capture import log_error
+ from .memory import _load, _store
+
+ log_error(config, f"notice ({kind}): {message}")
+ if not config.notify:
+ return False
+ data = _load(config)
+ last = (data.get("notified") or {}).get(kind)
+ if every and isinstance(last, (int, float)) and time.time() - last < every:
+ return False
+ argv = command(message)
+ if not argv:
+ return False
+ try:
+ subprocess.run(argv, capture_output=True, timeout=10, check=False)
+ except (OSError, subprocess.SubprocessError):
+ return False
+ data.setdefault("notified", {})[kind] = time.time()
+ _store(config, data)
+ return True
+
+
+def waiting(config, reason: str) -> bool:
+ """A fold was held: the models do not fit in memory now."""
+ return send(config,
+ f"Skill++ is waiting to fold your session: it {reason}. "
+ "It folds on its own once your computer is idle.",
+ kind="waiting", every=WAITING_EVERY_SECONDS)
+
+
+def stopped(config, reason: str, freed: int | None) -> bool:
+ """The watchdog stopped a fold and unloaded the models."""
+ from .memory import gb
+ what = f"stopped and freed {gb(freed)} GB" if freed else "stopped its local model"
+ return send(config,
+ f"Skill++ {what}: {reason}. "
+ "Your session is kept and folded later.",
+ kind="stopped")
diff --git a/skill_plus_plus/web.py b/skill_plus_plus/web.py
index 43dd0bb..c9dbad7 100644
--- a/skill_plus_plus/web.py
+++ b/skill_plus_plus/web.py
@@ -381,8 +381,13 @@ def summarise(config: Config, entry_id: str) -> dict:
before capture named them β then cached per entry and keyed on step count,
so the page never waits on a model twice and a candidate that grows is
described again.
+
+ Under the memory guard, like a fold: the model loads only if it fits. It is
+ left loaded for `memory.KEEP_ALIVE` rather than unloaded, because opening
+ one row is usually followed by opening the next.
"""
- from .local import LocalModelUnavailable
+ from .local import LocalModelUnavailable, MemoryShort
+ from .memory import BUSY, guarded
from .summary import name_and_sentence, store_summary
entry = Ledger(config).get(entry_id)
@@ -393,7 +398,11 @@ def summarise(config: Config, entry_id: str) -> dict:
return {"ok": True, "summary": cached}
try:
- _, text = name_and_sentence(config, entry)
+ with guarded(config, wait=False, unload_on_release=False):
+ _, text = name_and_sentence(config, entry)
+ except MemoryShort as exc:
+ return {"ok": False, "error": f"Skill++ {exc}" if str(exc) == BUSY
+ else f"not enough free memory: Skill++ {exc}"}
except LocalModelUnavailable as exc:
return {"ok": False, "error": f"no local model: {exc}"}
if not text:
@@ -427,6 +436,10 @@ def project_list(rows: list[dict], drafts: list[dict]) -> list[dict]:
def collect_state(config: Config) -> dict:
"""The rows the page lists. Reads files only: no model, no agent.
+ One exception, only while sessions are held for memory: the banner reads
+ the free memory, and the models' sizes from Ollama if what they take here
+ was never measured. Neither loads a model.
+
Every candidate, promoted and dismissed entry, most-seen first. `ready`
splits the page: recognized often enough to decide on, or still collecting.
"""
@@ -457,7 +470,46 @@ def collect_state(config: Config) -> dict:
projects = project_list(rows, drafts)
return {"threshold": threshold, "ttl": config.candidate_ttl_days,
"rows": rows, "drafts": drafts, "projects": projects,
- "capture": capture_status([p["path"] for p in projects])}
+ "capture": capture_status([p["path"] for p in projects]),
+ "memory": memory_waiting(config)}
+
+
+def memory_waiting(config: Config) -> dict:
+ """How many sessions the memory guard is holding, for the page's banner.
+
+ Counted from the session files. Only when some are waiting does it also
+ say from what free memory a fold starts, and how much is free now.
+ """
+ waiting = 0
+ for path in config.sessions_dir.glob("*.json"):
+ try:
+ held = json.loads(path.read_text(encoding="utf-8")).get("held")
+ except (OSError, json.JSONDecodeError):
+ continue
+ waiting += bool(isinstance(held, dict) and held.get("memory"))
+ if not waiting:
+ return {"waiting": 0}
+ from .memory import gb, status
+ s = status(config)
+ return {"waiting": waiting, "start_at_gb": gb(s["start_at"]),
+ "free_gb": gb(s["available"])}
+
+
+def fold_now(config: Config) -> dict:
+ """"Fold now" on the banner: `skill-plus-plus fold-pending --now`, if it fits.
+
+ The memory rule still holds: when the models do not fit, nothing loads and
+ the page says how much memory is missing.
+ """
+ from .memory import shortfall
+ short = shortfall(config)
+ if short:
+ return {"ok": False,
+ "error": f"Not yet: Skill++ {short}. Close a few apps and try again."}
+ proc = _run(config, "fold-pending", "--now")
+ if proc.returncode != 0:
+ return {"ok": False, "error": _tail(proc.stderr or proc.stdout)}
+ return {"ok": True}
def capture_status(project_paths: list[str], home: Path | None = None) -> dict:
@@ -1035,6 +1087,7 @@ def make_handler(config: Config, skills_dir: Path | None = None):
str(p.get("instruction", ""))),
"/api/answer": lambda p: answer_questions(config, str(p.get("id", "")),
p.get("answers") or []),
+ "/api/fold-now": lambda p: fold_now(config),
}
class Handler(BaseHTTPRequestHandler):
@@ -1211,6 +1264,9 @@ def serve(config: Config, skills_dir: Path | None = None, port: int = 8765,
.warn{margin:0 0 20px;padding:12px 16px;border:1px solid var(--noline);background:var(--nobg);
border-radius:8px;color:var(--fg);font-size:13px;line-height:1.6}
.warn strong{color:var(--no)}
+ .warn.wait{border-color:rgba(251,191,36,.4);background:rgba(251,191,36,.07)}
+ .warn.wait strong{color:#fbbf24}
+ .warn button.fold-now{margin-left:8px;color:#0c0d10;border-color:#fbbf24;background:#fbbf24;font-weight:600}
.cand{background:var(--panel);border:1px solid var(--line);border-radius:8px;margin-bottom:8px}
.cand.ready{border-left:3px solid #fbbf24;background:linear-gradient(90deg,rgba(251,191,36,.07),var(--panel) 40%)}
.cand.accepted{border-left:3px solid var(--go);background:linear-gradient(90deg,rgba(56,189,248,.07),var(--panel) 40%)}
@@ -1394,6 +1450,17 @@ def serve(config: Config, skills_dir: Path | None = None, port: int = 8765,
Run skill-plus-plus install --user --apply, then start a new Claude Code session;
skill-plus-plus doctor checks it. Installing β`;
}
+// Sessions the memory guard is holding: kept, not lost, and folded once the
+// computer is idle and the models fit, or now if memory allows.
+function memoryWarning(){
+ const m = S.memory || {waiting: 0};
+ if(!m.waiting) return "";
+ const n = m.waiting === 1 ? "1 session is" : `${m.waiting} sessions are`;
+ return `
Skill++ is watching your Claude Code sessions. A procedure shows up here once you've repeated it ${S.threshold}Γ.
` : ""; @@ -1941,7 +2008,7 @@ def serve(config: Config, skills_dir: Path | None = None, port: int = 8765, const promoted = S.rows.filter(r => !["undecided", "collecting", "dismissed"].includes(r.state) && !inDrafts(r)); const open_ = S.rows.filter(r => ["undecided", "collecting"].includes(r.state)); const ready = open_.filter(r => r.ready), collecting = open_.filter(r => !r.ready); - list.innerHTML = `${captureWarning()} + list.innerHTML = `${captureWarning()}${memoryWarning()} ${promoted.length ? `