diff --git a/CHANGELOG.md b/CHANGELOG.md index 28fc40f..4ff54aa 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,22 @@ All notable changes to Skill++. The format follows ## [Unreleased] +### Added + +- A memory guard for the local model. Measured on an 18 GB Mac, loading + `gemma4:e4b` and the embedder took 12.8 GB of free memory, and a fold with + other apps open ran at 92 % used with macOS swapping. There was no leak; the + model is simply large. Now a fold loads the models only when they fit with + 2 GB to spare, stops and unloads them at once if free memory falls below + that, and unloads them as soon as it ends instead of after Ollama's five + minutes. What the models take is measured on each computer; with + `gemma4:e4b-it-qat` a fold starts at about 9 GB free. A session that doesn't + fit is kept and folded later: when the computer is idle, at the next session + start, or with **Fold now** on the review page. A desktop notification says + when a fold waits or is stopped, and `doctor` and `stats` show the figures. + `SKILL_PLUS_PLUS_MEMORY_GUARD`, `SKILL_PLUS_PLUS_MEMORY_RESERVE_GB`, + `SKILL_PLUS_PLUS_IDLE_MINUTES` and `SKILL_PLUS_PLUS_NOTIFY` configure it. + ### Changed - The local model is `gemma4:e4b-it-qat`, a build of `gemma4:e4b` trained to diff --git a/README.md b/README.md index 689f02c..d90031f 100644 --- a/README.md +++ b/README.md @@ -153,8 +153,10 @@ new Code session in the desktop app) and work as usual. - `install` without `--apply` only shows what it would do, downloads included. Your settings file is backed up before it is changed. - `--user` instead of `--project` captures every project on the machine. -- The two models are about 6.4 GB on disk, and cutting a session needs about - 7 GB of free memory. `--no-models` leaves Ollama alone. +- The two models are about 6.4 GB on disk, and folding a session takes about + 7 GB of free memory. Skill++ only starts when that fits with 2 GB to spare, + and otherwise waits, so it never pushes your machine into swap + ([Memory](docs/usage.md#memory)). `--no-models` leaves Ollama alone. - `skill-plus-plus install --project ~/code/my-repo --remove --apply` takes the hooks and slash commands out again. - The latest `main`, before it is released: @@ -311,8 +313,9 @@ including sessions of our own work that we cannot publish. ## 🧭 When to use Β· when to skip **Good fit if you** repeat procedures in Claude Code (the CLI or the desktop -app's Code tab), on macOS or Linux, with about 7 GB of memory to spare for the -local model. +app's Code tab), on macOS or Linux, with about 9 GB of memory free now and +then: the local model takes about 7 GB while it folds a session, and Skill++ +waits until that fits. **Skip it if you** mostly do one-off work, run Windows, or use another agent: capture is Claude Code only for now. Drafting can use any agent CLI diff --git a/docs/usage.md b/docs/usage.md index b6aedc1..f368a27 100644 --- a/docs/usage.md +++ b/docs/usage.md @@ -178,6 +178,43 @@ edits settings or spends a model call is a dry run until you add `--apply`. | Skills you have | `lifecycle`, `tier`, `check`, `reconcile`, `bundle`, `expire`, `accuracy` | | Internal (run by the hooks) | `hook`, `fold-session`, `fold-pending` | +## Memory + +The local model is the one large thing Skill++ runs. On an 18 GB Mac, a fold +with `gemma4:e4b-it-qat` and `nomic-embed-text` takes about **7 GB** of free +memory, whatever the session: a one-word question costs as much as a whole +session. `gemma4:e4b`, the default before it, took 12.8 GB, and with other +apps open a fold ran at 92 % memory used while macOS swapped out 4.7 GB in +100 seconds. So Skill++ guards its memory use: + +- **It starts only if the models fit.** A fold loads the models only when + what they take still leaves 2 GB free: with the default models, from about + 9 GB free. What they take is measured on your computer: estimated from + their size at first, which asks for about 10 GB free, then the largest + amount a fold has actually used here. +- **It stops when memory runs short.** If free memory falls below 2 GB while a + fold runs, or macOS reports critical memory pressure, Skill++ stops and + unloads the models at once. +- **It lets go right away.** When a fold ends, the models are unloaded, rather + than staying in memory for Ollama's usual five minutes. +- **It never unloads what it didn't load.** A model you had loaded yourself + stays loaded. + +A session that doesn't fit isn't lost. It waits, and is folded: + +- **when your computer is idle,** five minutes without keyboard or mouse input, + once the models fit; +- **at the next session start,** retried at most every ten minutes; +- **when you ask:** **Fold now** on the review page's banner, or + `skill-plus-plus fold-pending --now`. The models still load only if they fit. + +When a fold waits or is stopped, a desktop notification says so: Notification +Center on macOS, `notify-send` on Linux. On a Mac the first one may ask you to +allow notifications for Script Editor, which is what shows them. +`skill-plus-plus doctor` shows what the models take on your computer and from +what free memory a fold starts; `skill-plus-plus stats` counts the sessions +waiting. + ## Configuration All settings are environment variables. @@ -198,6 +235,10 @@ All settings are environment variables. | `SKILL_PLUS_PLUS_MATCH` | `1` | `0` banks every task without comparing it (for measuring detection) | | `SKILL_PLUS_PLUS_DESCRIBE` | `0` | `1` asks the model to describe every tool call, inside the hook (slow) | | `SKILL_PLUS_PLUS_MAX_STEPS` / `SKILL_PLUS_PLUS_MAX_FIELD` | `500` / `2000` | caps per session and per captured field | +| `SKILL_PLUS_PLUS_MEMORY_GUARD` | `1` | `0` turns the [memory guard](#memory) off: folds load the models whatever is free | +| `SKILL_PLUS_PLUS_MEMORY_RESERVE_GB` | `2` | free memory a fold always leaves; it starts only if the models fit with this much to spare, and stops below it | +| `SKILL_PLUS_PLUS_IDLE_MINUTES` | `5` | how long nobody must use the keyboard or mouse before a session waiting for memory is folded | +| `SKILL_PLUS_PLUS_NOTIFY` | `1` | `0` turns off the desktop notices when a fold waits for memory or is stopped | | `SKILL_PLUS_PLUS_INTERNAL` | unset | set to anything to make the hooks do nothing, e.g. for one session | ## Troubleshooting @@ -225,6 +266,12 @@ running. Stop it with Ctrl-C in its terminal, or use `skill-plus-plus web --port **Tool calls feel slow.** Check that `SKILL_PLUS_PLUS_DESCRIBE` is not set to `1`; it asks the local model about every tool call inside the hook. +**Sessions keep waiting for memory.** `skill-plus-plus doctor` says from what +free memory a fold starts. Close a few apps and press **Fold now** on the +review page, or leave the computer idle for a few minutes. If `doctor` says the +models take more than the machine has, sessions never fold while the guard is +on; `SKILL_PLUS_PLUS_MEMORY_GUARD=0` folds them anyway, at the risk of swapping. + **Leave one session out.** Start it as `SKILL_PLUS_PLUS_INTERNAL=1 claude`, and the hooks do nothing for it. diff --git a/skill_plus_plus/capture.py b/skill_plus_plus/capture.py index 69a26aa..6b3b531 100644 --- a/skill_plus_plus/capture.py +++ b/skill_plus_plus/capture.py @@ -548,22 +548,47 @@ def handle_session_end(config: Config, payload: dict) -> dict: # keep path already attaches before judging. _attach_reply(session, payload.get("transcript_path") or session.get("transcript")) - if config.judge_boundaries and "folded" not in session: - try: - from .boundary import judge_session - judge_session(config, session) - except Exception as exc: # noqa: BLE001 - a hook never raises at a dev - log_error(config, f"boundary judge failed: {type(exc).__name__}: {exc}") - - # The last task's own completion report, said after its final tool call. - # Attached before folding so the episode carries it. - trailing = _trailing_narration(payload) - if trailing: - work = [s for s in session.get("steps", []) if not is_prompt(s)] - if work: - work[-1].setdefault("closing_note", trailing) - _save_session(config, session) - result = fold_session(config, session, persist=True) + # The memory guard (`skill_plus_plus.memory`) around everything that loads + # a model. The models load only if they fit and still leave the reserve + # free; otherwise the session is held and folded later, which costs + # nothing. A watchdog stops the fold and unloads them if memory runs short + # while it runs, and they are unloaded when it ends. + from . import memory + with memory.guarded(config) as guard: + short = guard.admit(memory.fold_models(config)) + if short: + return _hold_for_memory(config, session, short) + + if config.judge_boundaries and "folded" not in session: + try: + from .boundary import judge_session + judge_session(config, session) + except Exception as exc: # noqa: BLE001 - a hook never raises at a dev + log_error(config, f"boundary judge failed: {type(exc).__name__}: {exc}") + if guard.tripped: + # Stopped halfway through judging. The gaps judged so far carry + # verdicts and the rest carry `judge_session`'s up-front + # `False`, which reads as "no boundary here": banked, the + # session would merge tasks the judge never got to. Nothing is + # banked yet, so every verdict goes and the retry judges it all. + for step in session.get("steps", []): + step.pop("end", None) + return _hold_for_memory(config, session, guard.tripped, notify=False) + + # The last task's own completion report, said after its final tool call. + # Attached before folding so the episode carries it. + trailing = _trailing_narration(payload) + if trailing: + work = [s for s in session.get("steps", []) if not is_prompt(s)] + if work: + work[-1].setdefault("closing_note", trailing) + _save_session(config, session) + result = fold_session(config, session, persist=True) + if guard.tripped and result.get("status") == "offline": + # Stopped between episodes: `folded` keeps the ones banked, and the + # retry resumes after them, as it does when Ollama goes away. + result["reason"] = f"memory: {guard.tripped}" + result["memory"] = True # Offline: the session was never folded, so this file is the only copy of # the work. Losing the step is the one thing capture exists to prevent β€” # being offline costs the candidate, never the record. @@ -576,6 +601,8 @@ def handle_session_end(config: Config, payload: dict) -> dict: "at": datetime.now(timezone.utc).isoformat(timespec="seconds"), "reason": result.get("reason", ""), } + if result.get("memory"): + session["held"]["memory"] = True _save_session(config, session) else: try: @@ -585,6 +612,29 @@ def handle_session_end(config: Config, payload: dict) -> dict: return result +def _hold_for_memory(config: Config, session: dict, reason: str, *, + notify: bool = True) -> dict: + """Keep the session to fold later: the models do not fit in memory now. + + Held like a session no model answered, so `fold_pending` retries it and + `stats` counts it, with `memory` set so the retry waits for memory rather + than for Ollama. *notify* is off when the watchdog stopped the fold, which + has already said so. + """ + session["held"] = { + "at": datetime.now(timezone.utc).isoformat(timespec="seconds"), + "reason": f"memory: {reason}", + "memory": True, + } + _save_session(config, session) + if notify: + from .notify import waiting + waiting(config, reason) + return {"status": "offline", "steps": len(session.get("steps", [])), + "episodes": [], "flagged": 0, "reason": session["held"]["reason"], + "memory": True} + + def mark_ending(config: Config, session_id: str, transcript: str | None = None) -> dict: """Record that a session ended. The fold itself happens elsewhere. @@ -648,6 +698,19 @@ def fold_session_now(config: Config, session_id: str, *, # A ceiling on one session's fold. Bounds the damage a recycled pid can do: # past this the lock is stale whatever `os.kill` says. _FOLD_LOCK_SECONDS = 1800 +# How often the sweep retries a session held for memory (see `_is_pending`). +MEMORY_RETRY_SECONDS = 600.0 + + +def _age_seconds(stamp) -> float: + """Seconds since an ISO stamp. An unreadable one counts as long ago.""" + try: + at = datetime.fromisoformat(str(stamp)) + except ValueError: + return float("inf") + if at.tzinfo is None: + at = at.replace(tzinfo=timezone.utc) + return (datetime.now(timezone.utc) - at).total_seconds() def _lock_file(config: Config, session_id: str) -> Path: @@ -751,7 +814,8 @@ def _find_transcript(session_id: str) -> str | None: def fold_pending(config: Config, *, exclude: str | None = None, - idle_hours: float = PENDING_IDLE_HOURS) -> list[dict]: + idle_hours: float = PENDING_IDLE_HOURS, + now: bool = False) -> list[dict]: """Bank the sessions that ended without being banked. `SessionEnd` stamps `ending` and spawns a worker, so this sweep is the @@ -762,48 +826,115 @@ def fold_pending(config: Config, *, exclude: str | None = None, per-session lock and the held stamp are the ones a normal end uses, and `folded` keeps a partial retry from counting an episode twice. *exclude* is the session that is starting, which is live by definition. + + One memory guard covers the sweep, so the models load once for every + session in it and are unloaded at the end. The first session held for + memory ends the sweep: what did not fit for it will not fit for the next. + *now* skips the pause between memory retries, for the idle waiter and + "Fold now", which have just checked that the models fit. """ + from . import memory + results = [] with _locked(config.root / "fold-pending.lock", _PENDING_LOCK_SECONDS, "fold-pending") as got: if not got: return [{"status": "locked"}] - for path in sorted(config.sessions_dir.glob("*.json")): - sid = path.stem - if sid == exclude: - continue - try: - session = json.loads(path.read_text(encoding="utf-8")) - idle = (time.time() - path.stat().st_mtime) / 3600 - except (OSError, json.JSONDecodeError): - continue - if not _is_pending(session, idle, idle_hours): - results.append({"session": sid, "status": "live"}) - continue - # Checked before folding only so the outcome can say "folding" - # rather than "locked"; `fold_session_now` takes the lock itself, - # so nothing rests on this being race-free. - if _lock_alive(_lock_file(config, sid), _FOLD_LOCK_SECONDS): - results.append({"session": sid, "status": "folding"}) - continue - results.append({"session": sid, **fold_session_now(config, sid)}) + with memory.guarded(config): + for path in sorted(config.sessions_dir.glob("*.json")): + sid = path.stem + if sid == exclude: + continue + try: + session = json.loads(path.read_text(encoding="utf-8")) + idle = (time.time() - path.stat().st_mtime) / 3600 + except (OSError, json.JSONDecodeError): + continue + if not _is_pending(session, idle, idle_hours, now=now): + results.append({"session": sid, "status": "live"}) + continue + # Checked before folding only so the outcome can say "folding" + # rather than "locked"; `fold_session_now` takes the lock itself, + # so nothing rests on this being race-free. + if _lock_alive(_lock_file(config, sid), _FOLD_LOCK_SECONDS): + results.append({"session": sid, "status": "folding"}) + continue + result = fold_session_now(config, sid) + results.append({"session": sid, **result}) + if result.get("memory"): + break _sweep_orphan_locks(config) return results -def _is_pending(session: dict, idle: float, idle_hours: float) -> bool: +# How long a worker that held a session for memory waits for a better moment, +# and how often it looks. Past this, the next session start takes over. +MEMORY_WAIT_SECONDS = 12 * 3600 +MEMORY_POLL_SECONDS = 60.0 + + +def _held_for_memory(config: Config) -> bool: + for path in config.sessions_dir.glob("*.json"): + try: + if (json.loads(path.read_text(encoding="utf-8")).get("held") or {}).get("memory"): + return True + except (OSError, json.JSONDecodeError, AttributeError): + continue + return False + + +def wait_for_memory(config: Config, *, sleep=time.sleep, clock=time.time) -> list[dict]: + """Fold the sessions held for memory once the computer is idle and they fit. + + Run by the worker that held a session, after its own lock is released: + nobody waits on it, and it costs a sleeping process. One waiter at a time; + another worker that finds the lock taken leaves it to the one waiting. + Idle means nobody at the keyboard for `config.idle_minutes`, so the model + never competes with someone working. Where idle time cannot be read, the + waiter waits for memory alone. + """ + from . import memory + + with _locked(config.root / "memory-wait.lock", MEMORY_WAIT_SECONDS + 600, + "memory-wait") as got: + if not got: + return [] + deadline = clock() + MEMORY_WAIT_SECONDS + while clock() < deadline: + sleep(MEMORY_POLL_SECONDS) + if not _held_for_memory(config): + return [] + idle = memory.idle_seconds() + if idle is not None and idle < config.idle_minutes * 60: + continue + if memory.shortfall(config): + continue + results = fold_pending(config, now=True) + if not _held_for_memory(config): + return results + return [] + + +def _is_pending(session: dict, idle: float, idle_hours: float, *, + now: bool = False) -> bool: """Has this session ended without being banked? Three rules, covering three different failures, none of them redundant: - * ``held`` β€” a fold that ran and could not reach a model. + * ``held`` β€” a fold that ran and could not reach a model, or was held for + memory. Those are retried at most every `MEMORY_RETRY_SECONDS`, unless + *now*: a retry that passes the check and trips the watchdog again would + load and unload the models at every session start. * ``ending`` past the grace β€” a fold that was launched and died with it. Without the stamp it would look like a live session and wait out the idle rule. * idle β€” the only rule that catches a session where `SessionEnd` never fired at all: a window closed, a laptop shut down. """ - if session.get("held"): + held = session.get("held") + if held: + if isinstance(held, dict) and held.get("memory") and not now: + return _age_seconds(held.get("at")) >= MEMORY_RETRY_SECONDS return True ending = session.get("ending") if ending: diff --git a/skill_plus_plus/cli.py b/skill_plus_plus/cli.py index bb19439..206835b 100644 --- a/skill_plus_plus/cli.py +++ b/skill_plus_plus/cli.py @@ -555,6 +555,32 @@ def cmd_split(args: argparse.Namespace) -> int: return 0 +def _memory_guarded(command): + """Run a command that calls the local models under the memory guard. + + The same rule a fold follows (`skill_plus_plus.memory`): the models load + only if they fit and leave the reserve free, and are unloaded when the + command ends. + """ + def run(args: argparse.Namespace) -> int: + from .memory import guarded + with guarded(Config(args.root)): + return command(args) + run.__doc__ = command.__doc__ + return run + + +def _model_trouble(exc: Exception) -> str: + from .local import MemoryShort + from .memory import BUSY + if isinstance(exc, MemoryShort): + if str(exc) == BUSY: + return f"Skill++ {exc}" + return f"not enough free memory: Skill++ {exc}" + return f"no local model: {exc}" + + +@_memory_guarded def cmd_merge(args: argparse.Namespace) -> int: """Fold existing candidates that are the same procedure. @@ -590,7 +616,8 @@ def cmd_merge(args: argparse.Namespace) -> int: try: vectors = {e.id: vector_for(e, config, cache) for e in entries} except LocalModelUnavailable as exc: - print(f"No embedding model: {exc}. Nothing was compared.", file=sys.stderr) + why = _model_trouble(exc) + print(f"{why[:1].upper()}{why[1:]}. Nothing was compared.", file=sys.stderr) return 1 finally: save_cache(config, cache) @@ -690,6 +717,15 @@ def cmd_fold_session(args: argparse.Namespace) -> int: + (f" {result['id']}" if result.get("id") else "")) if args.verbose: print(json.dumps(result)) + if result.get("memory"): + # Held because the models did not fit. This worker is detached and + # nobody waits on it, so it stays to fold the held sessions once the + # computer is idle and they fit (`capture.wait_for_memory`). + from .capture import wait_for_memory + try: + wait_for_memory(config) + except Exception as exc: # noqa: BLE001 - detached; a traceback goes nowhere + log_error(config, f"memory wait failed: {type(exc).__name__}: {exc}") return 0 @@ -716,7 +752,8 @@ def cmd_fold_pending(args: argparse.Namespace) -> int: config = Config(args.root) config.ensure_dirs() - results = fold_pending(config, exclude=args.exclude, idle_hours=args.idle_hours) + results = fold_pending(config, exclude=args.exclude, idle_hours=args.idle_hours, + now=args.now) if results == [{"status": "locked"}]: print("Another fold-pending is running.") return 0 @@ -726,6 +763,8 @@ def cmd_fold_pending(args: argparse.Namespace) -> int: what = "still live, skipped" elif r["status"] == "folding": what = "a worker is folding it" + elif r.get("memory"): + what = f"waiting for memory: {r.get('reason', '').removeprefix('memory: ')}" elif r["status"] == "offline": what = f"held again: {r.get('reason', '')}" elif episodes: @@ -881,6 +920,7 @@ def cmd_reopen(args: argparse.Namespace) -> int: return 1 +@_memory_guarded def cmd_retitle(args: argparse.Namespace) -> int: """Name candidates the fold could not name, because no model answered. @@ -911,7 +951,7 @@ def cmd_retitle(args: argparse.Namespace) -> int: try: name, sentence = name_and_sentence(config, entry) except LocalModelUnavailable as exc: - print(f"no local model: {exc}") + print(_model_trouble(exc)) return 1 if sentence: store_summary(config, entry, sentence) @@ -1240,6 +1280,26 @@ def cmd_doctor(args: argparse.Namespace) -> int: here = _has_model(models, want) print(f" {'':8} {want} {'βœ“' if here else 'βœ— not pulled'}") + # What the models take, from what free memory a fold starts, and whether + # this machine can ever get there (`skill_plus_plus.memory`). + from .memory import gb, status + m = status(config) + if m["available"] is None: + print(f"memory {'':8} unknown on this system: the memory guard is off") + elif not config.memory_guard: + print(f"memory {'':8} guard off (SKILL_PLUS_PLUS_MEMORY_GUARD=0); " + f"{gb(m['available'])} GB free of {gb(m['total'])} GB") + else: + print(f"memory {'':8} {gb(m['total'])} GB RAM, {gb(m['available'])} GB free now") + print(f" {'':8} the models take about {gb(m['need'])} GB here " + f"({'measured' if m['measured'] else 'estimated from their size'}); " + f"a fold starts at {gb(m['start_at'])} GB free") + if m["total"] is not None and m["start_at"] > m["total"]: + print(f" {'':8} more than this machine has: sessions wait and " + "never fold while the guard is on") + print(f" {'':8} SKILL_PLUS_PLUS_MEMORY_GUARD=0 folds them anyway, " + "at the risk of swapping") + s = _waiting_sessions(config) pending = len(s["held"]) + len(s["waiting"]) if pending or s["folding"]: @@ -1279,18 +1339,27 @@ def cmd_stats(args: argparse.Namespace) -> int: # Only sessions `handle_session_end` stamped as held. A session still being # written has a file too, and counting it would report work lost from one # that is merely in flight. - held = [] + held, for_memory = [], [] for path in sorted(config.sessions_dir.glob("*.json")): try: - if json.loads(path.read_text(encoding="utf-8")).get("held"): - held.append(path) + stamp = json.loads(path.read_text(encoding="utf-8")).get("held") except (OSError, json.JSONDecodeError): continue + if stamp: + (for_memory if isinstance(stamp, dict) and stamp.get("memory") else held).append(path) if held: print(f"held {len(held)} session(s) not banked, kept in " f"{config.sessions_dir}") print(f" no local model answered ({config.local_model} at " f"{config.ollama_url}); the work is there, the candidates are not") + if for_memory: + from .memory import gb, status + s = status(config) + print(f"waiting {len(for_memory)} session(s) held for memory, kept in " + f"{config.sessions_dir}") + print(f" a fold starts when {gb(s['start_at'])} GB are free " + f"({gb(s['available'])} GB now); they fold once the computer is idle,") + print(" or run `skill-plus-plus fold-pending --now`") return 0 @@ -1734,6 +1803,9 @@ def build_parser() -> argparse.ArgumentParser: p.add_argument("--exclude", help="a live session to leave alone") p.add_argument("--idle-hours", type=float, default=12.0, help="treat a session untouched this long as ended") + p.add_argument("--now", action="store_true", + help="retry sessions held for memory now, not after the pause " + "(the models still load only if they fit)") p.set_defaults(func=cmd_fold_pending) p = sub.add_parser("web", help="browse the ledger in a local page") diff --git a/skill_plus_plus/config.py b/skill_plus_plus/config.py index 34c9cd6..b2b4e65 100644 --- a/skill_plus_plus/config.py +++ b/skill_plus_plus/config.py @@ -134,6 +134,28 @@ def __init__(self, root: str | Path | None = None) -> None: # A boundary that would close an episode smaller than this is ignored: # one step is not a workflow. self.min_episode_steps = _int_env("SKILL_PLUS_PLUS_MIN_EPISODE_STEPS", 2) + # The memory guard (`skill_plus_plus.memory`). The local model is not + # small: on an 18 GB Mac, loading gemma4:e4b and the embedder took 12.8 + # GB of available memory (gemma4:e4b-it-qat, the default after it, about + # 7), and a fold with other apps open ran at 92 % used with 4.7 GB + # swapped out in 100 s. So a fold starts only when the + # models fit with this much memory still free, and stops, unloading + # them, the moment less than this is left. An absolute figure, not a + # share of RAM: what keeps a machine out of swap is headroom in + # gigabytes. 2 GB is where that Mac turned: 2.0 GB available while + # loading, pressure normal and nothing swapped; 1.4-1.6 GB while + # folding, pressure warning and swapping. `SKILL_PLUS_PLUS_MEMORY_GUARD=0` + # turns the guard off. + self.memory_guard = _bool_env("SKILL_PLUS_PLUS_MEMORY_GUARD", True) + self.memory_reserve_gb = _float_env("SKILL_PLUS_PLUS_MEMORY_RESERVE_GB", 2.0) + # A session held for memory is folded once the computer has been idle + # this long and the models fit, so the model never competes with + # someone at the keyboard. Folding later costs nothing. + self.idle_minutes = _float_env("SKILL_PLUS_PLUS_IDLE_MINUTES", 5.0) + # Desktop notifications when the guard holds or stops a fold. The fold + # runs after the chat has ended, so the operating system is the one + # place the person will see it. `SKILL_PLUS_PLUS_NOTIFY=0` turns them off. + self.notify = _bool_env("SKILL_PLUS_PLUS_NOTIFY", True) @property def decisions_file(self) -> Path: @@ -168,6 +190,11 @@ def archive_dir(self) -> Path: def log_file(self) -> Path: return self.root / "skill-plus-plus.log" + @property + def memory_file(self) -> Path: + """What the models took on this computer, and when a notice was last sent.""" + return self.root / "memory.json" + def ensure_dirs(self) -> None: for d in (self.ledger_dir, self.sessions_dir, self.cold_dir, self.archive_dir): d.mkdir(parents=True, exist_ok=True) diff --git a/skill_plus_plus/local.py b/skill_plus_plus/local.py index 171a0a4..b7f39d8 100644 --- a/skill_plus_plus/local.py +++ b/skill_plus_plus/local.py @@ -43,6 +43,31 @@ class LocalModelUnavailable(RuntimeError): """Ollama is not reachable, or the model is not installed.""" +class MemoryShort(LocalModelUnavailable): + """The memory guard kept a model from loading, or stopped the fold. + + A subclass, so every caller that already treats an unreachable model as + "no opinion" and holds the session does the same here, with no new path. + """ + + +def _guard(model: str, payload: dict) -> None: + """Ask the open memory guard, if any, before *model* is called. + + Raises `MemoryShort` when the guard refuses. A model the guard loaded is + asked to leave memory soon after its last call (`memory.KEEP_ALIVE`), so a + fold that dies before its guard is released cannot hold it for Ollama's + five minutes. A model someone else loaded keeps its own `keep_alive`. + """ + from .memory import KEEP_ALIVE, active + guard = active() + if guard is None: + return + guard.before_call(model) + if guard.owns(model): + payload["keep_alive"] = KEEP_ALIVE + + def _num_ctx(prompt: str, reserve: int = 512) -> int: want = (len(prompt) // _CHARS_PER_TOKEN) + reserve return max(_CTX_FLOOR, min(_CTX_CEILING, want)) @@ -74,6 +99,7 @@ def ask(model: str, prompt: str, *, host: str = DEFAULT_HOST, } if think is not None: payload["think"] = think + _guard(model, payload) body = json.dumps(payload).encode("utf-8") request = urllib.request.Request( f"{host.rstrip('/')}/api/generate", data=body, @@ -146,8 +172,9 @@ def embed(text: str, *, model: str = DEFAULT_EMBED_MODEL, reached, so the caller can say so instead of matching on a fragment silently. """ - body = json.dumps({"model": model, "input": text, - "truncate": True}).encode("utf-8") + payload = {"model": model, "input": text, "truncate": True} + _guard(model, payload) + body = json.dumps(payload).encode("utf-8") request = urllib.request.Request( f"{host.rstrip('/')}/api/embed", data=body, headers={"Content-Type": "application/json"}) diff --git a/skill_plus_plus/memory.py b/skill_plus_plus/memory.py new file mode 100644 index 0000000..1cfcad6 --- /dev/null +++ b/skill_plus_plus/memory.py @@ -0,0 +1,568 @@ +"""Keep the local model from pushing the machine into swap. + +The local model is the one large thing Skill++ runs. Measured on an 18 GB Mac +(docs/usage.md, "Memory"): loading gemma4:e4b and nomic-embed-text took 12.8 +GB of available memory (a fold with gemma4:e4b-it-qat, the default after it, +about 7), and a fold with other apps open sat at 92 % used, macOS pressure at +warning, with 4.7 GB swapped out in 100 seconds. Nothing leaked: a +one-word question locks as much as a whole session. So the fix is not in how a +session is read but in when the models may load, and how long they stay. + +Three rules, in the order they act: + +* **Before loading.** A model that is not loaded yet loads only if what the + models need still leaves `config.memory_reserve_gb` free. Otherwise nothing + loads and the caller holds the session. Folding later costs nothing. +* **While folding.** A watchdog thread reads available memory twice a second. + Below the reserve, or at critical pressure, it trips: no further model call + starts, and the models this guard loaded are unloaded at once. +* **After folding.** Whatever this guard loaded is unloaded when the guard is + released, not five minutes later when Ollama's own `keep_alive` would. + +What the models need is measured on each computer: estimated from their size +on disk the first time, then the largest drop in available memory a fold has +seen here. Only models this guard loaded are ever unloaded; a model someone else +had loaded stays loaded. + +Every reader answers `None` when it cannot tell, and an unknown reading turns +the guard off rather than holding every session forever. +""" + +from __future__ import annotations + +import json +import os +import re +import socket +import subprocess +import sys +import threading +import time +import urllib.error +import urllib.request +from collections.abc import Iterable, Iterator +from contextlib import contextmanager +from pathlib import Path + +GB = 1024 ** 3 + +# Size on disk times this is the first guess at what loading takes. Measured in +# bytes on the calibration run: 9.88e9 on disk (gemma4:e4b and nomic-embed-text, +# as Ollama lists them) took 13.7e9 of available memory, which is 12.8 GiB: the +# weights plus the buffers and cache that loading brings along. It guesses high +# for gemma4:e4b-it-qat: 6.42e9 on disk, 8.3 GiB guessed, 7.0 GiB measured. The +# guess only decides the first fold of a model set on a computer; from then on +# the measured figure does (`need_bytes`). +NEED_FACTOR = 1.39 +# How long `unload` waits for Ollama to stop listing a model. It answers the +# unload at once and lets go a moment later; a check made in between counted a +# model that was leaving as loaded by someone else, and never unloaded it again. +_UNLOAD_WAIT_SECONDS = 15.0 +# After the watchdog trips, how long it keeps watching for a model that was +# still loading. Measured: a trip 4.3 s into a load found nothing listed to +# unload, and the load finished anyway, 9.3 GB that stayed for `KEEP_ALIVE`. +# A load of gemma4:e4b takes 10-14 s, so 20 s sees it land and unloads it. +_LINGER_SECONDS = 20.0 +# A drop larger than this multiple of the estimate is something else growing at +# the same time, a browser or a build, not the models. It is not stored. +_OUTLIER = 2.0 +# How often the watchdog reads memory. A load moves gigabytes in seconds. +WATCH_SECONDS = 0.5 +# macOS `kern.memorystatus_vm_pressure_level`: 1 normal, 2 warning, 4 critical. +_CRITICAL = 4 +# Linux PSI: every task stalled on memory for this share of the last ten +# seconds is a machine that is thrashing. +_PSI_FULL_CRITICAL = 10.0 +# How long a model this guard loaded stays after its last call if the process +# dies before `release()`. Ollama's own default is five minutes. +KEEP_ALIVE = "60s" +# How long after an unload the freed memory is read, for the notice. +_SETTLE_SECONDS = 1.0 + + +# -- readers ---------------------------------------------------------------- + +def _run(argv: list[str], timeout: float = 5.0) -> str | None: + try: + return subprocess.run(argv, capture_output=True, text=True, + timeout=timeout, check=False).stdout + except (OSError, subprocess.SubprocessError): + return None + + +def _sysctl(name: str) -> int | None: + out = _run(["sysctl", "-n", name]) + try: + return int(out.strip()) if out else None + except ValueError: + return None + + +def _meminfo() -> dict[str, int]: + try: + text = Path("/proc/meminfo").read_text(encoding="utf-8") + except OSError: + return {} + out = {} + for line in text.splitlines(): + m = re.match(r"(\w+):\s+(\d+)\s*kB", line) + if m: + out[m.group(1)] = int(m.group(2)) * 1024 + return out + + +def total_bytes() -> int | None: + if sys.platform == "darwin": + return _sysctl("hw.memsize") + if sys.platform.startswith("linux"): + return _meminfo().get("MemTotal") + return None + + +def available_bytes() -> int | None: + """Memory the system can hand out without swapping. Not "free". + + Free memory sits near zero on a Mac by design, because spare RAM holds + cache, and a check against it would never let a fold start: 0.3 GB free + while the kernel reported 81 % available. macOS says what it can hand out as + a percentage (`kern.memorystatus_level`); Linux as `MemAvailable`. + """ + if sys.platform == "darwin": + level, total = _sysctl("kern.memorystatus_level"), _sysctl("hw.memsize") + return total * level // 100 if level is not None and total else None + if sys.platform.startswith("linux"): + return _meminfo().get("MemAvailable") + return None + + +def pressure_critical() -> bool: + """Has the system itself declared a memory emergency?""" + if sys.platform == "darwin": + return _sysctl("kern.memorystatus_vm_pressure_level") == _CRITICAL + if sys.platform.startswith("linux"): + try: + text = Path("/proc/pressure/memory").read_text(encoding="utf-8") + except OSError: + return False + m = re.search(r"^full avg10=([\d.]+)", text, re.M) + return bool(m) and float(m.group(1)) >= _PSI_FULL_CRITICAL + return False + + +def idle_seconds() -> float | None: + """How long nobody has used the keyboard or mouse, or None if unknown.""" + if sys.platform == "darwin": + return mac_idle(_run(["ioreg", "-c", "IOHIDSystem", "-d", "4"])) + if sys.platform.startswith("linux"): + session = os.environ.get("XDG_SESSION_ID") + if session: + idle = loginctl_idle(_run(["loginctl", "show-session", session, + "-p", "IdleHint", "-p", "IdleSinceHint"])) + if idle is not None: + return idle + return xprintidle_idle(_run(["xprintidle"])) + return None + + +def mac_idle(text: str | None) -> float | None: + """`ioreg`'s `HIDIdleTime`, in nanoseconds, as seconds.""" + m = re.search(r'"HIDIdleTime"\s*=\s*(\d+)', text or "") + return int(m.group(1)) / 1e9 if m else None + + +def loginctl_idle(text: str | None, now: float | None = None) -> float | None: + """Seconds idle from `loginctl show-session`, or None if it does not say. + + `IdleSinceHint` is wall-clock microseconds. A desktop that never sets the + hint reports `IdleHint=no` forever; that reads as "not idle", which only + delays a fold, never forces one. + """ + values = dict(line.split("=", 1) for line in (text or "").splitlines() if "=" in line) + if "IdleHint" not in values: + return None + if values["IdleHint"].strip() != "yes": + return 0.0 + try: + since = int(values.get("IdleSinceHint", "0")) + except ValueError: + return None + if not since: + return None + return max(0.0, (time.time() if now is None else now) - since / 1e6) + + +def xprintidle_idle(text: str | None) -> float | None: + """`xprintidle` prints milliseconds.""" + try: + return int((text or "").strip()) / 1000 + except ValueError: + return None + + +# -- Ollama ----------------------------------------------------------------- + +def _ollama(config, path: str, payload: dict | None = None, + timeout: float = 10.0) -> dict | None: + url = f"{config.ollama_url.rstrip('/')}{path}" + data = json.dumps(payload).encode("utf-8") if payload is not None else None + request = urllib.request.Request(url, data=data, + headers={"Content-Type": "application/json"}) + try: + with urllib.request.urlopen(request, timeout=timeout) as response: + return json.loads(response.read().decode("utf-8")) + except (urllib.error.URLError, OSError, TimeoutError, ValueError): + return None + + +def model_name(name: str) -> str: + """`nomic-embed-text` and `nomic-embed-text:latest` are one model. + + Ollama lists the tag; the configuration usually leaves it off. + """ + name = str(name or "").strip() + return name if ":" in name.rsplit("/", 1)[-1] else f"{name}:latest" + + +def loaded_models(config) -> set[str] | None: + """What Ollama holds in memory now, or None if it cannot be asked.""" + reply = _ollama(config, "/api/ps") + if reply is None: + return None + return {model_name(m.get("name") or m.get("model") or "") + for m in reply.get("models") or []} + + +def model_sizes(config) -> dict[str, int]: + """Each installed model's size on disk, in bytes.""" + reply = _ollama(config, "/api/tags") or {} + return {model_name(m.get("name") or m.get("model") or ""): int(m.get("size") or 0) + for m in reply.get("models") or []} + + +def unload(config, models: Iterable[str], *, linger: float = 0.0) -> None: + """Drop *models* from memory, and return once Ollama has let them go. + + The documented way: `keep_alive` 0. `/api/generate` does it for an embedding + model too; it answers with `done_reason: "unload"` and generates nothing. + It answers before the model is gone, though, so this waits until `/api/ps` + stops listing it (see `_UNLOAD_WAIT_SECONDS`). + + *linger* keeps watching that long even once nothing is listed, and unloads + again whatever appears: a model still loading is not listed yet, ignores + the unload, and lands a few seconds later. + """ + names = {model_name(m) for m in models} + if not names: + return + + def ask(which): + for name in sorted(which): + _ollama(config, "/api/generate", {"model": name, "keep_alive": 0}, timeout=30.0) + + ask(names) + asked = time.time() + settle = asked + linger + deadline = settle + _UNLOAD_WAIT_SECONDS + while time.time() < deadline: + loaded = loaded_models(config) + if loaded is None: + return + listed = names & loaded + if not listed and time.time() >= settle: + return + if listed and time.time() - asked >= 2.0: + ask(listed) + asked = time.time() + time.sleep(0.25) + + +def fold_models(config) -> list[str]: + """The models a fold may load: the judge and namer, and the embedder.""" + models = [] + if config.judge_boundaries or config.name_candidates: + models.append(config.local_model) + if config.match_candidates: + models.append(config.embed_model) + return models + + +# -- what the models take on this computer ----------------------------------- + +def _load(config) -> dict: + try: + data = json.loads(config.memory_file.read_text(encoding="utf-8")) + except (OSError, json.JSONDecodeError): + return {} + return data if isinstance(data, dict) else {} + + +def _store(config, data: dict) -> None: + try: + config.root.mkdir(parents=True, exist_ok=True) + tmp = config.memory_file.with_suffix(".json.tmp") + tmp.write_text(json.dumps(data, indent=2, sort_keys=True), encoding="utf-8") + tmp.replace(config.memory_file) + except OSError: + pass + + +def _key(models: Iterable[str]) -> str: + return "+".join(sorted(model_name(m) for m in models)) + + +def need_bytes(config, models: Iterable[str], + sizes: dict[str, int] | None = None) -> tuple[int, bool]: + """What loading *models* takes on this computer, and whether it was measured.""" + models = [model_name(m) for m in models] + if not models: + return 0, True + stored = (_load(config).get("need", {}).get(socket.gethostname(), {}) + .get(_key(models))) + if isinstance(stored, int) and stored > 0: + return stored, True + sizes = model_sizes(config) if sizes is None else sizes + return int(sum(sizes.get(m, 0) for m in models) * NEED_FACTOR), False + + +def record_need(config, models: Iterable[str], drop: int, estimate: int) -> None: + """Keep the largest drop a load of *models* caused here, within reason.""" + models = list(models) + if not models or drop <= 0 or estimate <= 0 or drop > estimate * _OUTLIER: + return + data = _load(config) + here = data.setdefault("need", {}).setdefault(socket.gethostname(), {}) + key = _key(models) + here[key] = max(int(here.get(key) or 0), int(drop)) + _store(config, data) + + +def gb(n: int | None) -> str: + return "?" if n is None else f"{n / GB:.1f}" + + +def shortfall(config, models: Iterable[str] | None = None) -> str: + """Why a fold could not load its models now, or "" if it could. + + A dry run of the first rule for the idle waiter and "Fold now": it loads + nothing and starts no watchdog. + """ + return Guard(config).check(fold_models(config) if models is None else models) + + +def status(config) -> dict: + """RAM, free memory, what a fold needs here and the free memory it starts at.""" + need, measured = need_bytes(config, fold_models(config)) + reserve = int(config.memory_reserve_gb * GB) + return {"total": total_bytes(), "available": available_bytes(), "need": need, + "measured": measured, "reserve": reserve, "start_at": need + reserve} + + +# -- the guard -------------------------------------------------------------- + +class Guard: + """One fold's, sweep's or page request's use of the local models.""" + + def __init__(self, config, *, unload_on_release: bool = True) -> None: + self.config = config + self.reserve = int(config.memory_reserve_gb * GB) + self.unload_on_release = unload_on_release + self.enabled = bool(config.memory_guard) and available_bytes() is not None + self.ours: set[str] = set() # loaded by this guard, still loaded + self.approved: set[str] = set() # may be called without asking again + self.tripped = "" + self._loaded: set[str] = set() # everything this guard loaded + self._before: int | None = None + self._low: int | None = None + self._estimate = 0 + self._stop = threading.Event() + self._thread: threading.Thread | None = None + self._lock = threading.Lock() + + def _missing(self, models: Iterable[str]) -> tuple[set[str], set[str]] | None: + """The requested models not yet approved, and those of them not loaded.""" + wanted = {model_name(m) for m in models if m} - self.approved + if not wanted: + return set(), set() + loaded = loaded_models(self.config) + if loaded is None: + # Ollama cannot be asked, so the call will fail on its own and the + # session is held as offline. Nothing here can make that worse. + return None + return wanted, wanted - loaded + + def check(self, models: Iterable[str]) -> str: + """Why *models* may not load now, or "". Changes nothing.""" + if not self.enabled: + return "" + found = self._missing(models) + if not found or not found[1]: + return "" + return self._room(found[1], model_sizes(self.config))[0] + + def _room(self, missing: set[str], sizes: dict[str, int]) -> tuple[str, int | None]: + need, _ = need_bytes(self.config, missing, sizes) + available = available_bytes() + if available is None or available - need >= self.reserve: + return "", available + return (f"needs {gb(need + self.reserve)} GB of free memory, " + f"{gb(available)} GB free now"), available + + def admit(self, models: Iterable[str]) -> str: + """Approve *models* for this guard, or say why not. Starts the watchdog.""" + if not self.enabled: + return "" + if self.tripped: + return self.tripped + found = self._missing(models) + if found is None: + return "" + wanted, missing = found + if missing: + sizes = model_sizes(self.config) + reason, available = self._room(missing, sizes) + if reason: + return reason + with self._lock: + if self._before is None and available is not None: + self._before = self._low = available + self._estimate += int(sum(sizes.get(m, 0) for m in missing) * NEED_FACTOR) + self.ours |= missing + self._loaded |= missing + self.approved |= wanted + self._watch() + return "" + + def before_call(self, model: str) -> None: + """Raise `MemoryShort` unless *model* may be called now.""" + from .local import MemoryShort + reason = self.tripped or self.admit([model]) + if reason: + raise MemoryShort(reason) + + def owns(self, model: str) -> bool: + return model_name(model) in self.ours + + def _watch(self) -> None: + if self._thread is not None or not self.ours: + return + self._thread = threading.Thread(target=self._run, daemon=True, + name="skill-plus-plus-memory") + self._thread.start() + + def _run(self) -> None: + while not self._stop.wait(WATCH_SECONDS): + available = available_bytes() + if available is None: + continue + with self._lock: + self._low = available if self._low is None else min(self._low, available) + if available < self.reserve: + self.trip(f"free memory fell below {gb(self.reserve)} GB") + return + if pressure_critical(): + self.trip("memory pressure turned critical") + return + + def trip(self, reason: str) -> None: + """Stop: no further model call, and unload what this guard loaded, now. + + Everything it loaded, including a model still loading when this + tripped: that one is not listed yet, so the unload lingers until it + lands (`_LINGER_SECONDS`). + """ + with self._lock: + if self.tripped: + return + self.tripped = reason + loaded = sorted(self._loaded) + before = available_bytes() + unload(self.config, loaded, linger=_LINGER_SECONDS if loaded else 0.0) + time.sleep(_SETTLE_SECONDS) + after = available_bytes() + freed = after - before if before is not None and after is not None else None + from .notify import stopped + stopped(self.config, reason, freed if freed and freed > 0 else None) + + def release(self) -> None: + """End of the fold: remember what the load took, and unload our models. + + A trip in progress finishes first, so a worker cannot exit while the + watchdog is still unloading. + """ + self._stop.set() + if self._thread is not None: + self._thread.join() + with self._lock: + loaded, before, low = set(self._loaded), self._before, self._low + estimate = self._estimate + if loaded and before is not None and low is not None: + record_need(self.config, loaded, before - low, estimate) + if loaded and (self.unload_on_release or self.tripped): + unload(self.config, sorted(loaded)) + + +_active: Guard | None = None +_active_lock = threading.Lock() + +# One Skill++ process uses the models at a time. Two folds at once, two chats +# closed together or a session start's sweep beside a session end, each loaded +# the models for itself, and each unloaded them from under the other. Measured +# the same way: a fold that started while the one before was still unloading +# counted the leaving model as someone else's and left it loaded. +_MODELS_LOCK_SECONDS = 3600.0 +_MODELS_LOCK_POLL = 2.0 +BUSY = "is already folding a session; try again in a minute" + + +def active() -> Guard | None: + """The guard `local.ask` and `local.embed` answer to, if one is open.""" + return _active + + +def _take_models_lock(path: Path, wait: bool) -> bool: + from .capture import _acquire_lock + deadline = time.time() + _MODELS_LOCK_SECONDS + while True: + if _acquire_lock(path, _MODELS_LOCK_SECONDS, "models"): + return True + if not wait or time.time() > deadline: + return False + time.sleep(_MODELS_LOCK_POLL) + + +@contextmanager +def guarded(config, *, wait: bool = True, **kwargs) -> Iterator[Guard]: + """The guard for one fold, sweep or page request. + + Nested uses share the outermost: a `fold-pending` sweep keeps the models + loaded from one session to the next and unloads them once, at the end. + Across processes, one guard at a time holds the models: a fold waits for + the one before it, and with *wait* off (the review page, which must not + hang) the guard refuses every call instead. + """ + global _active + with _active_lock: + outer = _active + if outer is None: + _active = guard = Guard(config, **kwargs) + if outer is not None: + yield outer + return + lock = config.root / "models.lock" + held = False + try: + if guard.enabled: + config.root.mkdir(parents=True, exist_ok=True) + held = _take_models_lock(lock, wait) + if not held: + guard.tripped = BUSY + yield guard + finally: + with _active_lock: + _active = None + guard.release() + if held: + try: + lock.unlink() + except OSError: + pass diff --git a/skill_plus_plus/notify.py b/skill_plus_plus/notify.py new file mode 100644 index 0000000..c60a5b5 --- /dev/null +++ b/skill_plus_plus/notify.py @@ -0,0 +1,81 @@ +"""Tell the person when the memory guard holds or stops a fold. + +The fold runs detached after the chat has ended, so a message in the session +would reach nobody, and a SessionStart hook's output goes to the model, not to +the person (Claude Code hooks reference). The operating system's own +notifications are the one place they will see it: Notification Center through +`osascript` on macOS, `notify-send` on Linux. Anywhere else, and whenever those +fail, the notice reaches the log only. + +Nothing here raises. A missing notifier costs the notice, never the fold. +""" + +from __future__ import annotations + +import shutil +import subprocess +import sys +import time + +TITLE = "Skill++" +# "Waiting" repeats while sessions wait, so it is sent at most this often. +# "Stopped" is sent every time: it only happens when memory ran short. +WAITING_EVERY_SECONDS = 3600 + + +def command(message: str, platform: str | None = None) -> list[str] | None: + """The command that shows *message* on this system, or None.""" + platform = platform or sys.platform + if platform == "darwin": + return ["osascript", "-e", + f"display notification {_apple(message)} with title {_apple(TITLE)}"] + if platform.startswith("linux") and shutil.which("notify-send"): + return ["notify-send", "--app-name", TITLE, TITLE, message] + return None + + +def _apple(text: str) -> str: + """An AppleScript string literal.""" + return '"' + str(text).replace("\\", "\\\\").replace('"', '\\"') + '"' + + +def send(config, message: str, *, kind: str, every: float = 0.0) -> bool: + """Show *message*, at most once per *every* seconds for this *kind*.""" + from .capture import log_error + from .memory import _load, _store + + log_error(config, f"notice ({kind}): {message}") + if not config.notify: + return False + data = _load(config) + last = (data.get("notified") or {}).get(kind) + if every and isinstance(last, (int, float)) and time.time() - last < every: + return False + argv = command(message) + if not argv: + return False + try: + subprocess.run(argv, capture_output=True, timeout=10, check=False) + except (OSError, subprocess.SubprocessError): + return False + data.setdefault("notified", {})[kind] = time.time() + _store(config, data) + return True + + +def waiting(config, reason: str) -> bool: + """A fold was held: the models do not fit in memory now.""" + return send(config, + f"Skill++ is waiting to fold your session: it {reason}. " + "It folds on its own once your computer is idle.", + kind="waiting", every=WAITING_EVERY_SECONDS) + + +def stopped(config, reason: str, freed: int | None) -> bool: + """The watchdog stopped a fold and unloaded the models.""" + from .memory import gb + what = f"stopped and freed {gb(freed)} GB" if freed else "stopped its local model" + return send(config, + f"Skill++ {what}: {reason}. " + "Your session is kept and folded later.", + kind="stopped") diff --git a/skill_plus_plus/web.py b/skill_plus_plus/web.py index 43dd0bb..c9dbad7 100644 --- a/skill_plus_plus/web.py +++ b/skill_plus_plus/web.py @@ -381,8 +381,13 @@ def summarise(config: Config, entry_id: str) -> dict: before capture named them β€” then cached per entry and keyed on step count, so the page never waits on a model twice and a candidate that grows is described again. + + Under the memory guard, like a fold: the model loads only if it fits. It is + left loaded for `memory.KEEP_ALIVE` rather than unloaded, because opening + one row is usually followed by opening the next. """ - from .local import LocalModelUnavailable + from .local import LocalModelUnavailable, MemoryShort + from .memory import BUSY, guarded from .summary import name_and_sentence, store_summary entry = Ledger(config).get(entry_id) @@ -393,7 +398,11 @@ def summarise(config: Config, entry_id: str) -> dict: return {"ok": True, "summary": cached} try: - _, text = name_and_sentence(config, entry) + with guarded(config, wait=False, unload_on_release=False): + _, text = name_and_sentence(config, entry) + except MemoryShort as exc: + return {"ok": False, "error": f"Skill++ {exc}" if str(exc) == BUSY + else f"not enough free memory: Skill++ {exc}"} except LocalModelUnavailable as exc: return {"ok": False, "error": f"no local model: {exc}"} if not text: @@ -427,6 +436,10 @@ def project_list(rows: list[dict], drafts: list[dict]) -> list[dict]: def collect_state(config: Config) -> dict: """The rows the page lists. Reads files only: no model, no agent. + One exception, only while sessions are held for memory: the banner reads + the free memory, and the models' sizes from Ollama if what they take here + was never measured. Neither loads a model. + Every candidate, promoted and dismissed entry, most-seen first. `ready` splits the page: recognized often enough to decide on, or still collecting. """ @@ -457,7 +470,46 @@ def collect_state(config: Config) -> dict: projects = project_list(rows, drafts) return {"threshold": threshold, "ttl": config.candidate_ttl_days, "rows": rows, "drafts": drafts, "projects": projects, - "capture": capture_status([p["path"] for p in projects])} + "capture": capture_status([p["path"] for p in projects]), + "memory": memory_waiting(config)} + + +def memory_waiting(config: Config) -> dict: + """How many sessions the memory guard is holding, for the page's banner. + + Counted from the session files. Only when some are waiting does it also + say from what free memory a fold starts, and how much is free now. + """ + waiting = 0 + for path in config.sessions_dir.glob("*.json"): + try: + held = json.loads(path.read_text(encoding="utf-8")).get("held") + except (OSError, json.JSONDecodeError): + continue + waiting += bool(isinstance(held, dict) and held.get("memory")) + if not waiting: + return {"waiting": 0} + from .memory import gb, status + s = status(config) + return {"waiting": waiting, "start_at_gb": gb(s["start_at"]), + "free_gb": gb(s["available"])} + + +def fold_now(config: Config) -> dict: + """"Fold now" on the banner: `skill-plus-plus fold-pending --now`, if it fits. + + The memory rule still holds: when the models do not fit, nothing loads and + the page says how much memory is missing. + """ + from .memory import shortfall + short = shortfall(config) + if short: + return {"ok": False, + "error": f"Not yet: Skill++ {short}. Close a few apps and try again."} + proc = _run(config, "fold-pending", "--now") + if proc.returncode != 0: + return {"ok": False, "error": _tail(proc.stderr or proc.stdout)} + return {"ok": True} def capture_status(project_paths: list[str], home: Path | None = None) -> dict: @@ -1035,6 +1087,7 @@ def make_handler(config: Config, skills_dir: Path | None = None): str(p.get("instruction", ""))), "/api/answer": lambda p: answer_questions(config, str(p.get("id", "")), p.get("answers") or []), + "/api/fold-now": lambda p: fold_now(config), } class Handler(BaseHTTPRequestHandler): @@ -1211,6 +1264,9 @@ def serve(config: Config, skills_dir: Path | None = None, port: int = 8765, .warn{margin:0 0 20px;padding:12px 16px;border:1px solid var(--noline);background:var(--nobg); border-radius:8px;color:var(--fg);font-size:13px;line-height:1.6} .warn strong{color:var(--no)} + .warn.wait{border-color:rgba(251,191,36,.4);background:rgba(251,191,36,.07)} + .warn.wait strong{color:#fbbf24} + .warn button.fold-now{margin-left:8px;color:#0c0d10;border-color:#fbbf24;background:#fbbf24;font-weight:600} .cand{background:var(--panel);border:1px solid var(--line);border-radius:8px;margin-bottom:8px} .cand.ready{border-left:3px solid #fbbf24;background:linear-gradient(90deg,rgba(251,191,36,.07),var(--panel) 40%)} .cand.accepted{border-left:3px solid var(--go);background:linear-gradient(90deg,rgba(56,189,248,.07),var(--panel) 40%)} @@ -1394,6 +1450,17 @@ def serve(config: Config, skills_dir: Path | None = None, port: int = 8765, Run skill-plus-plus install --user --apply, then start a new Claude Code session; skill-plus-plus doctor checks it. Installing β†’`; } +// Sessions the memory guard is holding: kept, not lost, and folded once the +// computer is idle and the models fit, or now if memory allows. +function memoryWarning(){ + const m = S.memory || {waiting: 0}; + if(!m.waiting) return ""; + const n = m.waiting === 1 ? "1 session is" : `${m.waiting} sessions are`; + return `
${n} waiting for free memory. + Skill++ folds only when ${esc(m.start_at_gb)} GB are free, so your computer never runs short (${esc(m.free_gb)} GB free now). + They fold on their own once your computer is idle. +
`; +} function emptyCandidates(){ const watching = (S.capture || {}).state === "wired" ? `

Skill++ is watching your Claude Code sessions. A procedure shows up here once you've repeated it ${S.threshold}Γ—.

` : ""; @@ -1941,7 +2008,7 @@ def serve(config: Config, skills_dir: Path | None = None, port: int = 8765, const promoted = S.rows.filter(r => !["undecided", "collecting", "dismissed"].includes(r.state) && !inDrafts(r)); const open_ = S.rows.filter(r => ["undecided", "collecting"].includes(r.state)); const ready = open_.filter(r => r.ready), collecting = open_.filter(r => !r.ready); - list.innerHTML = `${captureWarning()} + list.innerHTML = `${captureWarning()}${memoryWarning()} ${promoted.length ? `

Promoted

Candidates you promoted. Draft a skill from them; the draft appears in the Drafts tab.
${promoted.map(card).join("")}` : ""}

Still collecting

Work Skill++ saw you repeat. Once something is seen ${S.threshold}Γ—, you can promote or ignore it.
diff --git a/tests/test_skill_plus_plus.py b/tests/test_skill_plus_plus.py index 48514fd..db8b6e2 100644 --- a/tests/test_skill_plus_plus.py +++ b/tests/test_skill_plus_plus.py @@ -158,6 +158,16 @@ def setUpModule() -> None: def _no_pull(config, name): raise AssertionError(f"a test tried to pull {name}") cli._pull_model = _no_pull + # And for the memory guard, which reads this machine's memory and asks + # Ollama what is loaded before every fold. Left on, a test would pass or + # hold depending on what else the machine running it has open. The tests + # of the guard turn it on themselves, with memory readings they script. + # Notifications likewise: a test run must not post to Notification Center. + global _REAL_ENV + _REAL_ENV = {k: os.environ.get(k) for k in ("SKILL_PLUS_PLUS_MEMORY_GUARD", + "SKILL_PLUS_PLUS_NOTIFY")} + os.environ["SKILL_PLUS_PLUS_MEMORY_GUARD"] = "0" + os.environ["SKILL_PLUS_PLUS_NOTIFY"] = "0" def tearDownModule() -> None: @@ -170,6 +180,11 @@ def tearDownModule() -> None: capture._name_from_model = _REAL_NAME import skill_plus_plus.cli as cli cli._available_models, cli._pull_model = _REAL_MODELS, _REAL_PULL + for key, value in _REAL_ENV.items(): + if value is None: + os.environ.pop(key, None) + else: + os.environ[key] = value def _marker_judge(config, session, verdict=is_marker): @@ -7099,3 +7114,497 @@ def test_merge_covers_a_candidate_a_skill_already_does(self): self.assertEqual(covered.status, STATUS_COVERED) self.assertNotIn("cand1", [c.id for c in led.candidates()], "work a skill already does is not a proposal") + + +GB = 1024 ** 3 + + +class _ScriptedMemory: + """This machine's memory and Ollama, as a memory-guard test scripts them. + + `available` is what the kernel says it can hand out, and `load` takes a + model's cost from it the way a real load does. Nothing reaches Ollama: + which models are loaded, their sizes and every unload live here, and the + two notices are recorded rather than posted. + """ + + def __init__(self, test, *, available_gb, total_gb=18.0, loaded=()): + from skill_plus_plus import memory, notify + self.memory = memory + self.available = int(available_gb * GB) + self.total = int(total_gb * GB) + self.loaded = {memory.model_name(m) for m in loaded} + self.sizes = {"gemma4:e4b": int(9.6 * GB), "nomic-embed-text:latest": int(0.27 * GB)} + self.unloaded: list[str] = [] + self.critical = False + self.notices: list[str] = [] + scripted = { + "available_bytes": lambda: self.available, + "total_bytes": lambda: self.total, + "pressure_critical": lambda: self.critical, + "loaded_models": lambda config: set(self.loaded), + "model_sizes": lambda config: dict(self.sizes), + "unload": self._unload, + "WATCH_SECONDS": 0.01, + "_SETTLE_SECONDS": 0.0, + } + for name, value in scripted.items(): + patch = mock.patch.object(memory, name, value) + patch.start() + test.addCleanup(patch.stop) + for kind in ("waiting", "stopped"): + patch = mock.patch.object( + notify, kind, lambda config, *args, _kind=kind: self.notices.append(_kind) or True) + patch.start() + test.addCleanup(patch.stop) + + def _unload(self, config, models, **kwargs): + for name in models: + self.unloaded.append(name) + self.loaded.discard(name) + + def load(self, model, cost_gb): + """What Ollama does on a model's first call: it is in memory now.""" + self.loaded.add(self.memory.model_name(model)) + self.available -= int(cost_gb * GB) + + +class TestMemoryGuard(TempRoot): + """The local model never takes a machine below 2 GB of free memory. + + Measured on an 18 GB Mac: gemma4:e4b and the embedder took 12.8 GB of + available memory, and a fold with other apps open ran at 92 % used with + 4.7 GB swapped out in 100 s. Nothing leaked; the model is simply large. So + the models load only when they fit, a watchdog unloads them if memory runs + short mid-fold, and a session that does not fit waits to be folded later. + """ + + FOLD = ["gemma4:e4b", "nomic-embed-text"] + TWO_TASKS = ("npm test", "git commit -m 'fix'", "cargo build", "git commit -m 'feat'") + + def setUp(self) -> None: + # The scripted machine runs gemma4:e4b, whatever the default model is: + # the sizes and the figures these tests expect are that model's. + pinned = mock.patch.dict(os.environ, {"SKILL_PLUS_PLUS_LOCAL_MODEL": self.FOLD[0]}) + pinned.start() + self.addCleanup(pinned.stop) + super().setUp() + self.config.memory_guard = True + + def _wait(self, condition, seconds=2.0): + deadline = time.time() + seconds + while not condition(): + if time.time() > deadline: + self.fail("timed out waiting") + time.sleep(0.01) + + def _session_on_disk(self, sid="mem", commands=TWO_TASKS): + handle_prompt(self.config, {"session_id": sid, "cwd": "/r", "prompt": "ship it"}) + for command in commands: + handle_tool(self.config, {"session_id": sid, "cwd": "/r", "tool_name": "Bash", + "tool_input": {"command": command}}) + + def _on_disk(self, sid="mem"): + from skill_plus_plus.capture import _session_file + return json.loads(_session_file(self.config, sid).read_text(encoding="utf-8")) + + # -- the check before loading -------------------------------------------- + + def test_models_that_do_not_fit_are_not_loaded(self): + from skill_plus_plus.memory import Guard + _ScriptedMemory(self, available_gb=9.1) + guard = Guard(self.config) + # (9.6 + 0.27) GiB on disk x 1.39 = 13.7 GiB, plus the 2 GiB reserve. + self.assertEqual(guard.admit(self.FOLD), + "needs 15.7 GB of free memory, 9.1 GB free now") + self.assertEqual(guard.ours, set()) + + def test_models_already_loaded_need_no_room_and_are_never_unloaded(self): + from skill_plus_plus.memory import Guard + mem = _ScriptedMemory(self, available_gb=1.0, + loaded=["gemma4:e4b", "nomic-embed-text:latest"]) + guard = Guard(self.config) + self.assertEqual(guard.admit(self.FOLD), "") + guard.release() + self.assertEqual(mem.unloaded, [], "someone else loaded them") + + # -- the watchdog --------------------------------------------------------- + + def test_the_watchdog_trips_below_the_reserve_and_unloads_only_its_own(self): + from skill_plus_plus.memory import Guard + mem = _ScriptedMemory(self, available_gb=16.0, loaded=["nomic-embed-text:latest"]) + guard = Guard(self.config) + self.assertEqual(guard.admit(self.FOLD), "") + mem.load("gemma4:e4b", 12.6) # 3.4 GB left: still fine + time.sleep(0.05) + self.assertEqual(guard.tripped, "") + mem.available = int(1.5 * GB) # something else grows + self._wait(lambda: guard.tripped) + guard.release() + self.assertEqual(guard.tripped, "free memory fell below 2.0 GB") + self.assertEqual(set(mem.unloaded), {"gemma4:e4b"}, "not the one someone else loaded") + self.assertEqual(mem.notices, ["stopped"]) + + def test_critical_pressure_trips_the_watchdog_too(self): + from skill_plus_plus.memory import Guard + mem = _ScriptedMemory(self, available_gb=16.0) + guard = Guard(self.config) + guard.admit(self.FOLD) + mem.critical = True + self._wait(lambda: guard.tripped) + guard.release() + self.assertEqual(guard.tripped, "memory pressure turned critical") + + def test_a_tripped_guard_makes_no_call(self): + from skill_plus_plus import local, memory + _ScriptedMemory(self, available_gb=16.0) + with memory.guarded(self.config) as guard: + guard.tripped = "free memory fell below 2.0 GB" + with mock.patch.object(local.urllib.request, "urlopen", + side_effect=AssertionError("Ollama was called")): + with self.assertRaises(local.MemoryShort): + local.ask("gemma4:e4b", "hi") + with self.assertRaises(local.MemoryShort): + local.embed("hi") + + def test_a_model_the_guard_loads_leaves_soon_and_is_unloaded_at_the_end(self): + import io + from skill_plus_plus import local, memory + mem = _ScriptedMemory(self, available_gb=16.0) + sent = [] + + def ollama(request, timeout=None): + sent.append(json.loads(request.data)) + return io.BytesIO(b'{"response": "yes"}') + + with memory.guarded(self.config): + with mock.patch.object(local.urllib.request, "urlopen", ollama): + self.assertEqual(local.ask("gemma4:e4b", "hi"), "yes") + self.assertEqual(sent[0]["keep_alive"], memory.KEEP_ALIVE) + self.assertEqual(mem.unloaded, ["gemma4:e4b"]) + + def test_a_model_someone_else_loaded_keeps_its_own_keep_alive(self): + import io + from skill_plus_plus import local, memory + mem = _ScriptedMemory(self, available_gb=16.0, loaded=["gemma4:e4b"]) + sent = [] + + def ollama(request, timeout=None): + sent.append(json.loads(request.data)) + return io.BytesIO(b'{"response": "yes"}') + + with memory.guarded(self.config): + with mock.patch.object(local.urllib.request, "urlopen", ollama): + local.ask("gemma4:e4b", "hi") + self.assertNotIn("keep_alive", sent[0]) + self.assertEqual(mem.unloaded, []) + + # -- what the models take on this computer --------------------------------- + + def test_what_a_load_took_is_remembered_and_the_highest_kept(self): + from skill_plus_plus.memory import Guard, need_bytes + mem = _ScriptedMemory(self, available_gb=16.0) + for cost in (12.8, 12.1): + mem.available, mem.loaded = int(16 * GB), set() + guard = Guard(self.config) + guard.admit(self.FOLD) + mem.load("gemma4:e4b", cost) + self._wait(lambda: guard._low == mem.available) + guard.release() + need, measured = need_bytes(self.config, self.FOLD) + self.assertTrue(measured) + self.assertAlmostEqual(need / GB, 12.8, places=1) + + def test_a_drop_far_beyond_the_estimate_is_not_kept(self): + """Something else grew at the same time; it is not what the models took.""" + from skill_plus_plus.memory import Guard, need_bytes + mem = _ScriptedMemory(self, available_gb=40.0, total_gb=64.0) + guard = Guard(self.config) + guard.admit(self.FOLD) + mem.load("gemma4:e4b", 30.0) + self._wait(lambda: guard._low == mem.available) + guard.release() + self.assertFalse(need_bytes(self.config, self.FOLD)[1]) + + # -- a session's fold ------------------------------------------------------ + + def test_a_session_that_does_not_fit_is_held_and_nothing_is_judged(self): + mem = _ScriptedMemory(self, available_gb=9.1) + self._session_on_disk() + result = handle_session_end(self.config, {"session_id": "mem"}) + self.assertTrue(result["memory"]) + held = self._on_disk()["held"] + self.assertTrue(held["memory"]) + self.assertEqual(held["reason"], + "memory: needs 15.7 GB of free memory, 9.1 GB free now") + self.assertEqual(self.judged, [], "the judge was asked with nothing loaded") + self.assertEqual(list(Ledger(self.config).all()), []) + self.assertEqual(mem.notices, ["waiting"]) + + def test_a_stop_while_judging_drops_every_verdict(self): + """Half-judged is worse than unjudged: the rest would read as one task.""" + import skill_plus_plus.boundary as boundary + from skill_plus_plus import memory + mem = _ScriptedMemory(self, available_gb=16.0) + + def judge_then_trip(config, session): + _marker_judge(config, session) + memory.active().trip("free memory fell below 2.0 GB") + + self._session_on_disk() + with mock.patch.object(boundary, "judge_session", judge_then_trip): + result = handle_session_end(self.config, {"session_id": "mem"}) + self.assertTrue(result["memory"]) + self.assertFalse([s for s in self._on_disk()["steps"] if "end" in s]) + self.assertEqual(list(Ledger(self.config).all()), []) + self.assertEqual(mem.notices, ["stopped"], "one notice, not a second") + + def test_a_stop_between_episodes_keeps_what_was_banked_and_resumes(self): + from skill_plus_plus import capture, memory + from skill_plus_plus.local import MemoryShort + _ScriptedMemory(self, available_gb=16.0) + real, calls = capture._fold_steps, [] + + def second_trips(config, session, steps, **kwargs): + calls.append(1) + if len(calls) == 2: + memory.active().trip("free memory fell below 2.0 GB") + raise MemoryShort("free memory fell below 2.0 GB") + return real(config, session, steps, **kwargs) + + self._session_on_disk() + with mock.patch.object(capture, "_fold_steps", second_trips): + first = handle_session_end(self.config, {"session_id": "mem"}) + self.assertTrue(first["memory"]) + doc = self._on_disk() + self.assertEqual(doc["folded"], [0]) + self.assertTrue([s for s in doc["steps"] if "end" in s], + "folding had begun: the verdicts stay") + + again = handle_session_end(self.config, {"session_id": "mem"}) + self.assertNotEqual(again["status"], "offline") + self.assertEqual(sum(e.occurrences for e in Ledger(self.config).all()), 2, + "each task counted once, the banked one not twice") + + # -- folding later --------------------------------------------------------- + + def _held_for_memory(self, sid, minutes_ago): + from datetime import datetime, timedelta, timezone + from skill_plus_plus.capture import _session_file + at = (datetime.now(timezone.utc) - timedelta(minutes=minutes_ago)).isoformat( + timespec="seconds") + steps = [{"tool": "UserPrompt", "input": {"text": f"ship {sid}"}}, + {"tool": "Bash", "input": {"command": "npm test"}}, + {"tool": "Bash", "input": {"command": "git commit -m x"}}] + _session_file(self.config, sid).write_text(json.dumps( + {"session_id": sid, "cwd": "/r", "prompts": [], "steps": steps, + "held": {"at": at, "reason": "memory: needs 15.7 GB", "memory": True}}), + encoding="utf-8") + + def test_a_session_held_for_memory_is_retried_every_ten_minutes(self): + from datetime import datetime, timedelta, timezone + from skill_plus_plus.capture import _is_pending + + def held(minutes_ago, memory=True): + at = datetime.now(timezone.utc) - timedelta(minutes=minutes_ago) + return {"held": {"at": at.isoformat(timespec="seconds"), "memory": memory}} + + self.assertFalse(_is_pending(held(1), 0.0, 12.0)) + self.assertTrue(_is_pending(held(1), 0.0, 12.0, now=True)) + self.assertTrue(_is_pending(held(11), 0.0, 12.0)) + self.assertTrue(_is_pending(held(1, memory=False), 0.0, 12.0), + "a session no model answered waits for nothing") + + def test_the_sweep_stops_at_the_first_session_that_does_not_fit(self): + from skill_plus_plus.capture import fold_pending + mem = _ScriptedMemory(self, available_gb=9.1) + for sid in ("a1", "b2", "c3"): + self._held_for_memory(sid, minutes_ago=11) + results = fold_pending(self.config) + self.assertEqual([r["session"] for r in results], ["a1"]) + self.assertTrue(results[0]["memory"]) + self.assertEqual(mem.notices, ["waiting"]) + + def test_the_idle_waiter_folds_once_nobody_is_at_the_keyboard_and_it_fits(self): + from skill_plus_plus import memory + from skill_plus_plus.capture import _session_file, wait_for_memory + mem = _ScriptedMemory(self, available_gb=9.1) + self._held_for_memory("w1", minutes_ago=1) + polls, now = [], [0.0] + + def idle(): + polls.append(1) + if len(polls) == 3: + mem.available = int(16 * GB) # apps closed + return 30.0 if len(polls) == 1 else 600.0 # busy, then idle + + with mock.patch.object(memory, "idle_seconds", idle): + wait_for_memory(self.config, sleep=lambda s: now.__setitem__(0, now[0] + s), + clock=lambda: now[0]) + self.assertFalse(_session_file(self.config, "w1").exists()) + self.assertEqual(len(list(Ledger(self.config).all())), 1) + self.assertEqual(len(polls), 3, "busy, idle but short, then folded") + + def test_only_one_idle_waiter(self): + import socket + from skill_plus_plus.capture import wait_for_memory + _ScriptedMemory(self, available_gb=16.0) + self._held_for_memory("w1", minutes_ago=1) + (self.config.root / "memory-wait.lock").write_text( + json.dumps({"pid": os.getpid(), "host": socket.gethostname()})) + self.assertEqual(wait_for_memory(self.config, sleep=lambda s: self.fail("waited")), []) + + def test_idle_time_is_read_on_each_system(self): + from skill_plus_plus.memory import loginctl_idle, mac_idle, xprintidle_idle + self.assertEqual(mac_idle(' | | "HIDIdleTime" = 125000000000\n'), 125.0) + self.assertIsNone(mac_idle("no such key")) + self.assertEqual(loginctl_idle("IdleHint=no\nIdleSinceHint=0\n"), 0.0) + self.assertAlmostEqual( + loginctl_idle("IdleHint=yes\nIdleSinceHint=1000000000\n", now=1600.0), 600.0) + self.assertIsNone(loginctl_idle("")) + self.assertEqual(xprintidle_idle("45000\n"), 45.0) + self.assertIsNone(xprintidle_idle("xprintidle: command not found")) + + def test_fold_now_says_how_much_memory_is_missing(self): + from skill_plus_plus.web import fold_now + _ScriptedMemory(self, available_gb=9.1) + with mock.patch("skill_plus_plus.web._run", side_effect=AssertionError("folded anyway")): + result = fold_now(self.config) + self.assertFalse(result["ok"]) + self.assertIn("needs 15.7 GB of free memory, 9.1 GB free now", result["error"]) + + def test_fold_now_folds_when_it_fits(self): + import subprocess + from skill_plus_plus.web import fold_now + _ScriptedMemory(self, available_gb=16.0) + ran = [] + with mock.patch("skill_plus_plus.web._run", side_effect=lambda config, *args: ran.append( + args) or subprocess.CompletedProcess(args, 0, "", "")): + self.assertEqual(fold_now(self.config), {"ok": True}) + self.assertEqual(ran, [("fold-pending", "--now")]) + + def test_stats_counts_the_sessions_waiting_for_memory(self): + import argparse + import contextlib + import io + from skill_plus_plus.cli import cmd_stats + _ScriptedMemory(self, available_gb=9.1) + self._held_for_memory("w1", minutes_ago=1) + out = io.StringIO() + with contextlib.redirect_stdout(out): + cmd_stats(argparse.Namespace(root=self.config.root)) + self.assertIn("waiting 1 session(s) held for memory", out.getvalue()) + self.assertIn("a fold starts when 15.7 GB are free (9.1 GB now)", out.getvalue()) + + +class TestNotify(TempRoot): + """The notice reaches the person, or at least the log, and never raises.""" + + def setUp(self) -> None: + super().setUp() + self.config.notify = True + + def test_the_command_on_each_system(self): + from skill_plus_plus.notify import command + mac = command('say "hi"', platform="darwin") + self.assertEqual(mac[:2], ["osascript", "-e"]) + self.assertIn('display notification "say \\"hi\\""', mac[2]) + with mock.patch("skill_plus_plus.notify.shutil.which", return_value="/usr/bin/notify-send"): + self.assertEqual(command("hi", platform="linux")[0], "notify-send") + with mock.patch("skill_plus_plus.notify.shutil.which", return_value=None): + self.assertIsNone(command("hi", platform="linux")) + self.assertIsNone(command("hi", platform="win32")) + + def test_waiting_at_most_hourly_and_stopped_every_time(self): + from skill_plus_plus import notify + sent = [] + with mock.patch.object(notify, "command", return_value=["true"]), \ + mock.patch.object(notify.subprocess, "run", + side_effect=lambda argv, **kwargs: sent.append(argv)): + self.assertTrue(notify.waiting(self.config, "needs 15.7 GB of free memory")) + self.assertFalse(notify.waiting(self.config, "needs 15.7 GB of free memory")) + self.assertTrue(notify.stopped(self.config, "free memory fell below 2.0 GB", 12 * GB)) + self.assertTrue(notify.stopped(self.config, "free memory fell below 2.0 GB", None)) + self.assertEqual(len(sent), 3) + + def test_without_a_notifier_the_notice_goes_to_the_log(self): + from skill_plus_plus import notify + with mock.patch.object(notify, "command", return_value=None): + self.assertFalse(notify.stopped(self.config, "free memory fell below 2.0 GB", None)) + self.assertIn("notice (stopped)", self.config.log_file.read_text(encoding="utf-8")) + + def test_off_means_log_only(self): + from skill_plus_plus import notify + self.config.notify = False + with mock.patch.object(notify.subprocess, "run", side_effect=AssertionError("posted")): + self.assertFalse(notify.waiting(self.config, "needs 15.7 GB of free memory")) + + +class TestMemoryGuardAcrossProcesses(TempRoot): + """Found by measuring, not by the tests above. + + Replaying the demo recordings back to back, a fold started 0.2 s after the + one before had unloaded the models. Ollama was still listing gemma4, so the + new guard counted it as someone else's: it loaded it again, the watchdog + tripped at 9 % free, and unloaded only the embedder, leaving gemma4 in + memory for Ollama's five minutes. + """ + + def setUp(self) -> None: + super().setUp() + self.config.memory_guard = True + + def test_unload_returns_once_ollama_has_let_go(self): + from skill_plus_plus import memory + listings = [{"gemma4:e4b"}, {"gemma4:e4b"}, set()] + asked = [] + with mock.patch.object(memory, "_ollama", lambda config, path, payload=None, **kw: + asked.append((path, payload)) or {}), \ + mock.patch.object(memory, "loaded_models", lambda config: listings.pop(0)), \ + mock.patch.object(memory.time, "sleep", lambda s: None): + memory.unload(self.config, ["gemma4:e4b"]) + self.assertEqual(asked, [("/api/generate", {"model": "gemma4:e4b", "keep_alive": 0})]) + self.assertEqual(listings, [], "returned before Ollama stopped listing it") + + def test_one_process_at_a_time_holds_the_models(self): + import socket + from skill_plus_plus import local, memory + _ScriptedMemory(self, available_gb=16.0) + (self.config.root / "models.lock").write_text( + json.dumps({"pid": os.getpid(), "host": socket.gethostname()})) + with memory.guarded(self.config, wait=False) as guard: + self.assertEqual(guard.tripped, memory.BUSY) + with self.assertRaises(local.MemoryShort): + local.ask("gemma4:e4b", "hi") + self.assertTrue((self.config.root / "models.lock").exists(), + "another process's lock is not ours to remove") + + def test_a_lock_left_by_a_dead_process_is_taken_over(self): + import socket + from skill_plus_plus import memory + _ScriptedMemory(self, available_gb=16.0) + (self.config.root / "models.lock").write_text( + json.dumps({"pid": 2 ** 22 + 12345, "host": socket.gethostname()})) + with memory.guarded(self.config, wait=False) as guard: + self.assertEqual(guard.tripped, "") + self.assertFalse((self.config.root / "models.lock").exists()) + + def test_a_trip_during_a_load_unloads_the_model_once_it_lands(self): + """Not listed while it loads, so the first unload finds nothing.""" + from types import SimpleNamespace + from skill_plus_plus import memory + now, asked = [0.0], [] + clock = SimpleNamespace(time=lambda: now[0], + sleep=lambda s: now.__setitem__(0, now[0] + s)) + + def listed(config): + if now[0] < 6.0: + return set() # still loading + return set() if len(asked) >= 2 else {"gemma4:e4b"} + + with mock.patch.object(memory, "time", clock), \ + mock.patch.object(memory, "loaded_models", listed), \ + mock.patch.object(memory, "_ollama", lambda config, path, payload=None, **kw: + asked.append(payload["model"]) or {}): + memory.unload(self.config, ["gemma4:e4b"], linger=20.0) + self.assertEqual(asked, ["gemma4:e4b", "gemma4:e4b"])