Skip to content

feat: a memory guard for the local model - #25

Merged
lucazagaia merged 4 commits into
mainfrom
feat/memory-guard
Sep 29, 2026
Merged

lucazagaia merged 4 commits into
mainfrom
feat/memory-guard

Conversation

@lucazagaia

@lucazagaia lucazagaia commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Why

A colleague's laptop nearly ran out of memory while running the demo. Measured with 0.1.1 from PyPI, replaying the 8 demo recordings on an 18 GB Mac:

  • No leak. Skill++'s own process stays at 20–34 MB, and model memory is flat from one round to the next.
  • The cost is the model, not the session. A one-word question locks as much memory as the whole demo (+10.1 GB against +10.4 GB). Every prompt uses a 4,096-token context.
  • With other apps open, a fold ran at 92 % used, macOS pressure at warning, with 4.7 GB swapped out in 100 s. Loading the two models took 12.8 GB of free memory.
  • Ollama keeps the model 5 minutes after the last call, so the memory stays taken after the fold.

What it does

  • Check before loading. The models load only when what they take still leaves 2 GB free. What they take is measured on each computer: estimated from their size on disk at first, then the largest drop a fold has seen there.
  • Watchdog. While a fold runs, free memory is read twice a second. Below 2 GB, or at critical macOS pressure, no further model call starts, and the models Skill++ loaded are unloaded at once, including one still loading when it tripped. A half-judged session drops its verdicts; one stopped between episodes resumes after the banked ones (folded).
  • Release right away. The models are unloaded when the fold ends, not 5 minutes later. keep_alive: 60s covers a fold that dies first.
  • Only its own models, and one Skill++ process holds them at a time (models.lock).
  • Folding later. A held session is folded when the computer is idle (macOS HIDIdleTime, Linux loginctl or xprintidle) and the models fit; at the next session start, at most every 10 minutes; or with Fold now on the review page's banner, or fold-pending --now.
  • Telling people. A desktop notification (osascript or notify-send) when a fold waits (at most hourly) or is stopped. doctor and stats show the figures.
  • Settings: SKILL_PLUS_PLUS_MEMORY_GUARD, SKILL_PLUS_PLUS_MEMORY_RESERVE_GB, SKILL_PLUS_PLUS_IDLE_MINUTES, SKILL_PLUS_PLUS_NOTIFY.

Measured on the same Mac, with this branch

Run Result
Guard on, 2 GB reserve, 13.5 GB free All 8 sessions held before loading. Model memory stayed at 0.0 GB, free memory never below 74 %, one notification.
Guard on with room (0.25 GB reserve) 7 folded, 1 rightly held (12.4 GB free against 12.5 needed), then folded with fold-pending --now: 5 entries, as in 0.1.1. The models were gone 0.0–0.5 s after every fold, and the need measured here was 12.2 GiB.
Forced stop (planted need too low, 5 GB reserve) Tripped 7.9 s in at 14 % free, during the load. gemma4 was gone 2 s later, and free memory back to 72 %. Notice: "stopped and freed 9.5 GB". Sessions 2–8 held without reloading. Resumed later: the same ledger as the uninterrupted run, occurrences [7, 2, 1, 1, 1].

Measuring found three things the unit tests could not:

  • Ollama lists a model for a moment after unloading it, so a fold starting then counted gemma4 as someone else's and left it loaded. unload now waits until /api/ps stops listing it, and one process holds the models at a time.
  • A trip during a load found nothing listed to unload, and the load finished anyway. The unload now lingers 20 s and unloads what lands.
  • The first estimate mixed decimal and binary gigabytes. The factor is 1.39 in bytes.

With gemma4:e4b-it-qat

The QAT default comes from #26: with one line added to the judge's question, gemma4:e4b-it-qat gives the same verdict as gemma4:e4b at all 158 recorded gaps. The docs in this PR state its figures, so merge #26 first. Both branches reword the README's install note; the one merged second keeps this PR's wording.

Local model A fold needs A fold starts at
gemma4:e4b 12.2 GiB 14.2 GiB free
gemma4:e4b-it-qat 7.0 GiB 9.0 GiB free, 10.3 GiB before a computer's first fold (estimated from size on disk until measured)

Folding the 8 demo recordings with gemma4:e4b-it-qat under the guard: lowest free memory 7.2 GiB, pressure normal throughout.

Tests

  • 569 tests pass: 28 new, 541 unchanged. The existing suite runs with the guard off, set in setUpModule like the other model stubs, and the new tests script the memory readings.
  • scripts/leak_guard.py: 0 hits.

Not verified here

  • Linux: the readers and notify-send are tested with recorded output only, not on a Linux machine.
  • The notification banners: whether they appear on screen needs a person to look.

A colleague's laptop nearly ran out of memory during the demo. Measured with
0.1.1 on an 18 GB Mac: no leak, Skill++ itself stays under 35 MB. The local
models take about 12.8 GB of free memory, a one-word question as much as a
whole session, and a fold with other apps open ran at 92 % used with 4.7 GB
swapped out in 100 s.

Now a fold loads the models only when they fit with 2 GB to spare, a watchdog
stops it and unloads them at once if free memory falls below that, and they
are unloaded as soon as the fold ends instead of after Ollama's five minutes.
What the models take is measured on each computer. One Skill++ process holds
the models at a time. A session that does not fit is held and folded later:
when the computer is idle, at the next session start, or with Fold now on the
review page. A desktop notification says when a fold waits or is stopped, and
doctor and stats show the figures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With gemma4:e4b-it-qat as the default local model (branch
experiment/judge-completion-flag), a fold takes about 7 GB of free memory
instead of 13 and starts at about 9 GB free instead of 15. The README,
usage.md and the CHANGELOG now say so; the measurement that started the
guard stays, named as gemma4:e4b's. The first guess at what loading takes
(size on disk x 1.39) runs high for the new model, 8.3 GiB against 7.0
measured, and only decides a computer's first fold.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The memory guard tests script a machine running gemma4:e4b: its sizes, and
the figures the tests expect, are that model's. They took the model from
the default, so once #26 makes gemma4:e4b-it-qat the default, the scripted
sizes no longer matched what a fold loads and five tests failed. The class
now sets SKILL_PLUS_PLUS_LOCAL_MODEL for itself, so it holds whatever the
default is and whichever of the two PRs merges first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Brings in #26, gemma4:e4b-it-qat as the default local model. The README's
install note and its "Good fit" paragraph conflicted: both branches rewrote
them for the new model. This branch's wording is kept: the same figures,
plus what the memory guard does.
@lucazagaia
lucazagaia merged commit 2271766 into main Sep 29, 2026
11 checks passed
@lucazagaia
lucazagaia deleted the feat/memory-guard branch September 29, 2026 15:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants