Repository navigation
feat: a memory guard for the local model - #25
Merged
Merged
Conversation
A colleague's laptop nearly ran out of memory during the demo. Measured with 0.1.1 on an 18 GB Mac: no leak, Skill++ itself stays under 35 MB. The local models take about 12.8 GB of free memory, a one-word question as much as a whole session, and a fold with other apps open ran at 92 % used with 4.7 GB swapped out in 100 s. Now a fold loads the models only when they fit with 2 GB to spare, a watchdog stops it and unloads them at once if free memory falls below that, and they are unloaded as soon as the fold ends instead of after Ollama's five minutes. What the models take is measured on each computer. One Skill++ process holds the models at a time. A session that does not fit is held and folded later: when the computer is idle, at the next session start, or with Fold now on the review page. A desktop notification says when a fold waits or is stopped, and doctor and stats show the figures. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With gemma4:e4b-it-qat as the default local model (branch experiment/judge-completion-flag), a fold takes about 7 GB of free memory instead of 13 and starts at about 9 GB free instead of 15. The README, usage.md and the CHANGELOG now say so; the measurement that started the guard stays, named as gemma4:e4b's. The first guess at what loading takes (size on disk x 1.39) runs high for the new model, 8.3 GiB against 7.0 measured, and only decides a computer's first fold. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The memory guard tests script a machine running gemma4:e4b: its sizes, and the figures the tests expect, are that model's. They took the model from the default, so once #26 makes gemma4:e4b-it-qat the default, the scripted sizes no longer matched what a fold loads and five tests failed. The class now sets SKILL_PLUS_PLUS_LOCAL_MODEL for itself, so it holds whatever the default is and whichever of the two PRs merges first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
himanshu096
approved these changes
Sep 29, 2026
Brings in #26, gemma4:e4b-it-qat as the default local model. The README's install note and its "Good fit" paragraph conflicted: both branches rewrote them for the new model. This branch's wording is kept: the same figures, plus what the memory guard does.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
A colleague's laptop nearly ran out of memory while running the demo. Measured with 0.1.1 from PyPI, replaying the 8 demo recordings on an 18 GB Mac:
What it does
folded).keep_alive: 60scovers a fold that dies first.models.lock).HIDIdleTime, Linuxloginctlorxprintidle) and the models fit; at the next session start, at most every 10 minutes; or with Fold now on the review page's banner, orfold-pending --now.osascriptornotify-send) when a fold waits (at most hourly) or is stopped.doctorandstatsshow the figures.SKILL_PLUS_PLUS_MEMORY_GUARD,SKILL_PLUS_PLUS_MEMORY_RESERVE_GB,SKILL_PLUS_PLUS_IDLE_MINUTES,SKILL_PLUS_PLUS_NOTIFY.Measured on the same Mac, with this branch
fold-pending --now: 5 entries, as in 0.1.1. The models were gone 0.0–0.5 s after every fold, and the need measured here was 12.2 GiB.Measuring found three things the unit tests could not:
unloadnow waits until/api/psstops listing it, and one process holds the models at a time.With
gemma4:e4b-it-qatThe QAT default comes from #26: with one line added to the judge's question,
gemma4:e4b-it-qatgives the same verdict asgemma4:e4bat all 158 recorded gaps. The docs in this PR state its figures, so merge #26 first. Both branches reword the README's install note; the one merged second keeps this PR's wording.gemma4:e4bgemma4:e4b-it-qatFolding the 8 demo recordings with
gemma4:e4b-it-qatunder the guard: lowest free memory 7.2 GiB, pressure normal throughout.Tests
setUpModulelike the other model stubs, and the new tests script the memory readings.scripts/leak_guard.py: 0 hits.Not verified here
notify-sendare tested with recorded output only, not on a Linux machine.