Why does a 7B model require some 12 GB of memory, and a 4B model requires more than twice (causing pagination)? #29031
frx-wintermute
started this conversation in
General
Replies: 2 comments
|
I have also tried another 4B model: My box began swapping again: I'm more and more puzzled... |
0 replies
|
@frx-wintermute this might be a hot take, but dense models are rubbish. I have a laptop with 48gb, and for coding I run Qwen3.8-Flash-Next alongside Qwen3.6-35B-A3B, occasionally with a third model. MoE is basically magic and in my opinion Qwen does it best. How to run it well depends on exactly what hardware you have. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
$ llama-cli --version version: 0.4.1-dev (build 10964, commit Debian) built with GNU 16.2.0 for Linux x86_64Hi,
I tested the following 7B model:
$ llama-cli --multiline-input --temp 0.2 --top-k 8 \ --output $(mktemp log_XXXXXX.md) \ -hf MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUFand saw that it requires some 12 GB or memory (as checkd with
top):Then I tried to load the following 4B model:
$ llama-cli --multiline-input --temp 0.2 --top-k 8 \ --output $(mktemp log_XXXXXX.md) \ -hf lm-kit/phi-3.5-mini-3.8b-instruct-ggufand my box (with 32 GiB of memory) began swapping heavily... It seems to me that this model, despite having less parameters, requires more than twice as memory:
I must be misunderstanding something. Could you please explain what's going on?
Thanks for your time and dedication!
P.S.: Talking about open-weights models released under an open source license, which models would you suggest for code generation (and reasoning/engineering/science)? Is there anything that can be reasonably run on CPU with 32 GiB of memory? Thank you!
All reactions