Replies: 3 comments
Try running with |
|
Thanks a lot. I actually found that when keeping the VRAM usage below ~82% with weights and context, the VRAM usage does not grow. Above that, it steadily grows until OOM. You can play around with --fit-target, but I have no clue as to why this threshold would have any significance. |
|
@Te-eMster - The ~7.5 GB swell in the "unaccounted" column is driver/runtime memory fragmentation inside the HIP/ROCm layer, outside of GGML's tracked allocators:
|
Uh oh!
There was an error while loading. Please reload this page.
I am running llama-server on a dual MI50 (32GB) setup. I am using the HIP-backend because for gfx906 it seems to be more reliable for long context performance than Vulkan.
I am mostly using it to run models to work with Roo Code in VS Code. I am often running orchestrator-prompts which create several subtasks autonomously.
However, I am noticing that over time, the memory usage on the GPUs steadily increases until the server eventually runs out of memory. If I manage to shut down the server manually before the crash, in the memory breakdown there is a large amount of memory in the "unaccounted" column.
Fresh start:

Just before crash:

It is happening on MiniMax M2.5/M2.7 (running mixed with some MoE-layers on CPU) and also on Gemma 4 31B (running purely on GPU). Both models have fairly large KV-Cache requirements.
Qwen 3.5 27B doesn't seem to be affected as much by the memory growth - it also reserves very little KV-Cache for max context length.
Can anyone relate and explain what's happening? Is there any runtime argument that can prevent this from happening?
All reactions