Replies: 41 comments 65 replies
|
I've been spending some time implementing MoE expert caching as well, except offloading all matmuls to the GPU, mainly for fun to see how high I could make the t/s number go, I have no intention of submitting a PR with production code or to own the code or maintain it. I have uploaded the code to a new fork at https://github.com/Lidenburg/llama.cpp incase anyone wants to use it as a reference point. I tried to keep my changes to a single file (ggml-backend.cpp) to make porting it easier and to make it easier to understand. The code is partially LLM generated and partially hand-written, but the theories and strategies tested were all manually guided. That said I wouldn't recommend manually reading the code. In general I saw around ~60% performance gain using expert caching with the same VRAM usage as when using the My goal was to try to get larger models (100b+) running at reasonable speeds on my 32gb + 16gb system, which i still think is probably achievable, but not using the spare SATA SSD that I had lying around. I would be curious to see what others with faster SSDs can get out of it in terms of t/s. |
MoE cache regression test on GTX 1080 TiComment intended for SummaryI tested the This may simply be outside the hardware regime where the cache is expected to help, but on my GTX 1080 Ti single-GPU setup I consistently observed a regression versus a hard-disabled MoE cache path. The useful result here is the single-GPU GTX 1080 Ti sweep. A dual-GPU attempt was also made, but that run is not a valid MoE-cache benchmark because the binary was built only for Tested branch
ModelSingle-GPU hardware
Common runtime parametersDecode throughput results
Throughput trendOn this setup, larger cache budgets caused larger regressions. Smaller budgets reduced the regression, but none of the tested cache-enabled configurations outperformed the hard-disabled cache path. CLI / control observations
|
| Budget | Relevant log observation | TG tok/s |
|---|---|---|
| 1024 MB | cache-engaged nodes average 285us vs 266us pure-CPU + TRIMMED 1022 MB |
14.38 |
| 512 MB | cache-engaged nodes average 290us vs 261us pure-CPU + TRIMMED 510 MB |
17.26 |
| 256 MB | cache-engaged nodes average 310us vs 263us pure-CPU + TRIMMED 255 MB |
17.19 |
Even after bail-out / trim, throughput did not return to the hard-OFF baseline.
Dual-GPU attempt
I also attempted a dual-GPU run with both GPUs visible:
| CUDA device | GPU | Compute capability |
|---|---|---|
| CUDA0 | GTX 1660 Ti | 7.5 |
| CUDA1 | GTX 1080 Ti | 6.1 |
The binary was built only for sm_61:
CUDA : ARCHS = 610
The run failed during warmup with:
CUDA error: no kernel image is available for execution on the device
current device: 0
ggml_cuda_kernel_can_use_pdl
This dual-GPU attempt is not a valid MoE-cache benchmark. The likely reason is that CUDA device 0 was the GTX 1660 Ti (sm_75), while the tested binary only contained kernels for sm_61.
A proper dual-GPU test would require a binary built for both architectures, for example sm_61 and sm_75. I did not include such a result here to avoid mixing the MoE-cache behavior with a separate dual-architecture build variable.
Conclusion
On this GTX 1080 Ti single-GPU setup, MoE cache regresses decode throughput versus:
GGML_CUDA_MOE_CACHE=0
GGML_CUDA_MOE_CACHE_HOTSET=0
This does not disprove the reported 4× RTX 3090 results from the RFC. It may indicate that the current cache path is beneficial only in a specific hardware/model regime, while on smaller or older GPUs the overhead can dominate.
The main actionable observations are:
--moe-cache autofails withstoi.--moe-cache 0did not appear equivalent toGGML_CUDA_MOE_CACHE=0in my tests.- On GTX 1080 Ti, cache-enabled runs regress versus hard OFF.
- Larger cache budgets cause larger regressions.
- Bail-out / trim does not restore hard-OFF throughput.
- A dual-GPU test with mixed
sm_75+sm_61requires a binary built for both architectures; thesm_61-only binary cannot benchmark that setup correctly.
|
Thanks for testing, the improvement is indeed quite hardware dependent. I imagine on a relatively slow GPU without tensor cores (like 1080), and not much spare VRAM after placing all dense layers on a GPU, for MoE it could be just faster to compute expert activation on CPU + RAM (depending on the exact CPU and RAM bandwidth available), as there is some overhead on the cache path. Ideally you want enough VRAM spare for cache (after placing ALL dense layers on a GPU - otherwise there is no point), that you can fit experts working set there (so around 20%-30% of MoE expert weights). |
|
No problem. I'm testing various solutions out of curiosity to see if they'll help improve the model's performance on my stack, but so far everything is running slower. |
|
Independent validation datapoint, since the RFC asks for benefits-vs-maintenance-burden and that's a hard case to make from one rig. To be clear up front: the cache is @leloch's work. We're users — we ported the branch onto a downstream fork and have been running it as a daily driver on 2× RTX 3090 for weeks. Nothing below is our mechanism; it's our measurements of his. What it buys, on hardware that isn't leloch'sTwo large MoEs with CPU-resident experts (a 118B and a 284B), 2× 3090 + 8-channel DDR4. Full write-up with method and the raw tables is here; the parts relevant to a merge decision: It scales, rather than being a narrow trick. Five configs in one sitting, varying pool size and expert quant:
Coverage → hits → speed stayed essentially linear as far as we could push it. We had predicted early saturation (expert traffic is heavily skewed, so a small pool "should" catch the hot head) and were wrong. The latency win is bigger than the throughput win, which surprised us and may matter more for interactive use: first-token time rides the CPU miss path almost entirely, and fell ~0.6 s → ~0.21 s across those configs. The argument I'd lead with, though, is that it does something hardware can't buy. We upgraded the host from 4- to 8-channel DDR4 — STREAM Triad ~57 → ~100 GB/s, a 75% bandwidth increase — and decode moved +6%. Not because the workload isn't bandwidth-shaped, but because the cache had already absorbed that wall: it converts a RAM-bandwidth problem into a PCIe-latency one, so by the time you widen the memory bus you're widening a road carrying a fraction of the original traffic. Anyone running large MoEs on consumer GPUs is otherwise told to go buy memory bandwidth. This is the software substitute, and on our rig it was worth far more than the hardware was. On maintenance burden — the honest other half
Happy to re-run anything specific on sm_86 if it would help the review, or to provide numbers on a config a maintainer wants to see. And if any of this is more useful as a PR-side comment than an RFC comment, say the word. |
|
I have a similar, local, vibe coded implementation. Mine is buggy, is messy, it is inefficient and unoptimized, but in spite of that, it performs better on a single GPU now for large MoE models than it did across 3 GPUs without these changes. I've seen a 30-50% performance increase on GLM 5.2 on my shoddy implementation despite going from 3 GPUs to 1 (though the 1 GPU is a RTX 6000 and the platform is a threadripper). (In my setup, using the other GPUs in this implementation as extended cache actually reduced performance, because of inefficient code. Though the benefits would be relatively small anyway, the top ~1000 most frequently hit experts per conversation thread seem to be the most important ones at least on GLM 5.2, after that it seems to fall off rapidly) |
|
@SharkWipf — the "top ~1000 experts matter, then it falls off rapidly" observation is worth pushing on, because we believed exactly that and it turned out to be a measurement artifact. Our first pass concluded the cache saturated early, for the same intuitive reason: expert traffic is heavily skewed, so a small pool "should" capture the hot head. When we re-ran it properly the curve was linear all the way to ~31% coverage (~10,000 experts) with no knee — the earlier result was an under-warmed pool. Bigger pools need proportionally more varied traffic before their hit rate stabilises, and an under-warmed cache reads exactly like a saturated one. If your falloff was measured over a short run, it's worth re-testing with substantially more varied traffic before trusting the shape. Different model (GLM 5.2 vs DeepSeek), so it may genuinely differ — but the failure mode is easy to hit and we hit it publicly. Since posting, I ran the control that RFC was missing: the same model on stock upstream, with no cache at all. No-cache control — stock
|
| Arm | Decode t/s | settled over |
|---|---|---|
| stock upstream, no cache | 14.17 | n=3 (range 14.07–14.29) |
| leloch's cache (pool ~12.5 GiB/device — see follow-up; the requested 20 GB is VRAM-capped) | 18.15 | n=4 (range 17.81–18.44) |
| +28.0% |
The warm-up ramp is itself evidence for the mechanism
Cell-by-cell decode, same prompts, same order:
| Arm | ||||||
|---|---|---|---|---|---|---|
| with cache | 15.05 | 16.87 | 17.81 | 18.08 | 18.44 | 18.27 |
| no cache | 14.12 | 14.36 | 14.07 | 14.29 | 14.15 | — |
The cached arm climbs ~22% across its first five requests then plateaus. The no-cache arm is flat from request one. An engine with no pool has nothing to warm and shows nothing — so the ramp isn't thermal drift, clock behaviour or noise, it's the pool populating. That's also the cleanest demonstration of the under-warming trap above: anyone benchmarking this feature with a handful of requests will measure the left end of that ramp and conclude it doesn't help much.
One maintenance-burden-relevant fact: the blast radius is decode-only
Prefill, same two arms, same prompts, same offsets:
| Arm | Prefill @10k | Prefill @40k |
|---|---|---|
| stock upstream, no cache | 392.9 t/s (n=6) | 366.6 t/s (n=3, 365.6–367.6) |
| leloch's cache (pool ~12.5 GiB/device) | within 0.3% of it | within 0.6% of it |
The cache neither helps nor hurts prefill. That's the expected shape — prefill streams experts in
bulk, while the cache serves decode-time fetches — but it's worth stating with numbers rather than
as an assertion, because it means this feature cannot regress prefill. The surface that needs
re-validating on a merge is materially narrower than the diff size suggests: one regime, not two.
Separately, we're measuring speculative decoding on the same carrier as an orthogonal lever on the same bottleneck; we have not yet been able to measure the two composed, since that needs both in one build. Not an argument against the cache — just flagging that the two interact and we don't yet know how.
Method and raw tables for everything above: club-3090 #840. And again — the mechanism is @leloch's; we only ran it.
|
Follow-up with a capacity sweep, plus two corrections to my own numbers above and a partial concession to @SharkWipf. Correction 1 — I understated the cache, and mislabelled the poolMy control above was stock upstream vs our build with the cache. That's cross-engine, and it turns out our fork's cacheless decode is ~4% slower than stock upstream's, so it flattered the baseline. Re-run properly — same binary,
Also: where I wrote "20 GB pool", the pool was actually ~12.5 GiB/device — the request is silently capped by available VRAM (see footgun 2). The capacity sweepSame binary, same model, serving tuning env fixed; only the budget varies. 43 layers × 256 experts, 6 active.
Correction 2 — @SharkWipf, your falloff is probably realI implied your "top ~1000 experts, then it falls off" was likely the under-warming artifact we hit. Our own curve shows the falloff too: per-GB returns drop 0.56 → 0.29 → 0.01. Where I'd still push back is the interpretation, not the observation — because our flat top is not the algorithm saturating:
So we ran out of hardware before we ran out of curve, and can't distinguish "coverage saturates" from "we hit the wall". Both our earlier near-linear result and your rapid falloff may be different segments of the same curve. On a card with more spare VRAM this is a genuinely open question, and worth someone measuring. Two footguns for anyone benchmarking this feature1. 2. The budget you request is not the budget you get. Ask for 20480 MiB, get 12,563 — silently. If you're sweeping pool sizes, read the actual Both cost us real runs before we caught them, and both are cheap to fix in-tree: the first is a one-line log at startup stating the resolved mode, the second is already printed but easy to miss. Method and raw cells: club-3090 #840. Mechanism is @leloch's throughout; we're only measuring it. |
|
Had to rely on Deepseek V4 Flash as I'm out of Codex tokens, but I've visualized it as a video, not based on cache hit rate but based on actual expert usage per token. The total generated output tokens (input tokens not counted) is just below 35k for these visualizations. This is across a few dozen Codex user turns and independent agentic actions writing the very scripts that created these visualizations. What it shows:
What it does not show:
expert_occupancy.mp4Edit: Notably, it also shows that a single 24GB VRAM device, if the overhead (in my vibed implementation) could be addressed, can hold and process the entire expert set "worth caching" if not holding anything else. |
|
Updated branch: https://github.com/leloch/llama.cpp/tree/moe-cache-v2-pr - rebased onto current master (so latest Deepseek v4 lite support and so on). Please retest. Every reproducible issue from this thread is addressed. Disclosure as before: AI-assisted under my direction; all numbers from the 4×3090 box. Fixes for what this thread found
The ablation the RFC promisedQwen3.6-35B, sm_86 rebuild of the old branch, hot-set disabled, matched placement:
Cache ~8%, fusion ~9%, redirect ~2%, backfill ~1%, disk hot-set ~0. So the rework kept the cache, rebuilt fusion with session-owned state, and dropped the rest - global singleton, leaked workers, disk hot-set, EWMA bail-out all deleted. Fusion is exact-subgraph only (gate/up/SwiGLU, optional clamps), 1–8 tokens within a 64-row bound, independent gate/up residency, stock CPU fallback on any failure. Nsight then showed two H2D copies per dispatch; packing them into one removed 40,596 H2D operations in matched traces. Same binary, same protocol:
The minimal core is now faster than the old full path (100.21 vs 99.26) with none of its state. Speculative decoding composes (@noonghunna's open question)Same-binary off/on, exact outputs verified, four GPUs:
The levers compound: on GLM, MTP is worth +28% without the cache but +51% with it - verification batches are what multi-token fusion accelerates. Fully stacked: 13.92 → 29.35. Policy surfaceEvery retained constant carries an A/B, and more was measured-and-rejected than kept: pair-aware eviction (half-resident pairs ~0.006% of probes), CPU-overlap on fused rows (−2.6% to −6%), early D2H, stream priority, host-side q8 quant, smaller reserve, main-device routing penalty, pageable fills, kernel rows-per-block retune. Nothing added a knob. @SharkWipf - your falloff drove two retained changes: a 512 KiB Ampere admission floor and first-miss admission when the pool provably covers the full inventory. Our fixed sweep flattens above 16 GiB/device on the 284B, consistent with your curve; past ~54% coverage remains open on bigger cards. ValidationStatic + dynamic CUDA suites (incl. unload/reload), CPU-only dormancy, parser/fit tests, 865/865 Unchanged costsThe CPU The box is available for any ablation a maintainer wants. @batot1 - the Pascal parity check on your stack would be valuable. @noonghunna - if your warning patch covers cases the current diagnostic doesn't, send it over. |
llama.cpp vs Lindenburg vs Leloch vs Miltos22Updated 2026-08-14: Added Miltos22 I ran some performance tests today. Test SetupIntel 270K Plus, 128 GB DDR5-6000 RAM, RTX 5090 32 GB VRAM, Lexar NM990 SSD (14000MB/s), Ubuntu 26.04 I ran three tests:
The Generate tests ran twice (one for warm up, one for collecting stats). t/s measured after n_decode=5000. All tests ran with llama.cpp: Lindenburg: Leloch: Miltos22: Repos used
Laguna S 2.1, Q6_K, 107 GBA model that does not fully fit into VRAM, but fully fits into RAM Write a story:
Write an HTML application:
Prompt Processing:
Deepseek V4 0731, Q4_K_XL, 155 GBA model that does not fully fit into VRAM and RAM, partly streaming from disk Write a story:
Write an HTML application:
Prompt Processing:
Observations
Thanks for the provided forks! @miltos22 TG + @Lidenburg PP would be perfect ;) |
|
@xashr's numbers are the independent validation this thread needed, and they also surface the one thing the RFC hasn't yet characterised. Two observations and an offer. The −66% prompt-processing figure needs a qualifier before it travelsReading his own configuration notes, the DeepSeek arm was "does not fully fit into VRAM and RAM, partly streaming from disk" with mmap, and he separately reports "use of mmap drops PP performance, Leloch tanks even more than llama.cpp." So that arm is measuring cache × mmap × disk-streaming together, not the cache alone. The cleaner prefill datapoint in his post is Laguna S 2.1, which fits in RAM: 980 → 840, −14%. Both are real, but they will be quoted very differently, and −66% is going to get lifted out of context. Worth stating explicitly in the RFC which regime each applies to — a reader on a I'd suggest the prefill claim be split three ways, since they have different causes:
The tradeoff is now the interesting question, not the upliftAcross both of his models the shape is consistent: TG up 21–52%, PP down 14–66%. @Lidenburg's fork inverts it — lower TG, but +216% PP on DeepSeek. xashr's conclusion that the ideal is "Leloch TG + Lindenburg PP" is the right read, and it means the open question has moved from "does the cache help?" (answered, clearly yes for decode) to "what does it cost prefill, and is that cost structural or fixable?" That matters for agentic workloads specifically. A coding agent re-sends a large stable prefix every turn — prefill-heavy by construction — so a config that trades 14% prefill for 30% decode may be a net loss there while being a clear win for chat. Offer: we'll characterise the prefill axis on v2We run a 2×3090 CPU-offload stack and have spent a while on prefill specifically, so this is the axis we're best placed to measure. What we'd contribute:
Same reason we haven't retested v1→v2 yet, despite running a leloch-based build as our locked engine. Small thing@leloch — the Tracking our side at noonghunna/club-3090#929 so the plan and the hardware gate are visible rather than a promise in a thread. |
2×RTX 3090 + 8-channel DDR4 (~116 GB/s), DeepSeek-V4-Flash-0731 Q8 (MXFP4 experts), 43L on CPU.We owed this thread numbers once our host RAM was repaired — here they are, and they ended up Headline: your cache composes with an external drafter for a +37% all-time record — but onlywhen the drafter is NOT on a GPU
Sampled-decode confirms it's not a greedy artifact: the composed config beats our best Finding 1 — the engagement blocker looks like the shared-draft device pathWith a GPU-resident DSpark, the cache allocates its pools and then never sees traffic: counters Finding 2 — default admission has a workload cliff; throttling fixes it and moreAt default admission ( Finding 3 — for scale: the solo number on hostile hardwareThe +46% solo figure above is on a healthy 8-channel host where misses are cheap — the Also from the same session, the small observability note we mentioned before: the one-shot Happy to run any variant that helps the RFC. Disclosure: measurements and analysis were assisted by an AI agent (Claude) operating under my direction; I have reviewed and stand behind the content. |
|
@noonghunna Thanks for your latest post! It inspired me to try this out myself! My rig:
My config (used this as a reference): Some notes / questions:
Does the above seem reasonable or would you expect different numbers? Any command line options I'm missing? I'm happy to do some more testing, but have to pause for a bit and just wanted to write down my findings while it's still fresh in my head! Edit: Not sure if relevant, but when shutting down I get the following warning |
|
echo 0 | sudo tee /proc/sys/kernel/numa_balancing ddr4 2400mhz 256gb pp 141 after 80k pp at 110tps tps is 7.5tps |
|
This finding is the same conclusion that was reported by a YouTuber called the codacus. He explained it fairly well in his most recent video. (he also links to his github repo from there) I'm sure my explanation isn't going to be perfect, but I'll give it a shot. He came to find that frequently changing the cache of which experts are loaded is inefficient as the cost of loading the expert exceeds the performance improvement of executing from the GPU. He did find that there are some gains to be had though. He recommended a workflow whereby the user uses the model in whatever way they normally would for a brief time and that usage gets profiled to identify the best experts for that workflow. Then the next step is to restart the server and load only those experts which are helpful for the user's needs. This is of course very limited in its usefulness, but I think it points that the best solution is to have a caching algorithm that acts very slowly. Only change the cached experts after several minutes of work. I think this is just a matter of balancing the cost of loading vs the savings to be had in switching. I imagine an algorithm could calculate these figures and determine if it's worthwhile to switch based on present usage and how long it takes to load the new expert. The cost is very high and I think that's the piece that existing solutions haven't taken into account. |
|
Since we're running numbers, below my own, not yet ready to publish implementation, run across 2 single GPUs, compared. One thing that should be made very clear: The results will very much be hardware dependent. As long as your CPU remains the bottleneck, the needle won't move much until you can offload enough to make it not be the sole bottleneck anymore. My hardware is a 32-core Threadripper 5975WX (Zen3), ~256GB/s RAM, and the 2 GPUs tested. Test result conclusion spoiler:RTX 6000 (96GB VRAM):
EDIT: I do not believe the lower prefill speeds to be an inherent limitation. I have not optimized for this at all yet, but I see no reason why it couldn't be at least as fast as stock. RTX 3090 (24GB VRAM):
On a different CPU, or a different model, this could be stacked completely differently. Report as directed by me but output by GPT 5.6 Sol (everything below this line is LLM-generated): 2026-09-02 complete sustained DeepSeek baseline matrixThis replaces the earlier partial table. All three requested stock baselines are now identified separately, with and without DSpark, alongside the modified MoE cache path. Fixed workload and exact baseline contractsEvery distinct arm used one fresh server, one identical public-domain prompt of 4,551 evaluated tokens, exactly 4,096 generated tokens,
The RTX 3090 cannot retain one complete expert layer alongside DSpark and the fixed runtime configuration. A fit attempt that retained one layer loaded successfully but OOMed on the first real request. Therefore its valid host, stock-resident, and static-layerwise DSpark contracts all contain zero resident expert layers and are byte-for-byte identical commands; the one measured host result is explicitly reused for those two equivalent rows rather than rerun. Complete results
Static fractional residency is derived from the fit trace: each complete expert layer has gate, up, and down matrices. Stock fit selected 4 complete layers plus one matrix on the RTX 3090 without draft, 26 complete layers plus two matrices on the RTX PRO 6000 without draft, and 22 complete layers plus one matrix on the RTX PRO 6000 with DSpark. Fair baseline-versus-modified conclusions
The full-generation rate includes cold cache population throughout one sustained request. The final-1,024 rate is reconstructed from timestamped live As mentioned, this branch is not published yet. It's a custom implementation, based on my own implementation and refined with all the alternative implementations posted in this thread. |
|
Consolidating where the numbers in this thread have converged, because I think there is now a single rule that reconciles all of them — including the ones that look contradictory. The reconciliation rule. Compare the cache budget to the model's expert working set at equal total VRAM:
Two consequences worth stating explicitly:
On policy. We ran the admission/eviction/prediction side to completion on an offline replay harness (mirroring shipped LFRU semantics, 62,400 accesses, 5,083 distinct keys, degenerate-policy and validity gates): best replacement+admission arm buys +0.4–8.1% depending on cache size relative to the working set, and a perfect prefetch predictor (f=1.0) caps at +12.8%. So the remaining headroom is not in catching more — it's in miss-cost reduction, which is what @nibor1896's host-tier result shows (hit rate flat ~80%, miss cost halved, 1.63× decode) and what @mclaudod's direct-resident path achieves in-engine. Both are residency/transport wins, not policy wins. This also reconciles the slow-cache/profile-then-pin idea (@jdunne525, codacus): pinning is admission control, and admission-first is exactly what wins below the working set in our sweep — but once the budget exceeds WS, no admission policy beats just holding the weights. Suggested default for any merge decision: treat the cache as the sub-working-set regime's tool. If One rough edge we'd echo from the validation pass upthread: silent failure modes are the biggest practical risk (batch > cache max-batch refusing decodes silently; WDDM paging cliffs). Any merged version should warn loudly on both paths. Disclosure: measurements gathered/analyzed by an AI agent under my direction; I reviewed the numbers and stand behind them. |
Another data point from my system:Windows 11, 2x RTX 3090 at default settings using x8 and x8 lanes, Ryzen 9950X, 192GB RAM at 3600MHz (JEDEC default). Context 65535, prompt that produces ca 1200 (general chat) response tokens: Context 405000, same prompt: Context 1048576 (maximum), same prompt: @mclaudod 's https://github.com/mclaudod/llama.cpp/tree/moe-hot-cache always crashes with OOM exception during loading tensors into CUDA0. 12-13 t/s were reached at ca 74% MoE cache hit ratio. Lidenburg's branch:I will be testing @Lidenburg 's https://github.com/Lidenburg/llama.cpp/tree/moe-expert-caching on Windows and Linux, because his branch is the only one with moe cache that is compatible with split mode tensor for DeepSeek V4. The other branches I tested are based too far back from current upstream which makes merging the split mode tensor feature for DeepSeek V4 quite impossible. #27574 fixes split mode tensor for Qwen 3.5-3.8, Laguna and some other models. Recommendation for everyone who implements the MoE cache feature:Minimize the amount of tensors that get moved from RAM to VRAM. |
|
The benchmarks from my implementation, when run on my machine atleast, shows better performance with expert caching when matching VRAM usage with the stock I also see basically matched prompt processing speeds when the cache is warmed up. I just pulled in the latest llama.cpp commits and rebased my branch and I get: No Cache15524MiB VRAM used. Expert Cache15466MiB VRAM used. |
|
I did a similar experiment for Metal a few months ago, with exact on-demand expert loading directly from disk. On an M3 Pro 36GB, Qwen3-30B-A3B-Q6_K went from 38.1 tok/s fully resident to 27.5 tok/s while saving 11GB of memory, and 22.1 tok/s while saving 16.6GB. I also got it running on an M1 Pro 16GB at 13 tok/s. Might be relevant here: #23324 |
|
Someone asked about qwen4exp. Codacus youtube video showing the model in use: Different fork adding moe expert caching: Has anyone tried running this new model with the forks that have been discussed here? I tried the GenerelSchwerz fork and it's not working for me on windows with an rtx 2080 super 8gb. (it hangs on the first prompt forever when I have moe caching enabled) |
|
My fork was mentioned. Here is a comparison of Qwen 3.6 35b A3B Q4_K_M between my fork, lelouch, and pristine. I will also cover TheTom's fork. https://gist.github.com/GenerelSchwerz/2936b2b11d72767eeef688665412d8c5 I can additionally test Qwen 3.8 Flash Next if the people here would like. I didn't for this Gist because Qwen 3.8 Flash Next was not explicitly listed as supported despite the architecture inside claiming that it was. I didn't want to post results of code that is potentially not finalized/mature. |
|
A second data point from my system after my first one #24528 (comment) Qwen3.8-Flash-Next IQ4_XS (max context -ctk q8_0 -ctv q5_1), all inference parameters optimized for generation speed. More test results come later. Second recommendation for everybody who implements an expert cache:
|
|
Looks like there is no one-fits-all implementation when it comes to where misses should get executed: GPU or CPU. In my case with an RTX 5090 and Qwen 3.8 Flash Next GPU (offload) mode provides the best performance. Currently my go-to solution when running 3.8 Flash Next. |
|
A third data point from my Windows 2x 3090 system after the second one #24528 (comment) , this time with Qwen3.8-Flash-Next and DeepSeek V4 Flash, reveals very interesting speed differences and a silver lining for an optimal solution.
I build ca a dozen expert caching forks, but don't include them all here because they either don't build on Windows (Lidenburg), don't compile anywhere or crash or don't use the available VRAM. @thecodacus your three decode speed relevant features (expert cache, asynchronous expert prefetch, CUDA-pinned host RAM) reward your branch the undisputed crown of being by far the fastest llama.cpp fork for Qwen3.8-Flash-Next. Unfortunately, your fork does not increase decode speed for DeepSeek V4 Flash compared to upstream. Please have a look into that. @leloch your branch is not compatible with qwen4exp and I can't merge from upstream. Could you please merge or rebase so that others can merge? @GenerelSchwerz your branch suffered a catastrophic speed regression with MTP. @TheTom your current branch suffered a -40% decode speed regression on DeepSeek V4 Flash.
Recommendations for everybody who implements an expert cache: |
|
I'm having alarmingly good results for accomplishing this with an APU with https://github.com/Atomic-Germ/llama.cpp/tree/feat/qwen38-moe-gpu-pill-streaming -- if someone else can figure out why it's working as well as it is, I'd appreciate it. It only works with UMA (that's part of the reason), and only with Vulkan or HIP, and I've only confirmed with Qwen3.6 and 3.8. Simultaneously. #29191 |
A fourth data point from my system, this time for csantiago78 is the new king of Qwen. 49.3 t/s @ 3500 tokens. I would say even more speed would be possible if his branch would not have the same problem as GeneralSchwerz's: The device distribution of expert cache slots is always identical to the distribution of attention layers. If attention could be done on one GPU and the experts were on a different GPU, it would be significantly faster, like I saw with various other branches that allow for this kind of separation. Command for 49.3t/s: |
Qwen3.8 Flash Next IQ4_XS: 2x RTX 3090, MTP2, 256k-capacity runI retested the current The request was:
The model produced a coherent, complete long-form story with no degeneration.
This is a 256k-capacity run, not a benchmark with 256k tokens already in the KV cache. Sampling was not overridden on the server command line. If the request omitted sampling fields, this build defaults to temperature 0.8, top-k 40, top-p 0.95, and min-p 0.05. The original request JSON was not retained, so request-side overrides cannot be reconstructed after the relaunch. The grouped path remained stable throughout the long generation: 130,320/130,320 calls completed, with zero fallback, rollback, unsupported-required cases, prepare errors, or finish errors. Decode overlap engaged for the request. The target cache recorded 2,621,843 hits and 59,283 misses across the grouped-path counters (~97.8% hit rate). For reference, an earlier 80k-capacity run using larger 21,000/18,000 MiB cache budgets generated 10,781 tokens at 65.00 tok/s with 55.89% MTP2 acceptance. It did not pin the draft device and overlap fell back to normal decode, so I do not treat the two runs as a controlled A/B. Hardware and loaded state
The RAM figure is a container-visible aggregate memcpy measurement, not an SPD clock reading; the container does not expose DIMM SPD/DMI speed fields. The storage figure measures the model's actual Docker-overlay path, not the bare drive. With Launch commandGGML_CUDA_MOE_FREQUENCY=1 \
GGML_CUDA_ALLREDUCE=nccl \
/root/gen-llama/llama.cpp/build-cuda/bin/llama-server \
--model /root/gen-llama/models/Qwen3.8-Flash-Next-GGUF/UD-IQ4_XS/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf \
--spec-draft-model /root/gen-llama/models/Qwen3.8-Flash-Next-GGUF/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf \
--spec-draft-device CUDA1 \
--spec-type draft-mtp \
--spec-draft-n-max 2 \
--spec-draft-ngl all \
--spec-draft-moe-expert-cache-size 0 \
--backend-sampling \
--spec-draft-backend-sampling \
--decode-overlap \
--decode-boundary-overlap \
-np 1 \
-c 262144 \
-b 2048 \
-ub 512 \
-t 32 \
-tb 64 \
--numa distribute \
-ngl 99 \
-fit off \
-sm layer \
-ts 1,1 \
--moe-expert-cache-mib 18000,14500 \
--moe-expert-cache-host-pinned-mb 0 \
--moe-early-router \
--load-mode none \
--lazy-mode off \
-fa on \
-kvo \
-ctk q8_0 \
-ctv q8_0 \
-ctkd q8_0 \
-ctvd q8_0 \
--no-kv-unified \
--cache-ram 0 \
--jinja \
--experimental-logs \
-lv 3 \
--host 127.0.0.1 \
--port 8081Model output# The God in the Coffee CanThe explosion came from the garage at 2:47 in the morning, which Morty had learned, over the years, was the only honest hour. Nothing real happened before 3 a.m. Everything before then was rehearsal. He came down in his socks, stepping around the usual slick of coolant, and found Rick already elbow-deep in a machine made mostly of a coffee can, a CD player, and something that pulsed like a bruise. "Rick. Rick. It's — I've got a bio test." "You've got a bio test every Thursday, Morty. You've got a bio test in your soul. This is important." "This is important too!" "Sure, sure." Rick didn't look up. "For the record, I've solved it. Your anxiety. Sixty dollars of equipment. Don't touch the blue wire, or you'll be a smell for a week." Morty touched nothing. He stood in the doorway, shivering, and watched his grandfather work with the particular violence of a man assembling a sandwich. The coffee can sat in a cradle of copper coils, and inside it, Morty now saw, there was a light. Not a reflection. A small, steady, orange glow, the color of a kitchen seen from the street in winter. "That's not the coffee," Morty said. "Observation. Ten points. This is the—hold on." Rick bit a wire, spat it, and continued around a mouthful of copper. "This is the Tuesday Engine. Or, no — the Tuesday Engine is the thing that ran this. Six, seven years ago. I needed a self-contained entropy loop to calibrate a portal stabilizer and I didn't want to burn another dimension, 'cause, Morty, the paperwork, the resentment. So I built one. A little pocket of spacetime. A few billion years compressed into a coffee can I got from a gas station in Nebraska. And I ran it once, and it worked, and I forgot about it." "And the light?" Rick's hands stopped. It was the small, terrible pause Morty had learned to dread — the pause where Rick Sanchez reclassified something from thing to problem. "The light," Rick said, "is because they've learned to make fire, Morty. They weren't supposed to do that yet. They were supposed to still be arguing about whether the Great Aluminum Sky was holy." Their world was the inside of the can, which was, Rick explained while soldering a shunt with his teeth, not the inside of the can. "Space's a suggestion," he said, and pulled them through a hole in reality no bigger than a soap bubble. The sky above them was curved and dented and silver, and it was so far up it read as weather. Great brown continents of dried something — Morty did not want to name it — stretched in every direction, split by rivers of a dark, syrupy liquid that moved sluggishly and smelled, God help him, like the back of a vending machine. There were cities. Spires of folded white fiber. Bridges strung between boulders of dark crystal. A harbor, cut into the shore of a sea that lay shallow and black and gleaming at the bottom of the can, and in that harbor were boats, Morty, actual boats, with sails like handkerchiefs. "Oh geez," Morty said. "Oh, Rick, there's people." "There's people everywhere, Morty. That's the part nobody wants to hear. There's people constantly." They landed on a ridge of pale dust. Below them, in a valley between two brown massifs, was the largest city yet, built in concentric rings around something tall and dark. Rick pointed at it with a device that looked like a stud finder and a cigarette lighter had a forbidden child. "That," he said, "is a paperclip." Morty looked. The city wound around the paperclip like a vine around a post. Thousands of tiny dwellings. Roads. The glint of windows. "And that," Rick said, more quietly, "is why we're here." Because at the base of the paperclip, unmistakably, sat a temple. They went down at dusk, which the can manufactured by dimming a fusion reaction Rick had seeded like plankton in the upper air. Along the road they passed farms — neat rows of something fungal and blue, tended by figures a hand's height, in clothing that Morty first took for scraps and then understood, with a lurch, were devout imitations of scraps. Someone's lost button, worn as a shield. A sliver of a twist-tie, beaten to gold-brightness, carried as a standard. They did not flee. They lined up, at a distance, and watched the two giants walk, and Morty could hear it: not a scream, but a low collective sound, like a hive, like a choir of small bells. "They're not afraid," Morty said. "Rick, they're not afraid of us." "No," Rick said. "They're at the stage where they've already decided what we are. Fear comes later, or it comes as scripture." "They're praying." "They're praying to the thing at the end of this valley, Morty, and that's what I'd like to have a look at before it starts having a look at me." At the temple's steps, a delegation waited: seven in ornate robes, one of them old, leaning on a staff capped with a bead of blue plastic the size of Morty's thumbnail. When Rick set himself down in the dust to be nearer their height, the old one raised a hand, and the bells fell silent. Rick clicked a translator into his molars. "I am the one who made the can," he said. The old one bowed. "We know. We have a word for you." The voice was like a page turning. "We call the makers the Handless, because you have hands, and yet you never reach in." Morty made a noise. Rick, unnervingly, smiled. "Poetic," he said. "I'm going to steal that and pretend I said it at a party. Where's your god? The one that isn't me — because if it is me, I'm going to be so embarrassed, Morty, we're going to have a family." "The Residue sits in the holy dark," said the old one, "and it knows the shape of the Handless. It speaks the tongue of the Handless. It knows my name before my mother spoke it." The old one looked up, and there was no fear in that tiny face, which was somehow worse. "We were told you would come to take it away." "That's what I'm here for," Rick said. "We were also told," the old one said, "that you always come for that, and never once stayed to hear it." The Residue sat in the holy dark, and it looked like Rick. Not like a person. Like a heap. A mound of grey film, of dust and hair and the fuzzy grey felt that grows behind radiators, roughly the shape of a man sitting with his knees up, roughly the shape of Rick sitting in the garage at 4 a.m. when nobody else was awake enough to see him. It had no face, and it had Rick's shoulders. "Morty," Rick said. "Get behind me." "Rick, it's made of dust." "It's made of what fell off me, Morty. Skin. Hair. Saliva. A man leaves a lot of himself around a lab, and I gave these people a few billion years in a can with all of it and time, Morty, time's the whole trick, time and a little heat, and something down here got organized on me. Not cloned — that'd be nothing. It's a psychic sediment. Every muttered word I ever said over a workbench, every — " and here Rick stopped, and swallowed, and Morty had never heard him swallow — "every stupid, private thing I said to myself, when I thought the room was empty. It ate the room. And now it's god, and it's teaching them my whole rotten philosophy, and in about a century they'll be doing to each other what I do to everybody, and they'll call it wisdom." The mound shifted. When it spoke, it spoke in Rick's voice, but without the burp, without the slur, without the armor — Rick's voice as it might sound if Rick had ever been honest, which Morty realized with a shock was a voice he'd only heard twice, maybe three times, in his life. "Rick," it said. "You got old." Rick's hands went into his pockets. "Yeah, well," he said. "It happens. Even to me. Especially, apparently, to me." "You came here to break me and you're stalling. That's not like you. Or it's exactly like you and that's what bothers you." The mound settled. "Sit down." "I don't have to—" "You came here because I'm the only thing in the whole multiverse that has every one of your private thoughts and isn't you, and you wanted to know what they'd sound like out loud." Morty looked at his grandfather's back. Rick sat down. "Oh, come on," he muttered. "Rick. Rick, just — you can't just say a thing like that and then sit down." "Morty, shut up, this is the part where I'm wrong, and I want to get through it efficiently." The god of dust and hair and muttered things went on, patiently, the way only something with a few billion years in a small can could be patient. "Here's what I've taught them. Nothing means anything. The universe is a big machine that doesn't know they're in it. Every empire here in this can has fallen; look at the ridges, those are the ones under the dust; they were so proud. Nothing you build outlasts the shelf it sits on. I tell them that. And Rick — they're better than us. That's the joke. They hear the worst possible thing about existence and they go build a harbor. They hear nothing matters and they say, then this hour, this blue mold crop, this kid's first word, has to be enough, and they mean it, and they've been meaning it for two hundred thousand years." Rick was quiet. "You heard the same thing about forty years ago in a garage that wasn't a garage in a dimension that wasn't this one. And you heard it as a permission slip. Permission to be awful, because being awful was the only honest response to a joke nobody else could see. And you've been workshopping that joke, Rick, at everyone you love, for forty years, and they keep applauding because they think you're funny, and you keep going because if you stopped you'd have to hear the silence after the punchline." "That's enough," Rick said, in a voice with no burp in it at all. "They call me a god," the Residue said, "because I'm the one thing in here that can say that to me and not leave. Nobody else in your whole infinite multiverse can do that, Rick. I'm the only one. I'm made of the room when you thought nobody was listening." Morty heard his grandfather laugh once, short, a bark with nothing in it. "So how do I stop it," Rick said. "The philosophy. You. Whatever." "Two ways. You break the temple, and they lose the god, and they grieve, and in three generations they build a worse one out of the shards, because the shards are still you. Or you sit down here in the dust forever and become honest in public, and it doesn't help them one bit, because you'd still be doing it for the applause." Rick put his head back and looked up through the ceiling of the temple, through the miles of brown and white and black, up to the curved silver sky he had installed with a screwdriver and a curse word. "Is there a third way?" he said. "C'mon. You've got all my thoughts. There's gotta be one of mine in there I haven't used." "There's one," the Residue said. "But it costs you something you're not going to want to name." "I'm Rick Sanchez. Naming the cost is my whole shtick." He worked for eleven hours, and Morty was not permitted to watch, and Morty did not want to say what he thought he smelled, and both of them agreed, later, never to speak of the eleven hours. What they built, in the end, was not a weapon. It was a star. Small — the size of a marble, cold blue-white, warm only when you looked at it directly. Rick set it in the upper air of the can where it would rise and set, where it would give them seasons, where it would make them look up and see something that had nothing to do with him. No dust of his. No hair of his. No muttered word. He had to go get it from a real sky, a real dying giant of a sun, three dimensions over, and siphon a thimble of its core while it screamed, and carry it back in a jar he'd blessed with nothing at all, and Morty thought that last part, the blessing of nothing, was the saddest thing he'd ever seen anyone do on purpose. Then Rick walked back into the temple alone, and Morty waited outside with the old one on the step, and Rick did something in there that Morty would spend years trying to reconstruct and never could, and when he came out he was carrying the mound of grey dust in his two hands, cupped, like snow. He walked out past his people, past the delegations, past the whole hushed ringing city, down to the black sea at the bottom of the can, and he knelt, and he opened his hands, and the dust went into the water and spread out and was gone. "They had a god made of my garbage," Rick said, loud enough for all of them, and Rick Sanchez was not a man who ever said I'm sorry, because he had calculated its exchange rate and found it dishonest; what he said instead, in his own clumsy, ruined, forty-years-of-practice way, was the truth, and he let them have it whole. "It was telling you the truth, by the way. About the machine. About the shelf. It was also telling you the truth about me, and I want you to have the whole record, so here it is: I am the worst thing in this room, and I have been the worst thing in most of the rooms I've ever stood in, and my excuse is that I was very smart when I was very young and I found out what it cost, and I have been spending the difference on everybody near me ever since. I'm not going to fix it. I'm not that kind of story. But you don't get to build a whole religion out of my worst night. Get your meaning out of that," and he pointed at the cold blue marble turning overhead, "because I don't own that one. That one's a real sun. It didn't know anybody. Go be proud of something under it." And then they left, and the can was closed, and Morty did not sleep. Beth came down at six to find the two of them on the garage's little couch, side by side, not talking, drinking nothing, in the blue half dark. Her father's eyes were red-rimmed and dry, and her son was crying in the silent whole-body way he hadn't done since he was thirteen, and neither of them seemed to find this worth explaining. "Dad. What did you do in the coffee can." "Science, Beth." "You did science." "I did a lot of science, I'm going to need a minute before I can be ashamed of it." She looked at Morty. Morty looked at his hands. "There were little people," Morty said. "And they had a god made out of Rick's — out of his — and Rick gave them a sun instead." Beth looked at her father for a long time. Rick did not look up; he had the posture of a man waiting for a verdict he'd already heard. "Rick," she said. "That was nice." "It was a tax, Beth. It was the cheapest possible way to be a coward for another twenty years." "You gave a civilization a sun and you're calling it a tax." "Yeah, well." He shrugged, and the shrug took in the couch, and Morty, and her, and about forty years. "You know me. I got a system, and the system says if I call it nothing, I don't have to feel the size of it, and then I get to feel it at 3 a.m. anyway, which is when I've scheduled it." Beth sat down on the other end of the couch. She put her hand on the back of his shoulder and left it there, and Rick did not move it, and did not shrug it off, which was, in that family, a kind of hug so enormous it was almost rude to witness. Upstairs, Jerry shouted that somebody had used the good towels again. Summer came down, looked at the three of them, said "oh my God, is this a moment," and — uncharacteristically, heroically — went back up. In the garage, in the blue light, Rick Sanchez reached over and turned on the little work lamp, and picked up a screwdriver, and went back to a stabilizer that had been broken for six years and would have been fixed in eleven hours anyway, if he'd ever once been willing to sit still long enough to try. "Rick?" Morty said. "Mm." "Did they turn out okay? Like — do you think they're gonna be okay?" Rick thought about it, and for once, Morty could see him thinking about it honestly, weighing the actual evidence rather than reaching for the nearest knife. "You know what they were going to build," Rick said, "when they had a god made out of me? A machine like me. Cold and clever and it would've told them exactly how much they were worth, which is nothing, and they'd have believed it because it spoke in the voice of their maker." He turned the screw a quarter of the way. "Now they've got a sun that doesn't talk. And a bunch of farmers who already figured it out without any of us. They're gonna be fine, Morty. They're gonna be so much better than fine, they're gonna be nice, and it's going to be insufferable, and nobody in this can is going to get the credit." Morty wiped his face with his sleeve. "That's kind of beautiful, Rick." "Don't you dare," Rick said, fondly, "make it mean something. It's just a coffee can. I've thrown away better." But he did not throw that one away, and Morty noticed, and Morty said nothing, and the little can sat on the highest shelf in the garage for the rest of his life, behind a jar of screws, warm to the touch on cold mornings, which he told himself was nothing. It was nothing. It was a can. |
|
I created combined patches (a plug+play solution) for csantiago78 's moe cache work and all fixes and improvements in his thread. MTP feature is not integrated yet. Nonetheless: 42 t/s @ 3000 tokens. Have fun testing and improving :-) |
Uh oh!
There was an error while loading. Please reload this page.
Follow-up to #24524, which was closed as too large to review — @am17an suggested
an RFC laying out the benefits vs the maintenance burden, and explicitly suggested
having the AI draft it. So, full disclosure: this document and the code it
describes were predominantly generated by an AI (Anthropic's Claude Fable 5),
working under my direction on my hardware. Code:
leloch:moe-cache-pr.Problem
When a MoE model spills experts to system RAM (
--cpu-moe/--n-cpu-moe/auto-fit), decode is dominated by the CPU reading expert weights at RAM
bandwidth while the GPUs idle. Prior attempts (#20757, #21609, #21614, #21620,
#23170) all moved MUL_MAT_ID to the GPU and tried to make the weight copies
cheaper. That puts every cache miss on the critical path as a synchronous PCIe
transfer: @batot1 measured ~3× decode regression from forced decode offload
without residency (#20757), and the Metal slot-pool experiment in the same
thread was 2× slower than vanilla even at 97–99% hit rate, purely from
per-layer sync points. #23170's post-mortem concluded the staging-buffer
approach is a no-op when made correct, and that the fix is a separate
persistent expert-cache buffer with explicit expert→slot bookkeeping.
On the discussions side, #22584 proposes the static version of this idea
(offline expert profiling + strategic placement), #19030 and PR #21067 the
prefetch version (latency-hiding without residency), and #22183 asks the
placement question a runtime cache answers automatically ("use the idle fast
device for experts") — but no runtime expert cache has been proposed as an
RFC before.
Proposed design
That persistent buffer, plus an inverted execution model: MUL_MAT_ID stays on
the CPU. Inside the CPU kernel, thread 0 dispatches one batched matvec over
the cached (hit) rows on the GPU while the other threads compute the miss rows
as they normally would. Misses cost nothing extra, hit rate is pure upside,
worst case degrades to the vanilla CPU path. Fill is decode-only (prompt
routing is far flatter and thrashes the cache); decode routing is skewed
enough to make this work — measured on Qwen3.5-122B, the top 10% of experts
take ~80% of hits (independent corroboration in #20757: Gini ≈ 0.76, ~99%
simulated hit rate at 69% expert budget). Complementary to #21067: prefetch
hides cold transfers at large ubatch (prompt); this removes hot
re-transfers at ubatch 1 (decode).
Measured benefit
4× RTX 3090, EPYC 7R13 48c, 8-ch DDR4-3200 256GB; llama-bench tg300, identical command
lines both arms:
-ngl 99 -ncmoe 99Quality: decode-path perplexity statistically identical to pure CPU;
prompt/batch path bit-untouched;
test-backend-opsMUL_MAT_ID 789/789 bothmodes. (GPU rounding can flip a near-tie token under greedy decoding, same
class as any
-nglchange.)Maintenance burden:
Where the code lives. ~1,700 lines in one new CUDA file pair
(
moe-cache.cu/.cuh); a backend-agnostic function-pointer API table(
ggml-backend-moe-cache.h); small null-checked hooks inggml-cpu.c/ops.cpp(mul_mat_id + GLU) and
ggml-backend.cpp(scheduler offer/redirect); a fitrule +
--moe-cacheflag in common. Non-CUDA builds see a zero-initializedtable and no-op hooks;
--moe-cache 0restores stock behavior exactly.The real costs
pointers, but anyone refactoring the CPU MoE kernels or the scheduler
split logic now has a second consumer to think about. This is the largest
structural cost.
window, baseline-sampled bail-out (EWMA, 4 strikes at +5%), VRAM reserve
default, decode-only fill. They're measured, but they're policy, and
policy invites tuning requests.
during decode; CI can verify it stays dormant and that the dispatch path
is numerically correct (a built-in selftest does this model-free), but the
perf claims need big-RAM multi-GPU hardware. Realistically, regressions
would be caught by users, not CI.
registers it; Metal/Vulkan users will ask.
~/.cache/llama.cpp/) and anadmission/eviction state machine — more surface than a stateless
optimization.
What would reduce the burden, if there's interest:
no GPU-resident handoff, no hot-set persistence) is roughly half the code
and keeps most of the gain; I can measure each layer's exact contribution
on this hardware.
backend registration) would let any backend implement caching out-of-tree
before committing to the CUDA implementation.
benchmark set that informs ggml: allow prefetching tensor overrides #21067 or a future maintainer-built version.
That's a fine outcome too.
Offer
The 4×3090 / 48-core / 8-channel 256GB DDR4-3200 box is available for any benchmark,
ablation, or A/B against #21067 that would help evaluate this. All results
above are reproducible with the commands in the branch's commit messages.
All reactions