Replies: 7 comments 41 replies
|
This repo ignores all GPUs except the first one. |
|
Update: prior art I missed when posting this RFC When I posted this RFC I was unaware of @csantiago78's #27861 ("GPU-resident LRU cache for host-offloaded MoE expert weights", opened Aug 28) — that's a miss on my part and worth correcting explicitly: the idea was already in flight. Having read it in full, the two are independent implementations of the same insight with different design points:
Their async-upload design and the per-step upload throttle (--moe-expert-cache-inserts) are worth adopting regardless of where this converges. One caution from my side for any remap-based design: the ids→slot mapping table must be anchored in the split that consumes it — ours initially consumed the previous ubatch's table through a view the scheduler's boundary check couldn't see, which is deterministic corruption (evidence chain linked in the post). The flag name now collides; whether to converge the implementations, keep one as reference, or pick a naming scheme is the maintainers' call — happy to rework either way. |
|
Glad the working-set rule made it into the design. Four things from the same single-5090 / Windows box, all measured with the July tooling and reproducible from it — the first one is for @Volunteer-1's report, the other three are for the flag itself. 0. "Slows down as context grows" is what a VRAM pool does when it eats the KV cache's headroom. @memoriaru is right that the per-token overhead is context-independent, so a slope with context points elsewhere — and on my box the "elsewhere" has a number. Two effects stack: attention cost grows with context in both arms (so the cache is not expected to invert that), and the pool takes VRAM that the KV cache would otherwise grow into, so the run reaches the paging line earlier. That line is absolute, not a fraction: stealing VRAM from a loaded server with an external tensor, 858 MiB free costs −1 % decode; 775 MiB free costs −11 % decode and 2.3× TTFT — 83 MiB separate "unnoticeable" from "broken". On a dense 27B with only 1. Size N under a VRAM budget, not a hit-rate target. With the #27861 build on Step-3.7-Flash, 2. Where the pool stops paying is predictable from the routing histogram, not from a sweep. Coder-Next puts ~80 % of hits in ~28 % of experts and gain tracks weighted-traffic coverage almost linearly; gpt-oss and Step-3.7 are flat and doubling coverage bought +2 % and +5–12 %. DeepSeek-V4-Flash is flatter still: per-layer routing entropy 6.3–7.9 bits out of 8.0 (256 experts) on a 139k-token code+prose profile, and 90 % of activation mass needs 5,854 of the 11,008 routed experts (~72 GiB at Q4). A pool that fits in one card's VRAM cannot reach the knee on this model, and a static hot list was worth ~nothing there (+3 %, within noise). That is consistent with the < 20 % @Volunteer-1 sees. A cheap entropy/Gini per layer from a short profile run (the new 3. Perplexity is not the acceptance test. At bias 0 the replica arms of my fork move the argmax on 2 % of code tokens and ~6 % of Portuguese prose tokens while PPL moves < 1 % — measured with Static split vs LRU pool on the same model, machine and slots (Step-3.7, #27861 build, 61k-token document): at depth the LRU matches the static split (+13 % vs +12 % over stock), and on a 512-token generation it pulls ahead (+20 % vs +13 %, single run) — the dynamic policy does earn its keep. Which is why I'd argue N should end up as an output, not an input: swap only when |
|
Hi, thanks for working on this! I did some tests but I did not get the speedups I was hoping for. First, the system specs: I am testing build 9ff189f (10721) using Qwen3.6-35B-A3B-UD-Q6_K_XL , all speeds are token generation. Measurements
40 means all experts are on the CPU.
160 needs slightly less VRAM than n_cpu_moe = 13, 196 is too much, but it still runs.
Observations:
Questions:
The commands I used:
|
|
Update: telemetry fix confirmed, a silent-pitfall fix, and two test-matrix corrections Thanks @TmDagger — the double-counting report was correct, and digging into it uncovered something much bigger downstream. What landed since the last update: 1. Telemetry double-count — confirmed and fixed (f75b069) The remap 2. Why the old test matrix never saw any of this Two blind spots, both fixed in the harness: the ids read-back consumed a scheduler copy that hadn't been uploaded yet for this ubatch, so the update always read stale routing; and all experts shared identical weights, so any mapping error was invisible to the bit-exact comparison. Weights are now distinct per expert. 3. The real find: Symptom: illegal-memory-access on CUDA with MXFP4, and — much worse — silent all-zero routing on Metal that still passed the tests. Mechanism, in three steps:
A control experiment (skip the reserve, keep everything else) flips all three symptoms off at once: correct ids read-back, table upload restored, anchor hits everywhere. Upstream's own call sequences avoid this (reserve uses a dedicated measure graph, decode goes through alloc+compute), but nothing warns you off Fixes:
Worth flagging for upstream: 4. A dispatch-matrix surprise Our 5. CUDA results The full 96-case matrix now also passes on CUDA (RTX 4090, container CUDA 13.1): 6 quant types × nt ∈ {1, 3, 8, 9} × cold/evict/hit/graph-rebuild rounds, including MXFP4 and the first real coverage of the MMQ ids path at nt=9. 6. For the field testers Since several of you have been running this on real workloads, a few notes for whenever you next update your builds — nothing urgent, just things worth a look:
(Clarifying the @JigSawPT bullet above: our CUDA matrix run is already in — what it can't answer is whether the Step-3.7-Flash crash is this same issue, since that path needed a real-model re-test on Blackwell.) |
|
Two things from the single-5090 Windows box, both on b0096df built with CUDA 13.0 ( 1. The Step-3.7-Flash crash no longer reproduces. Same configuration that segfaulted for me on 11/09 -- 2. Two pool symptoms on the same build, depending only on At So even when it runs, it finds nothing to manage -- with 34 layers' MoE weights on the CPU and the flag accepted. At
Ruled out from this end: ( Leaving this one with you -- I am deep in another piece of work and will not be chasing it further from this end. None of that changes the v2 write-up: I am still in, with the PCIe 5.0 data point, the routing histograms and the |
|
I've been investigating a related problem on a single Ryzen AI MAX+ 395 / 128 GB Strix Halo system, but at a larger model-to-memory ratio: host-managed routed-expert residency where the complete model cannot be resident. I have validated single-node execution of Qwen3-235B-A22B Q4_K_M and subsequently Kimi K2.5 UD_Q2_K_XL (~375 GB) using an NVMe-backed expert path. One thing your results reinforce strongly is that residency hit rate alone isn't enough — transfer/synchronisation overhead can dominate the useful expert compute. My implementation takes a CPU-authoritative approach and streams only routed expert records, with explicit residency/lease tracking. The Kimi experiment is slow rather than a throughput competitor, but it may be an interesting data point for the discussion because the model is roughly 3× the machine's physical memory rather than simply overflowing VRAM. I've published the methodology/results here if useful for comparison: https://zenodo.org/records/22730031 I'd be particularly interested in comparing routed-expert working-set distributions and cold/warm residency behaviour. |
Uh oh!
There was an error while loading. Please reload this page.
This follows up on #24528 (leloch's CUDA-side adaptive cache, whose PR #24524 was closed for scope) with a different design point: zero kernel changes via id remapping, plus a prefill fallback rule
RFC: Persistent expert slot pool for MoE CPU offload (
--moe-expert-cache)Summary
Add
--moe-expert-cache N(-mec N): for every MoE expert weight tensor that offloadingplaced in host memory, keep a persistent pool of
Nexpert slots in accelerator memory.Cache hits serve a decode step with zero host-to-device traffic; misses copy the
expert in and evict the least recently used slot. Expert ids are remapped to slot ids
through a per-tensor map table applied with
ggml_get_rows— no new ggml op and nobackend-specific code — so CUDA, Vulkan, Metal and CPU all work unchanged.
Measured on a single RTX 4090 24 GB, Qwen3.8-Flash-Next (qwen4exp, 48 MoE layers, all
experts on CPU, UD-Q3_K_XL, 8 K context, greedy):
On a real code-generation workload served through
llama-server(173-token prompt,512 generated tokens, warm rounds): 12.82 → 14.10 (+10%) at N=32 and → 17.14 (+34%)
at N=64; a long multi-task session (2561 tokens with pool-churning 1 K-token
generations interleaved) held anchor outputs byte-identical across four runs with
no throughput decay. Greedy decoding with the pool active is perplexity-equivalent
to the host-copy path (PPL 3.3096 ± 0.053 vs 3.3262 ± 0.053, deterministic reruns;
difference well inside the error bars).
Prefill is unchanged within noise by construction (below). A backend-level matrix test
passes bit-exact on CUDA (RTX 4090) and Metal (AMD dGPU): 5 quant types × dual pools
sharing one routing × ubatch sizes 1/3 × cold / full-eviction / hit+reload /
graph-rebuild. Server-level greedy A/B initially diverged deterministically from the
host-copy path; root cause was found and fixed (details in Validation — including one
invalid verification round we are disclosing rather than hiding).
Problem
With
--cpu-moe/--n-cpu-moe(or the auto-fitter), decode re-streams every selectedexpert over PCIe on every token: the selective expert copy added in #15346 removed the
unused experts but still copies the same hot experts token after token. Routing is
skewed, so for code-generation-style workloads the same small expert set is re-read
thousands of times. #20757 requests a cache for exactly this.
Why not the other approaches on the table
--prefetch-weights(#21067)overlaps next-layer transfers with compute but cannot know the next layer's routing;
for MoE it was measured moving 2.06× the bytes with +47.8% TTFT. Complementary, not
competing: prefetch targets dense/prefill transfer overlap, this targets decode
residency. The two can share the copy-stream infrastructure.
evict the decode hot set; a frequency-gated admission filter was needed to recover.
This design sidesteps the problem structurally: a ubatch that can touch more
distinct experts than the pool has slots never goes through the pool at all — it
takes the existing selective-copy path. Prefill therefore behaves exactly like
master (verified: pp512/pp2048 unchanged), and prefill never evicts decode state.
this branch is deliberately small in backend impact: zero kernel changes, one
scheduler hook, ~700 lines total of which ~350 are the scheduler core.
Design
ggml_backend_sched_register_expert_pool(sched, w, backend_id, n_slots, &table)allocates
n_slots * expert_size(+ a small NaN-safe tail) on the compute backendand a host-side I32
map_tableshaped[1, n_expert](flagged as a graph input).llm_graph_context::build_lora_mm_id— the single funnel for all MoE expertmatmuls — routes through the pool when one is registered for the tensor and the
ubatch cannot select more distinct experts than slots (
n_expert_used * n_tokens ≤ n_slots). The remap isget_rows(table, cont(ids)): existing ops only.so the remap
GET_ROWSalways anchors its own split. In that split's prologue thescheduler reads the (original) expert ids back to the host, updates LRU state,
issues async H2D copies for misses and rewrites the map table — before the same
split's input copies upload the fresh table. (This ordering is the fix described
below; anchoring the update anywhere later lets the remap consume a stale table.)
time (tensors beyond the budget silently keep the selective-copy path); pools are
disabled under pipeline parallelism; slot-overflow is a hard assert instead of
silent corruption.
Slot sizing:
Nmust cover the decode working set, not just top-k —N = top-kis afull-miss worst case and measurably slower than master (the per-layer id readback
synchronization has nothing to buy). Start at 2–4× top-k and scan; gains are
routing-skew dependent (code workloads benefit most, flat-routing models least).
Validation — including a bug we found, mis-verified once, then actually fixed
The full evidence chain (raw captures, the invalid round, the fix, the bypass control)
is archived in
memoriaru/llama-cpp-expert-pool-stale-table-fix.
tests/test-expert-pool.cppruns the pooledMUL_MAT_IDand the regular host-copy path in the same process — Q2_K/Q3_K/Q4_K/Q6_K/Q8_0 × two pools sharing one routing (fused gate_up + down shape) × ubatch 1/3
× cold/evict/hit/rebuild. 40/40 on CUDA and Metal.
CPU:
-mec 0is run-to-run byte-identical (three runs across days and codeversions);
-mec 32diverged at byte 17 on a coherent near-tie ("wants" vs "needs")and stayed deterministic across environments and code revisions.
showed mec0/mec32 agreeing was produced against the wrong server instance — the
-mec 32launch had failed (No such file or directoryinf32.err) and bothcurls hit the still-running
-mec 0server (visible retroactively in the responsetimings:cache_n = 830on what should have been a cold prompt). Lesson adoptedin the test guide: verify the instance (startup log, prompt cache counters) before
trusting an A/B pair.
GET_ROWSreads the map table through aRESHAPEview;views carry no buffer, so the pooled-split boundary check (which keys on src
buffers) could not see it, and the remap could land in an earlier split than the
pool-update prologue — consuming the previous ubatch's table. Every cache miss
then read whatever expert last occupied the remapped slot (~20–30% wrong weight
reads per step on a 512-expert model with 32-slot pools). Small enough to stay
coherent, deterministic enough to reproduce byte-for-byte.
[1, n_expert]and consumed directly (no view),its buffer is registered with the boundary check so the remap anchors its own
split, and the pool update runs at that remap, before the same split uploads the
fresh table. The bypass control — pools allocated but the graph on the host path —
is byte-identical to
-mec 0; with the pool active the A/B now agrees for 353bytes and then flips once on a near-tie (byte 354, deterministic), consistent with
llama.cpp's known sensitivity of kernel selection to buffer placement (the same
class of near-tie flips observed when changing
-ngl/tensor placement). Theperplexity check quantifies the residual: PPL 3.3096 ± 0.053 with the pool vs
3.3262 ± 0.053 without (8 × 4096-token chunks, deterministic reruns) — the
difference is well inside the error bars, i.e. quality-equivalent.
the LRU state between four replays of the same anchor prompt: anchor outputs
byte-identical (same SHA-1 four times), throughput stable (13.9–15.3 t/s).
Maintenance footprint
~700 lines:
ggml-backend.cpp(+~350: pool struct, register/update, split hooks),llama-graph.cpp(+~20 remap),llama-context.cpp(+~70 registration & budget),argument plumbing,
llama-benchflag, and the self-contained matrix test. No changesto any backend, no new ops, no allocator changes.
Limitations / future work
Validation §5); we propose treating it like the variance already accepted for
-nglchangescomposable later)
Reproduce
AI usage disclosure: this feature was developed with AI assistance (analysis, code and
benchmarks); the author ran, verified and debugged all results on their own hardware
and has reviewed every line — including writing the fix for the bug the validation
uncovered.
All reactions