Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
508 commits
Select commit Hold shift + click to select a range
127a069
design: VLLM_FQ_CAPACITY_UTILIZATION — gpu-max-util analog for tier h…
malaiwah Aug 10, 2026
b82d765
m4-swap: T3 PASS (maps read as data under CUDA graph — atomic path li…
malaiwah Aug 10, 2026
1b6b6af
loader-v2: trust + lazy-encode design/usage note — VLLM_FQ_SOURCES[_M…
malaiwah Aug 10, 2026
87c7914
0c: fix layer extraction in window publish (tr3 prefix collision)
malaiwah Aug 10, 2026
0d47739
0c: robust window publisher (fresh temp staging per K, no stale-glob …
malaiwah Aug 10, 2026
8547df7
0c: window-2 ring — capture layers 11-18 from preserved boundary, the…
malaiwah Aug 10, 2026
0d9fffb
docs: 14-build-findings — what the build taught us (addendum to 00-13)
malaiwah Aug 10, 2026
968ff1f
docs 00-13: additive build notes so the spec stops contradicting the …
malaiwah Aug 10, 2026
118cbed
runs: README as an evidence index; research README points at runs/ + …
malaiwah Aug 10, 2026
8a2e107
m0-seed: HF card corrected against the evidence; 14 picks up the fq_v…
malaiwah Aug 10, 2026
c7e942f
docs: normalize relative paths in the build notes; 11 picks up the ex…
malaiwah Aug 10, 2026
b0c65aa
0c: priming-report — window-1 community-quant extraction sealed
malaiwah Aug 10, 2026
923ad18
fq_verify: prove the community-reassembly claim (byte-identity + nume…
malaiwah Aug 10, 2026
cbb697d
test_fq_verify: probe for fq_assemble --trust-signer instead of assum…
malaiwah Aug 10, 2026
3848253
m4-swap: T5 PASS (no torn state observable, 6/6 abort points) + T6 PA…
malaiwah Aug 10, 2026
91556dc
card: prominent tools-repo + research links at the top; push the doc-…
malaiwah Aug 10, 2026
280ce6f
fq tools: peer-review findings 1, 2, 4, 6 — record for b29c0621f
malaiwah Aug 10, 2026
1714fd1
0c: K2-to-completion ring (operator priority) — layers 19-78 in 8-lay…
malaiwah Aug 10, 2026
5c98db2
card: all four tiers in the title (K4 arrives via community priming);…
malaiwah Aug 10, 2026
e4d5fb2
fq tools: mirror fq_fetch / fq_release / fq_trust from the public repo
malaiwah Aug 10, 2026
106d499
card: retire two stale known-issues (manifest last-writer-wins + K2/K…
malaiwah Aug 10, 2026
86c7170
M1/M2: overhead A/B + dryrun evidence — collector costs 3.7-5.0% on t…
malaiwah Aug 10, 2026
f53d3f6
0c: self-driving campaign supervisor — runs unattended for the life o…
malaiwah Aug 10, 2026
2bbd550
0c: supervisor single-owner lock + concurrent-encode guard
malaiwah Aug 10, 2026
ca5bcab
0c: campaign could not publish — supervisor ran the publisher without…
malaiwah Aug 10, 2026
40f93e3
0c: prune each window's capture right after ITS publish (was: only af…
malaiwah Aug 10, 2026
ce09fda
fq_assemble: --force can no longer eat a directory it did not create
claude Aug 10, 2026
35ad646
fq_verify: pin the signer, and stop calling "unverified" a pass
claude Aug 10, 2026
c29e306
fq_prime spot-check: absent evidence is a failure, and JSONL means al…
claude Aug 10, 2026
a333c56
m0-seed: HF card is the front door now — fix what was wrong on it, ad…
malaiwah Aug 10, 2026
9fd871e
fq tools: mirror the atomic fq_release publish mode from the public repo
malaiwah Aug 10, 2026
691e4a2
fq_fetch: authenticate the header the plan is built from
claude Aug 10, 2026
b2d56b0
0c: defer capture prune while any encode is reading it
malaiwah Aug 10, 2026
ff799b0
card: name the fq-manifest MISMATCH that incremental publishing causes
malaiwah Aug 10, 2026
fd89010
fq_eps: commit the budget-conserving uniform baseline (was uncommitted)
malaiwah Aug 10, 2026
0f8245c
fq_repack: sign the segment header digest (sync from public) — 194 pa…
malaiwah Aug 10, 2026
f4579dc
0c: per-tier encode guard — different tiers may run concurrently on d…
malaiwah Aug 10, 2026
a119b88
0c: defer capture prune per-window, not globally
malaiwah Aug 11, 2026
3671943
0c: GPU reservation file so a loading serve is not raided
malaiwah Aug 11, 2026
56c3838
m5: swap-evidence recorder (metrics timeline + domain-shift workload)
malaiwah Aug 11, 2026
dff36b9
m5: GLM-5.2 TP4 serve script (off/dryrun/live modes)
malaiwah Aug 11, 2026
3c51796
m5: dependency-free SVG evidence charts + structural tests
malaiwah Aug 11, 2026
afdc6b7
m5: evidence-campaign orchestrator
malaiwah Aug 11, 2026
7380fab
health: rewrite sweep for the current pipeline + growth detection
malaiwah Aug 11, 2026
b5c5d73
health: honest sweep labels (idle/exited/quiet vs STALLED)
malaiwah Aug 11, 2026
a299070
rebase: upstream overlap analysis for the PR duplicate-work check
malaiwah Aug 11, 2026
0ac6eae
m5: default serve window to 32k for honest eval numbers
malaiwah Aug 11, 2026
8a389f4
0c: never regenerate segments the remote already has
malaiwah Aug 11, 2026
9c946f6
docs(rebase): map vLLM fork hierarchy and GG v20 r34 provenance
malaiwah Aug 11, 2026
9bfd50f
rebase: decided PR target, and the pre-filled-base trap
malaiwah Aug 11, 2026
12fbf3d
health: track the newest m5 serve log, not a hardcoded tag
malaiwah Aug 11, 2026
8866774
m5: adversarial peer review of the evidence recorder + charts, and it…
malaiwah Aug 11, 2026
8e1810f
m5 review: MIN-8 (sweep's hardcoded m5 tag) was fixed concurrently in…
malaiwah Aug 11, 2026
454fe2e
m5-serve: reference Coder quant map + MTP78 corpus loader + convergen…
malaiwah Aug 11, 2026
6b00627
m5: assemble + verify pure-K3 and mixed-K3/K5 GLM-5.2 checkpoints fro…
malaiwah Aug 11, 2026
ba09cff
m5: fix four peer-review criticals before the evidence run
malaiwah Aug 11, 2026
dc0e45a
m5: M0 boot gate PASSES on the real GLM-5.2
malaiwah Aug 11, 2026
89e35b3
m5: convergence scorer — does runtime routing rediscover a human's qu…
malaiwah Aug 11, 2026
384ea5d
m5: K5 mixed tier exceeds SM120 shared-memory limit (measured)
malaiwah Aug 11, 2026
730e108
m5: scenario-1 boot policy (observe mode)
malaiwah Aug 11, 2026
b48409c
0c: pivot tier order to K2 -> K4 -> K5
malaiwah Aug 11, 2026
c889d7a
gg-env: deploy script — the serve runs the ROOTFS copy, not the sourc…
malaiwah Aug 11, 2026
26e50d0
m5-serve: spec the forced-retier admin API (design only)
malaiwah Aug 11, 2026
135a1b4
fq/m5: answer "is the fungible-quant design topology neutral?"
malaiwah Aug 11, 2026
d3d8aee
m5: first convergence number on real GLM-5.2 (preliminary)
malaiwah Aug 11, 2026
9bf45a5
m5: tensor-loader compatibility matrix for the fungible-quant integra…
malaiwah Aug 11, 2026
603f39a
m5: loader-compatibility — note the suite count drifts with concurren…
malaiwah Aug 11, 2026
47c3926
m5-serve: report — a missing K bpw never crashes the serve
malaiwah Aug 11, 2026
3aa5282
fq/m5: bind real gate mass in the collector; deploy base_router.py
malaiwah Aug 11, 2026
f1223db
m5: convergence on the real MTP78 corpus — 0.3938, 1.48x chance, 59% …
malaiwah Aug 11, 2026
4bd590d
m5: move raw routing dumps to HF, keep derived scores in git
malaiwah Aug 11, 2026
f2e288b
m5: RETRACT the corpus convergence numbers — driver replayed truncate…
malaiwah Aug 11, 2026
e1dd7ec
m5: demo-1 runner, per-expert K4 coverage discovery, honest scope
malaiwah Aug 11, 2026
c46f8b4
0c: disk guard must WAIT on its window, not advance past it
malaiwah Aug 11, 2026
b44a4f2
0c: MoE layers are 3-77, not 3-78 (layer 78 is MTP)
malaiwah Aug 11, 2026
415277a
health: coverage denominator is 75 layers (3-77), not 76
malaiwah Aug 11, 2026
f343cd0
m5-serve: admin API spec — "As implemented" section
malaiwah Aug 11, 2026
713e6c7
m5: corrected corpus convergence — 0.4176, 1.57x chance, 62% of human
malaiwah Aug 11, 2026
f2ac7f5
m5-serve/heatmap: design decision for the live MoE expert heatmap
malaiwah Aug 11, 2026
7679f39
FQ: activation-matrix endpoint spec (GET /fq/heatmap)
malaiwah Aug 11, 2026
0b2d8e5
0c: clear stale capture state when a window's data was pruned
malaiwah Aug 11, 2026
b88d7f4
FQ heatmap spec: reconcile the reset control with DESIGN.md
malaiwah Aug 11, 2026
d68323c
0c: recover from PARTIAL captures left by preemption
malaiwah Aug 11, 2026
6c9a313
m5: 4-axis corpus replay driver
malaiwah Aug 11, 2026
47d745b
m5: adversarial review of the demo-1 glue — 5 defects, fixed with tests
malaiwah Aug 11, 2026
e70896c
deploy-fq.sh: deploy entrypoints/serve/__init__.py too
malaiwah Aug 11, 2026
766770d
m5/heatmap: 4-panel per-axis activation figure + pairwise overlap matrix
malaiwah Aug 11, 2026
a3d2b73
m5/heatmap: one permutation row per line in the JSON sidecar
malaiwah Aug 11, 2026
37c6425
m5-serve/heatmap: operator heatmap page, self-contained, with maths +…
malaiwah Aug 11, 2026
fd3bb5e
m5-serve/heatmap: interop tests against the real endpoint's own encoder
malaiwah Aug 11, 2026
13ab2b1
m5: four-axis convergence — count beats mass, code axis is not special
malaiwah Aug 11, 2026
5975078
m5-serve/heatmap: adversarial review — three "pretty and wrong" fixes
malaiwah Aug 11, 2026
7f49140
m5: flagship 4-axis figure from REAL data + pairwise overlap matrix
malaiwah Aug 11, 2026
50a7b63
m5: GSM8K quality baseline — 89.2% on the assembled 3.0bpw checkpoint
malaiwah Aug 11, 2026
a5e0e91
0c: reclaim disk after encode, not only after publish
malaiwah Aug 11, 2026
8c0bca5
m5: demo-1 progressive boot — segments + policy, no assembled checkpoint
malaiwah Aug 11, 2026
4e7f46e
pr: upstream submission materials for fungible quant — not opened
malaiwah Aug 11, 2026
9ed57c8
pr: pin the test number to the committed tree, inventory the renders
malaiwah Aug 11, 2026
a8cad62
m5-serve/heatmap: first real renders + headless browser toolchain
malaiwah Aug 11, 2026
208cca6
m5: progressive boot degrades instead of dying; ranged reads retry
malaiwah Aug 11, 2026
62f5c1e
m5-serve: rewrite the HF model card for the FQ segments repo
malaiwah Aug 11, 2026
03dcf22
docs(fq): record upstream issues #282 and #283
malaiwah Aug 11, 2026
92c4fdd
m5: serve-demo1 deploys before booting — stale rootfs has cost 3 boots
malaiwah Aug 11, 2026
255096b
m5: refuse to boot onto occupied GPUs
malaiwah Aug 11, 2026
d148254
m5: reap stale GPU holders by DEVICE, not by argv pattern
malaiwah Aug 11, 2026
8b0b452
m5: enable hf_transfer for whole-segment prefetch
malaiwah Aug 11, 2026
09b5f82
m5-serve: Xet high-performance transfer, engine timeout, prefetch depth
malaiwah Aug 11, 2026
22dab4c
m5-serve: authenticate Hub downloads, report prefetch config at boot
malaiwah Aug 11, 2026
91392d8
m5-serve: document the progressive download path for the PR
malaiwah Aug 11, 2026
cf29bc8
Commit completed pre-preemption work: HF card edits, KV addendum, adv…
malaiwah Aug 11, 2026
7ceef25
m5-serve: record measured bulk-fetch throughput (~190x) and the defec…
malaiwah Aug 11, 2026
9375e2c
m5-serve: expose VLLM_FQ_KEEP_LAYERS and report it in the boot banner
malaiwah Aug 11, 2026
fabbb1f
m5-serve: record what a cold progressive boot actually costs
malaiwah Aug 11, 2026
226be64
gg-env: mirror fq_reload with the fq_converge_layers RPC
malaiwah Aug 11, 2026
f90c708
tools/fq_reload: fq_converge_layers RPC (and drop a duplicate copy)
malaiwah Aug 11, 2026
22dd6f1
tools/fq_reload: fq_converge_layers must not imply an install it neve…
malaiwah Aug 11, 2026
8cc73bc
m5-serve: floor-triggered fragment-cache guard
malaiwah Aug 11, 2026
28cc444
m5-serve: first complete progressive load, and the envelope that pric…
malaiwah Aug 11, 2026
35e1368
0c-campaign: clean pause once K4 reaches 75/75
malaiwah Aug 11, 2026
605fb32
m5-serve: battle-test plan, FQ_FAST loader-iteration mode, warm cache…
malaiwah Aug 11, 2026
44dee16
health/sweep: watch the log the serve actually writes, probe the real…
malaiwah Aug 11, 2026
4f5d3b8
prune-fragment-cache: tiered floor, absolute cutoff, no silent no-op
malaiwah Aug 11, 2026
68fd23a
health/sweep: count distinct layers on one rank, not matched lines
malaiwah Aug 11, 2026
626d3d9
m5-serve: bt_metrics — make the hot-restart claim falsifiable
malaiwah Aug 11, 2026
199587b
BT-1 PASSES: GLM-5.2 serves from Progressive Tensors segments
malaiwah Aug 11, 2026
30426dd
bt_metrics: count cache hits as their own segment origin
malaiwah Aug 11, 2026
c29decc
BT-2 PASSES + retract the allocator-residue claim
malaiwah Aug 11, 2026
b868012
BT-6 FAILS: the loop decides swaps it cannot apply
malaiwah Aug 11, 2026
ce5ee5f
preflight: enforce on calibrated projections, correct two bad constants
malaiwah Aug 11, 2026
b316193
BT-1-AND-2: correct the stale 8.06 GiB overhead figure to the measure…
malaiwah Aug 11, 2026
166441d
m5-serve: verify_retier — never trust the swap API's own answer
malaiwah Aug 11, 2026
df21879
verify_retier: pair the promotion with a demotion automatically
malaiwah Aug 11, 2026
55937bd
m5-serve: document the layer-78 asymmetry as one constraint, not four…
malaiwah Aug 11, 2026
6a51759
health/sweep: keep a progress signal on warm boots
malaiwah Aug 11, 2026
40dc302
BT-1-AND-2: the zero-residue result is now confirmed 16 times over
malaiwah Aug 11, 2026
eb5afb6
m5-serve: M4 swap verified on a live serve (evidence)
malaiwah Aug 11, 2026
887a962
growth_supported: design, and jaccard_trace to settle the guard question
malaiwah Aug 11, 2026
1cd3fc6
growth: adopt A+B — one slot per layer, one operation in flight, both…
malaiwah Aug 11, 2026
a8173d8
heatmap: decode fq-heatmap/1 properly, and a finding that bounds the …
malaiwah Aug 11, 2026
504cf06
FINDING: GLM-5.2's routing is too flat for online top-K re-tiering
malaiwah Aug 11, 2026
7476fc2
FLAT: the domain change is invisible in the top-K churn
malaiwah Aug 11, 2026
dd154e5
heatmaps: commit all 9 rendered samples for review
malaiwah Aug 11, 2026
7e42e09
FLAT-1 preliminary: 15x the sample buys +0.014 Jaccard
malaiwah Aug 11, 2026
a967b9b
SELECTION SIGNAL: routing frequency is right, and the encoder's data …
malaiwah Aug 11, 2026
eb57339
R10: the capture is documented, and the allocator supersedes my recon…
malaiwah Aug 11, 2026
6c690de
ROUTING-FLATNESS: retraction banner — the instability half was a guar…
malaiwah Aug 11, 2026
65383c8
serve: FQ_TAG/FQ_DEVICES_ENV so two instances can run concurrently
malaiwah Aug 11, 2026
5d22976
LANGUAGE MOVES THE EXPERT SET — 0.321 vs a 0.879 same-language control
malaiwah Aug 11, 2026
07e7c52
health/sweep: report BOTH instances, and whether apply actually insta…
malaiwah Aug 11, 2026
0e57191
reap-devices: kill by DEVICE, because argv matching has now failed bo…
malaiwah Aug 11, 2026
815cde1
heatmap format: benchmarked 7 encodings; the win was gzip, and it was…
malaiwah Aug 11, 2026
6f43252
The guard fix is vindicated live: jaccard 0.61 -> 0.931 and climbing
malaiwah Aug 11, 2026
ec803f2
serve: calibrate VLLM_FQ_JACCARD_FLOOR to 0.80 from measurement
malaiwah Aug 11, 2026
e909d90
Scenario 2: flat-K3 policy, and the reserve arithmetic that bounds it
malaiwah Aug 11, 2026
da132f8
serve: stop discarding an explicitly-passed CUDA_MODULE_LOADING
malaiwah Aug 11, 2026
ce8f19d
BT-6 PASSES: 64 swaps INSTALLED into a live model, verified from state
malaiwah Aug 11, 2026
969c01b
health/sweep: count installs by moment, not by rank log line
malaiwah Aug 11, 2026
8a35110
SESSION-STATE: resumable snapshot before context compaction
malaiwah Aug 11, 2026
57d5981
fq: tensor-level FQ feasibility — online per-slice re-tiering is not …
malaiwah Aug 11, 2026
5f314ef
fq: run artifacts through the 19:31 preemption
malaiwah Aug 11, 2026
68eeab8
fq: the two-tier limit is a code constant, not a hardware one — measured
malaiwah Aug 11, 2026
03e2bc6
fq: cold-start handoff doc, and the concurrent-boot OOM that explains…
malaiwah Aug 11, 2026
e3f3e19
fq: HANDOFF — the complete open-work register, all 39 items
malaiwah Aug 11, 2026
e215331
BT-6c PASSES: the loop installs repeatedly, invalid swap = 0
malaiwah Aug 11, 2026
3704d89
health: surface cgroup memory and per-worker RSS in the sweep
malaiwah Aug 11, 2026
1bb2e6b
fq: handoff final for this session + BT-6c run artifacts
malaiwah Aug 11, 2026
2c58139
fq: final BT-6c artifacts at handover
malaiwah Aug 11, 2026
6ef0d79
docs: add CHANGELOG.md summarizing all features and fixes on the branch
malaiwah Aug 11, 2026
28e81a5
poc: additive residual encoding for EXL3 trellis K2→K3→K4 — measured
malaiwah Aug 12, 2026
078e87e
poc: accurate additive residual measurements with real EXL3 trellis o…
malaiwah Aug 12, 2026
05cf24e
poc: optimization opportunities — Lloyd-Max, cross-expert SVD, codebo…
malaiwah Aug 12, 2026
fe9208c
chore: add .poc-venv to .gitignore
malaiwah Aug 12, 2026
ef04c46
research: comprehensive PoC v4 — all 6 optimization ideas measured
malaiwah Aug 12, 2026
a4b827f
research: v6/v7 — continuously variable per-expert sparse quantizatio…
malaiwah Aug 12, 2026
1ad9ef8
research: v8/v9 — BitsMoE decomposition + K4 tile K5 + full 3-tier (G…
malaiwah Aug 12, 2026
8d6005e
research: v10 — AlphaQ PL_Alpha_Hill allocation (doesn't help for GLM…
malaiwah Aug 12, 2026
f547a02
research: comprehensive findings — tile-level 3-tier K3/K4/K5 is best…
malaiwah Aug 12, 2026
038d6fd
research: v11 — DP-optimal tile assignment + bitmap entropy analysis
malaiwah Aug 12, 2026
07d7091
research: v12 — proxy-based tier assignment (variance as fast proxy)
malaiwah Aug 12, 2026
e7d131f
research: round 2-4 literature review — RRQ, HyperQuant, ICQuant, Q-P…
malaiwah Aug 12, 2026
0ee87e6
research: final summary — tile-level 3-tier K3/K4/K5 is best fungible…
malaiwah Aug 12, 2026
17f731b
research: v13 — cross-layer analysis (layers 10 & 40 are identical)
malaiwah Aug 12, 2026
c28c6eb
research: round 5 literature — MXFP4, CodeQuant, MoBiQuant, GAMMA, WUSH
malaiwah Aug 12, 2026
0d1c8c6
research: v14 — tile difficulty is spatially random (no clustering)
malaiwah Aug 12, 2026
8e32e49
research: round 5b — Leech lattice (LLVQ), entropy-constrained quanti…
malaiwah Aug 12, 2026
2db2ca3
research: v15 — complete smooth Pareto frontier (3.0-5.0 bpw, 0.1-bit…
malaiwah Aug 12, 2026
ff04f3e
research: v16 — extended Pareto to 6.0 bpw with 4-tier K3/K4/K5/K6
malaiwah Aug 12, 2026
8b839ef
research: final Pareto frontier — continuously variable 3.0-6.0 bpw
malaiwah Aug 12, 2026
6ef1dda
research: complete report — 4-tier tile-level K3/K4/K5/K6, 3.0-6.0 bp…
malaiwah Aug 12, 2026
5550194
research: round 6 literature — GLVQ, MoBiQuant, expert merging, tenso…
malaiwah Aug 12, 2026
97e1596
research: v17 — bitmap compression analysis
malaiwah Aug 12, 2026
d6a1913
research: v18 — progressive bitstream + quantized benefit (0% loss at…
malaiwah Aug 12, 2026
b598d5e
research: implementation guide — tile-level 4-tier fungible quantization
malaiwah Aug 12, 2026
81d72f1
research: round 7 — embedded TCQ theory, FLUTE kernel, final status
malaiwah Aug 12, 2026
09b9db2
research: v19 — per-tile rotation diversity (no effect)
malaiwah Aug 12, 2026
75062b7
research: round 8 — joint pruning+quantization, stochastic rounding, …
malaiwah Aug 12, 2026
3ce574f
research: round 9 — NVIDIA cuDNN Grouped GEMM+Quant (SM100+) hardware…
malaiwah Aug 12, 2026
d83348a
research: v20 — K5 residual decomposition comparison
malaiwah Aug 12, 2026
efaa094
research: round 10 — TurboQuant theory + final conclusions (45+ paper…
malaiwah Aug 12, 2026
8361b46
research: v21-v22b — corrected bpw labels + large Lloyd-Max codebooks
malaiwah Aug 12, 2026
10dd3a4
research: v23b — per-tile Lloyd-Max codebooks (17-41% better than glo…
malaiwah Aug 12, 2026
59517de
research: v24 — per-tile LM Pareto dominates global (1.6-70% improvem…
malaiwah Aug 12, 2026
f350e09
research: v25 — codebook clustering sweet spot at 64 clusters
malaiwah Aug 12, 2026
07252e1
research: Fruit model analysis — per-expert heterogeneity is tier art…
malaiwah Aug 12, 2026
14f4d81
research: v26 — universal normalized codebooks (cross-layer sharing =…
malaiwah Aug 12, 2026
cf8aeca
research: definitive report — 26 PoCs, 50+ papers, universal codebook…
malaiwah Aug 12, 2026
9c2b890
research: v27b-v28 — K2 base tier + reshape bug fix + cross-layer re-…
malaiwah Aug 12, 2026
2d566f9
research: v29 — entropy coding saves 5-13% LM bits, PQ worse than scalar
malaiwah Aug 12, 2026
4bb7815
research: v30 — c128 sweet spot, stacking always worse than single LM
malaiwah Aug 12, 2026
4901128
research: round 11 + definitive report updated (v27-v30, 55+ papers, …
malaiwah Aug 12, 2026
e44cbcf
research: v31 — definitive Pareto frontier (c128, 10 experts, gate+down)
malaiwah Aug 12, 2026
463185a
research: v32 — entropy-aware Pareto 14-49% better, BPDQ 4-23% worse
malaiwah Aug 12, 2026
97e3d20
research: round 12 — BPDQ tested, final conclusions (60+ papers, 32 P…
malaiwah Aug 12, 2026
c5af14b
research: v33 — sparse/adaptive/tier-specific clusters all worse than…
malaiwah Aug 12, 2026
2cea6d9
research: v34 — trellis-on-residual 2-34× WORSE than Lloyd-Max
malaiwah Aug 12, 2026
8a94c55
research: round 13 — HARP, MSQ, QTIP, Proteus, NanoQuant, ReSpinQuant…
malaiwah Aug 12, 2026
8e574fb
research: v35 — BREAKTHROUGH: rescaled trellis-on-residual 34-37% bet…
malaiwah Aug 12, 2026
0e51d54
research: v36 — definitive hybrid Pareto, K2+K4trsc 47% better at 6bpw
malaiwah Aug 12, 2026
a533376
research: updated definitive report with v35-v36 rescaled trellis bre…
malaiwah Aug 12, 2026
f28410f
research: v37 — K2+K5trsc new best at 7bpw, K2+K2trsc matches K4 at 4bpw
malaiwah Aug 12, 2026
0ea1896
research: round 14 — successive refinement of TCQ, Drop-by-Drop, nest…
malaiwah Aug 12, 2026
ec51d04
research: v38 — entropy-aware hybrid Pareto (LM entropy saves 0.56 bpw)
malaiwah Aug 12, 2026
8d90c55
research: updated definitive report — v37-v38, K2+K5trsc, entropy-awa…
malaiwah Aug 12, 2026
0c1d284
research: v39 — THREE breakthroughs: 2-stage trellis 17% better, K6tr…
malaiwah Aug 12, 2026
4086aac
research: v40 — 3-stage rescaled trellis 2.3× better than LM at 8bpw!
malaiwah Aug 12, 2026
c6ab377
research: v41 — DEFINITIVE: MSRT beats LM at ALL bitrates 5-10 bpw
malaiwah Aug 12, 2026
4308d56
research: DEFINITIVE REPORT — MSRT beats LM at all bitrates (41 PoCs,…
malaiwah Aug 12, 2026
d5881eb
research: v42 — K2 base confirmed best, MSRT+LM hybrid 6× worse, univ…
malaiwah Aug 12, 2026
f5377b9
research: round 16 — RateQuant reverse waterfilling explains MSRT opt…
malaiwah Aug 12, 2026
26e6ed2
research: v43 — tile-level MSRT Pareto, 0-1.4% improvement from mixing
malaiwah Aug 12, 2026
70bbfec
research: v44 — MSRT 68-101× better than RRQ, cross-layer validated
malaiwah Aug 12, 2026
afc961a
research: round 17 — RRQ 68-101× worse than MSRT, GSQ/Proteus reviewed
malaiwah Aug 12, 2026
b5b36e4
research: v45 — trellis indices are structured (16-bit), no entropy b…
malaiwah Aug 12, 2026
787fb3a
research: round 18 — ECTCQ, embedded E-ECTCQ; trellis indices not ent…
malaiwah Aug 12, 2026
026a03a
research: v46 — systematic allocation search confirms K1-first patter…
malaiwah Aug 12, 2026
3c6bc21
research: round 19 — Drop-by-Drop, gradient-guided allocation; revers…
malaiwah Aug 12, 2026
1cb9bcc
research: v47 — dithering 29-36% worse, per-tile K1 negligible
malaiwah Aug 12, 2026
a51e9bd
research: round 20 — dithering hurts TCQ, noise shaping N/A, interlea…
malaiwah Aug 12, 2026
4be4694
research: v48 — mul1 2.6× worse, up_proj universal, 70 experts CV=0.11%
malaiwah Aug 12, 2026
9923743
research: round 21 — LLVQ, BCJR-QAT, learned lattices; mcg codebook c…
malaiwah Aug 12, 2026
ef2f865
research: v49 — cross-layer validation + expert reordering (all ±0.03%)
malaiwah Aug 12, 2026
c69a80a
research: MSRT additive cartridge feasibility + implementation plan
malaiwah Aug 12, 2026
d5594f1
research: GLM-5.2 MSRT cost analysis — memory, performance, base K + …
malaiwah Aug 12, 2026
e630cda
fix: correct cost analysis — MSRT advantage starts at 5bpw, not 3bpw
malaiwah Aug 12, 2026
0d7bab7
research: MSRT cartridge via progressive-tensors + vLLM LoRA hot-swap
malaiwah Aug 12, 2026
296b392
research: v51 — MSRT cartridge matches native K4 within 3.1%
malaiwah Aug 12, 2026
f17cd8e
research: v52 — dual-cartridge MSRT (K2 base + tiered K1/K2/K3 cartri…
malaiwah Aug 12, 2026
b1515dd
feat: EXL3 LoRA cartridge support for MSRT additive quantization
malaiwah Aug 12, 2026
d06c871
fix: shard_map must map projection names (gate_proj→w1, up_proj→w3, d…
malaiwah Aug 12, 2026
bb99f51
fix: extend EXL3 rank-sliced bitrate to accept K2
malaiwah Aug 12, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
11 changes: 11 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -255,5 +255,16 @@ vllm/grpc/vllm_engine_pb2.py
vllm/grpc/vllm_engine_pb2_grpc.py
vllm/grpc/vllm_engine_pb2.pyi

# DGX Spark local build products
/.spark-artifacts/
/.spark-build/

# Ignore generated cpu headers
csrc/cpu/cpu_attn_dispatch_generated.h

# Raw routing-stat dumps: tens of MB each, published to HF instead
# (see results/k3-fq/CONVERGENCE-RESULTS.md for sha256 + location)
research/fungible-quant/runs/m5-serve/results/*/stats*.jsonl

# PoC virtualenv (local only, not for commit)
.poc-venv/
553 changes: 553 additions & 0 deletions CHANGELOG.md

Large diffs are not rendered by default.

12 changes: 12 additions & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -433,6 +433,9 @@ if(VLLM_GPU_LANG STREQUAL "CUDA" OR VLLM_GPU_LANG STREQUAL "HIP")
endif()

if(VLLM_GPU_LANG STREQUAL "CUDA")
list(APPEND VLLM_STABLE_EXT_SRC
"csrc/libtorch_stable/attention/mla/safe_query_bmm.cu")

SET(CUTLASS_ENABLE_HEADERS_ONLY ON CACHE BOOL "Enable only the header library")

# Set CUTLASS_REVISION. Used for FetchContent. Also fixes some bogus messages when building.
Expand Down Expand Up @@ -1106,6 +1109,15 @@ if(VLLM_GPU_LANG STREQUAL "CUDA" OR VLLM_GPU_LANG STREQUAL "HIP")
# Needed to use cuda/hip APIs from C-shim
if(VLLM_GPU_LANG STREQUAL "CUDA")
target_compile_definitions(_C_stable_libtorch PRIVATE USE_CUDA)
find_library(VLLM_CUBLAS_LIBRARY cublas
PATHS "/usr/local/cuda/lib64" "/usr/local/cuda/lib" NO_DEFAULT_PATH)
if(NOT VLLM_CUBLAS_LIBRARY)
find_library(VLLM_CUBLAS_LIBRARY cublas)
endif()
if(NOT VLLM_CUBLAS_LIBRARY)
message(FATAL_ERROR "safe_mla_query_bmm requires libcublas")
endif()
target_link_libraries(_C_stable_libtorch PRIVATE ${VLLM_CUBLAS_LIBRARY})
if(COOPERATIVE_TOPK_ARCHS)
target_compile_definitions(_C_stable_libtorch PRIVATE
VLLM_ENABLE_COOPERATIVE_TOPK=1)
Expand Down
165 changes: 165 additions & 0 deletions benchmarks/profile_dspark_sps_curve.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,165 @@
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
"""Profile the engine step-rate curve for the DSpark prefix scheduler.

Times the captured FULL cudagraph replays of the target verification step at
every captured batch token count and emits ``dspark_sps_curve`` breakpoints,
one per capture size. The scheduler linearly interpolates between
breakpoints, which amortizes cudagraph padding smoothly instead of
concentrating it into thresholds at capture-size boundaries.

Example:
python benchmarks/profile_dspark_sps_curve.py <target-model> \\
--speculative-config '{"method": "dspark", "model": "...", ...}' \\
--engine-args '{"tensor_parallel_size": 4, "max_num_seqs": 32}'

Paste the printed ``dspark_sps_curve`` entry into --speculative-config.

Caveats: replays run on whatever (dummy) buffer state capture left behind, so
data-dependent kernels (e.g. MoE routing) may be timed on unrepresentative
inputs, and per-step CPU/draft overhead is modeled only through the constant
``--overhead-ms``. Only the curve's shape matters to the scheduler.
"""

import argparse
import json


def _time_fullgraph_replays(worker, iters: int, warmup: int) -> dict[int, float]:
"""Worker-side: time FULL graph replay per batch token count (ms/step).

Runs on every TP rank via collective_rpc so the collectives captured in
the graphs stay matched; every rank replays the same descs in the same
sorted order. Before timing each descriptor the input buffers are
refreshed into the same coherent dummy state capture used, so replays
never read stale metadata.
"""
import torch

from vllm.v1.worker.gpu.cudagraph_utils import prepare_inputs_to_capture

runner = worker.model_runner
mgr = runner.cudagraph_manager
assert mgr is not None and mgr.graphs, (
"No FULL cudagraphs captured; run with a cudagraph_mode that captures "
"FULL decode graphs."
)
# Prefer varlen spec-decode descs; fall back to all captured graphs.
descs = [d for d in mgr.graphs if d.max_req_tokens is not None]
if not descs:
descs = list(mgr.graphs.keys())
# One desc per token count: the largest request count is the most
# representative shape under load.
by_tokens: dict[int, object] = {}
for d in descs:
cur = by_tokens.get(d.num_tokens)
if cur is None or (d.num_reqs or 0) > (cur.num_reqs or 0):
by_tokens[d.num_tokens] = d

results: dict[int, float] = {}
for num_tokens in sorted(by_tokens):
desc = by_tokens[num_tokens]
num_reqs = desc.num_reqs or min(num_tokens, mgr.max_num_reqs)
prepare_inputs_to_capture(
num_reqs,
num_tokens,
runner.model_state,
runner.input_buffers,
runner.block_tables,
runner.attn_groups,
runner.kv_cache_config,
max_req_tokens=desc.max_req_tokens,
)
graph = mgr.graphs[desc]
for _ in range(warmup):
graph.replay()
torch.accelerator.synchronize()
start = torch.Event(enable_timing=True)
end = torch.Event(enable_timing=True)
start.record()
for _ in range(iters):
graph.replay()
end.record()
torch.accelerator.synchronize()
results[num_tokens] = start.elapsed_time(end) / iters
return results


def curve_breakpoints(
ms_per_step: dict[int, float], overhead_ms: float
) -> list[list[float]]:
"""Convert per-capture-size step times into ``dspark_sps_curve``
breakpoints, one per capture size. The scheduler's table linearly
interpolates between them (and clamps at the ends)."""
return [
[size, round(1000.0 / (ms_per_step[size] + overhead_ms), 3)]
for size in sorted(ms_per_step)
]


def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("model", help="Target model (path or HF id)")
parser.add_argument(
"--speculative-config",
required=True,
help="JSON speculative config (same value you pass to vllm serve)",
)
parser.add_argument(
"--engine-args",
default="{}",
help="JSON dict of extra vllm.LLM kwargs "
'(e.g. \'{"tensor_parallel_size": 4, "max_num_seqs": 32}\')',
)
parser.add_argument("--iters", type=int, default=50)
parser.add_argument("--warmup", type=int, default=5)
parser.add_argument(
"--overhead-ms",
type=float,
default=0.0,
help="Constant per-step overhead (draft forward, sampling, CPU gap) "
"added to every measured step time before converting to a rate.",
)
parser.add_argument("--output", help="Write the curve JSON to this file")
args = parser.parse_args()
if args.iters <= 0:
parser.error("--iters must be greater than zero")
if args.warmup < 0:
parser.error("--warmup must be non-negative")
if args.overhead_ms < 0:
parser.error("--overhead-ms must be non-negative")

# The timing callable is shipped to the workers via collective_rpc, which
# requires the pickle fallback. Local profiling tool, trusted input.
import os

os.environ.setdefault("VLLM_ALLOW_INSECURE_SERIALIZATION", "1")

from vllm import LLM

llm = LLM(
model=args.model,
speculative_config=json.loads(args.speculative_config),
**json.loads(args.engine_args),
)
per_rank = llm.collective_rpc(
_time_fullgraph_replays, kwargs={"iters": args.iters, "warmup": args.warmup}
)
ms_per_step = per_rank[0]

print("\nMeasured FULL-graph step times (rank 0):")
for size in sorted(ms_per_step):
print(f" B={size:5d} tokens: {ms_per_step[size]:8.3f} ms/step")

curve = curve_breakpoints(ms_per_step, args.overhead_ms)
entry = {"dspark_sps_curve": curve}
print("\nAdd to --speculative-config:")
print(json.dumps(entry))
if args.output:
with open(args.output, "w") as f:
json.dump(entry, f, indent=2)
print(f"\nWritten to {args.output}")


if __name__ == "__main__":
main()
5 changes: 5 additions & 0 deletions cmake/external_projects/vllm_flash_attn.cmake
Original file line number Diff line number Diff line change
Expand Up @@ -36,10 +36,15 @@ if(VLLM_FLASH_ATTN_SRC_DIR)
BINARY_DIR ${CMAKE_BINARY_DIR}/vllm-flash-attn
)
else()
set(VLLM_FLASH_ATTN_GIT_SUBMODULES "")
if(VLLM_GPU_LANG STREQUAL "CUDA")
set(VLLM_FLASH_ATTN_GIT_SUBMODULES "csrc/cutlass")
endif()
FetchContent_Declare(
vllm-flash-attn
GIT_REPOSITORY https://github.com/vllm-project/flash-attention.git
GIT_TAG caaa4eb59845388a20b1f435ecaafb4bd9517ad8
GIT_SUBMODULES ${VLLM_FLASH_ATTN_GIT_SUBMODULES}
GIT_PROGRESS TRUE
# Don't share the vllm-flash-attn build between build types
BINARY_DIR ${CMAKE_BINARY_DIR}/vllm-flash-attn
Expand Down
106 changes: 106 additions & 0 deletions csrc/libtorch_stable/attention/mla/safe_query_bmm.cu
Original file line number Diff line number Diff line change
@@ -0,0 +1,106 @@
#include <torch/csrc/stable/library.h>
#include <torch/csrc/stable/tensor.h>
#include <torch/headeronly/core/ScalarType.h>

#include "core/registration.h"
#include "libtorch_stable/torch_utils.h"

#include <cublas_v2.h>
#include <cuda_bf16.h>

#include <limits>
#include <sstream>

namespace {

void check_cublas(cublasStatus_t status, const char* operation) {
if (status == CUBLAS_STATUS_SUCCESS) {
return;
}
std::ostringstream error;
error << operation << " failed with cuBLAS status "
<< static_cast<int>(status);
STD_TORCH_CHECK(false, error.str());
}

int checked_int(int64_t value, const char* name) {
STD_TORCH_CHECK(value > 0 && value <= std::numeric_limits<int>::max(), name,
" is outside cuBLAS int range: ", value);
return static_cast<int>(value);
}

void check_bf16_cuda_3d(torch::stable::Tensor const& tensor,
const char* name) {
STD_TORCH_CHECK(tensor.device().is_cuda(), name, " must be on CUDA");
STD_TORCH_CHECK(tensor.dim() == 3, name, " must be a 3D tensor");
STD_TORCH_CHECK(
tensor.scalar_type() == torch::headeronly::ScalarType::BFloat16, name,
" must be BF16");
}

} // namespace

void safe_mla_query_bmm(torch::stable::Tensor const& query,
torch::stable::Tensor const& weight,
torch::stable::Tensor& output) {
check_bf16_cuda_3d(query, "query");
check_bf16_cuda_3d(weight, "weight");
check_bf16_cuda_3d(output, "output");
STD_TORCH_CHECK(query.get_device_index() == weight.get_device_index() &&
query.get_device_index() == output.get_device_index(),
"query, weight, and output must be on the same CUDA device");

const int64_t heads_i64 = query.size(0);
const int64_t tokens_i64 = query.size(1);
const int64_t q_dim_i64 = query.size(2);
const int64_t latent_dim_i64 = weight.size(2);

STD_TORCH_CHECK(weight.size(0) == heads_i64 && weight.size(1) == q_dim_i64,
"weight must have shape [heads, q_dim, latent_dim]");
STD_TORCH_CHECK(output.size(0) == heads_i64 &&
output.size(1) == tokens_i64 &&
output.size(2) == latent_dim_i64,
"output must have shape [heads, tokens, latent_dim]");
STD_TORCH_CHECK(query.stride(2) == 1,
"query q_dim must be contiguous for safe_mla_query_bmm");
STD_TORCH_CHECK(weight.stride(2) == 1,
"weight latent_dim must be contiguous for safe_mla_query_bmm");
STD_TORCH_CHECK(output.stride(2) == 1,
"output latent_dim must be contiguous for safe_mla_query_bmm");

const int heads = checked_int(heads_i64, "heads");
const int tokens = checked_int(tokens_i64, "tokens");
const int q_dim = checked_int(q_dim_i64, "q_dim");
const int latent_dim = checked_int(latent_dim_i64, "latent_dim");
const int query_ld = checked_int(query.stride(1), "query.stride(1)");
const int weight_ld = checked_int(weight.stride(1), "weight.stride(1)");
const int output_ld = checked_int(output.stride(1), "output.stride(1)");

const torch::stable::accelerator::DeviceGuard device_guard(
query.get_device_index());
cublasHandle_t handle = get_current_cuda_blas_handle();
check_cublas(cublasSetStream(handle, get_current_cuda_stream(
query.get_device_index())),
"cublasSetStream");

const float alpha = 1.0f;
const float beta = 0.0f;
check_cublas(
cublasGemmStridedBatchedEx(
handle, CUBLAS_OP_N, CUBLAS_OP_N, latent_dim, tokens, q_dim, &alpha,
weight.const_data_ptr(), CUDA_R_16BF, weight_ld,
static_cast<long long>(weight.stride(0)), query.const_data_ptr(),
CUDA_R_16BF, query_ld, static_cast<long long>(query.stride(0)),
&beta, output.mutable_data_ptr(), CUDA_R_16BF, output_ld,
static_cast<long long>(output.stride(0)), heads,
// The explicit operand order and leading dimensions provide the
// tight-query contract. Regular FP32 accumulation keeps tensor-core
// kernels eligible; PEDANTIC forces a much slower fallback for
// production prefill shapes without improving that contract.
CUBLAS_COMPUTE_32F, CUBLAS_GEMM_DEFAULT),
"cublasGemmStridedBatchedEx");
}

STABLE_TORCH_LIBRARY_IMPL(_C, CUDA, m) {
m.impl("safe_mla_query_bmm", TORCH_BOX(&safe_mla_query_bmm));
}
Loading