Skip to content

perf(gfx1201): add opt-in exact K5120 MQ4V2 gate/up decode - #824

Open
HUSRCF wants to merge 3 commits into
warpfront:betafrom
HUSRCF:experiment/gfx1201-k5120-20261006
Open

HUSRCF wants to merge 3 commits into
warpfront:betafrom
HUSRCF:experiment/gfx1201-k5120-20261006

Conversation

@HUSRCF

@HUSRCF HUSRCF commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Summary

Extend the exact MQ4V2 K5120 gate/up specialization to gfx1201 as an independent, default-off opt-in. This is a small exact-arithmetic specialization, not a new quantization format: keep the generic DEV/RT weight-load policy, block32/wave32 geometry, arithmetic order, and resident weight layout; specialize the group count to 20 only for gate/up=17408, K=5120.

Rebased onto beta 787827258. In particular, this does not change the newly default-on gfx1100 route from #816. Enable the new route with hipfire config set kernel.gfx12_mq4v2_gateup_k5120 true or HIPFIRE_GFX12_MQ4V2_GATEUP_K5120=1, then restart the daemon. Other architectures/shapes remain unchanged. Registry/compiler-pack and Redline kernarg/pointer-effect contracts include the new symbol.

Performance boundary

Radeon AI PRO R9700 / gfx1201, ROCm 7.14.60850, Qwen3.8-27B MQ4-XT, AR/no speculation, Q8 requested, VMM observed. Artifact SHA256: 80e7c624424fd1d363ba86681d3dc1e5ac5534e0e064306a32be204c4843d0f3.

The initial TG4096 ABBA gave 40.05 -> 40.20 tok/s median (+0.37%); the final rebased ABBA gave 40.10 -> 40.20 (+0.25%), all four outputs again identical. The separate gate/up diagnostic gave +0.56% with a 1.515GB weight ring and +0.69% with repeated single-layer weights. These are modest signals, not a statistically established serving speedup: each serving comparison has only two samples per arm and the daemon reports rates rounded to 0.1 tok/s. A CPU-only CI run overlapped the early rebased serving test; that limitation is disclosed in the report. No prefill speedup or broad gfx12 claim.

Full logits (993,280 FP32 values) were byte-identical; all four 4096-token serving outputs were identical. Both arms passed four-position HIP/PM4 replay and GDN-state parity, with 64 expected gate/up symbols in each capture. Resource audit: VGPR 76 -> 76, SGPR 22 -> 19, scratch 0 -> 0. No extra resident weights.

Raw results and reusable commands: evidence README, including final serving battery rows. This includes the final rebased validation, not just the initial experiment. The report explains the remote source-overlay provenance explicitly.

Which surface(s) does this touch?

  • kernel: rdna-compute dispatch, source specialization, registry and replay contracts
  • load: typed hipfire-config process setting only; no weight-loader changes
  • serve
  • arch crate(s)
  • quantization formats
  • control plane
  • docs / CI / scripts only (this does change a hardware route)
  • policy files

Test plan

  • ./scripts/no-gpu-ci.sh passes (final serialized run; compressed full log attached)
  • Release build: hipfire-cli, hipfire-daemon, and the standalone logits probe
  • cargo test --lib --workspace passes (not claimed)
  • Official serve_harness.py --mode battery via the checked-in long-AR wrapper, both arms, attached raw rows
  • KV fields inspected: VMM observed, kv_backend_legacy=false; no legacy override. Q8 explicitly requested, but kv_mode was not observed by the harness.
  • ./scripts/speed-gate.sh within +/-2% of locked baselines (not run; no global speed-gate claim)

Targeted K5120 tests: 6 passed. Config tests: 107 passed. Initial full serial rdna-compute GPU tests: 438 passed / 2 failed / 13 ignored; both failing GDN Q8 tests also reproduce on unmodified beta 5d172b663 (baseline log attached), and no GDN implementation changes are included. This does not claim the full GPU suite is green.

Architecture-trait change?

None. No Architecture trait or external inference ABI change. The new internal HIP symbol has the same 64-byte kernarg and pointer effects as the existing generic gate/up kernel.

@HUSRCF
HUSRCF requested a review from Kaden-Schutt as a code owner October 6, 2026 13:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant