Repository navigation
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Extend the exact MQ4V2 K5120 gate/up specialization to gfx1201 as an independent, default-off opt-in. This is a small exact-arithmetic specialization, not a new quantization format: keep the generic DEV/RT weight-load policy, block32/wave32 geometry, arithmetic order, and resident weight layout; specialize the group count to 20 only for gate/up=17408, K=5120.
Rebased onto beta
787827258. In particular, this does not change the newly default-on gfx1100 route from #816. Enable the new route withhipfire config set kernel.gfx12_mq4v2_gateup_k5120 trueorHIPFIRE_GFX12_MQ4V2_GATEUP_K5120=1, then restart the daemon. Other architectures/shapes remain unchanged. Registry/compiler-pack and Redline kernarg/pointer-effect contracts include the new symbol.Performance boundary
Radeon AI PRO R9700 / gfx1201, ROCm 7.14.60850, Qwen3.8-27B MQ4-XT, AR/no speculation, Q8 requested, VMM observed. Artifact SHA256:
80e7c624424fd1d363ba86681d3dc1e5ac5534e0e064306a32be204c4843d0f3.The initial TG4096 ABBA gave 40.05 -> 40.20 tok/s median (+0.37%); the final rebased ABBA gave 40.10 -> 40.20 (+0.25%), all four outputs again identical. The separate gate/up diagnostic gave +0.56% with a 1.515GB weight ring and +0.69% with repeated single-layer weights. These are modest signals, not a statistically established serving speedup: each serving comparison has only two samples per arm and the daemon reports rates rounded to 0.1 tok/s. A CPU-only CI run overlapped the early rebased serving test; that limitation is disclosed in the report. No prefill speedup or broad gfx12 claim.
Full logits (993,280 FP32 values) were byte-identical; all four 4096-token serving outputs were identical. Both arms passed four-position HIP/PM4 replay and GDN-state parity, with 64 expected gate/up symbols in each capture. Resource audit: VGPR 76 -> 76, SGPR 22 -> 19, scratch 0 -> 0. No extra resident weights.
Raw results and reusable commands: evidence README, including final serving battery rows. This includes the final rebased validation, not just the initial experiment. The report explains the remote source-overlay provenance explicitly.
Which surface(s) does this touch?
rdna-computedispatch, source specialization, registry and replay contractshipfire-configprocess setting only; no weight-loader changesTest plan
./scripts/no-gpu-ci.shpasses (final serialized run; compressed full log attached)hipfire-cli,hipfire-daemon, and the standalone logits probecargo test --lib --workspacepasses (not claimed)serve_harness.py --mode batteryvia the checked-in long-AR wrapper, both arms, attached raw rowskv_backend_legacy=false; no legacy override. Q8 explicitly requested, butkv_modewas not observed by the harness../scripts/speed-gate.shwithin +/-2% of locked baselines (not run; no global speed-gate claim)Targeted K5120 tests: 6 passed. Config tests: 107 passed. Initial full serial
rdna-computeGPU tests: 438 passed / 2 failed / 13 ignored; both failing GDN Q8 tests also reproduce on unmodified beta5d172b663(baseline log attached), and no GDN implementation changes are included. This does not claim the full GPU suite is green.Architecture-trait change?
None. No
Architecturetrait or external inference ABI change. The new internal HIP symbol has the same 64-byte kernarg and pointer effects as the existing generic gate/up kernel.