Skip to content

perf(gfx1100): speed up long-context DFlash verify 13.4% with FA2 split-KV and retain HipGraph/Redline PM4 - #817

Open
HUSRCF wants to merge 2 commits into
warpfront:betafrom
HUSRCF:perf/gfx1100-fa2-splitkv-retained
Open

HUSRCF wants to merge 2 commits into
warpfront:betafrom
HUSRCF:perf/gfx1100-fa2-splitkv-retained

Conversation

@HUSRCF

@HUSRCF HUSRCF commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Make the exact-gfx1100 FA2 split-KV S8 backend the default long-context Q8 DFlash verifier inside its measured envelope. The pinned product A/B improved from 42.6 to 48.3 tok/s (+13.38%) at unchanged tau and cycle count. The same fixed-grid route is retained through HipGraph and Redline/PM4 with live-context admission, compiler-free kernels, stable scratch, typed replay pointer effects, and fail-closed fallback.

This is the gfx1100 companion to the split-KV schedule evaluated in #760. It compares against—but does not claim authorship of—the existing R4/R8 multi-row verifier from #741. The implementation keeps hipfire's in-tree packed-Q8 KV path; it does not vendor or load CK source or a CK .so.

The branch is rebased onto current beta@212347998. The performance evidence below was collected on the same split kernel/route before the rebase at beta@d5305333d; the review follow-up changes admission, packaging, and fallback policy rather than kernel math.

Default and admission policy

The route now defaults on only on exact gfx1100. Either of these restores the established verifier path:

  • kernel.gfx1100_fa2_split_verify=false / HIPFIRE_GFX1100_FA2_SPLIT_VERIFY=0
  • parent switch kernel.verify_attn=false / HIPFIRE_VERIFY_ATTN=0

Admission is exact gfx1100, plain Q8 KV, dense Qwen H24/KV4/D256, sequential non-tree batches 4..32, and live logical context >4096. HIPFIRE_FA_PERTOKEN_MIN_CTX=0 disables the route; positive overrides below 4096 are clamped to the measured crossover. Preallocated max_seq does not influence admission.

HipGraph/retained capture additionally requires the complete precompiled split module, fixed-address Q16 scratch, and enough flash_partials capacity for the S8 layout. The real PM4 loader checks the target's actual workspace at fixed B=16. Unsupported shapes or not-ready resources fail closed to the established batched path; they cannot be recorded under a split identity.

Maintainer-review follow-up

  • Honored the stable VerifyAttn parent opt-out and added an exact-gfx1100 route-specific stable config/escape hatch.
  • Enforced dense/Q8/H24-KV4-D256/B4..32/live-context admission at both model and launcher boundaries.
  • Split ChainVerifySplit from ordinary ChainVerify, so capture-only F16/LDS helpers are not widened below the crossover or on fallback routes.
  • Kept retained recording fail-closed while preserving existing gfx1151/gfx1201 HipGraph behavior.
  • Added preconvert, partial, merge, and direct symbols to one compiler-free gfx1100 module.
  • Included the caller-owned S8 partial workspace in eager, HipGraph, Redline shadow, and product PM4 admission.

Which surface(s) does this touch?

  • kernel — kernels/, crates/rdna-compute
  • load — carries target scratch capacity into retained-route admission
  • serve — crates/hipfire-generate, speculative verify and Redline plumbing
  • arch crate(s): hipfire-arch-qwen35
  • crates/hipfire-quantize / quant formats
  • control plane — stable config key and environment alias
  • docs / CI / scripts only (no hardware route)
  • policy files

Performance

Optimization A/B: established R4/R8 vs FA2 split-KV

W7900, exact gfx1100, HIP 7.15, Q8 VMM KV, graph off, six fresh processes in order off,on,on,off,off,on, one unrecorded warmup per process:

verifier attention route samples (tok/s) median tau / cycles
established R4/R8 43.5, 42.6, 42.5 42.6 1.97 / 67
FA2 split-KV S8 48.5, 48.3, 47.8 48.3 (+13.38%) 1.97 / 67

The unchanged tau/cycles isolate the improvement to verifier execution rather than draft acceptance.

Batch-16 kernel screen (10 warmups, 30 measurements):

logical context R4/R8 split S8 speedup
8,192 368.80 us 177.36 us 2.079x
20,676 828.45 us 410.72 us 2.017x
32,768 1,369.01 us 648.21 us 2.112x

S1 is bit-identical to direct FA2. S8 relative L2 is about 2.46e-4 to 2.58e-4 with cosine about 0.999999970. Direct/partial/merge compile at 254/255/18 VGPR with zero spill and zero private scratch.

HipGraph capture neutrality: split route off vs on

Both arms below already use the new split-KV route; this is not the optimization A/B above. It checks only whether retaining that route in HipGraph changes output or speed.

split-KV execution decode result
graph off 4.8 tok/s tau 2.06, 200 generated tokens
graph on 4.9 tok/s tau 2.06, 200 generated tokens

The B16 graph captured 706 launch blobs. Transcripts were byte-identical, MD5 7b8dc5b28daef60f803fe2a466c888b2, with identical request MD5 8a54e9aa236f678362f89864bf000125. Both cold serve_harness.py runs ended inside the model's hidden thought channel and were flagged RUNAWAY,EMPTY, so these rates are not a performance or answer-quality claim. This run establishes only graph-off/on route parity; the 42.6→48.3 fresh-process native bench above is the performance A/B.

Redline/PM4

The four-arm daemon oracle covered HipAuto, capture-safe direct HIP, recorded HIP, and PM4 at 12 consecutive B16 windows (positions 8176..8352):

  • PASS; all arms exactly agreed on tokens, argmax, hidden staging/ring, final hidden, logits, active KV hashes, GDN intermediates, and recurrent state after forward and rollback.
  • 711 launches/dispatches, 18 unique typed AQL kernel contracts, one packet, one queue, and one phase; dispatch count matched launch count.
  • 17 successful PM4 replays; zero contract, prepare, or replay failures.
  • Five interleaved smoke windows: HipGraph 46.413 ms median, PM4 43.630 ms median (-6.00%, p95 delta -2.652 ms).

The five-window PM4 number is a retained-route smoke measurement, not a standalone headline throughput claim.

Full immutable record: docs/perf-checkpoints/2026-10-05-gfx1100-fa2-splitkv-verifier-graph-pm4.md.

Test plan

  • ./scripts/no-gpu-ci.sh — previous run reached unchanged beta Python scripts/hw-gate/select.py shadowing of stdlib select; this branch does not touch the failing files
  • cargo build --release clean for the measured route commit
  • Claim-scoped Rust tests on the measured route commit: rdna-compute 438/438, hipfire-arch-qwen35 247/247, hipfire-generate 54/54
  • W7900 test_kernels: 17/17
  • Native serve_harness.py product route and route-specific Redline four-arm harness on W7900
  • Exact KV fields inspected: kv_mode=q8, kv_backend=vmm, kv_backend_legacy=false, no legacy fallback reason
  • Rebased follow-up: all-target compile checks for rdna-compute, hipfire-arch-qwen35, hipfire-generate, and hipfire-loader
  • Targeted default/opt-out, threshold, compiler-free registry, capture fallback, and undersized-workspace tests
  • Crate maps, env/lifecycle inventory, changed-file rustfmt, and diff checks
  • ./scripts/speed-gate.sh — this route is not a locked speed-gate case; the claim uses the fresh-process A/B above
  • No cleanup-threshold ratchet is changed

No gfx12 hardware rerun is included: this admission is exact gfx1100, and gfx1201 keeps its separate existing split-KV implementation. A static negative admission test protects the cross-architecture boundary.

Current GitHub Actions disclosure for head 4f07ac0bd: rustfmt and cargo-deny licenses/bans/sources pass. Workspace build, workspace unit tests, and advisory clippy stop in unchanged crates/railgun/src/npu/mod.rs:24, where Rust 1.99 rejects the libc open runtime-symbol signature. The ratchet job separately reports daemon_lines=5392 above its 5377 ceiling; this branch changes neither crates/hipfire-daemon/src/main.rs nor scripts/leanup-ratchets.sh. Cargo-deny advisories now also reports the unchanged locked private-gemm-x86 v0.1.20 as yanked. None of these failures originate in the files changed by this PR.

Warpfront-beta validation used CLI MD5 9bbfbbaac68ed262867a6e7136485081 and daemon MD5 f0c79a77e2cb1b019ee58bbae96de113. Target SHA-256 was 9f91556f7e0431a077d03756a7102d0154108757289e6e5fe9a2d204c0c9eeb7; draft SHA-256 was d0a74a232a0e2166d889f823e91e0fbf778d21dd9668d7de055cdecb065401bc; committed prompt MD5 was b4d0b63cddcac872648ddf3cdd92cac2.

Architecture-trait change?

No Architecture trait change.

@HUSRCF
HUSRCF requested a review from Kaden-Schutt as a code owner October 5, 2026 11:51
@fivetide

fivetide commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

why opt in? sounds like no reason not to be default..

@HUSRCF
HUSRCF marked this pull request as draft October 5, 2026 15:24
@HUSRCF
HUSRCF force-pushed the perf/gfx1100-fa2-splitkv-retained branch from 07b8995 to 3ec998b Compare October 5, 2026 16:13
@HUSRCF HUSRCF changed the title perf(gfx1100): speed up long-context DFlash verify 12.4% with FA2 split-KV and retain HipGraph/Redline PM4 perf(gfx1100): speed up long-context DFlash verify 12.7% with FA2 split-KV and retain HipGraph/Redline PM4 Oct 5, 2026
@HUSRCF
HUSRCF changed the base branch from master to beta October 5, 2026 16:14
@HUSRCF
HUSRCF force-pushed the perf/gfx1100-fa2-splitkv-retained branch from 3ec998b to 30e7a8e Compare October 5, 2026 16:34
@HUSRCF HUSRCF changed the title perf(gfx1100): speed up long-context DFlash verify 12.7% with FA2 split-KV and retain HipGraph/Redline PM4 perf(gfx1100): speed up long-context DFlash verify 13.4% with FA2 split-KV and retain HipGraph/Redline PM4 Oct 5, 2026
@HUSRCF
HUSRCF force-pushed the perf/gfx1100-fa2-splitkv-retained branch from 30e7a8e to db3331c Compare October 5, 2026 16:40
@HUSRCF

HUSRCF commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

@fivetide Thanks — I agree that default-on is probably the right final policy within the admitted envelope. I initially kept it opt-in because the validation scope was narrow, but a follow-up audit found a few things I would like to fix before flipping the default:

  • HIPFIRE_VERIFY_ATTN=0 does not currently suppress the new split route, conflicting with the documented master opt-out.
  • HipGraph admission is broader than the launcher’s exact H24/KV4/HD256 envelope, so a not-ready or unsupported capture can fall into an insufficiently validated multi-row path instead of a guaranteed batched fallback.
  • The new partial/merge symbols are not yet present in the release kernel registry, leaving compiler-free cold installs uncovered.
  • The same flag also enables capture-time F16 projection/LDS-stage paths below the 4K attention threshold, which is broader than its documented meaning.
  • The PR says dense-only, but the eager/HipGraph admission does not currently enforce num_experts == 0.

Also, the measured kernel matrix starts at 8K, so I would like to pin the 4K crossover with a small boundary A/B before making it the product default.

I’ll rebase onto current beta now that the gfx1201 split-KV work has landed, tighten these guards, add the compiler-free and opt-out tests, and then make the exact admitted gfx1100 route default-on with HIPFIRE_GFX1100_FA2_SPLIT_VERIFY=0 retained as the escape hatch.

HUSRCF added 2 commits October 6, 2026 16:30
Replace the long-context Q8 DFlash R4/R8 attention step with an opt-in FA2 split-KV S8 backend on exact gfx1100. Keep the live-context crossover fail-closed and retain the same fixed-grid route through HipGraph and Redline/PM4 with typed pointer effects, stable scratch, and product-identical lm-head handling.

Current warpfront/beta@d5305333d W7900 fresh-process E2E: 42.6 to 48.3 tok/s (+13.38%) at unchanged tau=1.97 and 67 cycles. HipGraph off/on produced byte-identical transcripts. Redline four-arm parity passed 12/12 windows; PM4 was 6.00% faster than HipGraph in the five-window verify-body smoke.

Validation: release product build, test_kernels 17/17, rdna-compute 438/438, hipfire-arch-qwen35 247/247, hipfire-generate 54/54, crate maps, env/lifecycle inventory, rustfmt, and diff checks. Local no-gpu-ci reached the unchanged beta Python select.py shadowing failure after its Rust/main checks; GitHub beta-runner baseline failures are disclosed in the PR.
Make the exact dense Q8 H24/KV4/HD256 long-context route default-on while retaining the parent VerifyAttn and route-specific opt-outs. Align eager, HipGraph, and Redline admission; package every split symbol in the compiler-free registry; and fail closed when kernels, fixed Q16 scratch, or the caller's S8 partial workspace are not ready.\n\nThe PM4 loader now includes the target flash-partials capacity in its fixed-B16 admission, so a batched fallback cannot be mislabeled as a retained split tape.
@HUSRCF
HUSRCF force-pushed the perf/gfx1100-fa2-splitkv-retained branch from db3331c to 4f07ac0 Compare October 6, 2026 09:03
@HUSRCF

HUSRCF commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

@fivetide Thanks — I have now made the route default-on inside the exact measured gfx1100 envelope and rebased the PR onto current beta@212347998.

The follow-up is commit 4f07ac0bd. It now:

  • honors both kernel.verify_attn=false / HIPFIRE_VERIFY_ATTN=0 and the route-specific kernel.gfx1100_fa2_split_verify=false / HIPFIRE_GFX1100_FA2_SPLIT_VERIFY=0 escape hatch;
  • requires exact gfx1100, plain Q8 KV, dense H24/KV4/D256, sequential non-tree B4..32, and live context above the resolved 4K crossover;
  • separates ChainVerifySplit from ordinary ChainVerify, so the capture-only F16/LDS helpers are not enabled below the crossover or on fallback routes;
  • packages preconvert/partial/merge/direct in the compiler-free gfx1100 module;
  • keeps capture fail-closed without changing the existing gfx1151/gfx1201 HipGraph paths; and
  • checks the actual S8 flash_partials capacity in eager, HipGraph, Redline shadow, and product PM4 B16 admission, so a small workspace cannot record a batched fallback under a split identity.

The rebased head passes all-target compile checks for rdna-compute, hipfire-arch-qwen35, hipfire-generate, and hipfire-loader, plus the targeted default/opt-out, threshold, registry, capture fallback, and undersized-workspace tests; crate maps, env/lifecycle docs, changed-file rustfmt, and diff checks also pass. I did not rerun gfx12 hardware because this admission is exact gfx1100 and gfx1201 keeps its separate existing split-KV backend; a static negative admission test covers that boundary.

I also updated the PR body to distinguish the prior pinned performance evidence from the current-beta admission/packaging follow-up.

@HUSRCF
HUSRCF marked this pull request as ready for review October 6, 2026 11:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants