Repository navigation
Conversation
|
why opt in? sounds like no reason not to be default.. |
07b8995 to
3ec998b
Compare
3ec998b to
30e7a8e
Compare
30e7a8e to
db3331c
Compare
|
@fivetide Thanks — I agree that default-on is probably the right final policy within the admitted envelope. I initially kept it opt-in because the validation scope was narrow, but a follow-up audit found a few things I would like to fix before flipping the default:
Also, the measured kernel matrix starts at 8K, so I would like to pin the 4K crossover with a small boundary A/B before making it the product default. I’ll rebase onto current beta now that the gfx1201 split-KV work has landed, tighten these guards, add the compiler-free and opt-out tests, and then make the exact admitted gfx1100 route default-on with |
Replace the long-context Q8 DFlash R4/R8 attention step with an opt-in FA2 split-KV S8 backend on exact gfx1100. Keep the live-context crossover fail-closed and retain the same fixed-grid route through HipGraph and Redline/PM4 with typed pointer effects, stable scratch, and product-identical lm-head handling. Current warpfront/beta@d5305333d W7900 fresh-process E2E: 42.6 to 48.3 tok/s (+13.38%) at unchanged tau=1.97 and 67 cycles. HipGraph off/on produced byte-identical transcripts. Redline four-arm parity passed 12/12 windows; PM4 was 6.00% faster than HipGraph in the five-window verify-body smoke. Validation: release product build, test_kernels 17/17, rdna-compute 438/438, hipfire-arch-qwen35 247/247, hipfire-generate 54/54, crate maps, env/lifecycle inventory, rustfmt, and diff checks. Local no-gpu-ci reached the unchanged beta Python select.py shadowing failure after its Rust/main checks; GitHub beta-runner baseline failures are disclosed in the PR.
Make the exact dense Q8 H24/KV4/HD256 long-context route default-on while retaining the parent VerifyAttn and route-specific opt-outs. Align eager, HipGraph, and Redline admission; package every split symbol in the compiler-free registry; and fail closed when kernels, fixed Q16 scratch, or the caller's S8 partial workspace are not ready.\n\nThe PM4 loader now includes the target flash-partials capacity in its fixed-B16 admission, so a batched fallback cannot be mislabeled as a retained split tape.
db3331c to
4f07ac0
Compare
|
@fivetide Thanks — I have now made the route default-on inside the exact measured gfx1100 envelope and rebased the PR onto current The follow-up is commit
The rebased head passes all-target compile checks for I also updated the PR body to distinguish the prior pinned performance evidence from the current-beta admission/packaging follow-up. |
Summary
Make the exact-gfx1100 FA2 split-KV S8 backend the default long-context Q8 DFlash verifier inside its measured envelope. The pinned product A/B improved from 42.6 to 48.3 tok/s (+13.38%) at unchanged tau and cycle count. The same fixed-grid route is retained through HipGraph and Redline/PM4 with live-context admission, compiler-free kernels, stable scratch, typed replay pointer effects, and fail-closed fallback.
This is the gfx1100 companion to the split-KV schedule evaluated in #760. It compares against—but does not claim authorship of—the existing R4/R8 multi-row verifier from #741. The implementation keeps hipfire's in-tree packed-Q8 KV path; it does not vendor or load CK source or a CK
.so.The branch is rebased onto current
beta@212347998. The performance evidence below was collected on the same split kernel/route before the rebase atbeta@d5305333d; the review follow-up changes admission, packaging, and fallback policy rather than kernel math.Default and admission policy
The route now defaults on only on exact gfx1100. Either of these restores the established verifier path:
kernel.gfx1100_fa2_split_verify=false/HIPFIRE_GFX1100_FA2_SPLIT_VERIFY=0kernel.verify_attn=false/HIPFIRE_VERIFY_ATTN=0Admission is exact gfx1100, plain Q8 KV, dense Qwen H24/KV4/D256, sequential non-tree batches 4..32, and live logical context >4096.
HIPFIRE_FA_PERTOKEN_MIN_CTX=0disables the route; positive overrides below 4096 are clamped to the measured crossover. Preallocatedmax_seqdoes not influence admission.HipGraph/retained capture additionally requires the complete precompiled split module, fixed-address Q16 scratch, and enough
flash_partialscapacity for the S8 layout. The real PM4 loader checks the target's actual workspace at fixed B=16. Unsupported shapes or not-ready resources fail closed to the established batched path; they cannot be recorded under a split identity.Maintainer-review follow-up
VerifyAttnparent opt-out and added an exact-gfx1100 route-specific stable config/escape hatch.ChainVerifySplitfrom ordinaryChainVerify, so capture-only F16/LDS helpers are not widened below the crossover or on fallback routes.Which surface(s) does this touch?
kernels/,crates/rdna-computecrates/hipfire-generate, speculative verify and Redline plumbinghipfire-arch-qwen35crates/hipfire-quantize/ quant formatsPerformance
Optimization A/B: established R4/R8 vs FA2 split-KV
W7900, exact gfx1100, HIP 7.15, Q8 VMM KV, graph off, six fresh processes in order
off,on,on,off,off,on, one unrecorded warmup per process:The unchanged tau/cycles isolate the improvement to verifier execution rather than draft acceptance.
Batch-16 kernel screen (10 warmups, 30 measurements):
S1 is bit-identical to direct FA2. S8 relative L2 is about 2.46e-4 to 2.58e-4 with cosine about 0.999999970. Direct/partial/merge compile at 254/255/18 VGPR with zero spill and zero private scratch.
HipGraph capture neutrality: split route off vs on
Both arms below already use the new split-KV route; this is not the optimization A/B above. It checks only whether retaining that route in HipGraph changes output or speed.
The B16 graph captured 706 launch blobs. Transcripts were byte-identical, MD5
7b8dc5b28daef60f803fe2a466c888b2, with identical request MD58a54e9aa236f678362f89864bf000125. Both coldserve_harness.pyruns ended inside the model's hidden thought channel and were flaggedRUNAWAY,EMPTY, so these rates are not a performance or answer-quality claim. This run establishes only graph-off/on route parity; the 42.6→48.3 fresh-process native bench above is the performance A/B.Redline/PM4
The four-arm daemon oracle covered HipAuto, capture-safe direct HIP, recorded HIP, and PM4 at 12 consecutive B16 windows (positions 8176..8352):
The five-window PM4 number is a retained-route smoke measurement, not a standalone headline throughput claim.
Full immutable record:
docs/perf-checkpoints/2026-10-05-gfx1100-fa2-splitkv-verifier-graph-pm4.md.Test plan
./scripts/no-gpu-ci.sh— previous run reached unchanged beta Pythonscripts/hw-gate/select.pyshadowing of stdlibselect; this branch does not touch the failing filescargo build --releaseclean for the measured route commitrdna-compute438/438,hipfire-arch-qwen35247/247,hipfire-generate54/54test_kernels: 17/17serve_harness.pyproduct route and route-specific Redline four-arm harness on W7900kv_mode=q8,kv_backend=vmm,kv_backend_legacy=false, no legacy fallback reasonrdna-compute,hipfire-arch-qwen35,hipfire-generate, andhipfire-loader./scripts/speed-gate.sh— this route is not a locked speed-gate case; the claim uses the fresh-process A/B aboveNo gfx12 hardware rerun is included: this admission is exact gfx1100, and gfx1201 keeps its separate existing split-KV implementation. A static negative admission test protects the cross-architecture boundary.
Current GitHub Actions disclosure for head
4f07ac0bd: rustfmt and cargo-deny licenses/bans/sources pass. Workspace build, workspace unit tests, and advisory clippy stop in unchangedcrates/railgun/src/npu/mod.rs:24, where Rust 1.99 rejects the libcopenruntime-symbol signature. The ratchet job separately reportsdaemon_lines=5392above its 5377 ceiling; this branch changes neithercrates/hipfire-daemon/src/main.rsnorscripts/leanup-ratchets.sh. Cargo-deny advisories now also reports the unchanged lockedprivate-gemm-x86 v0.1.20as yanked. None of these failures originate in the files changed by this PR.Warpfront-beta validation used CLI MD5
9bbfbbaac68ed262867a6e7136485081and daemon MD5f0c79a77e2cb1b019ee58bbae96de113. Target SHA-256 was9f91556f7e0431a077d03756a7102d0154108757289e6e5fe9a2d204c0c9eeb7; draft SHA-256 wasd0a74a232a0e2166d889f823e91e0fbf778d21dd9668d7de055cdecb065401bc; committed prompt MD5 wasb4d0b63cddcac872648ddf3cdd92cac2.Architecture-trait change?
No
Architecturetrait change.