Repository navigation
feat(dflash): developer-only online draft tuning — tokens/window ×1.094, decode ×1.105 on gfx1151 (research draft, builds on #819) - #820
Draft
fivetide wants to merge 2 commits into
Conversation
added 2 commits
October 6, 2026 08:04
Blocks below the batched-verify minimum (n < 4) took the per-token fallback, which captured no hidden rows; GenericDflashSpeculator::finish_chain then sliced the empty buffer and the daemon panicked (dflash_generic.rs:473) whenever a request's remaining budget made the block 1-3 rows. forward_scratch_compute is forward_scratch_compute_capture(.., None), so non-capturing callers are unchanged. The GPU-resident sink is only passed for batched-eligible blocks and is filtered out of the per-token path.
…NE_TUNE)
Per-request learner re-ranks the DFlash draft's top-16 per block row from
tokens the target already verified (chain n-grams, suffix-match injection,
stutter, neighbour-row tokens/probabilities). Verify is unchanged. Wired into
Qwen3.5 chain DFlash (greedy batched LM-head path; top-16 via the re-grid
kernel, byte-identical on gfx1151) and the generic llama-family chain (greedy
and temp>0 naive-sampling verify). Off by default.
gfx1151, 20 cold requests vs shipping DFlash: tokens/window x1.094, decode
x1.105 (Qwen3.5-9B x1.044/x1.072, Qwen3-8B weak draft x1.146/x1.139).
Tooling: stats/sweep dump modes, examples/dflash_online_replay.rs
(--simulate own-start evaluator, --carry-loo), scripts/dflash_online_tune/.
Evidence: docs/perf-checkpoints/2026-10-0{5,6}-gfx1151-dflash-online-*.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Draft, research parking PR. Not for merge as-is. This is developer-only online DFlash draft tuning (
HIPFIRE_DFLASH_ONLINE_TUNE, off by default).A small per-request learner re-ranks the draft's top-16 candidates per block row, training on tokens the target has already verified. Verify is unchanged, so every emitted token is the target's.
This is not Path C. It trains no draft weights offline: it is a host-side linear re-ranker learned during the session, about 0.1 ms of host work per cycle.
Builds on #819 (that commit is included here; review that PR first).
Learner features, per candidate:
The model is linear,
score = scale·(l_j − l_0) + w·f_j, trained with Adagrad (lr 0.3, acc0 1) on softmax cross-entropy over the candidates. It trains only on rows a greedy chain reaches, uses one weight set for all row depths, and carries its weights across requests. Text statistics are reset per request.Results (gfx1151, HIP 7.2)
Online, 20 cold single requests (10 prompts × 2 pairs). Tuned is compared against cached shipping DFlash (tuner unset). Greedy decoding, so tokens/window is deterministic per build.
qwen35-9b-dflash-mq4(strong draft)qwen3-8b-dflash.hfq(weak, generic path)hipfire run -n 512): tokens/window ×1.104, faster on 10 of 10.Fixtures:
qwen3.5-9b.mq4, md5296092bf1e6a45d78c1acf815eb93366; draft md5590f35403cd7f1d634945233234a12b7.qwen3-8b.mq4, md58af0eed5b8d2287bb6a33e6c9d03d1f9; draft md5f4be14d4a6acfabbf4a01f3c79f63dd4.benchmarks/prompts/online_tune/*.txt(md5s in the perf checkpoint).Smoke on this branch rebased onto beta
060cadcd3(daemon md5ccb5e7a1ab05…): Qwen3.5-9Bcode_edit_typehints, shipping 796 tok / 80 windows (τ 8.94), tuned 796 tok / 78 windows (τ 9.21). The simulator reproduces ×1.0833 on this build.Full record (append-only):
docs/perf-checkpoints/2026-10-05-gfx1151-dflash-online-draft-tuning.mddocs/perf-checkpoints/2026-10-06-gfx1151-dflash-online-draft-tuning-amendment.md: second pair, replay bias, rejected changesdocs/perf-checkpoints/2026-10-06-gfx1151-dflash-online-draft-tuning-simulator.md: own-start simulator, final numbersHow to evaluate (read this before changing the learner)
Do not use fixed-start replay of argmax sessions to select changes. It scores a policy at the block starts of a plain-argmax session. A policy that accepts more starts its later blocks on harder positions, which replay cannot see, so it overstated changes by up to about 5%. On one prompt, replay scored τ 1.89 while online gave 1.43.
Use the own-start simulator:
HIPFIRE_DFLASH_ONLINE_TUNE=sweepdrafts a token the target never picks. Every cycle then commits one target token, and the dump (HIPFIRE_DFLASH_ONLINE_DUMP=<file>) holds the draft's top-16 at every position of the greedy text.examples/dflash_online_replay.rs --simulate DUMP...walks any policy over those dumps with its own block starts (accept the longest matching prefix, start after the bonus), against argmax on the same text.--carry-looestimates a warm daemon: each session runs after the others in one tuner.HIPFIRE_HOME(symlinks to~/.hipfire) whoseconfig.tomlsets[developer] dflash_draft = ".../qwen3-8b-dflash.hfq"andspeculation.dflash = "on", because the pair is not registry-managed.fiction_lighthouse,prose_river_short,mixed_code_then_prose,merge_sort_thinking_off,humaneval_0_has_close_elements,lru_cache_pep8_strict,coherence_lloyd_long,bare_factual). The final policy scores ×1.0186 there (×1.0102 for the start-of-segment policy). With--carry-looit scores q9 ×1.025 and q8 ×1.039.Knobs
HIPFIRE_DFLASH_ONLINE_TUNEonre-ranks;statskeeps argmax and prints acceptance-ceiling stats;sweepproduces evaluation dumps (always rejected; slow)HIPFIRE_DFLASH_ONLINE_HPlr=,acc0=,margin=,inject=,carry=0|1HIPFIRE_DFLASH_ONLINE_DUMPCode map
topk_values_batched_f32_verified_regrid: the shipping one-block-per-row top-K costs 1349 µs per cycle at 15 × 248320, K=16, which ate the gain on gfx1151. The re-grid kernel takes 63 µs and passes the existingtest_select_regridbyte-identity gate on gfx1151 (run with its arch check bypassed; 210 launches, ties/NaN/−inf included). The shipping selector gate (select_regrid_enabled, gfx1201 only) is untouched; only the tuner calls the new entry point.Tried and rejected
Each was measured with the methods above; details are in the checkpoints.
ORDERS,RECENT,MAX_OCC,MAXM) within ±0.2 pp.Open items / how to pick this up
--carry-loo). The online harness is one request per fresh daemon.印) and stops on these prompts with or without speculation. This is a separate issue.hipfire bench(native daemon protocol, noslots).serve_harnessevidence once item 5 is resolved.Research history: fork branch
fivetide:autoresearch/find-ways-to-improve-performance-in-gpu-npu-inte-20261005(head9a5e5e58a), with every experiment as a commit; it also carries the NPU interop work from #818. Dumps and caches on the dev box:~/.cache/hipfire-online/{sweep,sweep_ho,shipping,hf8home}. Regenerate them with the scripts; they are not portable.Which surface(s) does this touch?
crates/rdna-compute/src/select_regrid.rs(new entry point only; existing kernels unchanged)hipfire-arch-qwen35(hooks dormant unless the env var is set)scripts/dflash_online_tune/Test plan
cargo build --releaseclean (on beta060cadcd3+ fix(llama_spec): generic DFlash daemon panic on 1–3 row verify blocks #819)cargo test --release -p hipfire-runtime --lib: 974 passed, includingtraining_features_match_proposal_features, a no-peek invariant mutation-checked against two injected leaksscripts/check-lifecycle.py,scripts/check-env-docs.pypassHIPFIRE_DFLASH_ONLINE_TUNEunset leavesDflashScratch.online = None, so every hook is skipped; stats mode reproduces shipping greedy output byte for byteserve_harnessbattery: blocked on open item 5speed-gate.sh: n/a (off by default)Merge Danger
Door: two-way. Off by default, behind a developer env var, and adds no persistent state or format.
Blast Radius: none at default. When enabled: proposal choice and host time in DFlash chain steps, and one extra top-K launch per cycle on Qwen3.5.