Repository navigation
Conversation
Blocks below the batched-verify minimum (n < 4) took the per-token fallback, which captured no hidden rows; GenericDflashSpeculator::finish_chain then sliced the empty buffer and the daemon panicked (dflash_generic.rs:473) whenever a request's remaining budget made the block 1-3 rows. forward_scratch_compute is forward_scratch_compute_capture(.., None), so non-capturing callers are unchanged. The GPU-resident sink is only passed for batched-eligible blocks and is filtered out of the per-token path.
9 of 11 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Generic (llama-family) DFlash crashed the daemon whenever a request's remaining budget left a verify block of 1–3 rows. This PR fixes it by making the llama per-token verify fallback capture the drafter hidden rows, the same way per-token prefill does.
verify_block_logits_or_argmax(block, capture) eligible = n >= 4 && dtypes/KV batchable if eligible forward_prefill_batch_capture(.., capture) # n rows captured else - capture = None # 0 rows captured - forward_scratch_compute(..) + forward_scratch_compute_capture(.., host capture) # 1 row per tokenforward_scratch_computeisforward_scratch_compute_capture(.., None), so callers that pass no capture sink run the same code as before. The GPU-resident sink (hidden_gpu) is only ever passed for blocks that take the batched path, and the per-token path filters it out.Which surface(s) does this touch?
crates/hipfire-runtime/src/llama_spec.rs), generic DFlash verifycrates/hipfire-quantize/ quant formatsEvidence
Setup:
060cadcd3with only this commit applied.qwen3-8b.mq4, md58af0eed5b8d2287bb6a33e6c9d03d1f9).qwen3-8b-dflash.hfq(md5f4be14d4a6acfabbf4a01f3c79f63dd4), set viadeveloper.dflash_draft,speculation.dflash = "on".hipfire bench <model> --spec dflash --runs 1 --warmups 0 --max-tokens 64 --backend noslots --workload stateless --prompt-file <p>d7f9cc3a70ba…)online_tune/prose_letter.txtpanicked at crates/hipfire-runtime/src/dflash_generic.rs:473:45: range end index 40960 out of range for slice of length 0online_tune/mixed_tcp.txtrange end index 20480 …)44f382d92d58…)prose_lettermixed_tcpCode prompts with high acceptance (for example
hw-gate/load-code.txt) can skip past the small-block tail and finish on both builds. The crash needs a window that starts with 1–3 tokens of budget left.The prompt files are in the companion draft PR (
benchmarks/prompts/online_tune/), and the prose prompt is reproduced verbatim here:prose_letter.txt (md5 0ba3a438a580995f07a453cf2a495e43)
Test plan
cargo build --releasecleancargo test --release -p hipfire-runtime --lib: 973 passed, 0 failed./scripts/no-gpu-ci.sh: not run locally; relying on CI--model qwen3-8b.mq4 --draft qwen3-8b-dflash.hfq --dflash on --thinking off --max-tokens 64 --mode battery), but its serve-path proof fails on this pair even on the fixed build. It reportsdflash=on requested but no request-level DFlash execution evidence(tau=None, and every turn is think-only). Serve does not appear to run generic DFlash for this non-registry pair. That is a separate issue; the evidence above uses the native daemonbenchpath, which does run DFlash (drafter=dflashper request).device gfx1151 has no certified VMM KV path for the selected mode); not set explicitly.Merge Danger
Door: two-way. One function and no format or state changes.
Blast Radius: narrow. It affects
verify_block_logits_or_argmaxcallers (llama-familyverify_block/verify_block_logits) only for blocks below the batched minimum or with ineligible dtypes/KV. Those blocks now also download one residual row per extract layer per token, which is the same work per-token prefill does.Architecture-trait change?
No.
Found while working on the online draft-tuning draft PR (filed separately; it builds on this commit).