Skip to content

fix(llama_spec): generic DFlash daemon panic on 1–3 row verify blocks - #819

Open
fivetide wants to merge 1 commit into
warpfront:betafrom
fivetide:fix/generic-dflash-small-block-capture
Open

fivetide wants to merge 1 commit into
warpfront:betafrom
fivetide:fix/generic-dflash-small-block-capture

Conversation

@fivetide

@fivetide fivetide commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

Summary

Generic (llama-family) DFlash crashed the daemon whenever a request's remaining budget left a verify block of 1–3 rows. This PR fixes it by making the llama per-token verify fallback capture the drafter hidden rows, the same way per-token prefill does.

 verify_block_logits_or_argmax(block, capture)
   eligible = n >= 4 && dtypes/KV batchable
   if eligible
     forward_prefill_batch_capture(.., capture)       # n rows captured
   else
-    capture = None                                   # 0 rows captured
-    forward_scratch_compute(..)
+    forward_scratch_compute_capture(.., host capture) # 1 row per token
GenericDflashSpeculator::step / step_one_token
  b = min(block_size, max_emit)            # 1..3 near the end of a request
  target.verify_block(block, Some(&mut block_hidden))
  finish_chain
    block_hidden[..(accepted+1)*ne*h]      # panicked: block_hidden was empty

forward_scratch_compute is forward_scratch_compute_capture(.., None), so callers that pass no capture sink run the same code as before. The GPU-resident sink (hidden_gpu) is only ever passed for blocks that take the batched path, and the per-token path filters it out.

Which surface(s) does this touch?

  • kernel
  • load
  • serve — runtime spec (crates/hipfire-runtime/src/llama_spec.rs), generic DFlash verify
  • arch crate(s)
  • crates/hipfire-quantize / quant formats
  • control plane
  • docs / CI / scripts only
  • policy files

Evidence

Setup:

  • gfx1151, HIP 7.2. Beta 060cadcd3 with only this commit applied.
  • Model: Qwen3-8B MQ4 (qwen3-8b.mq4, md5 8af0eed5b8d2287bb6a33e6c9d03d1f9).
  • Draft: qwen3-8b-dflash.hfq (md5 f4be14d4a6acfabbf4a01f3c79f63dd4), set via developer.dflash_draft, speculation.dflash = "on".
  • Command: hipfire bench <model> --spec dflash --runs 1 --warmups 0 --max-tokens 64 --backend noslots --workload stateless --prompt-file <p>
daemon prompt result
before (beta, md5 d7f9cc3a70ba…) online_tune/prose_letter.txt panicked at crates/hipfire-runtime/src/dflash_generic.rs:473:45: range end index 40960 out of range for slice of length 0
before online_tune/mixed_tcp.txt same panic (range end index 20480 …)
after (md5 44f382d92d58…) prose_letter completes: 64 tok, 45 windows
after mixed_tcp completes: 64 tok, 44 windows

Code prompts with high acceptance (for example hw-gate/load-code.txt) can skip past the small-block tail and finish on both builds. The crash needs a window that starts with 1–3 tokens of budget left.

The prompt files are in the companion draft PR (benchmarks/prompts/online_tune/), and the prose prompt is reproduced verbatim here:

prose_letter.txt (md5 0ba3a438a580995f07a453cf2a495e43)
Write a long letter, about 1200 words, from a retired sea captain to his granddaughter who is about to leave home for the first time. He recalls his own first voyage, the people he met in distant ports, the mistakes he made, and the advice he wishes someone had given him. Write it as a warm, rambling personal letter in continuous paragraphs.

Test plan

  • cargo build --release clean
  • cargo test --release -p hipfire-runtime --lib: 973 passed, 0 failed
  • ./scripts/no-gpu-ci.sh: not run locally; relying on CI
  • serve_harness battery: attempted (--model qwen3-8b.mq4 --draft qwen3-8b-dflash.hfq --dflash on --thinking off --max-tokens 64 --mode battery), but its serve-path proof fails on this pair even on the fixed build. It reports dflash=on requested but no request-level DFlash execution evidence (tau=None, and every turn is think-only). Serve does not appear to run generic DFlash for this non-registry pair. That is a separate issue; the evidence above uses the native daemon bench path, which does run DFlash (drafter=dflash per request).
  • KV backend: legacy, automatic (device gfx1151 has no certified VMM KV path for the selected mode); not set explicitly.
  • Not perf-relevant: per-token blocks only occur at the end of a request.

Merge Danger

Door: two-way. One function and no format or state changes.

Blast Radius: narrow. It affects verify_block_logits_or_argmax callers (llama-family verify_block / verify_block_logits) only for blocks below the batched minimum or with ineligible dtypes/KV. Those blocks now also download one residual row per extract layer per token, which is the same work per-token prefill does.

Architecture-trait change?

No.

Found while working on the online draft-tuning draft PR (filed separately; it builds on this commit).

Blocks below the batched-verify minimum (n < 4) took the per-token fallback,
which captured no hidden rows; GenericDflashSpeculator::finish_chain then
sliced the empty buffer and the daemon panicked (dflash_generic.rs:473)
whenever a request's remaining budget made the block 1-3 rows.

forward_scratch_compute is forward_scratch_compute_capture(.., None), so
non-capturing callers are unchanged. The GPU-resident sink is only passed
for batched-eligible blocks and is filtered out of the per-token path.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant