Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,9 @@
- **Flash-Next default flips (QSA PM, Halo hyper units, symmetric IU4 on gfx1151, long-prefill expert staging)**, against `fe77c0837`: on Strix Halo the shipped asymmetric `.mq4` goes from 1,447.6 to 1,544.9 tok/s at pp8192 (+6.7 %), and the 16K logits are byte-identical. The symmetric GPTQ3 requant takes the IU4 route by default and reaches pp8192 1,928.5 tok/s, at a BF16-source KLD of 0.1027 (WikiText-2, 32 × 512 tokens). On the F16 route the same artifact's KLD is 0.066. gfx1201 decode (R9700, `auto` placement, tg64) stays at 34.6 / 32.4 tok/s at 512 / 8K context. All figures are medians of 3 fresh processes.

### Flash-Next performance
- **Qwen3.8-Flash-Next (Qwen4): engine-owned session cache keeps many prompts warm (`memory.session_cache_bytes`, `HIPFIRE_SESSION_CACHE_BYTES`, default 8 GiB, `0` disables).** It replaces the single whole-chunk prefix checkpoint, which kept only one prompt: the next request with another prefix, such as another session or a subagent, overwrote it. `hipfire_runtime::session_cache` now owns keying, planning, LRU eviction, the memory guard and placement. Each prefill snapshots the full state at every prefill-chunk boundary it crosses: GDN, PLE and hyper state plus the QSA K/V, raw and pooled rows below the boundary, and with MTP the head's state and draft policy. Snapshots are self-contained, so a reset no longer drops them. A snapshot is published only after the client commits the turn. Qwen4 no longer reads `HIPFIRE_QWEN_PROMPT_CACHE`.
- Strix Halo, canonical GPTQ3 artifact, serve, greedy, MTP: prompt A (21,550 tokens), then B (8,142 tokens), then A again. A's prefill-and-decode time went from 16.3 s cold to 4.9 s, with `cached_tokens` 16,384 and the same content. With `HIPFIRE_SESSION_CACHE_BYTES=0`, A reports `cached_tokens` 0 and the same content. One 8,192-token snapshot is 513 MiB.
- Identity: `session_cache_hw` (AR, F32 QSA) restores a two-chunk snapshot after an unrelated prompt: the final logits are byte-equal to the cold prefill's and 16 greedy ids match.
- **Qwen3.8-Flash-Next (Qwen4): a large PLE row request reads its SSD pages over up to 256 threads instead of 16 (`HIPFIRE_QWEN4_PLE_WIDE_READERS`, default on; `=0` keeps 16). Byte-identical; Ciru 8K cold prefill +5.6 %.** A cold prefill's first chunk fetches about 111K scattered 1,280-byte PLE pages at the layer-1 consumption point. No earlier chunk can warm them, so the GPU idles for the whole read (the 1.4 s gap at the first PLE gather in the Halo trace). Timestamps show the gap is the row-store read: about 40 ms of row ids, plan and cache scan, then the SSD reads, ~130 ms of copy and cache publish, and ~13 ms of staging and upload. Load, pinning and the lookahead are not involved. A request with more than 1,024 page reads now uses one reader per 64 reads, at most 256. Smaller requests (decode, MTP verify, short prompts) keep 16. Bytes, order and the row-store cache are unchanged.
- Strix Halo, canonical GPTQ3 artifact, `hipfire bench` protocol, 8K chunk 0: host wait 1,555 → 1,178 ms. The SSD reads go 1,428 → 1,022 ms, the device's ceiling: ~100K IOPS at queue depth 128–256 against ~37–50K at 16. Publish stays at ~130–145 ms. rocprofv3, load followed directly by the prompt: the chunk-0 gap goes 1,254 → 1,105 ms at 8K and 1,451 → 1,077 ms at 32K, and chunks 1+ stay at 6–13 ms.
- The read rate depends on the drive's state. After ≥ 30 s idle the same reads run at ~18K IOPS for any thread count, and chunk 0 then waits 2–7.5 s on both arms. That cost is device-bound and unchanged here. In that state the wider readers still finish the next chunk's lookahead warm within the current chunk: base stalled 3.0–3.2 s at chunk 1 in 2 of 5 32K runs, the candidate in 0 of 4. Load time is unchanged (both arms 47–71 s, the same bimodal spread).
Expand Down
6 changes: 3 additions & 3 deletions crates/hipfire-arch-qwen35/map.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,7 @@ _Generated by `scripts/check-crate-maps.py` from the tree — do not edit inside
| [`src/arch.rs`](src/arch.rs) | 103 | 1 | 0 |
| [`src/arch_model.rs`](src/arch_model.rs) | 101 | 0 | 0 |
| [`src/carrier.rs`](src/carrier.rs) | 854 | 4 | 7 |
| [`src/checkpoint.rs`](src/checkpoint.rs) | 108 | 3 | 0 |
| [`src/checkpoint.rs`](src/checkpoint.rs) | 106 | 3 | 0 |
| [`src/dflash_slot.rs`](src/dflash_slot.rs) | 783 | 12 | 0 |
| [`src/dflash_spec.rs`](src/dflash_spec.rs) | 2,048 | 19 | 16 |
| [`src/dflash_verify_pm4.rs`](src/dflash_verify_pm4.rs) | 744 | 36 | 9 |
Expand All @@ -50,7 +50,7 @@ _Generated by `scripts/check-crate-maps.py` from the tree — do not edit inside
| [`src/qwen35/prefill.rs`](src/qwen35/prefill.rs) | 17,159 | 18 | 70 |
| [`src/qwen35/weights.rs`](src/qwen35/weights.rs) | 2,991 | 43 | 11 |
| [`src/qwen35.rs`](src/qwen35.rs) | 69 | 8 | 0 |
| [`src/serve_engine.rs`](src/serve_engine.rs) | 7,434 | 12 | 33 |
| [`src/serve_engine.rs`](src/serve_engine.rs) | 7,431 | 12 | 33 |
| [`src/spec_emit.rs`](src/spec_emit.rs) | 1,042 | 5 | 14 |
| [`src/spec_impl.rs`](src/spec_impl.rs) | 643 | 1 | 0 |
| [`src/speculative.rs`](src/speculative.rs) | 9,932 | 87 | 14 |
Expand Down Expand Up @@ -102,6 +102,6 @@ _Generated by `scripts/check-crate-maps.py` from the tree — do not edit inside

### Totals

- 31 modules · 87,250 lines · 555 public items · 280 tests · 18 examples
- 31 modules · 87,245 lines · 555 public items · 280 tests · 18 examples

<!-- crate-map:generated:end -->
14 changes: 6 additions & 8 deletions crates/hipfire-arch-qwen35/src/checkpoint.rs
Original file line number Diff line number Diff line change
Expand Up @@ -6,13 +6,13 @@
//! GPU capture/restore over a live `DeltaNetState`, and re-exports so
//! existing `crate::checkpoint::…` paths keep working.

use crate::qwen35::DeltaNetState;
use crate::speculative::DeltaNetSnapshot;
use hip_bridge::{HipError, HipResult};
pub use hipfire_runtime::checkpoint_pool::{
plan_resume, prefix_fingerprint, CheckpointBlob, CheckpointId, QwenCheckpointPool,
plan_resume, prefix_fingerprint, CheckpointBlob, CheckpointId, CheckpointPool,
};
use hipfire_runtime::serve_contract::CacheDomain;
use crate::qwen35::DeltaNetState;
use hip_bridge::{HipError, HipResult};
use rdna_compute::Gpu;

impl CheckpointBlob for DeltaNetSnapshot {
Expand All @@ -38,13 +38,13 @@ impl CheckpointBlob for DeltaNetSnapshot {
/// path at page-aligned completed boundaries.
pub fn capture_checkpoint(
gpu: &mut Gpu,
pool: &mut QwenCheckpointPool<DeltaNetSnapshot>,
pool: &mut CheckpointPool<DeltaNetSnapshot>,
domain: &CacheDomain,
p: u64,
boundary_tokens: &[u32],
state: &DeltaNetState,
) -> HipResult<CheckpointId> {
if !QwenCheckpointPool::<DeltaNetSnapshot>::is_aligned(p) {
if !CheckpointPool::<DeltaNetSnapshot>::is_aligned(p) {
return Err(HipError::new(
0,
"capture_checkpoint: boundary not page-aligned",
Expand Down Expand Up @@ -89,7 +89,7 @@ pub fn capture_checkpoint(
/// private recurrent state for a running request.
pub fn restore_private(
gpu: &mut Gpu,
pool: &mut QwenCheckpointPool<DeltaNetSnapshot>,
pool: &mut CheckpointPool<DeltaNetSnapshot>,
domain: &CacheDomain,
p: u64,
fp: u64,
Expand All @@ -104,5 +104,3 @@ pub fn restore_private(
// ───────────────────────────────────────────────────────────────────────────
// Tests (host-only — no GPU/HIP required)
// ───────────────────────────────────────────────────────────────────────────


Loading
Loading