Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,9 @@
- **Flash-Next default flips (QSA PM, Halo hyper units, symmetric IU4 on gfx1151, long-prefill expert staging)**, against `fe77c0837`: on Strix Halo the shipped asymmetric `.mq4` goes from 1,447.6 to 1,544.9 tok/s at pp8192 (+6.7 %), and the 16K logits are byte-identical. The symmetric GPTQ3 requant takes the IU4 route by default and reaches pp8192 1,928.5 tok/s, at a BF16-source KLD of 0.1027 (WikiText-2, 32 × 512 tokens). On the F16 route the same artifact's KLD is 0.066. gfx1201 decode (R9700, `auto` placement, tg64) stays at 34.6 / 32.4 tok/s at 512 / 8K context. All figures are medians of 3 fresh processes.

### Flash-Next performance
- **Qwen3.8-Flash-Next (Qwen4): engine-owned session cache keeps many prompts warm (`memory.session_cache_bytes`, `HIPFIRE_SESSION_CACHE_BYTES`, default 8 GiB, `0` disables).** It replaces the single whole-chunk prefix checkpoint, which kept only one prompt: the next request with another prefix, such as another session or a subagent, overwrote it. `hipfire_runtime::session_cache` now owns keying, planning, LRU eviction, the memory guard and placement. Each prefill snapshots the state at every prefill-chunk boundary it crosses: the GDN, PLE and hyper state, the selections and, with MTP, the head's state and draft policy whole. Of the append-only QSA K/V, raw and pooled rows (and the head's) a snapshot holds only the rows above its parent, the deepest snapshot of the same prefix. A snapshot therefore costs one chunk of rows at any depth, and sessions that share a prefix (subagents with one system prompt) share its links. Restores walk the chain from the root. Snapshots never depend on live state, so a reset no longer drops them, and a snapshot with children is pinned, so eviction removes only leaves. A snapshot is published only after the client commits the turn. Qwen4 no longer reads `HIPFIRE_QWEN_PROMPT_CACHE`.
- Strix Halo, canonical GPTQ3 artifact, serve, greedy, MTP: prompt A (21,550 tokens), then B (8,142 tokens), then A again. A's prefill-and-decode time went from 16.3 s cold to 4.9 s, with `cached_tokens` 16,384 and the same content. With `HIPFIRE_SESSION_CACHE_BYTES=0`, A reports `cached_tokens` 0 and the same content. Each snapshot is 513 MiB at any depth; full copies would grow by 481 MiB per chunk of depth (a 64K session: about 4 GiB for all eight boundaries instead of about 17 GiB).
- Identity: `session_cache_hw` (F32 QSA, the serve default on gfx1151, Q8 GDN, VMM) prefills A (3 chunks + 300) and then B (A's first chunk + 8,392 other tokens) cold, then restores A through its three-link chain and B through the shared root, on the AR and on the native MTP route. Every state part (metadata, fixed parts, valid rows of every row stream, the MTP head and its draft policy included) is byte-equal to the cold prefill's, the AR final logits row is byte-equal with 16 equal greedy ids, the MTP seed is equal, and the cache holds exactly four one-chunk snapshots per route.
- **Qwen3.8-Flash-Next (Qwen4): a large PLE row request reads its SSD pages over up to 256 threads instead of 16 (`HIPFIRE_QWEN4_PLE_WIDE_READERS`, default on; `=0` keeps 16). Byte-identical; Ciru 8K cold prefill +5.6 %.** A cold prefill's first chunk fetches about 111K scattered 1,280-byte PLE pages at the layer-1 consumption point. No earlier chunk can warm them, so the GPU idles for the whole read (the 1.4 s gap at the first PLE gather in the Halo trace). Timestamps show the gap is the row-store read: about 40 ms of row ids, plan and cache scan, then the SSD reads, ~130 ms of copy and cache publish, and ~13 ms of staging and upload. Load, pinning and the lookahead are not involved. A request with more than 1,024 page reads now uses one reader per 64 reads, at most 256. Smaller requests (decode, MTP verify, short prompts) keep 16. Bytes, order and the row-store cache are unchanged.
- Strix Halo, canonical GPTQ3 artifact, `hipfire bench` protocol, 8K chunk 0: host wait 1,555 → 1,178 ms. The SSD reads go 1,428 → 1,022 ms, the device's ceiling: ~100K IOPS at queue depth 128–256 against ~37–50K at 16. Publish stays at ~130–145 ms. rocprofv3, load followed directly by the prompt: the chunk-0 gap goes 1,254 → 1,105 ms at 8K and 1,451 → 1,077 ms at 32K, and chunks 1+ stay at 6–13 ms.
- The read rate depends on the drive's state. After ≥ 30 s idle the same reads run at ~18K IOPS for any thread count, and chunk 0 then waits 2–7.5 s on both arms. That cost is device-bound and unchanged here. In that state the wider readers still finish the next chunk's lookahead warm within the current chunk: base stalled 3.0–3.2 s at chunk 1 in 2 of 5 32K runs, the candidate in 0 of 4. Load time is unchanged (both arms 47–71 s, the same bimodal spread).
Expand Down
8 changes: 4 additions & 4 deletions crates/hipfire-arch-qwen35/src/checkpoint.rs
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@

use crate::speculative::DeltaNetSnapshot;
pub use hipfire_runtime::checkpoint_pool::{
plan_resume, prefix_fingerprint, CheckpointBlob, CheckpointId, QwenCheckpointPool,
plan_resume, prefix_fingerprint, CheckpointBlob, CheckpointId, CheckpointPool,
};
use hipfire_runtime::serve_contract::CacheDomain;
use crate::qwen35::DeltaNetState;
Expand Down Expand Up @@ -38,13 +38,13 @@ impl CheckpointBlob for DeltaNetSnapshot {
/// path at page-aligned completed boundaries.
pub fn capture_checkpoint(
gpu: &mut Gpu,
pool: &mut QwenCheckpointPool<DeltaNetSnapshot>,
pool: &mut CheckpointPool<DeltaNetSnapshot>,
domain: &CacheDomain,
p: u64,
boundary_tokens: &[u32],
state: &DeltaNetState,
) -> HipResult<CheckpointId> {
if !QwenCheckpointPool::<DeltaNetSnapshot>::is_aligned(p) {
if !CheckpointPool::<DeltaNetSnapshot>::is_aligned(p) {
return Err(HipError::new(
0,
"capture_checkpoint: boundary not page-aligned",
Expand Down Expand Up @@ -89,7 +89,7 @@ pub fn capture_checkpoint(
/// private recurrent state for a running request.
pub fn restore_private(
gpu: &mut Gpu,
pool: &mut QwenCheckpointPool<DeltaNetSnapshot>,
pool: &mut CheckpointPool<DeltaNetSnapshot>,
domain: &CacheDomain,
p: u64,
fp: u64,
Expand Down
8 changes: 4 additions & 4 deletions crates/hipfire-arch-qwen35/src/serve_engine.rs
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@ use rdna_compute::sampling::SlotSampleParams;
use rdna_compute::slot_pool::{SlotId, SlotPool};
use rdna_compute::{DType, Gpu, GpuTensor};

use crate::checkpoint::{capture_checkpoint, plan_resume, CheckpointId, QwenCheckpointPool};
use crate::checkpoint::{capture_checkpoint, plan_resume, CheckpointId, CheckpointPool};
use crate::forward_slots::{forward_batch_slots_graphed_opts, SlotDecodeGraph, SlotDescStaging};
use crate::grammar;
use crate::qwen35::{
Expand Down Expand Up @@ -112,7 +112,7 @@ pub struct EngineConfig {
/// Cross-session prefix cache (radix index + checkpoint pool). Default
/// false: the existing session-local continuation path (convo hashes)
/// is unchanged. When true, cold admits consult a PrefixIndex for
/// reusable KV pages and a QwenCheckpointPool for recurrent state
/// reusable KV pages and a CheckpointPool for recurrent state
/// (spec §4.5–4.6). Requires paged KV (PagePool).
pub prefix_cache: bool,
/// Maximum device bytes for the checkpoint pool. 0 = pool that refuses
Expand Down Expand Up @@ -388,7 +388,7 @@ struct ModelRig {
vl_tower_jobs: Vec<Option<hipfire_arch_qwen35_vl::qwen35_vl::VisionTowerJob>>,
/// Cross-session recurrent-state checkpoint pool (spec §4.5 C5).
/// None when `prefix_cache` is off.
checkpoint_pool: Option<crate::checkpoint::QwenCheckpointPool<DeltaNetSnapshot>>,
checkpoint_pool: Option<crate::checkpoint::CheckpointPool<DeltaNetSnapshot>>,
}

impl ModelRig {
Expand Down Expand Up @@ -1546,7 +1546,7 @@ impl Rig {
// holds its own ceiling on snapshot bytes).
let mut idx = idx;
idx.set_max_retained_bytes(cfg.prefix_cache_max_bytes as usize);
let pool = crate::checkpoint::QwenCheckpointPool::<DeltaNetSnapshot>::new(
let pool = crate::checkpoint::CheckpointPool::<DeltaNetSnapshot>::new(
cfg.prefix_cache_max_bytes,
);
eprintln!(
Expand Down
16 changes: 8 additions & 8 deletions crates/hipfire-arch-qwen4/map.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,29 +26,29 @@ _Generated by `scripts/check-crate-maps.py` from the tree — do not edit inside
|---|---:|---:|---:|
| [`src/admission.rs`](src/admission.rs) | 684 | 23 | 4 |
| [`src/artifact.rs`](src/artifact.rs) | 1,130 | 3 | 5 |
| [`src/bundle.rs`](src/bundle.rs) | 1,955 | 42 | 1 |
| [`src/bundle.rs`](src/bundle.rs) | 1,881 | 32 | 1 |
| [`src/config.rs`](src/config.rs) | 981 | 27 | 7 |
| [`src/expert_residency.rs`](src/expert_residency.rs) | 849 | 25 | 11 |
| [`src/gpu_forward.rs`](src/gpu_forward.rs) | 3,840 | 16 | 0 |
| [`src/kv_backend.rs`](src/kv_backend.rs) | 214 | 10 | 2 |
| [`src/lib.rs`](src/lib.rs) | 53 | 18 | 0 |
| [`src/mtp_gpu.rs`](src/mtp_gpu.rs) | 2,839 | 4 | 7 |
| [`src/mtp_spec.rs`](src/mtp_spec.rs) | 1,944 | 16 | 7 |
| [`src/mtp_gpu.rs`](src/mtp_gpu.rs) | 2,857 | 4 | 7 |
| [`src/mtp_spec.rs`](src/mtp_spec.rs) | 1,935 | 16 | 7 |
| [`src/ops.rs`](src/ops.rs) | 780 | 23 | 3 |
| [`src/ple.rs`](src/ple.rs) | 752 | 30 | 7 |
| [`src/program.rs`](src/program.rs) | 840 | 26 | 0 |
| [`src/projection.rs`](src/projection.rs) | 145 | 0 | 2 |
| [`src/reference_forward.rs`](src/reference_forward.rs) | 1,091 | 22 | 1 |
| [`src/reference_mtp.rs`](src/reference_mtp.rs) | 1,327 | 51 | 4 |
| [`src/state.rs`](src/state.rs) | 2,495 | 28 | 7 |
| [`src/state_parity.rs`](src/state_parity.rs) | 2,502 | 8 | 1 |
| [`src/state.rs`](src/state.rs) | 2,527 | 27 | 7 |
| [`src/state_parity.rs`](src/state_parity.rs) | 2,505 | 8 | 1 |
| [`src/weights.rs`](src/weights.rs) | 2,723 | 33 | 10 |

### Public API surface

- [`src/admission.rs`](src/admission.rs): `InputModality`, `RequestKind`, `Qwen4Capability`, `Qwen4Capabilities`, `fn`, `EffectiveMesh`, `MeshShape`, `Qwen4ArtifactFormat`, `ArtifactFormat`, `QuantFormat`, `QuantTensorGeometry`, `payload_bytes`, +11 more
- [`src/artifact.rs`](src/artifact.rs): `Qwen4HfqmArtifact`, `Qwen4ArtifactError`, `admit_hfqm_artifact`
- [`src/bundle.rs`](src/bundle.rs): `Qwen4Bundle`, `QWEN4_PREFIX_CACHE_DEFAULT`, `prefix_cache_requested`, `prefix_cache_device_bytes`, `Qwen4PrefixMode`, `Qwen4PrefixPlan`, `assemble`, `assemble_with_metadata`, `manifest`, `external_descriptor`, `ple_descriptors`, `ple_rows`, +30 more
- [`src/bundle.rs`](src/bundle.rs): `Qwen4Bundle`, `session_snapshot_bytes`, `assemble`, `assemble_with_metadata`, `manifest`, `external_descriptor`, `ple_descriptors`, `ple_rows`, `ple_rows_mut`, `attached_origin`, `attached_inventory_len`, `attached_external_rows_len`, +20 more
- [`src/config.rs`](src/config.rs): `ARCH_ID`, `MODEL_TYPE`, `TEXT_MODEL_TYPE`, `ARCHITECTURE_NAME`, `LayerType`, `parse`, `fn`, `RecurrentStateDType`, `SourceDType`, `Qwen4MtpConfig`, `validate`, `Qwen4Config`, +15 more
- [`src/expert_residency.rs`](src/expert_residency.rs): `EXPERT_VRAM_LAYERS_ENV`, `AUTO_VRAM_RESERVE_BYTES`, `AUTO_VRAM_RESERVE_MAX_SEQ`, `AUTO_VRAM_RESERVE_CHUNK`, `HOST_RAM_HEADROOM_BYTES`, `HOST_MAPPED_PAD_BYTES`, `ExpertVramLayers`, `expert_vram_layers_from_env`, `resolve_expert_vram_layers`, `fits_fully_resident`, `keeps_host_memory_out_of_reclaim`, `routed_expert_bytes`, +13 more
- [`src/gpu_forward.rs`](src/gpu_forward.rs): `qwen4_prefill_chunk_default`, `qwen4_prefill_chunk_requested`, `qwen4_forward_device_bytes`, `QWEN4_EXPERT_STAGE_ENV`, `QWEN4_EXPERT_STAGE_MIN_ROWS_ENV`, `Qwen4OutputRows`, `Qwen4GpuForwardError`, `Qwen4GpuForwardScratch`, `new`, `Qwen4QsaTap`, `Qwen4GpuForward`, `ROUTE_TRACE_ENV`, +4 more
Expand All @@ -62,7 +62,7 @@ _Generated by `scripts/check-crate-maps.py` from the tree — do not edit inside
- [`src/projection.rs`](src/projection.rs): —
- [`src/reference_forward.rs`](src/reference_forward.rs): `ForwardError`, `CompactForwardConfig`, `validate`, `wide`, `q_width`, `kv_width`, `index_width`, `gdn_qk_width`, `gdn_v_width`, `selected_capacity`, `ReferenceHyperWeights`, `ReferenceGdnWeights`, +10 more
- [`src/reference_mtp.rs`](src/reference_mtp.rs): `MTP_LAYER_COUNT`, `MTP_BRANCHES`, `MTP_SELECTED_CAPACITY`, `MTP_TOP_K`, `MTP_RMS_EPS`, `MTP_Q_HEADS`, `MTP_KV_HEADS`, `MTP_HEAD_DIM`, `MTP_INDEX_HEADS`, `MTP_INDEX_DIM`, `MTP_INDEX_BUDGET`, `MTP_COMPRESS_RATIO`, +39 more
- [`src/state.rs`](src/state.rs): `ReferenceGdnLayerState`, `new`, `reset`, `conv_history_rows`, `push_conv_row`, `ReferenceQsaLayerState`, `ReferenceHyperConnectionState`, `GdnGpuState`, `QsaGpuState`, `resolve_qsa_format`, `resolve_gdn_format`, `Qwen4StateFormat`, +16 more
- [`src/state.rs`](src/state.rs): `ReferenceGdnLayerState`, `new`, `reset`, `conv_history_rows`, `push_conv_row`, `ReferenceQsaLayerState`, `ReferenceHyperConnectionState`, `GdnGpuState`, `QsaGpuState`, `resolve_qsa_format`, `resolve_gdn_format`, `Qwen4StateFormat`, +15 more
- [`src/state_parity.rs`](src/state_parity.rs): `run_compact`, `StateParityReport`, `into_json`, `write`, `ProfileReport`, `run_profile`, `run_state_parity`, `run_mtp_fill_digest`
- [`src/weights.rs`](src/weights.rs): `ROUTED_GATE_UP_DTYPE`, `ROUTED_DOWN_DTYPE`, `PLE_SHARD_ROWS`, `crate`, `PLE_SHARD_COUNT`, `TensorRole`, `TensorRef`, `new`, `MetadataTensor`, `Qwen4Manifest`, `build`, `entry`, +21 more

Expand All @@ -79,6 +79,6 @@ _Generated by `scripts/check-crate-maps.py` from the tree — do not edit inside

### Totals

- 19 modules · 27,144 lines · 405 public items · 86 tests · 5 examples
- 19 modules · 27,114 lines · 394 public items · 87 tests · 5 examples

<!-- crate-map:generated:end -->
Loading