Skip to content

feat(runtime): delta session snapshots — Flash-Next snapshot costs one chunk at any depth (builds on #825) - #826

Open
fivetide wants to merge 6 commits into
warpfront:betafrom
fivetide:feat/qwen4-session-cache-delta
Open

fivetide wants to merge 6 commits into
warpfront:betafrom
fivetide:feat/qwen4-session-cache-delta

Conversation

@fivetide

@fivetide fivetide commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

Summary

Builds on #825; until that merges, this PR also contains its commit. Review the delta commit 872b0a128 on its own.

Session snapshots become deltas. A snapshot stores the state that is overwritten in place whole, but of each append-only row stream only the rows above its parent: the deepest snapshot of the same prefix that existed when it was captured. A restore walks the chain from the root.

 trait SessionState
-  snapshot_parts → { meta, parts: Vec<StatePart> }
+  snapshot_parts → { meta, layout: StateLayout { fixed: Vec<StatePart>, rows: Vec<RowStream> } }
-  restore_parts  → Vec<StatePart>
+  restore_parts  → StateLayout      # destination, with the snapshot's row counts
at_boundary(p)
  parent = deepest snapshot of prefix[..b], b < p, published or pending
  store fixed parts whole + rows [parent.to, live.rows) per stream
  link(parent)  → pinned while it has children

begin(reused)
  chain = leaf → parent → … → root      (each get refreshes LRU)
  one copy_regions: leaf's fixed parts + every link's row segment at row offset `from`
  check: segments contiguous from 0, end at the destination's row counts

evict: pop_lru only sees unpinned entries → leaves only; dropping a leaf unlinks (maybe unpins) its parent
commit: parents before children; a refused parent drops its pending children
  • Qwen4 code: the target and MTP-head codecs now just label their QSA K/V, raw and pooled arenas as RowStreams.
  • Unchanged: the generate loop, MTP driver, daemon, loader and dispatch interpreter.
  • Cost:
per 8192-token Flash-Next MTP snapshot before after
snapshot at depth k chunks k × 481 + 32 MiB 513 MiB
one 64K session, all 8 boundaries ≈ 17 GiB ≈ 4 GiB
subagents sharing a system-prompt chunk each pays shared link

Which surface(s) does this touch?

  • serve — runtime session_cache
  • arch crate(s): hipfire-arch-qwen4 (codec layout only)

Evidence

Strix Halo gfx1151, HIP 7.2, qwen3.8-flash-next-gptq3.mq4 sha256 8b15b6fede7d7c5bfed0db4720a8295bedda51bc93e545fa242bd50d0f200972. QSA is F32 (the serve default here), GDN is Q8, the context backend is VMM.

  • Bit-exact through chains and branches (session_cache_hw, extended). A = 3 chunks + 300 tokens; B = A's first chunk + 8,392 other tokens. Both are prefilled cold first, so B's 2-chunk snapshot is a delta over A's root. Then A is restored through its 3-link chain and B through the shared root:
    A cold ids [248045, 248045, 248045, 248045, 248046, 198, 248045, 74455, …]
    A warm ids [248045, 248045, 248045, 248045, 248046, 198, 248045, 74455, …]
    B cold ids [5513, 248046, 198, 248045, 74455, 198, 248068, 271, …]
    B warm ids [5513, 248046, 198, 248045, 74455, 198, 248068, 271, …]
    test restored_prefill_matches_cold_on_flash_next ... ok
    
    Final logits are byte-equal for both prompts. This also proves the invariant the design depends on: rows below a boundary never change after it.
  • Unit test (GPU toy state). The state has a fixed part plus per-token rows that depend causally on the prefix. Checked: an exact restore through a 3-link chain and a branch; leaf-only eviction (the full budget evicts A's deepest leaf, A still restores at 2 chunks, the shared root stays pinned); pins released when an uncommitted turn is abandoned; snapshots over the budget are skipped.
  • Serve, greedy, MTP:
    A 21550 tok  15.9 s  cached 0
    B  8142 tok   6.2 s  cached 0
    C 14154 tok   5.9 s  cached 8192    # C = first 2/3 of A's text + another question; output == cache-off C (11.1 s)
    A 21550 tok   4.9 s  cached 16384   # 2-link MTP chain; output == A#1
    
  • Static checks: cargo build --release and clippy on the changed crates are clean; the hipfire-runtime session_cache and checkpoint_pool tests pass; layering and crate maps are clean.

Test plan

  • cargo build --release clean
  • session_cache_hw (ignored hardware test) and the runtime session_cache test pass on gfx1151
  • serve greedy MTP A/B/C/A matches cache-off and cold outputs
  • CI jobs green

Merge Danger

Door: two-way. Snapshots live only in process memory; a revert or memory.session_cache_bytes = 0 returns to the earlier behaviour.

Blast Radius: Flash-Next session cache.

  • Corruption propagates: a corrupted parent now breaks every descendant. The guards are the exact-logits chain and branch test, and the layout checks on restore (contiguous segments, row sizes, final row counts).
  • Pinned ancestors: they can hold budget while their leaves are hot. Eviction then skips the capture rather than oversubscribing.

…pts warm

Add hipfire_runtime::session_cache: one cache per loaded model that owns
keying (CacheDomain scoped per route/chunk/state format + prefix
fingerprint), planning, LRU eviction (CheckpointPool, renamed from
QwenCheckpointPool, + pop_lru), a memory guard (UMA: MemAvailable with
8 GiB headroom; discrete: free VRAM), placement (SnapshotLocation, device
only for now) and every copy. Architectures implement SessionState and only
describe their state parts; ArchModel gains session_cache_attached,
session_plan and session_commit.

Qwen4 (Flash-Next) moves onto it and drops its single whole-chunk prefix
checkpoint. Snapshots are self-contained (GDN, PLE, hyper state plus QSA
K/V, raw and pooled rows below the boundary; with MTP the head and draft
policy), captured at every prefill-chunk boundary and published only on
client commit, so several sessions/subagents stay warm and a reset no
longer drops them. New config memory.session_cache_bytes /
HIPFIRE_SESSION_CACHE_BYTES (default 8 GiB, 0 disables).

Verification (Strix Halo gfx1151, qwen3.8-flash-next-gptq3.mq4 8b15b6fe):
- session_cache_hw: two-chunk restore after an unrelated prompt, final
  logits byte-equal to cold, 16 greedy ids equal.
- serve A(21550 tok)/B(8142)/A, greedy MTP: A#2 cached_tokens 16384,
  16.3 s -> 4.9 s, identical content; cache off: cached 0, same content.
- qwen4_mtp_fill warm (300,4096 after 1 chunk): plan == 8192.
- serve_harness chain greedy, MTP on and off: 5/5 turns, identical text.
@fivetide
fivetide requested a review from Kaden-Schutt as a code owner October 6, 2026 14:38
fivetide pushed a commit to fivetide/hipfire that referenced this pull request Oct 6, 2026
…tness

From the warpfront#826 review:
- at_boundary: a failing free-memory query (dGPU get_vram_info) no longer
  returns early with the parent pinned by a phantom child; it skips the
  capture ("memory query failed") through the unlink path, without evicting,
  and the prefill continues.
- docs/env-vars.md: add HIPFIRE_SESSION_CACHE_MODEL (lifecycle gate).
- CheckpointPool::pop_lru: host test that it returns the oldest unpinned
  entry and never a pinned one.
- SessionCache::stored_bytes + Qwen4Bundle::session_cache accessors.
- session_cache_hw: compare a SHA-256 of every state part (meta, fixed
  parts, valid rows of every row stream) cold vs restored; add the native
  MTP route (target + draft head + policy, seed); assert B's snapshot is a
  delta over A's root (four one-chunk snapshots per route).

gfx1151, qwen3.8-flash-next-gptq3.mq4 8b15b6fe, F32 QSA / Q8 GDN / VMM:
AR and MTP, A via 3-link chain and B via shared root: all state parts
equal, AR logits byte-equal + 16 ids equal, MTP seeds equal; stored bytes
exactly 4 x 499,698,688 (AR) and 4 x 538,545,664 (MTP).
@fivetide

fivetide commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Review fixes pushed in 7b715e531:

# Finding Fix
1 HIPFIRE_SESSION_CACHE_MODEL missing from the env inventory (lifecycle gate failed) Regenerated with check-lifecycle.py --write; the gate passes.
2 Failed free-VRAM query left the parent pinned by a phantom child and aborted the prefill The query failure is now a capture skip (memory query failed) through the unlink path, with no eviction; the prefill continues.
3 Exactness checked only through final logits and 16 ids session_cache_hw compares a SHA-256 of every state part (metadata, fixed parts, valid rows of every row stream), cold vs restored.
4 No exactness check for the MTP draft head The test adds the native MTP route (target, draft head and its policy, seed token) over the same chain and branch.
5 B's link to A's root not asserted The test asserts each route stores exactly four one-chunk snapshots.
6 No GPU-free test for pop_lru Host test: it returns the oldest unpinned entry and never a pinned one.

gfx1151, qwen3.8-flash-next-gptq3.mq4 (8b15b6fe…), F32 QSA, Q8 GDN, VMM:

AR:  1998794752 bytes stored = 4 × 499698688 per one-chunk snapshot
AR A / AR B   all state parts equal, final logits byte-equal, 16 ids equal
MTP: 2154182656 bytes stored = 4 × 538545664
MTP A seed cold 248045 warm 248045, MTP B seed cold 5513 warm 5513, all state parts equal
test restored_prefill_matches_cold_on_flash_next ... ok

The runtime session_cache and checkpoint_pool tests pass (28). Lifecycle, env-docs, layering and crate-map checks pass, and cargo build --release is clean.

Bjoern Agent added 4 commits October 6, 2026 17:27
…; refresh crate maps and env inventory

The formatter rewrote all of hipfire-daemon/src/main.rs (5392 -> 5477
lines, over the 5377 daemon_lines ceiling). Keep only the 3-line
session-cache change. Regenerate the stale crate maps and add
HIPFIRE_SESSION_CACHE_MODEL (session_cache_hw) to docs/env-vars.md.
… chunk at any depth

Session snapshots now store the fixed (overwritten-in-place) state whole
but, of each append-only row stream, only the rows above their parent: the
deepest snapshot of the same prefix present at capture time. Restores walk
the chain from the root in one copy_regions call; prefixes shared between
sessions (subagents with one system prompt) share their links. A snapshot
with children is pinned, so eviction only removes leaves; abandoned or
refused snapshots unlink from their parents.

SessionState now returns a StateLayout { fixed: Vec<StatePart>,
rows: Vec<RowStream> }; the Qwen4 target and MTP head codecs mark their QSA
K/V, raw and pooled arenas as row streams. No consumer changes.

Per 8192-token Flash-Next MTP snapshot: 513 MiB at any depth (was
481 MiB per chunk of depth + 32 MiB; a 64K session's eight boundaries
~4 GiB instead of ~17 GiB).

Verification (gfx1151, qwen3.8-flash-next-gptq3.mq4 8b15b6fe):
- session_cache_hw: A (3 chunks + 300) and B (A's first chunk + other
  tokens) cold, then A through its 3-link chain and B through the shared
  root: final logits byte-equal to cold, 16 greedy ids equal.
- unit test: chain/branch restore exact, leaf-only LRU eviction, pins
  released by abandoned turns, over-budget skip.
- serve greedy MTP A/B/C/A: C (shares A's first chunk) cached 8192 and
  equals its cache-off output; A#2 cached 16384 (2-link chain), equal to A#1.
…tness

From the warpfront#826 review:
- at_boundary: a failing free-memory query (dGPU get_vram_info) no longer
  returns early with the parent pinned by a phantom child; it skips the
  capture ("memory query failed") through the unlink path, without evicting,
  and the prefill continues.
- docs/env-vars.md: add HIPFIRE_SESSION_CACHE_MODEL (lifecycle gate).
- CheckpointPool::pop_lru: host test that it returns the oldest unpinned
  entry and never a pinned one.
- SessionCache::stored_bytes + Qwen4Bundle::session_cache accessors.
- session_cache_hw: compare a SHA-256 of every state part (meta, fixed
  parts, valid rows of every row stream) cold vs restored; add the native
  MTP route (target + draft head + policy, seed); assert B's snapshot is a
  delta over A's root (four one-chunk snapshots per route).

gfx1151, qwen3.8-flash-next-gptq3.mq4 8b15b6fe, F32 QSA / Q8 GDN / VMM:
AR and MTP, A via 3-link chain and B via shared root: all state parts
equal, AR logits byte-equal + 16 ids equal, MTP seeds equal; stored bytes
exactly 4 x 499,698,688 (AR) and 4 x 538,545,664 (MTP).
@fivetide
fivetide force-pushed the feat/qwen4-session-cache-delta branch from 73bc246 to 3eeb950 Compare October 6, 2026 15:28
Formatter churn: restore the base formatting of every file this branch
touched for unrelated reasons (qwen35 serve_engine/checkpoint, config,
loader carriers, generate qwen, qwen4 state/mtp_gpu, runtime lib/arch_model/
checkpoint_pool). The rename to CheckpointPool and the session-cache code
are kept.

Session cache:
- SessionCache drops its `budget` copy and reads `pool.max_bytes()`.
- SessionState::snapshot_boundaries loses the unused route parameter and
  snapshot_parts the unused gpu parameter.
- restore/at_boundary drop explicit drop()s and a meta clone NLL makes
  unnecessary; the parent search looks each candidate up once.
- Qwen4 session meta drops the bundle header (route is already in the
  scope key; the target meta has a fixed length, session_meta_bytes) and
  the redundant EOS-id and QSA-count words.

Verified on halo (gfx1151): release build; hipfire-runtime session_cache +
checkpoint_pool tests; session_cache_hw on qwen3.8-flash-next-gptq3.mq4
(AR and MTP chain/branch restores byte-equal to cold, ids and MTP seeds
equal; one-chunk snapshots 499698688 B AR, 538545664 B MTP); serve A/A/A' on a
21K-token prompt: 16.7 s cold, 4.8 s warm with cached_tokens 16384 and
identical text.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant