Repository navigation
Conversation
…pts warm Add hipfire_runtime::session_cache: one cache per loaded model that owns keying (CacheDomain scoped per route/chunk/state format + prefix fingerprint), planning, LRU eviction (CheckpointPool, renamed from QwenCheckpointPool, + pop_lru), a memory guard (UMA: MemAvailable with 8 GiB headroom; discrete: free VRAM), placement (SnapshotLocation, device only for now) and every copy. Architectures implement SessionState and only describe their state parts; ArchModel gains session_cache_attached, session_plan and session_commit. Qwen4 (Flash-Next) moves onto it and drops its single whole-chunk prefix checkpoint. Snapshots are self-contained (GDN, PLE, hyper state plus QSA K/V, raw and pooled rows below the boundary; with MTP the head and draft policy), captured at every prefill-chunk boundary and published only on client commit, so several sessions/subagents stay warm and a reset no longer drops them. New config memory.session_cache_bytes / HIPFIRE_SESSION_CACHE_BYTES (default 8 GiB, 0 disables). Verification (Strix Halo gfx1151, qwen3.8-flash-next-gptq3.mq4 8b15b6fe): - session_cache_hw: two-chunk restore after an unrelated prompt, final logits byte-equal to cold, 16 greedy ids equal. - serve A(21550 tok)/B(8142)/A, greedy MTP: A#2 cached_tokens 16384, 16.3 s -> 4.9 s, identical content; cache off: cached 0, same content. - qwen4_mtp_fill warm (300,4096 after 1 chunk): plan == 8192. - serve_harness chain greedy, MTP on and off: 5/5 turns, identical text.
fivetide
pushed a commit
to fivetide/hipfire
that referenced
this pull request
Oct 6, 2026
…tness From the warpfront#826 review: - at_boundary: a failing free-memory query (dGPU get_vram_info) no longer returns early with the parent pinned by a phantom child; it skips the capture ("memory query failed") through the unlink path, without evicting, and the prefill continues. - docs/env-vars.md: add HIPFIRE_SESSION_CACHE_MODEL (lifecycle gate). - CheckpointPool::pop_lru: host test that it returns the oldest unpinned entry and never a pinned one. - SessionCache::stored_bytes + Qwen4Bundle::session_cache accessors. - session_cache_hw: compare a SHA-256 of every state part (meta, fixed parts, valid rows of every row stream) cold vs restored; add the native MTP route (target + draft head + policy, seed); assert B's snapshot is a delta over A's root (four one-chunk snapshots per route). gfx1151, qwen3.8-flash-next-gptq3.mq4 8b15b6fe, F32 QSA / Q8 GDN / VMM: AR and MTP, A via 3-link chain and B via shared root: all state parts equal, AR logits byte-equal + 16 ids equal, MTP seeds equal; stored bytes exactly 4 x 499,698,688 (AR) and 4 x 538,545,664 (MTP).
Collaborator
Author
|
Review fixes pushed in
gfx1151, The runtime |
added 4 commits
October 6, 2026 17:27
…; refresh crate maps and env inventory The formatter rewrote all of hipfire-daemon/src/main.rs (5392 -> 5477 lines, over the 5377 daemon_lines ceiling). Keep only the 3-line session-cache change. Regenerate the stale crate maps and add HIPFIRE_SESSION_CACHE_MODEL (session_cache_hw) to docs/env-vars.md.
… chunk at any depth
Session snapshots now store the fixed (overwritten-in-place) state whole
but, of each append-only row stream, only the rows above their parent: the
deepest snapshot of the same prefix present at capture time. Restores walk
the chain from the root in one copy_regions call; prefixes shared between
sessions (subagents with one system prompt) share their links. A snapshot
with children is pinned, so eviction only removes leaves; abandoned or
refused snapshots unlink from their parents.
SessionState now returns a StateLayout { fixed: Vec<StatePart>,
rows: Vec<RowStream> }; the Qwen4 target and MTP head codecs mark their QSA
K/V, raw and pooled arenas as row streams. No consumer changes.
Per 8192-token Flash-Next MTP snapshot: 513 MiB at any depth (was
481 MiB per chunk of depth + 32 MiB; a 64K session's eight boundaries
~4 GiB instead of ~17 GiB).
Verification (gfx1151, qwen3.8-flash-next-gptq3.mq4 8b15b6fe):
- session_cache_hw: A (3 chunks + 300) and B (A's first chunk + other
tokens) cold, then A through its 3-link chain and B through the shared
root: final logits byte-equal to cold, 16 greedy ids equal.
- unit test: chain/branch restore exact, leaf-only LRU eviction, pins
released by abandoned turns, over-budget skip.
- serve greedy MTP A/B/C/A: C (shares A's first chunk) cached 8192 and
equals its cache-off output; A#2 cached 16384 (2-link chain), equal to A#1.
…tness From the warpfront#826 review: - at_boundary: a failing free-memory query (dGPU get_vram_info) no longer returns early with the parent pinned by a phantom child; it skips the capture ("memory query failed") through the unlink path, without evicting, and the prefill continues. - docs/env-vars.md: add HIPFIRE_SESSION_CACHE_MODEL (lifecycle gate). - CheckpointPool::pop_lru: host test that it returns the oldest unpinned entry and never a pinned one. - SessionCache::stored_bytes + Qwen4Bundle::session_cache accessors. - session_cache_hw: compare a SHA-256 of every state part (meta, fixed parts, valid rows of every row stream) cold vs restored; add the native MTP route (target + draft head + policy, seed); assert B's snapshot is a delta over A's root (four one-chunk snapshots per route). gfx1151, qwen3.8-flash-next-gptq3.mq4 8b15b6fe, F32 QSA / Q8 GDN / VMM: AR and MTP, A via 3-link chain and B via shared root: all state parts equal, AR logits byte-equal + 16 ids equal, MTP seeds equal; stored bytes exactly 4 x 499,698,688 (AR) and 4 x 538,545,664 (MTP).
fivetide
force-pushed
the
feat/qwen4-session-cache-delta
branch
from
October 6, 2026 15:28
73bc246 to
3eeb950
Compare
Formatter churn: restore the base formatting of every file this branch touched for unrelated reasons (qwen35 serve_engine/checkpoint, config, loader carriers, generate qwen, qwen4 state/mtp_gpu, runtime lib/arch_model/ checkpoint_pool). The rename to CheckpointPool and the session-cache code are kept. Session cache: - SessionCache drops its `budget` copy and reads `pool.max_bytes()`. - SessionState::snapshot_boundaries loses the unused route parameter and snapshot_parts the unused gpu parameter. - restore/at_boundary drop explicit drop()s and a meta clone NLL makes unnecessary; the parent search looks each candidate up once. - Qwen4 session meta drops the bundle header (route is already in the scope key; the target meta has a fixed length, session_meta_bytes) and the redundant EOS-id and QSA-count words. Verified on halo (gfx1151): release build; hipfire-runtime session_cache + checkpoint_pool tests; session_cache_hw on qwen3.8-flash-next-gptq3.mq4 (AR and MTP chain/branch restores byte-equal to cold, ids and MTP seeds equal; one-chunk snapshots 499698688 B AR, 538545664 B MTP); serve A/A/A' on a 21K-token prompt: 16.7 s cold, 4.8 s warm with cached_tokens 16384 and identical text.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Builds on #825; until that merges, this PR also contains its commit. Review the delta commit
872b0a128on its own.Session snapshots become deltas. A snapshot stores the state that is overwritten in place whole, but of each append-only row stream only the rows above its parent: the deepest snapshot of the same prefix that existed when it was captured. A restore walks the chain from the root.
RowStreams.Which surface(s) does this touch?
session_cachehipfire-arch-qwen4(codec layout only)Evidence
Strix Halo gfx1151, HIP 7.2,
qwen3.8-flash-next-gptq3.mq4sha2568b15b6fede7d7c5bfed0db4720a8295bedda51bc93e545fa242bd50d0f200972. QSA is F32 (the serve default here), GDN is Q8, the context backend is VMM.session_cache_hw, extended). A = 3 chunks + 300 tokens; B = A's first chunk + 8,392 other tokens. Both are prefilled cold first, so B's 2-chunk snapshot is a delta over A's root. Then A is restored through its 3-link chain and B through the shared root:cargo build --releaseand clippy on the changed crates are clean; thehipfire-runtimesession_cacheandcheckpoint_pooltests pass; layering and crate maps are clean.Test plan
cargo build --releasecleansession_cache_hw(ignored hardware test) and the runtimesession_cachetest pass on gfx1151Merge Danger
Door: two-way. Snapshots live only in process memory; a revert or
memory.session_cache_bytes = 0returns to the earlier behaviour.Blast Radius: Flash-Next session cache.