Skip to content

fix(cli): wire serve --max-seq override - #828

Open
HUSRCF wants to merge 1 commit into
warpfront:betafrom
HUSRCF:fix/serve-max-seq-821
Open

HUSRCF wants to merge 1 commit into
warpfront:betafrom
HUSRCF:fix/serve-max-seq-821

Conversation

@HUSRCF

@HUSRCF HUSRCF commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Summary

Add the documented hipfire serve --max-seq <N> override and carry it through detached startup, every model load/switch, and multi-slot context sizing. This closes #821.

Which surface(s) does this touch?

  • kernel — kernels/, crates/rdna-compute, crates/hipfire-dispatch, crates/hip-bridge, crates/saddle-core
  • load — crates/hipfire-loader, crates/hipfire-daemon, runtime load path (model_load, hfq, loader_api, config, safetensors_source, weight_backend, multi_gpu), arch load*/weights*/carrier.rs, hipfire-config, hipfire-registry, registry/, Cargo manifests
  • serve — crates/hipfire-engine, crates/hipfire-generate, daemon slots/serve, runtime emit/eos/dflash/dspark/spec/reset/triattn
  • arch crate(s)
  • crates/hipfire-quantize / quant formats
  • control plane — hipfire-cli, hipfire-client, hipfire-tui
  • docs / CI / scripts only (no hardware route)
  • policy files

Behavior

  • Parse and validate serve --max-seq through the existing memory.max_seq schema contract.
  • Forward it when serve --detach launches the foreground child.
  • Apply it after model/config load parameters are resolved, so the explicit CLI value wins for prewarm, lazy load, reload, and request-driven model switches.
  • Use the same explicit value for multi_slot_ctx, avoiding a mismatch between advertised context and slot capacity.

Test plan

  • ./scripts/no-gpu-ci.sh passes, or the CI jobs are green
  • cargo build --release -p hipfire-cli
  • cargo test -p hipfire-cli — 337 passed, 0 failed
  • Focused fake-daemon test loads two different installed models in one serve process and observes params.max_seq=65536 on both load operations
  • hipfire serve --help advertises --max-seq <N>
  • Real gfx1100/W7900 smoke with matched CLI and daemon binaries: qwen3.5-0.8b.mq4, Q8 KV, --max-seq 65536; load reported physical_cap=65536, /health reported n_ctx=65536 and multi_slot_ctx=65536, and a 40,010-prompt-token request completed successfully
  • cargo test --lib --workspace passes
  • If perf-relevant: ./scripts/speed-gate.sh within ±2% of locked baselines

./scripts/no-gpu-ci.sh completed its Rust stages, then reported 541 Python tests passed and 21 failed in scripts/hw-gate/tests/test_review.py. The same representative failure was independently reproduced on an unmodified worktree at the parent upstream/beta commit: the local scripts/hw-gate/select.py shadows Python's standard-library select module while importing subprocess/selectors. This PR does not touch that path, so the checkbox remains unchecked pending upstream CI.

This is not performance-relevant and changes no GPU/kernel route.

Architecture-trait change?

No. This does not change Architecture or any architecture crate.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant