Skip to content

[AMD][MI35X] Bump Qwen3.5 MXFP4 MI355X SGLang AgentX to v0.5.18 and retune all-reduce, prefill, and CUDA graph - #2737

Open
yichiche wants to merge 6 commits into
mainfrom
amd/qwen3.5-fp4-mi355x-sglang-agentic-v0.5.18
Open

[AMD][MI35X] Bump Qwen3.5 MXFP4 MI355X SGLang AgentX to v0.5.18 and retune all-reduce, prefill, and CUDA graph#2737
yichiche wants to merge 6 commits into
mainfrom
amd/qwen3.5-fp4-mi355x-sglang-agentic-v0.5.18

Conversation

@yichiche

@yichiche yichiche commented Aug 26, 2026

Copy link
Copy Markdown
Collaborator

Motivation

qwen3.5-fp4-mi355x-sglang-agentic-mtp is still on lmsysorg/sglang-rocm:v0.5.17-rocm720-mi35x-20260818, and three of its launch knobs are out of step with the sibling recipes on the same cluster.

First, all-reduce. The script correctly omits --enable-aiter-allreduce-fusion (removed in #2562 for TP2/EP2 EAGLE rank consistency) but never sets ROCM_QUICK_REDUCE_QUANTIZATION, so multi-GPU collectives run unquantized custom all-reduce. The published SGLang cookbook recipe for MXFP4 on MI355X calls for INT8-quantized quick all-reduce.

Second, the prefill budget is double the B200 sibling: --max-prefill-tokens 32768 / --chunked-prefill-size 32768 here versus 16384 / 16384 in benchmarks/single_node/agentic/qwen3.5_fp4_b200_sglang_mtp.sh.

Third, the decode CUDA graph is undersized for the batch the scheduler actually builds. --max-running-requests is 2*CONC, but the graph was only captured to min(CONC, 64), so every decode batch above CONC fell onto the eager path.

Modifications

Bump image on qwen3.5-fp4-mi355x-sglang-agentic-mtp to lmsysorg/sglang-rocm:v0.5.18-rocm720-mi35x-20260825. The model, runner, and search space are untouched.

In benchmarks/single_node/agentic/qwen3.5_fp4_mi355x_sglang_mtp.sh, add export ROCM_QUICK_REDUCE_QUANTIZATION=INT8 alongside the existing aiter exports. The two all-reduce paths are mutually exclusive in SGLang: the AITER fused AR+RMSNorm path is gated on --enable-aiter-allreduce-fusion, and only with it off do collectives fall back to custom all-reduce, where the quick-reduce regime is read. Since this arm already omits the fusion flag, the env takes effect with no other launch change.

Lower --max-prefill-tokens and --chunked-prefill-size from 32768 to 16384, matching the B200 sibling.

Capture the decode CUDA graph to min(2*CONC, 128) instead of min(CONC, 64), so it covers --max-running-requests. The expression and the 128 cap are taken verbatim from the sibling MI355X AgentX recipe benchmarks/single_node/agentic/dsv4_fp4_mi355x_sglang_mtp.sh, which already runs this idiom on cluster:mi355x-amds.

Append the corresponding perf-changelog.yaml trigger.

Accuracy Tests

No accuracy-affecting logic changes in this repo. INT8 quick all-reduce quantizes the collective payload, so it is not bit-identical to the unquantized path; the AgentX eval rows continue to run real target-model verification and will show whether that matters. The prefill and CUDA-graph knobs change scheduling and kernel launch shape, not what is computed.

Benchmarking

Repo validation was run locally:

  • python3 -m pytest utils/matrix_logic/ -q → 232 passed.
  • bash -n benchmarks/single_node/agentic/qwen3.5_fp4_mi355x_sglang_mtp.sh → clean.
  • python3 utils/matrix_logic/generate_sweep_configs.py full-sweep --config-files configs/amd-master.yaml --model-prefix qwen3.5 --precision fp4 --scenario-type agentic-coding → 16 configs, all on v0.5.18-rocm720-mi35x-20260825.
  • Hand-evaluated the new CUDA-graph expression across the config's conc-list (TP2 1,4,8,12,16,20; TP4 1,4,8,12,16,20,24,28,32,40): it yields 2*CONC at every point, so the 128 cap never binds in the current sweep.

End-to-end MI355X AgentX numbers will come from the sweep triggered on this PR (full-sweep-fail-fast). Three knobs plus an image move together here, so the resulting numbers are not attributable to any single one of them; if the arm regresses, the all-reduce regime is the first knob to isolate.

Conflicts

#2693 touches the same script, the same config block, and also appends to perf-changelog.yaml, so the changelog append will collide and whichever lands second needs a rebase. The changes themselves are complementary.

Two notes for whoever reviews alongside #2693. That PR raises the TP4 concurrency ceiling to 64 via its HiCache arms — at CONC=64 the new min(2*CONC, 128) lands exactly on the 128 cap, which is the intended behaviour but worth knowing. And #2693's stated motivation is MI355X/B200 comparability: the B200 sibling still captures to min(CONC, 64), so this PR re-introduces a difference on that knob. It is harness tuning rather than a deployment-defining server arg, but it is a real difference and should not be discovered by surprise.


Note

Medium Risk
Benchmark harness and collective quantization changes affect published MI355X AgentX throughput/latency and may not be bit-identical to the prior unquantized all-reduce path; accuracy is validated only via existing eval runs.

Overview
Updates qwen3.5-fp4-mi355x-sglang-agentic-mtp to SGLang ROCm v0.5.18 and retunes the AgentX launch script so MI355X MXFP4 serving matches sibling recipes and the published cookbook.

The benchmark script now routes multi-GPU collectives through INT8 ROCm quick all-reduce (ROCM_QUICK_REDUCE_QUANTIZATION), pins EAGLE draft-extend off unified AITER attention (SGLANG_AITER_UNIFIED_DRAFT_EXTEND=0), halves prefill limits to 16384 (--max-prefill-tokens / --chunked-prefill-size, aligned with the B200 arm), and sizes decode CUDA graphs to min(2×CONC, 128) so graph capture covers --max-running-requests at 2×CONC instead of undersized min(CONC, 64) eager fallbacks.

configs/amd-master.yaml only changes the container image; perf-changelog.yaml documents the retune for sweep invalidation.

Reviewed by Cursor Bugbot for commit 7ad5061. Bugbot is set up for automated code reviews on this repo. Configure here.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@github-actions

This comment was marked as outdated.

1 similar comment
@github-actions

Copy link
Copy Markdown
Contributor

…MXFP4 MI355X AgentX arm

Route the EAGLE draft-extend step through the AITER unified attention
kernel, and drop the CUDA graph sizing comment now that the changelog
entry carries the rationale.
…x-sglang-agentic-v0.5.18

# Conflicts:
#	perf-changelog.yaml
@github-actions

This comment was marked as outdated.

@github-actions

This comment was marked as outdated.

…x-sglang-agentic-v0.5.18

# Conflicts:
#	perf-changelog.yaml
@github-actions

Copy link
Copy Markdown
Contributor

@yichiche
yichiche force-pushed the amd/qwen3.5-fp4-mi355x-sglang-agentic-v0.5.18 branch from 74e8ab5 to 61258ba Compare August 27, 2026 00:21
@github-actions

Copy link
Copy Markdown
Contributor

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want higher recall? High effort reviews run extra passes and find more bugs. A team admin can switch effort levels in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 7ad5061. Configure here.

Comment thread perf-changelog.yaml
- "Route only enabled recipes through the immutable producer fork and preserve non-power launcher revisions."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2688

- config-keys:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changelog breaks append-only bytes

Medium Severity

The new perf-changelog.yaml entry deletes historical blank-line separator bytes between existing entries instead of leaving the prior file as an exact prefix. That file is byte-sensitive, so the mutation breaks the append-only prefix that sweep gating and reuse checks require.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 7ad5061. Configure here.

@github-actions

Copy link
Copy Markdown
Contributor

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant