Skip to content

perf: add low-latency nz quantized decode matmul. - #42

Open
maojunx99 wants to merge 3 commits into
xLLM-AI:mainfrom
maojunx99:perf/qwen3vl-decode-nofusion
Open

perf: add low-latency nz quantized decode matmul.#42
maojunx99 wants to merge 3 commits into
xLLM-AI:mainfrom
maojunx99:perf/qwen3vl-decode-nofusion

Conversation

@maojunx99

Copy link
Copy Markdown
Contributor

No description provided.

@maojunx99
maojunx99 force-pushed the perf/qwen3vl-decode-nofusion branch from 69b6760 to a1a6de0 Compare August 17, 2026 03:06
- use a 1C1V kernel type for decode-specialized shapes\n- process the 320-column gate epilogue on one vector core
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant