Skip to content

上传hc_pre_sinkhorn 和dequant_swiglu_quant的优化代码 - #47

Open
hanrui74 wants to merge 3 commits into
xLLM-AI:mainfrom
hanrui74:main
Open

上传hc_pre_sinkhorn 和dequant_swiglu_quant的优化代码#47
hanrui74 wants to merge 3 commits into
xLLM-AI:mainfrom
hanrui74:main

Conversation

@hanrui74

@hanrui74 hanrui74 commented Aug 21, 2026

Copy link
Copy Markdown

hc_pre_sinkhorn 性能数据

case bs M iters 优化前 µs 优化后 µs 加速比 路径
1 1 4 20 3.155 3.142 1.00x 小 batch
2 64 4 20 4.602 4.546 1.01x 小 batch
3 512 4 20 10.780 5.512 1.96x SoA
4 1024 4 20 16.730 5.407 3.09x SoA
5 4096 4 20 53.710 6.978 7.70x SoA
6 16384 4 20 204.479 12.884 15.87x SoA
7 1024 2 20 14.609 4.640 3.15x SoA
8 1024 3 20 16.120 5.323 3.03x SoA
9 1024 6 20 18.805 18.812 1.00x AoS
10 1024 8 20 20.854 20.786 1.00x AoS
11 1024 12 20 24.392 24.255 1.01x AoS
12 1024 16 20 29.433 29.423 1.00x AoS
13 1024 4 1 5.199 4.635 1.12x SoA
14 1024 4 5 6.498 4.521 1.44x SoA
15 1024 4 40 30.409 6.498 4.68x SoA
16 1024 4 20 17.030 5.760 2.96x SoA
17 1024 4 20 16.805 5.376 3.13x SoA
18 1 2 20 3.014 3.070 0.98x 小 batch
19 1024 4 20 16.809 5.524 3.04x SoA
20 1024 4 20 17.056 5.724 2.98x SoA
geomean 2.15x

hcpresinkhorn最终优化报告.html

dequant_swiglu_quant 性能数据

一、cann-bench 20 case(基线 → 最终优化 Round 4)

Case dtype rows H 竞品baseline (us) 基线 (us) 最终优化 (us) 基线加速比 最终加速比 提升倍数 代码路径
1 fp16 512 1024 8.84 17.04 5.74 0.52x 1.54x 3.0x row
2 fp16 1024 2048 17.23 37.83 12.54 0.46x 1.37x 3.0x row
3 fp16 2048 4096 47.87 102.01 39.08 0.47x 1.23x 2.6x row
4 bf16 4096 2048 47.39 156.04 41.92 0.30x 1.13x 3.7x row
5 bf16 127 512 8.44 5.99 4.49 1.41x 1.88x 1.3x row
6 fp16 8192 512 27.62 224.03 31.94 0.12x 0.86x 7.0x vec
7 bf16 1023 2049 17.55 42.18 19.37 0.42x 0.91x 2.2x row
8 bf16 255 4097 11.38 16.78 8.84 0.68x 1.29x 1.9x row
9 int32 512 1024 8.78 24.01 7.09 0.37x 1.24x 3.4x row
10 int32 1024 2048 19.63 53.58 17.76 0.37x 1.11x 3.0x row
11 int32 2048 4096 56.57 77.43 58.42 0.73x 0.97x 1.3x row
12 int32 4096 2048 55.94 112.23 65.09 0.50x 0.86x 1.7x row
13 int32 127 512 6.61 6.25 4.68 1.06x 1.41x 1.3x row
14 int32 8192 512 33.35 168.04 32.34 0.20x 1.03x 5.2x vec
15 int32 1023 2049 19.88 33.27 22.43 0.60x 0.89x 1.5x row
16 int32 255 4097 13.99 14.04 11.12 1.00x 1.26x 1.3x row
17 int32 10007 32 12.24 154.48 9.59 0.08x 1.28x 16.1x vec
18 int32 32768 128 33.87 539.47 36.67 0.06x 0.92x 14.7x vec
19 int32 4001 1024 32.39 94.19 37.04 0.34x 0.87x 2.5x row
20 int32 16384 256 32.95 309.84 32.54 0.11x 1.01x 9.5x vec

汇总统计

指标 基线 最终优化 (Round 4)
平均加速比 0.49x 1.15x
精度 20/20 20/20

二、xllm-ops 10 case(基线 → 最终优化 Round 1)

Case rows H 基线 (us) 优化后 (us) 加速比 瓶颈组
1 1 256 2.171 1.810 1.20x SCALAR-S
2 2 512 2.453 2.366 1.04x SCALAR-S
3 4 128 2.583 2.348 1.10x SCALAR-S
4 8 1024 2.856 2.835 1.01x SCALAR-S
5 16 768 3.393 3.280 1.03x SCALAR-S
6 32 256 4.063 3.431 1.18x SCALAR-S
7 128 512 6.231 5.193 1.20x SCALAR+icache
8 3 320 2.348 2.595 0.90x SCALAR-S
9 7 2048 3.183 2.927 1.09x SCALAR-S
10 64 1536 6.183 4.850 1.27x MTE2+SCALAR

几何平均加速比:1.098x


dequant_swiglu_quant最终优化报告.html

@hanrui74 hanrui74 changed the title 上传hc_pre_sinkhorn的优化代码 上传hc_pre_sinkhorn 和dequant_swiglu_quant的优化代码 Aug 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants