Skip to content

Latest commit

 

History

History
26 lines (24 loc) · 168 KB

File metadata and controls

26 lines (24 loc) · 168 KB

Query: By Kernel Type

Auto-generated. Do not edit manually.

Kernel Type Pages
attention FlashAttention-4 Blog, FlashMLA upstream README, Gated Delta Networks reference repository, NVIDIA Qwen3-Next Architecture Announcement, DeepSeek-V3.2-Exp in vLLM: Fine-Grained Sparse Attention in Action, FlashInfer MLSys 2026 - Track B: Sparse Attention, CUTLASS Changelog: SM100/Blackwell Entries, FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling, K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model, Native Sparse Attention (NSA), Sync nv_dev with upstream #316 (Mega MoE optimizations & benchmarks), [TRTLLM-10022][feat] Add hopper xqa decode support for skip softmax attention, [None][feat] Optimize super-v3 nvfp4 for better perf, [TRTLLM-11092][feat] add support for visual gen FA4 attention backend, [TRTLLM-11119][feat] Blackwell SageAttention, Integrate into AttentionOp API, [None][feat] Add fused DiT QK Norm + RoPE CUDA kernel for FLUX, [TRTLLM-11540][feat] Add EAGLE3 dynamic tree speculative decoding support, [TRTLLM-11289][feat] Integrate CuteDSL's bf16 dense GEMMs, [None][feat] Support update weight for nvfp4, [https://nvbugs/5983390][perf] Kernel fusions in _gather_k_cache_for_chunk of Indexer in DSA, [None][feat] Temporally-Correlated Heuristic-guided Indexer TopK for Sparse Attention, [None][feat] Support sparse mqa/gqa attention, [https://nvbugs/5983390][perf] Multiple host perf optimizations for DSA part, [None][feat] Trtllm-gen FMHA JIT support, [None][feat] Add triton paged attention for AutoDeploy, [TRTLLM-11485][feat] Feature rework: Add SageAttention refreshed kernels (attentionOp only), [#12784][feat] AutoDeploy: Optimize DeepSeek-R1 model performance, [#12716][feat] Fused cross-head QK Norm + RoPE kernel for WAN, [TRTLLM-34871][feat] Add cute dsl FP8 paged MQA logits decode kernel, [None][feat] Integrate FP4 indexer for DSA on Blackwell, [None][perf] Scheme X L2-aware dispatcher and PDL launchers for sparse-attention GVR Top-K, [None][feat] Fuse FP8 1x128 quantize + UE8M0 scale pack on SM100, [#13580][fix] AutoDeploy: Support Gemma3n/4 E2B variants, [None][feat] Add DeepSeekV4 attention kernels, [None][perf] mHC fused_hc kernel optimizations + DS-V4 entry-boundary RMSNorm fold-in, [TRTLLM-35237][feat] Add cute dsl FP4 paged MQA logits decode kernel, [None][feat] Keep DSv4 o_a_proj as FP8, and port vLLM's fused_inv_rope_fp8_quant, [None][feat] Add chunked prefill support for Gemma4 (text + vision multimodal), [None][feat] DSv4: enable GVR Heuristic Top-K for compress_ratio=4, [None][feat] Update the logic of FMHA JIT path, [None][feat] GPT-OSS Sm120/Sm121 Support, [TRTLLM-8535][feat] Support DeepSeek V3.2 with FP8 + BF16 KV cache/NVFP4 + BF16 KV cache, [None][fix] Fix the performance issue of FP8 blockwise grouped GEMM when using attention DP, [None][feat] Fused kernels (qknormrope + moe routing) and two-model MTP support for glm4moe, [ex77] fix mla split; add fwd lse; add bwd varlen, Example 77 add blackwell fmha bwd for MLA shape, Add Blackwell MLA forward (shape: d=192, dv=128) implementation, fix gqa issue for blackwell fmha.py, Fp8 kernel with "in-kernel" transpose of V in producer, Add local attention in Hopper FAv3, Paged Attention support for FA3, FA3 paged attention: Readiness for Cutlass 3.6 / default value for block_table, Fix FA3 Varlen Performance regression, feat: Adding varlen support to cute-dsl sm80 bwd, [Cute,Fwd,Sm100] Implement SplitKV, [Cute] Block sparse support Sm100, [Cute,Fwd,Sm100] Support paged attention, Add blocksparse support for bwd on blackwell, [Cute,Fwd,Sm100] distributed offset calculation for paged KV, [Cute,Fwd,Sm100] fp8 e4m3 and e5m2 support, [Cute,Fwd,Sm100] support irregular qhead / kvhead ratios, Add SM120 varlen attention support, [Fwd,Sm90] Add paged KV attention support (tma and cp.async), Feat([FA4][CUTE DSL]) Add head_dim=256 support (forward + backward), [hd256] Improve forward kernel with exp2 FMA emulation (3% to 9% performance gain), feat: update decode attention APIs, misc: fix instrument code for mla profiler, add multi-item scoring, fix: add zero init for KV tiled copy, feat: add functional per-head FP8 quantization for FA3, [nvidia] initial support for blackwell kernels, [nvidia] Add Blackwell FMHA decode kernel from TRT-LLM, Fix KV chunking for POD. , bugfix: temporally disable split-kv in blackwell mla, Parameterize prefix mask call (needed by POD-Attention), bugfix: adding lse output to blackwell fmha kernels, bugfix: follow user-specified sm_scale for blackwell cutlass fmha, bugfix: fix fp8 attention kernels aot compilation issue, bugfix: host-precomuted plan function for blackwell fmha, hotfix: fix the blackwell fmha stream, [Feature] Support PDL for batch Prefill and Decode, [feat] add unified batch attention w/ correctness tests., Fix FA2 and FA3 multi-item scoring and cuda illegal memory access error, Add more logging to TRTLLM-GEN debug trace (NFC), update trtllm-gen decode attention kernel launcher, bugfix: fix blackwell fmha hanging issue for empty kv_len, [feat] optimize persistent batch attention perf., [fix] fix BatchAttention CTA_TILE_KV mask issue, feat: add trtllm-gen mla cubin, Fix missing hash in the cudnn cubin path, Add trtllm-gen attention mha kernel with FP8 Q/K/V and FP8 output, Bug fix: fix duplicate launch in POD, refactor: refactor trtllm-gen attention kernel integration code, [fix] fix integer overflow in FA2 customized_mask & add buffer overflow warning., refactor: Improved metainfo for trtllm-gen fmha, Fix the bug of the kernel-selection heuristic in trtllm-gen, feat: Add k_scale and v_scale to persistent attention , feat: Support logits_soft_cap for Persistent attn; fix kv split limit, support trtllm-gen prefill fp4 output, GPT-OSS Support: Add Blackwell MoE mxfp4 implementation from TRTLLM and Attention Sink, Remove getEnvEnablePDL in favor of enable_pdl parameter, feat: Support fp8 qkv, fp16/bf16 out MHA for trtllm-gen., feat: integrate xqa attention backend, backend: Refactor trtllm-gen fmha metainfo loading, bugfix: Fix Persistent kernel precision for masked output , feat: Integrate TRTLLM varlen kernel for deepseek R1 prefill , bugfix: fix persistent attention kernel correctness on blackwell, Backend: downgrade trtllm-gen kernel to cuda-12, feat: initial support for SM103, SM110, SM120, SM121, bugfix: fix merge_attention_state in BatchAttention w/ gqa-group-size in Qwen family, bugfix: trtllm-gen fmha sm101 and sm100 compatibility, perf&bugfix: skip kv-tile computation out of sliding window in FA2; fix __syncthreads in mergestate, TGV GEMM as a BF16 backend alternative to cuBLAS, feat: Add variant.OutputTransform() to decode kernels, feat: Batch-size invariant FA2 Prefill & Decode, perf: Port the separate reduce kernel mode from trtllm., feat: add xqa fp8 mha and fp8 kv cache, Bugfix: Fix data hazard in persistent reduce, Add head_dim=64 for blackwell cutlass fmha implementation, Bugfix: fix o_strides in persistent kernel , Tune kernel compilation parameters for https://github.com/flashinfer-ai/flashinfer/pull/1850 , MLA RoPE + quantization fused kernel: shape generalization for MHA / GQA, feat: add xqa backend and completes NHD/HND coverage for trtllm-gen/xqa backend, feat: Add flashinfer.rope.rope_quantize_fp8_append_paged_kv_cache (fused RoPE + Q + KV cache, supports MLA/GQA/MHA) , Rebase FP8 SM100 Cutlass FMHA Attention to main (original PR#1238), Fix: several bugs/issues with trtllm-gen attention kernels. , [Feature] Support batch prefill for POD Attention, enable xqa speculative decoding, add tensor scale input for xqa, refactor: update fa3 codebase and fix hopper unittest [part 1], feature: make the LSE returned by MLA support base 2 or e #2113, perf: bunch of features and optimizations for top-k (sampling + sparse attention), fix flaky xqa test, feat: add trtllm-gen per-tensor sparseMla kernels., feat: TRTLLM FMHAv2 backend for ctx attention, [feat] Integrate SGLang concat_mla_k kernel into flashinfer, [TRTLLM-Gen Fmha] add optimized trtllm-gen decode kernels for high throughput + speculative decoding, feat: add GDN Attention, feat: [Qwen3-Next] Add Cute DSL GDN decode kernel and tests, fix: ensure each CTA processes full numHeadsQPerKv for trtllm decode kernel, feat: Add TRTLLM fmha_v2 library for SM90 attention with Skip-Softmax , feat: Add TRTLLM-Gen Skip-Softmax kernels for prefill and decode, Ameyn/gdn decode cutedsl kernel, Feat/gdn decode pooled, Add NVFP4 KV cache quantization support for SM100, Mamba2 SSD Combined Forward Pass (Blackwell CuTe DSL Kernel), feat: Add DiT-oriented kernels where Qk (Bmm1) type can be reinterpreted into Int8 or BFloat16, Add cute dsl mla decode op, feat: FP8 output support for CUTLASS MLA paged attention, feat: Support padding tokens with seqlen=0 for rope+quant+kv cache update fusion kernel, [CuTe DSL] Add modular FMHA prefill and MLA decode attention kernels, [Fmha] Sparse MLA decode kernel selection heuristics, [Fmha] support nvfp4 output keepsMmaAb generation kernels, [feat] Add blackwell GDN prefill kernel, Support NVFP4 KV for prefill and batch attention kernels, feat: Enable FP8 (E4M3/E5M2) in concat_mla_k for optimize long-context prefill performance and refactor type dispatch for BF16/FP16, feat: DiT layer norm fusions for WAN: flashinfer.diffusion_ops, cute-dsl fmha prefill (cubin integration): remove front-padding, add attention_sink, and pdl support, Support Kimi K2.5 H64 CuTe DSL MLA decode, Add dynamic tokens-per-page TRTLLM-GEN GQA kernels, fix(fmha_v2): fix FP8 V-scratch pipeline and varlen scheduler on SM90, checkpointing_ssu kernel: fused replay + conditional state-write for Mamba2, perf: fix the iteration bound of SWA in FA2 prefill template, Align KV chunk size binary search with actual KV chunk splitting., feat: support deepseek prefill attention shape, perf: refactor fa2 prefill template, bugfix: MLA decode should multiply sm_scale by math::log2e, fix rope logic in mla decoding, feat: support f32 attention output in FA2 template, feat: apply sm_scale at logits instead of q in FA2 template, perf: memory efficient deepseek mla fused page-attention kernel, bugfix: mla page-attention kernel for different page sizes, feat: unlocking MLA for A100, feat: unlock MLA attention for sm89 (L40/L40s/4090), bugfix: bugfix on sm89 MLA, perf: MLA decode kernel implemented by CuTe targeted to SM80, Add POD-Attention to FlashInfer, perf: dynamic split-k for MLA, bugfix: fix the behavior of MLA kernel when kv-length is 0, Naive Support for Hopper FP8 Prefill Kernel with Per-Head Quantization, perf: FlashAttention-3 style MLA PageAttention, perf: fix MLA split-k performance bug, perf: tweak the pipeline design of mla kernel, feat: flashinfer intra-kernel profiler, bugfix: fix potential issues of FA3 template loading nans for PageAttention, perf: Use 2WG pipeline design for MLA implementation on Hopper, perf: prefetch page indices for mla kernel, 3rdparty: upgrade cutlass to 3.9, Disable kernel cutlass_mla_decode on SM103, feat: Add FP4 (E2M1) KV Cache Support with Quantization Utilities for MLA, [sgl-kernel] Optimize concat_mla_k kernel, [DeepseekV32] Enable flashmla_prefill kernel with fp8 kvcache, Use trtllm_mla decode kernel for draft extend in speculative decoding, support cutlass fp4 kernel in sm120, [Fix] concat_mla_absorb_q_kernel fails for long inputs, [DeepSeek v3.2] opt Context Parallelism: support fused moe, multi batch and fp8 kvcache, [bug fix] fix ima with get_mla_kv_buffer_kernel overflow, [Fix] add block size logic for sm120 smem size, [sgl-kernel][1/2] Fused qk_norm_rope for GLM4.6, [NVIDIA] upstream FA4, Optimize FP8 MLA KV cache writes with Triton kernel, Add SwapAB Optimization for triton fused_moe_kernel on SM90., [diffusion] model: support TurboWan2.1-T2V-1.3B/14B SLA, optimize get_topk_ragged by fusing get k and k_scale triton kernel, Move fa4 from sgl-kernel to jit kernel, Kernel: optimize decoding metadata in NSA multi-spec backend with fused kernels, feat: add FA4 SM90 paged KV decode support & update attention docs, Tilelang sparse decode fwd for dsv32 mi355, [diffusion] Diffusion norm fusion for z-image, [DeepSeek-V3.2][JIT-kernel] Support nsa fuse store indexer k cache, [diffusion][llm] macOS support, [AMD] Tilelang sparse fwd for dsv32 mi355/mi300, [JIT Kernel] Reland NVFP4 kernels to JIT, Support Triton MLA FP8 KV cache, [GDN] Fuse GDN kkt + solve_tril into one kernel, [Diffusion] Add qknorm rope fuse kernel, [AMD] Enable FP8 KV cache and FP8 attention kernel for NSA on MI300/MI355 with TileLang backend, Fused_qknorm_rope kernel optimization: up to 2.4× faster, [XPU] Enable qwen3.5 on XPU, [nvidia] Gemma4 nvfp4 fix, [feat] Init true on policy with qwen_dense, [KDA] Optimize prefill kernels with diagonal and recompute fuse, [Gemma4] Optimize Gemm4 with fused Q/K/V RMSNorm + per-expert FP8 ckpt loader, Amd/deepseek v4 rebase main 0509, [rebase]Deepseek_v4 support w4(mxfp4)a16 on hopper, [Intel GPU] Enable DeepSeek V4 Inference on XPU, [fp8] SM90 swap-AB scaled_mm dispatch (~1.16x kernel geomean, +5.8-18.5% end-to-end), amd/deepseek_v4 27/N [fix] Reduce Triton autotune configs for faster first-time server launch, [Refactor] Rename NSA → DSA: user-facing aliases, file/class/import rename, support cmake for sgl-kernel, Feat/support encoder model (like bert), Support MHA with chunked prefix cache for DeepSeek chunked prefill, Blackwell Cutlass MLA kernel, fix: solve cu118 issue for cutlass mla, [perf] introduce deep gemm group_gemm_masked as bmm, [2/2] Add python wrapper for CUTLASS FP8 Blockscale MoE Kernel. , Cutlass MLA: Disable split kv due to https://github.com/NVIDIA/cutlass/issues/2274, Fix bug of deepseek-v3 under DP+EP mode with large batchsize/seqlen, feat: integrate deepgemm into EPMoE, [perf][sgl-kernel] extend cutlass_mla_decode to support num_head < 128, Tiny fix cutlass_mla_get_workspace_size stub incorrect signature, feat: support DeepSeek-R1-W4AFP8 model with ep-moe mode, [Feature] Support cp.reduce.async.bulk.tensor, Add swizzle layout detection and automatic merging for layout conflicts, [CUDA] Support tcgen5mma gemm ts, [TIR][IR] Update to use tirx, [ROCm] Faster Custom Paged Attention kernels, [Attention] MLA decode optimizations, [Attention] MLA with chunked prefill, [Perf] Mem align KV caches for CUDA devices (MLA perf improvement), [Kernel] Make rotary_embedding ops more flexible with input shape, [core] Perf improvement for DSv3 on AMD GPUs, [Kernel] allow non-contiguous input for marlin kernel, [ROCM][KERNEL] Paged attention for V1, [NVIDIA] Support Cutlass MLA for Blackwell GPUs, [ROCM] Add gfx950 to the custom attention archs, [Kernel] support merge_attn_states CUDA kernel, 3x speedup, Allocate kv_cache with stride order, [Kernel] Unified Triton kernel that doesn't distinguish between prefill + decode, [ROCm][Kernel][V1] Enable AMD Radeon GPU Custom Paged Attention on v1, [ROCm][FP8][Kernel] FP8 quantization fused into Custom Paged Attention, [Hardware][AMD] integrate aiter chunked prefill into vllm, [BugFix] FA2 MLA Accuracy Issue, Enable V1 for Hybrid SSM/Attention Models, [Bugfix] Fix some narrowing conversion warnings, [Kernel] Optimize Prefill Attention in Unified Triton Attention Kernel, SM100 Cutlass MLA decode with unrestricted num_heads (< 128) for DeepSeek TP, [Kernel] Enable Hybrid Model Support in Triton Unified Attention Kernel, [v1] - Mamba1 Attention Metadata, Fp8 paged attention update, [BugFix] Fix triton compile error in kernel_unified_attention_2/3d caused by attention sinks, Optimize input preparation for FlashInfer [2/N], [Compile] Fix Compile Warning SM100 Cutlass MLA, [Bugfix] Fixing division by zero in triton_attn if query_heads/kv_heads > 16 , [Kernel] Support decode context parallelism on Blackwell with CUTLASS MLA, [Attention] Use sparse prefill kernel for fp8 kv-cache in DeepSeek-v3.2, [Model] Add support for openPangu moe model, bugfix: correct attn output with base 2 or e, Add llmcompressor fp8 kv-cache quant (per-tensor and per-attn_head), OffloadingConnector: Support kernel_block_size != block_size, [Bugfix] [Kernel] Triton attention kernels: mask out V blocks that fall outside sliding window, [Bugfix] Fix incorrect tiles creation for mm prefix triton attention, [Bugfix][ROCm]Fix Qwen3-Next-80B-A3B-Thinking inference and optimize non-standard block size (544) support under rocm_atten, [Spec Decode] Unified Parallel Drafting, [PERF] Change GDN Attention State Layout from [N, HV, K, V] to [N, HV, V, K], Triton MLA perf fixes, [Quantization] add humming quantization kernel, [Kernel] Add FP8 KV cache support to Triton MLA decode attention, [Attention][Perf][Kernel] Replace torch.cat with vectorized CUDA kernel MLA query concat - DeepSeek-V3.2, [BUGFIX][Mamba][Qwen3.5] Zero freed SSM cache blocks on GPU, [Attention][Perf] Optimize cp_gather_and_upconvert_fp8_kv_cache - DeepSeek-v3.2, [Kernel] Add fused_sigmoid_gating_delta_rule_update kernel for Qwen3 Next, [Kernel] Fuse FP8 output quantization into merge_attn_states, [Feat][Spec Decode] DFlash, [Perf] FP8 FlashInfer Attn for ViT, [Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity, [Refactor] Improve indexer decode path metadata preparation, [Bugfix] Fix broken explicit unquantized kv cache dtype support, [Perf][GDN] Align TMA usage with upstream FLA, [MLA] Optimize mla indexer prepare uniform decode for MTP > 1, [Performance][DSR1]: Fused RoPE+KVCache+q_concat for MLA, [Perf] Batch invariance with Cutlass fp8 support, 28.9% E2E latency improvement, [DSv4] Improved fused Indexer Q quant kernel, [feat] Add FP8 per-tensor Q scale support to Triton attention backend, [DSv4] Improved dequant gather K cache kernel, [6/n] Migrate activation kernels, gptq, gguf, non cutlass w8a8 to libtorch stable ABI (continued), [Perf] Add do_not_specialize in fused FP8 RoPE kernel, add cutedsl dsv4 indexer fp8 kernel, FlashAttention-4, FlashAttention SM100 MLA TopK Sparse Forward, FlashMLA attention kernels, Native Sparse Attention (NSA), Sparse MLA, TensorRT-LLM Blackwell FP4 DSA Indexer
batched-gemv Twelve Attempts at an FP4 Kernel, NVFP4 GEMV, Blackwell NVFP4 Kernel Hackathon Journey, GPU Mode NVFP4 Hackathon - Problem 1: Batched GEMV, NVFP4 batched GEMV
decode FlashMLA upstream README, DeepSeek-V3.2-Exp in vLLM: Fine-Grained Sparse Attention in Action, FlashInfer MLSys 2026 - Track C: Gated Delta Net, FlashMLA attention kernels, Sparse MLA
flash-attention FlashAttention-4 Blog, FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling, [None][feat] Optimize super-v3 nvfp4 for better perf, [TRTLLM-11092][feat] add support for visual gen FA4 attention backend, [TRTLLM-11119][feat] Blackwell SageAttention, Integrate into AttentionOp API, [TRTLLM-11540][feat] Add EAGLE3 dynamic tree speculative decoding support, [None][feat] Support sparse mqa/gqa attention, [None][feat] Trtllm-gen FMHA JIT support, [None][feat] Add DeepSeekV4 attention kernels, [None][feat] Update the logic of FMHA JIT path, [None][feat] GPT-OSS Sm120/Sm121 Support, Flash MLA support, [ex77] fix mla split; add fwd lse; add bwd varlen, Example 77 add blackwell fmha bwd for MLA shape, Add Blackwell MLA forward (shape: d=192, dv=128) implementation, fix gqa issue for blackwell fmha.py, Paged Attention support for FA3, FA3 paged attention: Readiness for Cutlass 3.6 / default value for block_table, Fix FA3 Varlen Performance regression, feat: Adding varlen support to cute-dsl sm80 bwd, Blackwell FlashAttention-BWD (v1.0), [Cute] Block sparse support Sm100, [Cute,Fwd,Sm100] Support paged attention, [Cute,Fwd,Sm100] fp8 e4m3 and e5m2 support, [Cute,Fwd,Sm100] support irregular qhead / kvhead ratios, Feat([FA4][CUTE DSL]) Add head_dim=256 support (forward + backward), [hd256] Improve forward kernel with exp2 FMA emulation (3% to 9% performance gain), [nvidia] initial support for blackwell kernels, [nvidia] Add Blackwell FMHA decode kernel from TRT-LLM, bugfix: adding lse output to blackwell fmha kernels, bugfix: follow user-specified sm_scale for blackwell cutlass fmha, bugfix: host-precomuted plan function for blackwell fmha, hotfix: fix the blackwell fmha stream, Add more logging to TRTLLM-GEN debug trace (NFC), update trtllm-gen decode attention kernel launcher, bugfix: fix blackwell fmha hanging issue for empty kv_len, feat: add trtllm-gen mla cubin, Fix missing hash in the cudnn cubin path, Add trtllm-gen attention mha kernel with FP8 Q/K/V and FP8 output, refactor: refactor trtllm-gen attention kernel integration code, refactor: Improved metainfo for trtllm-gen fmha, Fix the bug of the kernel-selection heuristic in trtllm-gen, support trtllm-gen prefill fp4 output, Remove getEnvEnablePDL in favor of enable_pdl parameter, feat: Support fp8 qkv, fp16/bf16 out MHA for trtllm-gen., backend: Refactor trtllm-gen fmha metainfo loading, feat: Integrate TRTLLM varlen kernel for deepseek R1 prefill , Backend: downgrade trtllm-gen kernel to cuda-12, feat: initial support for SM103, SM110, SM120, SM121, bugfix: trtllm-gen fmha sm101 and sm100 compatibility, perf: Port the separate reduce kernel mode from trtllm., Add head_dim=64 for blackwell cutlass fmha implementation, Rebase FP8 SM100 Cutlass FMHA Attention to main (original PR#1238), Fix: several bugs/issues with trtllm-gen attention kernels. , feat: add trtllm-gen per-tensor sparseMla kernels., feat: TRTLLM FMHAv2 backend for ctx attention, [TRTLLM-Gen Fmha] add optimized trtllm-gen decode kernels for high throughput + speculative decoding, fix: ensure each CTA processes full numHeadsQPerKv for trtllm decode kernel, feat: Add TRTLLM fmha_v2 library for SM90 attention with Skip-Softmax , feat: Add TRTLLM-Gen Skip-Softmax kernels for prefill and decode, Add NVFP4 KV cache quantization support for SM100, feat: Add DiT-oriented kernels where Qk (Bmm1) type can be reinterpreted into Int8 or BFloat16, [CuTe DSL] Add modular FMHA prefill and MLA decode attention kernels, [Fmha] Sparse MLA decode kernel selection heuristics, [Fmha] support nvfp4 output keepsMmaAb generation kernels, feat: Enable FP8 (E4M3/E5M2) in concat_mla_k for optimize long-context prefill performance and refactor type dispatch for BF16/FP16, cute-dsl fmha prefill (cubin integration): remove front-padding, add attention_sink, and pdl support, Add dynamic tokens-per-page TRTLLM-GEN GQA kernels, fix(fmha_v2): fix FP8 V-scratch pipeline and varlen scheduler on SM90, Naive Support for Hopper FP8 Prefill Kernel with Per-Head Quantization, perf: FlashAttention-3 style MLA PageAttention, [NVIDIA] upstream FA4, Move fa4 from sgl-kernel to jit kernel, feat: add FA4 SM90 paged KV decode support & update attention docs, [SGLang-Diffusion] Fix custom op fake impl missing eps default for torch.compile, Amd/deepseek v4 rebase main 0509, Support MHA with chunked prefix cache for DeepSeek chunked prefill, Blackwell Cutlass MLA kernel, [Feature] Support cp.reduce.async.bulk.tensor, Add swizzle layout detection and automatic merging for layout conflicts, [CUDA] Support tcgen5mma gemm ts, [NVIDIA] Support Cutlass MLA for Blackwell GPUs, [Kernel] Unified Triton kernel that doesn't distinguish between prefill + decode, [Kernel] Optimize Prefill Attention in Unified Triton Attention Kernel, SM100 Cutlass MLA decode with unrestricted num_heads (< 128) for DeepSeek TP, [Kernel] Enable Hybrid Model Support in Triton Unified Attention Kernel, [Compile] Fix Compile Warning SM100 Cutlass MLA, [Bugfix] Fixing division by zero in triton_attn if query_heads/kv_heads > 16 , Add llmcompressor fp8 kv-cache quant (per-tensor and per-attn_head), [BUGFIX][Mamba][Qwen3.5] Zero freed SSM cache blocks on GPU, [Feat][Spec Decode] DFlash, [Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity, FlashAttention-4, FlashAttention SM100 MLA TopK Sparse Forward
fused-kernel FlashInfer MLSys 2026 - Track A: Fused MoE FP8, Fused MoE — Expert GEMM and Adjacent Operations, Gated Dual GEMM (Gate-Up + Activation)
gated-delta-net Gated Delta Networks reference repository, NVIDIA Qwen3-Next Architecture Announcement, FlashInfer MLSys 2026 - Track C: Gated Delta Net, Gated Delta Network kernels
gated-dual-gemm GPU Mode NVFP4 Hackathon - Problem 3: Gated Dual GEMM, Fused MoE — Expert GEMM and Adjacent Operations, Gated Dual GEMM (Gate-Up + Activation)
gemm Colfax Article Source Kernels, Colfax CUTLASS Kernels, DeepGEMM tensor-core kernel library, Anatomy of a Reward Hack, Writing High-Performance Matrix Multiplication Kernels for Blackwell with JAX Pallas, Modular: Matrix Multiplication on Blackwell, Part 3, NVIDIA Developer Code Samples, simveit effective_transpose, simveit load_and_store, Tilus: A Tile-Level GPGPU Programming Language for Low-Precision Computation, GPU Mode NVFP4 Hackathon - Problem 2: NVFP4 GEMM, GPU Mode NVFP4 Hackathon - Problem 3: Gated Dual GEMM, GPU Mode NVFP4 Hackathon - Problem 4: Grouped GEMM, Microbenchmarking NVIDIA's Blackwell Architecture, cuTile Python Documentation, CUTLASS Changelog: SM100/Blackwell Entries, CUTLASS Blackwell Cluster Launch Control, CUTLASS CuTe DSL Documentation, K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model, Fix performance issue of m-grouped contiguous GEMMs., Fix multicast bug and optimize masked GEMM, [Public release 26/04] Introducing Mega MoE, FP4 Indexer and other features/fixes, Support TMA multicast on B with m_grouped_gemm_contiguous., [None][feat] CuteDSL MOE FC1 Enhancement, [TRTLLM-9457][feat] Add cute dsl fp8 gemm for Blackwell, [None][feat] sm100 weight-only kernel, [TRTLLM-9831][perf] Enable 2CTA with autotune for CuteDSL MoE and Grouped GEMM optimizations, [None] [feat] Add densegemm backend for MoE, [TRTLLM-9831][perf] Use TMA.RED to improve effective memory bandwidth, [None][feat] Optimize super-v3 nvfp4 for better perf, [TRTLLM-11092][feat] add support for visual gen FA4 attention backend, [TRTLLM-10990][feat] Fuse SwiGLU and quant into shared expert, [TRTLLM-11289][feat] Integrate CuteDSL's bf16 dense GEMMs, [None][feat] CuteDSL MOE: Add raster along M/N support for blockscaled contiguous backbone kernel, [None][feat] Add DWDP (Distributed Weight Data Parallelism) support for MoE inference, [None][feat] Add PDL support to CuTE DSL top-k kernels, [TRTLLM-11585][feat] Add CUTEDSL moe backend for nemotron-h, [None][feat] Add FP4 residual quantization kernel without channel reo…, [TRTLLM-34871][feat] Add cute dsl FP8 paged MQA logits decode kernel, [None][feat] Fuse FP8 1x128 quantize + UE8M0 scale pack on SM100, [https://nvbugs/6108841][fix] add hidden_dim=6144 router GEMM instantiation for GLM-5, [None][perf] FC2 DenseGEMM autotune: split-K, swap_ab, fine-grained tuning buckets, [None][feat] Keep DSv4 o_a_proj as FP8, and port vLLM's fused_inv_rope_fp8_quant, [None][feat] Update the logic of FMHA JIT path, feat: Add w4a8_mxfp4_fp8 quantization recipe., [OMNIML-2336][feat] Add NVFP4 x FP8, [None][feat] GPT-OSS Sm120/Sm121 Support, [None][fix] Fix the performance issue of FP8 blockwise grouped GEMM when using attention DP, [None][feat] Enable nvfp4 cuda core for sm120, [TRTLLM-9685] [feat] Add gather fc1 kernel by cuteDSL, [https://nvbugs/5726962][feat] Apply fusion for W4AFP8_AWQ MoE, Improve sm90 mixed dtype kernel, [EVT] Add support for Row/Col broadcast PtrArray, Groupwise scaling along M for FP8 gemm, Improvements for: Groupwise scaling along M for FP8 gemm, Hopper Grouped GEMM support for FP8 Accum, Flash MLA support, Blockwise and Groupwise GEMM for Blackwell and Improvements for Hopper, Blockwise Improvement and Programmatic Dependent Launch, hopper-blockwise-generalization-optimization, Fix epilogue::thread::Convert cannot be used with DefaultEpilogue, Add Blackwell MLA forward (shape: d=192, dv=128) implementation, DistGEMM bug fixes, Support PDL for SM90 Array TMA GEMM, Support for GEMM-K=0 for Blackwell Grouped GEMMs, Add tutorial fp16_gemm_1, Blockscaled Ragged Contiguous Grouped Gemm for MoEs, [Bug Fix]Bypass launch grids for SM120 Kernel with SM90 Mainloop & SM100 TileScheduler, new example with TMA prefetch feature targeting for DRAM latency boun…, [Bug Fix]Set NumSplitsM to 1 when TileShapeM < 128 in sm90 fp8 blockwise scaling CollectiveMma, [CuTeDSL] Fix: SM100 block-scale gemm overlapping accumulator, Replace std::min with cute::min in sm120 blockwise scaling device functions, [Hopper CuTeDSL] Add grouped GEMM kernel example, Support for Group GEMM in CUTLASS Profiler for GeForce and Spark, [CLI] add cutedsl fp16 gemm tutorial from 2 to 6, Update blackwell tutorial to be compatible with 4.5-dev version, Small Tile N BlockScaled GEMM + Grouped GEMM on SM12x, Fp8 kernel with "in-kernel" transpose of V in producer, FA3 FP8 qkv descales + restore max offset for h128 causal + added sync for producer WG, FA3 kvcache + split kv + gqa parallelization, Paged Attention support for FA3, FA3 paged attention: Readiness for Cutlass 3.6 / default value for block_table, [Cute,Fwd,Sm100] Implement SplitKV, Blackwell FlashAttention-BWD (v1.0), Feat([FA4][CUTE DSL]) Add head_dim=256 support (forward + backward), feat: add functional per-head FP8 quantization for FA3, perf: accelerate blackwell grouped gemm, Add CUTLASS fused moe kernels from TensorRT-LLM., bugfix: Fix test and output shape of fp4 quantize, feat: Support MXFP8 x MXFP4 CUTLASS grouped GEMM, Reduce the JIT compilation time of gen_gemm_sm100_module, Update cutlass fp4 moe kernels, add cutlass backend for mm_fp4, Add blockwise-scaled FP8 GEMM via TRTLLM-Gen., feat: masked layout fp4 gemm using cute-dsl, GPT-OSS Support: Add Blackwell MoE mxfp4 implementation from TRTLLM and Attention Sink, feature: add cutlass as bmm_fp8 backend., Remove getEnvEnablePDL in favor of enable_pdl parameter, Add python API for masked grouped gemm, fix: update cutedsl masked moe gemm, feat: scaling at fp4 gemm epilogue, feat: integrate xqa attention backend, bugfix: Fix stream handling in cutedsl gemm, refactor fp4 masked gemm cute-dsl implementation and add manual cache, feat: cutlass fp4 gemm bringup for SM120 & SM121, feat: cutlass fp8 gemm bringup for SM120 & SM121, TGV GEMM as a BF16 backend alternative to cuBLAS, Support output signals for overlapping for cutedsl gemm, Update TGV GEMM default kernel and TGV code cleanup., Support Kimi-K2 for TRT: templatize number of experts, TVM: support TVM binding for GroupedGemm, feat: add xqa fp8 mha and fp8 kv cache, feat:enable fp8 blockscale moe for fused cultass for sm90, feat: trtrllm-gen global scaled FP8 GEMMs, feat: Add FP4 TRTLLM-Gen throughput MOE batched gemms, silu_and_mul nvfp4 quanization fusion rework, Feature: Support Relu2 activation in fused MoE, Feature: Add support for L40 FusedMoE in cutlass path, [DSV3] Optimized Router Gemm, update trtllm cutlass moe , feat: BF16 GEMM using CUTLASS backend for SM100, make DeepGEMM swapAB available for linear gemm SM90, feat: MxInt4 x Bf16 TRT-LLM Gen MoE support, feat: Fused RMSNorm + FP4 Quantization Kernels in CuTe-DSL, Remove cudaStreamSynchronize from gemm_groupwise_sm120.cuh for CUDA graph compatibility, feat: add GDN Attention, [Perf][Feature] Add SM103-specific schedulers for NVFP4 CUTLASS kernels, feat: Support Fused MoE non gated Relu2 NVFP4 & FP8 and support Nemotron, [ML3] Optimized Router Gemm, [perf] Improve gemm_fp8_nt_groupwise (cutlass backend) by 10-40% for batch sizes <= 32, MTP for mamba , perf: add fp4 GEMM tile configs and streamK scheduler for SM120, feat: Add MXFP8 GEMM mm_mxfp8 (cutlass), refactor: Port upstream CUTLASS fixes and refactor grouped_gemm_nt_masked GEMM module location, feat: cute dsl mmfp4 for blackwell, Implement cutlass_fused_moe mxfp8, feat: trtllm tinygemm2 in flashinfer as bf16 routergemm, fix: add SM121 support to SM120 version guards, [feat] trtllm-gen mxfp8 gemm, feat: support mxfp4 & mxfp8 entrypoint for blackwell cutedsl dense gemm, Support for MXFP4 and NVFP4 group GEMMs on GeForce and Spark, feat: FP8 output support for CUTLASS MLA paged attention, [CuTe DSL] Add modular FMHA prefill and MLA decode attention kernels, feat: Add CuTe-DSL backend for NVFP4 quantization, feat: add MXFP8 GEMM support for SM120, [NVIDIA] fix(jit): enable GDC for CUTLASS fused MoE PDL — prevent random crashes on SM12x, feat: Add cuBLASLt backend for mm_bf16 and enable multi-tactic autotuning for FP8/MXFP8 runners, feat: Add CuTe DSL grouped-gemm + combine fusion support, perf: Optimize CUTLASS MoE helper kernels for small-batch decode workloads, Prevent MoE autotuner buffer overflow on large token buckets, perf: Port TRT-LLM SM120/SM121 FP4 CUTLASS GEMM optimizations. Add PDL, [feat] Trtllm-gen Per-token Nvfp4 MoE, feat: Add backend="b12x" for mm_fp4 on SM120, feat: Add b12x CuTe DSL fused MoE for SM120, perf: Add no-bias path for tinygemm_bf16, Integrate CUTLASS Small Tile N Blockscaled GEMMs/Grouped GEMMs for SM120 and SM121, feat: enable glm5 router gemm, fix(cute_dsl/moe): make autotuner bucket configuration adapt to runtime input, feat(moe): add SM120 W4A16 b12x kernels, checkpointing_ssu kernel: fused replay + conditional state-write for Mamba2, feat(cute_dsl/moe): add moe_output_memset_inplace dense memset wrapper, feat: support deepseek prefill attention shape, perf: MLA decode kernel implemented by CuTe targeted to SM80, Naive Support for Hopper FP8 Prefill Kernel with Per-Head Quantization, SM-constraint-GEMM by triton persistent kernel, Optimize nvfp4 block scaled gemm kernel when M is small., Update CUTLASS. Refine KernelSchedule for fp8 (grouped) gemm., Optimize cutlass int8 gemm kernel for large M on SM89 Ada GPU, [NVIDIA] Add new SMs support for Spark & Thor, [sgl-kernel][1/N]Support Expert Specialization Grouped GEMM, support cutlass fp4 kernel in sm120, [sgl-kernel][4/N]Support Expert Specialization Grouped GEMM, [NVIDIA] Fix CUDA arch requirement in nvfp4 cast, [sgl-kernel][Feat][B200][1/N]Support MXFP8 Grouped GEMM in Blackwell, [LoRA][III] Add LoRA support for MoE layers and enable TP, Add new moe wna16 marlin gemm, [sgl-kernel][Feat][B200][2/N] Support MXFP8 Grouped GEMM in Blackwell, [sgl-kernel][6/7]Support Expert Specialization Grouped GEMM, [JIT kernel] Apply jit per_tensor_quant_fp8 kernel, Move fa4 from sgl-kernel to jit kernel, Add mxfp8 support for online quantization, Triton dense linear, and CUTLASS MoE, Tilelang sparse decode fwd for dsv32 mi355, [Kernel Slimming] Migrate NVFP4 kernels to JIT, [Feature] NVFP4 Marlin fallback for non-Blackwell GPUs (SM75+), [AMD] Tilelang sparse fwd for dsv32 mi355/mi300, [JIT Kernel] Reland NVFP4 kernels to JIT, CUTLASS FP8 Blockwise GEMM improvement of SM120, CUTLASS NVFP4 GEMM improvement of SM120, [AMD] Enable FP8 KV cache and FP8 attention kernel for NSA on MI300/MI355 with TileLang backend, [Diffusion] Fix weight scale swizzle and add large-M kernel config for FLUX.2-dev-NVFP4, Enable PDL for various kernels in DSV32/GLM5, Port MXFP4 Marlin MoE support to JIT kernel path, Amd/deepseek v4 rebase main 0509, [rebase]Deepseek_v4 support w4(mxfp4)a16 on hopper, [Intel GPU] Enable DeepSeek V4 Inference on XPU, [fp8] SM90 swap-AB scaled_mm dispatch (~1.16x kernel geomean, +5.8-18.5% end-to-end), Support cutlass Int8 gemm, Support sm90 Int8 gemm, support w8a8 fp8 kernel with CUTLASS, add tensorrt_llm common and cutlass_extensions as 3rdparty, support blockwise fp8 matmul kernel, Support FP4 gemm (1/2), linear support deepgemm, Accelerate FP8 CUDA Kernel by 20-28%, fix per_token_group_quant_fp8 illegal memory when num_groups % 16 != 0, Support Blackwell Block Scale FP8 Gemm, Support fp8 gemm for blackwell, [Build] Fix cuda12.8 build error in nvfp4_scaled_mm_kernels.cu, Support MHA with chunked prefix cache for DeepSeek chunked prefill, [1/2] Add FP8 Blockscale MoE CUTLASS kernel for Blackwell, [perf] introduce deep gemm group_gemm_masked as bmm, [2/2] Add python wrapper for CUTLASS FP8 Blockscale MoE Kernel. , cutlass 3.9 supported to improve fp8_blockwise_gemm, chore: upgrade cutlass 3.9.2, [1/2] Add Kernel support for Cutlass based Fused FP4 MoE, Upgrade CUTLASS 4.0, Fix bug of deepseek-v3 under DP+EP mode with large batchsize/seqlen, feat: integrate deepgemm into EPMoE, Fix AWQ Dequant and Weight Loading of deepseek v2, Add a CUDA kernel for fusing mapping and weighted sum for MoE., Add CUTLASS FP8 Blockscale MoE kernel for Hopper architecture, Add dsv3 router gemm kernel, Add dsv3 fused a gemm to sgl-kernel, feat: support DeepSeek-R1-W4AFP8 model with ep-moe mode, [1/n]: add cutlass W4A8 moe kernel for hopper architecture, [feat] Support tp mode for DeepSeek-R1-W4AFP8, [Fix][Ready]Fix register spilling in cutlass nvfp4 gemm kernel on Blackwell, [sgl-kernel] Opt per_token_quant_fp8 with warp reduce, [Perf] Tunings for SM100 FP8 CUTLASS kernel, [NVIDA] [1/N] Nvfp4 Masked Gemm: Add quant op for the flashinfer grouped gemm, [fix]: fix cutlass moe ut and and Opt H20 cutlass groupGemm performance, [sgl-kernel] feat: Support sm120 cutlass fp8 gemm kernel, [NVIDIA] [2/N] Optimize silu_and_mul_scaled_fp4_grouped_quant perf, Update CUTLASS 4.2 & Enable K-Major Scale Factor for SM90 FP8 Blockwise Group GEMM, Make sm100 fp8 kernels available on sm103, Make fp4_quantize kernels work on sm103, CUTLASS fp8 blockwise gemm support of sm120, [WIP] support more dtypes for tcgen05, [Enhancement] add more dtype and fix mma.ws for fp16 for tcgen05, Add swizzle layout detection and automatic merging for layout conflicts, [CUDA] Support tcgen5mma gemm ts, [Feature] 2-SM support for TMA, TMEM and TCGEN5MMA on Blackwell, [Feature] Add Producer-Consumer Warp Specialization and T.tma_copy() API, [Feature] Block-scaled GEMM support for MXFP8 on Blackwell, [Feature] Support TMA store in T.tma_copy(), [Backend] Refactor gemm_sp, [CUDA] Support int4 T.gemm, [CUDA] Improve int4 GEMM lowering and packed codegen support, [codex] Split GEMM implementations by backend, [CUDA] Add native SM75 MMA GEMM support for FP16, INT8 and INT4, [Kernel]: Cutlass 2:4 Sparsity + FP8/Int8 Quant Support, [Kernel] Update cutlass_scaled_mm to support 2d group (blockwise) scaling, [ROCm] Faster Custom Paged Attention kernels, [NVIDIA] Support nvfp4 quantization, [Misc][Kernel]: Add GPTQAllSpark Quantization, [Kernel]Add streamK for block-quantized CUTLASS kernels, [Kernel] moe wna16 cuda kernel, [NVIDIA] Support nvfp4 cutlass gemm, add cutlass support for blackwell fp8 gemm, [Kernel] CUTLASS grouped gemm fp8 MoE kernel, Add cutlass support for blackwell fp8 blockwise gemm, permute/unpermute kernel for moe optimization, [NVIDIA] Support Cutlass MLA for Blackwell GPUs, [Hardware/NVIDIA/Kernel] [Functional Enablement] [1/N] Enable nvidia/DeepSeek-R1-FP4 Model, [NVIDIA] Support Cutlass w8a8 FP8 for Blackwell Geforce GPUs (sm120), Sm100 blockwise fp8 swap ab, [Perf] Tunings for SM100 FP8 CUTLASS kernel, [Kernel] Enable fp8 support for pplx and BatchedTritonExperts., [Hardware][NVIDIA][kernel] Fp4 MOE quant kernel optimization, [Perf] Further tunings for SM100 FP8 CUTLASS kernel, [feat]: CUTLASS block scaled group gemm for SM100, [Feature] Integrate SM100 DeepGEMM support, [Kernel] SM90 CUTLASS FP8 GEMM: add support for swap AB + kernel tuning, [feat]: add SM100 support for cutlass FP8 groupGEMM, [fix]: disable cutlass block scaled group gemm for EP, [Kernel] DeepGemm MoE : Integrate triton permute / unpermute kernels , [Perf] Add swap_ab to SM90 FP8 non-block CUTLASS moe grouped gemm, [Perf] Cuda Kernel for Per Token Group Quant, Support CUTLASS NVFP4 (w4a4) for Blackwell Geforce GPUs (SM120), [Kernel] Add support for block FP8 on SM120 (NVIDIA 5090 and RTX PRO 6000), [Fix] enable swap_ab for pplx problem size computation, [Kernel] CUTLASS MoE FP8: Integrate cuda moe permute/unpermute, [kernel] Support W4A8 on Hopper, [Perf] Use upstream CUTLASS for SM90 Block FP8 kernel, [NVIDIA] Support SiluMul + NVFP4 quant fusion, [Perf] SM100 - add swap AB optimization to CUTLASS FP8 GEMM, [Kernel] Add NVFP4 MoE CUTLASS support for SM120, [Kernel]Support W4A8 Grouped GEMM on Hopper, gptq marlin quantization support for fused moe with lora, [LoRA] Support Quantized Adapters, [Perf][Kernel] Optimize FP4 quantization kernels (SM100F), [Bugfix] Fix quant RMS norm fusion for quantization with TMA-aligned scales, [Kernel] Add enable_sm120_or_later for SM121 (DGX Spark) CUTLASS support, [Bugfix]fix output Nan/Inf in marlin if dtype=float16, [ModelBash][DSV3] Add TRTLLM DSV3 Router GEMM kernel (6% B1 Speedup), [Kernel] Integrate SM100 MXFP8 blockscaled grouped MM and quant kernels, [Model Bash] DeepSeek R1 BF16 Min Latency QKV A GEMM (0.5% E2E Speedup), [Kernel] Add gpt-oss Router GEMM kernel, [Kernel] Add non-gated support for NVFP4 CUTLASS MoE, [Kernel] Add MXFP4 W4A4 CUTLASS MoE kernel for SM100, [4/n] Migrate FP4/W4A8 CUTLASS kernels to torch stable ABI, [Kernel] Optimize SM120 CUTLASS blockwise FP8 GEMM, [Kernel] Add swapAB support for SM120 CUTLASS blockwise FP8 GEMM , [NVIDIA] Bugfix NVFP4 DGX Spark and RTX50, [Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity, [Kernel] (1/N) Machete - Hopper Optimized Mixed Precision Linear Kernel , DeepGEMM — runtime-JIT tensor-core kernels, FP8 block-scale GEMM, Gated Dual GEMM (Gate-Up + Activation), Grouped GEMM for MoE, NVFP4 GEMM, Small-M M-grouped GEMM on SM90, TensorRT-LLM Blackwell FP4 DSA Indexer
gemv Twelve Attempts at an FP4 Kernel, NVFP4 GEMV, Blackwell NVFP4 Kernel Hackathon Journey, GPU Mode NVFP4 Hackathon - Problem 1: Batched GEMV, [Intel GPU] Enable DeepSeek V4 Inference on XPU, [DSV4] Fuse norm and router for low latency scenario, NVFP4 batched GEMV
grouped-gemm DeepGEMM tensor-core kernel library, Anatomy of a Reward Hack, TFLOPS Gap: Why FP4 MoE Kernel Engineering Matters on Blackwell, GPU Mode NVFP4 Hackathon - Problem 4: Grouped GEMM, CUTLASS Changelog: SM100/Blackwell Entries, [Public release 26/04] Introducing Mega MoE, FP4 Indexer and other features/fixes, [None][perf] Add more optimization options for MOE CuteDSL finalized kernel, [TRTLLM-9831][perf] Enable 2CTA with autotune for CuteDSL MoE and Grouped GEMM optimizations, [None] [feat] Add densegemm backend for MoE, [None][feat] Add DWDP (Distributed Weight Data Parallelism) support for MoE inference, [TRTLLM-11585][feat] Add CUTEDSL moe backend for nemotron-h, [None][fix] Fix the performance issue of FP8 blockwise grouped GEMM when using attention DP, [TRTLLM-9685] [feat] Add gather fc1 kernel by cuteDSL, [EVT] Add support for Row/Col broadcast PtrArray, Hopper Grouped GEMM support for FP8 Accum, Blockwise and Groupwise GEMM for Blackwell and Improvements for Hopper, Support PDL for SM90 Array TMA GEMM, Blockscaled Ragged Contiguous Grouped Gemm for MoEs, [Hopper CuTeDSL] Add grouped GEMM kernel example, Small Tile N BlockScaled GEMM + Grouped GEMM on SM12x, Paged Attention support for FA3, FA3 paged attention: Readiness for Cutlass 3.6 / default value for block_table, perf: accelerate blackwell grouped gemm, feat: Support MXFP8 x MXFP4 CUTLASS grouped GEMM, add cutlass backend for mm_fp4, GPT-OSS Support: Add Blackwell MoE mxfp4 implementation from TRTLLM and Attention Sink, Add python API for masked grouped gemm, feat: cutlass fp8 gemm bringup for SM120 & SM121, Support Kimi-K2 for TRT: templatize number of experts, TVM: support TVM binding for GroupedGemm, feat:enable fp8 blockscale moe for fused cultass for sm90, silu_and_mul nvfp4 quanization fusion rework, Feature: Add support for L40 FusedMoE in cutlass path, feat: Add CuTe DSL grouped-gemm + combine fusion support, perf: Optimize CUTLASS MoE helper kernels for small-batch decode workloads, [sgl-kernel][1/N]Support Expert Specialization Grouped GEMM, [sgl-kernel][4/N]Support Expert Specialization Grouped GEMM, [sgl-kernel][Feat][B200][1/N]Support MXFP8 Grouped GEMM in Blackwell, [sgl-kernel][Feat][B200][2/N] Support MXFP8 Grouped GEMM in Blackwell, [sgl-kernel][6/7]Support Expert Specialization Grouped GEMM, [fp8] SM90 swap-AB scaled_mm dispatch (~1.16x kernel geomean, +5.8-18.5% end-to-end), feat: integrate deepgemm into EPMoE, feat: support DeepSeek-R1-W4AFP8 model with ep-moe mode, [1/n]: add cutlass W4A8 moe kernel for hopper architecture, [NVIDA] [1/N] Nvfp4 Masked Gemm: Add quant op for the flashinfer grouped gemm, [Kernel] CUTLASS grouped gemm fp8 MoE kernel, permute/unpermute kernel for moe optimization, [Kernel] DeepGemm MoE : Integrate triton permute / unpermute kernels , [Perf] Add swap_ab to SM90 FP8 non-block CUTLASS moe grouped gemm, [Kernel] CUTLASS MoE FP8: Integrate cuda moe permute/unpermute, [Kernel]Support W4A8 Grouped GEMM on Hopper, [LoRA] Support Quantized Adapters, [Kernel] Add MXFP4 W4A4 CUTLASS MoE kernel for SM100, [NVIDIA] Bugfix NVFP4 DGX Spark and RTX50, DeepGEMM — runtime-JIT tensor-core kernels, Fused MoE — Expert GEMM and Adjacent Operations, Grouped GEMM for MoE, Small-M M-grouped GEMM on SM90
linear-attention Gated Delta Networks reference repository, NVIDIA Qwen3-Next Architecture Announcement, FlashInfer MLSys 2026 - Track C: Gated Delta Net, Tiled Flash Linear Attention (TFLA), Gated Delta Network kernels
mla FlashMLA upstream README, DeepSeek-V3.2-Exp in vLLM: Fine-Grained Sparse Attention in Action, CUTLASS Changelog: SM100/Blackwell Entries, K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model, [TRTLLM-11289][feat] Integrate CuteDSL's bf16 dense GEMMs, [None][feat] Support sparse mqa/gqa attention, [None][feat] Trtllm-gen FMHA JIT support, [#12784][feat] AutoDeploy: Optimize DeepSeek-R1 model performance, [None][feat] Add DeepSeekV4 attention kernels, [None][feat] Keep DSv4 o_a_proj as FP8, and port vLLM's fused_inv_rope_fp8_quant, [None][feat] DSv4: enable GVR Heuristic Top-K for compress_ratio=4, Flash MLA support, Flash MLA Support - Step 2, [ex77] fix mla split; add fwd lse; add bwd varlen, Example 77 add blackwell fmha bwd for MLA shape, Add Blackwell MLA forward (shape: d=192, dv=128) implementation, [Cute,Sm100,Fwd] add MLA 64/512 with topk sparsity for MQA 128 heads, misc: fix instrument code for mla profiler, [nvidia] Add Blackwell FMHA decode kernel from TRT-LLM, bugfix: temporally disable split-kv in blackwell mla, [Feature] Support PDL for batch Prefill and Decode, update trtllm-gen decode attention kernel launcher, feat: add trtllm-gen mla cubin, refactor: refactor trtllm-gen attention kernel integration code, Fix the bug of the kernel-selection heuristic in trtllm-gen, feat: Fused rope fp8 quantize kernel for MLA, TGV GEMM as a BF16 backend alternative to cuBLAS, perf: Port the separate reduce kernel mode from trtllm., feat: add xqa fp8 mha and fp8 kv cache, MLA RoPE + quantization fused kernel: shape generalization for MHA / GQA, minor fix for xqa, feat: Add flashinfer.rope.rope_quantize_fp8_append_paged_kv_cache (fused RoPE + Q + KV cache, supports MLA/GQA/MHA) , feat: add xqa mla backend, Fix: several bugs/issues with trtllm-gen attention kernels. , feature: make the LSE returned by MLA support base 2 or e #2113, feat: add trtllm-gen per-tensor sparseMla kernels., [feat] Integrate SGLang concat_mla_k kernel into flashinfer, [TRTLLM-Gen Fmha] add optimized trtllm-gen decode kernels for high throughput + speculative decoding, Add cute dsl mla decode op, feat: FP8 output support for CUTLASS MLA paged attention, feat: Support padding tokens with seqlen=0 for rope+quant+kv cache update fusion kernel, [CuTe DSL] Add modular FMHA prefill and MLA decode attention kernels, [Fmha] Sparse MLA decode kernel selection heuristics, feat: add pdl support for cute dsl mla decode kernel support, feat: Enable FP8 (E4M3/E5M2) in concat_mla_k for optimize long-context prefill performance and refactor type dispatch for BF16/FP16, Support Kimi K2.5 H64 CuTe DSL MLA decode, Add dynamic tokens-per-page TRTLLM-GEN GQA kernels, feat: support deepseek prefill attention shape, bugfix: MLA decode should multiply sm_scale by math::log2e, fix rope logic in mla decoding, perf: memory efficient deepseek mla fused page-attention kernel, bugfix: mla page-attention kernel for different page sizes, feat: unlocking MLA for A100, feat: unlock MLA attention for sm89 (L40/L40s/4090), bugfix: bugfix on sm89 MLA, perf: MLA decode kernel implemented by CuTe targeted to SM80, perf: dynamic split-k for MLA, bugfix: fix the behavior of MLA kernel when kv-length is 0, perf: FlashAttention-3 style MLA PageAttention, feat - support mla kvcache store, perf: fix MLA split-k performance bug, perf: tweak the pipeline design of mla kernel, feat: flashinfer intra-kernel profiler, perf: Use 2WG pipeline design for MLA implementation on Hopper, perf: prefetch page indices for mla kernel, feat: Add FP4 (E2M1) KV Cache Support with Quantization Utilities for MLA, Optimize FP8 MLA KV cache writes with Triton kernel, [Move sgl-kernel Kernel to JIT] Add JIT concat MLA kernels, [Hicache & JIT_kernel] Support page first layout & mla jit kernel, Support Triton MLA FP8 KV cache, [AMD] Enable FP8 KV cache and FP8 attention kernel for NSA on MI300/MI355 with TileLang backend, [Intel GPU] Enable DeepSeek V4 Inference on XPU, amd/deepseek_v4 27/N [fix] Reduce Triton autotune configs for faster first-time server launch, Support MHA with chunked prefix cache for DeepSeek chunked prefill, Blackwell Cutlass MLA kernel, fix: solve cu118 issue for cutlass mla, [perf] experimental enhance fp8 per-tensor quant, [perf] introduce deep gemm group_gemm_masked as bmm, Fuse MLA set kv cache kernel, Cutlass MLA: Disable split kv due to https://github.com/NVIDIA/cutlass/issues/2274, [perf][sgl-kernel] extend cutlass_mla_decode to support num_head < 128, [Attention] MLA decode optimizations, [Attention] MLA with chunked prefill, [Perf] Mem align KV caches for CUDA devices (MLA perf improvement), [Kernel] Make rotary_embedding ops more flexible with input shape, [Kernel] moe wna16 cuda kernel, [core] Perf improvement for DSv3 on AMD GPUs, [Kernel] moe wna16 marlin kernel, [Kernel] allow non-contiguous input for marlin kernel, [NVIDIA] Support Cutlass MLA for Blackwell GPUs, [Kernel] support merge_attn_states CUDA kernel, 3x speedup, [Kernel] Have rotary embeddings support tensors, [BugFix] FA2 MLA Accuracy Issue, [Bugfix] Fix some narrowing conversion warnings, SM100 Cutlass MLA decode with unrestricted num_heads (< 128) for DeepSeek TP, [perf] Add fused MLA QKV + strided layernorm, [Compile] Fix Compile Warning SM100 Cutlass MLA, [Feature] Support Decode Context Parallel (DCP) for MLA, [Kernel] Support decode context parallelism on Blackwell with CUTLASS MLA, Fuse RoPE and MLA KV-cache write, [Attention] Use sparse prefill kernel for fp8 kv-cache in DeepSeek-v3.2, bugfix: correct attn output with base 2 or e, OffloadingConnector: Support kernel_block_size != block_size, Triton MLA perf fixes, [Kernel] Add FP8 KV cache support to Triton MLA decode attention, [Attention][Perf][Kernel] Replace torch.cat with vectorized CUDA kernel MLA query concat - DeepSeek-V3.2, [Bugfix] Fix DSV3 kernels breaking _C and _moe_C on unsupported arches, Add 320 dimension size support to MLA, [MTP][Sparse MLA] Take advantage of native MTP support in indexer when possible, [Refactor] Improve indexer decode path metadata preparation, [MLA] Optimize mla indexer prepare uniform decode for MTP > 1, [Performance][DSR1]: Fused RoPE+KVCache+q_concat for MLA, FlashAttention SM100 MLA TopK Sparse Forward, FlashMLA attention kernels, Sparse MLA
moe NVIDIA Qwen3-Next Architecture Announcement, TFLOPS Gap: Why FP4 MoE Kernel Engineering Matters on Blackwell, FlashInfer MLSys 2026 - Track A: Fused MoE FP8, CUTLASS Changelog: SM100/Blackwell Entries, K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model, [Public release 26/04] Introducing Mega MoE, FP4 Indexer and other features/fixes, Sync nv_dev with upstream #316 (Mega MoE optimizations & benchmarks), [None][perf] Add more optimization options for MOE CuteDSL finalized kernel, [TRTLLM-9992][perf] Enable PDL for CuteDSL kernels and overlap MoeOutputMemset, [None][feat] CuteDSL MOE FC1 Enhancement, [TRTLLM-9831][perf] Enable 2CTA with autotune for CuteDSL MoE and Grouped GEMM optimizations, [None] [feat] Add densegemm backend for MoE, [None][feat] fuse shared to sparse experts in TRT-LLM Gen MoE, [https://nvbugs/5799917][fix] Recover from CUTLASS MoE doActivation perf regression for MXFP4/NVFP4 dtype, [None][feat] TRT-LLM Gen MoE finalize kernel optimization, [None][feat] Add support for expert_number<=2048 and K<=32, [https://nvbugs/5799917][fix] Recover from CUTLASS MoE doActivation perf regression for MXFP4/NVFP4 dtype, [None][feat] CuteDSL MOE: Add raster along M/N support for blockscaled contiguous backbone kernel, [None][feat] Add DWDP (Distributed Weight Data Parallelism) support for MoE inference, [None][feat] Support update weight for nvfp4, [https://nvbugs/5983390][perf] Kernel fusions in _gather_k_cache_for_chunk of Indexer in DSA, [None][perf] add Dynamic SMEM block routing in MOE, [None][feat] Optimize mamba SSD prefill and extend flashinfer dispatch, [TRTLLM-11585][feat] Add CUTEDSL moe backend for nemotron-h, [#12784][feat] AutoDeploy: Optimize DeepSeek-R1 model performance, [None][feat] Update rms_norm + fp4_qaunt kernel supporting more dim, [None][perf] Extend customMoeRouting kernel to support Qwen3.5, [None][feat] Fuse FP8 1x128 quantize + UE8M0 scale pack on SM100, [https://nvbugs/6108841][fix] add hidden_dim=6144 router GEMM instantiation for GLM-5, [None][fix] Plumb swiglu_limit through DeepGEMM and TRTLLMGen FP8 fused MoE, [None][perf] FC2 DenseGEMM autotune: split-K, swap_ab, fine-grained tuning buckets, [None][perf] mHC fused_hc kernel optimizations + DS-V4 entry-boundary RMSNorm fold-in, [None][feat] Keep DSv4 o_a_proj as FP8, and port vLLM's fused_inv_rope_fp8_quant, [None][feat] Add chunked prefill support for Gemma4 (text + vision multimodal), [None][chore] Fix kernel launch param and add TRTLLM MoE backend test, [None][fix] Fix and add test for TRTLLM MoE backend, [TRTLLM-8637][feat] Optimize the routing kernel for DeepseekV3 (MoE CUTLASS backend); Add support for 384 experts (MoE TRTLLM backend), [TRTLLM-8535][feat] Support DeepSeek V3.2 with FP8 + BF16 KV cache/NVFP4 + BF16 KV cache, [None][fix] support topk autotuner input for expert slot per group larger than 32, [None][feat] TRT-LLM Gen MoE optimize DeepSeek Fp8 activation kernel, [https://nvbugs/5726962][feat] Apply fusion for W4AFP8_AWQ MoE, [None][feat] Fused kernels (qknormrope + moe routing) and two-model MTP support for glm4moe, perf: accelerate blackwell grouped gemm, Add CUTLASS fused moe kernels from TensorRT-LLM., feat: trtllm-gen fp8 moe kernels, Feature/sm100 low latency nvfp4 kernels, Update cutlass fp4 moe kernels, bugfix: fixed cutlass fused moe usage of FP4QuantizationSFLayout::SWIZZLED, GPT-OSS Support: Add Blackwell MoE mxfp4 implementation from TRTLLM and Attention Sink, gpt-oss: Add MXFP8 x MXFP4 CUTLASS MOE for SM100 and BF16 x MXFP4 CUTLASS for SM90 + SwigluBias Activation, tuner: Trtllm-gen Fp4 MoE Autotunner, fix: update cutedsl masked moe gemm, Add GeGLU support to trtllm-gen NVFP4 Fused MoE Kernel, fix: separate out fp4 lib into sm90 and sm100 versions, add oob checking in fused moe, Fix DeepSeek quality for TRTLLM fused MoE routing, feat:enable fp8 blockscale moe for fused cultass for sm90, Update the routing for TRTLLMGEN to support kimi k2 and qwen, feat: Add FP4 TRTLLM-Gen throughput MOE batched gemms, Feature: Support Relu2 activation in fused MoE, Update trtllm-gen fused moe routing kernel and add more kernels, Feature: Add support for L40 FusedMoE in cutlass path, [feat] Refactor trtllmgen MOE and add Bf16 trtllmgen moe, update trtllm cutlass moe , perf: Speed up fp4 quantization for small batch with swizzling for cutlass MoE, [BUG] Fix trtllm-gen fp4 moe renormalize routing, perf: TRT-LLM MoE Block-FP8 activation optimization, refactor: pass hopper deepgemm include directory through python, perf: TRT-LLM Gen finalize kernel optimization, enable sm103 moe dsl backend, feat: MxInt4 x Bf16 TRT-LLM Gen MoE support, feat: Support Fused MoE non gated Relu2 NVFP4 & FP8 and support Nemotron, [ML3] Optimized Router Gemm, feat: cuteDSL fp4 moe for better DSR1 performance., feat: Support Fused MoE non gated Relu2 NVFP4 & FP8 and support Nemotron, fixed, fix: W4A8 autotune crash in cutlass_fused_moe profiler workspace, Implement cutlass_fused_moe mxfp8, fix: cute dsl nvfp4 moe routing index error, [fp8_blockwise]Fix int32 overflow in TRTLLM fused MoE activation kernel, [feat] Add 2048 experts and 32 Top K , CuteDSL MoE fix redundant output buffer zeroing, [NVIDIA] fix(jit): enable GDC for CUTLASS fused MoE PDL — prevent random crashes on SM12x, perf: Optimize CUTLASS MoE helper kernels for small-batch decode workloads, [feat] Add routing_replay_out support to MoE kernels and Python API, Prevent MoE autotuner buffer overflow on large token buckets, [feat] Trtllm-gen Per-token Nvfp4 MoE, feat: Add b12x CuTe DSL fused MoE for SM120, Integrate CUTLASS Small Tile N Blockscaled GEMMs/Grouped GEMMs for SM120 and SM121, feat: enable glm5 router gemm, fix(sm12x): fix micro-kernel workspace sizing when routed_rows > num_local_experts, fix(cute_dsl/moe): make autotuner bucket configuration adapt to runtime input, fix(cute_dsl/moe): unbias autotuner profiling for tile_size enumeration, feat(moe): add SM120 W4A16 b12x kernels, feat(cute_dsl/moe): deterministic balanced autotune profile inputs, feat(cute_dsl/moe): add moe_output_memset_inplace dense memset wrapper, feat: Add FP4 (E2M1) KV Cache Support with Quantization Utilities for MLA, Fix correction bias undefined behavior for nvfp4 models, Update CUTLASS. Refine KernelSchedule for fp8 (grouped) gemm., [sgl-kernel][1/N]Support Expert Specialization Grouped GEMM, support cutlass fp4 kernel in sm120, [sgl-kernel][4/N]Support Expert Specialization Grouped GEMM, Support moe topk sigmoid kernel, [DeepSeek v3.2] opt Context Parallelism: support fused moe, multi batch and fp8 kvcache, [kernel][moe] add moe topk fast, [LoRA][III] Add LoRA support for MoE layers and enable TP, Add new moe wna16 marlin gemm, Opt moe align block size kernel, [sgl-kernel][1/2] Fused qk_norm_rope for GLM4.6, Fix warp illegal instruction in kimi k2 thinking PCG, MoE: Skip SiLU/GELU activation for masked experts, Add SwapAB Optimization for triton fused_moe_kernel on SM90., [Rework] Add SwapAB Optimization for triton fused_moe_kernel on SM90., Add mxfp8 support for online quantization, Triton dense linear, and CUTLASS MoE, feat: add FA4 SM90 paged KV decode support & update attention docs, [jit_kernel] Add fused_qknorm_rope JIT kernel, [Kernel Slimming] Migrate NVFP4 kernels to JIT, [Feature] NVFP4 Marlin fallback for non-Blackwell GPUs (SM75+), [JIT Kernel] Reland NVFP4 kernels to JIT, Deepseek_v4 support w4(mxfp4)a16 on hopper, [feat] Init true on policy with qwen_dense, [VLM] Optimize Gemma4 VLM with PCG and fuse RMSNorm + residual add + scalar, Port MXFP4 Marlin MoE support to JIT kernel path, [Gemma4] Optimize Gemm4 with fused Q/K/V RMSNorm + per-expert FP8 ckpt loader, [rebase]Deepseek_v4 support w4(mxfp4)a16 on hopper, [Intel GPU] Enable DeepSeek V4 Inference on XPU, Feature DeepSeek V3/R1 INT8 Quantization (block-wise), [Feature] DeepSeek V3/R1 INT8 Quantization (channel-wise) , Accelerate FP8 CUDA Kernel by 20-28%, Add deepseek style fused moe group gate selection kernel, support cmake for sgl-kernel, reduce moe_align_block_size_kernel small batch mode overhead, [1/2] Add FP8 Blockscale MoE CUTLASS kernel for Blackwell, [2/2] Add python wrapper for CUTLASS FP8 Blockscale MoE Kernel. , [1/2] Add Kernel support for Cutlass based Fused FP4 MoE, reduce torch.zeros overhead in moe align block size kernel, Fix bug of deepseek-v3 under DP+EP mode with large batchsize/seqlen, Refine pre_reorder_triton_kernel slightly to improve performance, [EP] Add cuda kernel for moe_ep_pre_reorder, Set num_fused_shared_experts as num_shared_experts when shared_experts fusion is not disabled, Support token-level quantization for EP MoE, feat: integrate deepgemm into EPMoE, [EP] Add cuda kernel for moe_ep_post_reorder, fix ep_moe_reorder kernel bugs, Add a CUDA kernel for fusing mapping and weighted sum for MoE., [sgl-kernel] Add cuda kernel for moe_ep_silu_and_mul, Fuse routed scaling factor in deepseek, Add CUTLASS FP8 Blockscale MoE kernel for Hopper architecture, Fuse sorted_token_ids padding to moe_align_block_size kernel, fix: fix apply_shuffle_mul_sum, feat: support DeepSeek-R1-W4AFP8 model with ep-moe mode, [1/n]: add cutlass W4A8 moe kernel for hopper architecture, [kernel] opt moe align block kernel by block/warp scan algorithm, [feat] Support tp mode for DeepSeek-R1-W4AFP8, [fix]: fix cutlass moe ut and and Opt H20 cutlass groupGemm performance, Optimize moe_sum_reduce_kernel, Update CUTLASS 4.2 & Enable K-Major Scale Factor for SM90 FP8 Blockwise Group GEMM, Make sm100 fp8 kernels available on sm103, [Model] Support Meituan LongCat-Flash && LongCat-Flash-MTP, [Kernel] add triton fused moe kernel for gptq/awq, [Kernel] port sgl moe_align_block_size kernels, Expert Parallelism (EP) Support for DeepSeek Models, Optimize moe_align_block_size for deepseek_v3, [Kernel] moe wna16 cuda kernel, [core] Perf improvement for DSv3 on AMD GPUs, [Kernel] CUTLASS grouped gemm fp8 MoE kernel, [Kernel] moe wna16 marlin kernel, permute/unpermute kernel for moe optimization, [Kernel] GGUF MoE kernel, [Kernel] Fix conflicting macro names for gguf kernels, Modularize fused experts and integrate PPLX kernels, [Hardware/NVIDIA/Kernel] [Functional Enablement] [1/N] Enable nvidia/DeepSeek-R1-FP4 Model, [Kernel] GGUF MoeVec kernel, [BugFix] Accuracy fix for llama4 int4 - improperly casted scales, [Kernel] some optimizations for dense marlin and moe marlin, [Kernel] Add expert_map support to Cutlass FP8 MOE, Fix numel() downcast in vllm/csrc/moe/moe_align_sum_kernels.cu +2, [Kernel] fp4 marlin kernel, [Kernel] Integrate CUTLASS MoE kernel with PPLX, [Hardware][NVIDIA] FP4 MoE kernel optimization, [Hardware][NVIDIA][kernel] Fp4 MOE quant kernel optimization, [feat]: CUTLASS block scaled group gemm for SM100, [Feature] Integrate SM100 DeepGEMM support, [Bugfix] Fix topk_ids indices_type for CUTLASS w8a8 FP8 MoE, [feat]: add SM100 support for cutlass FP8 groupGEMM, [Performance] Performance improvements in non-blockwise fp8 CUTLASS MoE, [fix]: disable cutlass block scaled group gemm for EP, [Kernel] DeepGemm MoE : Integrate triton permute / unpermute kernels , [Perf] Add swap_ab to SM90 FP8 non-block CUTLASS moe grouped gemm, [Feature][Kernel]FusedMoE LoRA, [Bug] Fix Compressed Tensor NVFP4 cutlass_fp4_group_mm illegal memory access, [Fix] enable swap_ab for pplx problem size computation, [Kernel] CUTLASS MoE FP8: Integrate cuda moe permute/unpermute, [Kernel] Add fused grouped_topk kernel for MoE, [Bugfix][Misc] Fix silu_and_mul_nvfp4_quant issue and extract common utils for nvfp4 kernel source files, [Model] Add LongCat-Flash , [Kernel][Quantization] add w4a8 support for marlin kernel, Update launch_bounds_utils.h for correct compile on Multiple Cuda Arch - PTXAS out of range Warning, [Attention] Use sparse prefill kernel for fp8 kv-cache in DeepSeek-v3.2, [Perf][DeepSeek] Add sigmoid+bias fusion to fused_grouped_topk from TRTLLM, [Model] Add support for openPangu moe model, [Kernel] Add NVFP4 MoE CUTLASS support for SM120, Lora MoE Align Improvements, Add unpermute-aware fused MoE path and small-batch fallback, [Kernel][MoE] optimize moe_align_block_size, [Kernel]Support W4A8 Grouped GEMM on Hopper, [Kernel][Quantization][MoE] add marlin kernel support for turing (sm75), gptq marlin quantization support for fused moe with lora, [NVFP4][Perf] Tune NVFP4 input quant kernel for small batch size, [Kernel] Add topk_sigmoid kernel, Add TMA support to fused_moe_lora kernel, [Perf][Kernel] Optimize FP4 quantization kernels (SM100F), [Bugfix]fix output Nan/Inf in marlin if dtype=float16, [Kernel] Optimize grouped topk kernel, [ModelBash][DSV3] Add TRTLLM DSV3 Router GEMM kernel (6% B1 Speedup), [Kernel] Integrate SM100 MXFP8 blockscaled grouped MM and quant kernels, [Quantization] add humming quantization kernel, [Bugfix] Fix expert_ids padding values in moe_align_block_size kernel, [Kernel] Add gpt-oss Router GEMM kernel, [Kernel] Add non-gated support for NVFP4 CUTLASS MoE, [Kernel] Add MXFP4 W4A4 CUTLASS MoE kernel for SM100, [NVIDIA] Bugfix NVFP4 DGX Spark and RTX50, fix: clamp NaN/Inf in topk_softmax to prevent duplicate expert IDs, [Bugfix] moe lora align kernel grid, [DSV4] Fuse norm and router for low latency scenario, [MoE] Move various experts classes to fused_moe/experts/, [6/n] Migrate activation kernels, gptq, gguf, non cutlass w8a8 to libtorch stable ABI (continued), [Kernel] (2/N) Machete - Integrate into CompressedTensorsWNA16 and GPTQMarlin, Fused MoE — Expert GEMM and Adjacent Operations, Grouped GEMM for MoE, Small-M M-grouped GEMM on SM90
prefill FlashMLA upstream README, DeepSeek-V3.2-Exp in vLLM: Fine-Grained Sparse Attention in Action, FlashInfer MLSys 2026 - Track C: Gated Delta Net, FlashMLA attention kernels, Sparse MLA
quantization NVFP4 Format Details, DeepSeek-V3.2-Exp in vLLM: Fine-Grained Sparse Attention in Action, [None][feat] CuteDSL MOE FC1 Enhancement, [None][feat] sm100 weight-only kernel, [None][fix] impl fused triton kernel for e8m0 resmooth to reduce memory footprint, [None] [feat] Add densegemm backend for MoE, [None][feat] Optimize super-v3 nvfp4 for better perf, [None][feat] Optimize by fuse nvfp4_quant to layernorm_gated for mamba2_mixer, [TRTLLM-11119][feat] Blackwell SageAttention, Integrate into AttentionOp API, [TRTLLM-10990][feat] Fuse SwiGLU and quant into shared expert, [TRTLLM-10421][perf] Add fused cat+fp8_quantize CUDA kernel for DSA indexer, [None][feat] Add DWDP (Distributed Weight Data Parallelism) support for MoE inference, [None][feat] Support update weight for nvfp4, [TRTLLM-10407][perf] Add cute dsl single pass multi cta cluster topk, [TRTLLM-11585][feat] Add CUTEDSL moe backend for nemotron-h, [TRTLLM-11485][feat] Feature rework: Add SageAttention refreshed kernels (attentionOp only), [None][feat] Update rms_norm + fp4_qaunt kernel supporting more dim, [None][feat] Add FP4 residual quantization kernel without channel reo…, [None][feat] Integrate FP4 indexer for DSA on Blackwell, [None][feat] Fuse FP8 1x128 quantize + UE8M0 scale pack on SM100, [None][feat] Add DeepSeekV4 attention kernels, [None][fix] Plumb swiglu_limit through DeepGEMM and TRTLLMGen FP8 fused MoE, [TRTLLM-35237][feat] Add cute dsl FP4 paged MQA logits decode kernel, [None][feat] Keep DSv4 o_a_proj as FP8, and port vLLM's fused_inv_rope_fp8_quant, feat: Add w4a8_mxfp4_fp8 quantization recipe., [OMNIML-2336][feat] Add NVFP4 x FP8, [https://nvbugs/5726962][feat] Apply fusion for W4AFP8_AWQ MoE, [None][feat] Port fp4 quantization kernel optimization from FlashInfer, [None][feat] Adding torch ext API for FusedAddRMSNormQuant kernel, Improve sm90 mixed dtype kernel, new example with TMA prefetch feature targeting for DRAM latency boun…, [CLI] add cutedsl fp16 gemm tutorial from 2 to 6, [Cute,Fwd,Sm100] fp8 e4m3 and e5m2 support, [hd256] Improve forward kernel with exp2 FMA emulation (3% to 9% performance gain), [hd256] Add TMA paged KV support to SM100 2CTA forward kernel, add multi-item scoring, feat: add functional per-head FP8 quantization for FA3, bugfix: fix fp8 attention kernels aot compilation issue, Add CUTLASS fused moe kernels from TensorRT-LLM., bugfix: Fix test and output shape of fp4 quantize, [Feature] Support PDL for batch Prefill and Decode, Feature/sm100 low latency nvfp4 kernels, Bug fix: guard fp8 e8m0 and e2m1 compile , [fix] fix integer overflow in FA2 customized_mask & add buffer overflow warning., Update cutlass fp4 moe kernels, add cutlass backend for mm_fp4, Add blockwise-scaled FP8 GEMM via TRTLLM-Gen., feat: Fused rope fp8 quantize kernel for MLA, feature: add fp4 mm using trtllm backend, GPT-OSS Support: Add Blackwell MoE mxfp4 implementation from TRTLLM and Attention Sink, gpt-oss: Add MXFP8 x MXFP4 CUTLASS MOE for SM100 and BF16 x MXFP4 CUTLASS for SM90 + SwigluBias Activation, Add alignment in MxFP8Quantization, Remove getEnvEnablePDL in favor of enable_pdl parameter, bugfix: Fix compile error for undefined swizzle enum., fix: separate out fp4 lib into sm90 and sm100 versions, add oob checking in fused moe, feat: cutlass fp4 gemm bringup for SM120 & SM121, bugfix: fix fp4 quantization with 8x4 scale factor layout, perf: Port the separate reduce kernel mode from trtllm., Masked batch nvfp4 quantization, [Quantization] Add per-expert global scaling factor for fp4 batched quantize, MLA RoPE + quantization fused kernel: shape generalization for MHA / GQA, Add layernorm op for inputs of mixed dtype, silu_and_mul nvfp4 quanization fusion rework, fix: correct PDL parameter handling in RopeQuantize kernel, update trtllm cutlass moe , perf: Speed up fp4 quantization for small batch with swizzling for cutlass MoE, feat: Add flashinfer.rope.rope_quantize_fp8_append_paged_kv_cache (fused RoPE + Q + KV cache, supports MLA/GQA/MHA) , [BUG] Fix trtllm-gen fp4 moe renormalize routing, perf: TRT-LLM MoE Block-FP8 activation optimization, add tensor scale input for xqa, refactor: update fa3 codebase and fix hopper unittest [part 1], feat: TRTLLM FMHAv2 backend for ctx attention, feat: MxInt4 x Bf16 TRT-LLM Gen MoE support, feat: Fused RMSNorm + FP4 Quantization Kernels in CuTe-DSL, feat: RMSNorm/Fused RMSNorm + FP8 Quantization kernels, fix: Add global scale support and optional output allocation for RMSNorm+FP4Quant fusion kernels, Optimize quantization function in large problem size, fix: In-place Residual Update for add_rmsnorm_fp4quant, feat: Add output_both_sf_layouts option to add_rmsnorm_fp4quant API, feat: cuteDSL fp4 moe for better DSR1 performance., refactor: simplify fp4 rmsnorm, refactor: refactoring cuda code to cute-dsl (part 1), fix: Fix NaN output in mxfp8_quantize for very small input values, Add cute-dsl backends to mxfp[8,4]_quantization for future refactor, feat: Add TRTLLM fmha_v2 library for SM90 attention with Skip-Softmax , feat: Add MXFP8 GEMM mm_mxfp8 (cutlass), Support NVFP4 KV cache decode on SM120, feat: cute dsl mmfp4 for blackwell, fix: W4A8 autotune crash in cutlass_fused_moe profiler workspace, Implement cutlass_fused_moe mxfp8, fix: cute dsl nvfp4 moe routing index error, int16 Block-Scaled State and Stochastic Rounding for SSU (mamba), [feat] trtllm-gen mxfp8 gemm, feat: support mxfp4 & mxfp8 entrypoint for blackwell cutedsl dense gemm, Add varlen and speculative decoding support to selective state update, Add NVFP4 KV cache quantization support for SM100, feat: Add FP4 KV cache quant/dequant kernels , perf: Performance tune cute dsl RMSNorm variants, feat: FP8 output support for CUTLASS MLA paged attention, feat: Support padding tokens with seqlen=0 for rope+quant+kv cache update fusion kernel, [CuTe DSL] Add modular FMHA prefill and MLA decode attention kernels, feat: Add CuTe-DSL backend for NVFP4 quantization, perf: Optimize CuTe-DSL fp4 and fp8 quantization kernels, Improved simple mamba SSU kernel , Add flashinfer.fused_rmsnorm_silu() with native kernel backend, fix: use sym_int64 for strides in rmsnorm CuTe DSL kernels to prevent int32 overflow, feat: add PDL support to rmsnorm_fp4quant and add_rmsnorm_fp4quant CuTe DSL kernels, perf: Optimize CUTLASS MoE helper kernels for small-batch decode workloads, Prevent MoE autotuner buffer overflow on large token buckets, [feat] Trtllm-gen Per-token Nvfp4 MoE, feat: Add backend="b12x" for mm_fp4 on SM120, feat: Add b12x CuTe DSL fused MoE for SM120, feat: Enable FP8 (E4M3/E5M2) in concat_mla_k for optimize long-context prefill performance and refactor type dispatch for BF16/FP16, feat: DiT layer norm fusions for WAN: flashinfer.diffusion_ops, perf: optimize per-token nvfp4 quantization kernel., feat(moe): add SM120 W4A16 b12x kernels, feat(cute_dsl/moe): deterministic balanced autotune profile inputs, checkpointing_ssu kernel: fused replay + conditional state-write for Mamba2, Naive Support for Hopper FP8 Prefill Kernel with Per-Head Quantization, feat: Add FP4 (E2M1) KV Cache Support with Quantization Utilities for MLA, [sgl-kernel] Optimize concat_mla_k kernel, [DeepseekV32] Enable flashmla_prefill kernel with fp8 kvcache, support cutlass fp4 kernel in sm120, [sgl-kernel][Feat][B200][1/N]Support MXFP8 Grouped GEMM in Blackwell, [LoRA][III] Add LoRA support for MoE layers and enable TP, Optimize FP8 MLA KV cache writes with Triton kernel, Add mxfp8 support for online quantization, Triton dense linear, and CUTLASS MoE, [DeepSeek-V3.2][JIT-kernel] Support nsa fuse store indexer k cache, [Kernel Slimming] Migrate NVFP4 kernels to JIT, [diffusion][llm] macOS support, [Feature] NVFP4 Marlin fallback for non-Blackwell GPUs (SM75+), [JIT Kernel] Reland NVFP4 kernels to JIT, CUTLASS NVFP4 GEMM improvement of SM120, [AMD] Enable FP8 KV cache and FP8 attention kernel for NSA on MI300/MI355 with TileLang backend, [Fix/Kernel] Add JIT rmsnorm_hf kernel to fix transformers backend MMLU accuracy regression , Port MXFP4 Marlin MoE support to JIT kernel path, [fp8] SM90 swap-AB scaled_mm dispatch (~1.16x kernel geomean, +5.8-18.5% end-to-end), Support cutlass Int8 gemm, Support sm90 Int8 gemm, support w8a8 fp8 kernel with CUTLASS, Apply sgl w8a8 fp8 kernel, add tensorrt_llm common and cutlass_extensions as 3rdparty, integrate blockwise fp8 kernel, Feature DeepSeek V3/R1 INT8 Quantization (block-wise), [Feature] DeepSeek V3/R1 INT8 Quantization (channel-wise) , Support FP4 gemm (1/2), linear support deepgemm, Accelerate FP8 CUDA Kernel by 20-28%, [1/2] Add FP8 Blockscale MoE CUTLASS kernel for Blackwell, [perf] experimental enhance fp8 per-tensor quant, [perf] introduce deep gemm group_gemm_masked as bmm, [1/2] Add Kernel support for Cutlass based Fused FP4 MoE, Support token-level quantization for EP MoE, Support new DeepGEMM, feat: support DeepSeek-R1-W4AFP8 model with ep-moe mode, [1/n]: add cutlass W4A8 moe kernel for hopper architecture, [sgl-kernel] Opt per_token_quant_fp8 with warp reduce, optimize: reduce shulffle and quantization overhead in cutlass_moe sm90, [NVIDA] [1/N] Nvfp4 Masked Gemm: Add quant op for the flashinfer grouped gemm, [sgl-kernel] feat: Support sm120 cutlass fp8 gemm kernel, [NVIDIA] [2/N] Optimize silu_and_mul_scaled_fp4_grouped_quant perf, [Feature] Block-scaled GEMM support for MXFP8 on Blackwell, [Kernel]: Cutlass 2:4 Sparsity + FP8/Int8 Quant Support, [Kernel] Update cutlass_scaled_mm to support 2d group (blockwise) scaling, [Kernel] add triton fused moe kernel for gptq/awq, [ROCm] Faster Custom Paged Attention kernels, [Kernel][Quantization] Integrate block-quantized CUTLASS kernels for DeepSeekV3, [NVIDIA] Support nvfp4 quantization, [Misc][Kernel]: Add GPTQAllSpark Quantization, [Kernel]Add streamK for block-quantized CUTLASS kernels, [NVIDIA] Support nvfp4 cutlass gemm, add cutlass support for blackwell fp8 gemm, [Kernel] CUTLASS grouped gemm fp8 MoE kernel, [Kernel] optimize performance of gptq marlin kernel when n is small, dynamic distpatch of fp8 kernels, Add cutlass support for blackwell fp8 blockwise gemm, [Kernel] moe wna16 marlin kernel, [Kernel] GGUF MoE kernel, [Kernel] allow non-contiguous input for marlin kernel, [Kernel] Fix conflicting macro names for gguf kernels, [Bugfix] fix use_atomic_add support of marlin kernel when using v1 engine, Modularize fused experts and integrate PPLX kernels, [NVIDIA] Support Cutlass MLA for Blackwell GPUs, [Hardware/NVIDIA/Kernel] [Functional Enablement] [1/N] Enable nvidia/DeepSeek-R1-FP4 Model, [Kernel] Support W8A8 channel-wise weights and per-token activations in triton fused_moe_kernel, [Kernel] GGUF MoeVec kernel, [Kernel] some optimizations for dense marlin and moe marlin, [Kernel] Add expert_map support to Cutlass FP8 MOE, [ROCm][FP8][Kernel] FP8 quantization fused into Custom Paged Attention, [NVIDIA] Support Cutlass w8a8 FP8 for Blackwell Geforce GPUs (sm120), [Kernel] fp4 marlin kernel, Sm100 blockwise fp8 swap ab, [Hardware][AMD] integrate aiter chunked prefill into vllm, [Kernel] Integrate CUTLASS MoE kernel with PPLX, [Perf] Tunings for SM100 FP8 CUTLASS kernel, [Kernel] Enable fp8 support for pplx and BatchedTritonExperts., [Hardware][NVIDIA] FP4 MoE kernel optimization, [Hardware][NVIDIA][kernel] Fp4 MOE quant kernel optimization, [Perf] Further tunings for SM100 FP8 CUTLASS kernel, [feat]: CUTLASS block scaled group gemm for SM100, [Feature] Integrate SM100 DeepGEMM support, [Bugfix] Fix some narrowing conversion warnings, [Bugfix] Fix topk_ids indices_type for CUTLASS w8a8 FP8 MoE, [Kernel][Bugfix] Fixup some warnings in nvfp4_blockwise_moe when CUDA < 12.8, [Kernel] SM90 CUTLASS FP8 GEMM: add support for swap AB + kernel tuning, [feat]: add SM100 support for cutlass FP8 groupGEMM, [fix]: disable cutlass block scaled group gemm for EP, [Perf] Use Triton instead of Torch for DeepGEMM Per Token Group Quant, [Kernel] DeepGemm MoE : Integrate triton permute / unpermute kernels , [Perf] Add swap_ab to SM90 FP8 non-block CUTLASS moe grouped gemm, [Perf] Cuda Kernel for Per Token Group Quant, [perf] Add fused MLA QKV + strided layernorm, [Feature][Kernel]FusedMoE LoRA, Support CUTLASS NVFP4 (w4a4) for Blackwell Geforce GPUs (SM120), [Bug] Fix Compressed Tensor NVFP4 cutlass_fp4_group_mm illegal memory access, [Kernel] Improve machete memory bound perf, [Kernel] Add support for block FP8 on SM120 (NVIDIA 5090 and RTX PRO 6000), Fp8 paged attention update, [Bug] Fix B200 DeepGEMM E8M0 Accuracy Issue, [Fix] enable swap_ab for pplx problem size computation, [Kernel] CUTLASS MoE FP8: Integrate cuda moe permute/unpermute, [kernel] Support W4A8 on Hopper, [Perf] Use upstream CUTLASS for SM90 Block FP8 kernel, [Compile] Fix Compile Warning for w4a8_mm_entry.cu, [NVIDIA] Support SiluMul + NVFP4 quant fusion, [Bugfix][Misc] Fix silu_and_mul_nvfp4_quant issue and extract common utils for nvfp4 kernel source files, [Kernel] Faster pre-processing time for W4A8, [Transform] [Quantization] Add QuTLASS support to vLLM, [NVIDIA] Blackwell Family, [Kernel][Quantization] add w4a8 support for marlin kernel, [Bugfix] Fix accuracy issue for silu_mul + nvfp4 quant fusion kernel, [Compile] Fix Compile Warning for Ignoring MIN_BLOCK_PER_SM, Fuse RoPE and MLA KV-cache write, Update launch_bounds_utils.h for correct compile on Multiple Cuda Arch - PTXAS out of range Warning, [Perf] SM100 - add swap AB optimization to CUTLASS FP8 GEMM, [Performance] Fused blockwise quant RMS norm, [Performance][B200] silu_mul_quant: pack scales in int32, [Kernel] Add NVFP4 MoE CUTLASS support for SM120, Add unpermute-aware fused MoE path and small-batch fallback, [Kernel]Support W4A8 Grouped GEMM on Hopper, [Kernel][Quantization][MoE] add marlin kernel support for turing (sm75), Add llmcompressor fp8 kv-cache quant (per-tensor and per-attn_head), gptq marlin quantization support for fused moe with lora, [LoRA] Support Quantized Adapters, [NVFP4][Perf] Tune NVFP4 input quant kernel for small batch size, [Perf] Fuse stride preparation for NVFP4 cutlass_moe, [Perf][Kernel] Optimize FP4 quantization kernels (SM100F), [Spec Decode] Unified Parallel Drafting, [Bugfix] Fix quant RMS norm fusion for quantization with TMA-aligned scales, [Kernel] Add enable_sm120_or_later for SM121 (DGX Spark) CUTLASS support, [Bugfix]fix output Nan/Inf in marlin if dtype=float16, [Custom Ops] Add functional + out variant for scaled_fp4_quant, [Kernel] Integrate SM100 MXFP8 blockscaled grouped MM and quant kernels, [Quantization] add humming quantization kernel, [BugFix] Fix fp4 quant kernel on CUDA 12.8, [Kernel] Fuse FP8 output quantization into merge_attn_states, [Kernel] Add non-gated support for NVFP4 CUTLASS MoE, Add nvfp4 support to reshape_and_cache_flash, [Kernel] Add MXFP4 W4A4 CUTLASS MoE kernel for SM100, [4/n] Migrate FP4/W4A8 CUTLASS kernels to torch stable ABI, [Kernel] Optimize SM120 CUTLASS blockwise FP8 GEMM, [Perf] FP8 FlashInfer Attn for ViT, [Kernel] Add swapAB support for SM120 CUTLASS blockwise FP8 GEMM , [NVIDIA] Bugfix NVFP4 DGX Spark and RTX50, [Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity, [Bugfix] Fix broken explicit unquantized kv cache dtype support, [Perf] Fuse Zero Initializer for FP8 DeepGemm Block Quant Kernel, [Bugfix] Guard mxfp4_experts_quant bindings on ENABLE_NVFP4_SM100, [Perf] Batch invariance with Cutlass fp8 support, 28.9% E2E latency improvement, Faster per-token fp8 group quant packed kernel for blackwell, [DSv4] Improved fused Indexer Q quant kernel, [MoE] Move various experts classes to fused_moe/experts/, [Bugfix] Add swiglu limits to deepgemm fp8 methods, [feat] Add FP8 per-tensor Q scale support to Triton attention backend, [Perf] Use 2D-grid to eliminate divmod in W8W8 group quant, [DSv4] Improved dequant gather K cache kernel, [6/n] Migrate activation kernels, gptq, gguf, non cutlass w8a8 to libtorch stable ABI (continued), [Perf] Padded nvfp4 quant kernel to remove additional copy, 2.4%~5.7% e2e performance improvement, add cutedsl dsv4 indexer fp8 kernel, [Kernel] (1/N) Machete - Hopper Optimized Mixed Precision Linear Kernel , [Kernel] (2/N) Machete - Integrate into CompressedTensorsWNA16 and GPTQMarlin
reduction [Public release 26/04] Introducing Mega MoE, FP4 Indexer and other features/fixes, [TRTLLM-10276][feat] Integrate cutedsl argmax kernel, [TRTLLM-9831][perf] Use TMA.RED to improve effective memory bandwidth, [None][feat] Optimize by fuse nvfp4_quant to layernorm_gated for mamba2_mixer, [None][feat] Add support for expert_number<=2048 and K<=32, [TRTLLM-11092][feat] add support for visual gen FA4 attention backend, [TRTLLM-11119][feat] Blackwell SageAttention, Integrate into AttentionOp API, [TRTLLM-10421][perf] Add fused cat+fp8_quantize CUDA kernel for DSA indexer, [TRTLLM-10407][perf] Enable CuteDSL indexer_top_k in model, [None][feat] Support update weight for nvfp4, [None][feat] Temporally-Correlated Heuristic-guided Indexer TopK for Sparse Attention, [None][feat] Add triton paged attention for AutoDeploy, [TRTLLM-11485][feat] Feature rework: Add SageAttention refreshed kernels (attentionOp only), [#12716][feat] Fused cross-head QK Norm + RoPE kernel for WAN, [None][feat] Integrate FP4 indexer for DSA on Blackwell, [None][perf] Scheme X L2-aware dispatcher and PDL launchers for sparse-attention GVR Top-K, [None][perf] FC2 DenseGEMM autotune: split-K, swap_ab, fine-grained tuning buckets, [None][perf] mHC fused_hc kernel optimizations + DS-V4 entry-boundary RMSNorm fold-in, [None][feat] DSv4: enable GVR Heuristic Top-K for compress_ratio=4, [None][feat] Update the logic of FMHA JIT path, [OMNIML-2336][feat] Add NVFP4 x FP8, [TRTLLM-8637][feat] Optimize the routing kernel for DeepseekV3 (MoE CUTLASS backend); Add support for 384 experts (MoE TRTLLM backend), fix thread-reduce performance regression, Add nondeterministic reduce that uses atomics, Combine block_reduce_warp_reduction_nondeterministic.cuh specialization with original deterministic one , Split fixed-size segmented reduce dispatch header, Use integer promotion for warp_reduce, Two-phase reduction for fixed size segmented reduction for very large segment sizes, Implement the new tuning API for deterministic (rfa) reduce dispatch, Add env SegmentedReduce (non fixed-size overloads), Implement the new tuning API for detail::reduce::dispatch_streaming_arg_reduce_t, Avoid passing uninitialized values to scan_op, simplify dispatch segmented reduce to use latest dispatch and new tunings API, [STF] Add per-handle exec_place stream resources, Groupwise scaling along M for FP8 gemm, Example 77 add blackwell fmha bwd for MLA shape, Feat([FA4][CUTE DSL]) Add head_dim=256 support (forward + backward), [feat] add unified batch attention w/ correctness tests., feat: Fused temperature online softmax kernel, [feat] optimize persistent batch attention perf., Feature/sm100 low latency nvfp4 kernels, Add blockwise-scaled FP8 GEMM via TRTLLM-Gen., GPT-OSS Support: Add Blackwell MoE mxfp4 implementation from TRTLLM and Attention Sink, fix: Replace cub Max/Min with cuda::maximum/minimum for cuda 13 compatibility, perf: Port the separate reduce kernel mode from trtllm., Update the routing for TRTLLMGEN to support kimi k2 and qwen, [DSV3] Optimized Router Gemm, perf: improve sampling/mask/softmax performance (part 1/2), perf: Optimize helper max/minmax function in sampling.cuh, perf: TRT-LLM MoE Block-FP8 activation optimization, perf: bunch of features and optimizations for top-k (sampling + sparse attention), feat: Fused RMSNorm + FP4 Quantization Kernels in CuTe-DSL, [TRTLLM-Gen Fmha] add optimized trtllm-gen decode kernels for high throughput + speculative decoding, feat: [Qwen3-Next] Add Cute DSL GDN decode kernel and tests, A Blackwell-optimized version of selective_state_update (mamba), perf: improve gdn decode cute-dsl kernels, refactor: simplify fp4 rmsnorm, Add cute-dsl backends to mxfp[8,4]_quantization for future refactor, Ameyn/gdn decode cutedsl kernel, Feat/gdn decode pooled, Ameyn/gdn bf16 tolerance parallel reduction, perf(gdn): optimize MTP kernel with ILP rows and SMEM v caching, int16 Block-Scaled State and Stochastic Rounding for SSU (mamba), [feat] trtllm-gen mxfp8 gemm, Add cute dsl mla decode op, [feat] Add 2048 experts and 32 Top K , perf: Performance tune cute dsl RMSNorm variants, feat: FP8 output support for CUTLASS MLA paged attention, [CuTe DSL] Add modular FMHA prefill and MLA decode attention kernels, [Fmha] Sparse MLA decode kernel selection heuristics, feat: Add CuTe-DSL backend for NVFP4 quantization, perf: Optimize CuTe-DSL fp4 and fp8 quantization kernels, feat: Add CuTe DSL grouped-gemm + combine fusion support, Add flashinfer.fused_rmsnorm_silu() with native kernel backend, fix: tinygemm2 hang issue due to barrier sync, feat: Add b12x CuTe DSL fused MoE for SM120, Support Kimi K2.5 H64 CuTe DSL MLA decode, perf: optimize per-token nvfp4 quantization kernel., feat(moe): add SM120 W4A16 b12x kernels, checkpointing_ssu kernel: fused replay + conditional state-write for Mamba2, Support moe topk sigmoid kernel, [LoRA][III] Add LoRA support for MoE layers and enable TP, Opt moe align block size kernel, Move fa4 from sgl-kernel to jit kernel, [Diffsuion & JIT_kernel] QKNorm cross heads kernel, Tilelang sparse decode fwd for dsv32 mi355, [DeepSeek-V3.2][JIT-kernel] Support nsa fuse store indexer k cache, [Kernel] Fuse temperature + softmax in sampling for decode speedup, [Diffusion] Add qknorm rope fuse kernel, Fused_qknorm_rope kernel optimization: up to 2.4× faster, [Fix/Kernel] Add JIT rmsnorm_hf kernel to fix transformers backend MMLU accuracy regression , [feat] Init true on policy with qwen_dense, [codex] Optimize hidden-size 512 RMSNorm dispatch, [Intel GPU] Enable DeepSeek V4 Inference on XPU, Accelerate FP8 CUDA Kernel by 20-28%, Add a CUDA kernel for fusing mapping and weighted sum for MoE., Add dsv3 router gemm kernel, Add dsv3 fused a gemm to sgl-kernel, feat: auto-vectorize bf16/fp16 reduce with packed add2 intrinsics, [ROCm] Faster Custom Paged Attention kernels, Modularize fused experts and integrate PPLX kernels, [ROCm][Kernel][V1] Enable AMD Radeon GPU Custom Paged Attention on v1, Fp8 paged attention update, [Kernel] Add fused grouped_topk kernel for MoE, [Kernel] Optimize grouped topk kernel, [ModelBash][DSV3] Add TRTLLM DSV3 Router GEMM kernel (6% B1 Speedup), [Model Bash] DeepSeek R1 BF16 Min Latency QKV A GEMM (0.5% E2E Speedup), [Perf] Batch KV cache swap copies via cuMemcpyBatchAsync, [Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity, [DSV4] Fuse norm and router for low latency scenario
scan [None][feat] Add support for expert_number<=2048 and K<=32, [TRTLLM-10407][feat] Integrate CuTE DSL top-k kernel for Blackwell, [None][perf] add Dynamic SMEM block routing in MOE, [None][feat] Trtllm-gen FMHA JIT support, [TRTLLM-8637][feat] Optimize the routing kernel for DeepseekV3 (MoE CUTLASS backend); Add support for 384 experts (MoE TRTLLM backend), Add dynamic CUB dispatch for segmented_sort, Integrate decoupled lookahead warpspeed scan, Radix-selection based BlockTopK specialization, Implement the new tuning API for DeviceRleDispatch, Implement the new tuning API for DispatchSegmentedRadixSort, Implement the new tuning API for DispatchSegmentedSort, Implement the new tuning API for DispatchTopK, Avoid passing uninitialized values to scan_op, Implement the new tuning API for DispatchSelectIf, [cub]: implement utilities for policy selection, Replace detail::scan::dispatch by CUB's public API, Fix Warpspeed scan shifted output store, feat: ragged tensor padding kernel for blackwell kernel alignment, feat: Softmax free sampling, bugfix: host-precomuted plan function for blackwell fmha, Feature/sm100 low latency nvfp4 kernels, GPT-OSS Support: Add Blackwell MoE mxfp4 implementation from TRTLLM and Attention Sink, Update the routing for TRTLLMGEN to support kimi k2 and qwen, feat: implement deterministic topk, Mamba2 SSD Combined Forward Pass (Blackwell CuTe DSL Kernel), [feat] Add 2048 experts and 32 Top K , [KDA] Optimize prefill kernels with diagonal and recompute fuse, [kernel] opt moe align block kernel by block/warp scan algorithm, Lora MoE Align Improvements, [Performance] Tune Mamba selective scan kernel for B200, [Perf][Kernel] Persistent TopK scheduler: unified CUDAGraph-safe kernel with dynamic per-row dispatch - DeepSeek-V3.2 DSA decode
sort [None][feat] Add support for expert_number<=2048 and K<=32, [TRTLLM-10407][feat] Integrate CuTE DSL top-k kernel for Blackwell, [TRTLLM-10407][perf] Enable CuteDSL indexer_top_k in model, [TRTLLM-10407][perf] Add cute dsl single pass multi cta cluster topk, [None][feat] Temporally-Correlated Heuristic-guided Indexer TopK for Sparse Attention, [None][perf] Scheme X L2-aware dispatcher and PDL launchers for sparse-attention GVR Top-K, [None][feat] Indexer topk opt, [None][feat] DSv4: enable GVR Heuristic Top-K for compress_ratio=4, Radix-selection based BlockTopK specialization, Optimized Device-to-Device Tensor Copy (cudax), Implement the new tuning API for DispatchSegmentedRadixSort, Implement the new tuning API for DispatchSegmentedSort, Use the new tuning API for detail::radix_sort::dispatch, Adds support for non-fundamental types via decomposer to DeviceTopK , Optimized Device-to-Device Tensor Copy (cudax) - Transpose Case, Replace detail::merge_sort::dispatch by CUB's public API, Add sorting and head swizzle to varlen scheduler, Feature/sm100 low latency nvfp4 kernels, Support Kimi-K2 for TRT: templatize number of experts, perf: improve sampling/mask/softmax performance (part 1/2), perf: bunch of features and optimizations for top-k (sampling + sparse attention), feat: further optimize top-k and add fused top-k page construction kernels for DSA, bugfix: fix multi-cta top-k implementation when k value is different for different row, [bugfix] Fix FilteredTopK overflow correctness, feat: implement deterministic topk, [feat] Add 2048 experts and 32 Top K , [feat] Add air top-p algorithm, feat: Add b12x CuTe DSL fused MoE for SM120, feat: Add FP4 (E2M1) KV Cache Support with Quantization Utilities for MLA, [sgl-kernel] Optimize concat_mla_k kernel, support cutlass fp4 kernel in sm120, [DeepSeek v3.2] opt Context Parallelism: support fused moe, multi batch and fp8 kvcache, [LoRA][III] Add LoRA support for MoE layers and enable TP, [bug fix] fix ima with get_mla_kv_buffer_kernel overflow, [jit-kernel] Add CuTe DSL GDN Decode Kernel, [AMD] Tilelang sparse fwd for dsv32 mi355/mi300, [GDN] Fuse GDN kkt + solve_tril into one kernel, [AMD] Enable FP8 KV cache and FP8 attention kernel for NSA on MI300/MI355 with TileLang backend, [XPU] Enable qwen3.5 on XPU, Amd/deepseek v4 rebase main 0509, [rebase]Deepseek_v4 support w4(mxfp4)a16 on hopper, Apply sgl w8a8 fp8 kernel, Support MHA with chunked prefix cache for DeepSeek chunked prefill, [perf] introduce deep gemm group_gemm_masked as bmm, Fix bug of deepseek-v3 under DP+EP mode with large batchsize/seqlen, feat: integrate deepgemm into EPMoE, feat: support DeepSeek-R1-W4AFP8 model with ep-moe mode, [Kernel] Add fused grouped_topk kernel for MoE, [Kernel] Optimize grouped topk kernel, [Perf][Kernel] Persistent TopK scheduler: unified CUDAGraph-safe kernel with dynamic per-row dispatch - DeepSeek-V3.2 DSA decode, fix: clamp NaN/Inf in topk_softmax to prevent duplicate expert IDs, [Bugfix] moe lora align kernel grid
sparse-attention FlashMLA upstream README, DeepSeek-V3.2-Exp in vLLM: Fine-Grained Sparse Attention in Action, FlashInfer MLSys 2026 - Track B: Sparse Attention, CUTLASS Changelog: SM100/Blackwell Entries, Native Sparse Attention (NSA), [TRTLLM-11092][feat] add support for visual gen FA4 attention backend, [https://nvbugs/5983390][perf] Kernel fusions in _gather_k_cache_for_chunk of Indexer in DSA, [None][feat] Temporally-Correlated Heuristic-guided Indexer TopK for Sparse Attention, [None][feat] Support sparse mqa/gqa attention, [https://nvbugs/5983390][perf] Multiple host perf optimizations for DSA part, [TRTLLM-34871][feat] Add cute dsl FP8 paged MQA logits decode kernel, [None][feat] Integrate FP4 indexer for DSA on Blackwell, [None][perf] Scheme X L2-aware dispatcher and PDL launchers for sparse-attention GVR Top-K, [None][feat] Add DeepSeekV4 attention kernels, [None][feat] DSv4: enable GVR Heuristic Top-K for compress_ratio=4, [TRTLLM-8535][feat] Support DeepSeek V3.2 with FP8 + BF16 KV cache/NVFP4 + BF16 KV cache, [Cute] Block sparse support Sm100, refactor: update fa3 codebase and fix hopper unittest [part 1], perf: bunch of features and optimizations for top-k (sampling + sparse attention), feat: add trtllm-gen per-tensor sparseMla kernels., Add dynamic tokens-per-page TRTLLM-GEN GQA kernels, [diffusion] model: support TurboWan2.1-T2V-1.3B/14B SLA, Move fa4 from sgl-kernel to jit kernel, Kernel: optimize decoding metadata in NSA multi-spec backend with fused kernels, [Refactor] Rename NSA → DSA: user-facing aliases, file/class/import rename, FlashAttention SM100 MLA TopK Sparse Forward, FlashMLA attention kernels, Native Sparse Attention (NSA), Sparse MLA
topk [Public release 26/04] Introducing Mega MoE, FP4 Indexer and other features/fixes, [None][perf] Add more optimization options for MOE CuteDSL finalized kernel, [TRTLLM-9831][perf] Use TMA.RED to improve effective memory bandwidth, [None][feat] fuse shared to sparse experts in TRT-LLM Gen MoE, [None][feat] Add support for expert_number<=2048 and K<=32, [TRTLLM-10407][feat] Integrate CuTE DSL top-k kernel for Blackwell, [TRTLLM-11540][feat] Add EAGLE3 dynamic tree speculative decoding support, [TRTLLM-10407][perf] Enable CuteDSL indexer_top_k in model, [TRTLLM-10407][perf] Add cute dsl single pass multi cta cluster topk, [None][feat] Temporally-Correlated Heuristic-guided Indexer TopK for Sparse Attention, [None][perf] add Dynamic SMEM block routing in MOE, [None][feat] Add PDL support to CuTE DSL top-k kernels, [None][feat] Optimize mamba SSD prefill and extend flashinfer dispatch, [None][feat] Integrate FP4 indexer for DSA on Blackwell, [None][perf] Extend customMoeRouting kernel to support Qwen3.5, [None][perf] Scheme X L2-aware dispatcher and PDL launchers for sparse-attention GVR Top-K, [None][feat] Add DeepSeekV4 attention kernels, [None][feat] Indexer topk opt, [None][feat] DSv4: enable GVR Heuristic Top-K for compress_ratio=4, [None][chore] Fix kernel launch param and add TRTLLM MoE backend test, [None][fix] Fix and add test for TRTLLM MoE backend, [TRTLLM-8637][feat] Optimize the routing kernel for DeepseekV3 (MoE CUTLASS backend); Add support for 384 experts (MoE TRTLLM backend), [None][fix] support topk autotuner input for expert slot per group larger than 32, [None][feat] TRT-LLM Gen MoE optimize DeepSeek Fp8 activation kernel, [TRTLLM-9685] [feat] Add gather fc1 kernel by cuteDSL, Fix the vectorized loading of BlockLoad, Radix-selection based BlockTopK specialization, Implement the new tuning API for DispatchTopK, Adds support for non-fundamental types via decomposer to DeviceTopK , Use the new tuning API internally for detail::topk::dispatch, [Cute,Sm100,Fwd] add MLA 64/512 with topk sparsity for MQA 128 heads, Add CUTLASS fused moe kernels from TensorRT-LLM., bugfix: softmax NaN results caused by large -inf masks, feat: trtllm-gen fp8 moe kernels, Feature/sm100 low latency nvfp4 kernels, GPT-OSS Support: Add Blackwell MoE mxfp4 implementation from TRTLLM and Attention Sink, Add GeGLU support to trtllm-gen NVFP4 Fused MoE Kernel, bugfix: Fix arg passing to TORCH_CHECK and TORCH_WARN macros, bugfix: fix the register overflow issue for topk renorm kernels on blackwell, Support Kimi-K2 for TRT: templatize number of experts, bugfix: partially fix tests/test_trtllm_gen_fused_moe.py unit test failure, Update the routing for TRTLLMGEN to support kimi k2 and qwen, feat: Add FP4 TRTLLM-Gen throughput MOE batched gemms, Add layernorm op for inputs of mixed dtype, perf: improve sampling/mask/softmax performance (part 1/2), perf: TRT-LLM MoE Block-FP8 activation optimization, perf: TRT-LLM Gen finalize kernel optimization, perf: bunch of features and optimizations for top-k (sampling + sparse attention), feat: add trtllm-gen per-tensor sparseMla kernels., feat: further optimize top-k and add fused top-k page construction kernels for DSA, feat: Support Fused MoE non gated Relu2 NVFP4 & FP8 and support Nemotron, Fix: FilteredTopKUnifiedKernel read value out of length, bugfix: fix multi-cta top-k implementation when k value is different for different row, feat: Support Fused MoE non gated Relu2 NVFP4 & FP8 and support Nemotron, fixed, Implement cutlass_fused_moe mxfp8, [bugfix] Fix FilteredTopK overflow correctness, fix: cute dsl nvfp4 moe routing index error, [fp8_blockwise]Fix int32 overflow in TRTLLM fused MoE activation kernel, feat: implement deterministic topk, [feat] Add 2048 experts and 32 Top K , [feat] Add air top-p algorithm, [Fmha] Sparse MLA decode kernel selection heuristics, feat: Add CuTe DSL grouped-gemm + combine fusion support, fix: use float instead of double in sampling binary search to avoid FP64 bottleneck on SM103, [feat] Add routing_replay_out support to MoE kernels and Python API, feat: Add b12x CuTe DSL fused MoE for SM120, fix(sm12x): fix micro-kernel workspace sizing when routed_rows > num_local_experts, feat(moe): add SM120 W4A16 b12x kernels, perf: dual pivot top-p/top-k renorm, [DeepseekV32] Enable flashmla_prefill kernel with fp8 kvcache, Support moe topk sigmoid kernel, [kernel][moe] add moe topk fast, Opt moe align block size kernel, [sgl-kernel][1/2] Fused qk_norm_rope for GLM4.6, Fix warp illegal instruction in kimi k2 thinking PCG, [diffusion] model: support TurboWan2.1-T2V-1.3B/14B SLA, Kernel: optimize decoding metadata in NSA multi-spec backend with fused kernels, Tilelang sparse decode fwd for dsv32 mi355, [AMD] Tilelang sparse fwd for dsv32 mi355/mi300, CUTLASS FP8 Blockwise GEMM improvement of SM120, CUTLASS NVFP4 GEMM improvement of SM120, [GDN] Fuse GDN kkt + solve_tril into one kernel, [AMD] Enable FP8 KV cache and FP8 attention kernel for NSA on MI300/MI355 with TileLang backend, Deepseek_v4 support w4(mxfp4)a16 on hopper, [feat] Init true on policy with qwen_dense, Port MXFP4 Marlin MoE support to JIT kernel path, [rebase]Deepseek_v4 support w4(mxfp4)a16 on hopper, [Intel GPU] Enable DeepSeek V4 Inference on XPU, amd/deepseek_v4 27/N [fix] Reduce Triton autotune configs for faster first-time server launch, add tensorrt_llm common and cutlass_extensions as 3rdparty, Add deepseek style fused moe group gate selection kernel, reduce moe_align_block_size_kernel small batch mode overhead, [2/2] Add python wrapper for CUTLASS FP8 Blockscale MoE Kernel. , [1/2] Add Kernel support for Cutlass based Fused FP4 MoE, [EP] Add cuda kernel for moe_ep_pre_reorder, Set num_fused_shared_experts as num_shared_experts when shared_experts fusion is not disabled, feat: integrate deepgemm into EPMoE, [EP] Add cuda kernel for moe_ep_post_reorder, fix ep_moe_reorder kernel bugs, Add a CUDA kernel for fusing mapping and weighted sum for MoE., Fuse sorted_token_ids padding to moe_align_block_size kernel, fix: fix apply_shuffle_mul_sum, feat: support DeepSeek-R1-W4AFP8 model with ep-moe mode, [1/n]: add cutlass W4A8 moe kernel for hopper architecture, [kernel] opt moe align block kernel by block/warp scan algorithm, optimize: reduce shulffle and quantization overhead in cutlass_moe sm90, [Kernel] add triton fused moe kernel for gptq/awq, [Kernel] CUTLASS grouped gemm fp8 MoE kernel, [Kernel] moe wna16 marlin kernel, permute/unpermute kernel for moe optimization, Modularize fused experts and integrate PPLX kernels, [Kernel] GGUF MoeVec kernel, [Hardware][NVIDIA] FP4 MoE kernel optimization, [Hardware][NVIDIA][kernel] Fp4 MOE quant kernel optimization, [Performance] Performance improvements in non-blockwise fp8 CUTLASS MoE, [Kernel] DeepGemm MoE : Integrate triton permute / unpermute kernels , [Perf] Add swap_ab to SM90 FP8 non-block CUTLASS moe grouped gemm, [Kernel] CUTLASS MoE FP8: Integrate cuda moe permute/unpermute, [Kernel] Add fused grouped_topk kernel for MoE, [Perf][DeepSeek] Add sigmoid+bias fusion to fused_grouped_topk from TRTLLM, Lora MoE Align Improvements, Add unpermute-aware fused MoE path and small-batch fallback, [Kernel][MoE] optimize moe_align_block_size, [Kernel] Add topk_sigmoid kernel, [Kernel] Optimize grouped topk kernel, [Quantization] add humming quantization kernel, [Perf][Kernel] Persistent TopK scheduler: unified CUDAGraph-safe kernel with dynamic per-row dispatch - DeepSeek-V3.2 DSA decode, [Refactor] Improve indexer decode path metadata preparation, fix: clamp NaN/Inf in topk_softmax to prevent duplicate expert IDs, [Kernel] Pack topk id/weights triton kernel, FlashAttention SM100 MLA TopK Sparse Forward, TensorRT-LLM Blackwell FP4 DSA Indexer