Comprehensive knowledge base for GPU kernel optimization on NVIDIA Blackwell (SM100) and Hopper (SM90). Optimized for LLM agent retrieval. See CLAUDE.md for schema and conventions. For Claude Code agents: this repository is a Claude Code skill — see SKILL.md.
python3 scripts/query.py "<natural language>" [--tag <t>] [--type <kernel|technique|pr|...>]
python3 scripts/get_page.py <page-id-or-path> [--follow-sources]
python3 scripts/grep_wiki.py "<regex>" [--only wiki|sources]See references/examples.md for 10 worked query patterns.
| I want to... | Go to |
|---|---|
| Browse exact, family-only, or unknown architecture evidence | queries/by-architecture.md |
| Fix a performance problem | queries/by-problem.md |
| Learn a specific technique | queries/by-technique.md |
| Use a hardware feature | queries/by-hardware-feature.md |
| See what a repo contributed | queries/by-repo.md |
| Write a specific kernel type | queries/by-kernel-type.md |
| Use a specific language/DSL | queries/by-language.md |
- hw-tcgen05-mma — Blackwell MMA instruction (replaces wgmma)
- hw-tmem — Tensor Memory (CTA-visible 128-lane × 512-column view)
- hw-clc — Cluster Launch Control (dynamic tile scheduling)
- hw-tma — Tensor Memory Accelerator (async bulk loads)
- hw-2sm-cooperative — Two-SM cooperative MMA
- hw-nvfp4 — NVFP4 and block-scaled narrow precision
- hw-pdl-gdc — Programmatic Dependent Launch / Grid Dependency Control
- technique-warp-specialization — Warp role assignment
- technique-persistent-kernels — Persistent kernel patterns with CLC
- technique-swizzling — Shared memory swizzling
- technique-pipeline-stages — Software pipelining
- technique-epilogue-fusion — Fusing epilogue with mainloop
- technique-tile-scheduling — Tile scheduling strategies
- technique-double-buffering — Double/multi-buffering
- technique-software-exp — Software-emulated exponential
- technique-fine-grained-quantization - Fine-grained FP8/FP4 quantization
- technique-vectorized-loads — Wide vectorized loads and cache policies
- kernel-flash-attention-4 — FlashAttention-4 (up to 1613 TFLOPS on B200 in the paper's benchmark sweep)
- kernel-deepgemm — DeepGEMM FP8 GEMM (1550 TFLOPS on H800)
- kernel-flashmla — FlashMLA sparse/dense MLA decoding
- kernel-nsa — Native Sparse Attention (9x fwd speedup)
- kernel-gated-delta-net — Gated Delta Net linear attention
- kernel-nvfp4-gemm — NVFP4 GEMM from GPU Mode hackathon
- kernel-nvfp4-gemv — NVFP4 batched GEMV optimization
- kernel-grouped-gemm — Grouped GEMM for MoE
- kernel-small-m-m-grouped-gemm — Small-M compact/ragged M-grouped GEMM optimization playbook
- kernel-fused-moe — Fused MoE with FP8
- pattern-low-sm-utilization — SM utilization is low
- pattern-memory-bound — Kernel is memory bandwidth limited
- pattern-register-pressure — Too many registers → low occupancy
- pattern-compute-bound — Not reaching peak FLOPS
- pattern-tail-effect — Last wave underutilizes GPU
- lang-cute-dsl — CuTe DSL for Blackwell
- lang-cuda-cpp — CUDA C++ with PTX inline
- lang-ptx — PTX instructions for SM100
- lang-triton — Triton on Blackwell
- migration-wgmma-to-tcgen05 — Hopper wgmma → Blackwell tcgen05
- migration-register-to-tmem — Register accumulators → TMEM
| Repository | Focus |
|---|---|
| NVIDIA/cutlass | CUTLASS 4.x Blackwell support |
| sgl-project/sglang | SGLang Blackwell integration |
| vllm-project/vllm | vLLM Blackwell support |
| flashinfer-ai/flashinfer | FlashInfer Blackwell kernels |
| pytorch/pytorch | PyTorch/Inductor Blackwell |
- GPU Mode NVFP4 Hackathon — 4 NVFP4 kernel challenges on B200
- FlashInfer MLSys 2026 — MoE, Sparse Attention, GatedDeltaNet