Skip to content

Latest commit

 

History

History
98 lines (76 loc) · 5.71 KB

File metadata and controls

98 lines (76 loc) · 5.71 KB

Blackwell Kernel Optimization Knowledge Base

Comprehensive knowledge base for GPU kernel optimization on NVIDIA Blackwell (SM100) and Hopper (SM90). Optimized for LLM agent retrieval. See CLAUDE.md for schema and conventions. For Claude Code agents: this repository is a Claude Code skill — see SKILL.md.

Recommended Query Tools (for LLM agents)

python3 scripts/query.py "<natural language>" [--tag <t>] [--type <kernel|technique|pr|...>]
python3 scripts/get_page.py <page-id-or-path> [--follow-sources]
python3 scripts/grep_wiki.py "<regex>" [--only wiki|sources]

See references/examples.md for 10 worked query patterns.

Quick Navigation

I want to... Go to
Browse exact, family-only, or unknown architecture evidence queries/by-architecture.md
Fix a performance problem queries/by-problem.md
Learn a specific technique queries/by-technique.md
Use a hardware feature queries/by-hardware-feature.md
See what a repo contributed queries/by-repo.md
Write a specific kernel type queries/by-kernel-type.md
Use a specific language/DSL queries/by-language.md

Hardware Features

  • hw-tcgen05-mma — Blackwell MMA instruction (replaces wgmma)
  • hw-tmem — Tensor Memory (CTA-visible 128-lane × 512-column view)
  • hw-clc — Cluster Launch Control (dynamic tile scheduling)
  • hw-tma — Tensor Memory Accelerator (async bulk loads)
  • hw-2sm-cooperative — Two-SM cooperative MMA
  • hw-nvfp4 — NVFP4 and block-scaled narrow precision
  • hw-pdl-gdc — Programmatic Dependent Launch / Grid Dependency Control

Optimization Techniques

Kernel Case Studies

Problem → Solution Patterns

Languages & DSLs

Migration Guides

Source Repositories

Repository Focus
NVIDIA/cutlass CUTLASS 4.x Blackwell support
sgl-project/sglang SGLang Blackwell integration
vllm-project/vllm vLLM Blackwell support
flashinfer-ai/flashinfer FlashInfer Blackwell kernels
pytorch/pytorch PyTorch/Inductor Blackwell

Competitions