- Attention Is All You Need
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Reducing Activation Recomputation in Large Transformer Models
- FP8 Formats for Deep Learning
- Stable and Scalable FP8 Deep Learning Training on Blackwell | GTC 2025
- Blackwell Numerics for AI | GTC 2025
- Building LLMs: Accelerating Pretraining of Foundational Models with FP8 Precision | GTC 2025
- From FP8 LLM Training to Inference: Language AI at Scale | GTC 2025
- What's New in Transformer Engine and FP8 Training | GTC 2024
- FP8 Training with Transformer Engine | GTC 2023
- FP8 for Deep Learning | GTC 2023
- Inside the Hopper Architecture | GTC 2022