This page archives Transformer Engine releases, technical articles, research, and community news. The README highlights only the most recent updates.
- [09/2026] Transformer Engine v2.19 adds Rubin support, hybrid quantization, MXFP8 EP communication, and expanded FP8 attention support.
- [09/2026] Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine.
- [08/2026] Transformer Engine v2.18.
- [07/2026] Transformer Engine v2.17.
- [06/2026] Boosting MoE Training Throughput with Advanced Fusion Kernels.
- [06/2026] Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning.
- [06/2026] Train Models Faster with JAX and MaxText Using NVFP4 on NVIDIA Blackwell.
- [04/2026] Run High-Throughput Reinforcement Learning Training with End-to-End FP8 Precision.
- [02/2026] Using NVFP4 Low-Precision Model Training for Higher Throughput Without Losing Accuracy.
- [12/2025] NVIDIA Nemotron 3: Efficient and Open Intelligence — trained with NVFP4 on Transformer Engine.
- [11/2025] NVIDIA Blackwell Architecture Sweeps MLPerf Training v5.1 Benchmarks.
- [11/2025] Scale Biology Transformer Models with PyTorch and NVIDIA BioNeMo Recipes.
- [11/2025] FP8 Training of Large-Scale RL Models.
- [09/2025] Pretraining Large Language Models with NVFP4.
- [09/2025] Native FP8 Mixed Precision Training for Ling 2.0, Open Sourced!.
- [09/2025] Faster Training Throughput in FP8 Precision with NVIDIA NeMo.
- [08/2025] How We Built DeepL's Next-Generation LLMs with FP8 for Training and Inference.
- [08/2025] NVFP4 Trains with Precision of 16-Bit and Speed and Efficiency of 4-Bit.
- [06/2025] Floating Point 8: An Introduction to Efficient, Lower-Precision AI Training.
- [05/2025] Advanced Optimization Strategies for LLM Training on NVIDIA Grace Hopper.
- [03/2025] Stable and Scalable FP8 Deep Learning Training on Blackwell.
- [03/2025] Measure and Improve AI Workload Performance with NVIDIA DGX Cloud Benchmarking.
- [02/2025] Understanding the Language of Life's Biomolecules Across Evolution at a New Scale with Evo 2.
- [02/2025] NVIDIA DGX Cloud Introduces Ready-To-Use Templates to Benchmark AI Platform Performance.
- [01/2025] Continued Pretraining of State-of-the-Art LLMs for Sovereign AI and Regulated Industries with iGenius and NVIDIA DGX Cloud.
- [11/2024] Developing a 172B LLM with Strong Japanese Capabilities Using NVIDIA Megatron-LM.
- [11/2024] How FP8 Boosts LLM Training by 18% on Amazon SageMaker P5 Instances.
- [11/2024] Efficiently Train Models with Large Sequence Lengths Using Amazon SageMaker Model Parallel.
- [09/2024] Reducing AI Large Model Training Costs by 30% Requires Just a Single Line of Code from FP8 Mixed Precision Training Upgrades.
- [05/2024] Accelerating Transformers with NVIDIA cuDNN 9.
- [03/2024] Turbocharged Training: Optimizing the Databricks Mosaic AI Stack with FP8.
- [03/2024] FP8 Training Support in SageMaker Model Parallelism Library.
- [12/2023] New NVIDIA NeMo Framework Features and NVIDIA H200.
- [11/2023] Inflection-2: The Next Step Up.
- [11/2023] Unleashing the Power of Transformers with NVIDIA Transformer Engine.
- [11/2023] Accelerating PyTorch Training Workloads with FP8.
- [09/2023] Transformer Engine Added to the AWS Deep Learning Container for PyTorch Training.
- [06/2023] Breaking MLPerf Training Records with NVIDIA H100 GPUs.
- [04/2023] Benchmarking Large Language Models on NVIDIA H100 GPUs with CoreWeave.