Skip to content

Repository files navigation

Transformer Engine

PyPI release Documentation License

Quick start | User guide | PyTorch API | JAX API | Examples | Releases

What is Transformer Engine?

Transformer Engine (TE) is an NVIDIA library for accelerating Transformer model training on NVIDIA GPUs. It combines optimized building blocks and fused kernels with automatic mixed-precision-style APIs for PyTorch and JAX, so low-precision training can be adopted without rewriting a training stack.

Transformer Engine manages the scaling factors, amax histories, and quantization metadata required by low-precision recipes. Its modules cover attention, linear layers, normalization, Mixture-of-Experts (MoE), and communication operations used in large-scale distributed training.

Highlights

  • FP8 training on NVIDIA Hopper, Ada, Blackwell, and Rubin GPUs.
  • MXFP8 and NVFP4 training on NVIDIA Blackwell GPUs.
  • Optimized attention, GEMM, normalization, quantization, and fused Transformer and MoE modules.
  • PyTorch and JAX APIs with autocast-style contexts and configurable low-precision recipes.
  • Support for tensor, sequence, context, and EP, including communication overlap.
  • FP16 and BF16 optimizations on NVIDIA Ampere architecture GPUs and later.

News

See the project updates archive for earlier news.

Quick start

Install the latest stable release for your framework:

# PyTorch
pip install --no-build-isolation "transformer_engine[pytorch]"

# JAX
pip install --no-build-isolation "transformer_engine[jax]"

For a ready-to-run environment, use an NVIDIA NGC framework container. Replace <YY.MM> with a container release listed in the NVIDIA Deep Learning Frameworks Support Matrix.

docker run --gpus all -it --rm nvcr.io/nvidia/pytorch:<YY.MM>-py3
docker run --gpus all -it --rm nvcr.io/nvidia/jax:<YY.MM>-py3

Continue with the PyTorch and JAX getting started guide. For prerequisites, source builds, environment variables, and troubleshooting, see the installation guide.

Integrations

Transformer Engine has been integrated with popular LLM frameworks such as:

See Ecosystem and historical integrations for additional community integrations and projects that have worked with Transformer Engine.

Contributing

We welcome contributions to Transformer Engine! To contribute to Transformer Engine and make pull requests, follow the guidelines outlined in the CONTRIBUTING.rst guide.

Technical deep dives

See Resources for the complete collection of papers and recorded talks.

About

A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.

Topics

Resources

Contributing

Security policy

Stars

3.6k stars

Watchers

37 watching

Forks

Releases

Packages

Used by

Contributors

Languages