transformer-engine
Transformer acceleration library
Decision gist · record as of 2026-08-14
Yes, if you have access to a compatible NVIDIA GPU (Hopper, Ada, Ampere, or Blackwell) and are training or deploying Transformer models where memory and compute efficiency matter. Low install friction, active maintenance, and zero runtime dependencies make adoption straightforward. Verify the license terms before use in proprietary projects, and confirm your GPU and CUDA versions meet the stated requirements.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires NVIDIA GPU (Hopper, Ada, Ampere, or Blackwell) with CUDA 12.1+ (12.8+ for Blackwell), cuDNN 9.3+, and GCC 9+ or Clang 10+ with C++17 support.
- Low install friction with a pure Python wheel distribution.
- Active maintenance with a release 3 days old as of the fact sheet date.
License · maintenance · safety
(unclear) — License treatment is unclear—the package description references a LICENSE file but the fact sheet provides no SPDX identifier or raw license text. Verify the actual license terms before use in proprietary or copyleft-sensitive contexts.
last release 2026-08-11 (3 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 147,015 downloads/mo, #11,084 on PyPI
Alternatives
Verify before relying
import transformer_engine.pytorch as te
from transformer_engine.common import recipe
model = te.Linear(768, 3072, bias=True)
fp8_recipe = recipe.DelayedScaling(margin=0, fp8_format=recipe.Format.E4M3)
with te.autocast(enabled=True, recipe=fp8_recipe):
out = model(inp)- Whether the license is open-source or proprietary—the fact sheet does not specify.
- Exact performance gains and memory savings for specific model sizes and GPU architectures.
- Compatibility with frameworks beyond PyTorch and JAX (e.g., TensorFlow).
- Whether the package works with Python versions below 3.10 despite requiring 3.10+.
What it is and what it does
Transformer Engine is a library that accelerates Transformer model training and inference on NVIDIA GPUs by providing optimized low-precision computation. It supports FP8 on Hopper, Ada, and Ampere GPUs, and extends to MXFP8 and NVFP4 formats on Blackwell GPUs. The library automatically manages scaling factors and precision conversion, allowing users to enable low-precision training through a simple autocast API.
The package provides Python modules for building Transformer layers with built-in FP8 support and fused kernels for common operations. It integrates with popular frameworks and is designed to work with advanced training techniques like tensor parallelism, sequence parallelism, and mixture-of-experts architectures. Installation requires a compatible NVIDIA GPU, CUDA 12.1 or later, cuDNN 9.3 or later, and a modern C++ compiler with C++17 support.
Use it for
- Train large language models with reduced memory footprint using low-precision formats on Hopper or Blackwell GPUs.
- Optimize inference latency for deployed Transformer models by leveraging low-precision computation and fused kernels.
- Build mixture-of-experts or other advanced Transformer architectures with automatic mixed-precision support.
- Integrate low-precision Transformer operations into custom deep learning frameworks via the C++ API.
- Experiment with NVFP4 or MXFP8 formats on Blackwell hardware for training efficiency.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you have access to a compatible NVIDIA GPU (Hopper, Ada, Ampere, or Blackwell) and are training or deploying Transformer models where memory and compute efficiency matter.
Low install friction, active maintenance, and zero runtime dependencies make adoption straightforward. Verify the license terms before use in proprietary projects, and confirm your GPU and CUDA versions meet the stated requirements.
Install
transformer-engine on PyPI
Before you install
Low install friction with a pure Python wheel distribution. Active maintenance with a release 3 days old as of the fact sheet date. No runtime dependencies to manage.
Requires NVIDIA GPU (Hopper, Ada, Ampere, or Blackwell) with CUDA 12.1+ (12.8+ for Blackwell), cuDNN 9.3+, and GCC 9+ or Clang 10+ with C++17 support.
License in practice
License treatment is unclear—the package description references a LICENSE file but the fact sheet provides no SPDX identifier or raw license text. Verify the actual license terms before use in proprietary or copyleft-sensitive contexts.
Quickstart
import transformer_engine.pytorch as te
from transformer_engine.common import recipe
model = te.Linear(768, 3072, bias=True)
fp8_recipe = recipe.DelayedScaling(margin=0, fp8_format=recipe.Format.E4M3)
with te.autocast(enabled=True, recipe=fp8_recipe):
out = model(inp)
Verify before relying
- Whether the license is open-source or proprietary—the fact sheet does not specify.
- Exact performance gains and memory savings for specific model sizes and GPU architectures.
- Compatibility with frameworks beyond PyTorch and JAX (e.g., TensorFlow).
- Whether the package works with Python versions below 3.10 despite requiring 3.10+.
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.10.0 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 3 days since the last release |
| First released | |
| Downloads | 147,015 / month, #11,084 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Programming Language :: Python :: 3 |
Evidence: transformer_engine-2.18.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “transformer training acceleration”
- transformer-engineTransformer Engine accelerates Transformer model training and…
- transformer-engine-cu12Accelerates Transformer model training and inference on NVIDIA GPUs…
- transformer-engine-cu13Accelerates Transformer model training and inference on NVIDIA GPUs…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also deepspeed · transformer-engine-cu12 · transformer-engine-cu13 · nvidia-modelopt · megatron-core · nvdlfw-inspect · accelforge · transformer-smaller-training-vocab · nvidia-cudnn-frontend · ctranslate2