transformer-engine-cu13
Transformer acceleration library
Decision gist · record as of 2026-08-14
Yes, if you are training or deploying Transformer models on supported NVIDIA GPUs (Hopper, Ada, Ampere, or Blackwell) and want to reduce memory and compute cost. Verify the unclear license terms and ensure your CUDA/cuDNN versions meet the minimum requirements (12.1+ CUDA, 9.3+ cuDNN). No known vulnerabilities. Active maintenance and recent release are positive signals.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires CUDA 12.1+ (12.8+ for Blackwell), cuDNN 9.3+, and an NVIDIA GPU (Hopper, Ada, Ampere, or Blackwell).
- Python 3.10+ required.
- Medium install friction due to platform-specific wheels (aarch64 and x86_64 manylinux) and CUDA 12.1+ requirement.
License · maintenance · safety
(unclear) — License treatment is unclear—no SPDX identifier or raw license text provided. Verify licensing terms before production use, particularly if integrating into proprietary systems.
last release 2026-08-11 (3 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 81,041 downloads/mo, #14,255 on PyPI
Alternatives
Verify before relying
pip install transformer-engine-cu13
import transformer_engine.pytorch as te
from transformer_engine.common import recipe
model = te.Linear(768, 3072, bias=True)
fp8_recipe = recipe.DelayedScaling(margin=0, fp8_format=recipe.Format.E4M3)
with te.autocast(enabled=True, recipe=fp8_recipe):
out = model(inp)- Whether the package includes pre-built binaries for all target architectures or requires compilation
- Specific NVIDIA GPU models and CUDA versions supported beyond the general hardware list
- Whether cuDNN 9.3+ is a hard runtime requirement or only needed at build time
- Performance benchmarks or throughput gains claimed in the documentation
What it is and what it does
Transformer Engine is a library that speeds up Transformer model training and inference on NVIDIA GPUs by using lower-precision arithmetic (FP8, MXFP8, NVFP4) instead of standard floating-point formats, reducing memory use and computation time. It provides modules for PyTorch and JAX/Flax that integrate FP8 support directly into Transformer layers, handling the scaling factors and precision conversions automatically. The library includes optimized C++ kernels and fused operations for common Transformer patterns, and works with advanced features like mixture-of-experts, tensor parallelism, and sequence parallelism.
The package targets researchers and engineers training or deploying large language models and multimodal Transformers on Hopper, Ada, Ampere, and Blackwell GPU architectures. It depends on packaging, pydantic, and importlib-metadata, and requires CUDA 12.1+ (12.8+ for Blackwell), cuDNN 9.3+, and Python 3.10+. Installation uses platform-specific wheels for x86_64 and aarch64 Linux.
Use it for
- Train large language models with FP8 precision to reduce memory footprint and training time on supported GPUs.
- Integrate low-precision inference into production pipelines to serve models faster with lower latency.
- Build mixture-of-experts models using optimized kernels for efficient sparse computation.
- Experiment with NVFP4 format on Blackwell GPUs to achieve efficiency gains while maintaining precision.
- Combine FP8 training with tensor or sequence parallelism for distributed training of large models.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are training or deploying Transformer models on supported NVIDIA GPUs (Hopper, Ada, Ampere, or Blackwell) and want to reduce memory and compute cost.
Verify the unclear license terms and ensure your CUDA/cuDNN versions meet the minimum requirements (12.1+ CUDA, 9.3+ cuDNN). No known vulnerabilities. Active maintenance and recent release are positive signals.
Install
transformer-engine-cu13 on PyPI
Before you install
Medium install friction due to platform-specific wheels (aarch64 and x86_64 manylinux) and CUDA 12.1+ requirement. Active maintenance with recent release (3 days old). Requires Python 3.10+.
Requires CUDA 12.1+ (12.8+ for Blackwell), cuDNN 9.3+, and an NVIDIA GPU (Hopper, Ada, Ampere, or Blackwell). Python 3.10+ required.
License in practice
License treatment is unclear—no SPDX identifier or raw license text provided. Verify licensing terms before production use, particularly if integrating into proprietary systems.
Quickstart
pip install transformer-engine-cu13
import transformer_engine.pytorch as te
from transformer_engine.common import recipe
model = te.Linear(768, 3072, bias=True)
fp8_recipe = recipe.DelayedScaling(margin=0, fp8_format=recipe.Format.E4M3)
with te.autocast(enabled=True, recipe=fp8_recipe):
out = model(inp)
Verify before relying
- Whether the package includes pre-built binaries for all target architectures or requires compilation
- Specific NVIDIA GPU models and CUDA versions supported beyond the general hardware list
- Whether cuDNN 9.3+ is a hard runtime requirement or only needed at build time
- Performance benchmarks or throughput gains claimed in the documentation
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.10.0 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 3 packagespackagingpydanticimportlib-metadata |
| Maintenance | Actively maintained 3 days since the last release |
| First released | |
| Downloads | 81,041 / month, #14,255 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Programming Language :: Python :: 3 |
Evidence: transformer_engine_cu13-2.18.0-py3-none-manylinux_2_28_aarch64.whl; transformer_engine_cu13-2.18.0-py3-none-manylinux_2_28_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “FP8 mixed precision training”
- transformer-engine-cu13Accelerates Transformer model training and inference on NVIDIA GPUs…
- transformer-engine-cu12Accelerates Transformer model training and inference on NVIDIA GPUs…
- transformer-engineTransformer Engine accelerates Transformer model training and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also transformer-engine · transformer-engine-cu12 · megatron-core · nvdlfw-inspect · accelforge · nvidia-modelopt · nvidia-cudnn-frontend · ctranslate2 · torch-directml · tensorrt-cu13-libs