$npx skillfedfor your agent

transformer-engine-cu12

Transformer acceleration library

With conditionsPyPI Artificial IntelligenceReleased Aug 202695.2K downloads / moPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — transformer_engine_cu12-2.18.0-py3-none-manylinux_2_28_aarch64.whl · transformer_engine_cu12-2.18.0-py3-none-manylinux_2_28_x86_64.whl
v2.18.0 · released 2026-08-11 · Python >=3.10.0 · 3 runtime deps: pydantic, packaging, importlib-metadata

Yes, if you are training or serving Transformer models on supported NVIDIA GPUs (Ampere or newer) and have the required CUDA/cuDNN stack. The library is actively maintained, has no known vulnerabilities, and offers significant performance and memory benefits through low-precision training. However, verify license terms before commercial use and ensure your system meets the strict hardware and software prerequisites (CUDA 12.1+, cuDNN 9.3+, C++17 compiler).AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires NVIDIA GPU (Hopper, Ada, Ampere, or Blackwell), CUDA 12.1+, cuDNN 9.3+, and Linux or WSL2.
  • Compiled wheels are architecture-specific; verify your system matches manylinux_2_28 with x86_64 or aarch64.
  • Medium install friction due to compiled wheels for specific GPU architectures (aarch64, x86_64) and CUDA 12 requirement.

License · maintenance · safety

(unclear) — License treatment is unclear—no SPDX identifier or raw license text provided in metadata. Verify licensing terms before use, particularly for commercial deployment.

last release 2026-08-11 (3 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 95,170 downloads/mo, #13,288 on PyPI

Verify before relying

pip install transformer-engine-cu12

import transformer_engine.pytorch as te
from transformer_engine.common import recipe

model = te.Linear(768, 3072, bias=True)
fp8_recipe = recipe.DelayedScaling(margin=0, fp8_format=recipe.Format.E4M3)

with te.autocast(enabled=True, recipe=fp8_recipe):
    out = model(inp)
  • Exact performance gains (speedup percentages, memory savings) on different GPU architectures and model sizes
  • Accuracy degradation (if any) when using FP8, MXFP8, or NVFP4 formats versus standard precision
  • Compatibility with specific framework versions beyond the stated Python 3.10+ requirement
  • Support status and maintenance timeline for older GPU architectures (Ampere) versus newer ones (Blackwell)
Same gist for agents: .md · .json

What it is and what it does

Transformer Engine is a library that accelerates Transformer model training and inference on NVIDIA GPUs by providing optimized kernels and low-precision arithmetic support. It enables 8-bit floating-point (FP8) training on Hopper, Ada, and Ampere GPUs, and adds support for MXFP8 and NVFP4 formats on Blackwell GPUs. The library handles scaling factors and precision management internally, allowing developers to use a simple autocast API similar to mixed-precision training frameworks.

The package integrates with popular frameworks through framework-specific modules and provides a C++ API for integration with other deep learning libraries. It includes fused operations, support for distributed training patterns (tensor/sequence/context parallelism), and Mixture-of-Experts (MoE) optimizations. Installation requires a compatible NVIDIA GPU, CUDA 12.1+, cuDNN 9.3+, and a C++ compiler with C++17 support; wheels are pre-compiled for specific architectures (x86_64, aarch64) on manylinux_2_28.

Use it for

  • Train large language models with FP8 precision to reduce memory footprint and increase throughput without accuracy loss
  • Accelerate Mixture-of-Experts (MoE) model training using fused kernels and low-precision formats
  • Run inference on Transformer models with reduced latency and memory using optimized GPU kernels
  • Implement mixed-precision training workflows with automatic scaling factor management
  • Deploy multimodal or biology-focused Transformer models on Blackwell GPUs using NVFP4 for maximum efficiency

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are training or serving Transformer models on supported NVIDIA GPUs (Ampere or newer) and have the required CUDA/cuDNN stack.

The library is actively maintained, has no known vulnerabilities, and offers significant performance and memory benefits through low-precision training. However, verify license terms before commercial use and ensure your system meets the strict hardware and software prerequisites (CUDA 12.1+, cuDNN 9.3+, C++17 compiler).

Install

transformer-engine-cu12 on PyPI

Before you install

Medium install friction due to compiled wheels for specific GPU architectures (aarch64, x86_64) and CUDA 12 requirement. Active maintenance with recent release (3 days old). Requires Python 3.10+, CUDA 12.1+ (or 12.8+ for Blackwell), cuDNN 9.3+, and GCC 9+ or Clang 10+ with C++17 support.

Requires NVIDIA GPU (Hopper, Ada, Ampere, or Blackwell), CUDA 12.1+, cuDNN 9.3+, and Linux or WSL2. Compiled wheels are architecture-specific; verify your system matches manylinux_2_28 with x86_64 or aarch64.

License in practice

License treatment is unclear—no SPDX identifier or raw license text provided in metadata. Verify licensing terms before use, particularly for commercial deployment.

Quickstart

pip install transformer-engine-cu12

import transformer_engine.pytorch as te
from transformer_engine.common import recipe

model = te.Linear(768, 3072, bias=True)
fp8_recipe = recipe.DelayedScaling(margin=0, fp8_format=recipe.Format.E4M3)

with te.autocast(enabled=True, recipe=fp8_recipe):
    out = model(inp)

Verify before relying

  • Exact performance gains (speedup percentages, memory savings) on different GPU architectures and model sizes
  • Accuracy degradation (if any) when using FP8, MXFP8, or NVFP4 formats versus standard precision
  • Compatibility with specific framework versions beyond the stated Python 3.10+ requirement
  • Support status and maintenance timeline for older GPU architectures (Ampere) versus newer ones (Blackwell)

Package facts

LicenseNot declared unclear
Python supportSupports the current Python release >=3.10.0
Install frictionMedium. Platform-specific wheel
Runtime dependencies
3 packages
pydanticpackagingimportlib-metadata
MaintenanceActively maintained 3 days since the last release
First released
Downloads95,170 / month, #13,288 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Programming Language :: Python :: 3

Evidence: transformer_engine_cu12-2.18.0-py3-none-manylinux_2_28_aarch64.whl; transformer_engine_cu12-2.18.0-py3-none-manylinux_2_28_x86_64.whl

Tags

Capabilities
transformer training accelerationFP8 mixed precision trainingNVIDIA GPU optimizationlow-precision model trainingtransformer inference optimizationlarge language model accelerationCUDA kernel fusion
Topics
gpu-accelerationlow-precision-trainingtransformer-models

Let your AI agent find packages like this

Example. Real query, live index.

An agent finds packages by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language. Give your agent the search over MCP.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also megatron-core · nvidia-cudnn-frontend · nvidia-modelopt · transformer-engine · transformer-engine-cu13 · nvdlfw-inspect · accelforge · liger-kernel · torch-directml · deepspeed

Further reading