$npx skillfedfor your agent

transformer-engine-cu13

Transformer acceleration library

With conditionsPyPI Artificial IntelligenceReleased Aug 202681.0K downloads / moPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — transformer_engine_cu13-2.18.0-py3-none-manylinux_2_28_aarch64.whl · transformer_engine_cu13-2.18.0-py3-none-manylinux_2_28_x86_64.whl
v2.18.0 · released 2026-08-11 · Python >=3.10.0 · 3 runtime deps: packaging, pydantic, importlib-metadata

Yes, if you are training or deploying Transformer models on supported NVIDIA GPUs (Hopper, Ada, Ampere, or Blackwell) and want to reduce memory and compute cost. Verify the unclear license terms and ensure your CUDA/cuDNN versions meet the minimum requirements (12.1+ CUDA, 9.3+ cuDNN). No known vulnerabilities. Active maintenance and recent release are positive signals.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires CUDA 12.1+ (12.8+ for Blackwell), cuDNN 9.3+, and an NVIDIA GPU (Hopper, Ada, Ampere, or Blackwell).
  • Python 3.10+ required.
  • Medium install friction due to platform-specific wheels (aarch64 and x86_64 manylinux) and CUDA 12.1+ requirement.

License · maintenance · safety

(unclear) — License treatment is unclear—no SPDX identifier or raw license text provided. Verify licensing terms before production use, particularly if integrating into proprietary systems.

last release 2026-08-11 (3 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 81,041 downloads/mo, #14,255 on PyPI

Verify before relying

pip install transformer-engine-cu13

import transformer_engine.pytorch as te
from transformer_engine.common import recipe

model = te.Linear(768, 3072, bias=True)
fp8_recipe = recipe.DelayedScaling(margin=0, fp8_format=recipe.Format.E4M3)

with te.autocast(enabled=True, recipe=fp8_recipe):
    out = model(inp)
  • Whether the package includes pre-built binaries for all target architectures or requires compilation
  • Specific NVIDIA GPU models and CUDA versions supported beyond the general hardware list
  • Whether cuDNN 9.3+ is a hard runtime requirement or only needed at build time
  • Performance benchmarks or throughput gains claimed in the documentation
Same gist for agents: .md · .json

What it is and what it does

Transformer Engine is a library that speeds up Transformer model training and inference on NVIDIA GPUs by using lower-precision arithmetic (FP8, MXFP8, NVFP4) instead of standard floating-point formats, reducing memory use and computation time. It provides modules for PyTorch and JAX/Flax that integrate FP8 support directly into Transformer layers, handling the scaling factors and precision conversions automatically. The library includes optimized C++ kernels and fused operations for common Transformer patterns, and works with advanced features like mixture-of-experts, tensor parallelism, and sequence parallelism.

The package targets researchers and engineers training or deploying large language models and multimodal Transformers on Hopper, Ada, Ampere, and Blackwell GPU architectures. It depends on packaging, pydantic, and importlib-metadata, and requires CUDA 12.1+ (12.8+ for Blackwell), cuDNN 9.3+, and Python 3.10+. Installation uses platform-specific wheels for x86_64 and aarch64 Linux.

Use it for

  • Train large language models with FP8 precision to reduce memory footprint and training time on supported GPUs.
  • Integrate low-precision inference into production pipelines to serve models faster with lower latency.
  • Build mixture-of-experts models using optimized kernels for efficient sparse computation.
  • Experiment with NVFP4 format on Blackwell GPUs to achieve efficiency gains while maintaining precision.
  • Combine FP8 training with tensor or sequence parallelism for distributed training of large models.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are training or deploying Transformer models on supported NVIDIA GPUs (Hopper, Ada, Ampere, or Blackwell) and want to reduce memory and compute cost.

Verify the unclear license terms and ensure your CUDA/cuDNN versions meet the minimum requirements (12.1+ CUDA, 9.3+ cuDNN). No known vulnerabilities. Active maintenance and recent release are positive signals.

Install

transformer-engine-cu13 on PyPI

Before you install

Medium install friction due to platform-specific wheels (aarch64 and x86_64 manylinux) and CUDA 12.1+ requirement. Active maintenance with recent release (3 days old). Requires Python 3.10+.

Requires CUDA 12.1+ (12.8+ for Blackwell), cuDNN 9.3+, and an NVIDIA GPU (Hopper, Ada, Ampere, or Blackwell). Python 3.10+ required.

License in practice

License treatment is unclear—no SPDX identifier or raw license text provided. Verify licensing terms before production use, particularly if integrating into proprietary systems.

Quickstart

pip install transformer-engine-cu13

import transformer_engine.pytorch as te
from transformer_engine.common import recipe

model = te.Linear(768, 3072, bias=True)
fp8_recipe = recipe.DelayedScaling(margin=0, fp8_format=recipe.Format.E4M3)

with te.autocast(enabled=True, recipe=fp8_recipe):
    out = model(inp)

Verify before relying

  • Whether the package includes pre-built binaries for all target architectures or requires compilation
  • Specific NVIDIA GPU models and CUDA versions supported beyond the general hardware list
  • Whether cuDNN 9.3+ is a hard runtime requirement or only needed at build time
  • Performance benchmarks or throughput gains claimed in the documentation

Package facts

LicenseNot declared unclear
Python supportSupports the current Python release >=3.10.0
Install frictionMedium. Platform-specific wheel
Runtime dependencies
3 packages
packagingpydanticimportlib-metadata
MaintenanceActively maintained 3 days since the last release
First released
Downloads81,041 / month, #14,255 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Programming Language :: Python :: 3

Evidence: transformer_engine_cu13-2.18.0-py3-none-manylinux_2_28_aarch64.whl; transformer_engine_cu13-2.18.0-py3-none-manylinux_2_28_x86_64.whl

Tags

Capabilities
transformer training accelerationFP8 mixed precision trainingNVIDIA GPU optimizationlow precision model trainingtransformer inference optimizationlarge language model accelerationGPU kernel fusion
Topics
gpu-accelerationmixed-precisionllm-training

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “FP8 mixed precision training”

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also transformer-engine · transformer-engine-cu12 · megatron-core · nvdlfw-inspect · accelforge · nvidia-modelopt · nvidia-cudnn-frontend · ctranslate2 · torch-directml · tensorrt-cu13-libs

Further reading