nvidia-cudnn-frontend
NVIDIA cuDNN Frontend — Python and C++ Graph API with SOTA attention (SDPA / Flash Attention), MoE grouped GEMM fusions, and FP8/MXFP8 kernels for Hopper and Blackwell GPUs.
Decision gist · record as of 2026-08-14
Yes. The package is actively maintained, dual-licensed under permissive terms, has no known vulnerabilities, and provides essential optimized kernels for modern GPU workloads. Install it if you are training or deploying deep learning models on NVIDIA Hopper or Blackwell GPUs and need high-performance attention, grouped GEMM, or quantized operations. The requirement for CUDA Toolkit and cuDNN 8.5.0+ is a prerequisite, not a drawback—it reflects the package's tight integration with NVIDIA's GPU stack.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires NVIDIA driver, CUDA Toolkit, and cuDNN 8.5.0 or later installed on the system; Python 3.9 or later.
- Installation is straightforward via pip with prebuilt wheels for Python 3.10–3.14 on Linux (x86_64 and aarch64) and Windows.
- The package is actively maintained with a recent release (8 days old) and no known vulnerabilities.
License · maintenance · safety
Apache-2.0 AND MIT (permissive) — Dual-licensed under Apache-2.0 and MIT, both permissive licenses. You are free to use, modify, and distribute the package in commercial and private projects with minimal restrictions.
last release 2026-08-06 (8 days) · last repo commit 2026-08-14 · 904 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 4,479,592 downloads/mo, #2,292 on PyPI
Alternatives
Verify before relying
pip install nvidia-cudnn-frontend
import cudnn
# Create a graph for scaled dot-product attention
graph = cudnn.pygraph.Graph()
# Configure and execute attention operations on Hopper/Blackwell GPUs- Whether the package's autotuning mechanism requires additional configuration or warm-up time for production workloads.
- Performance characteristics and memory overhead of the open-source FROST GEMM engine relative to backend-native plans.
- Compatibility and integration effort with existing PyTorch models beyond the stated torch.compile support.
What it is and what it does
nvidia-cudnn-frontend is NVIDIA's modern entry point to cuDNN, offering both a header-only C++ API and Python bindings that abstract the complexity of the cuDNN Graph API. It exposes state-of-the-art GPU kernels including scaled dot-product attention (SDPA/Flash Attention), grouped GEMM fusions for mixture-of-experts training, fused normalization and activation operations, and quantized matrix multiplication in FP8 and MXFP8 precision. The package targets NVIDIA's latest GPU architectures—Hopper (H100/H200) and Blackwell (B200/GB200/GB300)—and includes native PyTorch integration with torch.compile support.
The package ships with open-source kernel implementations (FROST GEMM engine, block-sparse attention, native sparse attention, and fused RMSNorm+SiLU) that developers can inspect, modify, and contribute to. Installation is simple via pip, though it requires NVIDIA driver, CUDA Toolkit, and cuDNN 8.5.0 or later on the system. The library is actively maintained with prebuilt wheels for modern Python versions and multiple architectures.
Use it for
- Accelerate transformer attention mechanisms in large language models using SDPA kernels optimized for Hopper and Blackwell GPUs.
- Train mixture-of-experts models efficiently with fused grouped GEMM operations that reduce memory bandwidth and kernel launch overhead.
- Deploy quantized deep learning models using FP8 and MXFP8 precision kernels for reduced memory footprint and faster inference.
- Build custom GPU kernels by inspecting and modifying open-source implementations like FROST GEMM and block-sparse attention.
- Integrate cuDNN-accelerated operations into PyTorch models with automatic differentiation and torch.compile support.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
The package is actively maintained, dual-licensed under permissive terms, has no known vulnerabilities, and provides essential optimized kernels for modern GPU workloads. Install it if you are training or deploying deep learning models on NVIDIA Hopper or Blackwell GPUs and need high-performance attention, grouped GEMM, or quantized operations. The requirement for CUDA Toolkit and cuDNN 8.5.0+ is a prerequisite, not a drawback—it reflects the package's tight integration with NVIDIA's GPU stack.
Install
nvidia-cudnn-frontend on PyPI
Before you install
Installation is straightforward via pip with prebuilt wheels for Python 3.10–3.14 on Linux (x86_64 and aarch64) and Windows. The package is actively maintained with a recent release (8 days old) and no known vulnerabilities. Medium install friction reflects the requirement for NVIDIA driver, CUDA Toolkit, and cuDNN 8.5.0 or later to be present on the system.
Requires NVIDIA driver, CUDA Toolkit, and cuDNN 8.5.0 or later installed on the system; Python 3.9 or later.
License in practice
Dual-licensed under Apache-2.0 and MIT, both permissive licenses. You are free to use, modify, and distribute the package in commercial and private projects with minimal restrictions.
Quickstart
pip install nvidia-cudnn-frontend
import cudnn
# Create a graph for scaled dot-product attention
graph = cudnn.pygraph.Graph()
# Configure and execute attention operations on Hopper/Blackwell GPUs
Verify before relying
- Whether the package's autotuning mechanism requires additional configuration or warm-up time for production workloads.
- Performance characteristics and memory overhead of the open-source FROST GEMM engine relative to backend-native plans.
- Compatibility and integration effort with existing PyTorch models beyond the stated torch.compile support.
Package facts
| License | Apache-2.0 AND MIT permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 8 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 4,479,592 / month, #2,292 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: GPU :: NVIDIA CUDAIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseLicense :: OSI Approved :: MIT LicenseOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: C++Programming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Software Development :: Libraries :: Python Modules |
Evidence: nvidia_cudnn_frontend-1.27.0-cp310-cp310-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; nvidia_cudnn_frontend-1.27.0-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; nvidia_cudnn_frontend-1.27.0-cp310-cp310-win_amd64.whl; nvidia_cudnn_frontend-1.27.0-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; nvidia_cudnn_frontend-1.27.0-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; nvidia_cudnn_frontend-1.27.0-cp311-cp311-win_amd64.whl; nvidia_cudnn_frontend-1.27.0-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; nvidia_cudnn_frontend-1.27.0-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; nvidia_cudnn_frontend-1.27.0-cp312-cp312-win_amd64.whl; nvidia_cudnn_frontend-1.27.0-cp313-cp313-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; nvidia_cudnn_frontend-1.27.0-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; nvidia_cudnn_frontend-1.27.0-cp313-cp313-win_amd64.whl; nvidia_cudnn_frontend-1.27.0-cp313-cp313-win_arm64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314t-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314t-win_amd64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314t-win_arm64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314-win_amd64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “cudnn python bindings”
- nvidia-cudnn-frontendProvides Python and C++ APIs to NVIDIA's cuDNN library, exposing…
- nvidia-cudnn-cu13Provides cuDNN runtime libraries for GPU-accelerated deep neural…
- nvidia-cudnn-cu11Provides cuDNN runtime libraries for GPU-accelerated deep neural…
Give your agent the search over MCP, or paste the wish link into any chat.
More Python Modules packages
Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.
Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.
Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.
PyYAML parses and emits YAML 1.1 data format, enabling serialization and deserialization of configuration files and Python objects to and from human-readable YAML text.
Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.
Provides reusable metadata objects for use with PEP-593 `typing.Annotated` to express common constraints like bounds, collection sizes, and predicates on types.
Install it if you use or build libraries that need to express type constraints in a standardized, inspectable way—or if you want to annotate your own types with…
Provides runtime tools to inspect and introspect Python type annotations, enabling programmatic examination of type hints at execution time.
See also fa3-fwd · transformer-engine-cu12 · tokenspeed-mla · humming-kernels · gram-newton-schulz · causal-conv1d · flash-attn-4 · flashinfer-cubin · comfy-kitchen · nvidia-cutlass-dsl-libs-base