$npx skillfedfor your agent

nvidia-cudnn-frontend

NVIDIA cuDNN Frontend — Python and C++ Graph API with SOTA attention (SDPA / Flash Attention), MoE grouped GEMM fusions, and FP8/MXFP8 kernels for Hopper and Blackwell GPUs.

Worth itPyPI Python ModulesReleased Aug 20264.5M downloads / moApache-2.0 AND MITPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — nvidia_cudnn_frontend-1.27.0-cp310-cp310-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl · nvidia_cudnn_frontend-1.27.0-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl · nvidia_cudnn_frontend-1.27.0-cp310-cp310-win_amd64.whl
v1.27.0 · released 2026-08-06 · Python >=3.9

Yes. The package is actively maintained, dual-licensed under permissive terms, has no known vulnerabilities, and provides essential optimized kernels for modern GPU workloads. Install it if you are training or deploying deep learning models on NVIDIA Hopper or Blackwell GPUs and need high-performance attention, grouped GEMM, or quantized operations. The requirement for CUDA Toolkit and cuDNN 8.5.0+ is a prerequisite, not a drawback—it reflects the package's tight integration with NVIDIA's GPU stack.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires NVIDIA driver, CUDA Toolkit, and cuDNN 8.5.0 or later installed on the system; Python 3.9 or later.
  • Installation is straightforward via pip with prebuilt wheels for Python 3.10–3.14 on Linux (x86_64 and aarch64) and Windows.
  • The package is actively maintained with a recent release (8 days old) and no known vulnerabilities.

License · maintenance · safety

Apache-2.0 AND MIT (permissive) — Dual-licensed under Apache-2.0 and MIT, both permissive licenses. You are free to use, modify, and distribute the package in commercial and private projects with minimal restrictions.

last release 2026-08-06 (8 days) · last repo commit 2026-08-14 · 904 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 4,479,592 downloads/mo, #2,292 on PyPI

Verify before relying

pip install nvidia-cudnn-frontend

import cudnn

# Create a graph for scaled dot-product attention
graph = cudnn.pygraph.Graph()
# Configure and execute attention operations on Hopper/Blackwell GPUs
  • Whether the package's autotuning mechanism requires additional configuration or warm-up time for production workloads.
  • Performance characteristics and memory overhead of the open-source FROST GEMM engine relative to backend-native plans.
  • Compatibility and integration effort with existing PyTorch models beyond the stated torch.compile support.
Same gist for agents: .md · .json

What it is and what it does

nvidia-cudnn-frontend is NVIDIA's modern entry point to cuDNN, offering both a header-only C++ API and Python bindings that abstract the complexity of the cuDNN Graph API. It exposes state-of-the-art GPU kernels including scaled dot-product attention (SDPA/Flash Attention), grouped GEMM fusions for mixture-of-experts training, fused normalization and activation operations, and quantized matrix multiplication in FP8 and MXFP8 precision. The package targets NVIDIA's latest GPU architectures—Hopper (H100/H200) and Blackwell (B200/GB200/GB300)—and includes native PyTorch integration with torch.compile support.

The package ships with open-source kernel implementations (FROST GEMM engine, block-sparse attention, native sparse attention, and fused RMSNorm+SiLU) that developers can inspect, modify, and contribute to. Installation is simple via pip, though it requires NVIDIA driver, CUDA Toolkit, and cuDNN 8.5.0 or later on the system. The library is actively maintained with prebuilt wheels for modern Python versions and multiple architectures.

Use it for

  • Accelerate transformer attention mechanisms in large language models using SDPA kernels optimized for Hopper and Blackwell GPUs.
  • Train mixture-of-experts models efficiently with fused grouped GEMM operations that reduce memory bandwidth and kernel launch overhead.
  • Deploy quantized deep learning models using FP8 and MXFP8 precision kernels for reduced memory footprint and faster inference.
  • Build custom GPU kernels by inspecting and modifying open-source implementations like FROST GEMM and block-sparse attention.
  • Integrate cuDNN-accelerated operations into PyTorch models with automatic differentiation and torch.compile support.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

The package is actively maintained, dual-licensed under permissive terms, has no known vulnerabilities, and provides essential optimized kernels for modern GPU workloads. Install it if you are training or deploying deep learning models on NVIDIA Hopper or Blackwell GPUs and need high-performance attention, grouped GEMM, or quantized operations. The requirement for CUDA Toolkit and cuDNN 8.5.0+ is a prerequisite, not a drawback—it reflects the package's tight integration with NVIDIA's GPU stack.

Install

nvidia-cudnn-frontend on PyPI

Before you install

Installation is straightforward via pip with prebuilt wheels for Python 3.10–3.14 on Linux (x86_64 and aarch64) and Windows. The package is actively maintained with a recent release (8 days old) and no known vulnerabilities. Medium install friction reflects the requirement for NVIDIA driver, CUDA Toolkit, and cuDNN 8.5.0 or later to be present on the system.

Requires NVIDIA driver, CUDA Toolkit, and cuDNN 8.5.0 or later installed on the system; Python 3.9 or later.

License in practice

Dual-licensed under Apache-2.0 and MIT, both permissive licenses. You are free to use, modify, and distribute the package in commercial and private projects with minimal restrictions.

Quickstart

pip install nvidia-cudnn-frontend

import cudnn

# Create a graph for scaled dot-product attention
graph = cudnn.pygraph.Graph()
# Configure and execute attention operations on Hopper/Blackwell GPUs

Verify before relying

  • Whether the package's autotuning mechanism requires additional configuration or warm-up time for production workloads.
  • Performance characteristics and memory overhead of the open-source FROST GEMM engine relative to backend-native plans.
  • Compatibility and integration effort with existing PyTorch models beyond the stated torch.compile support.

Package facts

LicenseApache-2.0 AND MIT permissive
Python supportSupports the current Python release >=3.9
Install frictionMedium. Platform-specific wheel
Runtime dependenciesNone
MaintenanceActively maintained 8 days since the last release
Last repo commit
First released
Downloads4,479,592 / month, #2,292 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableEnvironment :: GPU :: NVIDIA CUDAIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseLicense :: OSI Approved :: MIT LicenseOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: C++Programming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Software Development :: Libraries :: Python Modules

Evidence: nvidia_cudnn_frontend-1.27.0-cp310-cp310-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; nvidia_cudnn_frontend-1.27.0-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; nvidia_cudnn_frontend-1.27.0-cp310-cp310-win_amd64.whl; nvidia_cudnn_frontend-1.27.0-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; nvidia_cudnn_frontend-1.27.0-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; nvidia_cudnn_frontend-1.27.0-cp311-cp311-win_amd64.whl; nvidia_cudnn_frontend-1.27.0-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; nvidia_cudnn_frontend-1.27.0-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; nvidia_cudnn_frontend-1.27.0-cp312-cp312-win_amd64.whl; nvidia_cudnn_frontend-1.27.0-cp313-cp313-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; nvidia_cudnn_frontend-1.27.0-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; nvidia_cudnn_frontend-1.27.0-cp313-cp313-win_amd64.whl; nvidia_cudnn_frontend-1.27.0-cp313-cp313-win_arm64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314t-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314t-win_amd64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314t-win_arm64.whl; nvidia_cudnn_frontend-1.27.0-cp314-cp314-win_amd64.whl

Tags

Capabilities
cudnn python bindingsgpu attention kernelsflash attention implementationmixture of experts gemmnvidia cuda deep learningfp8 matrix multiplicationhopper blackwell gpu kernels
Topics
gpu-accelerationtransformer-kernelsquantization
PyPI keywords
cudnncudagpunvidiadeep-learningattentionsdpaflash-attentiontransformermoemixture-of-expertsgrouped-gemmfp8mxfp8blackwellhopperpytorchkernelgraph-api

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “cudnn python bindings”

Give your agent the search over MCP, or paste the wish link into any chat.

More Python Modules packages

idna Worth it
PyPI · Python Modules · released Jun 2026

Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.

Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.

BSD-3-Clausepure Python · 3.9+
1.8Bdownloads / mo
setuptools Worth it
PyPI · Python Modules · released Aug 2026

Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.

MITpure Python · 3.10+
1.6Bdownloads / mo
PyYAML Worth it
PyPI · Python Modules · released Sep 2025

PyYAML parses and emits YAML 1.1 data format, enabling serialization and deserialization of configuration files and Python objects to and from human-readable YAML text.

MITcompiled wheel · 3.8+
1.2Bdownloads / mo
pydantic Worth it
PyPI · Python Modules · released May 2026

Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.

MITpure Python · 3.9+
1.1Bdownloads / mo
annotated-types Worth it
PyPI · Python Modules · released Jul 2026

Provides reusable metadata objects for use with PEP-593 `typing.Annotated` to express common constraints like bounds, collection sizes, and predicates on types.

Install it if you use or build libraries that need to express type constraints in a standardized, inspectable way—or if you want to annotate your own types with…

MITpure Python · 3.10+
871.3Mdownloads / mo
typing-inspection Worth it
PyPI · Python Modules · released Aug 2026

Provides runtime tools to inspect and introspect Python type annotations, enabling programmatic examination of type hints at execution time.

MITpure Python · 3.10+
783.0Mdownloads / mo

See also fa3-fwd · transformer-engine-cu12 · tokenspeed-mla · humming-kernels · gram-newton-schulz · causal-conv1d · flash-attn-4 · flashinfer-cubin · comfy-kitchen · nvidia-cutlass-dsl-libs-base

Further reading