nvidia-cutlass-dsl
NVIDIA CUTLASS Python DSL
Decision gist · record as of 2026-08-14
Yes, if you need to write optimized CUDA kernels in Python and can tolerate beta-stage software. The low install friction, active maintenance, and clear focus on reducing GPU programming complexity make it valuable for researchers and performance engineers. However, verify NVIDIA's proprietary license terms for your use case, and be aware that the API may change before the summer 2026 beta graduation.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Linux, CUDA 12 or 13, Python 3.10 or later, and NVIDIA GPU with Ampere, Hopper, or Blackwell architecture to execute kernels.
- Low install friction with a pure-wheel distribution.
- The package is actively maintained with recent commits and is in public beta status, though it remains under active development.
License · maintenance · safety
(unclear) — Licensed under an unclear proprietary license with no SPDX identifier published. You should review NVIDIA's licensing terms directly before committing to production use.
last release 2026-08-05 (9 days) · last repo commit 2026-08-14 · 10,250 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 6,110,289 downloads/mo, #1,969 on PyPI
Alternatives
Verify before relying
pip install nvidia-cutlass-dsl
# Import and use the package with its runtime dependencies nvidia-cutlass-dsl-libs-base and nvidia-cutlass-dsl-libs-cu12- Whether the beta status implies breaking API changes before summer 2026 graduation
- What specific performance guarantees or benchmarks are documented for different kernel types
- Whether the two runtime library dependencies can coexist or require selection
- What the actual import path and public API surface are for the package
What it is and what it does
nvidia-cutlass-dsl is NVIDIA's Python domain-specific language for writing optimized CUDA kernels that run on modern GPU Tensor Cores. It exposes core CuTe concepts—layouts, tensors, hardware atoms, and thread/data hierarchy control—directly in Python, eliminating the need to write C++ glue code or possess deep GPU programming expertise. The package targets Ampere, Hopper, and Blackwell architectures and is designed to accelerate matrix multiply and linear algebra operations.
The package is currently in public beta and depends on two runtime libraries (nvidia-cutlass-dsl-libs-base and nvidia-cutlass-dsl-libs-cu12) that provide the underlying compiled components. It supports Python 3.10 through 3.14 on Linux with CUDA 12 or 13, and installs as a pure wheel with low friction. The project is actively maintained by NVIDIA with recent commits and is positioned as a tool for students, researchers, and performance engineers to prototype and deploy GPU kernels.
Use it for
- Rapidly prototype and optimize matrix multiplication kernels for deep learning workloads without writing C++ code
- Develop custom linear algebra operations targeting Tensor Cores on modern NVIDIA GPUs
- Integrate optimized GPU kernels directly into Python deep learning frameworks
- Learn GPU programming and Tensor Core concepts with a lower barrier to entry than C++ CUTLASS
- Benchmark and compare different kernel designs for high-throughput tensor operations
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need to write optimized CUDA kernels in Python and can tolerate beta-stage software.
The low install friction, active maintenance, and clear focus on reducing GPU programming complexity make it valuable for researchers and performance engineers. However, verify NVIDIA's proprietary license terms for your use case, and be aware that the API may change before the summer 2026 beta graduation.
Install
nvidia-cutlass-dsl on PyPI
Before you install
Low install friction with a pure-wheel distribution. The package is actively maintained with recent commits and is in public beta status, though it remains under active development.
Requires Linux, CUDA 12 or 13, Python 3.10 or later, and NVIDIA GPU with Ampere, Hopper, or Blackwell architecture to execute kernels.
License in practice
Licensed under an unclear proprietary license with no SPDX identifier published. You should review NVIDIA's licensing terms directly before committing to production use.
Quickstart
pip install nvidia-cutlass-dsl
# Import and use the package with its runtime dependencies nvidia-cutlass-dsl-libs-base and nvidia-cutlass-dsl-libs-cu12
Verify before relying
- Whether the beta status implies breaking API changes before summer 2026 graduation
- What specific performance guarantees or benchmarks are documented for different kernel types
- Whether the two runtime library dependencies can coexist or require selection
- What the actual import path and public API surface are for the package
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagesnvidia-cutlass-dsl-libs-basenvidia-cutlass-dsl-libs-cu12 |
| Maintenance | Actively maintained 9 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 6,110,289 / month, #1,969 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaEnvironment :: GPU :: NVIDIA CUDA :: 12Environment :: GPU :: NVIDIA CUDA :: 13License :: Other/Proprietary LicenseOperating System :: POSIX :: LinuxProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: Implementation :: CPython |
Evidence: nvidia_cutlass_dsl-4.7.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “nvidia tensor core optimization”
- nvidia-cutlass-dslProvides a Python native interface for writing high-performance CUDA…
- nvidia-cutlass-dsl-libs-baseProvides Python interfaces for writing high-performance CUDA kernels…
- nvidia-cutlass-dsl-libs-cu12Provides a Python DSL for writing high-performance CUDA kernels using…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also nvidia-cutlass-dsl-libs-core · nvidia-cutlass-dsl-libs-cu12 · nvidia-cutlass-dsl-libs-cu13 · nvidia-cutlass-dsl-libs-base · cuda-tile · flydsl · apache-tvm-ffi · quadrants · nvidia-cudnn-frontend · cutensor-cu13