nvidia-cutlass-dsl-libs-core
NVIDIA CUTLASS Python DSL
Decision gist · record as of 2026-08-14
Yes, if you have a CUDA 13 environment and need to write or prototype GPU kernels without C++ expertise. The low install friction, active maintenance, and recent release are positive signals. However, the public beta status and unclear license terms warrant caution for production use—verify licensing and test stability for your workload before committing.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires CUDA 13 environment and NVIDIA GPU; Python >=3.10; cuda-python and nvidia-cuda-nvdisasm must be available.
- Low install friction with a pure-wheel distribution.
- Active maintenance—released 9 days ago with recent commits.
License · maintenance · safety
(unclear) — License treatment is unclear; no SPDX identifier or raw license text is available. Classified as proprietary. Verify licensing terms before using in production or redistribution.
last release 2026-08-05 (9 days) · last repo commit 2026-08-14 · 10,250 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,145,267 downloads/mo, #4,300 on PyPI
Alternatives
Verify before relying
pip install nvidia-cutlass-dsl-libs-core
import nvidia_cutlass_dsl_libs_core
# Use CuTe DSL abstractions for kernel design- Exact scope of CuTe DSL API surface and whether it covers all linear algebra operations mentioned
- Performance benchmarks vs. native C++ CUTLASS on the same architectures
- Stability guarantees given the public beta status and planned graduation by summer 2026
- Whether deep learning framework integration is automatic or requires additional glue code
What it is and what it does
nvidia-cutlass-dsl-libs-core is a Python wrapper around NVIDIA's CUTLASS 4.x, exposing CuTe DSL—a domain-specific language for GPU kernel programming. It lets you write high-performance CUDA kernels in Python without C++ expertise, targeting Tensor Cores on modern NVIDIA GPUs. The package handles layouts, tensors, hardware atoms, and thread/data hierarchy control, with a focus on matrix multiply and linear algebra operations.
The package is currently in public beta and aims to flatten the GPU programming learning curve. It depends on numpy, protobuf, cuda-python, and nvidia-cuda-nvdisasm, so your environment must have CUDA 13 and an NVIDIA GPU. Compile times are claimed to be orders of magnitude faster than C++ alternatives, and it integrates natively with deep learning frameworks.
Use it for
- Prototype optimized matrix multiply kernels for Tensor Cores without writing C++
- Rapidly iterate on custom GPU kernels for deep learning research and experimentation
- Integrate hand-tuned CUDA kernels into deep learning workflows without C++ bindings
- Learn GPU programming and CuTe abstractions with a lower barrier to entry than C++
- Deploy production-grade linear algebra kernels targeting Ampere, Hopper, or Blackwell GPUs
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you have a CUDA 13 environment and need to write or prototype GPU kernels without C++ expertise.
The low install friction, active maintenance, and recent release are positive signals. However, the public beta status and unclear license terms warrant caution for production use—verify licensing and test stability for your workload before committing.
Install
nvidia-cutlass-dsl-libs-core on PyPI
Before you install
Low install friction with a pure-wheel distribution. Active maintenance—released 9 days ago with recent commits. Depends on cuda-python and nvidia-cuda-nvdisasm, which require NVIDIA GPU tooling; verify your CUDA environment before installing.
Requires CUDA 13 environment and NVIDIA GPU; Python >=3.10; cuda-python and nvidia-cuda-nvdisasm must be available.
License in practice
License treatment is unclear; no SPDX identifier or raw license text is available. Classified as proprietary. Verify licensing terms before using in production or redistribution.
Quickstart
pip install nvidia-cutlass-dsl-libs-core
import nvidia_cutlass_dsl_libs_core
# Use CuTe DSL abstractions for kernel design
Verify before relying
- Exact scope of CuTe DSL API surface and whether it covers all linear algebra operations mentioned
- Performance benchmarks vs. native C++ CUTLASS on the same architectures
- Stability guarantees given the public beta status and planned graduation by summer 2026
- Whether deep learning framework integration is automatic or requires additional glue code
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 6 packagesnumpytyping-extensionscuda-pythonbackports.strenumprotobufnvidia-cuda-nvdisasm |
| Maintenance | Actively maintained 9 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,145,267 / month, #4,300 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaEnvironment :: GPU :: NVIDIA CUDA :: 13License :: Other/Proprietary LicenseOperating System :: POSIX :: LinuxProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: Implementation :: CPython |
Evidence: nvidia_cutlass_dsl_libs_core-4.7.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “nvidia cutlass python bindings”
- nvidia-cutlass-dsl-libs-coreProvides a Python interface for writing high-performance CUDA kernels…
- nvidia-cutlass-dsl-libs-baseProvides Python interfaces for writing high-performance CUDA kernels…
- nvidia-cutlass-dsl-libs-cu12Provides a Python DSL for writing high-performance CUDA kernels using…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also nvidia-cutlass-dsl · nvidia-cutlass-dsl-libs-base · nvidia-cutlass-dsl-libs-cu12 · nvidia-cutlass-dsl-libs-cu13 · flydsl · nvidia-cudnn-frontend · apache-tvm-ffi · nvidia-cublas-cu11 · nvidia-cublas · nvidia-cublas-cu12