$npx skillfedfor your agent

nvidia-cutlass-dsl-libs-base

NVIDIA CUTLASS Python DSL

With conditionsPyPI Scientific/EngineeringReleased Aug 20266.3M downloads / moPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — nvidia_cutlass_dsl_libs_base-4.7.0-cp310-cp310-manylinux_2_28_aarch64.whl · nvidia_cutlass_dsl_libs_base-4.7.0-cp310-cp310-manylinux_2_28_x86_64.whl · nvidia_cutlass_dsl_libs_base-4.7.0-cp311-cp311-manylinux_2_28_aarch64.whl
v4.7.0 · released 2026-08-05 · Python >=3.10 · 7 runtime deps: numpy, typing-extensions, cuda-python, backports.strenum, protobuf, nvidia-cuda-nvdisasm, nvidia-cutlass-dsl-libs-core

Yes, if you are targeting NVIDIA Tensor Cores on Ampere/Hopper/Blackwell GPUs and want to prototype or deploy optimized kernels from Python. The active maintenance, zero known vulnerabilities, and large download volume indicate real adoption. However, verify the unclear license terms before production use, and be aware the package is in public beta—expect potential API changes before summer 2026 graduation.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires NVIDIA CUDA 13 environment, Linux (manylinux_2_28), and Python 3.10+.
  • Wheels available only for x86_64 and aarch64 architectures.
  • Medium install friction due to platform-specific wheels (manylinux_2_28, aarch64/x86_64 only) and multiple compiled dependencies including cuda-python and nvidia-cuda-nvdisasm.

License · maintenance · safety

(unclear) — License treatment is unclear—no SPDX identifier or raw license text provided. Verify licensing terms with NVIDIA before using in production or proprietary projects.

last release 2026-08-05 (9 days) · last repo commit 2026-08-14 · 10,250 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 6,269,249 downloads/mo, #1,944 on PyPI

Verify before relying

pip install nvidia-cutlass-dsl-libs-base
import cutlass
# Access CuTe DSL for kernel programming via CUTLASS Python interfaces
  • Exact API surface and available DSL constructs beyond CuTe DSL mentioned in description
  • Performance benchmarks or comparison to C++ CUTLASS implementations
  • Timeline and stability guarantees for beta-to-production graduation (stated as summer 2026)
  • Compatibility with specific DL frameworks (PyTorch, TensorFlow, etc.) mentioned in description
Same gist for agents: .md · .json

What it is and what it does

CUTLASS DSL provides a Python-native layer for writing optimized CUDA kernels without deep C++ expertise. The package exposes core CUTLASS and CuTe concepts—layouts, tensors, hardware atoms, and thread/data hierarchy control—allowing developers to target NVIDIA's Tensor Cores on modern GPU architectures. It aims to reduce the learning curve for GPU programming, speed up kernel prototyping, and integrate directly with deep learning frameworks.

The package depends on numpy, protobuf, cuda-python, and nvidia-cuda-nvdisasm, plus a core library (nvidia-cutlass-dsl-libs-core). It is currently in public beta and runs on Linux with Python 3.10 or later. Installation requires platform-specific wheels (x86_64 or aarch64), and a CUDA 13 environment is needed at runtime.

Use it for

  • Prototype and optimize matrix multiply kernels targeting Tensor Cores without writing C++
  • Develop custom linear algebra operations for deep learning workloads with native framework integration
  • Learn GPU programming and CUTLASS concepts with lower barrier to entry than C++ implementations
  • Rapidly iterate on kernel designs for performance engineering and research on modern NVIDIA GPUs
  • Integrate optimized CUDA kernels into production pipelines without glue code between Python and C++

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are targeting NVIDIA Tensor Cores on Ampere/Hopper/Blackwell GPUs and want to prototype or deploy optimized kernels from Python.

The active maintenance, zero known vulnerabilities, and large download volume indicate real adoption. However, verify the unclear license terms before production use, and be aware the package is in public beta—expect potential API changes before summer 2026 graduation.

Install

nvidia-cutlass-dsl-libs-base on PyPI

Before you install

Medium install friction due to platform-specific wheels (manylinux_2_28, aarch64/x86_64 only) and multiple compiled dependencies including cuda-python and nvidia-cuda-nvdisasm. Package is actively maintained with recent releases, but remains in public beta.

Requires NVIDIA CUDA 13 environment, Linux (manylinux_2_28), and Python 3.10+. Wheels available only for x86_64 and aarch64 architectures.

License in practice

License treatment is unclear—no SPDX identifier or raw license text provided. Verify licensing terms with NVIDIA before using in production or proprietary projects.

Quickstart

pip install nvidia-cutlass-dsl-libs-base
import cutlass
# Access CuTe DSL for kernel programming via CUTLASS Python interfaces

Verify before relying

  • Exact API surface and available DSL constructs beyond CuTe DSL mentioned in description
  • Performance benchmarks or comparison to C++ CUTLASS implementations
  • Timeline and stability guarantees for beta-to-production graduation (stated as summer 2026)
  • Compatibility with specific DL frameworks (PyTorch, TensorFlow, etc.) mentioned in description

Package facts

LicenseNot declared unclear
Python supportSupports the current Python release >=3.10
Install frictionMedium. Platform-specific wheel
Runtime dependencies
7 packages
numpytyping-extensionscuda-pythonbackports.strenumprotobufnvidia-cuda-nvdisasmnvidia-cutlass-dsl-libs-core
MaintenanceActively maintained 9 days since the last release
Last repo commit
First released
Downloads6,269,249 / month, #1,944 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaEnvironment :: GPU :: NVIDIA CUDA :: 13License :: Other/Proprietary LicenseOperating System :: POSIX :: LinuxProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: Implementation :: CPython

Evidence: nvidia_cutlass_dsl_libs_base-4.7.0-cp310-cp310-manylinux_2_28_aarch64.whl; nvidia_cutlass_dsl_libs_base-4.7.0-cp310-cp310-manylinux_2_28_x86_64.whl; nvidia_cutlass_dsl_libs_base-4.7.0-cp311-cp311-manylinux_2_28_aarch64.whl; nvidia_cutlass_dsl_libs_base-4.7.0-cp311-cp311-manylinux_2_28_x86_64.whl; nvidia_cutlass_dsl_libs_base-4.7.0-cp312-cp312-manylinux_2_28_aarch64.whl; nvidia_cutlass_dsl_libs_base-4.7.0-cp312-cp312-manylinux_2_28_x86_64.whl; nvidia_cutlass_dsl_libs_base-4.7.0-cp313-cp313-manylinux_2_28_aarch64.whl; nvidia_cutlass_dsl_libs_base-4.7.0-cp313-cp313-manylinux_2_28_x86_64.whl; nvidia_cutlass_dsl_libs_base-4.7.0-cp314-cp314-manylinux_2_28_aarch64.whl; nvidia_cutlass_dsl_libs_base-4.7.0-cp314-cp314-manylinux_2_28_x86_64.whl; nvidia_cutlass_dsl_libs_base-4.7.0-cp314-cp314t-manylinux_2_28_aarch64.whl; nvidia_cutlass_dsl_libs_base-4.7.0-cp314-cp314t-manylinux_2_28_x86_64.whl

Tags

Capabilities
python cuda kernel programmingcutlass dsl python interfacenvidia tensor core optimizationgpu kernel development pythoncute dsl matrix operationshigh-performance cuda kernelsnvidia gpu programming framework
Topics
gpu-programmingcuda-kernelstensor-cores

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “python cuda kernel programming”

Give your agent the search over MCP, or paste the wish link into any chat.

More Scientific/Engineering packages

numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
pandas Worth it
PyPI · Scientific/Engineering · released Jul 2026

pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.

BSD-3-Clausecompiled wheel · 3.11+
769.1Mdownloads / mo
scipy Worth it
PyPI · Libraries · released Jun 2026

scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.

BSD-3-Clausecompiled wheel · 3.12+
449.0Mdownloads / mo
scikit-learn Worth it
PyPI · Software Development · released Jun 2026

scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.

Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.

BSD-3-Clausecompiled wheel · 3.11+
235.5Mdownloads / mo
dill Worth it
PyPI · Software Development · released Jan 2026

dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.

BSD-3-Clausepure Python · 3.9+
208.1Mdownloads / mo
multiprocess Worth it
PyPI · Software Development · released Jan 2026

Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.

Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.

BSD-3-Clausepure Python · 3.9+
202.7Mdownloads / mo

See also apache-tvm-ffi · flydsl · nvidia-cutlass-dsl-libs-core · nvidia-cutlass-dsl-libs-cu12 · nvidia-cutlass-dsl · nvidia-cutlass-dsl-libs-cu13 · cuda-tile · nvidia-cudnn-frontend · nvidia-cublas-cu11 · dstack

Further reading