$npx skillfedfor your agent

nvidia-cutlass-dsl-libs-cu13

NVIDIA CUTLASS Python DSL

With conditionsPyPI Scientific/EngineeringReleased Aug 20262.1M downloads / moPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — nvidia_cutlass_dsl_libs_cu13-4.7.0-cp310-cp310-manylinux_2_28_aarch64.whl · nvidia_cutlass_dsl_libs_cu13-4.7.0-cp310-cp310-manylinux_2_28_x86_64.whl · nvidia_cutlass_dsl_libs_cu13-4.7.0-cp311-cp311-manylinux_2_28_aarch64.whl
v4.7.0 · released 2026-08-05 · Python >=3.10 · 7 runtime deps: numpy, typing-extensions, cuda-python, backports.strenum, protobuf, nvidia-cuda-nvdisasm, nvidia-cutlass-dsl-libs-base

Yes, if you target NVIDIA Ampere/Hopper/Blackwell GPUs on Linux and need to write or prototype optimized CUDA kernels in Python. The active maintenance, recent release, and strong upstream support are positive signals. However, the unclear license status and beta maturity (target graduation summer 2026) warrant verification of licensing terms and stability requirements before production deployment.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires CUDA 13 runtime, Linux (x86_64 or aarch64), Python 3.10 or later, and NVIDIA GPU hardware (Ampere, Hopper, or Blackwell architecture).
  • Medium install friction due to platform-specific wheels (x86_64 and aarch64 Linux only, Python 3.10–3.14) and a dependency chain including cuda-python and nvidia-cuda-nvdisasm.
  • Active maintenance with a recent release (9 days old) and strong upstream repository activity (10250 stars, last commit 2026-08-14).

License · maintenance · safety

(unclear) — License treatment is unclear—no SPDX identifier or raw license text is available. Users should verify licensing terms with NVIDIA before deploying in production or commercial contexts.

last release 2026-08-05 (9 days) · last repo commit 2026-08-14 · 10,250 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,134,076 downloads/mo, #3,265 on PyPI

Verify before relying

pip install nvidia-cutlass-dsl-libs-cu13
import nvidia_cutlass_dsl_libs_cu13
# Use CuTe DSL to define tensor layouts and kernels
  • Exact scope of CuTe DSL API surface and supported operations beyond matrix multiply.
  • Performance benchmarks comparing CuTe DSL kernels to hand-written CUDA C++.
  • Timeline and stability guarantees for beta-to-production graduation (stated target: summer 2026).
  • Compatibility with specific deep learning frameworks mentioned in the description.
Same gist for agents: .md · .json

What it is and what it does

CUTLASS DSL is NVIDIA's Python interface for writing optimized CUDA kernels using high-level abstractions (layouts, tensors, hardware atoms) instead of low-level C++. The first release, CuTe DSL, targets Tensor Core operations on Ampere, Hopper, and Blackwell GPUs, aiming to reduce the learning curve for GPU programming and speed up kernel prototyping. It is currently in public beta and depends on cuda-python, numpy, protobuf, and nvidia-cuda-nvdisasm.

The package is designed for students, researchers, and performance engineers who need to write efficient GPU code without deep C++ expertise. It promises faster compile times and native integration with deep learning frameworks. Installation is restricted to Linux (x86_64 and aarch64) with Python 3.10–3.14 and requires CUDA 13 runtime support.

Use it for

  • Prototyping optimized matrix multiply kernels for Tensor Cores without writing C++ code.
  • Teaching GPU programming concepts to students using a Python-native interface.
  • Rapidly iterating on custom CUDA kernel designs for deep learning workloads.
  • Integrating high-performance tensor operations directly into Python ML frameworks.
  • Benchmarking and optimizing linear algebra operations on modern NVIDIA GPUs.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you target NVIDIA Ampere/Hopper/Blackwell GPUs on Linux and need to write or prototype optimized CUDA kernels in Python.

The active maintenance, recent release, and strong upstream support are positive signals. However, the unclear license status and beta maturity (target graduation summer 2026) warrant verification of licensing terms and stability requirements before production deployment.

Install

nvidia-cutlass-dsl-libs-cu13 on PyPI

Before you install

Medium install friction due to platform-specific wheels (x86_64 and aarch64 Linux only, Python 3.10–3.14) and a dependency chain including cuda-python and nvidia-cuda-nvdisasm. Active maintenance with a recent release (9 days old) and strong upstream repository activity (10250 stars, last commit 2026-08-14).

Requires CUDA 13 runtime, Linux (x86_64 or aarch64), Python 3.10 or later, and NVIDIA GPU hardware (Ampere, Hopper, or Blackwell architecture).

License in practice

License treatment is unclear—no SPDX identifier or raw license text is available. Users should verify licensing terms with NVIDIA before deploying in production or commercial contexts.

Quickstart

pip install nvidia-cutlass-dsl-libs-cu13
import nvidia_cutlass_dsl_libs_cu13
# Use CuTe DSL to define tensor layouts and kernels

Verify before relying

  • Exact scope of CuTe DSL API surface and supported operations beyond matrix multiply.
  • Performance benchmarks comparing CuTe DSL kernels to hand-written CUDA C++.
  • Timeline and stability guarantees for beta-to-production graduation (stated target: summer 2026).
  • Compatibility with specific deep learning frameworks mentioned in the description.

Package facts

LicenseNot declared unclear
Python supportSupports the current Python release >=3.10
Install frictionMedium. Platform-specific wheel
Runtime dependencies
7 packages
numpytyping-extensionscuda-pythonbackports.strenumprotobufnvidia-cuda-nvdisasmnvidia-cutlass-dsl-libs-base
MaintenanceActively maintained 9 days since the last release
Last repo commit
First released
Downloads2,134,076 / month, #3,265 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaEnvironment :: GPU :: NVIDIA CUDA :: 13License :: Other/Proprietary LicenseOperating System :: POSIX :: LinuxProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: Implementation :: CPython

Evidence: nvidia_cutlass_dsl_libs_cu13-4.7.0-cp310-cp310-manylinux_2_28_aarch64.whl; nvidia_cutlass_dsl_libs_cu13-4.7.0-cp310-cp310-manylinux_2_28_x86_64.whl; nvidia_cutlass_dsl_libs_cu13-4.7.0-cp311-cp311-manylinux_2_28_aarch64.whl; nvidia_cutlass_dsl_libs_cu13-4.7.0-cp311-cp311-manylinux_2_28_x86_64.whl; nvidia_cutlass_dsl_libs_cu13-4.7.0-cp312-cp312-manylinux_2_28_aarch64.whl; nvidia_cutlass_dsl_libs_cu13-4.7.0-cp312-cp312-manylinux_2_28_x86_64.whl; nvidia_cutlass_dsl_libs_cu13-4.7.0-cp313-cp313-manylinux_2_28_aarch64.whl; nvidia_cutlass_dsl_libs_cu13-4.7.0-cp313-cp313-manylinux_2_28_x86_64.whl; nvidia_cutlass_dsl_libs_cu13-4.7.0-cp314-cp314-manylinux_2_28_aarch64.whl; nvidia_cutlass_dsl_libs_cu13-4.7.0-cp314-cp314-manylinux_2_28_x86_64.whl; nvidia_cutlass_dsl_libs_cu13-4.7.0-cp314-cp314t-manylinux_2_28_aarch64.whl; nvidia_cutlass_dsl_libs_cu13-4.7.0-cp314-cp314t-manylinux_2_28_x86_64.whl

Tags

Capabilities
python cuda kernel programmingcutlass dsl python interfacegpu tensor core optimizationcute dsl matrix multiplynvidia cuda kernel developmenthigh-performance gpu computingampere hopper blackwell kernels
Topics
gpu-computingcuda-kernelstensor-optimization

Let your AI agent find packages like this

Example. Real query, live index.

An agent finds packages by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language. Give your agent the search over MCP.

More Scientific/Engineering packages

numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
pandas Worth it
PyPI · Scientific/Engineering · released Jul 2026

pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.

BSD-3-Clausecompiled wheel · 3.11+
769.1Mdownloads / mo
scipy Worth it
PyPI · Libraries · released Jun 2026

scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.

BSD-3-Clausecompiled wheel · 3.12+
449.0Mdownloads / mo
scikit-learn Worth it
PyPI · Software Development · released Jun 2026

scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.

Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.

BSD-3-Clausecompiled wheel · 3.11+
235.5Mdownloads / mo
dill Worth it
PyPI · Software Development · released Jan 2026

dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.

BSD-3-Clausepure Python · 3.9+
208.1Mdownloads / mo
multiprocess Worth it
PyPI · Software Development · released Jan 2026

Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.

Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.

BSD-3-Clausepure Python · 3.9+
202.7Mdownloads / mo

See also nvidia-cutlass-dsl · flydsl · nvidia-cutlass-dsl-libs-cu12 · nvidia-cutlass-dsl-libs-core · nvidia-cutlass-dsl-libs-base · cuda-tile · nvidia-cudnn-frontend · transformer-engine-cu13 · nvidia-cublas-cu11 · cuda-python

Further reading