cutensor-cu13
NVIDIA cuTENSOR
Decision gist · record as of 2026-08-14
Yes, if you have an NVIDIA GPU with CUDA 13 and need direct access to high-performance tensor primitives. The library is actively maintained, has no known vulnerabilities, and is the standard CUDA tensor backend for many ML frameworks. However, the proprietary license requires verification of compliance with NVIDIA's terms, and installation friction is moderate due to GPU and CUDA runtime requirements.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires NVIDIA GPU with CUDA 13 support; CUDA toolkit and cuTENSOR runtime libraries must be installed on the system.
- Medium install friction: platform-specific wheels for Linux (aarch64, x86_64) and Windows only; requires CUDA 13 and compatible GPU hardware.
- Package is actively maintained with recent releases.
License · maintenance · safety
NVIDIA Proprietary Software (unclear) — Licensed under NVIDIA Proprietary Software with unclear treatment—not an open-source license. Users should verify compliance with NVIDIA's terms before deployment in production or commercial contexts.
last release 2026-06-15 (60 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 102,402 downloads/mo, #12,872 on PyPI
Alternatives
Verify before relying
pip install cutensor-cu13
import cutensor
# Use cutensor for tensor operations on NVIDIA GPU- Exact Python version support (requires_python is unspecified in metadata)
- Whether JIT kernel compilation requires additional build tools or CUDA headers
- Performance characteristics and typical use-case scale (tensor sizes, operation types)
What it is and what it does
cuTENSOR is NVIDIA's proprietary CUDA library for tensor primitives—the low-level building block for GPU-accelerated tensor computations. It provides direct tensor contractions (including JIT-compiled kernels), partial and full reductions, and element-wise operations (permutations, type conversions, activations) on tensors up to 64 dimensions. The library supports mixed-precision workflows: FP64 inputs with FP32 compute, FP32 inputs with FP16/BF16/TF32 compute, and complex-times-real operations.
Installation requires a compatible NVIDIA GPU and CUDA 13 runtime; wheels are provided for Linux (x86_64, aarch64) and Windows. The package has no Python runtime dependencies and is actively maintained. It is typically used as a backend for higher-level tensor frameworks or as a direct compute primitive in machine learning and scientific computing pipelines that need fine-grained control over GPU tensor operations.
Use it for
- Accelerate tensor contraction operations in machine learning models on NVIDIA GPUs with JIT kernel compilation.
- Perform mixed-precision tensor reductions and element-wise operations in scientific computing workflows.
- Implement custom tensor network algorithms requiring direct control over GPU tensor primitives and data layouts.
- Build tensor manipulation kernels for deep learning frameworks that need low-level CUDA tensor support.
- Execute arbitrary tensor permutations and type conversions on GPU with minimal overhead.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you have an NVIDIA GPU with CUDA 13 and need direct access to high-performance tensor primitives.
The library is actively maintained, has no known vulnerabilities, and is the standard CUDA tensor backend for many ML frameworks. However, the proprietary license requires verification of compliance with NVIDIA's terms, and installation friction is moderate due to GPU and CUDA runtime requirements.
Install
cutensor-cu13 on PyPI
Before you install
Medium install friction: platform-specific wheels for Linux (aarch64, x86_64) and Windows only; requires CUDA 13 and compatible GPU hardware. Package is actively maintained with recent releases.
Requires NVIDIA GPU with CUDA 13 support; CUDA toolkit and cuTENSOR runtime libraries must be installed on the system.
License in practice
Licensed under NVIDIA Proprietary Software with unclear treatment—not an open-source license. Users should verify compliance with NVIDIA's terms before deployment in production or commercial contexts.
Quickstart
pip install cutensor-cu13
import cutensor
# Use cutensor for tensor operations on NVIDIA GPU
Verify before relying
- Exact Python version support (requires_python is unspecified in metadata)
- Whether JIT kernel compilation requires additional build tools or CUDA headers
- Performance characteristics and typical use-case scale (tensor sizes, operation types)
Package facts
| License | NVIDIA Proprietary Software unclear |
| Python support | Not specified |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 60 days since the last release |
| First released | |
| Downloads | 102,402 / month, #12,872 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Environment :: GPU :: NVIDIA CUDAEnvironment :: GPU :: NVIDIA CUDA :: 13Topic :: Scientific/Engineering |
Evidence: cutensor_cu13-2.7.0-py3-none-manylinux2014_aarch64.whl; cutensor_cu13-2.7.0-py3-none-manylinux2014_x86_64.whl; cutensor_cu13-2.7.0-py3-none-win_amd64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “cuda tensor operations”
- cutensor-cu13cuTENSOR is a CUDA library for high-performance tensor operations,…
- cutensor-cu12cuTENSOR is a CUDA library for high-performance tensor operations on…
- cutensornet-cu13cuTensorNet is a GPU-accelerated library for tensor network…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also cutensor-cu12 · cutensornet-cu13 · nvidia-cusparselt-cu13 · nvidia-cusparselt-cu12 · nvidia-cutlass-dsl-libs-cu12 · nvidia-cutlass-dsl · nvidia-cusparse-cu12 · nvidia-cutlass-dsl-libs-cu13 · nvidia-cudnn-cu13 · nvidia-cusparse