cutensor-cu12
NVIDIA cuTENSOR
Decision gist · record as of 2026-08-14
Yes, if you have CUDA 12 and an NVIDIA GPU and need high-performance tensor operations. The package is actively maintained, has no security vulnerabilities, and fills a specialized role in GPU-accelerated computing. The proprietary license and CUDA 12 dependency are the main constraints—verify licensing terms for your use case and confirm CUDA 12 availability before installing.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires NVIDIA CUDA 12 and an NVIDIA GPU; not usable on CPU-only systems or with other CUDA versions.
- Medium install friction due to platform-specific wheels (x86_64, aarch64, Windows) and CUDA 12 dependency.
- Package is actively maintained with a recent release.
License · maintenance · safety
NVIDIA Proprietary Software (unclear) — Licensed under NVIDIA Proprietary Software with no SPDX identifier; license treatment is unclear, so review NVIDIA's terms before use in proprietary or redistributed software.
last release 2026-06-15 (60 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 82,929 downloads/mo, #14,120 on PyPI
Alternatives
Verify before relying
pip install cutensor-cu12
import cutensor- Whether Python version constraints exist (requires_python is unspecified in metadata)
- Exact scope of NVIDIA Proprietary Software license and redistribution restrictions
- Whether just-in-time kernel compilation requires additional build tools or CUDA toolkit installation
What it is and what it does
cuTENSOR is NVIDIA's high-performance CUDA library for tensor operations on GPUs. It provides optimized implementations of tensor contractions (including just-in-time kernel compilation), reductions, and element-wise operations, with support for mixed-precision compute (FP64 input with FP32 compute, FP32 input with FP16/BF16/TF32 compute) and tensors up to 64 dimensions. It handles arbitrary data layouts and supports operations like tensor permutations, type conversions, and various activation functions.
The package is distributed as platform-specific wheels for CUDA 12 on x86_64, aarch64, and Windows. It has no Python runtime dependencies and is actively maintained. Installation requires CUDA 12 and an NVIDIA GPU; it cannot run on CPU-only systems.
Use it for
- Accelerate tensor network simulations and quantum computing workloads on NVIDIA GPUs.
- Optimize mixed-precision deep learning operations requiring FP32 inputs with FP16 or BF16 compute.
- Perform high-dimensional tensor contractions in scientific computing and machine learning frameworks.
- Implement custom tensor operations with just-in-time kernel compilation for domain-specific algorithms.
- Execute partial tensor reductions and element-wise operations on large multi-dimensional arrays.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you have CUDA 12 and an NVIDIA GPU and need high-performance tensor operations.
The package is actively maintained, has no security vulnerabilities, and fills a specialized role in GPU-accelerated computing. The proprietary license and CUDA 12 dependency are the main constraints—verify licensing terms for your use case and confirm CUDA 12 availability before installing.
Install
cutensor-cu12 on PyPI
Before you install
Medium install friction due to platform-specific wheels (x86_64, aarch64, Windows) and CUDA 12 dependency. Package is actively maintained with a recent release.
Requires NVIDIA CUDA 12 and an NVIDIA GPU; not usable on CPU-only systems or with other CUDA versions.
License in practice
Licensed under NVIDIA Proprietary Software with no SPDX identifier; license treatment is unclear, so review NVIDIA's terms before use in proprietary or redistributed software.
Quickstart
pip install cutensor-cu12
import cutensor
Verify before relying
- Whether Python version constraints exist (requires_python is unspecified in metadata)
- Exact scope of NVIDIA Proprietary Software license and redistribution restrictions
- Whether just-in-time kernel compilation requires additional build tools or CUDA toolkit installation
Package facts
| License | NVIDIA Proprietary Software unclear |
| Python support | Not specified |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 60 days since the last release |
| First released | |
| Downloads | 82,929 / month, #14,120 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Environment :: GPU :: NVIDIA CUDAEnvironment :: GPU :: NVIDIA CUDA :: 12Topic :: Scientific/Engineering |
Evidence: cutensor_cu12-2.7.0-py3-none-manylinux2014_aarch64.whl; cutensor_cu12-2.7.0-py3-none-manylinux2014_x86_64.whl; cutensor_cu12-2.7.0-py3-none-win_amd64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “cuda tensor operations gpu”
- cutensor-cu12cuTENSOR is a CUDA library for high-performance tensor operations on…
- cutensor-cu13cuTENSOR is a CUDA library for high-performance tensor operations,…
- cutensornet-cu13cuTensorNet is a GPU-accelerated library for tensor network…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also cutensor-cu13 · cutensornet-cu13 · nvidia-cusparselt-cu12 · nvidia-cusparselt-cu13 · nvidia-cutlass-dsl-libs-cu12 · nvidia-cusparse-cu12 · nvidia-cusparse-cu11 · nvidia-cusparse · nvidia-cudnn-cu11 · nvidia-cutlass-dsl