rmm-cu12
rmm - RAPIDS Memory Manager
Decision gist · record as of 2026-08-14
Yes, if you are building GPU-accelerated applications with CUDA 12.2+ and need to optimize memory allocation behavior. The active maintenance, permissive license, and zero known vulnerabilities make it a safe choice. Install friction is moderate due to compiled dependencies and CUDA requirements, but this is expected for GPU libraries. Not necessary for simple GPU workloads that do not require custom memory strategies.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires CUDA 12.2+ and a compatible NVIDIA GPU (Volta architecture or newer).
- The cu12 variant is specific to CUDA 12.x; other CUDA versions require different rmm packages.
- Medium install friction due to compiled wheels and CUDA runtime dependencies.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions.
last release 2026-08-06 (8 days) · last repo commit 2026-08-14 · 709 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 429,944 downloads/mo, #6,736 on PyPI
Alternatives
Verify before relying
pip install rmm-cu12
import rmm
from rmm.allocators import CudaMemoryResource
# Use RMM's memory resource for GPU allocation
rmm.reinitialize(allocator='cuda')- Whether the package works with CUDA versions other than 12.x (the cu12 suffix suggests CUDA 12 specificity)
- Performance characteristics compared to default CUDA memory allocation
- Compatibility with non-Volta GPU architectures despite documentation stating Volta+ support
What it is and what it does
RMM is a memory management library for GPU-accelerated workflows that sits between your application and CUDA's memory allocator. It provides a pluggable interface for customizing how device memory and host memory are allocated, enabling strategies like device memory pooling (to reduce allocation overhead) and pinned host memory (for faster asynchronous transfers). The library is part of the RAPIDS ecosystem and is designed for applications that need fine-grained control over memory behavior to achieve optimal GPU performance.
The package depends on cuda-bindings, librmm-cu12 (the compiled C++ library), and numpy. Installation requires a CUDA 12.2+ environment and works with Python 3.11 through 3.14. It is actively maintained and has no known security vulnerabilities.
Use it for
- Reduce GPU memory allocation overhead in tight loops by using a device memory pool sub-allocator.
- Enable asynchronous host-to-device transfers by allocating pinned host memory through RMM.
- Integrate custom memory allocation strategies into RAPIDS libraries like cuDF or cuML.
- Profile and optimize memory usage patterns in GPU-accelerated data processing pipelines.
- Build C++ applications that need portable, customizable GPU memory management across different NVIDIA architectures.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are building GPU-accelerated applications with CUDA 12.2+ and need to optimize memory allocation behavior.
The active maintenance, permissive license, and zero known vulnerabilities make it a safe choice. Install friction is moderate due to compiled dependencies and CUDA requirements, but this is expected for GPU libraries. Not necessary for simple GPU workloads that do not require custom memory strategies.
Install
rmm-cu12 on PyPI
Before you install
Medium install friction due to compiled wheels and CUDA runtime dependencies. Active maintenance with recent releases (8 days since last update) and a stable repository.
Requires CUDA 12.2+ and a compatible NVIDIA GPU (Volta architecture or newer). The cu12 variant is specific to CUDA 12.x; other CUDA versions require different rmm packages.
License in practice
Apache-2.0 permissive license allows commercial and private use with minimal restrictions.
Quickstart
pip install rmm-cu12
import rmm
from rmm.allocators import CudaMemoryResource
# Use RMM's memory resource for GPU allocation
rmm.reinitialize(allocator='cuda')
Verify before relying
- Whether the package works with CUDA versions other than 12.x (the cu12 suffix suggests CUDA 12 specificity)
- Performance characteristics compared to default CUDA memory allocation
- Compatibility with non-Volta GPU architectures despite documentation stating Volta+ support
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.11 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 3 packagescuda-bindingslibrmm-cu12numpy |
| Maintenance | Actively maintained 8 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 429,944 / month, #6,736 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersProgramming Language :: PythonProgramming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: DatabaseTopic :: Scientific/Engineering |
Evidence: rmm_cu12-26.8.0-cp311-abi3-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl; rmm_cu12-26.8.0-cp311-abi3-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “gpu memory allocation”
- rmm-cu12RMM provides a common interface for customizing GPU and host memory…
- librmm-cu12librmm-cu12 provides GPU memory allocation and management for CUDA 12…
- py3nvmlProvides Python 3 bindings to query NVIDIA GPU state and manage GPU…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also cymem · librmm-cu12 · libraft-cu12 · umf · raft-dask-cu12 · libucx-cu12 · cpm-kernels · nvidia-cuda-cccl-cu12 · libcudf-cu12 · omnimalloc