dask-cudf-cu12
Utilities for Dask and cuDF interactions
Decision gist · record as of 2026-08-14
Yes, if you have NVIDIA GPUs and need to process large datasets faster than CPU Dask. The low install friction, active maintenance, permissive Apache-2.0 license, and zero known vulnerabilities make it a solid choice. Requires CUDA 12 runtime and GPU hardware; not suitable for CPU-only environments.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires NVIDIA GPU(s) with CUDA 12 support and compatible drivers; cudf-cu12 and cupy-cuda12x must be installed.
- Python >=3.11 required.
- Low install friction; pure Python wheel.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions.
last release 2026-08-06 (8 days) · last repo commit 2026-08-14 · 9,730 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 186,928 downloads/mo, #9,974 on PyPI
Alternatives
Verify before relying
pip install dask-cudf-cu12
import dask.dataframe as dd
from cupy import cuda
df = dd.read_parquet("/path/to/data/")
result = df.groupby('item')['price'].mean().compute()- Whether multi-GPU scaling on a single node works without additional cluster deployment tooling.
- Performance characteristics and memory overhead compared to single-GPU cuDF or CPU Dask DataFrame.
- Compatibility with specific NVIDIA GPU architectures and driver versions beyond CUDA 12.
What it is and what it does
Dask cuDF is a GPU-accelerated extension for Dask DataFrame that brings RAPIDS cuDF's pandas-like API to distributed GPU computing. It automatically registers as the 'cudf' backend for Dask, allowing you to write familiar pandas-style code that executes on GPUs instead of CPUs. The package depends on cudf-cu12, cupy-cuda12x, fsspec, numpy, nvidia-ml-py, pandas, and rapids-dask-dependency to provide GPU computation and memory management. It handles coordination between Dask's task scheduler and GPU operations, making it possible to process datasets larger than a single GPU's memory by spilling to host memory.
The package is actively maintained, supports Python 3.11–3.14, and carries no known security vulnerabilities. Single-node multi-GPU workflows are the primary use case. The description notes that multi-node execution requires deploying a distributed cluster separately.
Use it for
- Process multi-gigabyte Parquet or CSV datasets on a single machine with multiple GPUs faster than CPU-based Dask.
- Run groupby, join, and aggregation operations on GPU-resident data using familiar pandas-style syntax.
- Prototype data pipelines that will later scale to multi-node GPU clusters without rewriting core logic.
- Leverage GPU memory pools and spilling to host memory for workloads that exceed individual GPU VRAM.
- Integrate GPU-accelerated dataframe operations into existing Dask workflows.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you have NVIDIA GPUs and need to process large datasets faster than CPU Dask.
The low install friction, active maintenance, permissive Apache-2.0 license, and zero known vulnerabilities make it a solid choice. Requires CUDA 12 runtime and GPU hardware; not suitable for CPU-only environments.
Install
dask-cudf-cu12 on PyPI
Before you install
Low install friction; pure Python wheel. Active maintenance with recent releases. Requires CUDA 12 runtime and GPU libraries (cudf-cu12, cupy-cuda12x, nvidia-ml-py) as dependencies, which may require system-level NVIDIA driver setup.
Requires NVIDIA GPU(s) with CUDA 12 support and compatible drivers; cudf-cu12 and cupy-cuda12x must be installed. Python >=3.11 required.
License in practice
Apache-2.0 permissive license allows commercial and private use with minimal restrictions.
Quickstart
pip install dask-cudf-cu12
import dask.dataframe as dd
from cupy import cuda
df = dd.read_parquet("/path/to/data/")
result = df.groupby('item')['price'].mean().compute()
Verify before relying
- Whether multi-GPU scaling on a single node works without additional cluster deployment tooling.
- Performance characteristics and memory overhead compared to single-GPU cuDF or CPU Dask DataFrame.
- Compatibility with specific NVIDIA GPU architectures and driver versions beyond CUDA 12.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.11 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 7 packagescudf-cu12cupy-cuda12xfsspecnumpynvidia-ml-pypandasrapids-dask-dependency |
| Maintenance | Actively maintained 8 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 186,928 / month, #9,974 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersProgramming Language :: PythonProgramming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: DatabaseTopic :: Scientific/Engineering |
Evidence: dask_cudf_cu12-26.8.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “dask cuda dataframe”
- dask-cudf-cu12Dask cuDF extends Dask DataFrame with a GPU-accelerated backend,…
- cudf-cu12cuDF is a GPU-accelerated DataFrame library that provides pandas-like…
- dask-exprDask Expressions provides query optimization for Dask DataFrames by…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also cudf-cu12 · dask-cuda · libcudf-cu12 · pylibcudf-cu12 · raft-dask-cu12 · dask-expr · libraft-cu12 · libcuvs-cu12 · libcuml-cu12 · nvidia-cuda-cccl-cu12