cudf-cu12
cuDF - GPU Dataframe
Decision gist · record as of 2026-08-14
Yes, with conditions. cuDF is actively maintained, permissively licensed, and well-suited for GPU-accelerated tabular data processing. Install only if you have an NVIDIA GPU with compatible CUDA 12 drivers and can manage the 18 runtime dependencies (including CUDA toolkit and GPU libraries). The medium install friction and CUDA version matching requirement are the main barriers; once satisfied, it offers substantial performance gains for data processing workloads that fit GPU memory.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires NVIDIA GPU with compatible CUDA 12 driver and matching CUDA toolkit; CUDA version suffix (-cu12) must match your installed CUDA version.
- Medium install friction: requires 18 runtime dependencies including CUDA toolkit, cupy, numba, and GPU-specific libraries (libcudf-cu12, pylibcudf-cu12, rmm-cu12, nvidia-cufile-cu12).
- Active maintenance with recent release (8 days old) and strong repository signals (9730 stars, last commit 2026-08-14).
License · maintenance · safety
Apache-2.0 (permissive) — Apache 2.0 permissive license allows commercial and private use with minimal restrictions, typical for scientific and data processing libraries in the RAPIDS ecosystem.
last release 2026-08-06 (8 days) · last repo commit 2026-08-14 · 9,730 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 420,570 downloads/mo, #6,788 on PyPI
Alternatives
Verify before relying
pip install cudf-cu12
import cudf
df = cudf.read_parquet("data.parquet")
df.dropna().groupby(["A", "B"]).mean()- Performance gains over pandas for typical workloads and data sizes not quantified in fact sheet.
- Compatibility matrix with specific GPU models and CUDA minor versions beyond the cu12 designation.
- Memory overhead or limitations compared to pandas DataFrames on GPU.
What it is and what it does
cuDF is a GPU-accelerated DataFrame library that mirrors the pandas API, allowing you to work with tabular data on NVIDIA GPUs instead of CPUs. It's part of the RAPIDS suite and includes multiple components: libcudf (the core CUDA C++ library), pylibcudf (Cython bindings), the main cudf library (pandas-like interface), cudf-polars (GPU engine for Polars), and dask-cudf (GPU backend for Dask). The package depends on 18 runtime libraries including CUDA toolkit, cupy, numba, pandas, pyarrow, and GPU-specific NVIDIA libraries.
The primary use case is accelerating data processing workloads by offloading computation to GPU. cuDF offers a drop-in replacement for pandas code through cudf.pandas, which can be invoked with `python -m cudf.pandas` or a Jupyter magic command, requiring no code changes. It also integrates with Polars (via lazy evaluation with `engine="gpu"`) and works as a backend for distributed computing frameworks like Dask and Apache Spark (via Spark RAPIDS).
Use it for
- Accelerate existing pandas workflows on GPU by running `python -m cudf.pandas script.py` without modifying code.
- Process large parquet or CSV files with GPU-accelerated groupby, filtering, and aggregation operations.
- Build distributed GPU data pipelines using dask-cudf for multi-GPU or multi-node workloads.
- Run Polars queries on GPU by collecting lazy frames with `engine="gpu"` for performance-critical analytics.
- Integrate GPU-accelerated DataFrames into Spark jobs via Spark RAPIDS for hybrid CPU-GPU ETL.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with conditions.
cuDF is actively maintained, permissively licensed, and well-suited for GPU-accelerated tabular data processing. Install only if you have an NVIDIA GPU with compatible CUDA 12 drivers and can manage the 18 runtime dependencies (including CUDA toolkit and GPU libraries). The medium install friction and CUDA version matching requirement are the main barriers; once satisfied, it offers substantial performance gains for data processing workloads that fit GPU memory.
Install
cudf-cu12 on PyPI
Before you install
Medium install friction: requires 18 runtime dependencies including CUDA toolkit, cupy, numba, and GPU-specific libraries (libcudf-cu12, pylibcudf-cu12, rmm-cu12, nvidia-cufile-cu12). Active maintenance with recent release (8 days old) and strong repository signals (9730 stars, last commit 2026-08-14).
Requires NVIDIA GPU with compatible CUDA 12 driver and matching CUDA toolkit; CUDA version suffix (-cu12) must match your installed CUDA version.
License in practice
Apache 2.0 permissive license allows commercial and private use with minimal restrictions, typical for scientific and data processing libraries in the RAPIDS ecosystem.
Quickstart
pip install cudf-cu12
import cudf
df = cudf.read_parquet("data.parquet")
df.dropna().groupby(["A", "B"]).mean()
Verify before relying
- Performance gains over pandas for typical workloads and data sizes not quantified in fact sheet.
- Compatibility matrix with specific GPU models and CUDA minor versions beyond the cu12 designation.
- Memory overhead or limitations compared to pandas DataFrames on GPU.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.11 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 18 packagescachetoolscuda-bindingscuda-toolkitcupy-cuda12xfsspeclibcudf-cu12numba-cuda-mlirnumba-cudanumbanumpynvidia-cufile-cu12nvtxpackagingpandaspyarrowpylibcudf-cu12richrmm-cu12 |
| Maintenance | Actively maintained 8 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 420,570 / month, #6,788 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersProgramming Language :: PythonProgramming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: DatabaseTopic :: Scientific/Engineering |
Evidence: cudf_cu12-26.8.0-cp311-abi3-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl; cudf_cu12-26.8.0-cp311-abi3-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “gpu dataframe library”
- cudf-cu12cuDF is a GPU-accelerated DataFrame library that provides pandas-like…
- libcudf-cu12libcudf-cu12 is a GPU-accelerated C++ library providing Apache…
- pylibcudf-cu12pylibcudf-cu12 provides Python bindings for libcudf, a CUDA C++…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also dask-cudf-cu12 · pylibcudf-cu12 · libcudf-cu12 · libraft-cu12 · libcuml-cu12 · agate · swifter · nvidia-cudnn-cu12 · dask-cuda · narwhals