cuvs-cu12
cuVS: Vector Search on the GPU
Decision gist · record as of 2026-08-14
Yes, if you have NVIDIA GPU hardware and need fast vector search or clustering. The active maintenance, permissive Apache-2.0 license, and zero known vulnerabilities make it production-ready. Medium install friction (CUDA dependencies) is the main trade-off; ensure your environment has compatible NVIDIA hardware and CUDA 12 support before installing.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires NVIDIA GPU with CUDA 12 support and compatible CUDA runtime environment installed on the system.
- Medium install friction due to CUDA 12 GPU library dependencies (libcuvs-cu12, pylibraft-cu12, cuda-bindings, numpy).
- Active maintenance with recent releases.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions, making it suitable for production applications.
last release 2026-08-06 (8 days) · last repo commit 2026-08-14 · 833 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 115,545 downloads/mo, #12,244 on PyPI
Alternatives
Verify before relying
pip install cuvs-cu12
from cuvs.neighbors import cagra
index_params = cagra.IndexParams()
index = cagra.build(index_params, dataset)- Whether the package works on non-NVIDIA GPUs or requires NVIDIA-specific hardware
- Performance characteristics compared to CPU-based vector search libraries
- Specific NVIDIA GPU compute capability requirements beyond CUDA 12
- Typical dataset and query sizes the library is optimized for
What it is and what it does
cuVS is a GPU-accelerated library for vector search and clustering built on NVIDIA's RAPIDS RAFT primitives. It implements algorithms like CAGRA for approximate nearest neighbor search, enabling fast similarity queries on embedding collections. The library is designed to accelerate semantic search, recommendation systems, and clustering workloads by offloading computation to NVIDIA GPUs.
The package provides Python, C++, C, and Rust APIs. It depends on libcuvs-cu12 (the core CUDA library), pylibraft-cu12 (RAPIDS RAFT bindings), cuda-bindings, and numpy. Installation requires Python 3.11 or later and a compatible NVIDIA GPU with CUDA 12 support. The library is actively maintained and handles CUDA version compatibility automatically.
Use it for
- Build semantic search systems for retrieval-augmented generation (RAG) pipelines using GPU acceleration
- Implement GPU-accelerated k-nearest neighbor graph construction for clustering and visualization algorithms
- Deploy high-throughput embedding similarity search in recommendation systems or image/text search applications
- Accelerate data mining tasks like clustering and visualization by computing nearest neighbor graphs on GPU
- Integrate vector search into databases or applications that need low-latency similarity queries at scale
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you have NVIDIA GPU hardware and need fast vector search or clustering.
The active maintenance, permissive Apache-2.0 license, and zero known vulnerabilities make it production-ready. Medium install friction (CUDA dependencies) is the main trade-off; ensure your environment has compatible NVIDIA hardware and CUDA 12 support before installing.
Install
cuvs-cu12 on PyPI
Before you install
Medium install friction due to CUDA 12 GPU library dependencies (libcuvs-cu12, pylibraft-cu12, cuda-bindings, numpy). Active maintenance with recent releases.
Requires NVIDIA GPU with CUDA 12 support and compatible CUDA runtime environment installed on the system.
License in practice
Apache-2.0 permissive license allows commercial and private use with minimal restrictions, making it suitable for production applications.
Quickstart
pip install cuvs-cu12
from cuvs.neighbors import cagra
index_params = cagra.IndexParams()
index = cagra.build(index_params, dataset)
Verify before relying
- Whether the package works on non-NVIDIA GPUs or requires NVIDIA-specific hardware
- Performance characteristics compared to CPU-based vector search libraries
- Specific NVIDIA GPU compute capability requirements beyond CUDA 12
- Typical dataset and query sizes the library is optimized for
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.11 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 4 packagescuda-bindingslibcuvs-cu12numpypylibraft-cu12 |
| Maintenance | Actively maintained 8 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 115,545 / month, #12,244 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersProgramming Language :: PythonProgramming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14 |
Evidence: cuvs_cu12-26.8.1-cp311-abi3-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl; cuvs_cu12-26.8.1-cp311-abi3-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “GPU vector search”
- cuvs-cu12Provides GPU-accelerated approximate nearest neighbor search and…
- faiss-gpuFaiss provides GPU-accelerated similarity search and clustering for…
- libcuvs-cu12GPU-accelerated vector search and clustering library providing…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also libcuvs-cu12 · usearch · faiss-gpu · voyager · scann · pynndescent · hnswlib · fastcluster · cuml-cu12 · libcuml-cu12