datasketches
The Apache DataSketches Library for Python
Decision gist · record as of 2026-08-14
Yes, if you need approximate answers to expensive big-data queries with proven error bounds and can tolerate a 531-day release gap. The library is stable, well-maintained under Apache governance, has no known vulnerabilities, and offers a mature API. Install friction is moderate due to C++ compilation, but prebuilt wheels reduce friction for standard platforms. Not suitable if you require exact results or need active feature development.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires numpy as a runtime dependency; C++ bindings require a compatible platform (wheels provided for common architectures).
- Medium install friction due to compiled C++ bindings; prebuilt wheels available across macOS, Linux, and Windows.
- Last release 531 days ago; maintenance status is aging.
License · maintenance · safety
Apache License 2.0 (permissive) — Apache 2.0 permissive license; precompiled binaries include nanobind (BSD license). No restrictions on commercial use or modification.
last release 2025-03-01 (531 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,205,968 downloads/mo, #4,222 on PyPI
Alternatives
Verify before relying
pip install datasketches
import datasketches
sketch = datasketches.kll_floats_sketch()
for value in data:
sketch.update(value)
print(sketch.get_quantile(0.5))- Whether the 531-day gap since last release indicates active maintenance or dormancy under Apache governance.
- Exact Python version support (requires_python not specified in metadata).
- Performance characteristics and memory overhead compared to alternatives for specific sketch types.
What it is and what it does
DataSketches is a Python binding to the Apache DataSketches library, a collection of streaming algorithms designed to answer expensive big-data queries approximately but with mathematically proven error bounds. Instead of computing exact results (which may require prohibitive compute time and resources), sketches trade precision for speed—often delivering results orders of magnitude faster. The library includes algorithms for cardinality estimation (HLL, CPC), quantile approximation (KLL, REQ, Quantiles), frequent-item detection, set operations (Theta, Tuple), sampling (VarOpt, EBPPS), and specialized tools like count-min sketches and kernel density estimation.
The package wraps C++ implementations via nanobind, providing a Python API that closely mirrors the C++ interface. It depends on numpy and is distributed as precompiled wheels for modern Python versions and common platforms. The library is portable across languages when endianness matches, making it suitable for systems that need to serialize sketches across different environments. Use it when exact answers are infeasible but approximate results with error guarantees are acceptable—typical scenarios include real-time analytics, interactive queries on massive datasets, and distributed systems where communication overhead dominates.
Use it for
- Estimate cardinality (distinct count) in a data stream without storing all unique values.
- Compute approximate quantiles and percentiles from unbounded data streams with controlled error.
- Identify frequent items or heavy hitters in a stream with false-positive/negative guarantees.
- Perform approximate set operations (union, intersection) on large datasets with memory efficiency.
- Implement real-time analytics dashboards where exact results are unnecessary but speed is critical.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need approximate answers to expensive big-data queries with proven error bounds and can tolerate a 531-day release gap.
The library is stable, well-maintained under Apache governance, has no known vulnerabilities, and offers a mature API. Install friction is moderate due to C++ compilation, but prebuilt wheels reduce friction for standard platforms. Not suitable if you require exact results or need active feature development.
Install
datasketches on PyPI
Before you install
Medium install friction due to compiled C++ bindings; prebuilt wheels available across macOS, Linux, and Windows. Last release 531 days ago; maintenance status is aging.
Requires numpy as a runtime dependency; C++ bindings require a compatible platform (wheels provided for common architectures).
License in practice
Apache 2.0 permissive license; precompiled binaries include nanobind (BSD license). No restrictions on commercial use or modification.
Quickstart
pip install datasketches
import datasketches
sketch = datasketches.kll_floats_sketch()
for value in data:
sketch.update(value)
print(sketch.get_quantile(0.5))
Verify before relying
- Whether the 531-day gap since last release indicates active maintenance or dormancy under Apache governance.
- Exact Python version support (requires_python not specified in metadata).
- Performance characteristics and memory overhead compared to alternatives for specific sketch types.
Package facts
| License | Apache License 2.0 permissive |
| Python support | Not specified |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 1 packagenumpy |
| Maintenance | Aging 531 days since the last release |
| First released | |
| Downloads | 1,205,968 / month, #4,222 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: datasketches-5.2.0-cp310-cp310-macosx_10_14_x86_64.whl; datasketches-5.2.0-cp310-cp310-macosx_11_0_arm64.whl; datasketches-5.2.0-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; datasketches-5.2.0-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; datasketches-5.2.0-cp310-cp310-musllinux_1_2_aarch64.whl; datasketches-5.2.0-cp310-cp310-musllinux_1_2_x86_64.whl; datasketches-5.2.0-cp310-cp310-win_amd64.whl; datasketches-5.2.0-cp311-cp311-macosx_10_14_x86_64.whl; datasketches-5.2.0-cp311-cp311-macosx_11_0_arm64.whl; datasketches-5.2.0-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; datasketches-5.2.0-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; datasketches-5.2.0-cp311-cp311-musllinux_1_2_aarch64.whl; datasketches-5.2.0-cp311-cp311-musllinux_1_2_x86_64.whl; datasketches-5.2.0-cp311-cp311-win_amd64.whl; datasketches-5.2.0-cp312-cp312-macosx_10_14_x86_64.whl; datasketches-5.2.0-cp312-cp312-macosx_11_0_arm64.whl; datasketches-5.2.0-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; datasketches-5.2.0-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; datasketches-5.2.0-cp312-cp312-musllinux_1_2_aarch64.whl; datasketches-5.2.0-cp312-cp312-musllinux_1_2_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “cardinality estimation sketches”
- datasketchesProvides streaming algorithms (sketches) for approximate answers to…
- whylogs-sketchingProvides Python bindings to Apache DataSketches' core C++ sketching…
- datasketchProvides probabilistic data structures (MinHash, HyperLogLog, and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Information Analysis packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.
Install it if you use Snowflake and want to process data without moving it to your application layer.
See also datasketch · ddsketch · madoka · whylogs-sketching · crick · HLL · fig2sketch · tdigest · fastdigest · libcuvs-cu12