ddsketch
Distributed quantile sketches
Decision gist · record as of 2026-08-14
Yes. DDSketch is actively maintained, has no known vulnerabilities, installs with minimal friction, and solves a specific problem (distributed quantile estimation with error guarantees) that is difficult to implement correctly from scratch. Use it if you need accurate percentile estimates from streaming or distributed data without storing raw values.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.7 or later.
- Low friction: pure Python wheel with a single lightweight runtime dependency (six).
- Actively maintained with recent commits and no known vulnerabilities.
License · maintenance · safety
permissive license (permissive) — Permissive license allows commercial and private use with minimal restrictions.
last release 2024-04-01 (865 days) · last repo commit 2026-03-16 · 92 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 5,822,808 downloads/mo, #2,032 on PyPI
Alternatives
Verify before relying
pip install ddsketch
from ddsketch import DDSketch
sketch = DDSketch()
for value in [1.0, 2.5, 3.7, 5.2]:
sketch.add(value)
median = sketch.get_quantile_value(0.5)- Whether numpy is a runtime dependency or only a development/optional dependency (description mentions it but fact sheet lists only six).
- Performance characteristics for very large datasets or high-frequency streaming scenarios.
- Real-world use cases for tail latency monitoring and typical quantile queries in production systems.
What it is and what it does
DDSketch is a Python implementation of a distributed quantile sketch algorithm that estimates any quantile of a dataset while guaranteeing a bounded relative error. The default relative error is set to 0.01, meaning if the true quantile is x, the sketch returns a value y such that the relative difference stays within that bound. Instead of storing all data points, it maintains a compact sketch that grows predictably even for large datasets with sub-exponential tails.
The package is designed for distributed systems: you can create sketches on different nodes, add data to each independently, and then merge them into a central sketch to compute accurate quantiles across the combined dataset. It provides a standard DDSketch implementation plus variants (LogCollapsingLowestDenseDDSketch and LogCollapsingHighestDenseDDSketch) that trade accuracy guarantees for different quantile ranges, with configurable bin counts (default m = 2048) to tune memory versus precision.
Use it for
- Computing percentiles of request latencies across distributed services without storing all raw measurements.
- Estimating quantiles of financial data or sensor readings where you need bounded error but cannot afford to keep every data point.
- Merging quantile summaries from multiple data sources or time windows to compute aggregate statistics.
- Monitoring systems where you need to track tail latencies with minimal memory overhead.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
DDSketch is actively maintained, has no known vulnerabilities, installs with minimal friction, and solves a specific problem (distributed quantile estimation with error guarantees) that is difficult to implement correctly from scratch. Use it if you need accurate percentile estimates from streaming or distributed data without storing raw values.
Install
ddsketch on PyPI
Before you install
Low friction: pure Python wheel with a single lightweight runtime dependency (six). Actively maintained with recent commits and no known vulnerabilities.
Requires Python 3.7 or later.
License in practice
Permissive license allows commercial and private use with minimal restrictions.
Quickstart
pip install ddsketch
from ddsketch import DDSketch
sketch = DDSketch()
for value in [1.0, 2.5, 3.7, 5.2]:
sketch.add(value)
median = sketch.get_quantile_value(0.5)
Verify before relying
- Whether numpy is a runtime dependency or only a development/optional dependency (description mentions it but fact sheet lists only six).
- Performance characteristics for very large datasets or high-frequency streaming scenarios.
- Real-world use cases for tail latency monitoring and typical quantile queries in production systems.
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.7 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagesix |
| Maintenance | Actively maintained 865 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 5,822,808 / month, #2,032 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3 |
Evidence: ddsketch-3.0.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “quantile estimation streaming”
- ddsketchDDSketch computes quantiles (percentiles) of streaming or batch data…
- tdigestImplements Ted Dunning's t-digest data structure for efficient…
- fastdigestfastdigest provides a Rust-backed t-digest implementation for…
Give your agent the search over MCP, or paste the wish link into any chat.
More Information Analysis packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.
Install it if you use Snowflake and want to process data without moving it to your application layer.
See also datasketches · fig2sketch · fastdigest · tdigest · datasketch · quantile-forest · madoka · whylogs-sketching · ddtrace