tdigest
T-Digest data structure
Decision gist · record as of 2026-08-14
Yes, if you need percentile estimation on streaming or distributed data and can tolerate an abandoned package. The algorithm is well-established and the implementation is stable; no known vulnerabilities exist. However, expect no bug fixes or updates—verify that accumulation-tree and pyudorandom remain compatible with your environment, and test accuracy for your specific use case before relying on it in production.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with two pure-Python wheels available.
- However, the package is abandoned—last release was 2019-05-07 and last commit 2023-05-04—so no active maintenance or security updates should be expected.
License · maintenance · safety
MIT (permissive) — MIT license is permissive, allowing commercial and private use with minimal restrictions; you may use, modify, and distribute this package freely provided you include the license notice.
last release 2019-05-07 (2656 days) · last repo commit 2023-05-04 · 408 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 336,793 downloads/mo, #7,456 on PyPI
Alternatives
Verify before relying
from tdigest import TDigest
digest = TDigest()
digest.update(0.5)
digest.update(0.3)
print(digest.percentile(50)) # median- Whether the package's two runtime dependencies (accumulation-tree, pyudorandom) are actively maintained or have known issues.
- Whether accuracy of percentile estimates meets requirements for your specific use case and data distribution.
- Current compatibility with modern Python versions beyond what the classifiers indicate.
What it is and what it does
tdigest is a Python implementation of Ted Dunning's t-digest algorithm, a probabilistic data structure designed to compute accurate percentiles, quantiles, and trimmed means from streaming or distributed datasets. Rather than storing all raw data points, it maintains a compact summary using centroids, allowing it to serialize to under 10kB and merge results from multiple data sources—making it particularly useful in map-reduce and distributed computing contexts.
The package provides methods to update the digest sequentially or in batches, query percentiles and cumulative distribution functions, compress the internal structure to reduce memory, and serialize/deserialize to and from Python dictionaries for storage or transmission. It depends on accumulation-tree and pyudorandom for its core operations.
Use it for
- Computing percentiles on large streaming datasets without storing all raw values in memory.
- Aggregating statistics across distributed systems by merging multiple t-digests from different nodes.
- Estimating quantiles and trimmed means in map-reduce pipelines where data is too large to centralize.
- Serializing statistical summaries of datasets for transmission or storage with minimal overhead.
- Calculating medians and percentile ranges for real-time monitoring or analytics applications.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need percentile estimation on streaming or distributed data and can tolerate an abandoned package.
The algorithm is well-established and the implementation is stable; no known vulnerabilities exist. However, expect no bug fixes or updates—verify that accumulation-tree and pyudorandom remain compatible with your environment, and test accuracy for your specific use case before relying on it in production.
Install
tdigest on PyPI
Before you install
Low install friction with two pure-Python wheels available. However, the package is abandoned—last release was 2019-05-07 and last commit 2023-05-04—so no active maintenance or security updates should be expected.
License in practice
MIT license is permissive, allowing commercial and private use with minimal restrictions; you may use, modify, and distribute this package freely provided you include the license notice.
Quickstart
from tdigest import TDigest
digest = TDigest()
digest.update(0.5)
digest.update(0.3)
print(digest.percentile(50)) # median
Verify before relying
- Whether the package's two runtime dependencies (accumulation-tree, pyudorandom) are actively maintained or have known issues.
- Whether accuracy of percentile estimates meets requirements for your specific use case and data distribution.
- Current compatibility with modern Python versions beyond what the classifiers indicate.
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagesaccumulation-treepyudorandom |
| Maintenance | Abandoned 2,656 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 336,793 / month, #7,456 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaLicense :: OSI Approved :: MIT LicenseProgramming Language :: PythonProgramming Language :: Python :: 3Topic :: Scientific/Engineering |
Evidence: tdigest-0.5.2.2-py2.py3-none-any.whl; tdigest-0.5.2.2-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “percentile estimation streaming data”
- tdigestImplements Ted Dunning's t-digest data structure for efficient…
- ddsketchDDSketch computes quantiles (percentiles) of streaming or batch data…
- fastdigestfastdigest provides a Rust-backed t-digest implementation for…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also fastdigest · ddsketch · crick · datasketches · quantile-forest · accumulation-tree · py-multihash · hdrhistogram · pytensor-distributions