$npx skillfedfor your agent

HLL

Fast HyperLogLog for Python

With conditionsPyPI Scientific/EngineeringReleased Feb 2026118.7K downloads / moMITSource build

Decision gist · record as of 2026-08-14

sdist only — hll-3.0.0.tar.gz · builds from source
v3.0.0 · released 2026-02-25 · Python >=3.9

Yes, if you need cardinality estimation and can tolerate a build-time dependency. The package is actively maintained, production-stable, has no known vulnerabilities, and solves a specific algorithmic problem well. Install friction is real (requires C compilation and dev headers), but that is inherent to the algorithm's performance. Not worth installing if you need exact counts or cannot set up a build environment.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires C compiler and Python development headers to build the C extension.
  • High install friction: requires C compilation and Python development headers.
  • The package is actively maintained (last commit 2026-02-25) and production-stable, but installation will demand a build environment on your system.

License · maintenance · safety

MIT (permissive) — MIT license is permissive and places no restrictions on use, modification, or distribution in proprietary or open-source projects.

last release 2026-02-25 (170 days) · last repo commit 2026-02-25 · 111 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 118,735 downloads/mo, #12,104 on PyPI

Verify before relying

pip install HLL

from HLL import HyperLogLog

hll = HyperLogLog(10)  # 2^10 registers
hll.add('some data')
print(hll.cardinality())
  • Whether the sparse-to-dense conversion strategy in 3.0.0 materially improves memory usage for typical workloads compared to prior versions
  • Accuracy degradation specifics when Jaccard similarity falls below 0.05 in intersection_cardinality() estimates
Same gist for agents: .md · .json

What it is and what it does

HLL is a C-based Python module implementing the 64-bit HyperLogLog algorithm for cardinality estimation. It trades accuracy for memory efficiency, allowing you to estimate how many unique elements exist in a dataset without storing all of them—useful when datasets are too large to fit in memory or when you need fast approximate counts in streaming contexts.

The package uses a Murmur64A hash and stores registers in a hybrid sparse-dense representation: sparse when few registers are set (using a sorted dynamic array), switching to dense when memory would be wasted. Version 3.0.0 adds intersection cardinality estimation via Ertl's JMLE method, bulk insertion via add_range(), and fixes memory leaks and type errors from earlier releases. Requires Python >= 3.9 and a C compiler.

Use it for

  • Estimate unique visitor counts or unique IDs in high-volume streaming logs without storing all values
  • Merge cardinality estimates from multiple data sources to approximate total unique items across distributed systems
  • Estimate set intersection size (e.g., overlapping users between two datasets) using intersection_cardinality()
  • Profile memory usage of large datasets by approximating cardinality with minimal RAM overhead
  • Bulk-insert sequential integers efficiently using add_range() to avoid Python-C call overhead

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need cardinality estimation and can tolerate a build-time dependency.

The package is actively maintained, production-stable, has no known vulnerabilities, and solves a specific algorithmic problem well. Install friction is real (requires C compilation and dev headers), but that is inherent to the algorithm's performance. Not worth installing if you need exact counts or cannot set up a build environment.

Install

hll on PyPI

Before you install

High install friction: requires C compilation and Python development headers. The package is actively maintained (last commit 2026-02-25) and production-stable, but installation will demand a build environment on your system.

Requires C compiler and Python development headers to build the C extension.

License in practice

MIT license is permissive and places no restrictions on use, modification, or distribution in proprietary or open-source projects.

Quickstart

pip install HLL

from HLL import HyperLogLog

hll = HyperLogLog(10)  # 2^10 registers
hll.add('some data')
print(hll.cardinality())

Verify before relying

  • Whether the sparse-to-dense conversion strategy in 3.0.0 materially improves memory usage for typical workloads compared to prior versions
  • Accuracy degradation specifics when Jaccard similarity falls below 0.05 in intersection_cardinality() estimates

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.9
Install frictionHigh. Source build required
Runtime dependenciesNone
MaintenanceActively maintained 170 days since the last release
Last repo commit
First released
Downloads118,735 / month, #12,104 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseOperating System :: MacOSOperating System :: POSIX :: LinuxProgramming Language :: CProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.9Topic :: Scientific/Engineering

Evidence: hll-3.0.0.tar.gz

Tags

Capabilities
cardinality estimationhyperloglog algorithmapproximate unique countmemory efficient set sizeprobabilistic data structurestreaming cardinalitybig data counting
Topics
cardinality-estimationprobabilistic-data-structuresstreaming-algorithms
PyPI keywords
hyperloglogcardinalitycardinality estimateapproximate countingprobabilistic data structuressketchdata sciencebig datastreaming algorithmsmemory efficientset cardinalityunique count

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “cardinality estimation”

  • HLLEstimates the cardinality (unique count) of very large datasets using…
  • datasketchesProvides streaming algorithms (sketches) for approximate answers to…
  • datasketchProvides probabilistic data structures (MinHash, HyperLogLog, and…

Give your agent the search over MCP, or paste the wish link into any chat.

More Scientific/Engineering packages

numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
pandas Worth it
PyPI · Scientific/Engineering · released Jul 2026

pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.

BSD-3-Clausecompiled wheel · 3.11+
769.1Mdownloads / mo
scipy Worth it
PyPI · Libraries · released Jun 2026

scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.

BSD-3-Clausecompiled wheel · 3.12+
449.0Mdownloads / mo
scikit-learn Worth it
PyPI · Software Development · released Jun 2026

scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.

Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.

BSD-3-Clausecompiled wheel · 3.11+
235.5Mdownloads / mo
dill Worth it
PyPI · Software Development · released Jan 2026

dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.

BSD-3-Clausepure Python · 3.9+
208.1Mdownloads / mo
multiprocess Worth it
PyPI · Software Development · released Jan 2026

Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.

Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.

BSD-3-Clausepure Python · 3.9+
202.7Mdownloads / mo

See also datasketch · datasketches · intbitset · whylogs-sketching · bytesparse · crick · omnimalloc · array-record · fast-array-utils · bitarray-hardbyte