bbhash
A Python wrapper for the BBHash Minimal Perfect Hash Function
Decision gist · record as of 2026-08-14
Yes, with conditions. Install if you need minimal perfect hash functions for 64-bit hashes in a bioinformatics or k-mer-heavy workflow and can tolerate Cython compilation. The package is actively maintained and has no known vulnerabilities, but verify the unclear license status before use in proprietary contexts. High install friction and Python >=3.11 requirement may limit adoption in legacy environments.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Cython compilation during install (high friction).
- Python >=3.11 required.
- Depends on numpy.
License · maintenance · safety
(unclear) — License status is unclear—no SPDX identifier or raw license text is available in the package metadata. Verify the actual license before use in proprietary or restricted contexts.
last release 2025-10-26 (292 days) · last repo commit 2025-10-26 · 19 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 87,699 downloads/mo, #13,775 on PyPI
Alternatives
Verify before relying
import bbhash
uint_hashes = [10, 20, 50, 80]
mph = bbhash.PyMPHF(uint_hashes, len(uint_hashes), 1, 1.0)
for val in uint_hashes:
print('{} hashes to {}'.format(val, mph.lookup(val)))- Thread safety of PyMPHF and BBHashTable is noted as needing investigation by the maintainer.
- Whether the package is suitable for production use in bioinformatics pipelines beyond the original spacegraphcats use case.
- Performance characteristics and memory overhead compared to alternative hash table implementations.
What it is and what it does
pybbhash is a Cython wrapper around the BBHash C++ library for constructing minimal perfect hash functions optimized for 64-bit hash values. It provides two main interfaces: PyMPHF for building and querying a minimal perfect hash function, and BBHashTable for associating arbitrary values with hashes and retrieving them later. The package is designed primarily for bioinformatics workflows involving k-mer hashing, where you need to map large collections of hash values to compact integer identifiers or associated metadata.
The core use case is storing and querying relationships between hashes generated by tools like khmer or sourmash—for example, mapping k-mer hashes to De Bruijn graph node IDs. BBHashTable extends the basic MPHF by supporting lookups on hashes that were not part of the original construction, returning None for missing keys. Both modules support save/load to disk for persistence.
Use it for
- Map k-mer hashes to De Bruijn graph node identifiers in genome assembly workflows.
- Store and retrieve compact metadata associated with large collections of 64-bit hashes without full hash table overhead.
- Build minimal perfect hash functions for read deduplication or abundance tracking in sequencing pipelines.
- Query pre-built MPHF structures loaded from disk to avoid reconstruction overhead in repeated analyses.
- Associate arbitrary values with hashes in memory-constrained bioinformatics environments where space efficiency matters.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with conditions.
Install if you need minimal perfect hash functions for 64-bit hashes in a bioinformatics or k-mer-heavy workflow and can tolerate Cython compilation. The package is actively maintained and has no known vulnerabilities, but verify the unclear license status before use in proprietary contexts. High install friction and Python >=3.11 requirement may limit adoption in legacy environments.
Install
bbhash on PyPI
Before you install
High install friction due to Cython compilation requirement. Package is aging (292 days since last release) but repository remains active with a recent commit on 2025-10-26. Requires Python >=3.11.
Requires Cython compilation during install (high friction). Python >=3.11 required. Depends on numpy.
License in practice
License status is unclear—no SPDX identifier or raw license text is available in the package metadata. Verify the actual license before use in proprietary or restricted contexts.
Quickstart
import bbhash
uint_hashes = [10, 20, 50, 80]
mph = bbhash.PyMPHF(uint_hashes, len(uint_hashes), 1, 1.0)
for val in uint_hashes:
print('{} hashes to {}'.format(val, mph.lookup(val)))
Verify before relying
- Thread safety of PyMPHF and BBHashTable is noted as needing investigation by the maintainer.
- Whether the package is suitable for production use in bioinformatics pipelines beyond the original spacegraphcats use case.
- Performance characteristics and memory overhead compared to alternative hash table implementations.
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.11 |
| Install friction | High. Source build required |
| Runtime dependencies | 1 packagenumpy |
| Maintenance | Aging 292 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 87,699 / month, #13,775 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: bbhash-0.6.0.tar.gz
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “minimal perfect hash function”
- bbhashBuilds minimal perfect hash functions for 64-bit hashes using BBHash,…
- murmurhashProvides fast MurmurHash2 hashing through Cython bindings, enabling…
- pyfarmhashProvides Python bindings to Google's FarmHash, a fast…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also mmhash3 · mmh3 · py-multihash · pyfarmhash · xxhash · murmurhash · filehash · dict-hash · cityhash · blurhash-python