simhash
A Python implementation of Simhash Algorithm
Decision gist · record as of 2026-08-14
Yes, if you need Simhash specifically and can verify it works with your Python and numpy versions. The algorithm is well-established and the package is stable, but abandonment means no fixes for future compatibility issues. Suitable for projects where you control dependencies and can test thoroughly before deployment.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low friction install with only numpy as a runtime dependency.
- However, the package is abandoned—last release was 2022-03-03 and last commit 2022-03-24, with no active maintenance.
License · maintenance · safety
MIT License (permissive) — MIT License permits free use, modification, and distribution with minimal restrictions, making it safe to incorporate into most projects without licensing concerns.
last release 2022-03-03 (1625 days) · last repo commit 2022-03-24 · 1,038 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 325,911 downloads/mo, #7,581 on PyPI
Alternatives
Verify before relying
pip install simhash
from simhash import Simhash
hash1 = Simhash('This is a test document').value
hash2 = Simhash('This is a test document with minor changes').value- Whether the package works correctly with current numpy versions and modern Python releases
- Performance characteristics on large document collections or real-world deduplication workloads
What it is and what it does
Simhash is a Python implementation of Google's Simhash algorithm, which generates a compact fingerprint for text documents. Rather than comparing full documents, Simhash produces a fixed-size hash that allows fast detection of near-duplicate content—documents with small differences will have similar hashes. This is useful for deduplication in web crawlers, duplicate detection in databases, and content similarity analysis.
The package depends on numpy for numerical operations. It has been stable since its last release in 2022 but is no longer actively maintained. With 1038 GitHub stars and consistent monthly downloads, it remains a recognized tool for this specific task, though you should verify compatibility with your Python and numpy versions before relying on it in production.
Use it for
- Detect near-duplicate web pages or documents in a crawl or corpus to avoid storing redundant content.
- Identify similar text submissions in bulk uploads to flag potential plagiarism or spam.
- Deduplicate large document collections by grouping content with similar Simhash values.
- Build a similarity index for fast approximate matching of new documents against a known set.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need Simhash specifically and can verify it works with your Python and numpy versions.
The algorithm is well-established and the package is stable, but abandonment means no fixes for future compatibility issues. Suitable for projects where you control dependencies and can test thoroughly before deployment.
Install
simhash on PyPI
Before you install
Low friction install with only numpy as a runtime dependency. However, the package is abandoned—last release was 2022-03-03 and last commit 2022-03-24, with no active maintenance.
License in practice
MIT License permits free use, modification, and distribution with minimal restrictions, making it safe to incorporate into most projects without licensing concerns.
Quickstart
pip install simhash
from simhash import Simhash
hash1 = Simhash('This is a test document').value
hash2 = Simhash('This is a test document with minor changes').value
Verify before relying
- Whether the package works correctly with current numpy versions and modern Python releases
- Performance characteristics on large document collections or real-world deduplication workloads
Package facts
| License | MIT License permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagenumpy |
| Maintenance | Abandoned 1,625 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 325,911 / month, #7,581 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: simhash-2.1.2-py2-none-any.whl; simhash-2.1.2-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “simhash duplicate detection”
- simhashComputes Simhash fingerprints for text and documents to detect…
- django-queryinspectDjango middleware that logs and reports SQL query statistics for each…
- imagededupFinds exact and near-duplicate images in collections using hashing…
Give your agent the search over MCP, or paste the wish link into any chat.
More Information Analysis packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.
Install it if you use Snowflake and want to process data without moving it to your application layer.
See also ImageHash · ppdeep · imagededup · rensa · fnvhash · fingerprints · py-tlsh · semhash · mhfp · fnv-hash-fast