strsimpy
A library implementing different string similarity and distance measures
Decision gist · record as of 2026-08-14
Yes, if you need string similarity algorithms and can accept an abandoned package. The library is stable, has no known vulnerabilities, low install friction, and covers a comprehensive range of algorithms. However, do not install if you require active maintenance, bug fixes, or support for Python versions beyond 3.9—the last release was 2021-09-10 with no commits since November 2022.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with no runtime dependencies.
- However, the package is abandoned—last release was 2021-09-10 and last commit 2022-11-12—so no active maintenance or bug fixes should be expected.
License · maintenance · safety
MIT License (permissive) — MIT License permits commercial and private use with minimal restrictions, requiring only license and copyright notice retention.
last release 2021-09-10 (1799 days) · last repo commit 2022-11-12 · 1,018 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 513,251 downloads/mo, #6,248 on PyPI
Alternatives
Verify before relying
pip install strsimpy
from strsimpy.jaro_winkler import JaroWinkler
jw = JaroWinkler()
similarity = jw.similarity('string1', 'string2')- Whether the package works reliably with Python versions beyond 3.9 despite classifiers only listing up to 3.9
- Performance characteristics on very long strings or large-scale batch comparisons
- Whether all dozen algorithms mentioned in the description are fully implemented and documented
What it is and what it does
strsimpy is a Python port of a Java string-similarity library that provides implementations of multiple string comparison algorithms. It lets you measure how similar or different two strings are using different mathematical approaches—some compute edit distance (how many character changes are needed), others use n-gram profiles or set-based metrics. Each algorithm has different trade-offs: some are normalized to 0–1 ranges, some satisfy the triangle inequality property useful for indexing, and some are optimized for specific tasks like typo correction or diff utilities.
You import an algorithm class, instantiate it, and call its similarity() or distance() method on two strings. The library includes Levenshtein variants (standard, normalized, weighted, Damerau), Jaro-Winkler, Longest Common Subsequence, n-gram based methods (Q-Gram, cosine, Jaccard, Sorensen-Dice, overlap), and experimental SIFT4. It has no external dependencies and installs as a pure Python wheel.
Use it for
- Detect typos and suggest corrections by finding the most similar string from a dictionary using Jaro-Winkler
- Implement a fuzzy search feature that matches user input against product names or database records despite spelling variations
- Compare file contents or code diffs by computing Longest Common Subsequence distance
- Build a deduplication system that identifies near-duplicate records in a dataset using normalized Levenshtein
- Analyze OCR output by using weighted Levenshtein to account for character-specific error costs
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need string similarity algorithms and can accept an abandoned package.
The library is stable, has no known vulnerabilities, low install friction, and covers a comprehensive range of algorithms. However, do not install if you require active maintenance, bug fixes, or support for Python versions beyond 3.9—the last release was 2021-09-10 with no commits since November 2022.
Install
strsimpy on PyPI
Before you install
Low install friction with no runtime dependencies. However, the package is abandoned—last release was 2021-09-10 and last commit 2022-11-12—so no active maintenance or bug fixes should be expected.
License in practice
MIT License permits commercial and private use with minimal restrictions, requiring only license and copyright notice retention.
Quickstart
pip install strsimpy
from strsimpy.jaro_winkler import JaroWinkler
jw = JaroWinkler()
similarity = jw.similarity('string1', 'string2')
Verify before relying
- Whether the package works reliably with Python versions beyond 3.9 despite classifiers only listing up to 3.9
- Performance characteristics on very long strings or large-scale batch comparisons
- Whether all dozen algorithms mentioned in the description are fully implemented and documented
Package facts
| License | MIT License permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Abandoned 1,799 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 513,251 / month, #6,248 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.5Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9 |
Evidence: strsimpy-0.2.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “string similarity distance”
- strsimpyImplements a dozen string similarity and distance algorithms…
- python-LevenshteinComputes Levenshtein edit distance, string similarity, and…
- pylevComputes the Levenshtein distance between two strings—the minimum…
Give your agent the search over MCP, or paste the wish link into any chat.
More Text Processing packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.
Install it if you need to parse text or define grammars programmatically.
fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.
Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.
Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.
tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.
Install it if your project requires CSS tokenization or syntax manipulation.
See also pylcs · stringzilla · suffix-trees · textdistance · fuzzysearch · Distance · editdistpy · jaro-winkler · zss · ngram