textdistance
Compute distance between the two texts.
Decision gist · record as of 2026-08-14
Yes. The package is stable, permissively licensed, has no dependencies, and provides a comprehensive toolkit for a common task. The aging maintenance status (last release 759 days ago) is a minor concern but not a blocker—the library solves a well-defined problem with established algorithms, and the repository remains active and not archived. Install it if you need string distance or similarity computation.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low friction installation with no runtime dependencies.
- Maintenance status is aging—last release was 2024-07-16, with 759 days since then—but the package remains in production/stable state with an active repository (3538 stars, not archived).
License · maintenance · safety
MIT (permissive) — MIT license (permissive) means you can use this package freely in commercial and private projects with minimal restrictions.
last release 2024-07-16 (759 days) · last repo commit 2025-04-18 · 3,538 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,580,088 downloads/mo, #2,991 on PyPI
Alternatives
Verify before relying
pip install textdistance
from textdistance import Levenshtein
levenshtein = Levenshtein()
distance = levenshtein.distance('kitten', 'sitting')
print(distance) # 3- Whether optional numpy integration actually improves performance for the algorithms listed, and which ones benefit most
- Current state of the 'work in progress' compression-based algorithms (BZ2NCD, LZMANCD, ZLIBNCD) and whether they are production-ready
What it is and what it does
TextDistance is a pure-Python library that implements a large collection of algorithms for measuring how different two or more text sequences are from each other. It provides both distance metrics (how far apart sequences are) and similarity metrics (how alike they are), with support for edit-based algorithms like Levenshtein and Damerau-Levenshtein, token-based methods like Jaccard and Cosine similarity, sequence-based approaches like longest common subsequence, compression-based distance, and phonetic matching. The library has zero runtime dependencies by default and offers optional numpy integration for speed on specific algorithms.
You use it by instantiating an algorithm class or calling a function directly, then invoking methods like `.distance()`, `.similarity()`, or `.normalized_distance()` to compare sequences. It's designed for straightforward use cases where you need to measure text similarity—fuzzy matching, deduplication, spell-checking support, or record linkage—without needing to implement these algorithms yourself.
Use it for
- Fuzzy string matching to find similar product names or user entries despite typos or variations
- Deduplication of text records by computing similarity scores between candidates
- Spell-checking or autocorrect by ranking candidate corrections by edit distance
- Record linkage in data integration by comparing field values across datasets
- Phonetic matching for names that sound similar but are spelled differently
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
The package is stable, permissively licensed, has no dependencies, and provides a comprehensive toolkit for a common task. The aging maintenance status (last release 759 days ago) is a minor concern but not a blocker—the library solves a well-defined problem with established algorithms, and the repository remains active and not archived. Install it if you need string distance or similarity computation.
Install
textdistance on PyPI
Before you install
Low friction installation with no runtime dependencies. Maintenance status is aging—last release was 2024-07-16, with 759 days since then—but the package remains in production/stable state with an active repository (3538 stars, not archived).
License in practice
MIT license (permissive) means you can use this package freely in commercial and private projects with minimal restrictions.
Quickstart
pip install textdistance
from textdistance import Levenshtein
levenshtein = Levenshtein()
distance = levenshtein.distance('kitten', 'sitting')
print(distance) # 3
Verify before relying
- Whether optional numpy integration actually improves performance for the algorithms listed, and which ones benefit most
- Current state of the 'work in progress' compression-based algorithms (BZ2NCD, LZMANCD, ZLIBNCD) and whether they are production-ready
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.5 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Aging 759 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 2,580,088 / month, #2,991 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: PluginsIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: PythonTopic :: Scientific/Engineering :: Human Machine Interfaces |
Evidence: textdistance-4.6.3-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “edit distance algorithms”
- textdistanceComputes distance and similarity between text sequences using 30+…
- strsimpyImplements a dozen string similarity and distance algorithms…
- kaldialignComputes edit distance, alignment, and word error rate (WER) between…
Give your agent the search over MCP, or paste the wish link into any chat.
More Human Machine Interfaces packages
NLTK is a Python library for natural language processing tasks including tokenization, parsing, tagging, and linguistic analysis, with built-in datasets and educational resources.
Install it if you need foundational NLP tools, linguistic datasets, or are learning the field; consider specialized libraries (spaCy, transformers) if you need…
Formats and parses numbers, file sizes, timespans, and other values into human-readable text; provides terminal interaction utilities including ANSI text styling and user prompts.
Adds colored output to Python's standard logging module using ANSI escape sequences, with automatic fallback support for Windows terminals.
However, it is abandoned and receives no maintenance, so it will not adapt to future Python changes or platform updates.
Converts written-out number words (like "twenty one") into numeric digits (21), supporting positive integers and decimals up to 999,999,999,999.
No—install only if you have a legacy codebase already using it.
Detects voiced versus unvoiced segments in audio by wrapping Google's WebRTC Voice Activity Detector with pre-built binary wheels for Windows, macOS, and Linux.
Install it if you need reliable voice activity detection in Python.
Provides a Python interface to Google's WebRTC Voice Activity Detector, classifying audio frames as voiced or unvoiced for speech recognition and telephony applications.
However, high install friction (compiled extension), dormancy since 2017-01-07, and uncertainty about modern Python compatibility mean you should verify it builds on…
See also Distance · strsimpy · fuzzysearch · editdistance · pylev · jaro-winkler · pyjarowinkler · apted · jarowinkler · zss