wordsegment
English word segmentation.
Decision gist · record as of 2026-08-14
Yes, if you need English word segmentation and can tolerate an unmaintained package. The library is stable, has no dependencies, and works reliably for its narrow use case. However, the 2960-day gap since the last release means no updates for modern Python versions, security patches, or corpus improvements. Install only if the 2018-era corpus data and Python 3.6 compatibility are sufficient for your task.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- The load() function must be called once to read corpus data from disk before segmenting; this is a one-time setup cost per process.
- Installation is straightforward with no runtime dependencies.
- However, the package has been abandoned since 2018 (2960 days without release) and is no longer maintained, so expect no bug fixes or updates.
License · maintenance · safety
Apache 2.0 (permissive) — Licensed under Apache 2.0 (permissive), allowing commercial and private use with minimal restrictions—you may use, modify, and distribute freely provided you include the license notice.
last release 2018-07-07 (2960 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 123,744 downloads/mo, #11,901 on PyPI
Alternatives
Verify before relying
pip install wordsegment
from wordsegment import load, segment
load()
result = segment('thisisatest')
print(result) # ['this', 'is', 'a', 'test']- Whether the corpus data (333,000 unigrams, 250,000 bigrams) remains adequate for modern English or specialized domains.
- Performance characteristics on very long texts or high-volume batch processing.
- Compatibility with Python versions beyond 3.6, given the package's age and lack of maintenance.
What it is and what it does
WordSegment is a pure-Python library that reconstructs word boundaries in English text where spaces have been removed or are ambiguous. It uses unigram and bigram frequency data derived from the Google Web Trillion Word Corpus to score candidate segmentations and select the most likely split. The package includes 333,000 unigrams and 250,000 bigrams, all lowercased with punctuation removed, and provides both a programmatic API and a command-line interface for batch processing.
The library works by loading corpus frequency tables into memory and then using dynamic programming to find the segmentation that maximizes the product of word frequencies. It handles the canonical form of input (lowercasing, punctuation removal) automatically via a clean() function. Maximum word length is 24 characters, matching the corpus data. The package has no external dependencies and runs on both Python 2 and 3, though it has not been updated since 2018.
Use it for
- Recover readable words from concatenated text like 'thequickbrownfox' in OCR or data-cleaning pipelines.
- Process hashtags or domain names where word boundaries are unclear, splitting '#pythonrocks' into candidate words.
- Batch-segment large files of unsegmented English text via the command-line interface.
- Analyze corpus frequency data directly (UNIGRAMS and BIGRAMS dictionaries) for linguistic research or spelling preference studies.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need English word segmentation and can tolerate an unmaintained package.
The library is stable, has no dependencies, and works reliably for its narrow use case. However, the 2960-day gap since the last release means no updates for modern Python versions, security patches, or corpus improvements. Install only if the 2018-era corpus data and Python 3.6 compatibility are sufficient for your task.
Install
wordsegment on PyPI
Before you install
Installation is straightforward with no runtime dependencies. However, the package has been abandoned since 2018 (2960 days without release) and is no longer maintained, so expect no bug fixes or updates.
The load() function must be called once to read corpus data from disk before segmenting; this is a one-time setup cost per process.
License in practice
Licensed under Apache 2.0 (permissive), allowing commercial and private use with minimal restrictions—you may use, modify, and distribute freely provided you include the license notice.
Quickstart
pip install wordsegment
from wordsegment import load, segment
load()
result = segment('thisisatest')
print(result) # ['this', 'is', 'a', 'test']
Verify before relying
- Whether the corpus data (333,000 unigrams, 250,000 bigrams) remains adequate for modern English or specialized domains.
- Performance characteristics on very long texts or high-volume batch processing.
- Compatibility with Python versions beyond 3.6, given the package's age and lack of maintenance.
Package facts
| License | Apache 2.0 permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Abandoned 2,960 days since the last release |
| First released | |
| Downloads | 123,744 / month, #11,901 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: Apache Software LicenseNatural Language :: EnglishProgramming Language :: PythonProgramming Language :: Python :: 2Programming Language :: Python :: 2.6Programming Language :: Python :: 2.7Programming Language :: Python :: 3Programming Language :: Python :: 3.2Programming Language :: Python :: 3.3Programming Language :: Python :: 3.4Programming Language :: Python :: 3.5Programming Language :: Python :: 3.6 |
Evidence: wordsegment-1.3.1-py2.py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “unsegmented text parsing”
- wordsegmentSplits unsegmented English text into individual words using a…
- symspellpysymspellpy is a Python port of SymSpell v6.7.2 that performs fast…
- ttpTTP is a Python library for parsing semi-structured text into…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also segtok · jieba3k · curated-tokenizers · jieba · english-words · segments · wordninja · uniseg · unicode-segmentation-rs