segtok
sentence segmentation and word tokenization tools
Decision gist · record as of 2026-08-14
No—the description explicitly recommends syntok as a successor that fixes known issues with sentence splitting. segtok itself is abandoned (last commit 2021-12-15), and while it has low install friction and no known vulnerabilities, choosing an unmaintained package when a maintained successor exists is not justified.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- The regex dependency requires python-dev or python3-dev headers on Linux systems to compile.
- Low friction install with a single regex dependency.
- However, the package is abandoned—last release was 2021-12-15 and no commits since.
License · maintenance · safety
MIT (permissive) — MIT license is permissive and imposes no restrictions on use, modification, or distribution in proprietary or open-source projects.
last release 2021-12-15 (1703 days) · last repo commit 2021-12-15 · 171 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 767,397 downloads/mo, #5,116 on PyPI
Alternatives
Verify before relying
pip install segtok
from segtok.segmenter import split_sentences
from segtok.tokenizer import web_tokenizer
text = "Hello world. This is a test."
for sentence in split_sentences(text):
tokens = list(web_tokenizer(sentence))
print(tokens)- Whether the known issues mentioned in the description (sentence splitting with terminals not followed by spaces) affect your use case
- Current compatibility with Python versions beyond 3.8, given the package is abandoned and classifiers list only up to 3.8
What it is and what it does
segtok provides two core modules for breaking down text: segtok.segmenter splits text into sentences, and segtok.tokenizer breaks sentences into words and symbols. It is designed for Indo-European languages (English, Spanish, German) and includes command-line tools for processing plain-text files. The package depends only on regex and installs with low friction.
The segmenter handles sentence boundaries including abbreviations, numbers, and edge cases like terminals followed by invalid characters. The tokenizer offers multiple strategies, with web_tokenizer providing semantic splitting while preserving URLs and email addresses. It also includes utilities for handling English contractions and possessive markers. However, the package has been abandoned since 2021-12-15, and the description explicitly recommends syntok as a successor that fixes tricky splitting issues.
Use it for
- Preprocessing text corpora for NLP pipelines that require sentence and word-level boundaries before downstream processing
- Batch processing plain-text documents via command-line to normalize and split text for indexing or analysis
- Extracting tokens from multilingual documents where Indo-European language support is sufficient and rule-based segmentation is preferred
- Splitting English text with contractions and possessives using built-in pattern matching and splitting functions
- Normalizing line breaks and sentence boundaries in documents with irregular formatting before further processing
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
No—the description explicitly recommends syntok as a successor that fixes known issues with sentence splitting.
segtok itself is abandoned (last commit 2021-12-15), and while it has low install friction and no known vulnerabilities, choosing an unmaintained package when a maintained successor exists is not justified.
Install
segtok on PyPI
Before you install
Low friction install with a single regex dependency. However, the package is abandoned—last release was 2021-12-15 and no commits since.
The regex dependency requires python-dev or python3-dev headers on Linux systems to compile.
License in practice
MIT license is permissive and imposes no restrictions on use, modification, or distribution in proprietary or open-source projects.
Quickstart
pip install segtok
from segtok.segmenter import split_sentences
from segtok.tokenizer import web_tokenizer
text = "Hello world. This is a test."
for sentence in split_sentences(text):
tokens = list(web_tokenizer(sentence))
print(tokens)
Verify before relying
- Whether the known issues mentioned in the description (sentence splitting with terminals not followed by spaces) affect your use case
- Current compatibility with Python versions beyond 3.8, given the package is abandoned and classifiers list only up to 3.8
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packageregex |
| Maintenance | Abandoned 1,703 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 767,397 / month, #5,116 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 2Programming Language :: Python :: 2.7Programming Language :: Python :: 3Programming Language :: Python :: 3.5Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Topic :: Scientific/Engineering :: Information AnalysisTopic :: Software Development :: LibrariesTopic :: Text ProcessingTopic :: Text Processing :: Linguistic |
Evidence: segtok-1.5.11-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
An agent finds packages by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language. Give your agent the search over MCP.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also minisbd · razdel · tokenizer · wordsegment · sentence-stream · pysbd · PyRuSH · tinysegmenter · segments · jieba3k