$npx skillfedfor your agent

segtok

sentence segmentation and word tokenization tools

SkipPyPI LibrariesReleased Dec 2021767.4K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — segtok-1.5.11-py3-none-any.whl
v1.5.11 · released 2021-12-15 · 1 runtime deps: regex

No—the description explicitly recommends syntok as a successor that fixes known issues with sentence splitting. segtok itself is abandoned (last commit 2021-12-15), and while it has low install friction and no known vulnerabilities, choosing an unmaintained package when a maintained successor exists is not justified.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • The regex dependency requires python-dev or python3-dev headers on Linux systems to compile.
  • Low friction install with a single regex dependency.
  • However, the package is abandoned—last release was 2021-12-15 and no commits since.

License · maintenance · safety

MIT (permissive) — MIT license is permissive and imposes no restrictions on use, modification, or distribution in proprietary or open-source projects.

last release 2021-12-15 (1703 days) · last repo commit 2021-12-15 · 171 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 767,397 downloads/mo, #5,116 on PyPI

Verify before relying

pip install segtok

from segtok.segmenter import split_sentences
from segtok.tokenizer import web_tokenizer

text = "Hello world. This is a test."
for sentence in split_sentences(text):
    tokens = list(web_tokenizer(sentence))
    print(tokens)
  • Whether the known issues mentioned in the description (sentence splitting with terminals not followed by spaces) affect your use case
  • Current compatibility with Python versions beyond 3.8, given the package is abandoned and classifiers list only up to 3.8
Same gist for agents: .md · .json

What it is and what it does

segtok provides two core modules for breaking down text: segtok.segmenter splits text into sentences, and segtok.tokenizer breaks sentences into words and symbols. It is designed for Indo-European languages (English, Spanish, German) and includes command-line tools for processing plain-text files. The package depends only on regex and installs with low friction.

The segmenter handles sentence boundaries including abbreviations, numbers, and edge cases like terminals followed by invalid characters. The tokenizer offers multiple strategies, with web_tokenizer providing semantic splitting while preserving URLs and email addresses. It also includes utilities for handling English contractions and possessive markers. However, the package has been abandoned since 2021-12-15, and the description explicitly recommends syntok as a successor that fixes tricky splitting issues.

Use it for

  • Preprocessing text corpora for NLP pipelines that require sentence and word-level boundaries before downstream processing
  • Batch processing plain-text documents via command-line to normalize and split text for indexing or analysis
  • Extracting tokens from multilingual documents where Indo-European language support is sufficient and rule-based segmentation is preferred
  • Splitting English text with contractions and possessives using built-in pattern matching and splitting functions
  • Normalizing line breaks and sentence boundaries in documents with irregular formatting before further processing

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Skip

No—the description explicitly recommends syntok as a successor that fixes known issues with sentence splitting.

segtok itself is abandoned (last commit 2021-12-15), and while it has low install friction and no known vulnerabilities, choosing an unmaintained package when a maintained successor exists is not justified.

Install

segtok on PyPI

Before you install

Low friction install with a single regex dependency. However, the package is abandoned—last release was 2021-12-15 and no commits since.

The regex dependency requires python-dev or python3-dev headers on Linux systems to compile.

License in practice

MIT license is permissive and imposes no restrictions on use, modification, or distribution in proprietary or open-source projects.

Quickstart

pip install segtok

from segtok.segmenter import split_sentences
from segtok.tokenizer import web_tokenizer

text = "Hello world. This is a test."
for sentence in split_sentences(text):
    tokens = list(web_tokenizer(sentence))
    print(tokens)

Verify before relying

  • Whether the known issues mentioned in the description (sentence splitting with terminals not followed by spaces) affect your use case
  • Current compatibility with Python versions beyond 3.8, given the package is abandoned and classifiers list only up to 3.8

Package facts

LicenseMIT permissive
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependencies
1 package
regex
MaintenanceAbandoned 1,703 days since the last release
Last repo commit
First released
Downloads767,397 / month, #5,116 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 3 - AlphaLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 2Programming Language :: Python :: 2.7Programming Language :: Python :: 3Programming Language :: Python :: 3.5Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Topic :: Scientific/Engineering :: Information AnalysisTopic :: Software Development :: LibrariesTopic :: Text ProcessingTopic :: Text Processing :: Linguistic

Evidence: segtok-1.5.11-py3-none-any.whl

Tags

Capabilities
sentence segmentationword tokenizationtext splittingNLP preprocessingsentence splittertoken extractiontext segmentation
Topics
sentence-segmentationtokenizationtext-preprocessing
PyPI keywords
sentencesegmentersplittersplitwordtokenizertoken

Let your AI agent find packages like this

Example. Real query, live index.

An agent finds packages by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language. Give your agent the search over MCP.

More Libraries packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
python-dateutil Worth it
PyPI · Libraries · released Mar 2024

Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.

Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.

Apache-2.0pure Python
1.2Bdownloads / mo
six With conditions
PyPI · Libraries · released Dec 2024

Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.

MITpure Python
1.2Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo

See also minisbd · razdel · tokenizer · wordsegment · sentence-stream · pysbd · PyRuSH · tinysegmenter · segments · jieba3k