$npx skillfedfor your agent

razdel

Splits russian text into tokens, sentences, section. Rule-based

With conditionsPyPI LinguisticReleased Mar 2020130.5K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — razdel-0.5.0-py3-none-any.whl
v0.5.0 · released 2020-03-26

Yes, if your text is primarily Russian news or fiction. The package is lightweight, actively maintained, permissively licensed, and offers competitive accuracy on its target domains. Install with caution if your text is from social media, legal documents, or scientific articles—the description explicitly warns performance may degrade outside news/fiction.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Low friction: pure Python wheel with no runtime dependencies.
  • Actively maintained as of April 2026 with 286 repository stars.

License · maintenance · safety

MIT (permissive) — MIT license permits unrestricted use, modification, and distribution with minimal attribution requirements.

last release 2020-03-26 (2332 days) · last repo commit 2026-04-13 · 286 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 130,482 downloads/mo, #11,639 on PyPI

Verify before relying

from razdel import tokenize, sentenize

tokens = list(tokenize('Кружка-термос на 0.5л'))
sents = list(sentenize('Так в чем же дело? Не радуют.'))
  • Whether performance on non-news/fiction domains (social media, legal, scientific) is documented or measurable
  • Exact Python version support range beyond the '3.5+' claim in the description
Same gist for agents: .md · .json

What it is and what it does

razdel is a rule-based tokenizer and sentence splitter designed specifically for Russian text. It provides two main functions: tokenize() breaks text into words and punctuation with character offsets, and sentenize() splits text into sentences. The package returns Substring objects that preserve the original text span, making it easy to map results back to the source.

The library is optimized against four Russian corpora (SynTagRus, OpenCorpora, GICRYA, RNC) consisting mainly of news and fiction. It trades off perfect accuracy for practical performance: the description acknowledges that tokenization has no single correct answer and documents its segmentation choices explicitly. Benchmarks show it achieves competitive error rates on these domains while maintaining reasonable speed.

Use it for

  • Preprocess Russian news articles or literary texts before feeding into downstream NLP models
  • Extract individual words and punctuation from Russian text while preserving their exact positions for annotation or highlighting
  • Split Russian documents into sentences for batch processing or analysis at the sentence level
  • Build Russian text search or indexing pipelines that require accurate token boundaries
  • Evaluate or compare tokenization quality on Russian corpora using the included benchmarking tools

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if your text is primarily Russian news or fiction.

The package is lightweight, actively maintained, permissively licensed, and offers competitive accuracy on its target domains. Install with caution if your text is from social media, legal documents, or scientific articles—the description explicitly warns performance may degrade outside news/fiction.

Install

razdel on PyPI

Before you install

Low friction: pure Python wheel with no runtime dependencies. Actively maintained as of April 2026 with 286 repository stars.

License in practice

MIT license permits unrestricted use, modification, and distribution with minimal attribution requirements.

Quickstart

from razdel import tokenize, sentenize

tokens = list(tokenize('Кружка-термос на 0.5л'))
sents = list(sentenize('Так в чем же дело? Не радуют.'))

Verify before relying

  • Whether performance on non-news/fiction domains (social media, legal, scientific) is documented or measurable
  • Exact Python version support range beyond the '3.5+' claim in the description

Package facts

LicenseMIT permissive
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependenciesNone
MaintenanceActively maintained 2,332 days since the last release
Last repo commit
First released
Downloads130,482 / month, #11,639 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
License :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3

Evidence: razdel-0.5.0-py3-none-any.whl

Tags

Capabilities
russian text tokenizationrussian sentence segmentationcyrillic word tokenizerrussian nlp preprocessingrussian text splittingmorphological tokenization russiansentence boundary detection russian
Topics
russian-nlprule-basedtokenization
PyPI keywords
nlpnatural language processingrussiantokensentencetokenize

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “russian text tokenization”

  • razdelSplits Russian text into sentences and tokens using rule-based…
  • pymorphy2-dicts-ruProvides Russian morphological dictionaries for the pymorphy2…
  • pymorphy3Morphological analyzer and POS tagger for Russian and Ukrainian text…

Give your agent the search over MCP, or paste the wish link into any chat.

More Linguistic packages

charset-normalizer Worth it
PyPI · Utilities · released Aug 2026

Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.

permissive licensepure Python · 3.7+
1.7Bdownloads / mo
tiktoken Worth it
PyPI · Linguistic · released May 2026

tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.

Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.

permissive licensecompiled wheel · 3.9+
233.0Mdownloads / mo
chardet Worth it
PyPI · Python Modules · released Aug 2026

Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.

Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.

0BSDpure Python · 3.10+
199.0Mdownloads / mo
text-unidecode With conditions
PyPI · Python Modules · released Aug 2019

Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.

However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.

GPL-2.0-or-laterpure Pythonabandoned
89.0Mdownloads / mo
lark Worth it
PyPI · Python Modules · released Oct 2025

Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.

MITpure Python · 3.8+
79.7Mdownloads / mo
tree-sitter Worth it
PyPI · Linguistic · released Jun 2026

Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.

MITcompiled wheel · 3.10+
79.0Mdownloads / mo

See also segtok · PyRuSH · tokenizer · pysbd · segments · sacremoses · sentence-stream · jieba3k · hassil · pymorphy3-dicts-ru