razdel
Splits russian text into tokens, sentences, section. Rule-based
Decision gist · record as of 2026-08-14
Yes, if your text is primarily Russian news or fiction. The package is lightweight, actively maintained, permissively licensed, and offers competitive accuracy on its target domains. Install with caution if your text is from social media, legal documents, or scientific articles—the description explicitly warns performance may degrade outside news/fiction.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low friction: pure Python wheel with no runtime dependencies.
- Actively maintained as of April 2026 with 286 repository stars.
License · maintenance · safety
MIT (permissive) — MIT license permits unrestricted use, modification, and distribution with minimal attribution requirements.
last release 2020-03-26 (2332 days) · last repo commit 2026-04-13 · 286 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 130,482 downloads/mo, #11,639 on PyPI
Alternatives
Verify before relying
from razdel import tokenize, sentenize
tokens = list(tokenize('Кружка-термос на 0.5л'))
sents = list(sentenize('Так в чем же дело? Не радуют.'))- Whether performance on non-news/fiction domains (social media, legal, scientific) is documented or measurable
- Exact Python version support range beyond the '3.5+' claim in the description
What it is and what it does
razdel is a rule-based tokenizer and sentence splitter designed specifically for Russian text. It provides two main functions: tokenize() breaks text into words and punctuation with character offsets, and sentenize() splits text into sentences. The package returns Substring objects that preserve the original text span, making it easy to map results back to the source.
The library is optimized against four Russian corpora (SynTagRus, OpenCorpora, GICRYA, RNC) consisting mainly of news and fiction. It trades off perfect accuracy for practical performance: the description acknowledges that tokenization has no single correct answer and documents its segmentation choices explicitly. Benchmarks show it achieves competitive error rates on these domains while maintaining reasonable speed.
Use it for
- Preprocess Russian news articles or literary texts before feeding into downstream NLP models
- Extract individual words and punctuation from Russian text while preserving their exact positions for annotation or highlighting
- Split Russian documents into sentences for batch processing or analysis at the sentence level
- Build Russian text search or indexing pipelines that require accurate token boundaries
- Evaluate or compare tokenization quality on Russian corpora using the included benchmarking tools
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if your text is primarily Russian news or fiction.
The package is lightweight, actively maintained, permissively licensed, and offers competitive accuracy on its target domains. Install with caution if your text is from social media, legal documents, or scientific articles—the description explicitly warns performance may degrade outside news/fiction.
Install
razdel on PyPI
Before you install
Low friction: pure Python wheel with no runtime dependencies. Actively maintained as of April 2026 with 286 repository stars.
License in practice
MIT license permits unrestricted use, modification, and distribution with minimal attribution requirements.
Quickstart
from razdel import tokenize, sentenize
tokens = list(tokenize('Кружка-термос на 0.5л'))
sents = list(sentenize('Так в чем же дело? Не радуют.'))
Verify before relying
- Whether performance on non-news/fiction domains (social media, legal, scientific) is documented or measurable
- Exact Python version support range beyond the '3.5+' claim in the description
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 2,332 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 130,482 / month, #11,639 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3 |
Evidence: razdel-0.5.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “russian text tokenization”
- razdelSplits Russian text into sentences and tokens using rule-based…
- pymorphy2-dicts-ruProvides Russian morphological dictionaries for the pymorphy2…
- pymorphy3Morphological analyzer and POS tagger for Russian and Ukrainian text…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also segtok · PyRuSH · tokenizer · pysbd · segments · sacremoses · sentence-stream · jieba3k · hassil · pymorphy3-dicts-ru