unicode-segmentation-rs
Unicode segmentation and width for Python using Rust
Decision gist · record as of 2026-08-14
Yes, with a license caveat. The package is actively maintained, has no known vulnerabilities, installs cleanly on common platforms, and solves a real problem (correct Unicode text segmentation) that pure Python solutions handle poorly. However, verify the license status before use in proprietary projects, since the metadata does not declare one.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later; precompiled wheels available for most platforms, but source builds require maturin and Rust.
- Medium install friction due to compiled wheels; however, prebuilt binaries are available for common platforms (x86_64, ARM, PowerPC, s390x on Linux/macOS/Windows).
- Package is actively maintained with recent releases.
License · maintenance · safety
(unclear) — License status is unclear—no SPDX identifier or raw license text is recorded in the package metadata. Verify the actual license before use in proprietary or restricted contexts.
last release 2026-08-08 (6 days) · last repo commit 2026-08-14 · 2 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 265,647 downloads/mo, #8,320 on PyPI
Alternatives
Verify before relying
pip install unicode-segmentation-rs
import unicode_segmentation_rs
text = "Hello 👨👩👧👦 World"
clusters = unicode_segmentation_rs.graphemes(text, is_extended=True)
print(clusters)- Whether the package is licensed under an open-source or proprietary license (metadata shows 'unclear').
- Performance characteristics and memory overhead for very large text inputs.
- Whether PyPy support (listed in classifiers) is fully tested and stable.
What it is and what it does
unicode-segmentation-rs is a Python library that wraps Rust implementations of Unicode text processing algorithms. It splits text into grapheme clusters (user-perceived characters, handling complex emojis and combining marks), words, sentences, and line-break opportunities according to Unicode standards. It also calculates display width for terminal or monospace rendering, and provides gettext PO file wrapping with proper handling of escape sequences and CJK characters.
The package has no Python runtime dependencies and installs as a compiled extension, making it fast for text processing tasks common in localization, terminal UI, and multilingual applications. It supports modern Python versions (3.10+) and is actively maintained.
Use it for
- Split text into grapheme clusters to correctly count user-perceived characters in strings with complex emojis or combining diacritics.
- Find word and sentence boundaries for text analysis, search indexing, or natural language processing pipelines.
- Calculate display width of text for terminal wrapping, monospace layout, or UI rendering without measuring individual glyphs.
- Wrap localization strings for gettext PO files while preserving escape sequences and respecting CJK character boundaries.
- Identify legal line-break opportunities in multilingual text for text editors or document formatters.
- Process Arabic, Japanese, Chinese, and other non-Latin scripts with Unicode-compliant segmentation rules.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with a license caveat.
The package is actively maintained, has no known vulnerabilities, installs cleanly on common platforms, and solves a real problem (correct Unicode text segmentation) that pure Python solutions handle poorly. However, verify the license status before use in proprietary projects, since the metadata does not declare one.
Install
unicode-segmentation-rs on PyPI
Before you install
Medium install friction due to compiled wheels; however, prebuilt binaries are available for common platforms (x86_64, ARM, PowerPC, s390x on Linux/macOS/Windows). Package is actively maintained with recent releases.
Requires Python 3.10 or later; precompiled wheels available for most platforms, but source builds require maturin and Rust.
License in practice
License status is unclear—no SPDX identifier or raw license text is recorded in the package metadata. Verify the actual license before use in proprietary or restricted contexts.
Quickstart
pip install unicode-segmentation-rs
import unicode_segmentation_rs
text = "Hello 👨👩👧👦 World"
clusters = unicode_segmentation_rs.graphemes(text, is_extended=True)
print(clusters)
Verify before relying
- Whether the package is licensed under an open-source or proprietary license (metadata shows 'unclear').
- Performance characteristics and memory overhead for very large text inputs.
- Whether PyPy support (listed in classifiers) is fully tested and stable.
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 6 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 265,647 / month, #8,320 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.15Programming Language :: Python :: Free Threading :: 3 - StableProgramming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPyProgramming Language :: RustTopic :: Software Development :: Libraries :: Python Modules |
Evidence: unicode_segmentation_rs-0.3.3-cp310-abi3-macosx_10_12_x86_64.whl; unicode_segmentation_rs-0.3.3-cp310-abi3-macosx_11_0_arm64.whl; unicode_segmentation_rs-0.3.3-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; unicode_segmentation_rs-0.3.3-cp310-abi3-manylinux_2_17_armv7l.manylinux2014_armv7l.whl; unicode_segmentation_rs-0.3.3-cp310-abi3-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl; unicode_segmentation_rs-0.3.3-cp310-abi3-manylinux_2_17_s390x.manylinux2014_s390x.whl; unicode_segmentation_rs-0.3.3-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; unicode_segmentation_rs-0.3.3-cp310-abi3-manylinux_2_5_i686.manylinux1_i686.whl; unicode_segmentation_rs-0.3.3-cp310-abi3-musllinux_1_2_aarch64.whl; unicode_segmentation_rs-0.3.3-cp310-abi3-musllinux_1_2_armv7l.whl; unicode_segmentation_rs-0.3.3-cp310-abi3-musllinux_1_2_i686.whl; unicode_segmentation_rs-0.3.3-cp310-abi3-musllinux_1_2_x86_64.whl; unicode_segmentation_rs-0.3.3-cp310-abi3-win32.whl; unicode_segmentation_rs-0.3.3-cp310-abi3-win_amd64.whl; unicode_segmentation_rs-0.3.3-cp314-cp314t-macosx_10_12_x86_64.whl; unicode_segmentation_rs-0.3.3-cp314-cp314t-macosx_11_0_arm64.whl; unicode_segmentation_rs-0.3.3-cp314-cp314t-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; unicode_segmentation_rs-0.3.3-cp314-cp314t-manylinux_2_17_armv7l.manylinux2014_armv7l.whl; unicode_segmentation_rs-0.3.3-cp314-cp314t-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl; unicode_segmentation_rs-0.3.3-cp314-cp314t-manylinux_2_17_s390x.manylinux2014_s390x.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “grapheme cluster splitting”
- unicode-segmentation-rsProvides Unicode text segmentation (graphemes, words, sentences, line…
- graphemeProvides string manipulation functions that work with grapheme…
- unisegDetermines Unicode text segmentation boundaries—grapheme clusters,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Python Modules packages
Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.
Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.
Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.
PyYAML parses and emits YAML 1.1 data format, enabling serialization and deserialization of configuration files and Python objects to and from human-readable YAML text.
Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.
Provides reusable metadata objects for use with PEP-593 `typing.Annotated` to express common constraints like bounds, collection sizes, and predicates on types.
Install it if you use or build libraries that need to express type constraints in a standardized, inspectable way—or if you want to annotate your own types with…
Provides runtime tools to inspect and introspect Python type annotations, enabling programmatic examination of type hints at execution time.
See also uniseg · segments · graphemeu · grapheme · segtok · unicodedataplus · translation-finder · tokenizer · weblate-fonts · py-rust-stemmers