grapheme
Unicode grapheme helpers
Decision gist · record as of 2026-08-14
No. The package is abandoned (last release 2020-03-07, last commit 2022-03-21) and carries high installation friction due to compilation requirements. While it solves a real problem—correct grapheme handling—the lack of maintenance means compatibility issues with newer Python versions or Unicode standards will not be fixed. Consider it only if you are locked into an older Python environment and have no alternative; otherwise, seek an actively maintained grapheme library or implement grapheme logic inline if your use case is narrow.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires compilation during installation; no explicit Python version requirement stated, but classifiers list support through Python 3.8.
- Installation requires compilation (high friction).
- The package is abandoned—last release was 2020-03-07, last commit 2022-03-21—and classifiers indicate Alpha status.
License · maintenance · safety
MIT (permissive) — MIT license is permissive and poses no restrictions on use, modification, or distribution.
last release 2020-03-07 (2351 days) · last repo commit 2022-03-21 · 116 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,192,529 downloads/mo, #4,237 on PyPI
Alternatives
Verify before relying
pip install grapheme
import grapheme
string = 'u̲n̲d̲e̲r̲l̲i̲n̲e̲d̲'
print(grapheme.length(string)) # 10 (user-perceived characters)
print(grapheme.substr(string, 0, 3)) # 'u̲n̲d̲'- Whether the package works reliably with Python versions beyond 3.8 (classifiers stop there, but no explicit upper bound is documented).
- Current Unicode Standard Annex #29 compliance status—the package targets Unicode 13.0.0, but no statement on whether later Unicode versions are supported.
What it is and what it does
grapheme is a Python library for working with grapheme clusters—the user-perceived characters that the Unicode Standard defines—rather than raw Unicode code points. Standard Python string functions treat each Unicode code point as a separate unit, which breaks strings containing combining marks (like underlines or accents), emoji with skin-tone modifiers, Korean Hangul, and other multi-codepoint sequences. This library implements the Unicode default rules for extended grapheme clusters and provides functions like `length()`, `substr()`, `slice()`, and `contains()` that operate on graphemes instead.
The package is useful when you need to count, truncate, or format text the way users actually see it—for example, when building text-based tables in monospaced fonts or ensuring that user input doesn't corrupt multi-codepoint characters. Performance scales linearly with string length, and the library is designed for short strings or the beginning of long strings; the documentation notes that grapheme calculation is notably slower than counting code points and recommends using standard Python functions when performance is prioritized over correctness.
Use it for
- Count user-perceived character length in strings with combining marks or emoji modifiers without overcounting code points.
- Truncate or slice text at user-perceived boundaries to avoid splitting multi-codepoint characters and corrupting display.
- Format text-based tables or monospaced output by actual visible character width rather than Unicode code point count.
- Validate user input length constraints based on what users actually see rather than internal Unicode representation.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
No.
The package is abandoned (last release 2020-03-07, last commit 2022-03-21) and carries high installation friction due to compilation requirements. While it solves a real problem—correct grapheme handling—the lack of maintenance means compatibility issues with newer Python versions or Unicode standards will not be fixed. Consider it only if you are locked into an older Python environment and have no alternative; otherwise, seek an actively maintained grapheme library or implement grapheme logic inline if your use case is narrow.
Install
grapheme on PyPI
Before you install
Installation requires compilation (high friction). The package is abandoned—last release was 2020-03-07, last commit 2022-03-21—and classifiers indicate Alpha status. No runtime dependencies, but no active maintenance means security or compatibility issues will not be addressed.
Requires compilation during installation; no explicit Python version requirement stated, but classifiers list support through Python 3.8.
License in practice
MIT license is permissive and poses no restrictions on use, modification, or distribution.
Quickstart
pip install grapheme
import grapheme
string = 'u̲n̲d̲e̲r̲l̲i̲n̲e̲d̲'
print(grapheme.length(string)) # 10 (user-perceived characters)
print(grapheme.substr(string, 0, 3)) # 'u̲n̲d̲'
Verify before relying
- Whether the package works reliably with Python versions beyond 3.8 (classifiers stop there, but no explicit upper bound is documented).
- Current Unicode Standard Annex #29 compliance status—the package targets Unicode 13.0.0, but no statement on whether later Unicode versions are supported.
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | High. Source build required |
| Runtime dependencies | None |
| Maintenance | Abandoned 2,351 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,192,529 / month, #4,237 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaIntended Audience :: DevelopersProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.5Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8 |
Evidence: grapheme-0.6.0.tar.gz
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “grapheme cluster string handling”
- graphemeProvides string manipulation functions that work with grapheme…
- graphemeuProvides string manipulation functions that work with user-perceived…
- wcwidthMeasures the displayed width of Unicode strings in terminals,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also graphemeu · uniseg · unicode-segmentation-rs · emoji · anyascii · unicodedataplus · confusables · demoji · segments · confusable-homoglyphs