$npx skillfedfor your agent

zhon

Zhon provides constants used in Chinese text processing.

Worth itPyPI Python ModulesReleased Nov 2024571.8K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — zhon-2.1.1-py3-none-any.whl
v2.1.1 · released 2024-11-20 · Python >=3.9

Yes. Zhon is a stable, dependency-free reference library for Chinese text processing with no known vulnerabilities. Install it if you're building tools that need to identify, validate, or extract Chinese characters, Pinyin, or Zhuyin. The dormant maintenance status is not a concern for a constants-only package—the character sets it exports are unlikely to change frequently.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Low install friction with no runtime dependencies.
  • Dormant maintenance status (last release 2024-11-20, last commit 2024-12-11) but stable and widely used; no active development but no signs of abandonment.

License · maintenance · safety

MIT (permissive) — MIT license is permissive; you can use, modify, and distribute this package freely with minimal restrictions.

last release 2024-11-20 (632 days) · last repo commit 2024-12-11 · 386 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 571,763 downloads/mo, #5,949 on PyPI

Verify before relying

import re
import zhon.hanzi

# Find CJK characters in a string
re.findall('[{}]'.format(zhon.hanzi.characters), 'I broke a plate: 我打破了一个盘子.')
# Returns: ['我', '打', '破', '了', '一', '个', '盘', '子']
  • Whether the package's character sets and patterns are current with Unicode standards for modern Chinese text
Same gist for agents: .md · .json

What it is and what it does

Zhon is a reference library of constants for Chinese language processing. It exports pre-defined character sets (CJK characters, radicals, punctuation), Pinyin validation patterns (vowels, consonants, syllable/word/sentence regex), Zhuyin marks, and CC-CEDICT character references. The package has no runtime dependencies and is designed to be imported and used with Python's re module or other text processing tools.

Typical use is to extract or validate Chinese text components: finding all hanzi characters in a mixed-language string, matching Pinyin syllables or words with regex, or validating sentence structure. It's a data reference, not an algorithm library—you provide the processing logic (usually regex) and zhon provides the character classes and patterns.

Use it for

  • Extract all CJK characters from mixed-language text using zhon.hanzi.characters with regex
  • Validate and extract Pinyin syllables or words from romanized Chinese using zhon.pinyin patterns
  • Build Chinese text segmentation or tokenization tools that rely on character class definitions
  • Identify Chinese punctuation marks in text for parsing or cleaning tasks
  • Validate Zhuyin (bopomofo) syllables in phonetic Chinese text processing

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Zhon is a stable, dependency-free reference library for Chinese text processing with no known vulnerabilities. Install it if you're building tools that need to identify, validate, or extract Chinese characters, Pinyin, or Zhuyin. The dormant maintenance status is not a concern for a constants-only package—the character sets it exports are unlikely to change frequently.

Install

zhon on PyPI

Before you install

Low install friction with no runtime dependencies. Dormant maintenance status (last release 2024-11-20, last commit 2024-12-11) but stable and widely used; no active development but no signs of abandonment.

License in practice

MIT license is permissive; you can use, modify, and distribute this package freely with minimal restrictions.

Quickstart

import re
import zhon.hanzi

# Find CJK characters in a string
re.findall('[{}]'.format(zhon.hanzi.characters), 'I broke a plate: 我打破了一个盘子.')
# Returns: ['我', '打', '破', '了', '一', '个', '盘', '子']

Verify before relying

  • Whether the package's character sets and patterns are current with Unicode standards for modern Chinese text

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.9
Install frictionLow. Pure-Python wheel
Runtime dependenciesNone
MaintenanceDormant 632 days since the last release
Last repo commit
First released
Downloads571,763 / month, #5,949 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.9Topic :: Software Development :: Libraries :: Python ModulesTopic :: Text Processing :: Linguistic

Evidence: zhon-2.1.1-py3-none-any.whl

Tags

Capabilities
chinese text processing constantscjk character patternspinyin regex validationhanzi character setschinese punctuation markszhuyin syllable patternsmandarin text utilities
Topics
chinese-nlpcharacter-dataregex-patterns
PyPI keywords
cc-cedictcedictcharacterschinesecjkhanhanzimandarinpinyinpunctuationradicalssegmentationsimplifiedtokenizationtraditionalunicodezhuyin

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “chinese text processing constants”

  • zhonZhon provides character constants and regular expression patterns for…
  • hanzidentifierIdentifies whether a string contains Simplified Chinese, Traditional…
  • jiebaJieba segments Chinese text into words using multiple algorithms…

Give your agent the search over MCP, or paste the wish link into any chat.

More Python Modules packages

idna Worth it
PyPI · Python Modules · released Jun 2026

Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.

Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.

BSD-3-Clausepure Python · 3.9+
1.8Bdownloads / mo
setuptools Worth it
PyPI · Python Modules · released Aug 2026

Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.

MITpure Python · 3.10+
1.6Bdownloads / mo
PyYAML Worth it
PyPI · Python Modules · released Sep 2025

PyYAML parses and emits YAML 1.1 data format, enabling serialization and deserialization of configuration files and Python objects to and from human-readable YAML text.

MITcompiled wheel · 3.8+
1.2Bdownloads / mo
pydantic Worth it
PyPI · Python Modules · released May 2026

Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.

MITpure Python · 3.9+
1.1Bdownloads / mo
annotated-types Worth it
PyPI · Python Modules · released Jul 2026

Provides reusable metadata objects for use with PEP-593 `typing.Annotated` to express common constraints like bounds, collection sizes, and predicates on types.

Install it if you use or build libraries that need to express type constraints in a standardized, inspectable way—or if you want to annotate your own types with…

MITpure Python · 3.10+
871.3Mdownloads / mo
typing-inspection Worth it
PyPI · Python Modules · released Aug 2026

Provides runtime tools to inspect and introspect Python type annotations, enabling programmatic examination of type hints at execution time.

MITpure Python · 3.10+
783.0Mdownloads / mo

See also hanzidentifier · nameparser · pypinyin · xpinyin · pypinyin-dict · pinyin · jamo · jieba · OpenCC · mplfonts