jieba3k
Chinese Words Segementation Utilities
Decision gist · record as of 2026-08-14
No. The package is abandoned (last release 2014-11-15), has unclear licensing, unspecified Python support, and high install friction. Modern Chinese NLP needs are better served by actively maintained alternatives. Install only if you are maintaining legacy code that already depends on it.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Package is abandoned and may not install cleanly on modern Python versions; Python version support is unspecified.
- High install friction due to distribution as a .zip file rather than a standard wheel or source distribution.
- The package is abandoned—last release was 2014-11-15, over 4290 days ago—with no maintenance activity or repository information available.
License · maintenance · safety
UNKNOWN (unclear) — License status is unclear; the package lists no SPDX identifier and raw license information is unknown, making it difficult to assess legal compatibility or usage restrictions.
last release 2014-11-15 (4290 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 465,476 downloads/mo, #6,507 on PyPI
Alternatives
Verify before relying
pip install jieba3k
import jieba3k
result = jieba3k.cut('Chinese text here')- Whether the package works on current Python versions (support unspecified in metadata)
- Actual license terms and compatibility (listed as UNKNOWN)
- Whether the .zip distribution unpacks and installs without manual intervention
- Whether jieba3k is a fork or variant of another jieba package and how they differ
What it is and what it does
jieba3k is a Chinese word segmentation library that tokenizes Chinese text into individual words or meaningful units. It is designed for natural language processing workflows where raw Chinese text must be split into analyzable tokens before further processing.
The package has been abandoned since late 2014 and receives no maintenance. It has no declared runtime dependencies and is distributed as a .zip file, creating installation friction on modern systems. Python version support is unspecified, and the license terms are unknown, raising compatibility and legal concerns for new projects.
Use it for
- Tokenizing Chinese text for machine learning feature extraction or text classification pipelines
- Preprocessing Chinese documents before feeding them into NLP models or search indexing systems
- Building Chinese language text analysis tools where word boundaries must be identified
- Legacy system maintenance where existing code depends on jieba3k's segmentation output
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
No.
The package is abandoned (last release 2014-11-15), has unclear licensing, unspecified Python support, and high install friction. Modern Chinese NLP needs are better served by actively maintained alternatives. Install only if you are maintaining legacy code that already depends on it.
Install
jieba3k on PyPI
Before you install
High install friction due to distribution as a .zip file rather than a standard wheel or source distribution. The package is abandoned—last release was 2014-11-15, over 4290 days ago—with no maintenance activity or repository information available.
Package is abandoned and may not install cleanly on modern Python versions; Python version support is unspecified.
License in practice
License status is unclear; the package lists no SPDX identifier and raw license information is unknown, making it difficult to assess legal compatibility or usage restrictions.
Quickstart
pip install jieba3k
import jieba3k
result = jieba3k.cut('Chinese text here')
Verify before relying
- Whether the package works on current Python versions (support unspecified in metadata)
- Actual license terms and compatibility (listed as UNKNOWN)
- Whether the .zip distribution unpacks and installs without manual intervention
- Whether jieba3k is a fork or variant of another jieba package and how they differ
Package facts
| License | UNKNOWN unclear |
| Python support | Not specified |
| Install friction | High. Source build required |
| Runtime dependencies | None |
| Maintenance | Abandoned 4,290 days since the last release |
| First released | |
| Downloads | 465,476 / month, #6,507 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: jieba3k-0.35.1.zip
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “chinese nlp tokenizer”
- jieba3kPerforms Chinese word segmentation, breaking Chinese text into…
- rjiebaA Python binding to the Rust-based jieba-rs Chinese text segmentation…
- rouge-chineseComputes ROUGE evaluation metrics for Chinese text summarization and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also jieba · rjieba · wordninja · nagisa · soynlp · spacy-pkuseg · segtok · wordsegment · sacremoses · tokenizer