$npx skillfedfor your agent

spacy-pkuseg

Chinese word segmentation toolkit for spaCy (fork of pkuseg-python)

With conditionsPyPI Scientific/EngineeringReleased Jul 2025702.7K downloads / moMITPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — spacy_pkuseg-1.0.1-cp310-cp310-macosx_10_9_x86_64.whl · spacy_pkuseg-1.0.1-cp310-cp310-macosx_11_0_arm64.whl · spacy_pkuseg-1.0.1-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
v1.0.1 · released 2025-07-14 · Python >=3.9 · 2 runtime deps: numpy, srsly

Yes, if you need Chinese word segmentation in a spaCy pipeline or want domain-specific accuracy. The package is actively maintained, has no known vulnerabilities, supports modern Python versions, and offers pretrained models for multiple domains. Install friction is moderate but manageable. Not necessary if you only need generic Chinese tokenization or are not using spaCy.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.9 or later; precompiled wheels available for Linux, macOS, and Windows x86_64/ARM64.
  • Medium install friction due to compiled wheels, but well-supported across Python 3.9–3.13 and major platforms (Linux, macOS, Windows).
  • Active maintenance with recent commits and stable release history.

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; no copyleft obligations.

last release 2025-07-14 (396 days) · last repo commit 2026-03-27 · 71 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 702,717 downloads/mo, #5,283 on PyPI

Verify before relying

pip install spacy-pkuseg
import spacy_pkuseg
seg = spacy_pkuseg.pkuseg()
text = seg.cut('我爱北京天安门')
print(text)
  • Whether spaCy itself must be installed separately or is pulled in as a transitive dependency.
  • Performance benchmarks comparing this fork to the original pkuseg-python on the same hardware.
  • Whether domain models (medicine, tourism, etc.) auto-download or require manual setup when used via spaCy.
Same gist for agents: .md · .json

What it is and what it does

spacy-pkuseg is a spaCy-integrated fork of the pkuseg Chinese word segmentation toolkit. It provides domain-aware tokenization for Chinese text across five pretrained models: a mixed-domain default, plus specialized models for news, web, medicine, and tourism text. The package wraps the underlying segmentation engine (unmodified from the original pkuseg) and simplifies both installation and model serialization for spaCy workflows.

The toolkit supports optional part-of-speech tagging alongside segmentation, batch processing of files with multiprocessing, and user-defined custom dictionaries. It depends on numpy and srsly for numerical and serialization operations. Installation is straightforward via pip with precompiled wheels for modern Python versions and common platforms, though it carries medium friction due to compiled components.

Use it for

  • Segment Chinese news articles or web text where domain-specific accuracy matters more than generic tokenization.
  • Build spaCy NLP pipelines for Chinese that need both word boundaries and part-of-speech labels in one step.
  • Process medical or tourism domain Chinese documents with models trained on domain-specific corpora.
  • Batch-process large Chinese text files with multiprocessing to split into words and optionally tag parts of speech.
  • Extend spaCy's Chinese support with a higher-accuracy alternative to generic tokenizers when domain context is known.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need Chinese word segmentation in a spaCy pipeline or want domain-specific accuracy.

The package is actively maintained, has no known vulnerabilities, supports modern Python versions, and offers pretrained models for multiple domains. Install friction is moderate but manageable. Not necessary if you only need generic Chinese tokenization or are not using spaCy.

Install

spacy-pkuseg on PyPI

Before you install

Medium install friction due to compiled wheels, but well-supported across Python 3.9–3.13 and major platforms (Linux, macOS, Windows). Active maintenance with recent commits and stable release history.

Requires Python 3.9 or later; precompiled wheels available for Linux, macOS, and Windows x86_64/ARM64.

License in practice

MIT license permits commercial and private use with minimal restrictions; no copyleft obligations.

Quickstart

pip install spacy-pkuseg
import spacy_pkuseg
seg = spacy_pkuseg.pkuseg()
text = seg.cut('我爱北京天安门')
print(text)

Verify before relying

  • Whether spaCy itself must be installed separately or is pulled in as a transitive dependency.
  • Performance benchmarks comparing this fork to the original pkuseg-python on the same hardware.
  • Whether domain models (medicine, tourism, etc.) auto-download or require manual setup when used via spaCy.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.9
Install frictionMedium. Platform-specific wheel
Runtime dependencies
2 packages
numpysrsly
MaintenanceActively maintained 396 days since the last release
Last repo commit
First released
Downloads702,717 / month, #5,283 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseOperating System :: MacOS :: MacOS XOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: CythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.9Topic :: Scientific/Engineering

Evidence: spacy_pkuseg-1.0.1-cp310-cp310-macosx_10_9_x86_64.whl; spacy_pkuseg-1.0.1-cp310-cp310-macosx_11_0_arm64.whl; spacy_pkuseg-1.0.1-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; spacy_pkuseg-1.0.1-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; spacy_pkuseg-1.0.1-cp310-cp310-musllinux_1_2_aarch64.whl; spacy_pkuseg-1.0.1-cp310-cp310-musllinux_1_2_x86_64.whl; spacy_pkuseg-1.0.1-cp310-cp310-win_amd64.whl; spacy_pkuseg-1.0.1-cp311-cp311-macosx_10_9_x86_64.whl; spacy_pkuseg-1.0.1-cp311-cp311-macosx_11_0_arm64.whl; spacy_pkuseg-1.0.1-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; spacy_pkuseg-1.0.1-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; spacy_pkuseg-1.0.1-cp311-cp311-musllinux_1_2_aarch64.whl; spacy_pkuseg-1.0.1-cp311-cp311-musllinux_1_2_x86_64.whl; spacy_pkuseg-1.0.1-cp311-cp311-win_amd64.whl; spacy_pkuseg-1.0.1-cp312-cp312-macosx_10_13_x86_64.whl; spacy_pkuseg-1.0.1-cp312-cp312-macosx_11_0_arm64.whl; spacy_pkuseg-1.0.1-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; spacy_pkuseg-1.0.1-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; spacy_pkuseg-1.0.1-cp312-cp312-musllinux_1_2_aarch64.whl; spacy_pkuseg-1.0.1-cp312-cp312-musllinux_1_2_x86_64.whl

Tags

Capabilities
chinese word segmentationchinese nlp tokenizationspacy chinese segmentationpkuseg domain segmentationchinese text processingmulti-domain word segmentationchinese pos tagging
Topics
chinese-nlpspacy-integrationdomain-segmentation

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “chinese word segmentation”

  • spacy-pkusegChinese word segmentation for spaCy with domain-specific models…
  • jieba3kPerforms Chinese word segmentation, breaking Chinese text into…
  • jiebaJieba segments Chinese text into words using multiple algorithms…

Give your agent the search over MCP, or paste the wish link into any chat.

More Scientific/Engineering packages

numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
pandas Worth it
PyPI · Scientific/Engineering · released Jul 2026

pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.

BSD-3-Clausecompiled wheel · 3.11+
769.1Mdownloads / mo
scipy Worth it
PyPI · Libraries · released Jun 2026

scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.

BSD-3-Clausecompiled wheel · 3.12+
449.0Mdownloads / mo
scikit-learn Worth it
PyPI · Software Development · released Jun 2026

scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.

Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.

BSD-3-Clausecompiled wheel · 3.11+
235.5Mdownloads / mo
dill Worth it
PyPI · Software Development · released Jan 2026

dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.

BSD-3-Clausepure Python · 3.9+
208.1Mdownloads / mo
multiprocess Worth it
PyPI · Software Development · released Jan 2026

Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.

Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.

BSD-3-Clausepure Python · 3.9+
202.7Mdownloads / mo

See also jieba · jieba3k · spacy · rjieba · nagisa · segtok · rouge-chinese · bpemb · PyRuSH · pyvi