$npx skillfedfor your agent

SudachiPy

Python version of Sudachi, the Japanese Morphological Analyzer

With conditionsPyPI LinguisticReleased Apr 20262.2M downloads / moApache-2.0Platform wheel

Decision gist · record as of 2026-08-14

platform wheels — sudachipy-0.6.11-cp310-cp310-macosx_10_12_universal2.whl · sudachipy-0.6.11-cp310-cp310-macosx_10_12_x86_64.whl · sudachipy-0.6.11-cp310-cp310-macosx_11_0_arm64.whl
v0.6.11 · released 2026-04-13

Yes, if you need Japanese morphological analysis. SudachiPy is actively maintained, has no security vulnerabilities, uses a permissive license, and provides both CLI and Python API interfaces. The main gotcha is the separate dictionary dependency and medium install friction from compiled wheels, but prebuilt binaries cover common platforms. Suitable for production use in Japanese NLP workflows.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires a separate dictionary package to be installed; dictionary files download on first use (approximately 70MB for the core edition).
  • Medium install friction due to compiled binary wheels.
  • Prebuilt wheels are provided for macOS (10.14+), Windows, and Linux x86_64/aarch64, but ARM-based macOS requires Rust toolchain and Cargo.

License · maintenance · safety

Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions.

last release 2026-04-13 (123 days) · last repo commit 2026-06-29 · 467 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,232,850 downloads/mo, #3,197 on PyPI

Verify before relying

pip install sudachipy

from sudachipy import Dictionary, SplitMode

tokenizer = Dictionary().create()
morphemes = tokenizer.tokenize("国会議事堂前駅", SplitMode.A)
print([m.surface() for m in morphemes])
  • Whether Python version constraints exist beyond what the fact sheet indicates
  • Performance characteristics for large-scale tokenization workloads
  • Whether user dictionaries can be easily integrated for domain-specific terms
  • Specific dictionary package names and their installation requirements
Same gist for agents: .md · .json

What it is and what it does

SudachiPy is a Python binding to Sudachi.rs, a Japanese morphological analyzer that breaks Japanese text into individual morphemes (words/word components) and annotates each with grammatical information. It provides three tokenization granularity modes (A, B, C) to control how aggressively text is split, and returns morpheme objects with methods to access surface forms, dictionary forms, reading forms, part-of-speech tags, and normalized variants. The package includes both a command-line interface for shell pipelines and a Python API for programmatic use.

The package is distributed as precompiled wheels for modern Python versions on macOS, Windows, and Linux (x86_64 and aarch64). It has no pure-Python runtime dependencies—the heavy lifting is done by the compiled Sudachi.rs backend. A dictionary must be installed separately as a companion package. The project is actively maintained and widely used in Japanese NLP workflows.

Use it for

  • Tokenize Japanese text for downstream NLP tasks like named entity recognition or sentiment analysis
  • Build a command-line pipeline to process Japanese documents with multiple granularity levels
  • Extract normalized forms and reading annotations for Japanese text in search or indexing systems
  • Integrate Japanese morphological analysis into a larger text processing or machine learning pipeline

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need Japanese morphological analysis.

SudachiPy is actively maintained, has no security vulnerabilities, uses a permissive license, and provides both CLI and Python API interfaces. The main gotcha is the separate dictionary dependency and medium install friction from compiled wheels, but prebuilt binaries cover common platforms. Suitable for production use in Japanese NLP workflows.

Install

sudachipy on PyPI

Before you install

Medium install friction due to compiled binary wheels. Prebuilt wheels are provided for macOS (10.14+), Windows, and Linux x86_64/aarch64, but ARM-based macOS requires Rust toolchain and Cargo. Package is actively maintained with recent releases.

Requires a separate dictionary package to be installed; dictionary files download on first use (approximately 70MB for the core edition).

License in practice

Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions.

Quickstart

pip install sudachipy

from sudachipy import Dictionary, SplitMode

tokenizer = Dictionary().create()
morphemes = tokenizer.tokenize("国会議事堂前駅", SplitMode.A)
print([m.surface() for m in morphemes])

Verify before relying

  • Whether Python version constraints exist beyond what the fact sheet indicates
  • Performance characteristics for large-scale tokenization workloads
  • Whether user dictionaries can be easily integrated for domain-specific terms
  • Specific dictionary package names and their installation requirements

Package facts

LicenseApache-2.0 permissive
Python supportNot specified
Install frictionMedium. Platform-specific wheel
Runtime dependenciesNone
MaintenanceActively maintained 123 days since the last release
Last repo commit
First released
Downloads2,232,850 / month, #3,197 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: sudachipy-0.6.11-cp310-cp310-macosx_10_12_universal2.whl; sudachipy-0.6.11-cp310-cp310-macosx_10_12_x86_64.whl; sudachipy-0.6.11-cp310-cp310-macosx_11_0_arm64.whl; sudachipy-0.6.11-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.whl; sudachipy-0.6.11-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.whl; sudachipy-0.6.11-cp310-cp310-win_amd64.whl; sudachipy-0.6.11-cp311-cp311-macosx_10_12_universal2.whl; sudachipy-0.6.11-cp311-cp311-macosx_10_12_x86_64.whl; sudachipy-0.6.11-cp311-cp311-macosx_11_0_arm64.whl; sudachipy-0.6.11-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.whl; sudachipy-0.6.11-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.whl; sudachipy-0.6.11-cp311-cp311-win_amd64.whl; sudachipy-0.6.11-cp312-cp312-macosx_10_13_universal2.whl; sudachipy-0.6.11-cp312-cp312-macosx_10_13_x86_64.whl; sudachipy-0.6.11-cp312-cp312-macosx_11_0_arm64.whl; sudachipy-0.6.11-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.whl; sudachipy-0.6.11-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.whl; sudachipy-0.6.11-cp312-cp312-win_amd64.whl; sudachipy-0.6.11-cp313-cp313-macosx_10_13_universal2.whl; sudachipy-0.6.11-cp313-cp313-macosx_10_13_x86_64.whl

Tags

Capabilities
japanese text tokenizationmorphological analysis japanesejapanese nlp tokenizersudachi morphological analyzerjapanese word segmentationjapanese language processingmorpheme extraction japanese
Topics
japanese-nlpmorphological-analysistokenization

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “morpheme extraction japanese”

  • SudachiPySudachiPy is a Python binding for Sudachi.rs, a Japanese…
  • python-mecab-koProvides Python bindings for MeCab-ko, a morphological analyzer for…
  • pyopenjtalkWraps OpenJTalk to provide Japanese text-to-speech synthesis,…

Give your agent the search over MCP, or paste the wish link into any chat.

More Linguistic packages

charset-normalizer Worth it
PyPI · Utilities · released Aug 2026

Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.

permissive licensepure Python · 3.7+
1.7Bdownloads / mo
tiktoken Worth it
PyPI · Linguistic · released May 2026

tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.

Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.

permissive licensecompiled wheel · 3.9+
233.0Mdownloads / mo
chardet Worth it
PyPI · Python Modules · released Aug 2026

Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.

Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.

0BSDpure Python · 3.10+
199.0Mdownloads / mo
text-unidecode With conditions
PyPI · Python Modules · released Aug 2019

Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.

However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.

GPL-2.0-or-laterpure Pythonabandoned
89.0Mdownloads / mo
lark Worth it
PyPI · Python Modules · released Oct 2025

Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.

MITpure Python · 3.8+
79.7Mdownloads / mo
tree-sitter Worth it
PyPI · Linguistic · released Jun 2026

Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.

MITcompiled wheel · 3.10+
79.0Mdownloads / mo

See also SudachiDict-core · SudachiDict-full · rhoknp · Janome · SudachiDict-small · mecab · mecab-python3 · nagisa · ginza · ja-ginza