SudachiPy
Python version of Sudachi, the Japanese Morphological Analyzer
Decision gist · record as of 2026-08-14
Yes, if you need Japanese morphological analysis. SudachiPy is actively maintained, has no security vulnerabilities, uses a permissive license, and provides both CLI and Python API interfaces. The main gotcha is the separate dictionary dependency and medium install friction from compiled wheels, but prebuilt binaries cover common platforms. Suitable for production use in Japanese NLP workflows.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires a separate dictionary package to be installed; dictionary files download on first use (approximately 70MB for the core edition).
- Medium install friction due to compiled binary wheels.
- Prebuilt wheels are provided for macOS (10.14+), Windows, and Linux x86_64/aarch64, but ARM-based macOS requires Rust toolchain and Cargo.
License · maintenance · safety
Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions.
last release 2026-04-13 (123 days) · last repo commit 2026-06-29 · 467 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,232,850 downloads/mo, #3,197 on PyPI
Alternatives
Verify before relying
pip install sudachipy
from sudachipy import Dictionary, SplitMode
tokenizer = Dictionary().create()
morphemes = tokenizer.tokenize("国会議事堂前駅", SplitMode.A)
print([m.surface() for m in morphemes])- Whether Python version constraints exist beyond what the fact sheet indicates
- Performance characteristics for large-scale tokenization workloads
- Whether user dictionaries can be easily integrated for domain-specific terms
- Specific dictionary package names and their installation requirements
What it is and what it does
SudachiPy is a Python binding to Sudachi.rs, a Japanese morphological analyzer that breaks Japanese text into individual morphemes (words/word components) and annotates each with grammatical information. It provides three tokenization granularity modes (A, B, C) to control how aggressively text is split, and returns morpheme objects with methods to access surface forms, dictionary forms, reading forms, part-of-speech tags, and normalized variants. The package includes both a command-line interface for shell pipelines and a Python API for programmatic use.
The package is distributed as precompiled wheels for modern Python versions on macOS, Windows, and Linux (x86_64 and aarch64). It has no pure-Python runtime dependencies—the heavy lifting is done by the compiled Sudachi.rs backend. A dictionary must be installed separately as a companion package. The project is actively maintained and widely used in Japanese NLP workflows.
Use it for
- Tokenize Japanese text for downstream NLP tasks like named entity recognition or sentiment analysis
- Build a command-line pipeline to process Japanese documents with multiple granularity levels
- Extract normalized forms and reading annotations for Japanese text in search or indexing systems
- Integrate Japanese morphological analysis into a larger text processing or machine learning pipeline
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need Japanese morphological analysis.
SudachiPy is actively maintained, has no security vulnerabilities, uses a permissive license, and provides both CLI and Python API interfaces. The main gotcha is the separate dictionary dependency and medium install friction from compiled wheels, but prebuilt binaries cover common platforms. Suitable for production use in Japanese NLP workflows.
Install
sudachipy on PyPI
Before you install
Medium install friction due to compiled binary wheels. Prebuilt wheels are provided for macOS (10.14+), Windows, and Linux x86_64/aarch64, but ARM-based macOS requires Rust toolchain and Cargo. Package is actively maintained with recent releases.
Requires a separate dictionary package to be installed; dictionary files download on first use (approximately 70MB for the core edition).
License in practice
Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions.
Quickstart
pip install sudachipy
from sudachipy import Dictionary, SplitMode
tokenizer = Dictionary().create()
morphemes = tokenizer.tokenize("国会議事堂前駅", SplitMode.A)
print([m.surface() for m in morphemes])
Verify before relying
- Whether Python version constraints exist beyond what the fact sheet indicates
- Performance characteristics for large-scale tokenization workloads
- Whether user dictionaries can be easily integrated for domain-specific terms
- Specific dictionary package names and their installation requirements
Package facts
| License | Apache-2.0 permissive |
| Python support | Not specified |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 123 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 2,232,850 / month, #3,197 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: sudachipy-0.6.11-cp310-cp310-macosx_10_12_universal2.whl; sudachipy-0.6.11-cp310-cp310-macosx_10_12_x86_64.whl; sudachipy-0.6.11-cp310-cp310-macosx_11_0_arm64.whl; sudachipy-0.6.11-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.whl; sudachipy-0.6.11-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.whl; sudachipy-0.6.11-cp310-cp310-win_amd64.whl; sudachipy-0.6.11-cp311-cp311-macosx_10_12_universal2.whl; sudachipy-0.6.11-cp311-cp311-macosx_10_12_x86_64.whl; sudachipy-0.6.11-cp311-cp311-macosx_11_0_arm64.whl; sudachipy-0.6.11-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.whl; sudachipy-0.6.11-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.whl; sudachipy-0.6.11-cp311-cp311-win_amd64.whl; sudachipy-0.6.11-cp312-cp312-macosx_10_13_universal2.whl; sudachipy-0.6.11-cp312-cp312-macosx_10_13_x86_64.whl; sudachipy-0.6.11-cp312-cp312-macosx_11_0_arm64.whl; sudachipy-0.6.11-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.whl; sudachipy-0.6.11-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.whl; sudachipy-0.6.11-cp312-cp312-win_amd64.whl; sudachipy-0.6.11-cp313-cp313-macosx_10_13_universal2.whl; sudachipy-0.6.11-cp313-cp313-macosx_10_13_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “morpheme extraction japanese”
- SudachiPySudachiPy is a Python binding for Sudachi.rs, a Japanese…
- python-mecab-koProvides Python bindings for MeCab-ko, a morphological analyzer for…
- pyopenjtalkWraps OpenJTalk to provide Japanese text-to-speech synthesis,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also SudachiDict-core · SudachiDict-full · rhoknp · Janome · SudachiDict-small · mecab · mecab-python3 · nagisa · ginza · ja-ginza