SudachiDict-small
Sudachi Dictionary for SudachiPy - Small Edition
Decision gist · record as of 2026-08-14
Yes, if you are using SudachiPy and need Japanese morphological analysis. The package is actively maintained, has low install friction, carries a permissive license, and has no known vulnerabilities. Choose the small edition when dictionary size matters; otherwise, evaluate whether core or full editions better suit your linguistic coverage needs.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires SudachiPy to be installed separately; the dictionary is downloaded during installation, so an internet connection is needed at setup time.
- Installation is low-friction; the package is a pure Python wheel that downloads its dictionary data on setup.
- Maintenance is active with a recent release and ongoing repository activity.
License · maintenance · safety
Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing use in commercial and private projects with minimal restrictions.
last release 2026-07-24 (21 days) · last repo commit 2026-07-24 · 305 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 99,927 downloads/mo, #13,008 on PyPI
Alternatives
Verify before relying
pip install sudachidict_small
from sudachipy.tokenizer import Tokenizer
from sudachidict_small import DICT_SMALL_PATH
tokenizer = Tokenizer(dictionary=DICT_SMALL_PATH)
tokens = tokenizer.tokenize("テキスト")- Whether the small edition is suitable for production use or if core/full editions are recommended for specific applications.
- Memory footprint and performance characteristics compared to core and full dictionary editions.
- Compatibility with SudachiPy versions prior to v0.5.2 and the migration path for users on older versions.
What it is and what it does
SudachiDict-small is a packaged dictionary resource for SudachiPy, a Japanese morphological analyzer. It bundles the small edition of the Sudachi dictionary, which is automatically downloaded and installed as a Python package. The package acts as a data dependency—it does not provide parsing or tokenization logic itself, but rather supplies the linguistic data that SudachiPy uses to analyze and segment Japanese text.
The small edition is the most compact of three available dictionary sizes (small, core, full), making it suitable for environments where disk space or download time is a constraint. Once installed, the dictionary is referenced by SudachiPy via a path or configuration, allowing developers to tokenize and parse Japanese text without manually managing dictionary files.
Use it for
- Tokenizing Japanese text in a lightweight application where minimal dictionary size is preferred over comprehensive coverage.
- Setting up a SudachiPy-based NLP pipeline in resource-constrained environments such as embedded systems or serverless functions.
- Providing Japanese morphological analysis in a Python project without the overhead of larger dictionary editions.
- Packaging a complete Japanese text processing solution as a Python dependency for distribution via pip.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are using SudachiPy and need Japanese morphological analysis.
The package is actively maintained, has low install friction, carries a permissive license, and has no known vulnerabilities. Choose the small edition when dictionary size matters; otherwise, evaluate whether core or full editions better suit your linguistic coverage needs.
Install
sudachidict-small on PyPI
Before you install
Installation is low-friction; the package is a pure Python wheel that downloads its dictionary data on setup. Maintenance is active with a recent release and ongoing repository activity.
Requires SudachiPy to be installed separately; the dictionary is downloaded during installation, so an internet connection is needed at setup time.
License in practice
Licensed under Apache-2.0 (permissive), allowing use in commercial and private projects with minimal restrictions.
Quickstart
pip install sudachidict_small
from sudachipy.tokenizer import Tokenizer
from sudachidict_small import DICT_SMALL_PATH
tokenizer = Tokenizer(dictionary=DICT_SMALL_PATH)
tokens = tokenizer.tokenize("テキスト")
Verify before relying
- Whether the small edition is suitable for production use or if core/full editions are recommended for specific applications.
- Memory footprint and performance characteristics compared to core and full dictionary editions.
- Compatibility with SudachiPy versions prior to v0.5.2 and the migration path for users on older versions.
Package facts
| License | Apache-2.0 permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packageSudachiPy |
| Maintenance | Actively maintained 21 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 99,927 / month, #13,008 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: sudachidict_small-20260723-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “sudachi dictionary small”
- SudachiDict-smallProvides the small-edition Sudachi dictionary resource for SudachiPy,…
- SudachiDict-fullProvides the full-edition Sudachi dictionary resource for Japanese…
- SudachiDict-coreProvides the core edition of the Sudachi morphological analyzer…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also SudachiDict-core · SudachiDict-full · SudachiPy · ginza · ja-ginza · mecab-python3 · ipadic · tinysegmenter · unidic-lite · mecab