fugashi
Cython MeCab wrapper for fast, pythonic Japanese tokenization.
Decision gist · record as of 2026-08-14
Yes, if you need to process Japanese text. This is the standard Python wrapper for MeCab, has no known vulnerabilities, and prebuilt wheels make installation straightforward on common platforms. The aging maintenance status is not a blocker—the package is stable and widely used—but verify that the MeCab and dictionary versions meet your production requirements. Avoid on musl-based systems (Alpine Linux) or 32-bit Windows without manual MeCab compilation.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires MeCab system library; on platforms without prebuilt wheels (musl-based distros, PowerPC, Windows 32-bit), you must install MeCab from source first.
- A dictionary must also be installed.
- Medium install friction due to compiled C extensions; prebuilt wheels are provided for Linux, macOS (Intel and ARM), and Windows x64, but requires MeCab system library on unsupported platforms (musl-based distros, PowerPC, Windows 32-bit).
License · maintenance · safety
MIT AND BSD-3-Clause (permissive) — Dual-licensed under MIT and BSD-3-Clause (permissive). The package itself is MIT; included MeCab binaries in wheels are BSD-licensed. Both are permissive and suitable for commercial use with attribution.
last release 2025-10-24 (294 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,003,315 downloads/mo, #4,532 on PyPI
Alternatives
Verify before relying
pip install 'fugashi[unidic-lite]'
from fugashi import Tagger
tagger = Tagger('-Owakati')
for word in tagger('麩菓子は、麩を主材料とした日本の菓子。'):
print(word, word.feature.lemma, word.pos, sep='\t')- Performance characteristics (speed, memory usage) compared to alternative Japanese tokenizers
- Whether the aging maintenance status reflects active development or dormancy
- Compatibility with latest versions of MeCab and UniDic beyond what the fact sheet indicates
What it is and what it does
This package wraps MeCab, a mature Japanese tokenizer and morphological analyzer, making it accessible from Python. It parses Japanese text (which has no spaces) into individual words and returns grammatical features like part-of-speech, lemma, and other linguistic attributes as named tuples. The package ships with prebuilt wheels for common platforms, eliminating the need to compile MeCab yourself on Linux, macOS, and Windows x64.
You typically install a dictionary alongside the package—unidic-lite for quick testing or unidic for production work—then create a Tagger instance and call it on Japanese text. The package supports both simple tokenization (splitting text into words) and detailed morphological analysis (extracting grammatical information). It also allows custom dictionaries and feature wrappers for non-Unidic use cases.
Use it for
- Tokenizing Japanese text in NLP pipelines where word boundaries and morphological features are needed
- Extracting lemmas and parts-of-speech from Japanese documents for linguistic analysis or search indexing
- Building Japanese language processing applications that require accurate tokenization without spaces
- Preparing Japanese text for downstream tasks like machine translation, sentiment analysis, or named-entity recognition
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need to process Japanese text.
This is the standard Python wrapper for MeCab, has no known vulnerabilities, and prebuilt wheels make installation straightforward on common platforms. The aging maintenance status is not a blocker—the package is stable and widely used—but verify that the MeCab and dictionary versions meet your production requirements. Avoid on musl-based systems (Alpine Linux) or 32-bit Windows without manual MeCab compilation.
Install
fugashi on PyPI
Before you install
Medium install friction due to compiled C extensions; prebuilt wheels are provided for Linux, macOS (Intel and ARM), and Windows x64, but requires MeCab system library on unsupported platforms (musl-based distros, PowerPC, Windows 32-bit). Package is aging but has no known vulnerabilities.
Requires MeCab system library; on platforms without prebuilt wheels (musl-based distros, PowerPC, Windows 32-bit), you must install MeCab from source first. A dictionary must also be installed.
License in practice
Dual-licensed under MIT and BSD-3-Clause (permissive). The package itself is MIT; included MeCab binaries in wheels are BSD-licensed. Both are permissive and suitable for commercial use with attribution.
Quickstart
pip install 'fugashi[unidic-lite]'
from fugashi import Tagger
tagger = Tagger('-Owakati')
for word in tagger('麩菓子は、麩を主材料とした日本の菓子。'):
print(word, word.feature.lemma, word.pos, sep='\t')
Verify before relying
- Performance characteristics (speed, memory usage) compared to alternative Japanese tokenizers
- Whether the aging maintenance status reflects active development or dormancy
- Compatibility with latest versions of MeCab and UniDic beyond what the fact sheet indicates
Package facts
| License | MIT AND BSD-3-Clause permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Aging 294 days since the last release |
| First released | |
| Downloads | 1,003,315 / month, #4,532 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Environment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Science/ResearchNatural Language :: JapaneseOperating System :: MacOS :: MacOS XOperating System :: POSIX :: LinuxProgramming Language :: CythonProgramming Language :: Python :: 3Topic :: Text Processing :: Linguistic |
Evidence: fugashi-1.5.2-cp310-cp310-macosx_10_9_universal2.whl; fugashi-1.5.2-cp310-cp310-macosx_10_9_x86_64.whl; fugashi-1.5.2-cp310-cp310-macosx_11_0_arm64.whl; fugashi-1.5.2-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.whl; fugashi-1.5.2-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.whl; fugashi-1.5.2-cp310-cp310-win_amd64.whl; fugashi-1.5.2-cp311-cp311-macosx_10_9_universal2.whl; fugashi-1.5.2-cp311-cp311-macosx_10_9_x86_64.whl; fugashi-1.5.2-cp311-cp311-macosx_11_0_arm64.whl; fugashi-1.5.2-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.whl; fugashi-1.5.2-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.whl; fugashi-1.5.2-cp311-cp311-win_amd64.whl; fugashi-1.5.2-cp312-cp312-macosx_10_13_universal2.whl; fugashi-1.5.2-cp312-cp312-macosx_10_13_x86_64.whl; fugashi-1.5.2-cp312-cp312-macosx_11_0_arm64.whl; fugashi-1.5.2-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.whl; fugashi-1.5.2-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.whl; fugashi-1.5.2-cp312-cp312-win_amd64.whl; fugashi-1.5.2-cp313-cp313-macosx_10_13_universal2.whl; fugashi-1.5.2-cp313-cp313-macosx_10_13_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “japanese tokenization”
- fugashiA Cython wrapper for MeCab that tokenizes and performs morphological…
- konohaKonoha provides a unified Python interface to multiple Japanese…
- mecab-python3Python wrapper for MeCab, a morphological analyzer that tokenizes and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also mecab-python3 · ipadic · unidic-lite · mecab · konoha · mecab-ko-dic · mecab-ko · unidic · nagisa · tinysegmenter