konoha
Add your description here
Decision gist · record as of 2026-08-14
Yes, if you work with Japanese text and want flexibility in tokenizer choice. The low install friction, active maintenance, MIT license, and zero known vulnerabilities make it a safe dependency. Install with a specific tokenizer extra (e.g., `konoha[mecab]`) unless you plan to choose at runtime; the base package alone won't tokenize without an underlying tokenizer installed.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- The underlying tokenizer (MeCab, Janome, etc.) must be installed separately or via extras (e.g., pip install 'konoha[mecab]').
- Low friction: pure Python wheel with only requests as a runtime dependency.
License · maintenance · safety
MIT (permissive) — MIT license permits free use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.
last release 2026-03-01 (166 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 156,907 downloads/mo, #10,772 on PyPI
Alternatives
Verify before relying
pip install konoha
from konoha import WordTokenizer
tokenizer = WordTokenizer('MeCab')
print(tokenizer.tokenize('自然言語処理を勉強しています'))- Whether all advertised tokenizers (MeCab, Janome, Sentencepiece) are equally well-maintained and tested.
- Performance characteristics and tokenization accuracy compared to using tokenizers directly.
- Whether the Docker-based API server is actively maintained and production-ready.
What it is and what it does
Konoha is a wrapper library that abstracts away the differences between multiple Japanese tokenizers, allowing you to write tokenization code once and swap tokenizers by changing a single parameter. It supports word-level tokenization (via MeCab, Janome, Sentencepiece, and others), sentence-level splitting with customizable delimiters and bracket handling, and rule-based tokenizers for simple cases. The library also includes optional remote file support for loading dictionaries and models from Amazon S3.
You use it by instantiating a WordTokenizer or SentenceTokenizer with your chosen backend, then calling tokenize() on your input text. It's designed for preprocessing pipelines in Japanese NLP tasks where you might want to experiment with different tokenizers or deploy with a specific one. The package also exposes a REST API via Docker for tokenization as a service.
Use it for
- Switching between MeCab and Janome during development to compare tokenization quality without rewriting preprocessing code.
- Building a Japanese NLP preprocessing pipeline that can use different tokenizers in different environments (dev, test, production).
- Sentence-level splitting of Japanese text with custom punctuation and bracket rules for downstream analysis.
- Deploying a tokenization microservice via Docker for multiple applications to call over HTTP.
- Loading tokenizer models and dictionaries from S3 in cloud-based NLP workflows.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you work with Japanese text and want flexibility in tokenizer choice.
The low install friction, active maintenance, MIT license, and zero known vulnerabilities make it a safe dependency. Install with a specific tokenizer extra (e.g., `konoha[mecab]`) unless you plan to choose at runtime; the base package alone won't tokenize without an underlying tokenizer installed.
Install
konoha on PyPI
Before you install
Low friction: pure Python wheel with only requests as a runtime dependency. Actively maintained as of March 2026. Requires Python 3.10 or later.
Requires Python 3.10 or later. The underlying tokenizer (MeCab, Janome, etc.) must be installed separately or via extras (e.g., pip install 'konoha[mecab]').
License in practice
MIT license permits free use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.
Quickstart
pip install konoha
from konoha import WordTokenizer
tokenizer = WordTokenizer('MeCab')
print(tokenizer.tokenize('自然言語処理を勉強しています'))
Verify before relying
- Whether all advertised tokenizers (MeCab, Janome, Sentencepiece) are equally well-maintained and tested.
- Performance characteristics and tokenization accuracy compared to using tokenizers directly.
- Whether the Docker-based API server is actively maintained and production-ready.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagerequests |
| Maintenance | Actively maintained 166 days since the last release |
| First released | |
| Downloads | 156,907 / month, #10,772 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: konoha-5.7.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “switch between japanese tokenizers”
- konohaKonoha provides a unified Python interface to multiple Japanese…
- unidicProvides the UniDic 2.3.0 Japanese morphological dictionary for use…
- mplfontsManages Matplotlib fonts and solves CJK (Chinese, Japanese, Korean)…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also Janome · mecab · fugashi · mecab-python3 · ipadic · mecab-ko-dic · curated-tokenizers · unidic · SudachiPy · sentencepiece