mecab-ko-dic
mecab-ko-dic packaged for Python
Decision gist · record as of 2026-08-14
No. The package is abandoned (last update 2021-02-19, repository archived), has high install friction due to compiled dependencies, and the maintainers themselves recommend other tools for Korean NLP unless MeCab interoperability is a hard requirement. License metadata is also unclear. Only install if you have an existing MeCab-based pipeline that specifically needs Korean support and cannot migrate.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires MeCab and mecab-python3 (or fugashi) to be installed and properly configured on the system; MeCab is a compiled C++ library with system-level dependencies.
- High install friction due to compiled dependencies (mecab-ko-dic-1.0.0.tar.gz).
- Package is abandoned as of 2021-02-19 with no updates since initial release; the repository is archived.
License · maintenance · safety
(unclear) — License treatment is unclear in the metadata, though the description states the dictionary data and code are released under Apache License 2.0. Verify the actual license terms before use in proprietary projects.
last release 2021-02-19 (2002 days) · last repo commit 2021-02-19 · 10 stars · archived
0 known vulnerabilities (OSV.dev, 2026-08-14) · 99,170 downloads/mo, #13,038 on PyPI
Alternatives
Verify before relying
pip install mecab-ko-dic
import MeCab
import mecab_ko_dic
tagger = MeCab.Tagger(mecab_ko_dic.MECAB_ARGS)
print(tagger.parse("안녕하세요세계."))- Whether the package works with current versions of mecab-python3 or fugashi given the 2021 freeze.
- Actual license metadata clarity—description claims Apache 2.0 but license_treatment is marked unclear.
- Whether Korean NLP has better maintained alternatives for modern use cases.
What it is and what it does
mecab-ko-dic is a Korean morphological dictionary packaged for Python to work with the MeCab tokenizer. It enables splitting and analyzing Korean text into meaningful tokens (morphemes) rather than simple space-based word boundaries. The package wraps dictionary data created by Yongwoon Lee and Yungho Yu, originally designed for use with MeCab's Japanese support but adapted for Korean.
The package is a thin wrapper that provides the dictionary and configuration arguments needed to initialize MeCab for Korean text processing. It is intended primarily for projects that already use MeCab (for example, in Japanese NLP pipelines) and need consistent Korean tokenization interoperability. The maintainers acknowledge MeCab was not originally designed for Korean and suggest other tools may be better suited for Korean-only NLP tasks.
Use it for
- Tokenizing Korean text in a multilingual NLP pipeline that already uses MeCab for Japanese.
- Extracting consistent morphological tokens from Korean text for frequency analysis or vocabulary lookup.
- Integrating Korean language support into wordfreq or similar frequency-based text analysis tools.
- Processing Korean text in legacy systems built around MeCab that require interoperability.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
No.
The package is abandoned (last update 2021-02-19, repository archived), has high install friction due to compiled dependencies, and the maintainers themselves recommend other tools for Korean NLP unless MeCab interoperability is a hard requirement. License metadata is also unclear. Only install if you have an existing MeCab-based pipeline that specifically needs Korean support and cannot migrate.
Install
mecab-ko-dic on PyPI
Before you install
High install friction due to compiled dependencies (mecab-ko-dic-1.0.0.tar.gz). Package is abandoned as of 2021-02-19 with no updates since initial release; the repository is archived.
Requires MeCab and mecab-python3 (or fugashi) to be installed and properly configured on the system; MeCab is a compiled C++ library with system-level dependencies.
License in practice
License treatment is unclear in the metadata, though the description states the dictionary data and code are released under Apache License 2.0. Verify the actual license terms before use in proprietary projects.
Quickstart
pip install mecab-ko-dic
import MeCab
import mecab_ko_dic
tagger = MeCab.Tagger(mecab_ko_dic.MECAB_ARGS)
print(tagger.parse("안녕하세요세계."))
Verify before relying
- Whether the package works with current versions of mecab-python3 or fugashi given the 2021 freeze.
- Actual license metadata clarity—description claims Apache 2.0 but license_treatment is marked unclear.
- Whether Korean NLP has better maintained alternatives for modern use cases.
Package facts
| License | Not declared unclear |
| Python support | Not specified |
| Install friction | High. Source build required |
| Runtime dependencies | None |
| Maintenance | Abandoned 2,002 days since the last release |
| Last repo commit | repository archived |
| First released | |
| Downloads | 99,170 / month, #13,038 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Natural Language :: KoreanOperating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: mecab-ko-dic-1.0.0.tar.gz
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “korean text tokenization”
- mecab-ko-dicProvides a Korean dictionary for MeCab tokenization, enabling…
- mecab-koPython wrapper for MeCab-ko, a morphological analyzer that tokenizes…
- python-mecab-koProvides Python bindings for MeCab-ko, a morphological analyzer for…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also python-mecab-ko-dic · mecab-ko · ipadic · kiwipiepy · unidic-lite · kiwipiepy-model · soynlp · mecab-python3 · fugashi · konlpy