$npx skillfedfor your agent

unidic

UniDic packaged for Python

With conditionsPyPI LinguisticReleased Oct 2021416.4K downloads / moMITSource build

Decision gist · record as of 2026-08-14

sdist only — unidic-1.1.0.tar.gz · builds from source
v1.1.0 · released 2021-10-10 · Python >=3.5

Yes, if you need full Japanese morphological analysis and are willing to accept the 1GB disk footprint and the two-step installation process (pip install + manual download). The dictionary is comprehensive and well-maintained by NINJAL. No, if you want a lightweight solution—the package explicitly recommends unidic-lite as an alternative. The aging maintenance status (last release 2021-10-10) is a minor concern but not a blocker, since the dictionary data is stable and the code is minimal.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires a MeCab installation and a MeCab-based tokenizer library as a separate runtime dependency to actually tokenize text; unidic alone provides only the dictionary data and the DICDIR path.
  • Installation has high friction: the package itself is small, but the dictionary data must be downloaded separately via `python -m unidic download` after pip install, and the full dictionary occupies approximately 1GB on disk.
  • The project is aging (last release 2021-10-10, 1769 days ago), though the repository remains active with recent commits.

License · maintenance · safety

MIT (permissive) — MIT license on the code is permissive and poses no restrictions. The UniDic dictionary data itself is available under GPL, LGPL, or BSD (user's choice per the UniDic Consortium); the package distributes it under BSD terms. No licensing barrier to use.

last release 2021-10-10 (1769 days) · last repo commit 2025-02-26 · 113 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 416,399 downloads/mo, #6,820 on PyPI

Verify before relying

pip install unidic
python -m unidic download

import unidic
dicdir = unidic.DICDIR
  • Whether the 1GB disk requirement is accurate for the current 2.3.0 dictionary version or has changed.
  • Current state of the AWS Open Data mirror hosting the dictionary—whether downloads remain reliable.
  • Whether the package works with Python versions newer than 3.5 without issues, given the aging maintenance status.
  • Actual integration behavior with MeCab-based tokenizers when installed alongside them.
Same gist for agents: .md · .json

What it is and what it does

unidic-py packages the UniDic 2.3.0 Japanese morphological dictionary for pip installation. UniDic is a comprehensive lexical resource maintained by NINJAL (the National Institute for Japanese Language and Linguistics) that includes detailed linguistic annotations for Japanese words: part-of-speech tags, conjugation types, lemmas, pronunciations, etymological categories, and accent information. The package itself is a thin wrapper; after installation, you must run `python -m unidic download` to fetch the dictionary data from AWS, which takes up approximately 1GB of disk space.

Once installed, unidic exposes the DICDIR constant to locate the dictionary. The package includes minor modifications from the official UniDic release (additions for 令和, removal of single-character numeric/alphabetic entries, and changes to unknown-punctuation handling) to improve usability in Python workflows. It is designed to work with MeCab-based tokenizers that can accept a dictionary path argument.

Use it for

  • Japanese NLP pipelines that need detailed morphological analysis beyond basic part-of-speech tags, such as lemmatization or conjugation-type identification.
  • Building Japanese text processing systems that require the full UniDic annotation set for linguistic research or production NLP.
  • Japanese language learning or corpus analysis tools that benefit from rich lemma and pronunciation data.
  • Accent analysis and standard-language pronunciation research using the aType and kana fields in UniDic.
  • Counting-expression parsing in Japanese text, using the specialized iConType and fConType fields for numeric and counter contexts.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need full Japanese morphological analysis and are willing to accept the 1GB disk footprint and the two-step installation process (pip install + manual download).

The dictionary is comprehensive and well-maintained by NINJAL. No, if you want a lightweight solution—the package explicitly recommends unidic-lite as an alternative. The aging maintenance status (last release 2021-10-10) is a minor concern but not a blocker, since the dictionary data is stable and the code is minimal.

Install

unidic on PyPI

Before you install

Installation has high friction: the package itself is small, but the dictionary data must be downloaded separately via `python -m unidic download` after pip install, and the full dictionary occupies approximately 1GB on disk. The project is aging (last release 2021-10-10, 1769 days ago), though the repository remains active with recent commits.

Requires a MeCab installation and a MeCab-based tokenizer library as a separate runtime dependency to actually tokenize text; unidic alone provides only the dictionary data and the DICDIR path.

License in practice

MIT license on the code is permissive and poses no restrictions. The UniDic dictionary data itself is available under GPL, LGPL, or BSD (user's choice per the UniDic Consortium); the package distributes it under BSD terms. No licensing barrier to use.

Quickstart

pip install unidic
python -m unidic download

import unidic
dicdir = unidic.DICDIR

Verify before relying

  • Whether the 1GB disk requirement is accurate for the current 2.3.0 dictionary version or has changed.
  • Current state of the AWS Open Data mirror hosting the dictionary—whether downloads remain reliable.
  • Whether the package works with Python versions newer than 3.5 without issues, given the aging maintenance status.
  • Actual integration behavior with MeCab-based tokenizers when installed alongside them.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.5
Install frictionHigh. Source build required
Runtime dependenciesNone
MaintenanceAging 1,769 days since the last release
Last repo commit
First released
Downloads416,399 / month, #6,820 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
License :: OSI Approved :: MIT LicenseNatural Language :: Japanese

Evidence: unidic-1.1.0.tar.gz

Tags

Capabilities
japanese morphological dictionaryunidic mecab tokenizerjapanese pos taggingjapanese nlp dictionaryjapanese lemmatizationmecab dictionary japanesejapanese linguistic analysis
Topics
japanese-nlpmorphological-analysismecab-dictionary

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “unidic mecab tokenizer”

  • unidicProvides the UniDic 2.3.0 Japanese morphological dictionary for use…
  • fugashiA Cython wrapper for MeCab that tokenizes and performs morphological…
  • unidic-liteProvides a pip-installable Japanese morphological analysis dictionary…

Give your agent the search over MCP, or paste the wish link into any chat.

More Linguistic packages

charset-normalizer Worth it
PyPI · Utilities · released Aug 2026

Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.

permissive licensepure Python · 3.7+
1.7Bdownloads / mo
tiktoken Worth it
PyPI · Linguistic · released May 2026

tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.

Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.

permissive licensecompiled wheel · 3.9+
233.0Mdownloads / mo
chardet Worth it
PyPI · Python Modules · released Aug 2026

Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.

Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.

0BSDpure Python · 3.10+
199.0Mdownloads / mo
text-unidecode With conditions
PyPI · Python Modules · released Aug 2019

Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.

However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.

GPL-2.0-or-laterpure Pythonabandoned
89.0Mdownloads / mo
lark Worth it
PyPI · Python Modules · released Oct 2025

Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.

MITpure Python · 3.8+
79.7Mdownloads / mo
tree-sitter Worth it
PyPI · Linguistic · released Jun 2026

Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.

MITcompiled wheel · 3.10+
79.0Mdownloads / mo

See also unidic-lite · ipadic · mecab-python3 · mecab-ko-dic · fugashi · mecab · mojimoji · ginza · konoha · mecab-ko