kiwipiepy
Kiwi, the Korean Tokenizer for Python
Decision gist · record as of 2026-08-14
Yes, if you need Korean morphological analysis. The package is actively maintained, has no known vulnerabilities, and offers prebuilt wheels for common platforms. LGPL v3 licensing requires careful review for proprietary use. Medium install friction is acceptable given the compiled nature of the underlying C++ library. Suitable for both research and production Korean NLP pipelines.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires cmake 3.12+ if installing from source on platforms without prebuilt wheels; binary distributions available for Python 3.9+ on macOS (10.14+), Linux (manylinux2014), and Windows (Vista+).
- Medium install friction due to compiled C++ components; prebuilt wheels available for Python 3.9+ on major platforms (macOS, Linux, Windows), but source builds require cmake 3.12+ and a C++17 compiler.
- Active maintenance with recent releases.
License · maintenance · safety
LGPL v3 License (copyleft) — Licensed under LGPL v3 (copyleft); derivative works and modifications must be distributed under the same license. Suitable for open-source projects but requires careful review for proprietary or closed-source use.
last release 2026-06-11 (64 days) · last repo commit 2026-08-14 · 396 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 346,086 downloads/mo, #7,359 on PyPI
Alternatives
Verify before relying
pip install kiwipiepy
from kiwipiepy import Kiwi
kiwi = Kiwi()
tokens = kiwi.tokenize("안녕하세요 형태소 분석기 키위입니다.")
print(tokens)- Performance characteristics (speed, memory usage) on large-scale Korean text processing
- Accuracy metrics or benchmarks against other Korean tokenizers
- Whether the interactive test mode (python -m kiwipiepy) is suitable for production debugging
What it is and what it does
Kiwipiepy is a Python binding for Kiwi, a Korean morphological analyzer that breaks Korean text into morphemes (smallest meaningful units) and assigns part-of-speech tags based on the Sejong Corpus tagset. It provides tokenization, sentence splitting, stopword filtering, and user dictionary management for customizing analysis on domain-specific vocabulary.
The package ships with prebuilt wheels for modern Python versions (3.9+) on common platforms, reducing installation friction for most users. It depends on numpy, tqdm, and a model data package (kiwipiepy_model) that is automatically installed. The library is actively maintained and supports interactive testing via command-line interface, making it accessible for quick experimentation with Korean text analysis.
Use it for
- Tokenize Korean documents for natural language processing pipelines and text classification
- Extract morphemes and part-of-speech tags for linguistic analysis or corpus studies
- Split multi-sentence Korean text into individual sentences for batch processing
- Filter stopwords from Korean text before downstream NLP tasks like topic modeling
- Add domain-specific terminology to the analyzer via user dictionary for improved accuracy on specialized corpora
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need Korean morphological analysis.
The package is actively maintained, has no known vulnerabilities, and offers prebuilt wheels for common platforms. LGPL v3 licensing requires careful review for proprietary use. Medium install friction is acceptable given the compiled nature of the underlying C++ library. Suitable for both research and production Korean NLP pipelines.
Install
kiwipiepy on PyPI
Before you install
Medium install friction due to compiled C++ components; prebuilt wheels available for Python 3.9+ on major platforms (macOS, Linux, Windows), but source builds require cmake 3.12+ and a C++17 compiler. Active maintenance with recent releases.
Requires cmake 3.12+ if installing from source on platforms without prebuilt wheels; binary distributions available for Python 3.9+ on macOS (10.14+), Linux (manylinux2014), and Windows (Vista+).
License in practice
Licensed under LGPL v3 (copyleft); derivative works and modifications must be distributed under the same license. Suitable for open-source projects but requires careful review for proprietary or closed-source use.
Quickstart
pip install kiwipiepy
from kiwipiepy import Kiwi
kiwi = Kiwi()
tokens = kiwi.tokenize("안녕하세요 형태소 분석기 키위입니다.")
print(tokens)
Verify before relying
- Performance characteristics (speed, memory usage) on large-scale Korean text processing
- Accuracy metrics or benchmarks against other Korean tokenizers
- Whether the interactive test mode (python -m kiwipiepy) is suitable for production debugging
Package facts
| License | LGPL v3 License copyleft |
| Python support | Not specified |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 4 packagesdataclasseskiwipiepy_modelnumpytqdm |
| Maintenance | Actively maintained 64 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 346,086 / month, #7,359 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: GNU Lesser General Public License v3 (LGPLv3)Programming Language :: C++Programming Language :: Python :: 3Topic :: Software Development :: LibrariesTopic :: Text Processing :: Linguistic |
Evidence: kiwipiepy-0.23.2-cp314-cp314t-macosx_10_15_x86_64.whl; kiwipiepy-0.23.2-cp314-cp314t-macosx_11_0_arm64.whl; kiwipiepy-0.23.2-cp314-cp314t-manylinux2014_aarch64.manylinux_2_17_aarch64.whl; kiwipiepy-0.23.2-cp314-cp314t-manylinux2014_x86_64.manylinux_2_17_x86_64.whl; kiwipiepy-0.23.2-cp314-cp314t-win_amd64.whl; kiwipiepy-0.23.2-cp39-abi3-macosx_10_14_x86_64.whl; kiwipiepy-0.23.2-cp39-abi3-macosx_11_0_arm64.whl; kiwipiepy-0.23.2-cp39-abi3-manylinux2014_aarch64.manylinux_2_17_aarch64.whl; kiwipiepy-0.23.2-cp39-abi3-manylinux2014_x86_64.manylinux_2_17_x86_64.whl; kiwipiepy-0.23.2-cp39-abi3-win_amd64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “korean nlp library”
- kiwipiepyKiwipiepy tokenizes and analyzes Korean text into morphemes with…
- konlpyKoNLPy provides Korean natural language processing tools including…
- mecab-ko-dicProvides a Korean dictionary for MeCab tokenization, enabling…
Give your agent the search over MCP, or paste the wish link into any chat.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also kiwipiepy-model · mecab-ko · python-mecab-ko · mecab-ko-dic · rhoknp · konlpy · soynlp · python-mecab-ko-dic · g2pkk · pymorphy3