pyvi
Python Vietnamese Toolkit
Decision gist · record as of 2026-08-14
Yes, if you need Vietnamese NLP and can tolerate dormant maintenance. The package has no known vulnerabilities, low install friction, and a permissive MIT license. However, verify that scikit-learn and sklearn-crfsuite versions remain compatible with your Python environment, and be aware that no updates have been released since 2021.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with a pure-Python wheel.
- Maintenance is dormant—last release was 2021-06-30 and the repository shows no recent activity, though it remains unarchived with a final commit on 2024-09-26.
License · maintenance · safety
MIT (permissive) — MIT license is permissive, allowing commercial and private use with minimal restrictions.
last release 2021-06-30 (1871 days) · last repo commit 2024-09-26 · 278 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 165,976 downloads/mo, #10,512 on PyPI
Alternatives
Verify before relying
pip install pyvi
from pyvi import ViTokenizer, ViPosTagger, ViUtils
ViTokenizer.tokenize(u"Trường đại học bách khoa hà nội")
ViPosTagger.postagging(ViTokenizer.tokenize(u"Trường đại học Bách Khoa Hà Nội"))
ViUtils.remove_accents(u"Trường đại học bách khoa hà nội")- Whether scikit-learn and sklearn-crfsuite are downloaded as pre-trained models or require separate setup.
- Current compatibility with Python versions beyond 3.x, given classifiers list Python 2.6 and 2.7.
What it is and what it does
Pyvi is a Vietnamese natural language processing toolkit that handles tokenization, part-of-speech tagging, and accent manipulation. It uses conditional random fields as its underlying algorithm and reports an F1 score of 0.985 for tokenization and 0.925 for POS tagging. The package depends on scikit-learn and sklearn-crfsuite for its machine learning operations.
The toolkit is designed for developers working with Vietnamese text who need to split sentences into tokens, identify grammatical roles (adjectives, nouns, verbs, etc.), or normalize accents. It integrates with spacy.io and handles common Vietnamese text issues like redundant spacing. However, the package has been dormant since mid-2021, with no recent updates or active maintenance.
Use it for
- Tokenize Vietnamese sentences into words for downstream NLP pipelines or text analysis.
- Tag Vietnamese words with their grammatical roles (noun, verb, adjective, etc.) for linguistic analysis.
- Remove or add diacritical marks to Vietnamese text for normalization or text processing workflows.
- Prepare Vietnamese text for machine learning models that require tokenized and tagged input.
- Integrate Vietnamese language support into spacy.io-based NLP applications.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need Vietnamese NLP and can tolerate dormant maintenance.
The package has no known vulnerabilities, low install friction, and a permissive MIT license. However, verify that scikit-learn and sklearn-crfsuite versions remain compatible with your Python environment, and be aware that no updates have been released since 2021.
Install
pyvi on PyPI
Before you install
Low install friction with a pure-Python wheel. Maintenance is dormant—last release was 2021-06-30 and the repository shows no recent activity, though it remains unarchived with a final commit on 2024-09-26.
License in practice
MIT license is permissive, allowing commercial and private use with minimal restrictions.
Quickstart
pip install pyvi
from pyvi import ViTokenizer, ViPosTagger, ViUtils
ViTokenizer.tokenize(u"Trường đại học bách khoa hà nội")
ViPosTagger.postagging(ViTokenizer.tokenize(u"Trường đại học Bách Khoa Hà Nội"))
ViUtils.remove_accents(u"Trường đại học bách khoa hà nội")
Verify before relying
- Whether scikit-learn and sklearn-crfsuite are downloaded as pre-trained models or require separate setup.
- Current compatibility with Python versions beyond 3.x, given classifiers list Python 2.6 and 2.7.
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagesscikit-learnsklearn-crfsuite |
| Maintenance | Dormant 1,871 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 165,976 / month, #10,512 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseNatural Language :: VietnameseProgramming Language :: Python :: 2Programming Language :: Python :: 2.6Programming Language :: Python :: 2.7Programming Language :: Python :: 3 |
Evidence: pyvi-0.1.1-py2.py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “vietnamese tokenization”
- pyviProvides Vietnamese language processing tools including tokenization,…
- misakiConverts written text to phonetic representations…
- vnaiVnAI provides a Python interface for accessing Vietnamese stock…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also vnstock · sea-g2p · textblob · nagisa · soynlp · polyglot · pythainlp · jieba3k · urduhack · vnstock-ezchart