bpemb
Byte-pair embeddings in 275 languages
Decision gist · record as of 2026-08-14
Yes, if you need multilingual subword embeddings and can accept a dormant package. The low install friction, permissive license, and broad language coverage make it useful for research and prototyping. However, the 682-day gap since the last release means no bug fixes or updates are forthcoming; verify that the remote model repository remains accessible before relying on it in production.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Models and embeddings download automatically from nlp.h-its.org on first use; requires internet access and disk space for language-specific files.
- Low install friction with a pure Python wheel and five common dependencies (gensim, numpy, requests, sentencepiece, tqdm).
- The package is dormant—last release 682 days ago—so expect no active maintenance or updates.
License · maintenance · safety
MIT (permissive) — MIT license is permissive; you can use BPEmb freely in commercial and proprietary projects with minimal restrictions.
last release 2024-10-01 (682 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 124,620 downloads/mo, #11,860 on PyPI
Alternatives
Verify before relying
pip install bpemb
from bpemb import BPEmb
bpemb_en = BPEmb(lang="en", dim=50)
print(bpemb_en.encode("Stratford"))
# Models and embeddings download automatically on first use- Whether the remote model repository (nlp.h-its.org) remains stable and available long-term given the package's dormant status.
- Performance characteristics when working with very large texts or batch processing across many languages.
What it is and what it does
BPEmb is a collection of pre-trained subword embeddings covering 275 languages, built on Wikipedia using Byte-Pair Encoding (BPE). It wraps gensim's KeyedVectors to provide both subword tokenization and embedding vectors for use in neural NLP pipelines.
The package serves two main purposes: splitting text into subwords according to a learned vocabulary (with configurable vocabulary sizes from 1000 to 200000 tokens), and providing dense vector representations for those subwords. You load a language-specific model, encode text into subword tokens or token IDs, and either retrieve embeddings directly or use the built-in embed method. Models and embeddings download automatically on first use.
Use it for
- Tokenize text into subwords for morphologically rich or low-resource languages where word-level tokenization is insufficient.
- Initialize neural sequence models with pre-trained multilingual embeddings without training your own.
- Build cross-lingual NLP systems that share a common subword representation across languages.
- Experiment with different vocabulary sizes to balance granularity and sparsity for a specific language.
- Embed text fragments for similarity search or clustering in languages beyond English.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need multilingual subword embeddings and can accept a dormant package.
The low install friction, permissive license, and broad language coverage make it useful for research and prototyping. However, the 682-day gap since the last release means no bug fixes or updates are forthcoming; verify that the remote model repository remains accessible before relying on it in production.
Install
bpemb on PyPI
Before you install
Low install friction with a pure Python wheel and five common dependencies (gensim, numpy, requests, sentencepiece, tqdm). The package is dormant—last release 682 days ago—so expect no active maintenance or updates.
Models and embeddings download automatically from nlp.h-its.org on first use; requires internet access and disk space for language-specific files.
License in practice
MIT license is permissive; you can use BPEmb freely in commercial and proprietary projects with minimal restrictions.
Quickstart
pip install bpemb
from bpemb import BPEmb
bpemb_en = BPEmb(lang="en", dim=50)
print(bpemb_en.encode("Stratford"))
# Models and embeddings download automatically on first use
Verify before relying
- Whether the remote model repository (nlp.h-its.org) remains stable and available long-term given the package's dormant status.
- Performance characteristics when working with very large texts or batch processing across many languages.
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 5 packagesgensimnumpyrequestssentencepiecetqdm |
| Maintenance | Dormant 682 days since the last release |
| First released | |
| Downloads | 124,620 / month, #11,860 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: bpemb-0.3.6-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “multilingual subword embeddings”
- bpembBPEmb provides pre-trained subword embeddings in 275 languages based…
- floretfloret trains compact word embeddings using fastText's subword…
- fasttextfastText is a library for learning word embeddings and training text…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also curated-tokenizers · jieba3k · jieba · num2words · floret · spacy-pkuseg · model2vec · minisbd · text2num · sentence-transformers