py3langid
Fork of the language identification tool langid.py, featuring a modernized codebase and faster execution times.
Decision gist · record as of 2026-08-14
Yes, if you need fast language detection across 97 languages with minimal dependencies. The dormant maintenance status is not a blocker—the package is stable, has no known vulnerabilities, and the last commit is recent enough. Install if you want a lightweight, performant drop-in replacement for the original langid.py; skip if you need active development or support for languages outside the pre-trained set.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >= 3.8 and numpy; model loading on first use may take a moment.
- Low friction: pure Python wheel with only numpy as a runtime dependency.
- Dormant maintenance (last release 787 days ago, last commit 2024-11-22), but stable and archived repository shows no active development pressure.
License · maintenance · safety
BSD (permissive) — BSD-3-Clause permissive license allows commercial and private use with minimal restrictions; attribution and license notice required in distributions.
last release 2024-06-18 (787 days) · last repo commit 2024-11-22 · 63 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 437,924 downloads/mo, #6,663 on PyPI
Alternatives
Verify before relying
pip install py3langid
import py3langid as langid
lang, prob = langid.classify('This text is in English.')
print(lang, prob) # ('en', -56.77429)- Whether the 30% import speedup and 25-30x model loading speedup claims are reproducible in current environments.
- Accuracy consistency across the 97 supported languages and whether results degrade on mixed-language or code-heavy text.
What it is and what it does
py3langid is a modernized fork of langid.py that detects which of 97 languages a piece of text is written in. It returns both the ISO 639-1 language code and a confidence score. The package is designed to be a drop-in replacement for the original, with significant performance improvements: imports are about 30% faster, model loading is 25-30x faster, and classification itself is 5-6x faster on paragraphs.
The tool works as both a Python library and a command-line utility. It uses a pre-trained statistical model built from diverse sources (JRC-Acquis, ClueWeb 09, Wikipedia, Reuters RCV2, Debian i18n) and requires only numpy as a dependency. It supports language subsetting, probability normalization, and can be deployed as a WSGI web service.
Use it for
- Automatically categorize user-generated content by language for multilingual applications or content management systems.
- Filter or route text documents to language-specific processing pipelines in NLP workflows.
- Identify the dominant language in mixed-language documents or social media posts for downstream analysis.
- Build language detection into a command-line data pipeline using the standalone langid tool.
- Normalize confidence scores to probabilities for machine learning feature engineering or confidence thresholding.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need fast language detection across 97 languages with minimal dependencies.
The dormant maintenance status is not a blocker—the package is stable, has no known vulnerabilities, and the last commit is recent enough. Install if you want a lightweight, performant drop-in replacement for the original langid.py; skip if you need active development or support for languages outside the pre-trained set.
Install
py3langid on PyPI
Before you install
Low friction: pure Python wheel with only numpy as a runtime dependency. Dormant maintenance (last release 787 days ago, last commit 2024-11-22), but stable and archived repository shows no active development pressure.
Requires Python >= 3.8 and numpy; model loading on first use may take a moment.
License in practice
BSD-3-Clause permissive license allows commercial and private use with minimal restrictions; attribution and license notice required in distributions.
Quickstart
pip install py3langid
import py3langid as langid
lang, prob = langid.classify('This text is in English.')
print(lang, prob) # ('en', -56.77429)
Verify before relying
- Whether the 30% import speedup and 25-30x model loading speedup claims are reproducible in current environments.
- Accuracy consistency across the 97 supported languages and whether results degrade on mixed-language or code-heavy text.
Package facts
| License | BSD permissive |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagenumpy |
| Maintenance | Dormant 787 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 437,924 / month, #6,663 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchLicense :: OSI Approved :: BSD LicenseOperating System :: MacOS :: MacOS XOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Text Processing :: Linguistic |
Evidence: py3langid-0.3.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “langid fork python”
- py3langidIdentifies the language of text in one of 97 languages using a…
- langidIdentifies the language of text input across 97 languages using a…
- subprocess32A backport of Python 3's subprocess module to Python 2, with…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also langid · langdetect · fasttext-langdetect · lingua-language-detector · gcld3 · pycld2 · identify · pycld3 · qwen-asr · chardet