langid
langid.py is a standalone Language Identification (LangID) tool.
Decision gist · record as of 2026-08-14
Yes, if you need offline language detection across many languages and can accept an unmaintained codebase. The package is stable, carries no known vulnerabilities, and has a permissive license. However, do not use it if you require active maintenance, support for modern Python versions beyond basic compatibility, or confidence that the model reflects current language patterns. For new projects, consider actively maintained alternatives.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires numpy; Python 2.7+ or Python 3 support is present but training tools are Python 2-only.
- High install friction due to numpy dependency and no active maintenance since 2016.
- The package is marked abandoned with no commits since 2020-01-01, though it remains in production/stable status.
License · maintenance · safety
BSD (permissive) — BSD license is permissive, allowing commercial and private use with minimal restrictions—suitable for most projects that can accept a mature, unmaintained codebase.
last release 2016-04-05 (3783 days) · last repo commit 2020-01-01 · 2,463 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 713,451 downloads/mo, #5,253 on PyPI
Alternatives
Verify before relying
pip install langid
import langid
result = langid.classify("This is a test")
print(result) # ('en', -54.41310358047485)- Whether the pre-trained model on 97 languages remains accurate for modern text and evolving language usage patterns.
- Performance characteristics (speed, memory) on large-scale batch processing or real-time web service deployments.
- Compatibility with current Python 3 versions and whether the package works reliably on modern systems.
What it is and what it does
langid.py is a standalone language identification tool that classifies text into one of 97 pre-trained languages using a statistical model. It operates in three modes: as a command-line tool for interactive or batch processing, as a Python library for programmatic use, and as a WSGI-compliant web service. The model was trained on diverse sources including Wikipedia, Reuters, and ClueWeb data, and is designed to be fast and robust to domain-specific features like HTML markup.
The package requires only Python 2.7+ (or Python 3) and numpy. It returns both a language code (ISO 639-1) and a confidence score in log-probability space, with optional normalization to the 0–1 range. You can constrain predictions to a specific set of languages and process input from stdin, files, or HTTP requests. However, the project has been abandoned since 2016 with no active maintenance, so it may not reflect modern language usage or receive bug fixes.
Use it for
- Classify user-submitted text by language in a web application or API without external service calls.
- Batch process document collections to sort or filter by detected language across 97 supported languages.
- Deploy as a lightweight WSGI web service for language detection in microservice architectures.
- Constrain language detection to a known subset (e.g., en, de, fr) to improve accuracy in multilingual but bounded scenarios.
- Identify language in log files or text streams line-by-line for monitoring or data pipeline preprocessing.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need offline language detection across many languages and can accept an unmaintained codebase.
The package is stable, carries no known vulnerabilities, and has a permissive license. However, do not use it if you require active maintenance, support for modern Python versions beyond basic compatibility, or confidence that the model reflects current language patterns. For new projects, consider actively maintained alternatives.
Install
langid on PyPI
Before you install
High install friction due to numpy dependency and no active maintenance since 2016. The package is marked abandoned with no commits since 2020-01-01, though it remains in production/stable status.
Requires numpy; Python 2.7+ or Python 3 support is present but training tools are Python 2-only.
License in practice
BSD license is permissive, allowing commercial and private use with minimal restrictions—suitable for most projects that can accept a mature, unmaintained codebase.
Quickstart
pip install langid
import langid
result = langid.classify("This is a test")
print(result) # ('en', -54.41310358047485)
Verify before relying
- Whether the pre-trained model on 97 languages remains accurate for modern text and evolving language usage patterns.
- Performance characteristics (speed, memory) on large-scale batch processing or real-time web service deployments.
- Compatibility with current Python 3 versions and whether the package works reliably on modern systems.
Package facts
| License | BSD permissive |
| Python support | Not specified |
| Install friction | High. Source build required |
| Runtime dependencies | None |
| Maintenance | Abandoned 3,783 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 713,451 / month, #5,253 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: BSD LicenseProgramming Language :: Python :: 2Programming Language :: Python :: 2.7Programming Language :: Python :: 3Topic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Text Processing :: Linguistic |
Evidence: langid-1.1.6.tar.gz
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “language detection”
- langidIdentifies the language of text input across 97 languages using a…
- lingua-language-detectorDetects which language a text is written in, supporting 75 languages…
- langdetectDetects the language of text input, supporting 55 languages and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also py3langid · langdetect · fasttext-langdetect · lingua-language-detector · stopwordsiso · pycld3 · chardet · qwen-asr · pycld2 · langcodes