pycld2
Python bindings around Google Chromium's embedded compact language detection library (CLD2)
Decision gist · record as of 2026-08-14
Yes, if you need lightweight, fast language detection across many languages without external dependencies or neural models. The prebuilt wheels make installation straightforward on common platforms. However, consider the aging maintenance status (last release 504 days ago): verify that it works reliably with your target Python version and that CLD2's detection quality meets your accuracy requirements. If you need state-of-the-art accuracy, CLD3 or neural alternatives may be better; if you need speed and simplicity, pycld2 is a solid choice.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Input must be UTF-8 encoded (str or UTF-8 bytes); non-UTF-8 bytes will raise pycld2.error.
- Medium install friction due to compiled C++ bindings; prebuilt wheels are available for Python 3.8–3.12 on macOS, Linux (x86_64, i686, musl), and Windows.
- Last release was 504 days ago; the project is aging but the repository remains active.
License · maintenance · safety
Apache2 (permissive) — Licensed under Apache 2.0 (permissive). Safe for commercial and proprietary use with minimal restrictions; attribution and license inclusion are required.
last release 2025-03-28 (504 days) · last repo commit 2025-03-28 · 179 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 774,217 downloads/mo, #5,096 on PyPI
Alternatives
Verify before relying
import pycld2 as cld2
isReliable, textBytesFound, details = cld2.detect(
"а неправильный формат идентификатора дн назад"
)
print(details[0])
# ('RUSSIAN', 'ru', 98, 404.0)- Whether the aging maintenance status (504 days since last release) affects reliability or security for new Python versions beyond 3.12.
- Performance characteristics and memory footprint for large-scale batch language detection.
- Accuracy comparison with CLD3 or other modern language detection libraries.
What it is and what it does
pycld2 is a Python wrapper around Google's Compact Language Detect 2 (CLD2), a C++ library that identifies the language of text. It consolidates the upstream CLD2 library with its bindings into a single pip-installable package, supporting detection across over 165 languages by default. The package exposes a single primary function, `detect()`, which takes UTF-8 text (as str or bytes) and returns a reliability flag, byte count, and a list of detected languages with confidence scores. Optionally, it can return byte-range vectors showing which language was detected in which part of the input, and accepts hints (top-level domain, language preference, encoding) to influence detection.
The package is built on precompiled wheels for Python 3.8–3.12 across macOS, Linux, and Windows, eliminating the need to compile C++ locally in most cases. It has no runtime dependencies beyond the C++ bindings. The API is minimal and straightforward: call `detect()` with your text and optional parameters, and receive structured results. It is useful for applications that need fast, lightweight language identification without the overhead of neural models, though the aging maintenance status (last release 504 days ago) means it may not have been tested against the very latest Python patch releases.
Use it for
- Automatically categorize user-generated content by language in multilingual web applications or content management systems.
- Pre-process text corpora for natural language processing pipelines by filtering or routing documents to language-specific models.
- Detect language in HTTP requests or HTML documents, using optional hints from domain or HTTP headers to improve accuracy.
- Identify the primary language in mixed-language documents and extract byte ranges showing where each language appears.
- Triage support tickets or chat messages by language to route them to appropriate language-specific teams or handlers.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need lightweight, fast language detection across many languages without external dependencies or neural models.
The prebuilt wheels make installation straightforward on common platforms. However, consider the aging maintenance status (last release 504 days ago): verify that it works reliably with your target Python version and that CLD2's detection quality meets your accuracy requirements. If you need state-of-the-art accuracy, CLD3 or neural alternatives may be better; if you need speed and simplicity, pycld2 is a solid choice.
Install
pycld2 on PyPI
Before you install
Medium install friction due to compiled C++ bindings; prebuilt wheels are available for Python 3.8–3.12 on macOS, Linux (x86_64, i686, musl), and Windows. Last release was 504 days ago; the project is aging but the repository remains active.
Input must be UTF-8 encoded (str or UTF-8 bytes); non-UTF-8 bytes will raise pycld2.error.
License in practice
Licensed under Apache 2.0 (permissive). Safe for commercial and proprietary use with minimal restrictions; attribution and license inclusion are required.
Quickstart
import pycld2 as cld2
isReliable, textBytesFound, details = cld2.detect(
"а неправильный формат идентификатора дн назад"
)
print(details[0])
# ('RUSSIAN', 'ru', 98, 404.0)
Verify before relying
- Whether the aging maintenance status (504 days since last release) affects reliability or security for new Python versions beyond 3.12.
- Performance characteristics and memory footprint for large-scale batch language detection.
- Accuracy comparison with CLD3 or other modern language detection libraries.
Package facts
| License | Apache2 permissive |
| Python support | Not specified |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Aging 504 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 774,217 / month, #5,096 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: Apache Software LicenseOperating System :: MacOS :: MacOS XOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: C++Programming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Text Processing :: Linguistic |
Evidence: pycld2-0.42-cp310-cp310-macosx_11_0_arm64.whl; pycld2-0.42-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; pycld2-0.42-cp310-cp310-manylinux_2_5_i686.manylinux1_i686.manylinux_2_17_i686.manylinux2014_i686.whl; pycld2-0.42-cp310-cp310-musllinux_1_2_i686.whl; pycld2-0.42-cp310-cp310-musllinux_1_2_x86_64.whl; pycld2-0.42-cp310-cp310-win32.whl; pycld2-0.42-cp310-cp310-win_amd64.whl; pycld2-0.42-cp311-cp311-macosx_11_0_arm64.whl; pycld2-0.42-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; pycld2-0.42-cp311-cp311-manylinux_2_5_i686.manylinux1_i686.manylinux_2_17_i686.manylinux2014_i686.whl; pycld2-0.42-cp311-cp311-musllinux_1_2_i686.whl; pycld2-0.42-cp311-cp311-musllinux_1_2_x86_64.whl; pycld2-0.42-cp311-cp311-win32.whl; pycld2-0.42-cp311-cp311-win_amd64.whl; pycld2-0.42-cp312-cp312-macosx_11_0_arm64.whl; pycld2-0.42-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; pycld2-0.42-cp312-cp312-manylinux_2_5_i686.manylinux1_i686.manylinux_2_17_i686.manylinux2014_i686.whl; pycld2-0.42-cp312-cp312-musllinux_1_2_i686.whl; pycld2-0.42-cp312-cp312-musllinux_1_2_x86_64.whl; pycld2-0.42-cp312-cp312-win32.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “detect text language”
- pycld2Detects the language of text using Google's CLD2 library, supporting…
- pycld3Detects the language of text using Google's Compact Language Detector…
- fast-langdetectDetects the language of text using FastText models, returning…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also pycld3 · spacy-language-detection · fasttext-langdetect · cld3-py · gcld3 · lingua-language-detector · py3langid · langdetect · mbstrdecoder · cnstd