gcld3
CLD3 is a neural network model for language identification.
Decision gist · record as of 2026-08-14
Yes, if you need lightweight language detection for a stable, well-defined use case and can accept that the package will not receive updates. The model is frozen and the package has no runtime dependencies, so it will continue to work predictably. No, if you require support for modern Python versions, active maintenance, or security patches—use an actively maintained alternative instead.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires prebuilt wheels for specific Python versions on macOS or manylinux2010; building from source requires Chromium repository and build tools.
- Medium install friction due to compiled wheels for specific Python versions and platforms (macOS, manylinux2010).
- Package is abandoned as of 2023-05-24 with no active maintenance.
License · maintenance · safety
(unclear) — License status is unclear—no SPDX identifier or raw license text provided. Verify licensing terms before use in proprietary or redistributed projects.
last release 2020-08-19 (2186 days) · last repo commit 2023-05-24 · 885 stars · archived
0 known vulnerabilities (OSV.dev, 2026-08-14) · 109,802 downloads/mo, #12,498 on PyPI
Alternatives
Verify before relying
pip install gcld3
import gcld3
detector = gcld3.NNetLanguageIdentifier(min_num_bytes=0, max_num_bytes=0)
result = detector.FindLanguage(text="Hello world")
print(result.language, result.probability)- Whether the package works with Python versions beyond those with prebuilt wheels
- Current stability and accuracy of the trained model given no updates since 2020-08-19
- Whether the archived repository will accept security patches or bug fixes
What it is and what it does
gcld3 is a Python wrapper around Google's Compact Language Detector v3, a neural network model trained to identify the language of text. It extracts character n-grams from input, hashes them to embedding vectors, and passes the averaged embeddings through a small neural network to predict language. The model outputs BCP-47 language codes and can distinguish between different scripts of the same language.
The package ships with precompiled inference code and a trained model, so no training or model download is required at runtime. It has no runtime dependencies. However, the package is abandoned—the repository was archived and the last release was 2020-08-19, meaning no bug fixes, security updates, or support for newer Python versions are expected.
Use it for
- Automatically categorize user-generated content by language for multilingual applications or content moderation pipelines.
- Detect the language of incoming text in chatbots or support systems to route to appropriate handlers.
- Identify the script used in text for rendering or font selection in international user interfaces.
- Preprocess text corpora by filtering or grouping documents by detected language before NLP analysis.
- Validate or correct language metadata in multilingual datasets or document stores.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need lightweight language detection for a stable, well-defined use case and can accept that the package will not receive updates.
The model is frozen and the package has no runtime dependencies, so it will continue to work predictably. No, if you require support for modern Python versions, active maintenance, or security patches—use an actively maintained alternative instead.
Install
gcld3 on PyPI
Before you install
Medium install friction due to compiled wheels for specific Python versions and platforms (macOS, manylinux2010). Package is abandoned as of 2023-05-24 with no active maintenance.
Requires prebuilt wheels for specific Python versions on macOS or manylinux2010; building from source requires Chromium repository and build tools.
License in practice
License status is unclear—no SPDX identifier or raw license text provided. Verify licensing terms before use in proprietary or redistributed projects.
Quickstart
pip install gcld3
import gcld3
detector = gcld3.NNetLanguageIdentifier(min_num_bytes=0, max_num_bytes=0)
result = detector.FindLanguage(text="Hello world")
print(result.language, result.probability)
Verify before relying
- Whether the package works with Python versions beyond those with prebuilt wheels
- Current stability and accuracy of the trained model given no updates since 2020-08-19
- Whether the archived repository will accept security patches or bug fixes
Package facts
| License | Not declared unclear |
| Python support | Not specified |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Abandoned 2,186 days since the last release |
| Last repo commit | repository archived |
| First released | |
| Downloads | 109,802 / month, #12,498 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: gcld3-3.0.13-cp36-cp36m-macosx_10_9_x86_64.whl; gcld3-3.0.13-cp36-cp36m-manylinux2010_x86_64.whl; gcld3-3.0.13-cp38-cp38-macosx_10_9_x86_64.whl; gcld3-3.0.13-cp38-cp38-manylinux2010_x86_64.whl; gcld3-3.0.13-pp36-pypy36_pp73-macosx_10_9_x86_64.whl; gcld3-3.0.13-pp36-pypy36_pp73-manylinux2010_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “language code prediction”
- gcld3Identifies the language of input text using a neural network model,…
- fair-esmProvides pre-trained transformer protein language models (ESM-2,…
- llama-index-llms-anthropicIntegrates Anthropic's Claude language models into LlamaIndex…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also cld3-py · pycld3 · lingua-language-detector · py3langid · pycld2 · fasttext-langdetect · langdetect · language-data · clabe · django-zxcvbn-password-validator