charset-normalizer
The Real First Universal Charset Detector. Open, modern and actively maintained alternative to Chardet.
Decision gist · record as of 2026-08-14
Yes. Charset-normalizer is actively maintained, has no dependencies, installs easily, carries a permissive MIT license, and ranks in the top 100 most-downloaded PyPI packages. It solves a real problem—unknown text encoding—with a pure-Python implementation that works across Python 3.7 through 3.15. Use it whenever you need to detect or normalize text from untrusted or legacy sources.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with a pure-Python wheel distribution.
- Active maintenance with a release within the last 2 days and no runtime dependencies to manage.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; you must include a copy of the license and copyright notice in distributions.
last release 2026-08-12 (2 days) · last repo commit 2026-08-13 · 787 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,674,555,757 downloads/mo, #8 on PyPI
Alternatives
Verify before relying
pip install charset-normalizer
from charset_normalizer import from_bytes
result = from_bytes(b'\xc3\xa9').best()
print(result.encoding) # utf-8
print(str(result)) # é- Whether the 98% accuracy figure cited in the description applies to your specific text sources and encoding mix.
- Performance characteristics on your actual workload—the benchmarks shown use a specific dataset of 477 files.
- Whether custom codec registration capability is documented with examples for your use case.
What it is and what it does
Charset-normalizer is a library for detecting the encoding of text when you don't know it in advance. It reads bytes and attempts to identify which character set was used to encode them, then can normalize or decode the text accordingly. The package supports all IANA character set names that Python's standard library can handle, and lets you register additional custom codecs if needed.
Unlike some alternatives, it's written in pure Python with no compiled dependencies, making it portable across platforms and Python implementations. It aims to be faster and more reliable than earlier detection approaches, particularly in edge cases where encoding is ambiguous. The library handles UnicodeDecodeError safety and can also detect spoken language from text, though its primary purpose is solving the charset detection problem.
Use it for
- Automatically detect encoding when reading files from unknown sources or legacy systems before processing their text.
- Normalize text from web scraping or API responses where the charset header is missing or incorrect.
- Build a data pipeline that ingests files in mixed encodings and standardizes them to UTF-8 for downstream processing.
- Handle user-uploaded files in a web application without requiring them to specify their encoding manually.
- Process log files or database exports where the original encoding is undocumented or inconsistent across records.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Charset-normalizer is actively maintained, has no dependencies, installs easily, carries a permissive MIT license, and ranks in the top 100 most-downloaded PyPI packages. It solves a real problem—unknown text encoding—with a pure-Python implementation that works across Python 3.7 through 3.15. Use it whenever you need to detect or normalize text from untrusted or legacy sources.
Install
charset-normalizer on PyPI
Before you install
Low install friction with a pure-Python wheel distribution. Active maintenance with a release within the last 2 days and no runtime dependencies to manage.
License in practice
MIT license permits commercial and private use with minimal restrictions; you must include a copy of the license and copyright notice in distributions.
Quickstart
pip install charset-normalizer
from charset_normalizer import from_bytes
result = from_bytes(b'\xc3\xa9').best()
print(result.encoding) # utf-8
print(str(result)) # é
Verify before relying
- Whether the 98% accuracy figure cited in the description applies to your specific text sources and encoding mix.
- Performance characteristics on your actual workload—the benchmarks shown use a specific dataset of 477 files.
- Whether custom codec registration capability is documented with examples for your use case.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.7 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 2 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,674,555,757 / month, #8 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: WebAssemblyEnvironment :: WebAssembly :: EmscriptenIntended Audience :: DevelopersOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.15Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Programming Language :: Python :: Free Threading :: 4 - ResilientProgramming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPyTopic :: Text Processing :: LinguisticTopic :: UtilitiesTyping :: Typed |
Evidence: charset_normalizer-3.5.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “charset detection”
- charset-normalizerDetects and normalizes text encoding from unknown or ambiguous…
- webencodingsImplements the WHATWG Encoding standard to map legacy web character…
- cchardetcchardet detects the character encoding of byte strings using a C…
Give your agent the search over MCP, or paste the wish link into any chat.
More Utilities packages
Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.
Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.
Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Pygments is a syntax highlighter that colorizes source code and text in over 500 languages and formats, outputting to HTML, LaTeX, RTF, SVG, images, or ANSI terminal sequences.
Install it if you need to display or transform source code.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also chardet · encutils · mutf8 · cchardet · faust-cchardet · webencodings · mbstrdecoder · bnunicodenormalizer · gibberish-detector · wetext