$npx skillfedfor your agent

charset-normalizer

The Real First Universal Charset Detector. Open, modern and actively maintained alternative to Chardet.

Worth itPyPI UtilitiesReleased Aug 20261.7B downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — charset_normalizer-3.5.0-py3-none-any.whl
v3.5.0 · released 2026-08-12 · Python >=3.7

Yes. Charset-normalizer is actively maintained, has no dependencies, installs easily, carries a permissive MIT license, and ranks in the top 100 most-downloaded PyPI packages. It solves a real problem—unknown text encoding—with a pure-Python implementation that works across Python 3.7 through 3.15. Use it whenever you need to detect or normalize text from untrusted or legacy sources.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Low install friction with a pure-Python wheel distribution.
  • Active maintenance with a release within the last 2 days and no runtime dependencies to manage.

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; you must include a copy of the license and copyright notice in distributions.

last release 2026-08-12 (2 days) · last repo commit 2026-08-13 · 787 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,674,555,757 downloads/mo, #8 on PyPI

Verify before relying

pip install charset-normalizer

from charset_normalizer import from_bytes

result = from_bytes(b'\xc3\xa9').best()
print(result.encoding)  # utf-8
print(str(result))      # é
  • Whether the 98% accuracy figure cited in the description applies to your specific text sources and encoding mix.
  • Performance characteristics on your actual workload—the benchmarks shown use a specific dataset of 477 files.
  • Whether custom codec registration capability is documented with examples for your use case.
Same gist for agents: .md · .json

What it is and what it does

Charset-normalizer is a library for detecting the encoding of text when you don't know it in advance. It reads bytes and attempts to identify which character set was used to encode them, then can normalize or decode the text accordingly. The package supports all IANA character set names that Python's standard library can handle, and lets you register additional custom codecs if needed.

Unlike some alternatives, it's written in pure Python with no compiled dependencies, making it portable across platforms and Python implementations. It aims to be faster and more reliable than earlier detection approaches, particularly in edge cases where encoding is ambiguous. The library handles UnicodeDecodeError safety and can also detect spoken language from text, though its primary purpose is solving the charset detection problem.

Use it for

  • Automatically detect encoding when reading files from unknown sources or legacy systems before processing their text.
  • Normalize text from web scraping or API responses where the charset header is missing or incorrect.
  • Build a data pipeline that ingests files in mixed encodings and standardizes them to UTF-8 for downstream processing.
  • Handle user-uploaded files in a web application without requiring them to specify their encoding manually.
  • Process log files or database exports where the original encoding is undocumented or inconsistent across records.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Charset-normalizer is actively maintained, has no dependencies, installs easily, carries a permissive MIT license, and ranks in the top 100 most-downloaded PyPI packages. It solves a real problem—unknown text encoding—with a pure-Python implementation that works across Python 3.7 through 3.15. Use it whenever you need to detect or normalize text from untrusted or legacy sources.

Install

charset-normalizer on PyPI

Before you install

Low install friction with a pure-Python wheel distribution. Active maintenance with a release within the last 2 days and no runtime dependencies to manage.

License in practice

MIT license permits commercial and private use with minimal restrictions; you must include a copy of the license and copyright notice in distributions.

Quickstart

pip install charset-normalizer

from charset_normalizer import from_bytes

result = from_bytes(b'\xc3\xa9').best()
print(result.encoding)  # utf-8
print(str(result))      # é

Verify before relying

  • Whether the 98% accuracy figure cited in the description applies to your specific text sources and encoding mix.
  • Performance characteristics on your actual workload—the benchmarks shown use a specific dataset of 477 files.
  • Whether custom codec registration capability is documented with examples for your use case.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.7
Install frictionLow. Pure-Python wheel
Runtime dependenciesNone
MaintenanceActively maintained 2 days since the last release
Last repo commit
First released
Downloads1,674,555,757 / month, #8 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableEnvironment :: WebAssemblyEnvironment :: WebAssembly :: EmscriptenIntended Audience :: DevelopersOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.15Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Programming Language :: Python :: Free Threading :: 4 - ResilientProgramming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPyTopic :: Text Processing :: LinguisticTopic :: UtilitiesTyping :: Typed

Evidence: charset_normalizer-3.5.0-py3-none-any.whl

Tags

Capabilities
charset detectionencoding detectiontext encoding identifiercharacter set detectorunicode normalizationchardet alternativedetect text encoding
Topics
encoding-detectiontext-processingcharset-handling
PyPI keywords
encodingcharsetcharset-detectordetectornormalizationunicodechardetdetect

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “charset detection”

  • charset-normalizerDetects and normalizes text encoding from unknown or ambiguous…
  • webencodingsImplements the WHATWG Encoding standard to map legacy web character…
  • cchardetcchardet detects the character encoding of byte strings using a C…

Give your agent the search over MCP, or paste the wish link into any chat.

More Utilities packages

idna Worth it
PyPI · Python Modules · released Jun 2026

Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.

Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.

BSD-3-Clausepure Python · 3.9+
1.8Bdownloads / mo
setuptools Worth it
PyPI · Python Modules · released Aug 2026

Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.

MITpure Python · 3.10+
1.6Bdownloads / mo
pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
Pygments Worth it
PyPI · Utilities · released Mar 2026

Pygments is a syntax highlighter that colorizes source code and text in over 500 languages and formats, outputting to HTML, LaTeX, RTF, SVG, images, or ANSI terminal sequences.

Install it if you need to display or transform source code.

BSD-2-Clausepure Python · 3.9+
1.3Bdownloads / mo
six With conditions
PyPI · Libraries · released Dec 2024

Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.

MITpure Python
1.2Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo

See also chardet · encutils · mutf8 · cchardet · faust-cchardet · webencodings · mbstrdecoder · bnunicodenormalizer · gibberish-detector · wetext