$npx skillfedfor your agent

ftfy

Fixes mojibake and other problems with Unicode, after the fact

With conditionsPyPI Text ProcessingReleased Oct 202414.5M downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — ftfy-6.3.1-py3-none-any.whl
v6.3.1 · released 2024-10-26 · Python >=3.9 · 1 runtime deps: wcwidth

Yes, if you work with text from diverse or legacy sources. ftfy solves a real, hard problem (mojibake recovery) that few other tools address. Low install friction, no security issues, and active maintenance make it a safe dependency. The Apache license requires attribution but is otherwise permissive. Install it when text corruption is a known issue in your pipeline; skip it if your text is already clean.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.9 or later.
  • Low friction: pure Python wheel with a single runtime dependency (wcwidth).
  • Last release was 657 days ago; repo remains active with recent commits and no archived status, though maintenance is dormant.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license requires attribution to Robyn Speer. The package explicitly prohibits use in AI training datasets or derived works that obscure authorship; violators may be notified and required to remedy or delete copies.

last release 2024-10-26 (657 days) · last repo commit 2024-10-30 · 4,055 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 14,452,490 downloads/mo, #1,228 on PyPI

Verify before relying

pip install ftfy

from ftfy import fix_text
print(fix_text('âœ" No problems'))
# Output: ✔ No problems
  • Whether the package handles all real-world encoding scenarios or only a documented subset of common mojibake patterns.
  • Performance characteristics on very large text volumes or streaming input.
Same gist for agents: .md · .json

What it is and what it does

ftfy detects and repairs mojibake—text that was encoded as UTF-8 but decoded as a different encoding (or multiple times in succession)—by recognizing telltale byte patterns and recovering the original string. It handles complex cases including multiple layers of corruption, curly quotes applied over mojibake, non-breaking spaces mangled into regular spaces, and incorrectly capitalized HTML entities. The package is conservative: it avoids false positives by refusing to "fix" text that is already sensible, even if it could theoretically be reinterpreted as mojibake.

The library is used as a data-cleaning step in NLP research and text processing pipelines. It exposes a simple API (primarily `fix_text()` and `fix_encoding()`) and includes command-line tools. It depends only on wcwidth for character width calculations and supports current Python versions (3.9+).

Use it for

  • Clean scraped web content or user-generated text that has been corrupted by encoding mismatches during storage or transmission.
  • Preprocess text datasets for NLP research or machine learning to remove mojibake before training.
  • Repair legacy data imported from systems that mixed character encodings (e.g., UTF-8 decoded as Latin-1).
  • Decode HTML entities that appear outside HTML context, including non-standard capitalizations.
  • Fix text with multiple overlapping encoding errors that cannot be solved by a single decode operation.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you work with text from diverse or legacy sources.

ftfy solves a real, hard problem (mojibake recovery) that few other tools address. Low install friction, no security issues, and active maintenance make it a safe dependency. The Apache license requires attribution but is otherwise permissive. Install it when text corruption is a known issue in your pipeline; skip it if your text is already clean.

Install

ftfy on PyPI

Before you install

Low friction: pure Python wheel with a single runtime dependency (wcwidth). Last release was 657 days ago; repo remains active with recent commits and no archived status, though maintenance is dormant.

Requires Python 3.9 or later.

License in practice

Apache-2.0 permissive license requires attribution to Robyn Speer. The package explicitly prohibits use in AI training datasets or derived works that obscure authorship; violators may be notified and required to remedy or delete copies.

Quickstart

pip install ftfy

from ftfy import fix_text
print(fix_text('âœ" No problems'))
# Output: ✔ No problems

Verify before relying

  • Whether the package handles all real-world encoding scenarios or only a documented subset of common mojibake patterns.
  • Performance characteristics on very large text volumes or streaming input.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.9
Install frictionLow. Pure-Python wheel
Runtime dependencies
1 package
wcwidth
MaintenanceDormant 657 days since the last release
Last repo commit
First released
Downloads14,452,490 / month, #1,228 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: ftfy-6.3.1-py3-none-any.whl

Tags

Capabilities
fix mojibake unicodeencoding corruption repairgarbled text recoveryunicode text cleaningcharacter encoding fixestext decoding errorshtml entity decoding
Topics
text-repairencoding-recoverynlp-preprocessing

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “fix mojibake unicode”

  • ftfyDetects and fixes mojibake (garbled Unicode text caused by encoding…
  • win_unicode_consoleFixes Unicode input and display in Python when running from the…
  • mplfontsManages Matplotlib fonts and solves CJK (Chinese, Japanese, Korean)…

Give your agent the search over MCP, or paste the wish link into any chat.

More Text Processing packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyparsing Worth it
PyPI · Text Processing · released Jan 2026

pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.

Install it if you need to parse text or define grammars programmatically.

MITpure Python · 3.9+
412.7Mdownloads / mo
fonttools Worth it
PyPI · Text Processing · released May 2026

fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.

Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.

permissive licensepure Python · 3.10+
235.9Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
RapidFuzz Worth it
PyPI · Text Processing · released Apr 2026

RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.

Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.

MITcompiled wheel · 3.10+
184.2Mdownloads / mo
tinycss2 Worth it
PyPI · Text Processing · released Nov 2025

tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.

Install it if your project requires CSS tokenization or syntax manipulation.

BSD-3-Clausepure Python · 3.10+
113.2Mdownloads / mo

See also chardet · fold-to-ascii · morphys · latexcodec · confusable-homoglyphs · mbstrdecoder · confusables · webencodings · anyascii · Unidecode