$npx skillfedfor your agent

clean-text

Functions to preprocess and normalize text.

With conditionsPyPI LinguisticReleased Jan 2026317.6K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — clean_text-0.7.1-py3-none-any.whl
v0.7.1 · released 2026-01-28 · Python >=3.9 · 2 runtime deps: emoji, ftfy

Yes, if you need to clean messy user-generated text before NLP or indexing. The package is stable, permissively licensed, and has low install friction. The aging maintenance status (198 days since last release) is a minor concern but not a blocker for a mature text-cleaning utility. Consider it a solid choice for preprocessing pipelines, especially if you already use scikit-learn.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.9 or later.
  • Optional unidecode dependency (installed via clean-text[gpl]) provides better transliteration but carries GPL licensing; without it, falls back to Python's built-in unicodedata with slightly different output.
  • Low friction install with two lightweight runtime dependencies (emoji and ftfy).

License · maintenance · safety

Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing commercial and private use with minimal restrictions. No GPL obligation unless you explicitly install the optional unidecode extra.

last release 2026-01-28 (198 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 317,577 downloads/mo, #7,659 on PyPI

Verify before relying

pip install clean-text

from cleantext import clean

result = clean("Yóù àré right!", fix_unicode=True, to_ascii=True, lower=True)
  • Whether the aging maintenance status (198 days since last release) affects bug fixes or feature requests going forward.
  • Real-world performance of parallel processing (n_jobs parameter) on typical workloads.
  • Consistency and quality differences between transliteration with and without unidecode in production use.
Same gist for agents: .md · .json

What it is and what it does

clean-text is a text preprocessing library that normalizes messy user-generated content from the web and social media. It fixes unicode encoding errors, removes or replaces unwanted patterns (URLs, emails, phone numbers, IP addresses, code snippets), handles transliteration to ASCII, and supports language-specific rules for English and German. The package wraps ftfy for unicode repair and optionally uses unidecode for transliteration, falling back to Python's built-in unicodedata when unidecode is unavailable.

The library offers both a simple functional API (clean() for single strings, clean_texts() for batch processing with optional multiprocessing) and a scikit-learn compatible transformer for integration into ML pipelines. You can preserve specific text patterns using regex-based exceptions to prevent them from being modified during cleaning, and customize replacement tokens for different content types.

Use it for

  • Normalize scraped web or social media text before feeding it into NLP models or search indexing.
  • Batch-clean large text datasets in parallel using n_jobs to speed up preprocessing on multi-core systems.
  • Standardize user input in web applications by removing malformed unicode, URLs, and email addresses while preserving legitimate content.
  • Prepare text for machine learning pipelines via the scikit-learn CleanTransformer for reproducible preprocessing steps.
  • Transliterate accented or non-ASCII characters to ASCII equivalents for systems that require ASCII-only text.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need to clean messy user-generated text before NLP or indexing.

The package is stable, permissively licensed, and has low install friction. The aging maintenance status (198 days since last release) is a minor concern but not a blocker for a mature text-cleaning utility. Consider it a solid choice for preprocessing pipelines, especially if you already use scikit-learn.

Install

clean-text on PyPI

Before you install

Low friction install with two lightweight runtime dependencies (emoji and ftfy). Last release was 198 days ago; maintenance status is aging but the package remains functional for its core use case.

Requires Python 3.9 or later. Optional unidecode dependency (installed via clean-text[gpl]) provides better transliteration but carries GPL licensing; without it, falls back to Python's built-in unicodedata with slightly different output.

License in practice

Licensed under Apache-2.0 (permissive), allowing commercial and private use with minimal restrictions. No GPL obligation unless you explicitly install the optional unidecode extra.

Quickstart

pip install clean-text

from cleantext import clean

result = clean("Yóù àré right!", fix_unicode=True, to_ascii=True, lower=True)

Verify before relying

  • Whether the aging maintenance status (198 days since last release) affects bug fixes or feature requests going forward.
  • Real-world performance of parallel processing (n_jobs parameter) on typical workloads.
  • Consistency and quality differences between transliteration with and without unidecode in production use.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.9
Install frictionLow. Pure-Python wheel
Runtime dependencies
2 packages
emojiftfy
MaintenanceAging 198 days since the last release
First released
Downloads317,577 / month, #7,659 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
License :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.9

Evidence: clean_text-0.7.1-py3-none-any.whl

Tags

Capabilities
text normalization preprocessingclean user generated contentunicode text fixingremove urls emails from texttext sanitization pipelinetransliterate unicode to asciibatch text cleaning
Topics
text-preprocessingnlp-utilitiesdata-cleaning
PyPI keywords
natural-language-processingtext-cleaningtext-preprocessingtext-normalizationuser-generated-content

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “clean user generated content”

  • clean-textPreprocesses and normalizes user-generated text by fixing unicode…
  • nh3nh3 sanitizes HTML by removing unsafe tags and attributes, exposing…
  • langchain-exaIntegrates Exa's web search API with LangChain, enabling AI…

Give your agent the search over MCP, or paste the wish link into any chat.

More Linguistic packages

charset-normalizer Worth it
PyPI · Utilities · released Aug 2026

Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.

permissive licensepure Python · 3.7+
1.7Bdownloads / mo
tiktoken Worth it
PyPI · Linguistic · released May 2026

tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.

Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.

permissive licensecompiled wheel · 3.9+
233.0Mdownloads / mo
chardet Worth it
PyPI · Python Modules · released Aug 2026

Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.

Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.

0BSDpure Python · 3.10+
199.0Mdownloads / mo
text-unidecode With conditions
PyPI · Python Modules · released Aug 2019

Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.

However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.

GPL-2.0-or-laterpure Pythonabandoned
89.0Mdownloads / mo
lark Worth it
PyPI · Python Modules · released Oct 2025

Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.

MITpure Python · 3.8+
79.7Mdownloads / mo
tree-sitter Worth it
PyPI · Linguistic · released Jun 2026

Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.

MITcompiled wheel · 3.10+
79.0Mdownloads / mo

See also python-slugify · Unidecode · proces · unicode-slugify · anyascii · normality · bnunicodenormalizer · fold-to-ascii · rigour