$npx skillfedfor your agent

lxml-html-clean

HTML cleaner from lxml project

With conditionsPyPI HTMLReleased May 202620.5M downloads / moBSD-3-ClausePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — lxml_html_clean-0.4.5-py3-none-any.whl
v0.4.5 · released 2026-05-20 · 1 runtime deps: lxml

Yes, if you need basic HTML sanitization in a non-security-critical context. The package is actively maintained, has low install friction, and carries a permissive license. However, do not use it for security-sensitive applications—the maintainers explicitly recommend alternatives like nh3 for those cases. Verify that blocklist-based filtering meets your specific requirements.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires lxml, which may need compilation on some systems.
  • Not recommended for security-sensitive applications—see package documentation for alternatives.
  • Low friction: pure Python wheel with a single runtime dependency on lxml.

License · maintenance · safety

BSD-3-Clause (permissive) — BSD-3-Clause permissive license allows commercial and private use with minimal restrictions; attribution required.

last release 2026-05-20 (86 days) · last repo commit 2026-05-20 · 13 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 20,471,755 downloads/mo, #1,036 on PyPI

Verify before relying

pip install lxml_html_clean

from lxml_html_clean import Cleaner

cleaner = Cleaner()
clean_html = cleaner.clean_html('<p>Hello <script>alert(1)</script></p>')
  • Whether the blocklist approach provides adequate protection for your specific use case, given the package's own warning against security-sensitive environments
  • Whether URL parsing via urllib.parse's non-validating functions poses a risk for your allowed-hosts configuration
Same gist for agents: .md · .json

What it is and what it does

lxml_html_clean is a standalone HTML sanitization library extracted from lxml's original cleaner module. It uses a blocklist-based approach to remove unwanted HTML tags and attributes from user-supplied or untrusted content. The package depends only on lxml and installs as a pure Python wheel.

The package explicitly warns that it is not suitable for security-sensitive environments. Its URL parsing relies on Python's urllib.parse, which does not validate inputs, and a maliciously crafted URL could bypass the allowed-hosts check. If you need robust HTML sanitization for high-security contexts, the maintainers recommend alternatives like nh3.

Use it for

  • Sanitize user-submitted HTML in a blog or comment system where moderate filtering suffices
  • Remove script tags and event handlers from HTML copied from untrusted web sources
  • Clean up HTML content for display in a non-security-critical web application
  • Strip formatting and styling from HTML while preserving structure for content migration

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need basic HTML sanitization in a non-security-critical context.

The package is actively maintained, has low install friction, and carries a permissive license. However, do not use it for security-sensitive applications—the maintainers explicitly recommend alternatives like nh3 for those cases. Verify that blocklist-based filtering meets your specific requirements.

Install

lxml-html-clean on PyPI

Before you install

Low friction: pure Python wheel with a single runtime dependency on lxml. Actively maintained as of 2026-05-20 with recent releases.

Requires lxml, which may need compilation on some systems. Not recommended for security-sensitive applications—see package documentation for alternatives.

License in practice

BSD-3-Clause permissive license allows commercial and private use with minimal restrictions; attribution required.

Quickstart

pip install lxml_html_clean

from lxml_html_clean import Cleaner

cleaner = Cleaner()
clean_html = cleaner.clean_html('<p>Hello <script>alert(1)</script></p>')

Verify before relying

  • Whether the blocklist approach provides adequate protection for your specific use case, given the package's own warning against security-sensitive environments
  • Whether URL parsing via urllib.parse's non-validating functions poses a risk for your allowed-hosts configuration

Package facts

LicenseBSD-3-Clause permissive
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependencies
1 package
lxml
MaintenanceActively maintained 86 days since the last release
Last repo commit
First released
Downloads20,471,755 / month, #1,036 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Programming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9

Evidence: lxml_html_clean-0.4.5-py3-none-any.whl

Tags

Capabilities
html sanitization blocklistremove html tags safelyhtml cleaner utilitystrip dangerous html elementshtml content filtering
Topics
html-sanitizationblocklist-based

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “html sanitization blocklist”

  • lxml-html-cleanCleans and sanitizes HTML by removing unwanted tags and attributes…
  • bleach-allowlistProvides curated allowlists of HTML tags, attributes, and CSS styles…
  • nh3nh3 sanitizes HTML by removing unsafe tags and attributes, exposing…

Give your agent the search over MCP, or paste the wish link into any chat.

More HTML packages

MarkupSafe Worth it
PyPI · Dynamic Content · released Sep 2025

MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.

BSD-3-Clausecompiled wheel · 3.9+aging
797.1Mdownloads / mo
Jinja2 Worth it
PyPI · Dynamic Content · released Mar 2025

Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.

BSD-3-Clausepure Python · 3.7+aging
718.6Mdownloads / mo
beautifulsoup4 Worth it
PyPI · Python Modules · released Jun 2026

Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.

Install it if you need to parse or extract data from markup documents.

MITpure Python · 3.7.0+
432.1Mdownloads / mo
lxml Worth it
PyPI · Python Modules · released May 2026

lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.

Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.

permissive licensecompiled wheel · 3.8+
416.8Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
Markdown Worth it
PyPI · Python Modules · released Jul 2026

Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.

Install it if you need to parse Markdown in Python.

BSD-3-Clausepure Python · 3.10+
121.7Mdownloads / mo

See also django-bleach · nh3 · bleach · bleach-allowlist · py_svg_hush · html-sanitizer · disposable-email-domains · git-filter-repo · htmlmin2