lxml-html-clean
HTML cleaner from lxml project
Decision gist · record as of 2026-08-14
Yes, if you need basic HTML sanitization in a non-security-critical context. The package is actively maintained, has low install friction, and carries a permissive license. However, do not use it for security-sensitive applications—the maintainers explicitly recommend alternatives like nh3 for those cases. Verify that blocklist-based filtering meets your specific requirements.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires lxml, which may need compilation on some systems.
- Not recommended for security-sensitive applications—see package documentation for alternatives.
- Low friction: pure Python wheel with a single runtime dependency on lxml.
License · maintenance · safety
BSD-3-Clause (permissive) — BSD-3-Clause permissive license allows commercial and private use with minimal restrictions; attribution required.
last release 2026-05-20 (86 days) · last repo commit 2026-05-20 · 13 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 20,471,755 downloads/mo, #1,036 on PyPI
Alternatives
Verify before relying
pip install lxml_html_clean
from lxml_html_clean import Cleaner
cleaner = Cleaner()
clean_html = cleaner.clean_html('<p>Hello <script>alert(1)</script></p>')- Whether the blocklist approach provides adequate protection for your specific use case, given the package's own warning against security-sensitive environments
- Whether URL parsing via urllib.parse's non-validating functions poses a risk for your allowed-hosts configuration
What it is and what it does
lxml_html_clean is a standalone HTML sanitization library extracted from lxml's original cleaner module. It uses a blocklist-based approach to remove unwanted HTML tags and attributes from user-supplied or untrusted content. The package depends only on lxml and installs as a pure Python wheel.
The package explicitly warns that it is not suitable for security-sensitive environments. Its URL parsing relies on Python's urllib.parse, which does not validate inputs, and a maliciously crafted URL could bypass the allowed-hosts check. If you need robust HTML sanitization for high-security contexts, the maintainers recommend alternatives like nh3.
Use it for
- Sanitize user-submitted HTML in a blog or comment system where moderate filtering suffices
- Remove script tags and event handlers from HTML copied from untrusted web sources
- Clean up HTML content for display in a non-security-critical web application
- Strip formatting and styling from HTML while preserving structure for content migration
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need basic HTML sanitization in a non-security-critical context.
The package is actively maintained, has low install friction, and carries a permissive license. However, do not use it for security-sensitive applications—the maintainers explicitly recommend alternatives like nh3 for those cases. Verify that blocklist-based filtering meets your specific requirements.
Install
lxml-html-clean on PyPI
Before you install
Low friction: pure Python wheel with a single runtime dependency on lxml. Actively maintained as of 2026-05-20 with recent releases.
Requires lxml, which may need compilation on some systems. Not recommended for security-sensitive applications—see package documentation for alternatives.
License in practice
BSD-3-Clause permissive license allows commercial and private use with minimal restrictions; attribution required.
Quickstart
pip install lxml_html_clean
from lxml_html_clean import Cleaner
cleaner = Cleaner()
clean_html = cleaner.clean_html('<p>Hello <script>alert(1)</script></p>')
Verify before relying
- Whether the blocklist approach provides adequate protection for your specific use case, given the package's own warning against security-sensitive environments
- Whether URL parsing via urllib.parse's non-validating functions poses a risk for your allowed-hosts configuration
Package facts
| License | BSD-3-Clause permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagelxml |
| Maintenance | Actively maintained 86 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 20,471,755 / month, #1,036 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Programming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9 |
Evidence: lxml_html_clean-0.4.5-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “html sanitization blocklist”
- lxml-html-cleanCleans and sanitizes HTML by removing unwanted tags and attributes…
- bleach-allowlistProvides curated allowlists of HTML tags, attributes, and CSS styles…
- nh3nh3 sanitizes HTML by removing unsafe tags and attributes, exposing…
Give your agent the search over MCP, or paste the wish link into any chat.
More HTML packages
MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.
Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.
Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.
Install it if you need to parse or extract data from markup documents.
lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.
Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.
Install it if you need to parse Markdown in Python.
See also django-bleach · nh3 · bleach · bleach-allowlist · py_svg_hush · html-sanitizer · disposable-email-domains · git-filter-repo · htmlmin2