html5lib
HTML parser based on the WHATWG HTML specification
Decision gist · record as of 2026-08-14
Yes. html5lib is a mature, actively maintained library with low install friction, permissive licensing, no known vulnerabilities, and broad adoption (top 1000 PyPI packages). Install it when you need standards-compliant HTML parsing that matches browser behavior. Avoid it only if you need a faster, more lenient parser for malformed HTML where spec compliance is not required.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with only two runtime dependencies (six and webencodings).
- The package is actively maintained with a recent commit on 2026-04-21 and has been in production use since its first release on 2012-02-11.
License · maintenance · safety
MIT License (permissive) — MIT License (permissive) allows use in commercial and proprietary projects with minimal restrictions—only requiring attribution and inclusion of the license text.
last release 2020-06-22 (2244 days) · last repo commit 2026-04-21 · 1,224 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 38,714,729 downloads/mo, #710 on PyPI
Alternatives
Verify before relying
import html5lib
# Parse from file
with open("mydocument.html", "rb") as f:
document = html5lib.parse(f)
# Or parse from string
document = html5lib.parse("Hello World!")- Whether lxml integration works reliably on PyPy (description notes segfaults are known to occur)
- Current status of optional dependencies (genshi, chardet, lxml) and their compatibility with modern versions
- Whether the deprecated sanitizer has been removed or if migration to Bleach is required for existing code
What it is and what it does
html5lib is a pure-Python HTML parser designed to conform to the WHATWG HTML specification as implemented by all major web browsers. It parses HTML from files, strings, or file-like objects and builds a tree representation using xml.etree.ElementTree by default, with optional support for lxml.etree and xml.dom.minidom tree formats. The parser handles character encoding detection and can accept explicit encoding hints from HTTP headers or other sources.
The package is useful when you need standards-compliant HTML parsing that matches browser behavior rather than lenient tag-soup parsing. It depends on six and webencodings for Python 2/3 compatibility and encoding detection. Optional dependencies like lxml, genshi, and chardet can extend functionality for alternative tree formats and character encoding fallbacks, though lxml is not recommended on PyPy due to known segfault issues.
Use it for
- Parse HTML documents from web scraping or file I/O and convert them to a queryable tree structure
- Build HTML processing pipelines that need standards-compliant parsing matching browser behavior
- Extract and manipulate HTML content while preserving document structure according to WHATWG rules
- Integrate HTML parsing into web frameworks or content management systems requiring spec compliance
- Serialize parsed HTML back to strings with configurable quote handling and sanitization options
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
html5lib is a mature, actively maintained library with low install friction, permissive licensing, no known vulnerabilities, and broad adoption (top 1000 PyPI packages). Install it when you need standards-compliant HTML parsing that matches browser behavior. Avoid it only if you need a faster, more lenient parser for malformed HTML where spec compliance is not required.
Install
html5lib on PyPI
Before you install
Low install friction with only two runtime dependencies (six and webencodings). The package is actively maintained with a recent commit on 2026-04-21 and has been in production use since its first release on 2012-02-11.
License in practice
MIT License (permissive) allows use in commercial and proprietary projects with minimal restrictions—only requiring attribution and inclusion of the license text.
Quickstart
import html5lib
# Parse from file
with open("mydocument.html", "rb") as f:
document = html5lib.parse(f)
# Or parse from string
document = html5lib.parse("Hello World!")
Verify before relying
- Whether lxml integration works reliably on PyPy (description notes segfaults are known to occur)
- Current status of optional dependencies (genshi, chardet, lxml) and their compatibility with modern versions
- Whether the deprecated sanitizer has been removed or if migration to Bleach is required for existing code
Package facts
| License | MIT License permissive |
| Python support | Supports the current Python release >=2.7, !=3.0.*, !=3.1.*, !=3.2.*, !=3.3.*, !=3.4.* |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagessixwebencodings |
| Maintenance | Actively maintained 2,244 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 38,714,729 / month, #710 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 2Programming Language :: Python :: 2.7Programming Language :: Python :: 3Programming Language :: Python :: 3.5Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPyTopic :: Software Development :: Libraries :: Python ModulesTopic :: Text Processing :: Markup :: HTML |
Evidence: html5lib-1.1-py2.py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “whatwg html specification”
- html5libhtml5lib parses HTML documents into a tree structure conforming to…
- html5rdfParses HTML documents into tree structures compatible with RDFLib,…
- webencodingsImplements the WHATWG Encoding standard to map legacy web character…
Give your agent the search over MCP, or paste the wish link into any chat.
More Python Modules packages
Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.
Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.
Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.
PyYAML parses and emits YAML 1.1 data format, enabling serialization and deserialization of configuration files and Python objects to and from human-readable YAML text.
Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.
Provides reusable metadata objects for use with PEP-593 `typing.Annotated` to express common constraints like bounds, collection sizes, and predicates on types.
Install it if you use or build libraries that need to express type constraints in a standardized, inspectable way—or if you want to annotate your own types with…
Provides runtime tools to inspect and introspect Python type annotations, enabling programmatic examination of type hints at execution time.
See also html5rdf · tinyhtml5 · webencodings · turbohtml · types-html5lib · can-ada · pyquery · cssselect2 · lxml · bleach