Packages
MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.
Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.
Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.
Install it if you need to parse or extract data from markup documents.
lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.
Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.
Install it if you need to parse Markdown in Python.
Sphinx generates professional documentation from reStructuredText source files, producing HTML, PDF, EPUB, and other formats with automatic cross-references, code highlighting, and hierarchical navigation.
Provides a Pygments syntax highlighting theme that applies JupyterLab's CSS variables to code displayed in HTML, enabling consistent visual styling across JupyterLab and Pygments-generated output.
Install it if visual consistency between Jupyter and static HTML output matters to your workflow.
Safely renders README files in Markdown, reStructuredText, and plain text formats to HTML, designed for use in package repositories like PyPI.
Install it if you need to render README files safely to HTML or validate package metadata formatting—it's the standard tool for this task in the Python packaging…
WeasyPrint converts HTML and CSS into PDF documents using a Python-based rendering engine designed for pagination and document generation.
html5lib parses HTML documents into a tree structure conforming to the WHATWG HTML specification, the standard implemented by all major web browsers.
Parses HTML5 documents—including malformed ones—into an ElementTree structure, providing a lightweight alternative to full-featured HTML parsers.
Pymdown Extensions provides a collection of additional extensions for Python Markdown that add syntax support for advanced formatting, including tabbed content, superfences with syntax highlighting, emoji, and other markup enhancements.
Install it if you need markdown features beyond the standard library.
cssselect parses CSS3 selectors and translates them to XPath 1.0 expressions for use with XPath engines like lxml to query XML and HTML documents.
Install it if you need to translate CSS selectors to XPath for lxml or similar engines.
Cleans and sanitizes HTML by removing unwanted tags and attributes using a blocklist approach, extracted from lxml's original HTML cleaner module.
However, do not use it for security-sensitive applications—the maintainers explicitly recommend alternatives like nh3 for those cases.
Material for MkDocs is a professional theme and extension for MkDocs that transforms Markdown documentation into a responsive, searchable static site with minimal configuration.
Install it if you need to publish Markdown-based documentation as a professional static site without writing custom HTML or CSS.
Converts Microsoft Word .docx documents to clean HTML or Markdown, using document styles to generate semantic markup rather than attempting to replicate visual formatting.
Converts HTML to clean, readable plain text or Markdown-formatted output, with options to customize link handling, escaping, and code block marking.
Extracts original and updated publication dates from web pages by parsing HTML markup, metadata, and text content, with both Python API and command-line interfaces.
Install it if you need reliable date extraction from web pages; the fast mode offers good speed and the extensive mode provides high recall when accuracy matters most.
Provides emoji and SVG icon extensions for MkDocs Material, enabling insertion of Material, FontAwesome, and Octicons into Markdown using emoji syntax.
Converts tabular data into formatted tables for output in multiple text, binary, and application-specific formats including Markdown, HTML, CSV, JSON, Excel, LaTeX, and SQLite.
Install it if you need to generate formatted tables in any of its supported output formats; skip it only if your use case is limited to a single format where a…
Trafilatura extracts main text, metadata, and structured content from web pages and HTML, converting raw HTML into clean, usable data in multiple output formats.
jusText removes boilerplate content (navigation, headers, footers) from HTML pages while preserving text with full sentences, useful for extracting clean content for linguistic analysis and web corpora.
However, maintenance is aging (535 days since last release), so it is best suited for established use cases where the algorithm is known to work well for your content…
Provides type hints for the html5lib HTML parser, enabling static type checkers like mypy and pyright to validate code that uses html5lib.
A Python 3 port of the deprecated stdlib sgmllib module for parsing SGML and HTML documents using an event-driven interface.
Install only if you are maintaining legacy code with a hard dependency on sgmllib3k and cannot refactor to a maintained alternative.
Converts Markdown text to HTML using a fast, complete Python implementation of the Markdown specification.
Install it if you need reliable Markdown rendering without external dependencies.
cssutils parses and builds CSS stylesheets as a DOM structure without rendering—it reads, modifies, and serializes CSS according to W3C specifications.
Python wrapper for wkhtmltopdf that converts HTML to PDF using the Webkit rendering engine, supporting URLs, files, and HTML strings as input.
Extracts encapsulated HTML and plain text content from RTF bodies in .msg email files, reversing the RTF encapsulation that Microsoft Exchange applies.
selectolax is a fast HTML5 parser with CSS selector support, written in Cython and backed by either the Modest or Lexbor C parsing engines.
Provides type hints for beautifulsoup4 code via PEP 561 stubs, enabling static type checkers to validate code that uses beautifulsoup4.
Install only if your project targets beautifulsoup4 4.12; uninstall if you upgrade to 4.13.0 or later, which includes native type annotations.
Parsel extracts data from HTML, JSON, and XML documents using CSS selectors, XPath expressions, JMESPath queries, and regular expressions.
Install it if you need to extract data from HTML, XML, or JSON documents in a Python application.
Converts LaTeX mathematical expressions to MathML format, available as a Python library and command-line tool.
Install it if you need to convert LaTeX math to MathML; the main unknown is the breadth of LaTeX syntax it handles, which you should verify against your specific use…
Converts HTML and CSS to PDF documents using Python, enabling developers with web skills to generate PDF templates without learning specialized PDF libraries.
However, the aging maintenance status (last release 537 days ago) warrants caution in production: verify that its rendering quality and feature set meet your specific…
djlint formats and lints HTML templates with embedded template syntax (Django, Jinja, Twig, Nunjucks, Handlebars, Liquid, Go templates, and others), fixing indentation, tag case, spacing, and catching structural issues that generic HTML tools miss.
Inlines CSS from style and link tags directly into HTML element style attributes, transforming external stylesheets into inline styles suitable for email and embedded HTML contexts.
Provides an async Python API to index and search HTML content, enabling full-text search indexing of static sites and custom content.
Install it if you need to index HTML or custom content for search—particularly valuable for static site generators and documentation platforms where you want search…
Converts terminal output with ANSI color codes to HTML or LaTeX, preserving colors and formatting for display in browsers or documents.
Install it if you need to convert ANSI-colored terminal output to HTML or LaTeX.
Dominate is a Python library for writing HTML documents programmatically using a DOM API, letting you generate HTML pages in pure Python without learning a template language.
Install it if you prefer writing HTML in Python over templates or string manipulation.
pyquery lets you query and manipulate XML and HTML documents using a jQuery-like API, built on lxml for fast parsing and transformation.