html-to-markdown
High-performance HTML to Markdown converter
Decision gist · record as of 2026-08-14
Yes. The package is actively maintained, has no known vulnerabilities, supports current Python versions (3.10–3.14), and solves a real problem—converting messy HTML to clean Markdown without manual intervention. Medium install friction is acceptable for a compiled extension with pre-built wheels. Use it if you need robust HTML-to-Markdown conversion in production; skip it only if you need Python <3.10 or have a simpler, pure-Python alternative already in place.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later; compiled extension wheels are platform-specific but pre-built for x86_64, aarch64, macOS, and Windows.
- Medium install friction due to compiled wheels (cp310-abi3 binaries for multiple platforms), but pre-built for common architectures (x86_64, aarch64, macOS, Windows).
- Active maintenance with a release 9 days old and 845 repository stars.
License · maintenance · safety
MIT (permissive) — MIT license permits unrestricted use, modification, and distribution in commercial and private projects with minimal attribution requirements.
last release 2026-08-05 (9 days) · last repo commit 2026-08-14 · 845 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 829,991 downloads/mo, #4,952 on PyPI
Alternatives
Verify before relying
pip install html-to-markdown
from html_to_markdown import convert
result = convert('<p>Hello <b>world</b></p>')
print(result['content']) # Hello **world**- Whether the visitor API and metadata extraction features are exposed in the Python binding or require direct Rust access.
- Performance characteristics (19–116 MB/s cited for corpus) on typical document sizes in production Python workflows.
- Whether Djot output format is fully supported in the Python package or limited to CommonMark.
What it is and what it does
html-to-markdown is a Python binding to a Rust-based HTML-to-Markdown converter designed to handle messy, real-world HTML without requiring manual parsing strategy selection or tuning. It accepts unclosed tags, malformed entities, CDATA, custom elements, nested tables, and mixed encodings, then returns clean Markdown output in a single `convert()` call. The package exposes a structured result containing the converted content, warnings, and optionally extracted metadata (Open Graph, Twitter, JSON-LD, microdata).
The converter uses a tiered dispatch strategy (byte scanner → DOM walker → html5ever repair) to achieve byte-equal output across different parsing paths, ensuring consistent results regardless of input complexity. It supports both CommonMark and Djot output formats, handles GitHub-flavored Markdown tables with alignment and cell padding, and can optionally mirror inline images. The implementation is fast enough for whole-corpus jobs and requires Python 3.10 or later; it ships as a compiled extension with pre-built wheels for major platforms.
Use it for
- Extract clean Markdown from web scraped HTML without manual cleanup or tag-balancing logic.
- Convert email HTML bodies or rich-text editor output to Markdown for storage or downstream processing.
- Parse and extract structured metadata (Open Graph, JSON-LD) from web pages in the same pass as content conversion.
- Build content pipelines that accept arbitrary HTML and produce consistent, lossless Markdown without tuning per-source.
- Generate Markdown documentation from HTML-based CMS exports or legacy web content archives.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
The package is actively maintained, has no known vulnerabilities, supports current Python versions (3.10–3.14), and solves a real problem—converting messy HTML to clean Markdown without manual intervention. Medium install friction is acceptable for a compiled extension with pre-built wheels. Use it if you need robust HTML-to-Markdown conversion in production; skip it only if you need Python <3.10 or have a simpler, pure-Python alternative already in place.
Install
html-to-markdown on PyPI
Before you install
Medium install friction due to compiled wheels (cp310-abi3 binaries for multiple platforms), but pre-built for common architectures (x86_64, aarch64, macOS, Windows). Active maintenance with a release 9 days old and 845 repository stars.
Requires Python 3.10 or later; compiled extension wheels are platform-specific but pre-built for x86_64, aarch64, macOS, and Windows.
License in practice
MIT license permits unrestricted use, modification, and distribution in commercial and private projects with minimal attribution requirements.
Quickstart
pip install html-to-markdown
from html_to_markdown import convert
result = convert('<p>Hello <b>world</b></p>')
print(result['content']) # Hello **world**
Verify before relying
- Whether the visitor API and metadata extraction features are exposed in the Python binding or require direct Rust access.
- Performance characteristics (19–116 MB/s cited for corpus) on typical document sizes in production Python workflows.
- Whether Djot output format is fully supported in the Python package or limited to CommonMark.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 9 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 829,991 / month, #4,952 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14 |
Evidence: html_to_markdown-3.10.6-cp310-abi3-macosx_10_12_x86_64.whl; html_to_markdown-3.10.6-cp310-abi3-macosx_11_0_arm64.whl; html_to_markdown-3.10.6-cp310-abi3-manylinux2014_aarch64.manylinux_2_17_aarch64.whl; html_to_markdown-3.10.6-cp310-abi3-manylinux2014_x86_64.manylinux_2_17_x86_64.whl; html_to_markdown-3.10.6-cp310-abi3-win_amd64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “html to markdown converter”
- html-to-markdownConverts real-world HTML—including malformed tags, broken entities,…
- pyromarkpyromark is a CommonMark-compliant Markdown parser that converts…
- gh-md-to-htmlConverts GitHub-flavored Markdown to HTML with support for offline…
Give your agent the search over MCP, or paste the wish link into any chat.
More HTML packages
MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.
Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.
Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.
Install it if you need to parse or extract data from markup documents.
lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.
Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.
Install it if you need to parse Markdown in Python.
See also jira2markdown · tree-sitter-markdown · mistune · markdown-it-pyrs · comrak · multimark · commonmark · markdown2 · markdownify · docstring-to-markdown