$npx skillfedfor your agent

html2text

Turn HTML into equivalent Markdown-structured text.

With conditionsPyPI HTMLReleased Apr 202515.2M downloads / moGPL-3.0-or-laterPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — html2text-2025.4.15-py3-none-any.whl
v2025.4.15 · released 2025-04-15 · Python >=3.9

Yes, if you need HTML-to-text or HTML-to-Markdown conversion without external dependencies. The package is stable, widely used (top 5000 on PyPI), and has no known vulnerabilities. The aging maintenance status (last release 486 days ago) is a minor concern for new feature requests but does not block routine use. Avoid if you require active, frequent updates or have strict copyleft license restrictions in your project.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.9 or later.
  • Low install friction with no runtime dependencies.
  • Maintenance status is aging—last release was 486 days ago—but the repository remains active and the package is marked Production/Stable with broad Python version support (3.9–3.13).

License · maintenance · safety

GPL-3.0-or-later (copyleft) — Licensed under GPL-3.0-or-later (copyleft). Any software that incorporates or distributes this package must also be released under a compatible copyleft license; proprietary or closed-source projects cannot use it without legal review.

last release 2025-04-15 (486 days) · last repo commit 2025-10-28 · 2,168 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 15,159,925 downloads/mo, #1,197 on PyPI

Verify before relying

pip install html2text

import html2text
h = html2text.HTML2Text()
h.ignore_links = False
print(h.handle("<p>Hello, <b>world</b>!</p>"))
  • Whether the aging maintenance status (486 days since last release) affects bug fixes or compatibility with newer Python minor versions beyond 3.13.
  • Performance characteristics on large HTML documents or batch conversion workloads.
Same gist for agents: .md · .json

What it is and what it does

html2text is a command-line tool and Python library that transforms HTML into plain ASCII text formatted as valid Markdown. It strips HTML markup while preserving document structure—bold, italic, links, code blocks, and other semantic elements are converted to their Markdown equivalents. The tool offers fine-grained control over output through options like link handling (ignore, reference-style, or inline), special character escaping, and code block marking.

The package has zero runtime dependencies and works as both a standalone script (invoked with `html2text [filename]`) and as an importable Python module. It supports modern Python versions (3.9 through 3.13) on CPython and PyPy, making it suitable for integration into text processing pipelines, documentation generators, or any workflow that needs to extract readable content from HTML.

Use it for

  • Extract readable text from web pages or HTML emails for archival or plain-text distribution.
  • Convert HTML documentation to Markdown for version control and static site generators.
  • Batch-process HTML files into text format for downstream NLP or text analysis tasks.
  • Generate reference-style Markdown from HTML with customizable link formatting.
  • Strip HTML markup while preserving semantic structure for accessibility or readability tools.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need HTML-to-text or HTML-to-Markdown conversion without external dependencies.

The package is stable, widely used (top 5000 on PyPI), and has no known vulnerabilities. The aging maintenance status (last release 486 days ago) is a minor concern for new feature requests but does not block routine use. Avoid if you require active, frequent updates or have strict copyleft license restrictions in your project.

Install

html2text on PyPI

Before you install

Low install friction with no runtime dependencies. Maintenance status is aging—last release was 486 days ago—but the repository remains active and the package is marked Production/Stable with broad Python version support (3.9–3.13).

Requires Python 3.9 or later.

License in practice

Licensed under GPL-3.0-or-later (copyleft). Any software that incorporates or distributes this package must also be released under a compatible copyleft license; proprietary or closed-source projects cannot use it without legal review.

Quickstart

pip install html2text

import html2text
h = html2text.HTML2Text()
h.ignore_links = False
print(h.handle("<p>Hello, <b>world</b>!</p>"))

Verify before relying

  • Whether the aging maintenance status (486 days since last release) affects bug fixes or compatibility with newer Python minor versions beyond 3.13.
  • Performance characteristics on large HTML documents or batch conversion workloads.

Package facts

LicenseGPL-3.0-or-later copyleft
Python supportSupports the current Python release >=3.9
Install frictionLow. Pure-Python wheel
Runtime dependenciesNone
MaintenanceAging 486 days since the last release
Last repo commit
First released
Downloads15,159,925 / month, #1,197 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPy

Evidence: html2text-2025.4.15-py3-none-any.whl

Tags

Capabilities
html to markdown conversionhtml to plain textconvert html documentsextract text from htmlhtml markdown converterstrip html formattinghtml parser text extraction
Topics
html-parsingmarkdown-generationtext-extraction

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “extract text from html”

  • html2textConverts HTML to clean, readable plain text or Markdown-formatted…
  • html-textExtracts plain text from HTML while filtering out styles, scripts,…
  • inscriptisConverts HTML documents to plain text while preserving layout,…

Give your agent the search over MCP, or paste the wish link into any chat.

More HTML packages

MarkupSafe Worth it
PyPI · Dynamic Content · released Sep 2025

MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.

BSD-3-Clausecompiled wheel · 3.9+aging
797.1Mdownloads / mo
Jinja2 Worth it
PyPI · Dynamic Content · released Mar 2025

Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.

BSD-3-Clausepure Python · 3.7+aging
718.6Mdownloads / mo
beautifulsoup4 Worth it
PyPI · Python Modules · released Jun 2026

Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.

Install it if you need to parse or extract data from markup documents.

MITpure Python · 3.7.0+
432.1Mdownloads / mo
lxml Worth it
PyPI · Python Modules · released May 2026

lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.

Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.

permissive licensecompiled wheel · 3.8+
416.8Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
Markdown Worth it
PyPI · Python Modules · released Jul 2026

Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.

Install it if you need to parse Markdown in Python.

BSD-3-Clausepure Python · 3.10+
121.7Mdownloads / mo

See also markdown2 · strip-markdown · mdit-plain · jusText · junit2html · mistune · ansi2txt · telegramify-markdown · readme-renderer