html-for-docx
Convert HTML to Docx easily and fastly
Decision gist · record as of 2026-08-14
Yes. The package is actively maintained, has low install friction, carries a permissive MIT license, supports Python 3.7 through 3.14, and solves a concrete problem with a straightforward API. No known vulnerabilities. Install it if you need to convert HTML to Word documents; the main limitation is fidelity of complex CSS styling, which is inherent to the Word format itself.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low friction install with two stable runtime dependencies.
- Actively maintained with recent release activity.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; you may use, modify, and distribute freely provided you include the license notice.
last release 2026-07-31 (14 days) · last repo commit 2026-08-01 · 64 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 344,972 downloads/mo, #7,370 on PyPI
Alternatives
Verify before relying
pip install html-for-docx
from html_for_docx import HtmlToDocx
from docx import Document
parser = HtmlToDocx()
document = Document()
parser.add_html_to_document('Hello world', document)
document.save('output.docx')- Extent of HTML/CSS feature coverage beyond what the description excerpt shows (e.g., form elements, media queries, complex layouts).
- Performance characteristics with large HTML documents or complex nested structures.
- Fidelity of style conversion from CSS to Word's native style system in edge cases.
What it is and what it does
html-for-docx is a Python library that transforms HTML markup into Word documents (.docx format). It wraps python-docx and beautifulsoup4 to parse HTML and render it as native Word content, preserving text, tables, images, and styling where possible. The library supports inline CSS, custom style mappings from HTML classes to Word styles, tag-level style overrides, and document metadata manipulation.
You use it by instantiating a parser, then either adding HTML snippets to an existing Word document incrementally, converting entire HTML files, or parsing HTML strings directly to new documents. It handles both file-based and in-memory workflows, offers granular control over which HTML features to process (images, tables, styles, comments), and lets you apply Word template styles to preserve formatting across conversions.
Use it for
- Generate Word reports from HTML templates or web content without manual formatting.
- Batch-convert HTML files to .docx for archival, distribution, or further editing.
- Embed HTML-to-docx conversion in a web application to let users download formatted documents.
- Combine markdown-to-HTML pipelines with this library to produce Word documents from markdown source.
- Populate Word documents with styled HTML content from a CMS or database dynamically.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
The package is actively maintained, has low install friction, carries a permissive MIT license, supports Python 3.7 through 3.14, and solves a concrete problem with a straightforward API. No known vulnerabilities. Install it if you need to convert HTML to Word documents; the main limitation is fidelity of complex CSS styling, which is inherent to the Word format itself.
Install
html-for-docx on PyPI
Before you install
Low friction install with two stable runtime dependencies. Actively maintained with recent release activity.
License in practice
MIT license permits commercial and private use with minimal restrictions; you may use, modify, and distribute freely provided you include the license notice.
Quickstart
pip install html-for-docx
from html_for_docx import HtmlToDocx
from docx import Document
parser = HtmlToDocx()
document = Document()
parser.add_html_to_document('Hello world', document)
document.save('output.docx')
Verify before relying
- Extent of HTML/CSS feature coverage beyond what the description excerpt shows (e.g., form elements, media queries, complex layouts).
- Performance characteristics with large HTML documents or complex nested structures.
- Fidelity of style conversion from CSS to Word's native style system in edge cases.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.7 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagespython-docxbeautifulsoup4 |
| Maintenance | Actively maintained 14 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 344,972 / month, #7,370 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Software Development :: Build ToolsTopic :: Software Development :: LibrariesTopic :: Text ProcessingTopic :: Text Processing :: Markup :: HTMLTopic :: Utilities |
Evidence: html_for_docx-1.1.7-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “html parsing to docx”
- html-for-docxConverts HTML content to Word documents (.docx), with support for…
- htmldocxConverts HTML content into Microsoft Word documents (.docx format),…
- docling-slimDocling Slim is a lightweight, modular SDK for parsing and converting…
Give your agent the search over MCP, or paste the wish link into any chat.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also html-to-json · html2docx · htmldocx · mammoth · docxtpl · docxcompose · spire-doc · docx2python · MarkupPy · python-docx-ml6