mammoth
Convert Word documents from docx to simple and clean HTML and Markdown
Decision gist · record as of 2026-08-14
Yes. Mammoth is actively maintained, has low install friction, carries a permissive license, and solves a well-defined problem—converting .docx to clean HTML using semantic information. It is suitable for most document conversion workflows, but avoid it for untrusted input without additional sanitization and be aware that complex document layouts may not convert perfectly.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Input .docx files must be opened in binary mode; conversion works best when documents use semantic styles rather than direct formatting.
- Low friction: pure Python wheel with a single runtime dependency (cobble).
- Active maintenance with a release 5 days ago and 1114 repository stars.
License · maintenance · safety
BSD-2-Clause (permissive) — BSD-2-Clause permissive license allows commercial and private use with minimal restrictions, though the package performs no sanitization of untrusted input.
last release 2026-08-09 (5 days) · last repo commit 2026-08-09 · 1,114 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 16,076,406 downloads/mo, #1,163 on PyPI
Alternatives
Verify before relying
pip install mammoth
import mammoth
with open("document.docx", "rb") as docx_file:
result = mammoth.convert_to_html(docx_file)
html = result.value- Whether cobble dependency is lightweight and well-maintained (not disclosed in fact sheet).
- Performance characteristics for large documents or batch conversions.
- Fidelity of table formatting preservation beyond text styling.
What it is and what it does
Mammoth is a .docx-to-HTML converter designed to produce clean, semantic markup from Word documents created by Microsoft Word, Google Docs, or LibreOffice. Rather than attempting to replicate visual styling, it interprets document structure—converting a paragraph styled as "Heading 1" to an h1 element, for example—and ignores formatting details like fonts and colors. This approach works well for documents that use styles semantically but may lose fidelity on complex layouts.
The package supports headings, lists, tables, footnotes, images, text formatting (bold, italic, strikethrough, superscript, subscript), links, comments, and custom style mappings. It can be used as a CLI tool or imported as a library. A critical caveat: Mammoth performs no sanitization of source documents, so it must be used carefully with untrusted input. The library also includes text extraction and experimental Markdown output (deprecated in favor of generating HTML first).
Use it for
- Batch convert Word documents to HTML for web publishing or content management systems.
- Extract semantic structure from styled Word documents for downstream processing or archival.
- Build document pipelines that map custom Word styles to specific HTML classes or elements.
- Preserve document content and structure when migrating from Word to web-based authoring platforms.
- Extract raw text from .docx files while ignoring all formatting and styling.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Mammoth is actively maintained, has low install friction, carries a permissive license, and solves a well-defined problem—converting .docx to clean HTML using semantic information. It is suitable for most document conversion workflows, but avoid it for untrusted input without additional sanitization and be aware that complex document layouts may not convert perfectly.
Install
mammoth on PyPI
Before you install
Low friction: pure Python wheel with a single runtime dependency (cobble). Active maintenance with a release 5 days ago and 1114 repository stars. Supports Python 3.7 through 3.12.
Input .docx files must be opened in binary mode; conversion works best when documents use semantic styles rather than direct formatting.
License in practice
BSD-2-Clause permissive license allows commercial and private use with minimal restrictions, though the package performs no sanitization of untrusted input.
Quickstart
pip install mammoth
import mammoth
with open("document.docx", "rb") as docx_file:
result = mammoth.convert_to_html(docx_file)
html = result.value
Verify before relying
- Whether cobble dependency is lightweight and well-maintained (not disclosed in fact sheet).
- Performance characteristics for large documents or batch conversions.
- Fidelity of table formatting preservation beyond text styling.
Package facts
| License | BSD-2-Clause permissive |
| Python support | Supports the current Python release >=3.7 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagecobble |
| Maintenance | Actively maintained 5 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 16,076,406 / month, #1,163 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: BSD LicenseProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9 |
Evidence: mammoth-1.12.1-py2.py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “docx to html converter”
- mammothConverts Microsoft Word .docx documents to clean HTML or Markdown,…
- html-for-docxConverts HTML content to Word documents (.docx), with support for…
- pypandocPypandoc wraps pandoc, a universal document converter, letting you…
Give your agent the search over MCP, or paste the wish link into any chat.
More HTML packages
MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.
Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.
Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.
Install it if you need to parse or extract data from markup documents.
lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.
Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.
Install it if you need to parse Markdown in Python.
See also draftjs_exporter · html-for-docx · textile · htmldocx · html2docx · spire-doc · markdown2 · doc2docx · quill-delta · docxtpl