$npx skillfedfor your agent

docx2python

Extract content from docx files

Worth itPyPI Text ProcessingReleased Aug 2026335.0K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — docx2python-3.7.1-py3-none-any.whl
v3.7.1 · released 2026-08-06 · Python >=3.10 · 3 runtime deps: lxml, paragraphs, typing-extensions

Yes. The package is actively maintained, has low install friction, carries no known vulnerabilities, uses a permissive MIT license, and solves a common problem (reading .docx files) with a well-structured API. It is suitable for production use in both open-source and commercial contexts. Install if you need to programmatically extract content from Word documents.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later; the .docx file must be a valid ZIP-based Office Open XML document.
  • Low install friction with a pure-Python wheel.
  • Actively maintained with a release 8 days ago.

License · maintenance · safety

MIT (permissive) — MIT license permits use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.

last release 2026-08-06 (8 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 334,965 downloads/mo, #7,483 on PyPI

Verify before relying

from docx2python import docx2python

with docx2python('path/to/file.docx') as docx_content:
    print(docx_content.text)
    print(docx_content.properties)
    print(docx_content.images)
  • Whether the package handles corrupted or malformed .docx files gracefully
  • Performance characteristics on very large documents or batch processing scenarios
  • Extent of support for Word's advanced formatting features beyond those listed
Same gist for agents: .md · .json

What it is and what it does

docx2python reads Microsoft Word .docx files (which are ZIP archives containing XML) and exposes their content as Python objects. It extracts text, images, headers, footers, footnotes, endnotes, document properties, comments, and paragraph metadata—including styles (e.g., Heading 1, Subtitle), formatting runs (bold, italic, underline, color, size), and position within nested lists. Tables are normalized to n×m grids and can be identified without guessing. The package optionally converts formatting to HTML tags and can write extracted images to disk.

The library uses lxml to parse the underlying XML and exposes a DocxContent object with separate attributes for header, footer, body, footnotes, endnotes, and document-level properties. Paragraphs are flattened to a consistent depth and enriched with metadata (style, lineage, list position, runs with formatting). It supports both strict and superset Open Office XML namespaces and works with Python 3.10+.

Use it for

  • Extract text and structure from Word documents for indexing, search, or content migration to other formats
  • Automate document processing pipelines that need to read .docx files and convert them to markdown, HTML, or database records
  • Parse Word documents to identify and extract headings, lists, and formatted text for document analysis or summarization
  • Batch extract images embedded in .docx files and save them to a directory for asset management
  • Build tools that read Word document properties (creator, modification date, etc.) for metadata extraction or audit trails

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

The package is actively maintained, has low install friction, carries no known vulnerabilities, uses a permissive MIT license, and solves a common problem (reading .docx files) with a well-structured API. It is suitable for production use in both open-source and commercial contexts. Install if you need to programmatically extract content from Word documents.

Install

docx2python on PyPI

Before you install

Low install friction with a pure-Python wheel. Actively maintained with a release 8 days ago. Requires Python 3.10 or later and three runtime dependencies (lxml, paragraphs, typing-extensions), all common and lightweight.

Requires Python 3.10 or later; the .docx file must be a valid ZIP-based Office Open XML document.

License in practice

MIT license permits use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.

Quickstart

from docx2python import docx2python

with docx2python('path/to/file.docx') as docx_content:
    print(docx_content.text)
    print(docx_content.properties)
    print(docx_content.images)

Verify before relying

  • Whether the package handles corrupted or malformed .docx files gracefully
  • Performance characteristics on very large documents or batch processing scenarios
  • Extent of support for Word's advanced formatting features beyond those listed

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
3 packages
lxmlparagraphstyping-extensions
MaintenanceActively maintained 8 days since the last release
First released
Downloads334,965 / month, #7,483 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Typing :: Typed

Evidence: docx2python-3.7.1-py3-none-any.whl

Tags

Capabilities
extract text from docxparse word documentsdocx to pythonread docx filesextract docx contentword document parserdocx image extraction
Topics
document-parsingoffice-formats

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “parse word documents”

  • docx2pythonExtracts text, images, headers, footers, footnotes, endnotes,…
  • docxCreates, reads, and writes Microsoft Office Word 2007 docx files in…
  • python-docx-ml6Reads, creates, and updates Microsoft Word 2007+ (.docx) files from…

Give your agent the search over MCP, or paste the wish link into any chat.

More Text Processing packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyparsing Worth it
PyPI · Text Processing · released Jan 2026

pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.

Install it if you need to parse text or define grammars programmatically.

MITpure Python · 3.9+
412.7Mdownloads / mo
fonttools Worth it
PyPI · Text Processing · released May 2026

fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.

Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.

permissive licensepure Python · 3.10+
235.9Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
RapidFuzz Worth it
PyPI · Text Processing · released Apr 2026

RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.

Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.

MITcompiled wheel · 3.10+
184.2Mdownloads / mo
tinycss2 Worth it
PyPI · Text Processing · released Nov 2025

tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.

Install it if your project requires CSS tokenization or syntax manipulation.

BSD-3-Clausepure Python · 3.10+
113.2Mdownloads / mo

See also docx2txt · office-word-mcp-server · docx · html-for-docx · python-docx · html2docx · htmldocx · docx-mailmerge2 · docxtpl · wordcloud