$npx skillfedfor your agent

docx2txt

A pure python-based utility to extract text and images from docx files.

With conditionsPyPI Text ProcessingReleased Mar 20258.6M downloads / moPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — docx2txt-0.9-py3-none-any.whl
v0.9 · released 2025-03-24

Yes, if you need straightforward .docx text and image extraction and can verify the license terms. The package is mature, has no dependencies, and handles the core task well. The aging maintenance status and unclear license are minor concerns—check the repository for the actual license before using in production, and test with your specific .docx variants to confirm compatibility.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Installation is straightforward with no runtime dependencies.
  • The package is aging (last release 508 days ago) but the repository remains active and unarchived, suggesting maintenance is infrequent rather than abandoned.

License · maintenance · safety

(unclear) — License status is unclear—no SPDX identifier or raw license text is available in the package metadata. You should verify the actual license terms in the repository before relying on this package in a commercial or copyleft-sensitive context.

last release 2025-03-24 (508 days) · last repo commit 2025-03-24 · 585 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 8,581,646 downloads/mo, #1,604 on PyPI

Verify before relying

pip install docx2txt

import docx2txt

# Extract text from a .docx file
text = docx2txt.process("file.docx")

# Extract text and save images to a directory
text = docx2txt.process("file.docx", "/tmp/img_dir")
  • Whether the package handles all .docx file variants and edge cases reliably.
  • Exact license terms and any attribution requirements from the adapted code.
  • Python version compatibility (requires_python is unspecified in metadata).
  • Performance characteristics with large or complex .docx documents.
Same gist for agents: .md · .json

What it is and what it does

docx2txt is a pure-Python utility that reads Microsoft Word .docx files and extracts their text content, along with headers, footers, hyperlinks, and embedded images. It provides both a command-line tool and a Python API, making it usable in scripts or as part of a larger application. The package has no runtime dependencies, so installation is lightweight and friction-free.

The code is adapted from existing docx tooling but focuses specifically on text and image extraction rather than document manipulation. It's positioned as a simpler alternative when you only need to pull content out of .docx files rather than create or modify them. With significant real-world use, though maintenance is infrequent (last update 508 days ago).

Use it for

  • Batch-process Word documents to extract text for indexing, search, or archival systems.
  • Automate extraction of images embedded in .docx files for asset management or document scanning workflows.
  • Build a document ingestion pipeline that converts .docx content into plain text for NLP or analysis tasks.
  • Extract header and footer content from formal documents for metadata or compliance auditing.
  • Command-line tool for one-off conversion of .docx files to text without opening Word.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need straightforward .docx text and image extraction and can verify the license terms.

The package is mature, has no dependencies, and handles the core task well. The aging maintenance status and unclear license are minor concerns—check the repository for the actual license before using in production, and test with your specific .docx variants to confirm compatibility.

Install

docx2txt on PyPI

Before you install

Installation is straightforward with no runtime dependencies. The package is aging (last release 508 days ago) but the repository remains active and unarchived, suggesting maintenance is infrequent rather than abandoned.

License in practice

License status is unclear—no SPDX identifier or raw license text is available in the package metadata. You should verify the actual license terms in the repository before relying on this package in a commercial or copyleft-sensitive context.

Quickstart

pip install docx2txt

import docx2txt

# Extract text from a .docx file
text = docx2txt.process("file.docx")

# Extract text and save images to a directory
text = docx2txt.process("file.docx", "/tmp/img_dir")

Verify before relying

  • Whether the package handles all .docx file variants and edge cases reliably.
  • Exact license terms and any attribution requirements from the adapted code.
  • Python version compatibility (requires_python is unspecified in metadata).
  • Performance characteristics with large or complex .docx documents.

Package facts

LicenseNot declared unclear
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependenciesNone
MaintenanceAging 508 days since the last release
Last repo commit
First released
Downloads8,581,646 / month, #1,604 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: docx2txt-0.9-py3-none-any.whl

Tags

Capabilities
extract text from docxdocx to text converterword document text extractiondocx image extractionpython docx parserextract images from word filesdocx header footer extraction
Topics
document-extractionoffice-formatscli-tool
PyPI keywords
pythondocxtextimagesextract

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “extract text from docx”

  • docx2txtExtracts text, headers, footers, hyperlinks, and images from…
  • docx2pythonExtracts text, images, headers, footers, footnotes, endnotes,…
  • docxCreates, reads, and writes Microsoft Office Word 2007 docx files in…

Give your agent the search over MCP, or paste the wish link into any chat.

More Text Processing packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyparsing Worth it
PyPI · Text Processing · released Jan 2026

pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.

Install it if you need to parse text or define grammars programmatically.

MITpure Python · 3.9+
412.7Mdownloads / mo
fonttools Worth it
PyPI · Text Processing · released May 2026

fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.

Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.

permissive licensepure Python · 3.10+
235.9Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
RapidFuzz Worth it
PyPI · Text Processing · released Apr 2026

RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.

Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.

MITcompiled wheel · 3.10+
184.2Mdownloads / mo
tinycss2 Worth it
PyPI · Text Processing · released Nov 2025

tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.

Install it if your project requires CSS tokenization or syntax manipulation.

BSD-3-Clausepure Python · 3.10+
113.2Mdownloads / mo

See also docx2python · htmldocx · docx · spire-doc · pdf2docx · textract · html-for-docx · docxtpl · html2docx · python-docx