$npx skillfedfor your agent

pdf2docx

Open source Python library converting pdf to docx.

With conditionsPyPI Text ProcessingReleased May 20261.2M downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — pdf2docx-0.5.13-py3-none-any.whl
v0.5.13 · released 2026-05-01 · Python >=3.10 · 6 runtime deps: PyMuPDF, python-docx, fonttools, numpy, opencv-python-headless, fire

Yes, with conditions. The library is straightforward to install and suitable for basic PDF-to-Word conversion tasks. However, be aware that active maintenance has ended and the project is now community-supported. If you need robust, actively-maintained PDF handling, the description recommends PyMuPDF as an alternative. No known security vulnerabilities are reported.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python >= 3.10.
  • opencv-python-headless requires a system with image processing libraries available.
  • Low friction installation with a pure-Python wheel.

License · maintenance · safety

MIT (permissive) — MIT license permits free use, modification, and distribution with minimal restrictions, making it suitable for both open-source and commercial projects.

last release 2026-05-01 (105 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,151,092 downloads/mo, #4,297 on PyPI

Verify before relying

pip install pdf2docx

from pdf2docx import Converter

converter = Converter('input.pdf')
converter.convert('output.docx')
converter.close()
  • Conversion quality and accuracy across different PDF types and layouts
  • Performance characteristics on large or complex PDF documents
  • Extent of table extraction accuracy and edge-case handling
Same gist for agents: .md · .json

What it is and what it does

pdf2docx is a Python library that converts PDF documents into editable Word (.docx) files. It uses PyMuPDF for PDF parsing and python-docx for document generation, with additional support for table extraction and layout analysis via opencv-python-headless and numpy. The library provides both programmatic and command-line interfaces for conversion tasks.

The project was originally maintained by Artifex but is now community-driven under the MIT license. It depends on several image processing and document manipulation libraries to reconstruct PDF content as structured Word documents. Users should be aware that active maintenance has transitioned to the community, and the description suggests considering PyMuPDF directly for more comprehensive PDF processing needs.

Use it for

  • Convert scanned or digital PDFs to editable Word documents for further editing
  • Extract tables from PDF reports and import them into Word format
  • Batch-convert multiple PDFs to .docx for document management workflows
  • Automate PDF-to-Word conversion in document processing pipelines

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, with conditions.

The library is straightforward to install and suitable for basic PDF-to-Word conversion tasks. However, be aware that active maintenance has ended and the project is now community-supported. If you need robust, actively-maintained PDF handling, the description recommends PyMuPDF as an alternative. No known security vulnerabilities are reported.

Install

pdf2docx on PyPI

Before you install

Low friction installation with a pure-Python wheel. Maintenance is active, though the project description notes it is no longer actively maintained by Artifex and relies on community contributions.

Requires Python >= 3.10. opencv-python-headless requires a system with image processing libraries available.

License in practice

MIT license permits free use, modification, and distribution with minimal restrictions, making it suitable for both open-source and commercial projects.

Quickstart

pip install pdf2docx

from pdf2docx import Converter

converter = Converter('input.pdf')
converter.convert('output.docx')
converter.close()

Verify before relying

  • Conversion quality and accuracy across different PDF types and layouts
  • Performance characteristics on large or complex PDF documents
  • Extent of table extraction accuracy and edge-case handling

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
6 packages
PyMuPDFpython-docxfonttoolsnumpyopencv-python-headlessfire
MaintenanceActively maintained 105 days since the last release
First released
Downloads1,151,092 / month, #4,297 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: pdf2docx-0.5.13-py3-none-any.whl

Tags

Capabilities
pdf to docx conversionpdf to word pythonextract pdf to documentpdf table extractionbatch pdf conversionpdf layout preservation
Topics
document-conversionpdf-processing
PyPI keywords
pdf-to-wordpdf-to-docx

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “pdf layout preservation”

  • pdf2docxConverts PDF files to Word documents (.docx format), with support for…
  • pdfminerExtracts text and layout information from PDF documents, including…
  • pdfminer.sixExtracts text, images, and layout information from PDF documents by…

Give your agent the search over MCP, or paste the wish link into any chat.

More Text Processing packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyparsing Worth it
PyPI · Text Processing · released Jan 2026

pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.

Install it if you need to parse text or define grammars programmatically.

MITpure Python · 3.9+
412.7Mdownloads / mo
fonttools Worth it
PyPI · Text Processing · released May 2026

fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.

Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.

permissive licensepure Python · 3.10+
235.9Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
RapidFuzz Worth it
PyPI · Text Processing · released Apr 2026

RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.

Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.

MITcompiled wheel · 3.10+
184.2Mdownloads / mo
tinycss2 Worth it
PyPI · Text Processing · released Nov 2025

tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.

Install it if your project requires CSS tokenization or syntax manipulation.

BSD-3-Clausepure Python · 3.10+
113.2Mdownloads / mo

See also html2docx · docx2pdf · pymupdfpro · html-for-docx · docling · docx2txt · spire-doc · aspose-words · doc2docx · htmldocx