$npx skillfedfor your agent

pymupdf

A high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.

Worth itPyPI LibrariesReleased Aug 2026114.9M downloads / moPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — pymupdf-1.28.2-cp310-abi3-macosx_10_15_x86_64.whl · pymupdf-1.28.2-cp310-abi3-macosx_11_0_arm64.whl · pymupdf-1.28.2-cp310-abi3-manylinux_2_28_aarch64.whl
v1.28.2 · released 2026-08-06 · Python >=3.10

Yes. PyMuPDF is a mature, actively maintained library with no known vulnerabilities, strong community adoption, and zero mandatory runtime dependencies. The AGPL license requires careful review if you are building proprietary software—commercial licensing is available from Artifex. For open-source projects, data extraction pipelines, and AI workflows, it is a solid choice. Install friction is low on supported platforms; source compilation on unsupported platforms requires a C/C++ toolchain.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later.
  • On platforms without pre-built wheels, a C/C++ compiler and build toolchain are needed to compile from source.
  • Installation is straightforward via pip with pre-built wheels for Windows, macOS, and Linux on Python 3.10–3.14; no mandatory runtime dependencies.

License · maintenance · safety

(agpl) — PyMuPDF is dual-licensed under GNU AGPL 3.0 or Artifex Commercial License. The AGPL treatment means open-source projects using it must comply with copyleft obligations; proprietary or closed-source use requires a commercial license from Artifex.

last release 2026-08-06 (8 days) · last repo commit 2026-08-13 · 10,472 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 114,904,797 downloads/mo, #310 on PyPI

Verify before relying

pip install pymupdf

import pymupdf

doc = pymupdf.open("document.pdf")
for page in doc:
    print(page.get_text())
  • Whether OCR via Tesseract integration requires Tesseract to be pre-installed on the system (the description mentions separate installation but does not confirm if it is a hard blocker for the feature)
  • Whether pymupdf.pro (Office document support) is available as a separate package or requires a commercial license key
  • Performance benchmarks ('10–50× speed improvements', '100× or more' for rendering) are claimed in the description but not independently verified
  • Actual monthly download volume and community adoption metrics beyond the tier classification
Same gist for agents: .md · .json

What it is and what it does

PyMuPDF is a Python wrapper around MuPDF, a fast C-based PDF rendering engine. It provides both low-level control and high-level convenience APIs for reading, writing, and transforming PDF and related document formats. The library handles text extraction with font and color metadata, table detection and export, image extraction and rendering, annotations, redaction, form filling, encryption, and document merging—all with no mandatory external dependencies beyond the compiled MuPDF core.

The package is designed for data extraction pipelines, document automation, and AI workflows. It supports input from PDF, XPS, EPUB, CBZ, MOBI, FB2, SVG, TXT, MD, images, and (via the Pro extension) Microsoft Office and Korean Office formats. Output can be rendered to images at any DPI, converted to Markdown or JSON for LLM consumption, or exported as plain text or structured data. Optional companions like pymupdf4llm provide LLM-ready extraction, and pymupdf-fonts extends font support.

Use it for

  • Extract structured text with layout and font metadata from PDFs for document analysis or archival systems.
  • Detect and export tables from PDF reports as Markdown or Pandas DataFrames for data pipeline ingestion.
  • Render PDF pages to high-resolution images for thumbnail generation, preview systems, or downstream image processing.
  • Convert scanned PDFs to searchable text using Tesseract OCR integration for digitization workflows.
  • Prepare PDF or Office documents as LLM-ready Markdown for retrieval-augmented generation (RAG) and AI pipelines.
  • Automate document redaction and annotation workflows for compliance, legal review, or sensitive data handling.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

PyMuPDF is a mature, actively maintained library with no known vulnerabilities, strong community adoption, and zero mandatory runtime dependencies. The AGPL license requires careful review if you are building proprietary software—commercial licensing is available from Artifex. For open-source projects, data extraction pipelines, and AI workflows, it is a solid choice. Install friction is low on supported platforms; source compilation on unsupported platforms requires a C/C++ toolchain.

Install

pymupdf on PyPI

Before you install

Installation is straightforward via pip with pre-built wheels for Windows, macOS, and Linux on Python 3.10–3.14; no mandatory runtime dependencies. Medium friction reflects that source compilation requires a C/C++ toolchain on unsupported platforms. The project is actively maintained with recent releases and strong community engagement.

Requires Python 3.10 or later. On platforms without pre-built wheels, a C/C++ compiler and build toolchain are needed to compile from source.

License in practice

PyMuPDF is dual-licensed under GNU AGPL 3.0 or Artifex Commercial License. The AGPL treatment means open-source projects using it must comply with copyleft obligations; proprietary or closed-source use requires a commercial license from Artifex.

Quickstart

pip install pymupdf

import pymupdf

doc = pymupdf.open("document.pdf")
for page in doc:
    print(page.get_text())

Verify before relying

  • Whether OCR via Tesseract integration requires Tesseract to be pre-installed on the system (the description mentions separate installation but does not confirm if it is a hard blocker for the feature)
  • Whether pymupdf.pro (Office document support) is available as a separate package or requires a commercial license key
  • Performance benchmarks ('10–50× speed improvements', '100× or more' for rendering) are claimed in the description but not independently verified
  • Actual monthly download volume and community adoption metrics beyond the tier classification

Package facts

LicenseNot declared agpl
Python supportSupports the current Python release >=3.10
Install frictionMedium. Platform-specific wheel
Runtime dependenciesNone
MaintenanceActively maintained 8 days since the last release
Last repo commit
First released
Downloads114,904,797 / month, #310 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Information TechnologyOperating System :: MacOSOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: CProgramming Language :: C++Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: Implementation :: CPythonTopic :: Multimedia :: GraphicsTopic :: Software Development :: LibrariesTopic :: Utilities

Evidence: pymupdf-1.28.2-cp310-abi3-macosx_10_15_x86_64.whl; pymupdf-1.28.2-cp310-abi3-macosx_11_0_arm64.whl; pymupdf-1.28.2-cp310-abi3-manylinux_2_28_aarch64.whl; pymupdf-1.28.2-cp310-abi3-manylinux_2_28_x86_64.whl; pymupdf-1.28.2-cp310-abi3-musllinux_1_2_x86_64.whl; pymupdf-1.28.2-cp310-abi3-win32.whl; pymupdf-1.28.2-cp310-abi3-win_amd64.whl; pymupdf-1.28.2-cp310-abi3-win_arm64.whl; pymupdf-1.28.2-cp313-abi3-pyemscripten_2025_0_wasm32.whl; pymupdf-1.28.2-cp314-cp314t-manylinux_2_28_x86_64.whl

Tags

Capabilities
pdf text extraction pythonpdf rendering and conversiondocument data extraction librarypdf table detectionpdf to markdown conversionpdf annotation and redactionocr and document processing
Topics
pdf-processingdocument-extractionllm-ready

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “pdf annotation and redaction”

  • pymupdfPyMuPDF extracts, renders, converts, and manipulates PDF and other…
  • PyMuPDFbPyMuPDFb provides the compiled MuPDF C engine binaries that power…
  • streamlit-pdf-viewerStreamlit component that renders PDF documents in web applications…

Give your agent the search over MCP, or paste the wish link into any chat.

More Libraries packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
python-dateutil Worth it
PyPI · Libraries · released Mar 2024

Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.

Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.

Apache-2.0pure Python
1.2Bdownloads / mo
six With conditions
PyPI · Libraries · released Dec 2024

Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.

MITpure Python
1.2Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo

See also pdf-oxide · pymupdf-fonts · pymupdf-layout · PyMuPDFb · pymupdfpro · pymupdf4llm · unPDF · camelot-py · kreuzberg · pdftext