pymupdf
A high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Decision gist · record as of 2026-08-14
Yes. PyMuPDF is a mature, actively maintained library with no known vulnerabilities, strong community adoption, and zero mandatory runtime dependencies. The AGPL license requires careful review if you are building proprietary software—commercial licensing is available from Artifex. For open-source projects, data extraction pipelines, and AI workflows, it is a solid choice. Install friction is low on supported platforms; source compilation on unsupported platforms requires a C/C++ toolchain.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- On platforms without pre-built wheels, a C/C++ compiler and build toolchain are needed to compile from source.
- Installation is straightforward via pip with pre-built wheels for Windows, macOS, and Linux on Python 3.10–3.14; no mandatory runtime dependencies.
License · maintenance · safety
(agpl) — PyMuPDF is dual-licensed under GNU AGPL 3.0 or Artifex Commercial License. The AGPL treatment means open-source projects using it must comply with copyleft obligations; proprietary or closed-source use requires a commercial license from Artifex.
last release 2026-08-06 (8 days) · last repo commit 2026-08-13 · 10,472 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 114,904,797 downloads/mo, #310 on PyPI
Alternatives
Verify before relying
pip install pymupdf
import pymupdf
doc = pymupdf.open("document.pdf")
for page in doc:
print(page.get_text())- Whether OCR via Tesseract integration requires Tesseract to be pre-installed on the system (the description mentions separate installation but does not confirm if it is a hard blocker for the feature)
- Whether pymupdf.pro (Office document support) is available as a separate package or requires a commercial license key
- Performance benchmarks ('10–50× speed improvements', '100× or more' for rendering) are claimed in the description but not independently verified
- Actual monthly download volume and community adoption metrics beyond the tier classification
What it is and what it does
PyMuPDF is a Python wrapper around MuPDF, a fast C-based PDF rendering engine. It provides both low-level control and high-level convenience APIs for reading, writing, and transforming PDF and related document formats. The library handles text extraction with font and color metadata, table detection and export, image extraction and rendering, annotations, redaction, form filling, encryption, and document merging—all with no mandatory external dependencies beyond the compiled MuPDF core.
The package is designed for data extraction pipelines, document automation, and AI workflows. It supports input from PDF, XPS, EPUB, CBZ, MOBI, FB2, SVG, TXT, MD, images, and (via the Pro extension) Microsoft Office and Korean Office formats. Output can be rendered to images at any DPI, converted to Markdown or JSON for LLM consumption, or exported as plain text or structured data. Optional companions like pymupdf4llm provide LLM-ready extraction, and pymupdf-fonts extends font support.
Use it for
- Extract structured text with layout and font metadata from PDFs for document analysis or archival systems.
- Detect and export tables from PDF reports as Markdown or Pandas DataFrames for data pipeline ingestion.
- Render PDF pages to high-resolution images for thumbnail generation, preview systems, or downstream image processing.
- Convert scanned PDFs to searchable text using Tesseract OCR integration for digitization workflows.
- Prepare PDF or Office documents as LLM-ready Markdown for retrieval-augmented generation (RAG) and AI pipelines.
- Automate document redaction and annotation workflows for compliance, legal review, or sensitive data handling.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
PyMuPDF is a mature, actively maintained library with no known vulnerabilities, strong community adoption, and zero mandatory runtime dependencies. The AGPL license requires careful review if you are building proprietary software—commercial licensing is available from Artifex. For open-source projects, data extraction pipelines, and AI workflows, it is a solid choice. Install friction is low on supported platforms; source compilation on unsupported platforms requires a C/C++ toolchain.
Install
pymupdf on PyPI
Before you install
Installation is straightforward via pip with pre-built wheels for Windows, macOS, and Linux on Python 3.10–3.14; no mandatory runtime dependencies. Medium friction reflects that source compilation requires a C/C++ toolchain on unsupported platforms. The project is actively maintained with recent releases and strong community engagement.
Requires Python 3.10 or later. On platforms without pre-built wheels, a C/C++ compiler and build toolchain are needed to compile from source.
License in practice
PyMuPDF is dual-licensed under GNU AGPL 3.0 or Artifex Commercial License. The AGPL treatment means open-source projects using it must comply with copyleft obligations; proprietary or closed-source use requires a commercial license from Artifex.
Quickstart
pip install pymupdf
import pymupdf
doc = pymupdf.open("document.pdf")
for page in doc:
print(page.get_text())
Verify before relying
- Whether OCR via Tesseract integration requires Tesseract to be pre-installed on the system (the description mentions separate installation but does not confirm if it is a hard blocker for the feature)
- Whether pymupdf.pro (Office document support) is available as a separate package or requires a commercial license key
- Performance benchmarks ('10–50× speed improvements', '100× or more' for rendering) are claimed in the description but not independently verified
- Actual monthly download volume and community adoption metrics beyond the tier classification
Package facts
| License | Not declared agpl |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 8 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 114,904,797 / month, #310 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Information TechnologyOperating System :: MacOSOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: CProgramming Language :: C++Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: Implementation :: CPythonTopic :: Multimedia :: GraphicsTopic :: Software Development :: LibrariesTopic :: Utilities |
Evidence: pymupdf-1.28.2-cp310-abi3-macosx_10_15_x86_64.whl; pymupdf-1.28.2-cp310-abi3-macosx_11_0_arm64.whl; pymupdf-1.28.2-cp310-abi3-manylinux_2_28_aarch64.whl; pymupdf-1.28.2-cp310-abi3-manylinux_2_28_x86_64.whl; pymupdf-1.28.2-cp310-abi3-musllinux_1_2_x86_64.whl; pymupdf-1.28.2-cp310-abi3-win32.whl; pymupdf-1.28.2-cp310-abi3-win_amd64.whl; pymupdf-1.28.2-cp310-abi3-win_arm64.whl; pymupdf-1.28.2-cp313-abi3-pyemscripten_2025_0_wasm32.whl; pymupdf-1.28.2-cp314-cp314t-manylinux_2_28_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “pdf annotation and redaction”
- pymupdfPyMuPDF extracts, renders, converts, and manipulates PDF and other…
- PyMuPDFbPyMuPDFb provides the compiled MuPDF C engine binaries that power…
- streamlit-pdf-viewerStreamlit component that renders PDF documents in web applications…
Give your agent the search over MCP, or paste the wish link into any chat.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also pdf-oxide · pymupdf-fonts · pymupdf-layout · PyMuPDFb · pymupdfpro · pymupdf4llm · unPDF · camelot-py · kreuzberg · pdftext