pymupdf4llm
PyMuPDF Utilities for LLM/RAG
Decision gist · record as of 2026-08-14
Yes. Active maintenance, no known vulnerabilities, low install friction, and strong fit for LLM/RAG workflows. AGPL licensing is a real constraint for proprietary use—verify your project's license compatibility before committing. If you need to extract documents for LLM ingestion and can work under AGPL (or purchase a commercial license), this is a mature, well-maintained choice.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10+.
- OCR features require Tesseract to be installed separately on the system.
- Low friction: pure Python wheel with four straightforward runtime dependencies (pymupdf, pymupdf_layout, tabulate, psutil).
License · maintenance · safety
(agpl) — Dual-licensed under GNU AGPL 3.0 or Artifex Commercial License. AGPL applies to open-source use; derivative works or distribution require source disclosure. Commercial license available for proprietary projects.
last release 2026-08-06 (8 days) · last repo commit 2026-08-12 · 2,094 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 23,462,774 downloads/mo, #943 on PyPI
Alternatives
Verify before relying
pip install pymupdf4llm
import pymupdf4llm
md = pymupdf4llm.to_markdown("document.pdf")
print(md)- Whether Tesseract installation is automatic or manual-only on different platforms
- Performance baseline for typical document sizes and OCR scenarios
- Compatibility with PyMuPDF Pro for Office document support beyond the base install
What it is and what it does
PyMuPDF4LLM is a lightweight wrapper around PyMuPDF that transforms documents into LLM-ready structured text. It handles multi-column layouts, table detection, header hierarchy, inline formatting, and embedded images—reconstructing natural reading order and metadata without requiring cloud services or GPU. The package includes three output formats (Markdown, JSON, plain text) and integrates with LlamaIndex and LangChain for direct RAG pipeline use.
Its core strength is hybrid OCR: it analyzes each page to decide whether OCR is needed, applying it only to illegible regions or image-covered areas, typically reducing OCR time by around 50% compared to full-document processing. Configuration options let you force OCR on specific pages, set resolution and language, or bring your own OCR function. Page chunking with metadata is available for vector store ingestion.
Use it for
- Extract research papers and technical documents into Markdown for prompt context in LLM applications
- Build RAG pipelines that ingest PDFs as pre-chunked, metadata-rich JSON for vector databases
- Process mixed documents (clean text + scanned pages) with selective OCR to recover all readable content
- Convert multi-column layouts and tables into structured text while preserving reading order
- Index documents for semantic search by converting to plain text or embeddings-ready formats
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Active maintenance, no known vulnerabilities, low install friction, and strong fit for LLM/RAG workflows. AGPL licensing is a real constraint for proprietary use—verify your project's license compatibility before committing. If you need to extract documents for LLM ingestion and can work under AGPL (or purchase a commercial license), this is a mature, well-maintained choice.
Install
pymupdf4llm on PyPI
Before you install
Low friction: pure Python wheel with four straightforward runtime dependencies (pymupdf, pymupdf_layout, tabulate, psutil). Active maintenance—last commit 2026-08-12, release 8 days old. Requires Python 3.10+.
Requires Python 3.10+. OCR features require Tesseract to be installed separately on the system.
License in practice
Dual-licensed under GNU AGPL 3.0 or Artifex Commercial License. AGPL applies to open-source use; derivative works or distribution require source disclosure. Commercial license available for proprietary projects.
Quickstart
pip install pymupdf4llm
import pymupdf4llm
md = pymupdf4llm.to_markdown("document.pdf")
print(md)
Verify before relying
- Whether Tesseract installation is automatic or manual-only on different platforms
- Performance baseline for typical document sizes and OCR scenarios
- Compatibility with PyMuPDF Pro for Office document support beyond the base install
Package facts
| License | Not declared agpl |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 4 packagespymupdfpymupdf_layouttabulatepsutil |
| Maintenance | Actively maintained 8 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 23,462,774 / month, #943 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersProgramming Language :: Python :: 3Topic :: Utilities |
Evidence: pymupdf4llm-1.28.2-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “pdf to markdown for llm”
- pymupdf4llmConverts PDFs and documents into clean, structured Markdown, JSON, or…
- marker-pdfMarker converts PDFs, images, and other document formats (PPTX, DOCX,…
- pdf-oxideExtracts text, images, and metadata from PDFs and converts them to…
Give your agent the search over MCP, or paste the wish link into any chat.
More Utilities packages
Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.
Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Pygments is a syntax highlighter that colorizes source code and text in over 500 languages and formats, outputting to HTML, LaTeX, RTF, SVG, images, or ANSI terminal sequences.
Install it if you need to display or transform source code.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
See also liteparse · opendataloader-pdf · pymupdf-layout · pymupdf · mineru · PyMuPDFb · unstructured · llama-parse · pdf-oxide · MainContentExtractor