$npx skillfedfor your agent

pymupdf4llm

PyMuPDF Utilities for LLM/RAG

Worth itPyPI UtilitiesReleased Aug 202623.5M downloads / moPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — pymupdf4llm-1.28.2-py3-none-any.whl
v1.28.2 · released 2026-08-06 · Python >=3.10 · 4 runtime deps: pymupdf, pymupdf_layout, tabulate, psutil

Yes. Active maintenance, no known vulnerabilities, low install friction, and strong fit for LLM/RAG workflows. AGPL licensing is a real constraint for proprietary use—verify your project's license compatibility before committing. If you need to extract documents for LLM ingestion and can work under AGPL (or purchase a commercial license), this is a mature, well-maintained choice.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10+.
  • OCR features require Tesseract to be installed separately on the system.
  • Low friction: pure Python wheel with four straightforward runtime dependencies (pymupdf, pymupdf_layout, tabulate, psutil).

License · maintenance · safety

(agpl) — Dual-licensed under GNU AGPL 3.0 or Artifex Commercial License. AGPL applies to open-source use; derivative works or distribution require source disclosure. Commercial license available for proprietary projects.

last release 2026-08-06 (8 days) · last repo commit 2026-08-12 · 2,094 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 23,462,774 downloads/mo, #943 on PyPI

Verify before relying

pip install pymupdf4llm

import pymupdf4llm

md = pymupdf4llm.to_markdown("document.pdf")
print(md)
  • Whether Tesseract installation is automatic or manual-only on different platforms
  • Performance baseline for typical document sizes and OCR scenarios
  • Compatibility with PyMuPDF Pro for Office document support beyond the base install
Same gist for agents: .md · .json

What it is and what it does

PyMuPDF4LLM is a lightweight wrapper around PyMuPDF that transforms documents into LLM-ready structured text. It handles multi-column layouts, table detection, header hierarchy, inline formatting, and embedded images—reconstructing natural reading order and metadata without requiring cloud services or GPU. The package includes three output formats (Markdown, JSON, plain text) and integrates with LlamaIndex and LangChain for direct RAG pipeline use.

Its core strength is hybrid OCR: it analyzes each page to decide whether OCR is needed, applying it only to illegible regions or image-covered areas, typically reducing OCR time by around 50% compared to full-document processing. Configuration options let you force OCR on specific pages, set resolution and language, or bring your own OCR function. Page chunking with metadata is available for vector store ingestion.

Use it for

  • Extract research papers and technical documents into Markdown for prompt context in LLM applications
  • Build RAG pipelines that ingest PDFs as pre-chunked, metadata-rich JSON for vector databases
  • Process mixed documents (clean text + scanned pages) with selective OCR to recover all readable content
  • Convert multi-column layouts and tables into structured text while preserving reading order
  • Index documents for semantic search by converting to plain text or embeddings-ready formats

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Active maintenance, no known vulnerabilities, low install friction, and strong fit for LLM/RAG workflows. AGPL licensing is a real constraint for proprietary use—verify your project's license compatibility before committing. If you need to extract documents for LLM ingestion and can work under AGPL (or purchase a commercial license), this is a mature, well-maintained choice.

Install

pymupdf4llm on PyPI

Before you install

Low friction: pure Python wheel with four straightforward runtime dependencies (pymupdf, pymupdf_layout, tabulate, psutil). Active maintenance—last commit 2026-08-12, release 8 days old. Requires Python 3.10+.

Requires Python 3.10+. OCR features require Tesseract to be installed separately on the system.

License in practice

Dual-licensed under GNU AGPL 3.0 or Artifex Commercial License. AGPL applies to open-source use; derivative works or distribution require source disclosure. Commercial license available for proprietary projects.

Quickstart

pip install pymupdf4llm

import pymupdf4llm

md = pymupdf4llm.to_markdown("document.pdf")
print(md)

Verify before relying

  • Whether Tesseract installation is automatic or manual-only on different platforms
  • Performance baseline for typical document sizes and OCR scenarios
  • Compatibility with PyMuPDF Pro for Office document support beyond the base install

Package facts

LicenseNot declared agpl
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
4 packages
pymupdfpymupdf_layouttabulatepsutil
MaintenanceActively maintained 8 days since the last release
Last repo commit
First released
Downloads23,462,774 / month, #943 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersProgramming Language :: Python :: 3Topic :: Utilities

Evidence: pymupdf4llm-1.28.2-py3-none-any.whl

Tags

Capabilities
pdf to markdown for llmdocument extraction rag pipelinepdf text extraction with layoutocr for llm ingestionpdf to json structured datadocument parsing vector embeddingspdf layout analysis
Topics
llm-ragdocument-extractionocr

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “pdf to markdown for llm”

  • pymupdf4llmConverts PDFs and documents into clean, structured Markdown, JSON, or…
  • marker-pdfMarker converts PDFs, images, and other document formats (PPTX, DOCX,…
  • pdf-oxideExtracts text, images, and metadata from PDFs and converts them to…

Give your agent the search over MCP, or paste the wish link into any chat.

More Utilities packages

idna Worth it
PyPI · Python Modules · released Jun 2026

Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.

Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.

BSD-3-Clausepure Python · 3.9+
1.8Bdownloads / mo
charset-normalizer Worth it
PyPI · Utilities · released Aug 2026

Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.

permissive licensepure Python · 3.7+
1.7Bdownloads / mo
setuptools Worth it
PyPI · Python Modules · released Aug 2026

Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.

MITpure Python · 3.10+
1.6Bdownloads / mo
pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
Pygments Worth it
PyPI · Utilities · released Mar 2026

Pygments is a syntax highlighter that colorizes source code and text in over 500 languages and formats, outputting to HTML, LaTeX, RTF, SVG, images, or ANSI terminal sequences.

Install it if you need to display or transform source code.

BSD-2-Clausepure Python · 3.9+
1.3Bdownloads / mo
six With conditions
PyPI · Libraries · released Dec 2024

Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.

MITpure Python
1.2Bdownloads / mo

See also liteparse · opendataloader-pdf · pymupdf-layout · pymupdf · mineru · PyMuPDFb · unstructured · llama-parse · pdf-oxide · MainContentExtractor