pymupdf-layout
PyMuPDF Layout turns PDFs into structured data 10× faster than vision-based tools using AI trained on PDF internals, not images. CPU-only. No GPU required.
Decision gist · record as of 2026-08-14
Yes, with conditions. The package is actively maintained, recently released, and offers a genuine alternative to slower vision-based PDF analysis. However, AGPL licensing means you must either open-source derivative works or obtain a commercial license from Artifex. Install if your use case permits AGPL compliance or if you can license commercially; avoid if proprietary closed-source deployment is required.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >=3.10.
- Compiled wheels available for macOS (x86_64, arm64), Linux (x86_64, aarch64), and Windows (x86_64); other platforms may require building from source.
- Medium install friction due to compiled wheels (C/C++ components) for multiple platforms; active maintenance with recent release (8 days old) and ongoing development.
License · maintenance · safety
(agpl) — Dual-licensed under GNU AFFERO GPL 3.0 or Artifex Commercial License. AGPL terms require that derivative works and network use be licensed under compatible terms; commercial licensing is available from Artifex if proprietary use is needed.
last release 2026-08-06 (8 days) · last repo commit 2026-08-06
0 known vulnerabilities (OSV.dev, 2026-08-14) · 23,361,308 downloads/mo, #947 on PyPI
Alternatives
Verify before relying
pip install pymupdf-layout
from pymupdf_layout import analyze_layout
import pymupdf
doc = pymupdf.open('document.pdf')
result = analyze_layout(doc)- Whether the package's Graph Neural Network model is pre-trained and included, or requires separate download/training.
- CPU performance characteristics and typical analysis time for documents of various sizes.
- Whether onnxruntime can run on CPU-only systems without additional configuration or if GPU acceleration is optional.
What it is and what it does
PyMuPDF Layout is a document analysis library that extracts semantic structure from PDFs by training Graph Neural Networks on PDF internals rather than rendered images. It identifies and labels document elements—titles, headings, headers, footers, tables, images, and text styling—and outputs the results in structured formats (Markdown, JSON, TXT). The package integrates with PyMuPDF and is designed to be fast and CPU-only, avoiding the overhead of vision-based machine learning models.
The library is part of the PyMuPDF ecosystem and serves as a core component of PyMuPDF4LLM for document preparation. It depends on PyMuPDF for PDF access, onnxruntime for neural network inference, networkx for graph operations, numpy for numerical work, and pyyaml for configuration. It targets developers who need to extract and structure document content programmatically without GPU resources.
Use it for
- Convert PDF documents to clean Markdown or JSON for ingestion into language models or downstream processing pipelines.
- Automatically detect and separate header and footer content from main document text across multiple pages.
- Extract and label semantic document structure (titles, sections, tables) for content management or archival systems.
- Prepare PDF documents for LLM-based analysis by providing pre-parsed, semantically annotated content.
- Build document processing workflows that identify and isolate specific element types (tables, images, text blocks) without manual annotation.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with conditions.
The package is actively maintained, recently released, and offers a genuine alternative to slower vision-based PDF analysis. However, AGPL licensing means you must either open-source derivative works or obtain a commercial license from Artifex. Install if your use case permits AGPL compliance or if you can license commercially; avoid if proprietary closed-source deployment is required.
Install
pymupdf-layout on PyPI
Before you install
Medium install friction due to compiled wheels (C/C++ components) for multiple platforms; active maintenance with recent release (8 days old) and ongoing development. Requires Python >=3.10 and five runtime dependencies including PyMuPDF, onnxruntime, and networkx.
Requires Python >=3.10. Compiled wheels available for macOS (x86_64, arm64), Linux (x86_64, aarch64), and Windows (x86_64); other platforms may require building from source.
License in practice
Dual-licensed under GNU AFFERO GPL 3.0 or Artifex Commercial License. AGPL terms require that derivative works and network use be licensed under compatible terms; commercial licensing is available from Artifex if proprietary use is needed.
Quickstart
pip install pymupdf-layout
from pymupdf_layout import analyze_layout
import pymupdf
doc = pymupdf.open('document.pdf')
result = analyze_layout(doc)
Verify before relying
- Whether the package's Graph Neural Network model is pre-trained and included, or requires separate download/training.
- CPU performance characteristics and typical analysis time for documents of various sizes.
- Whether onnxruntime can run on CPU-only systems without additional configuration or if GPU acceleration is optional.
Package facts
| License | Not declared agpl |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 5 packagesPyMuPDFpyyamlnumpyonnxruntimenetworkx |
| Maintenance | Actively maintained 8 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 23,361,308 / month, #947 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Information TechnologyLicense :: Other/Proprietary LicenseOperating System :: MacOSOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: CProgramming Language :: C++Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: Implementation :: CPythonTopic :: Multimedia :: GraphicsTopic :: Software Development :: LibrariesTopic :: Utilities |
Evidence: pymupdf_layout-1.28.2-cp310-abi3-macosx_10_9_x86_64.whl; pymupdf_layout-1.28.2-cp310-abi3-macosx_11_0_arm64.whl; pymupdf_layout-1.28.2-cp310-abi3-manylinux_2_28_aarch64.whl; pymupdf_layout-1.28.2-cp310-abi3-manylinux_2_28_x86_64.whl; pymupdf_layout-1.28.2-cp310-abi3-win_amd64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “extract pdf structure”
- pymupdf-layoutPyMuPDF Layout analyzes PDF structure and content using Graph Neural…
- playa-pdfPlaya-pdf reads PDF files and exposes their internal structure—pages,…
- pdfminerExtracts text and layout information from PDF documents, including…
Give your agent the search over MCP, or paste the wish link into any chat.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also pymupdf · pymupdf4llm · pdfid · PyMuPDFb · pdftext · opendataloader-pdf · pdfminer.six · marker-pdf · markdown-pdf · skan