$npx skillfedfor your agent

pymupdf-layout

PyMuPDF Layout turns PDFs into structured data 10× faster than vision-based tools using AI trained on PDF internals, not images. CPU-only. No GPU required.

With conditionsPyPI LibrariesReleased Aug 202623.4M downloads / moPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — pymupdf_layout-1.28.2-cp310-abi3-macosx_10_9_x86_64.whl · pymupdf_layout-1.28.2-cp310-abi3-macosx_11_0_arm64.whl · pymupdf_layout-1.28.2-cp310-abi3-manylinux_2_28_aarch64.whl
v1.28.2 · released 2026-08-06 · Python >=3.10 · 5 runtime deps: PyMuPDF, pyyaml, numpy, onnxruntime, networkx

Yes, with conditions. The package is actively maintained, recently released, and offers a genuine alternative to slower vision-based PDF analysis. However, AGPL licensing means you must either open-source derivative works or obtain a commercial license from Artifex. Install if your use case permits AGPL compliance or if you can license commercially; avoid if proprietary closed-source deployment is required.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python >=3.10.
  • Compiled wheels available for macOS (x86_64, arm64), Linux (x86_64, aarch64), and Windows (x86_64); other platforms may require building from source.
  • Medium install friction due to compiled wheels (C/C++ components) for multiple platforms; active maintenance with recent release (8 days old) and ongoing development.

License · maintenance · safety

(agpl) — Dual-licensed under GNU AFFERO GPL 3.0 or Artifex Commercial License. AGPL terms require that derivative works and network use be licensed under compatible terms; commercial licensing is available from Artifex if proprietary use is needed.

last release 2026-08-06 (8 days) · last repo commit 2026-08-06

0 known vulnerabilities (OSV.dev, 2026-08-14) · 23,361,308 downloads/mo, #947 on PyPI

Verify before relying

pip install pymupdf-layout

from pymupdf_layout import analyze_layout
import pymupdf

doc = pymupdf.open('document.pdf')
result = analyze_layout(doc)
  • Whether the package's Graph Neural Network model is pre-trained and included, or requires separate download/training.
  • CPU performance characteristics and typical analysis time for documents of various sizes.
  • Whether onnxruntime can run on CPU-only systems without additional configuration or if GPU acceleration is optional.
Same gist for agents: .md · .json

What it is and what it does

PyMuPDF Layout is a document analysis library that extracts semantic structure from PDFs by training Graph Neural Networks on PDF internals rather than rendered images. It identifies and labels document elements—titles, headings, headers, footers, tables, images, and text styling—and outputs the results in structured formats (Markdown, JSON, TXT). The package integrates with PyMuPDF and is designed to be fast and CPU-only, avoiding the overhead of vision-based machine learning models.

The library is part of the PyMuPDF ecosystem and serves as a core component of PyMuPDF4LLM for document preparation. It depends on PyMuPDF for PDF access, onnxruntime for neural network inference, networkx for graph operations, numpy for numerical work, and pyyaml for configuration. It targets developers who need to extract and structure document content programmatically without GPU resources.

Use it for

  • Convert PDF documents to clean Markdown or JSON for ingestion into language models or downstream processing pipelines.
  • Automatically detect and separate header and footer content from main document text across multiple pages.
  • Extract and label semantic document structure (titles, sections, tables) for content management or archival systems.
  • Prepare PDF documents for LLM-based analysis by providing pre-parsed, semantically annotated content.
  • Build document processing workflows that identify and isolate specific element types (tables, images, text blocks) without manual annotation.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, with conditions.

The package is actively maintained, recently released, and offers a genuine alternative to slower vision-based PDF analysis. However, AGPL licensing means you must either open-source derivative works or obtain a commercial license from Artifex. Install if your use case permits AGPL compliance or if you can license commercially; avoid if proprietary closed-source deployment is required.

Install

pymupdf-layout on PyPI

Before you install

Medium install friction due to compiled wheels (C/C++ components) for multiple platforms; active maintenance with recent release (8 days old) and ongoing development. Requires Python >=3.10 and five runtime dependencies including PyMuPDF, onnxruntime, and networkx.

Requires Python >=3.10. Compiled wheels available for macOS (x86_64, arm64), Linux (x86_64, aarch64), and Windows (x86_64); other platforms may require building from source.

License in practice

Dual-licensed under GNU AFFERO GPL 3.0 or Artifex Commercial License. AGPL terms require that derivative works and network use be licensed under compatible terms; commercial licensing is available from Artifex if proprietary use is needed.

Quickstart

pip install pymupdf-layout

from pymupdf_layout import analyze_layout
import pymupdf

doc = pymupdf.open('document.pdf')
result = analyze_layout(doc)

Verify before relying

  • Whether the package's Graph Neural Network model is pre-trained and included, or requires separate download/training.
  • CPU performance characteristics and typical analysis time for documents of various sizes.
  • Whether onnxruntime can run on CPU-only systems without additional configuration or if GPU acceleration is optional.

Package facts

LicenseNot declared agpl
Python supportSupports the current Python release >=3.10
Install frictionMedium. Platform-specific wheel
Runtime dependencies
5 packages
PyMuPDFpyyamlnumpyonnxruntimenetworkx
MaintenanceActively maintained 8 days since the last release
Last repo commit
First released
Downloads23,361,308 / month, #947 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Information TechnologyLicense :: Other/Proprietary LicenseOperating System :: MacOSOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: CProgramming Language :: C++Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: Implementation :: CPythonTopic :: Multimedia :: GraphicsTopic :: Software Development :: LibrariesTopic :: Utilities

Evidence: pymupdf_layout-1.28.2-cp310-abi3-macosx_10_9_x86_64.whl; pymupdf_layout-1.28.2-cp310-abi3-macosx_11_0_arm64.whl; pymupdf_layout-1.28.2-cp310-abi3-manylinux_2_28_aarch64.whl; pymupdf_layout-1.28.2-cp310-abi3-manylinux_2_28_x86_64.whl; pymupdf_layout-1.28.2-cp310-abi3-win_amd64.whl

Tags

Capabilities
pdf layout analysisextract pdf structurepdf to markdown jsondocument semantic understandingpdf content extractionstructured pdf parsingpdf document analysis
Topics
pdf-processinglayout-analysisdocument-extraction

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “extract pdf structure”

  • pymupdf-layoutPyMuPDF Layout analyzes PDF structure and content using Graph Neural…
  • playa-pdfPlaya-pdf reads PDF files and exposes their internal structure—pages,…
  • pdfminerExtracts text and layout information from PDF documents, including…

Give your agent the search over MCP, or paste the wish link into any chat.

More Libraries packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
python-dateutil Worth it
PyPI · Libraries · released Mar 2024

Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.

Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.

Apache-2.0pure Python
1.2Bdownloads / mo
six With conditions
PyPI · Libraries · released Dec 2024

Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.

MITpure Python
1.2Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo

See also pymupdf · pymupdf4llm · pdfid · PyMuPDFb · pdftext · opendataloader-pdf · pdfminer.six · marker-pdf · markdown-pdf · skan