$npx skillfedfor your agent

marker-pdf

Convert documents to markdown with high speed and accuracy.

With conditionsPyPI MarkupReleased Jul 2026551.3K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — marker_pdf-2.0.0-py3-none-any.whl
v2.0.0 · released 2026-07-20 · Python <4,>=3.10 · 21 runtime deps: anthropic, click, filetype, ftfy, google-genai, markdown2, markdownify, openai

Yes, with conditions. Marker is actively maintained, well-starred, and solves a real problem—document-to-markdown conversion at scale with layout awareness. The Apache 2.0 license is permissive for code. However, the 21 runtime dependencies (especially torch and transformers) create significant installation and memory overhead. The model weights carry a commercial license restriction for companies over $5M revenue. Install if you need production-grade document parsing; skip if you need lightweight PDF text extraction.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Python 3.10+ required.
  • For non-PDF formats, install marker-pdf[full].
  • Local inference server (vLLM or llama.cpp) spawns automatically; GPU recommended for balanced mode.

License · maintenance · safety

Apache-2.0 (permissive) — Apache 2.0 permissive license allows free commercial use of the code. Model weights use a modified AI Pubs Open Rail-M license (free for research and startups under $5M funding/revenue); commercial use beyond that threshold requires a license.

last release 2026-07-20 (25 days) · last repo commit 2026-08-07 · 38,743 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 551,298 downloads/mo, #6,050 on PyPI

Verify before relying

pip install marker-pdf

from marker.converters.pdf import PdfConverter
from marker.models import create_model_dict

model_dict = create_model_dict()
converter = PdfConverter(model_dict=model_dict)
result = converter("/path/to/document.pdf")
print(result.markdown)
  • Whether the local inference server (vLLM or llama.cpp) spawning is reliable across different system configurations.
  • Performance characteristics on CPU-only systems, since balanced mode defaults to GPU.
  • Whether the LLM integration (Gemini, Claude, OpenAI, etc.) incurs API costs by default or if a free tier is available.
  • Actual memory and disk footprint of the 21 runtime dependencies in a typical installation.
Same gist for agents: .md · .json

What it is and what it does

Marker is a document-to-markdown converter built on vision language models that handles PDFs, images, and office documents across all languages. It reconstructs layout, tables, equations, and inline math, removes headers and footers, and extracts images—all with optional LLM post-processing for higher accuracy. The package runs in balanced mode (GPU-optimized, full-page OCR) or fast mode (CPU-optimized, minimal VLM calls), and can disable OCR entirely for pure text-layer extraction.

The package depends on a large stack: torch, transformers, surya-ocr, anthropic, google-genai, openai, and others. It spawns a local inference server automatically (vLLM on NVIDIA GPUs, llama.cpp elsewhere) unless you point it at an existing one. Installation requires Python 3.10+; the full extras install adds support for non-PDF formats.

Use it for

  • Convert academic papers or textbooks to searchable markdown while preserving tables, equations, and multi-column layout.
  • Extract structured data (tables, forms, values) from scanned or digital PDFs using optional LLM refinement.
  • Batch-process large document collections to markdown for indexing, RAG pipelines, or downstream NLP tasks.
  • Convert presentations (PPTX) and spreadsheets (XLSX) to markdown or JSON for archival or content migration.
  • Build a document ingestion pipeline that handles mixed formats (PDF, image, DOCX, EPUB) in a single workflow.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, with conditions.

Marker is actively maintained, well-starred, and solves a real problem—document-to-markdown conversion at scale with layout awareness. The Apache 2.0 license is permissive for code. However, the 21 runtime dependencies (especially torch and transformers) create significant installation and memory overhead. The model weights carry a commercial license restriction for companies over $5M revenue. Install if you need production-grade document parsing; skip if you need lightweight PDF text extraction.

Install

marker-pdf on PyPI

Before you install

Low friction: pure Python wheel. Active maintenance—last commit 2026-08-07, 38743 stars. However, 21 runtime dependencies including torch, transformers, and surya-ocr add significant disk and memory overhead.

Python 3.10+ required. For non-PDF formats, install marker-pdf[full]. Local inference server (vLLM or llama.cpp) spawns automatically; GPU recommended for balanced mode.

License in practice

Apache 2.0 permissive license allows free commercial use of the code. Model weights use a modified AI Pubs Open Rail-M license (free for research and startups under $5M funding/revenue); commercial use beyond that threshold requires a license.

Quickstart

pip install marker-pdf

from marker.converters.pdf import PdfConverter
from marker.models import create_model_dict

model_dict = create_model_dict()
converter = PdfConverter(model_dict=model_dict)
result = converter("/path/to/document.pdf")
print(result.markdown)

Verify before relying

  • Whether the local inference server (vLLM or llama.cpp) spawning is reliable across different system configurations.
  • Performance characteristics on CPU-only systems, since balanced mode defaults to GPU.
  • Whether the LLM integration (Gemini, Claude, OpenAI, etc.) incurs API costs by default or if a free tier is available.
  • Actual memory and disk footprint of the 21 runtime dependencies in a typical installation.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release <4,>=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
21 packages
anthropicclickfiletypeftfygoogle-genaimarkdown2markdownifyopenaipdftextpillowpsutilpydantic-settingspydanticpython-dotenvrapidfuzzregexscikit-learnsurya-ocrtorchtqdmtransformers
MaintenanceActively maintained 25 days since the last release
Last repo commit
First released
Downloads551,298 / month, #6,050 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: marker_pdf-2.0.0-py3-none-any.whl

Tags

Capabilities
pdf to markdown conversiondocument ocr and extractionpdf layout analysistable extraction from pdfdocument intelligencepdf text extractionmultiformat document parsing
Topics
document-intelligenceocrlayout-analysis
PyPI keywords
markdownnlpocrpdf

Let your AI agent find packages like this

Example. Real query, live index.

An agent finds packages by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language. Give your agent the search over MCP.

More Markup packages

PyYAML Worth it
PyPI · Python Modules · released Sep 2025

PyYAML parses and emits YAML 1.1 data format, enabling serialization and deserialization of configuration files and Python objects to and from human-readable YAML text.

MITcompiled wheel · 3.8+
1.2Bdownloads / mo
markdown-it-py Worth it
PyPI · Python Modules · released May 2026

A Python markdown parser that converts markdown to HTML following the CommonMark specification, with support for plugins and custom syntax rules.

Install it if you need reliable markdown-to-HTML conversion.

MITpure Python · 3.10+
613.8Mdownloads / mo
beautifulsoup4 Worth it
PyPI · Python Modules · released Jun 2026

Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.

Install it if you need to parse or extract data from markup documents.

MITpure Python · 3.7.0+
432.1Mdownloads / mo
et-xmlfile With conditions
PyPI · Markup · released Oct 2024

et_xmlfile writes large XML files with minimal memory overhead by serializing elements to disk as they are created, rather than holding the entire tree in memory.

Install it if incremental XML writing fits your use case; skip it if your XML documents are small or you already use lxml.

MITpure Python · 3.8+dormant
343.3Mdownloads / mo
tomlkit Worth it
PyPI · Markup · released Jul 2026

Parses and edits TOML files while preserving formatting, comments, and structure, then serializes them back with layout intact.

Install it if you're building tools that touch TOML files and user readability of the source matters.

MITpure Python · 3.9+
338.0Mdownloads / mo
docstring-parser Worth it
PyPI · Python Modules · released Apr 2026

Parses Python docstrings in ReST, Google, Numpydoc, and Epydoc formats, extracting structured information like descriptions, parameters, return types, and exceptions.

Install it if you need to programmatically read and extract structured data from Python docstrings.

MITpure Python · 3.8+
274.9Mdownloads / mo

See also datalab-python-sdk · mineru · surya-ocr · landingai-ade · opendataloader-pdf · docling · paddleocr · docling-ibm-models · pdf-oxide · ocrmypdf