$npx skillfedfor your agent

opendataloader-pdf

A Python wrapper for the opendataloader-pdf Java CLI.

Worth itPyPI Text ProcessingReleased Jul 2026184.2K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — opendataloader_pdf-2.5.0-py3-none-any.whl
v2.5.0 · released 2026-07-14 · Python >=3.10

Yes. Active, well-maintained project with strong benchmark performance (0.907 overall accuracy, 0.928 table extraction), no security vulnerabilities, and low install friction. Apache-2.0 license covers core features. Java 11+ system dependency is the only real prerequisite. Suitable for production RAG pipelines and accessibility automation workflows; hybrid mode adds cost (API calls) but handles complex documents well.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Java 11+ installed and available on PATH before running Python code.
  • Low friction: pure Python wheel with no runtime dependencies.
  • Active maintenance (last commit 2026-08-13, 28406 stars) and recent releases.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license. Core extraction and auto-tagging are free and open-source; PDF/UA export and accessibility studio are enterprise add-ons.

last release 2026-07-14 (31 days) · last repo commit 2026-08-13 · 28,406 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 184,207 downloads/mo, #10,044 on PyPI

Verify before relying

pip install opendataloader-pdf

import opendataloader_pdf

opendataloader_pdf.convert(
    input_path=["file.pdf"],
    output_dir="output/",
    format="markdown,json"
)
  • Whether the hybrid mode (AI-assisted extraction for complex tables and scanned PDFs) requires external API calls or runs locally.
  • Performance characteristics and resource usage when processing large batches or high-volume PDF collections.
  • Whether OCR in hybrid mode supports all 80+ claimed languages equally or has language-specific accuracy variations.
Same gist for agents: .md · .json

What it is and what it does

OpenDataLoader PDF is a Python wrapper around a Java-based PDF parser that extracts text, tables, images, and structural metadata from PDFs into multiple formats (Markdown, JSON with bounding boxes, HTML). It offers two modes: a deterministic local mode for standard digital PDFs, and a hybrid mode that routes complex pages (nested tables, scanned documents, formulas, charts) to an AI backend. The package also includes auto-tagging functionality to convert untagged PDFs into Tagged PDF format, which is the foundation for PDF accessibility compliance.

The tool is designed for two main workflows: building AI-ready datasets for RAG and LLM pipelines (with structured output and source citations via bounding boxes), and automating PDF accessibility remediation at scale. It requires Java 11+ as a system dependency but has no Python runtime dependencies, making installation straightforward. The core extraction and auto-tagging features are open-source under Apache-2.0; PDF/UA export and visual editing tools are enterprise add-ons.

Use it for

  • Extract structured Markdown and JSON from PDFs for RAG pipelines, with bounding boxes for source citation and retrieval.
  • Auto-tag untagged PDFs into screen-reader-ready Tagged PDF format to meet accessibility regulations (EAA, ADA, Section 508) at scale.
  • Parse complex or scanned PDFs (tables, formulas, charts, poor-quality scans) using hybrid mode for accurate structured output.
  • Build document processing workflows that preserve reading order, heading hierarchy, and list structure for downstream NLP tasks.
  • Extract tables and images with precise coordinates for document layout reconstruction or visual document analysis.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Active, well-maintained project with strong benchmark performance (0.907 overall accuracy, 0.928 table extraction), no security vulnerabilities, and low install friction. Apache-2.0 license covers core features. Java 11+ system dependency is the only real prerequisite. Suitable for production RAG pipelines and accessibility automation workflows; hybrid mode adds cost (API calls) but handles complex documents well.

Install

opendataloader-pdf on PyPI

Before you install

Low friction: pure Python wheel with no runtime dependencies. Active maintenance (last commit 2026-08-13, 28406 stars) and recent releases. Requires Java 11+ as a system prerequisite, not a Python dependency.

Requires Java 11+ installed and available on PATH before running Python code.

License in practice

Apache-2.0 permissive license. Core extraction and auto-tagging are free and open-source; PDF/UA export and accessibility studio are enterprise add-ons.

Quickstart

pip install opendataloader-pdf

import opendataloader_pdf

opendataloader_pdf.convert(
    input_path=["file.pdf"],
    output_dir="output/",
    format="markdown,json"
)

Verify before relying

  • Whether the hybrid mode (AI-assisted extraction for complex tables and scanned PDFs) requires external API calls or runs locally.
  • Performance characteristics and resource usage when processing large batches or high-volume PDF collections.
  • Whether OCR in hybrid mode supports all 80+ claimed languages equally or has language-specific accuracy variations.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependenciesNone
MaintenanceActively maintained 31 days since the last release
Last repo commit
First released
Downloads184,207 / month, #10,044 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Operating System :: OS IndependentProgramming Language :: Python :: 3

Evidence: opendataloader_pdf-2.5.0-py3-none-any.whl

Tags

Capabilities
pdf to markdown json extractionpdf data extraction for ragpdf accessibility auto-taggingpdf parser with bounding boxespdf to structured datatagged pdf generationpdf layout analysis
Topics
pdf-extractionrag-pipelineaccessibility

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “pdf to markdown json extraction”

  • opendataloader-pdfExtracts structured data (Markdown, JSON, HTML) from PDFs with…
  • marker-pdfMarker converts PDFs, images, and other document formats (PPTX, DOCX,…
  • pymupdf-layoutPyMuPDF Layout analyzes PDF structure and content using Graph Neural…

Give your agent the search over MCP, or paste the wish link into any chat.

More Text Processing packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyparsing Worth it
PyPI · Text Processing · released Jan 2026

pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.

Install it if you need to parse text or define grammars programmatically.

MITpure Python · 3.9+
412.7Mdownloads / mo
fonttools Worth it
PyPI · Text Processing · released May 2026

fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.

Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.

permissive licensepure Python · 3.10+
235.9Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
RapidFuzz Worth it
PyPI · Text Processing · released Apr 2026

RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.

Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.

MITcompiled wheel · 3.10+
184.2Mdownloads / mo
tinycss2 Worth it
PyPI · Text Processing · released Nov 2025

tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.

Install it if your project requires CSS tokenization or syntax manipulation.

BSD-3-Clausepure Python · 3.10+
113.2Mdownloads / mo

See also landingai-ade · pymupdf4llm · marker-pdf · camelot-py · pymupdf-layout · liteparse · unstructured-inference · playa-pdf · pdftext · unPDF