$npx skillfedfor your agent

docling-parse

Simple package to extract text with coordinates from programmatic PDFs

Worth itPyPI Text ProcessingReleased Aug 20264.6M downloads / moMITPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — docling_parse-7.13.0-cp310-cp310-macosx_14_0_arm64.whl · docling_parse-7.13.0-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl · docling_parse-7.13.0-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
v7.13.0 · released 2026-08-14 · Python >=3.10 · 4 runtime deps: pillow, pydantic, docling-core, pywin32

Yes. Docling Parse is actively maintained, permissively licensed, and offers a well-designed API for structured PDF extraction with multi-threaded support. Install friction is moderate due to compiled components, but pre-built wheels cover all major platforms and Python versions. Suitable for production document processing workflows.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python >=3.10; compiled wheels depend on system C++ runtime libraries.
  • Medium install friction due to compiled C++ components with pre-built wheels for Python 3.10–3.14 across macOS, Linux, and Windows.
  • Active maintenance with a release on 2026-08-14 and 326 repository stars.

License · maintenance · safety

MIT (permissive) — MIT license permits unrestricted use, modification, and distribution in both open-source and commercial projects.

last release 2026-08-14 (0 days) · last repo commit 2026-08-14 · 326 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 4,605,553 downloads/mo, #2,275 on PyPI

Verify before relying

pip install docling-parse

from docling_parse.pdf_parser import DoclingPdfParser, DecodeConfig, ContentConfig, ContentLevel

parser = DoclingPdfParser(loglevel="fatal")
pdf_doc = parser.load(
    path_or_stream="file.pdf",
    decode_config=DecodeConfig(do_sanitization=True),
    content_config=ContentConfig(
        word_cells_content_level=ContentLevel.COMPUTE_AND_MATERIALIZE,
    ),
)

for page_no, page in pdf_doc.iterate_pages():
    for word in page.iterate_cells():
        print(word.rect, word.text)
  • Whether the package handles encrypted or password-protected PDFs beyond what the CLI suggests.
  • Performance characteristics on very large PDFs or batch workloads compared to alternatives.
  • Memory footprint when materializing all cell levels for high-page-count documents.
Same gist for agents: .md · .json

What it is and what it does

Docling Parse is a Python wrapper around a C++ PDF parser that extracts structured text, geometric coordinates, and images from programmatic PDFs. It splits parsing into two phases: a fixed `DecodeConfig` applied at document open time (controlling sanitization and glyph handling) and a per-page `ContentConfig` that determines what to compute and materialize (character cells, word cells, line cells, shapes, bitmaps). This separation allows cheap initial loading and selective enrichment on demand—if you request richer output later, the page is re-decoded automatically.

The package supports both sequential parsing (one PDF at a time) and parallel multi-threaded parsing with backpressure control. It includes a CLI for single-file processing and integrates with the broader Docling PDF conversion ecosystem. The library is actively maintained, supports Python 3.10–3.14 across major platforms, and provides performance benchmarks against other PDF packages.

Use it for

  • Extract word-level bounding boxes and text from PDFs for document layout analysis or OCR validation.
  • Batch-process multiple PDFs in parallel with configurable thread pools and result backpressure.
  • Render pages as images with overlaid cell boundaries (character, word, or line level) for debugging or visualization.
  • Selectively materialize only the content levels needed per page to optimize memory and CPU in large-scale workflows.
  • Integrate PDF parsing into document conversion pipelines that require both text and spatial metadata.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Docling Parse is actively maintained, permissively licensed, and offers a well-designed API for structured PDF extraction with multi-threaded support. Install friction is moderate due to compiled components, but pre-built wheels cover all major platforms and Python versions. Suitable for production document processing workflows.

Install

docling-parse on PyPI

Before you install

Medium install friction due to compiled C++ components with pre-built wheels for Python 3.10–3.14 across macOS, Linux, and Windows. Active maintenance with a release on 2026-08-14 and 326 repository stars.

Requires Python >=3.10; compiled wheels depend on system C++ runtime libraries.

License in practice

MIT license permits unrestricted use, modification, and distribution in both open-source and commercial projects.

Quickstart

pip install docling-parse

from docling_parse.pdf_parser import DoclingPdfParser, DecodeConfig, ContentConfig, ContentLevel

parser = DoclingPdfParser(loglevel="fatal")
pdf_doc = parser.load(
    path_or_stream="file.pdf",
    decode_config=DecodeConfig(do_sanitization=True),
    content_config=ContentConfig(
        word_cells_content_level=ContentLevel.COMPUTE_AND_MATERIALIZE,
    ),
)

for page_no, page in pdf_doc.iterate_pages():
    for word in page.iterate_cells():
        print(word.rect, word.text)

Verify before relying

  • Whether the package handles encrypted or password-protected PDFs beyond what the CLI suggests.
  • Performance characteristics on very large PDFs or batch workloads compared to alternatives.
  • Memory footprint when materializing all cell levels for high-page-count documents.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.10
Install frictionMedium. Platform-specific wheel
Runtime dependencies
4 packages
pillowpydanticdocling-corepywin32
MaintenanceActively maintained 0 days since the last release
Last repo commit
First released
Downloads4,605,553 / month, #2,275 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Science/ResearchOperating System :: MacOS :: MacOS XOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: C++Programming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14

Evidence: docling_parse-7.13.0-cp310-cp310-macosx_14_0_arm64.whl; docling_parse-7.13.0-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; docling_parse-7.13.0-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; docling_parse-7.13.0-cp310-cp310-win_amd64.whl; docling_parse-7.13.0-cp310-cp310-win_arm64.whl; docling_parse-7.13.0-cp311-cp311-macosx_14_0_arm64.whl; docling_parse-7.13.0-cp311-cp311-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl; docling_parse-7.13.0-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; docling_parse-7.13.0-cp311-cp311-win_amd64.whl; docling_parse-7.13.0-cp311-cp311-win_arm64.whl; docling_parse-7.13.0-cp312-cp312-macosx_14_0_arm64.whl; docling_parse-7.13.0-cp312-cp312-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl; docling_parse-7.13.0-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; docling_parse-7.13.0-cp312-cp312-win_amd64.whl; docling_parse-7.13.0-cp312-cp312-win_arm64.whl; docling_parse-7.13.0-cp313-cp313-macosx_14_0_arm64.whl; docling_parse-7.13.0-cp313-cp313-manylinux_2_26_aarch64.manylinux_2_28_aarch64.whl; docling_parse-7.13.0-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; docling_parse-7.13.0-cp313-cp313-win_amd64.whl; docling_parse-7.13.0-cp313-cp313-win_arm64.whl

Tags

Capabilities
pdf text extraction with coordinatespdf parsing pythonextract text from pdfpdf document parserprogrammatic pdf processingpdf image extractionbatch pdf parsing
Topics
pdf-parsingdocument-extractionmulti-threaded
PyPI keywords
doclingpdfparser

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “pdf text extraction with coordinates”

  • docling-parseExtracts text, coordinates, and bitmap images from programmatic PDFs…
  • opendataloader-pdfExtracts structured data (Markdown, JSON, HTML) from PDFs with…
  • pdftextExtracts plain text or structured blocks, lines, and spans from PDFs…

Give your agent the search over MCP, or paste the wish link into any chat.

More Text Processing packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyparsing Worth it
PyPI · Text Processing · released Jan 2026

pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.

Install it if you need to parse text or define grammars programmatically.

MITpure Python · 3.9+
412.7Mdownloads / mo
fonttools Worth it
PyPI · Text Processing · released May 2026

fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.

Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.

permissive licensepure Python · 3.10+
235.9Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
RapidFuzz Worth it
PyPI · Text Processing · released Apr 2026

RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.

Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.

MITcompiled wheel · 3.10+
184.2Mdownloads / mo
tinycss2 Worth it
PyPI · Text Processing · released Nov 2025

tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.

Install it if your project requires CSS tokenization or syntax manipulation.

BSD-3-Clausepure Python · 3.10+
113.2Mdownloads / mo

See also docling · textract · docling-core · unPDF · docling-slim · marker-pdf · pdftotext · pdftext · docling-ibm-models · langchain-docling