$npx skillfedfor your agent

playa-pdf

Parallel and LazY Analyzer for PDFs

With conditionsPyPI Text ProcessingReleased Mar 2026324.0K downloads / moMITPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — playa_pdf-1.1.0-cp310-cp310-macosx_10_9_x86_64.whl · playa_pdf-1.1.0-cp310-cp310-macosx_11_0_arm64.whl · playa_pdf-1.1.0-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl
v1.1.0 · released 2026-03-09 · Python >=3.8 · 1 runtime deps: mypy-extensions

Yes, if you need low-level PDF structure access, metadata inspection, or parallel batch processing. The pure-Python, dependency-light design and MIT license make it a solid choice for those use cases. No, if your only goal is fast text extraction—use pypdfium2 or pypdf instead. Medium friction on install due to compiled wheels, but active maintenance and zero known vulnerabilities reduce risk.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.8 or newer.
  • Optional: install playa-pdf[crypto] to read encrypted PDFs.
  • Medium install friction due to compiled wheels for multiple Python versions and platforms (cp310–cp313, macOS/Linux/Windows).

License · maintenance · safety

MIT (permissive) — MIT license is permissive; you can use, modify, and distribute this package freely in commercial and private projects with minimal restrictions.

last release 2026-03-09 (158 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 323,987 downloads/mo, #7,600 on PyPI

Verify before relying

pip install playa-pdf

import playa
pdf = playa.open("document.pdf")
for page in pdf.pages:
    print(f"Page {page.label}: {page.width} x {page.height}")
  • Whether the layout analysis algorithm implementation is materially faster than pdfminer.six in typical workflows outside the author's benchmarks.
  • Whether text extraction quality and completeness match or exceed other pure-Python PDF libraries for real-world documents.
  • Scope and stability of the command-line interface and its output formats across releases.
Same gist for agents: .md · .json

What it is and what it does

Playa-pdf is a pure-Python PDF reader designed to expose the internals of PDF files—pages, content streams, fonts, images, annotations, document outlines, and logical structure trees—through a lazy, parallelizable interface. It implements the layout analysis algorithm from pdfminer.six and offers both a Python API and a command-line tool for dumping PDF metadata and content. The package is not primarily a text extraction tool; its main strength is providing low-level access to PDF structure and metadata, with optional parallelization across multiple CPUs.

The library supports Python 3.8 through 3.13 and has no C++ dependencies, relying only on mypy-extensions at runtime. It is MIT licensed and actively maintained. While text extraction is possible, the documentation explicitly recommends other tools (pypdfium2, pypdf) for that use case alone. Playa-pdf is most useful when you need to inspect or manipulate PDF internals, extract structured metadata, or process large batches of PDFs in parallel.

Use it for

  • Extract and analyze document outlines, page trees, and logical structure trees from tagged PDFs for accessibility or content mapping.
  • Dump all PDF operators and content streams from a document for low-level analysis or debugging.
  • Batch-extract images and fonts from multiple PDFs in parallel using the lazy, parallelizable API.
  • Access absolute positions and attributes of text, lines, paths, and images on each page for layout analysis.
  • Read encrypted PDFs (with the crypto add-on) and inspect their metadata without full decryption overhead.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need low-level PDF structure access, metadata inspection, or parallel batch processing.

The pure-Python, dependency-light design and MIT license make it a solid choice for those use cases. No, if your only goal is fast text extraction—use pypdfium2 or pypdf instead. Medium friction on install due to compiled wheels, but active maintenance and zero known vulnerabilities reduce risk.

Install

playa-pdf on PyPI

Before you install

Medium install friction due to compiled wheels for multiple Python versions and platforms (cp310–cp313, macOS/Linux/Windows). Active maintenance with recent releases. Single runtime dependency (mypy-extensions) keeps the dependency tree lean.

Requires Python 3.8 or newer. Optional: install playa-pdf[crypto] to read encrypted PDFs.

License in practice

MIT license is permissive; you can use, modify, and distribute this package freely in commercial and private projects with minimal restrictions.

Quickstart

pip install playa-pdf

import playa
pdf = playa.open("document.pdf")
for page in pdf.pages:
    print(f"Page {page.label}: {page.width} x {page.height}")

Verify before relying

  • Whether the layout analysis algorithm implementation is materially faster than pdfminer.six in typical workflows outside the author's benchmarks.
  • Whether text extraction quality and completeness match or exceed other pure-Python PDF libraries for real-world documents.
  • Scope and stability of the command-line interface and its output formats across releases.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.8
Install frictionMedium. Platform-specific wheel
Runtime dependencies
1 package
mypy-extensions
MaintenanceActively maintained 158 days since the last release
First released
Downloads323,987 / month, #7,600 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Science/ResearchProgramming Language :: PythonProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPyTopic :: Text Processing

Evidence: playa_pdf-1.1.0-cp310-cp310-macosx_10_9_x86_64.whl; playa_pdf-1.1.0-cp310-cp310-macosx_11_0_arm64.whl; playa_pdf-1.1.0-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; playa_pdf-1.1.0-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; playa_pdf-1.1.0-cp310-cp310-win_amd64.whl; playa_pdf-1.1.0-cp311-cp311-macosx_10_9_x86_64.whl; playa_pdf-1.1.0-cp311-cp311-macosx_11_0_arm64.whl; playa_pdf-1.1.0-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; playa_pdf-1.1.0-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; playa_pdf-1.1.0-cp311-cp311-win_amd64.whl; playa_pdf-1.1.0-cp312-cp312-macosx_10_13_x86_64.whl; playa_pdf-1.1.0-cp312-cp312-macosx_11_0_arm64.whl; playa_pdf-1.1.0-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; playa_pdf-1.1.0-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; playa_pdf-1.1.0-cp312-cp312-win_amd64.whl; playa_pdf-1.1.0-cp313-cp313-macosx_10_13_x86_64.whl; playa_pdf-1.1.0-cp313-cp313-macosx_11_0_arm64.whl; playa_pdf-1.1.0-cp313-cp313-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; playa_pdf-1.1.0-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; playa_pdf-1.1.0-cp313-cp313-win_amd64.whl

Tags

Capabilities
pdf parsing and extractionpdf metadata and structuretext and image extraction from pdfpdf content stream analysisparallel pdf processingpdf logical structure treelow-level pdf access
Topics
pdf-parsingparallel-processingcli-tool
PyPI keywords
pdf parsertext mining

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “pdf metadata and structure”

  • playa-pdfPlaya-pdf reads PDF files and exposes their internal structure—pages,…
  • borbborb reads, writes, and manipulates PDF files using a pure Python…
  • pdfrw2pdfrw2 reads, writes, and manipulates PDF files with operations…

Give your agent the search over MCP, or paste the wish link into any chat.

More Text Processing packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyparsing Worth it
PyPI · Text Processing · released Jan 2026

pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.

Install it if you need to parse text or define grammars programmatically.

MITpure Python · 3.9+
412.7Mdownloads / mo
fonttools Worth it
PyPI · Text Processing · released May 2026

fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.

Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.

permissive licensepure Python · 3.10+
235.9Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
RapidFuzz Worth it
PyPI · Text Processing · released Apr 2026

RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.

Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.

MITcompiled wheel · 3.10+
184.2Mdownloads / mo
tinycss2 Worth it
PyPI · Text Processing · released Nov 2025

tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.

Install it if your project requires CSS tokenization or syntax manipulation.

BSD-3-Clausepure Python · 3.10+
113.2Mdownloads / mo

See also pdfminer.six · pdfminer · pdftext · pdfrw · pdfplumber · unPDF · pdfrw2 · opendataloader-pdf · fillpdf · pdftotext