$npx skillfedfor your agent

unPDF

Quickly extract text characters and character metadata from pdfs using pdfium.

With conditionsPyPI Text ProcessingReleased Jul 2025102.0K downloads / moApache-2.0Platform wheel

Decision gist · record as of 2026-08-14

platform wheels — unpdf-1.0.0-cp310-cp310-macosx_11_0_arm64.whl · unpdf-1.0.0-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl · unpdf-1.0.0-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl
v1.0.0 · released 2025-07-10 · Python >=3.9

Yes, if you need character-level PDF metadata and positioning. The zero-dependency design and precompiled wheels make installation straightforward. However, the aging maintenance status (400 days since release) and lack of visible repository activity suggest limited ongoing support—verify stability for production use before committing.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python >=3.9; PyArrow optional but recommended for data analysis workflows.
  • Medium install friction due to compiled wheels for multiple Python versions and platforms (cp310–cp313, macOS/Linux/Windows).
  • Aging maintenance status (400 days since release) suggests limited ongoing support.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; attribution required.

last release 2025-07-10 (400 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 101,999 downloads/mo, #12,899 on PyPI

Verify before relying

pip install unpdf

from unpdf import extract

pages, chars, text_objs, fonts = extract("document.pdf")
text = ''.join(chr(c) for c in chars.arrays['char'])
print(text)
  • Whether PDFium is bundled in wheels or requires separate system installation.
  • Performance characteristics on large PDFs or high-volume extraction tasks.
  • Stability and bug-fix frequency given aging maintenance status.
Same gist for agents: .md · .json

What it is and what it does

unPDF is a Python library that extracts text and metadata from PDF files at the character level using PDFium. It returns four table objects—pages, characters, text objects, and fonts—each containing structured data: character positions and bounding boxes, font sizes and names, RGBA color values, and transformation matrices. The library ships with no runtime dependencies and optionally integrates with PyArrow for conversion to pandas or polars DataFrames.

The package is designed for developers who need precise, granular access to PDF content rather than simple text extraction. It exposes low-level PDFium data structures directly, making it suitable for document analysis, layout reconstruction, and metadata-driven workflows. Installation requires Python >=3.9 and uses precompiled wheels for common platforms.

Use it for

  • Extract character positions and bounding boxes for document layout analysis or OCR validation.
  • Retrieve font metadata and color information to reconstruct visual styling or detect formatting changes.
  • Convert PDF character data to pandas/polars DataFrames for statistical analysis or data science pipelines.
  • Build custom text reconstruction logic that preserves spatial relationships and transformation matrices.
  • Analyze multi-language PDFs by accessing unicode character codes and mapping errors.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need character-level PDF metadata and positioning.

The zero-dependency design and precompiled wheels make installation straightforward. However, the aging maintenance status (400 days since release) and lack of visible repository activity suggest limited ongoing support—verify stability for production use before committing.

Install

unpdf on PyPI

Before you install

Medium install friction due to compiled wheels for multiple Python versions and platforms (cp310–cp313, macOS/Linux/Windows). Aging maintenance status (400 days since release) suggests limited ongoing support.

Requires Python >=3.9; PyArrow optional but recommended for data analysis workflows.

License in practice

Apache-2.0 permissive license allows commercial and private use with minimal restrictions; attribution required.

Quickstart

pip install unpdf

from unpdf import extract

pages, chars, text_objs, fonts = extract("document.pdf")
text = ''.join(chr(c) for c in chars.arrays['char'])
print(text)

Verify before relying

  • Whether PDFium is bundled in wheels or requires separate system installation.
  • Performance characteristics on large PDFs or high-volume extraction tasks.
  • Stability and bug-fix frequency given aging maintenance status.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.9
Install frictionMedium. Platform-specific wheel
Runtime dependenciesNone
MaintenanceAging 400 days since the last release
First released
Downloads101,999 / month, #12,899 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: unpdf-1.0.0-cp310-cp310-macosx_11_0_arm64.whl; unpdf-1.0.0-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; unpdf-1.0.0-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; unpdf-1.0.0-cp310-cp310-win_amd64.whl; unpdf-1.0.0-cp310-cp310-win_arm64.whl; unpdf-1.0.0-cp311-cp311-macosx_11_0_arm64.whl; unpdf-1.0.0-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; unpdf-1.0.0-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; unpdf-1.0.0-cp311-cp311-musllinux_1_2_aarch64.whl; unpdf-1.0.0-cp311-cp311-musllinux_1_2_x86_64.whl; unpdf-1.0.0-cp311-cp311-win_amd64.whl; unpdf-1.0.0-cp311-cp311-win_arm64.whl; unpdf-1.0.0-cp312-cp312-macosx_11_0_arm64.whl; unpdf-1.0.0-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; unpdf-1.0.0-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; unpdf-1.0.0-cp312-cp312-musllinux_1_2_aarch64.whl; unpdf-1.0.0-cp312-cp312-musllinux_1_2_x86_64.whl; unpdf-1.0.0-cp312-cp312-win_amd64.whl; unpdf-1.0.0-cp312-cp312-win_arm64.whl; unpdf-1.0.0-cp313-cp313-macosx_11_0_arm64.whl

Tags

Capabilities
pdf text extractioncharacter-level pdf parsingpdf metadata extractionpdf font and color infopdfium python wrapperpdf bounding box extractionpdf character positioning
Topics
pdf-extractiondocument-analysis

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “character-level pdf parsing”

  • unPDFExtracts individual characters and metadata from PDF files using…
  • unstructured.pytesseractPython wrapper for Google's Tesseract OCR engine that extracts text…
  • docling-parseExtracts text, coordinates, and bitmap images from programmatic PDFs…

Give your agent the search over MCP, or paste the wish link into any chat.

More Text Processing packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyparsing Worth it
PyPI · Text Processing · released Jan 2026

pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.

Install it if you need to parse text or define grammars programmatically.

MITpure Python · 3.9+
412.7Mdownloads / mo
fonttools Worth it
PyPI · Text Processing · released May 2026

fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.

Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.

permissive licensepure Python · 3.10+
235.9Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
RapidFuzz Worth it
PyPI · Text Processing · released Apr 2026

RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.

Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.

MITcompiled wheel · 3.10+
184.2Mdownloads / mo
tinycss2 Worth it
PyPI · Text Processing · released Nov 2025

tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.

Install it if your project requires CSS tokenization or syntax manipulation.

BSD-3-Clausepure Python · 3.10+
113.2Mdownloads / mo

See also pdfplumber · pdftext · pymupdf · pdfminer · pypdfium2 · pdfminer.six · fillpdf · pymupdf-fonts · pdf-oxide · playa-pdf