$npx skillfedfor your agent

pdfminer

PDF parser and analyzer

SkipPyPI Text ProcessingReleased Nov 2019229.5K downloads / moMITSource build

Decision gist · record as of 2026-08-14

sdist only — pdfminer-20191125.tar.gz · builds from source
v20191125 · released 2019-11-25 · Python >=3.6

No—not recommended for new projects. The package is abandoned (last release 2019-11-25, repository archived 2022), receives no maintenance or security updates, and has high install friction due to source-only distribution. Use pdfminer.six instead, which is actively maintained and provides the same core functionality with ongoing support.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.6 or above; Python 2 is not supported.
  • Source distribution requires build tools to compile.
  • High install friction due to source-only distribution (pdfminer-20191125.tar.gz).

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions, requiring only attribution and inclusion of the license text.

last release 2019-11-25 (2454 days) · last repo commit 2022-12-07 · 5,275 stars · archived

0 known vulnerabilities (OSV.dev, 2026-08-14) · 229,484 downloads/mo, #9,131 on PyPI

Verify before relying

pip install pdfminer
python -m pdfminer.six samples/simple1.pdf
# or via command line:
pdf2txt.py samples/simple1.pdf
  • Whether the archived repository still accepts security patches or community contributions
  • Current compatibility with modern PDF specifications beyond PDF-1.7
  • Performance characteristics on large or complex PDF files
Same gist for agents: .md · .json

What it is and what it does

PDFMiner is a pure Python PDF parser and text extraction tool that reads PDF documents and extracts rendered text along with precise layout metadata—font names, sizes, positions, and writing direction. It performs automatic layout analysis to reconstruct document structure and can output results as plain text, HTML, XML, or tagged content. It handles encrypted PDFs (RC4 and AES), multiple font types (Type1, TrueType, Type3, CID), and CJK languages with vertical writing support.

The package provides both a programmatic API for integration into Python applications and command-line tools (pdf2txt.py for extraction, dumppdf.py for debugging). However, it is no longer maintained—the repository was archived in 2022 with the last commit in December of that year, and no updates have been released since November 2019. While it remains functional for basic PDF text extraction tasks, it receives no security updates or bug fixes.

Use it for

  • Extract text and position data from PDF documents for document processing or data mining workflows
  • Convert PDFs to HTML or XML for downstream analysis or republishing
  • Debug PDF structure and internal content using dumppdf.py for troubleshooting
  • Parse encrypted PDFs with password protection to access restricted content
  • Analyze document layout and reconstruct reading order from complex multi-column or figure-heavy PDFs

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Skip

No—not recommended for new projects.

The package is abandoned (last release 2019-11-25, repository archived 2022), receives no maintenance or security updates, and has high install friction due to source-only distribution. Use pdfminer.six instead, which is actively maintained and provides the same core functionality with ongoing support.

Install

pdfminer on PyPI

Before you install

High install friction due to source-only distribution (pdfminer-20191125.tar.gz). The project is archived and abandoned as of 2022-12-07, with no maintenance since version 20191125 released 2019-11-25. Consider pdfminer.six if ongoing support is needed.

Requires Python 3.6 or above; Python 2 is not supported. Source distribution requires build tools to compile.

License in practice

MIT license permits commercial and private use with minimal restrictions, requiring only attribution and inclusion of the license text.

Quickstart

pip install pdfminer
python -m pdfminer.six samples/simple1.pdf
# or via command line:
pdf2txt.py samples/simple1.pdf

Verify before relying

  • Whether the archived repository still accepts security patches or community contributions
  • Current compatibility with modern PDF specifications beyond PDF-1.7
  • Performance characteristics on large or complex PDF files

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.6
Install frictionHigh. Source build required
Runtime dependenciesNone
MaintenanceAbandoned 2,454 days since the last release
Last repo commit repository archived
First released
Downloads229,484 / month, #9,131 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseTopic :: Text Processing

Evidence: pdfminer-20191125.tar.gz

Tags

Capabilities
pdf text extractionpdf parser pythonextract text from pdfpdf layout analysispdf to text converterpdf content analysispdf document parsing
Topics
pdf-parsingabandoned
PyPI keywords
pdf parserpdf converterlayout analysistext mining

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “pdf layout analysis”

  • pdfminerExtracts text and layout information from PDF documents, including…
  • pdfminer.sixExtracts text, images, and layout information from PDF documents by…
  • pymupdf-layoutPyMuPDF Layout analyzes PDF structure and content using Graph Neural…

Give your agent the search over MCP, or paste the wish link into any chat.

More Text Processing packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyparsing Worth it
PyPI · Text Processing · released Jan 2026

pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.

Install it if you need to parse text or define grammars programmatically.

MITpure Python · 3.9+
412.7Mdownloads / mo
fonttools Worth it
PyPI · Text Processing · released May 2026

fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.

Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.

permissive licensepure Python · 3.10+
235.9Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
RapidFuzz Worth it
PyPI · Text Processing · released Apr 2026

RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.

Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.

MITcompiled wheel · 3.10+
184.2Mdownloads / mo
tinycss2 Worth it
PyPI · Text Processing · released Nov 2025

tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.

Install it if your project requires CSS tokenization or syntax manipulation.

BSD-3-Clausepure Python · 3.10+
113.2Mdownloads / mo

See also pdfminer.six · playa-pdf · unPDF · pdftext · textract · pdfplumber · pdftotext · pymupdf · pymupdf-layout · drafthorse