$npx skillfedfor your agent

pdfminer.six

PDF parser and analyzer

Worth itPyPI Text ProcessingReleased Jan 202672.5M downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — pdfminer_six-20260107-py3-none-any.whl
v20260107 · released 2026-01-07 · Python >=3.10 · 2 runtime deps: charset-normalizer, cryptography

Yes. pdfminer.six is a mature, actively maintained library with low install friction, permissive licensing, no known vulnerabilities, and broad community adoption. It is the standard choice for PDF text extraction in Python when you need reliable parsing of modern PDF specifications and layout analysis.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or newer.
  • Low friction installation with only two runtime dependencies (charset-normalizer, cryptography).
  • Actively maintained with recent commits and a large community presence (7017 stars).

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions—you may use, modify, and distribute the package freely provided you include the license notice.

last release 2026-01-07 (219 days) · last repo commit 2026-03-13 · 7,017 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 72,508,554 downloads/mo, #460 on PyPI

Verify before relying

pip install pdfminer.six

from pdfminer.high_level import extract_text

text = extract_text("example.pdf")
print(text)
  • Performance characteristics when processing large or complex PDF files
  • Accuracy of text extraction for PDFs with unusual encodings or corrupted sections
  • Whether the 'image' extra dependency adds significant install friction
Same gist for agents: .md · .json

What it is and what it does

pdfminer.six is a Python library for extracting and analyzing content from PDF documents. It parses the PDF source code directly to retrieve text, images, and layout metadata—including font, color, and exact position information. The library supports modern PDF specifications, CJK languages, multiple font types, various compression schemes, and encryption methods. It is built modularly, allowing you to replace components or implement custom interpreters for specialized use cases.

The package comes with a command-line tool (pdf2txt.py) for quick text extraction and a Python API for programmatic access. It requires Python 3.10 or newer and depends only on charset-normalizer and cryptography. An optional 'image' extra adds dependencies for embedded image extraction.

Use it for

  • Extract plain text from PDF files for indexing, search, or natural language processing pipelines
  • Analyze PDF layout and retrieve the precise location and formatting of text elements
  • Parse structured PDFs (forms, tables) to extract data programmatically
  • Convert PDF documents to alternative formats (HTML, hOCR) for downstream processing
  • Build document analysis tools that need to handle CJK text or multiple font encodings

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

pdfminer.six is a mature, actively maintained library with low install friction, permissive licensing, no known vulnerabilities, and broad community adoption. It is the standard choice for PDF text extraction in Python when you need reliable parsing of modern PDF specifications and layout analysis.

Install

pdfminer-six on PyPI

Before you install

Low friction installation with only two runtime dependencies (charset-normalizer, cryptography). Actively maintained with recent commits and a large community presence (7017 stars). Requires Python 3.10 or newer.

Requires Python 3.10 or newer.

License in practice

MIT license permits commercial and private use with minimal restrictions—you may use, modify, and distribute the package freely provided you include the license notice.

Quickstart

pip install pdfminer.six

from pdfminer.high_level import extract_text

text = extract_text("example.pdf")
print(text)

Verify before relying

  • Performance characteristics when processing large or complex PDF files
  • Accuracy of text extraction for PDFs with unusual encodings or corrupted sections
  • Whether the 'image' extra dependency adds significant install friction

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
2 packages
charset-normalizercryptography
MaintenanceActively maintained 219 days since the last release
Last repo commit
First released
Downloads72,508,554 / month, #460 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Science/ResearchProgramming Language :: PythonProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Text Processing

Evidence: pdfminer_six-20260107-py3-none-any.whl

Tags

Capabilities
pdf text extractionpdf parsing pythonextract text from pdfpdf layout analysispdf content extractionpdf parser librarypdf document analysis
Topics
pdf-parsingtext-extractiondocument-analysis
PyPI keywords
layout analysispdf converterpdf parsertext mining

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “extract text from pdf”

  • pdfminer.sixExtracts text, images, and layout information from PDF documents by…
  • pdftotextExtracts text from PDF files, including password-protected documents,…
  • pdftextExtracts plain text or structured blocks, lines, and spans from PDFs…

Give your agent the search over MCP, or paste the wish link into any chat.

More Text Processing packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyparsing Worth it
PyPI · Text Processing · released Jan 2026

pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.

Install it if you need to parse text or define grammars programmatically.

MITpure Python · 3.9+
412.7Mdownloads / mo
fonttools Worth it
PyPI · Text Processing · released May 2026

fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.

Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.

permissive licensepure Python · 3.10+
235.9Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
RapidFuzz Worth it
PyPI · Text Processing · released Apr 2026

RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.

Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.

MITcompiled wheel · 3.10+
184.2Mdownloads / mo
tinycss2 Worth it
PyPI · Text Processing · released Nov 2025

tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.

Install it if your project requires CSS tokenization or syntax manipulation.

BSD-3-Clausepure Python · 3.10+
113.2Mdownloads / mo

See also pdfminer · pdfplumber · playa-pdf · pdftext · unPDF · pdfid · pymupdf-layout · textract · pymupdf · eyecite