pdfminer.six
PDF parser and analyzer
Decision gist · record as of 2026-08-14
Yes. pdfminer.six is a mature, actively maintained library with low install friction, permissive licensing, no known vulnerabilities, and broad community adoption. It is the standard choice for PDF text extraction in Python when you need reliable parsing of modern PDF specifications and layout analysis.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or newer.
- Low friction installation with only two runtime dependencies (charset-normalizer, cryptography).
- Actively maintained with recent commits and a large community presence (7017 stars).
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions—you may use, modify, and distribute the package freely provided you include the license notice.
last release 2026-01-07 (219 days) · last repo commit 2026-03-13 · 7,017 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 72,508,554 downloads/mo, #460 on PyPI
Alternatives
Verify before relying
pip install pdfminer.six
from pdfminer.high_level import extract_text
text = extract_text("example.pdf")
print(text)- Performance characteristics when processing large or complex PDF files
- Accuracy of text extraction for PDFs with unusual encodings or corrupted sections
- Whether the 'image' extra dependency adds significant install friction
What it is and what it does
pdfminer.six is a Python library for extracting and analyzing content from PDF documents. It parses the PDF source code directly to retrieve text, images, and layout metadata—including font, color, and exact position information. The library supports modern PDF specifications, CJK languages, multiple font types, various compression schemes, and encryption methods. It is built modularly, allowing you to replace components or implement custom interpreters for specialized use cases.
The package comes with a command-line tool (pdf2txt.py) for quick text extraction and a Python API for programmatic access. It requires Python 3.10 or newer and depends only on charset-normalizer and cryptography. An optional 'image' extra adds dependencies for embedded image extraction.
Use it for
- Extract plain text from PDF files for indexing, search, or natural language processing pipelines
- Analyze PDF layout and retrieve the precise location and formatting of text elements
- Parse structured PDFs (forms, tables) to extract data programmatically
- Convert PDF documents to alternative formats (HTML, hOCR) for downstream processing
- Build document analysis tools that need to handle CJK text or multiple font encodings
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
pdfminer.six is a mature, actively maintained library with low install friction, permissive licensing, no known vulnerabilities, and broad community adoption. It is the standard choice for PDF text extraction in Python when you need reliable parsing of modern PDF specifications and layout analysis.
Install
pdfminer-six on PyPI
Before you install
Low friction installation with only two runtime dependencies (charset-normalizer, cryptography). Actively maintained with recent commits and a large community presence (7017 stars). Requires Python 3.10 or newer.
Requires Python 3.10 or newer.
License in practice
MIT license permits commercial and private use with minimal restrictions—you may use, modify, and distribute the package freely provided you include the license notice.
Quickstart
pip install pdfminer.six
from pdfminer.high_level import extract_text
text = extract_text("example.pdf")
print(text)
Verify before relying
- Performance characteristics when processing large or complex PDF files
- Accuracy of text extraction for PDFs with unusual encodings or corrupted sections
- Whether the 'image' extra dependency adds significant install friction
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagescharset-normalizercryptography |
| Maintenance | Actively maintained 219 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 72,508,554 / month, #460 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Science/ResearchProgramming Language :: PythonProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Text Processing |
Evidence: pdfminer_six-20260107-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “extract text from pdf”
- pdfminer.sixExtracts text, images, and layout information from PDF documents by…
- pdftotextExtracts text from PDF files, including password-protected documents,…
- pdftextExtracts plain text or structured blocks, lines, and spans from PDFs…
Give your agent the search over MCP, or paste the wish link into any chat.
More Text Processing packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.
Install it if you need to parse text or define grammars programmatically.
fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.
Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.
Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.
tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.
Install it if your project requires CSS tokenization or syntax manipulation.
See also pdfminer · pdfplumber · playa-pdf · pdftext · unPDF · pdfid · pymupdf-layout · textract · pymupdf · eyecite