tabula-py
Simple wrapper for tabula-java, read tables from PDF into DataFrame
Decision gist · record as of 2026-08-14
Yes, with conditions. The package is stable, marked Production/Stable, and carries no security vulnerabilities. However, dormant development since 2024-10-17 means new features or breaking-change fixes are unlikely. Install only if you have Java 8+ available on your system and need straightforward PDF table extraction; if you require active development or support for cutting-edge PDF formats, evaluate alternatives or plan for maintenance yourself.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Java 8 or later installed and available on the system PATH; Python 3.9 or higher.
- Low install friction with a pure-Python wheel.
- Maintenance is dormant with no releases since 2024-10-17, though the repository remains active with a recent commit on 2024-12-05 and 2315 stars.
License · maintenance · safety
permissive license (permissive) — MIT license permits unrestricted use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.
last release 2024-10-17 (666 days) · last repo commit 2024-12-05 · 2,315 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 5,658,034 downloads/mo, #2,060 on PyPI
Alternatives
Verify before relying
pip install tabula-py
import tabula
dfs = tabula.read_pdf("test.pdf", pages='all')
tabula.convert_into("test.pdf", "output.csv", output_format="csv", pages='all')- Stability and compatibility of the Java backend with recent Java versions and modern PDF formats.
- Performance characteristics when processing large PDFs or batch operations.
- Accuracy of table detection and extraction for complex or non-standard PDF layouts.
What it is and what it does
tabula-py is a Python wrapper around tabula-java that extracts structured table data from PDF files. It reads tables directly into pandas DataFrames, making the extracted data immediately usable for analysis and manipulation. The package also supports batch conversion of PDFs to CSV, TSV, or JSON formats. It depends on pandas, numpy, and distro, and requires a Java 8+ runtime on the system—the Java dependency is the primary installation consideration.
The package is stable and widely used (top 5000 on PyPI with monthly downloads of 5658034), though development is dormant with no releases since 2024-10-17. It supports Python 3.9, 3.10, 3.11, 3.12, and 3.13 and carries no known security vulnerabilities. The MIT license places no restrictions on use.
Use it for
- Extract financial or statistical tables from PDF reports into DataFrames for analysis.
- Batch convert a directory of PDFs containing tabular data into CSV files for data pipeline ingestion.
- Automate scraping of structured data from archived or published PDF documents.
- Convert PDF-based forms or invoices with table layouts into machine-readable formats.
- Build data collection workflows that read tables from remote PDFs via URL.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with conditions.
The package is stable, marked Production/Stable, and carries no security vulnerabilities. However, dormant development since 2024-10-17 means new features or breaking-change fixes are unlikely. Install only if you have Java 8+ available on your system and need straightforward PDF table extraction; if you require active development or support for cutting-edge PDF formats, evaluate alternatives or plan for maintenance yourself.
Install
tabula-py on PyPI
Before you install
Low install friction with a pure-Python wheel. Maintenance is dormant with no releases since 2024-10-17, though the repository remains active with a recent commit on 2024-12-05 and 2315 stars. The package is marked Production/Stable and supports current Python versions.
Requires Java 8 or later installed and available on the system PATH; Python 3.9 or higher.
License in practice
MIT license permits unrestricted use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.
Quickstart
pip install tabula-py
import tabula
dfs = tabula.read_pdf("test.pdf", pages='all')
tabula.convert_into("test.pdf", "output.csv", output_format="csv", pages='all')
Verify before relying
- Stability and compatibility of the Java backend with recent Java versions and modern PDF formats.
- Performance characteristics when processing large PDFs or batch operations.
- Accuracy of table detection and extraction for complex or non-standard PDF layouts.
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packagespandasnumpydistro |
| Maintenance | Dormant 666 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 5,658,034 / month, #2,060 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.9Topic :: Text Processing :: General |
Evidence: tabula_py-2.10.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “extract tables from pdf”
- tabula-pyExtracts tables from PDF files and converts them into pandas…
- camelot-pyExtracts tables from PDFs into pandas DataFrames using multiple…
- img2tableIdentifies and extracts tables from images and PDF files using…
Give your agent the search over MCP, or paste the wish link into any chat.
More General packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Sphinx generates professional documentation from reStructuredText source files, producing HTML, PDF, EPUB, and other formats with automatic cross-references, code highlighting, and hierarchical navigation.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
NLTK is a Python library for natural language processing tasks including tokenization, parsing, tagging, and linguistic analysis, with built-in datasets and educational resources.
Install it if you need foundational NLP tools, linguistic datasets, or are learning the field; consider specialized libraries (spaCy, transformers) if you need…
Converts numbers, dates, times, and file sizes into human-readable text formats, with support for fuzzy durations like "3 minutes ago" and localization to multiple languages.
Install it if you need to display human-readable numbers, durations, or sizes to end users.
See also camelot-py · tablib · pantab · unPDF · aspose-cells · pdftext · pdftotext · img2table · itables