$npx skillfedfor your agent

img2table

img2table is a table identification and extraction Python Library for PDF and images, based on OpenCV image processing

With conditionsPyPI Information AnalysisReleased May 2026218.6K downloads / moMITPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — img2table-2.0.0-cp310-cp310-macosx_10_9_x86_64.whl · img2table-2.0.0-cp310-cp310-macosx_11_0_arm64.whl · img2table-2.0.0-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl
v2.0.0 · released 2026-05-10 · Python <3.15,>=3.10 · 5 runtime deps: numpy, pypdfium2, opencv-contrib-python, beautifulsoup4, xlsxwriter

Yes, if you need lightweight table extraction from images or PDFs without deep-learning overhead. The MIT license, active maintenance, and broad OCR integration make it practical for production use. Install friction is moderate due to compiled dependencies, but pre-built wheels are available. No known vulnerabilities. Suitable for CPU-constrained environments and batch document processing.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Tesseract-OCR must be installed separately on your system; alternatively, use one of the optional OCR backends (PaddleOCR, EasyOCR, etc.) via extras like `pip install img2table[paddle]`.
  • Medium install friction due to compiled dependencies (opencv-contrib-python, pypdfium2).
  • The package is actively maintained with recent releases and supports modern Python versions (3.10–3.13).

License · maintenance · safety

MIT (permissive) — MIT license is permissive; you may use, modify, and distribute this package freely with minimal restrictions, making it suitable for both open-source and commercial projects.

last release 2026-05-10 (96 days) · last repo commit 2026-07-12 · 890 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 218,583 downloads/mo, #9,338 on PyPI

Verify before relying

pip install img2table

from img2table.document import Image
from img2table.ocr import TesseractOCR

image = Image("path/to/image.png")
ocr = TesseractOCR()
tables = image.extract_tables(ocr=ocr)
  • Whether the package handles scanned documents with poor image quality or heavy distortion reliably
  • Performance characteristics when processing large PDF files or high-resolution images
  • Accuracy of table structure detection for complex layouts (nested tables, irregular cells)
Same gist for agents: .md · .json

What it is and what it does

img2table is a Python library for identifying and extracting tables from images and PDF files. It uses OpenCV-based image processing to detect table boundaries and cell structures without relying on neural networks, making it lighter and more CPU-friendly than deep-learning alternatives. The library supports common image formats and PDFs, handles complex table structures like merged cells, and returns results as structured objects with Pandas DataFrame representations.

The package integrates with multiple OCR services—Tesseract, PaddleOCR, EasyOCR, docTR, RapidOCR, Surya, Google Vision, and AWS Textract—to extract text from detected table cells. Extracted tables can be exported to Excel while preserving their original structure. It requires numpy, pypdfium2, opencv-contrib-python, beautifulsoup4, and xlsxwriter as runtime dependencies.

Use it for

  • Batch extract tables from scanned documents or screenshots for data analysis without manual re-entry
  • Convert PDF reports containing tabular data into structured Excel files for downstream processing
  • Automate table detection in document pipelines where neural network overhead is impractical
  • Parse invoice or receipt images to extract line-item tables for accounting systems
  • Build document ingestion workflows that preserve table formatting when converting to structured formats

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need lightweight table extraction from images or PDFs without deep-learning overhead.

The MIT license, active maintenance, and broad OCR integration make it practical for production use. Install friction is moderate due to compiled dependencies, but pre-built wheels are available. No known vulnerabilities. Suitable for CPU-constrained environments and batch document processing.

Install

img2table on PyPI

Before you install

Medium install friction due to compiled dependencies (opencv-contrib-python, pypdfium2). The package is actively maintained with recent releases and supports modern Python versions (3.10–3.13). Pre-built wheels available for common platforms reduce friction.

Tesseract-OCR must be installed separately on your system; alternatively, use one of the optional OCR backends (PaddleOCR, EasyOCR, etc.) via extras like `pip install img2table[paddle]`.

License in practice

MIT license is permissive; you may use, modify, and distribute this package freely with minimal restrictions, making it suitable for both open-source and commercial projects.

Quickstart

pip install img2table

from img2table.document import Image
from img2table.ocr import TesseractOCR

image = Image("path/to/image.png")
ocr = TesseractOCR()
tables = image.extract_tables(ocr=ocr)

Verify before relying

  • Whether the package handles scanned documents with poor image quality or heavy distortion reliably
  • Performance characteristics when processing large PDF files or high-resolution images
  • Accuracy of table structure detection for complex layouts (nested tables, irregular cells)

Package facts

LicenseMIT permissive
Python supportSupports the current Python release <3.15,>=3.10
Install frictionMedium. Platform-specific wheel
Runtime dependencies
5 packages
numpypypdfium2opencv-contrib-pythonbeautifulsoup4xlsxwriter
MaintenanceActively maintained 96 days since the last release
Last repo commit
First released
Downloads218,583 / month, #9,338 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Operating System :: OS IndependentProgramming Language :: Python :: 3 :: Only

Evidence: img2table-2.0.0-cp310-cp310-macosx_10_9_x86_64.whl; img2table-2.0.0-cp310-cp310-macosx_11_0_arm64.whl; img2table-2.0.0-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; img2table-2.0.0-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; img2table-2.0.0-cp310-cp310-win32.whl; img2table-2.0.0-cp310-cp310-win_amd64.whl; img2table-2.0.0-cp311-cp311-macosx_10_9_x86_64.whl; img2table-2.0.0-cp311-cp311-macosx_11_0_arm64.whl; img2table-2.0.0-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; img2table-2.0.0-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; img2table-2.0.0-cp311-cp311-win32.whl; img2table-2.0.0-cp311-cp311-win_amd64.whl; img2table-2.0.0-cp312-cp312-macosx_10_13_x86_64.whl; img2table-2.0.0-cp312-cp312-macosx_11_0_arm64.whl; img2table-2.0.0-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; img2table-2.0.0-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; img2table-2.0.0-cp312-cp312-win32.whl; img2table-2.0.0-cp312-cp312-win_amd64.whl; img2table-2.0.0-cp313-cp313-macosx_10_13_x86_64.whl; img2table-2.0.0-cp313-cp313-macosx_11_0_arm64.whl

Tags

Capabilities
table extraction from imagespdf table detectionextract tables from documentstable identification opencvocr table parsingimage to structured datadocument table recognition
Topics
document-processingtable-extractionocr-integration

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “table extraction from images”

  • img2tableIdentifies and extracts tables from images and PDF files using…
  • kreuzbergExtracts text, tables, images, and metadata from 91+ file formats…
  • amazon-textract-textractorTextractor wraps Amazon Textract APIs to extract text, tables, forms,…

Give your agent the search over MCP, or paste the wish link into any chat.

More Information Analysis packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyarrow Worth it
PyPI · Information Analysis · released Aug 2026

pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.

Apache-2.0compiled wheel · 3.10+
432.9Mdownloads / mo
networkx Worth it
PyPI · Python Modules · released Dec 2025

NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.

BSD-3-Clausepure Python
290.9Mdownloads / mo
snowflake-connector-python Worth it
PyPI · Software Development · released Aug 2026

Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.

Apache-2.0compiled wheel · 3.10+
193.6Mdownloads / mo
contourpy Worth it
PyPI · Information Analysis · released Jul 2025

ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.

BSD-3-Clausecompiled wheel · 3.11+
191.2Mdownloads / mo
snowflake-snowpark-python Worth it
PyPI · Software Development · released Jul 2026

Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.

Install it if you use Snowflake and want to process data without moving it to your application layer.

Apache-2.0pure Python
100.7Mdownloads / mo

See also kreuzberg · surya-ocr · camelot-py · unstructured.pytesseract · marker-pdf · rapidocr · paddleocr · amazon-textract-textractor · pytesseract · python-doctr