pyocr
A Python wrapper for OCR engines (Tesseract, Cuneiform, etc)
Decision gist · record as of 2026-08-14
Yes, if you need a stable, lightweight OCR wrapper and have an OCR engine already installed. The low install friction and lack of vulnerabilities make it straightforward to add to a project. However, the dormant maintenance status means you should verify compatibility with your target OCR engine version and Python version before committing to it in new projects. The GPL-3.0-or-later license is a hard blocker for proprietary software.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires an external OCR engine (Tesseract, libtesseract, or Cuneiform) to be installed and available on the system; PyOCR is a wrapper only and will not function without one.
- Low install friction with a single runtime dependency (Pillow).
- Maintenance is dormant—last release was in 2023-09-17—so expect no active bug fixes or feature updates, though the package remains functional for stable OCR engine backends.
License · maintenance · safety
GPL-3.0-or-later (copyleft) — Licensed under GPL-3.0-or-later (copyleft). Any derivative work or distribution must also be open-source under a compatible GPL license, which may restrict use in proprietary or closed-source projects.
last release 2023-09-17 (1062 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 77,433 downloads/mo, #14,525 on PyPI
Alternatives
Verify before relying
pip install pyocr
import pyocr
import pyocr.builders
tools = pyocr.get_available_tools()
if tools:
tool = tools[0]
txt = tool.image_to_string(
Image.open('test.png'),
builder=pyocr.builders.TextBuilder()
)- Current compatibility with modern Python versions (3.4+ stated in description, but dormant since 2023)
- Compatibility with recent Tesseract or libtesseract versions
- Whether Windows and macOS support (noted as uncertain in description) has improved
- Whether Image.open() usage requires explicit Pillow import in user code
What it is and what it does
PyOCR is a Python wrapper that provides a unified interface to multiple OCR engines—Tesseract, libtesseract, and Cuneiform—allowing you to extract text and spatial information from images without writing engine-specific code. It handles image input through Pillow, supporting all formats Pillow supports (jpeg, png, gif, bmp, tiff and others), and returns results in multiple formats: plain text strings, word or line bounding boxes with pixel coordinates, or hOCR markup for structured document representation.
The package is designed for straightforward OCR workflows: detect available engines on the system, choose one, specify a language, and call image_to_string() with your image and desired output builder. It also supports orientation detection and digit-only recognition (Tesseract and libtesseract only) and can generate PDF files from images. However, the project is dormant—last updated 2023-09-17—so it receives no active maintenance, and compatibility with the latest OCR engine versions or Python releases is unverified.
Use it for
- Extract text from scanned documents or photographs for indexing or archival
- Build a document digitization pipeline that converts images to searchable text or hOCR
- Detect and correct page orientation in batch OCR workflows
- Extract structured data (word positions, confidence scores) for layout analysis or form processing
- Generate searchable PDF files from image scans using libtesseract
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need a stable, lightweight OCR wrapper and have an OCR engine already installed.
The low install friction and lack of vulnerabilities make it straightforward to add to a project. However, the dormant maintenance status means you should verify compatibility with your target OCR engine version and Python version before committing to it in new projects. The GPL-3.0-or-later license is a hard blocker for proprietary software.
Install
pyocr on PyPI
Before you install
Low install friction with a single runtime dependency (Pillow). Maintenance is dormant—last release was in 2023-09-17—so expect no active bug fixes or feature updates, though the package remains functional for stable OCR engine backends.
Requires an external OCR engine (Tesseract, libtesseract, or Cuneiform) to be installed and available on the system; PyOCR is a wrapper only and will not function without one.
License in practice
Licensed under GPL-3.0-or-later (copyleft). Any derivative work or distribution must also be open-source under a compatible GPL license, which may restrict use in proprietary or closed-source projects.
Quickstart
pip install pyocr
import pyocr
import pyocr.builders
tools = pyocr.get_available_tools()
if tools:
tool = tools[0]
txt = tool.image_to_string(
Image.open('test.png'),
builder=pyocr.builders.TextBuilder()
)
Verify before relying
- Current compatibility with modern Python versions (3.4+ stated in description, but dormant since 2023)
- Compatibility with recent Tesseract or libtesseract versions
- Whether Windows and macOS support (noted as uncertain in description) has improved
- Whether Image.open() usage requires explicit Pillow import in user code
Package facts
| License | GPL-3.0-or-later copyleft |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagePillow |
| Maintenance | Dormant 1,062 days since the last release |
| First released | |
| Downloads | 77,433 / month, #14,525 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: pyocr-0.8.5-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “ocr bounding boxes”
- pyocrPyOCR wraps multiple OCR engines (Tesseract, Cuneiform, libtesseract)…
- pytesseractPytesseract wraps Google's Tesseract-OCR engine to extract text from…
- unstructured.pytesseractPython wrapper for Google's Tesseract OCR engine that extracts text…
Give your agent the search over MCP, or paste the wish link into any chat.
More Text Processing packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.
Install it if you need to parse text or define grammars programmatically.
fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.
Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.
Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.
tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.
Install it if your project requires CSS tokenization or syntax manipulation.
See also pytesseract · unstructured.pytesseract · tesserocr · easyocr · ocrmac · keras-ocr · img2table · python-doctr · perceptron · amazon-textract-response-parser