presidio-image-redactor
Presidio image redactor package
Decision gist · record as of 2026-08-14
Yes, if you need to redact PII text from images or DICOM files. The package is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and integrates with standard image libraries. The main gotcha is the Tesseract OCR system dependency—plan for that installation before deploying. Not suitable if you need metadata-level DICOM anonymization without a separate tool.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Tesseract OCR must be installed on your system (not via pip).
- Documentation recommends v5.2.0.
- The spacy model en_core_web_lg must be downloaded after package installation.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions—suitable for most production deployments.
last release 2026-07-22 (23 days) · last repo commit 2026-08-11 · 10,486 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 168,350 downloads/mo, #10,450 on PyPI
Alternatives
Verify before relying
pip install presidio-image-redactor
python -m spacy download en_core_web_lg
from presidio_image_redactor import ImageRedactorEngine
from pillow import Image
image = Image.open("image.png")
engine = ImageRedactorEngine()
redacted = engine.redact(image, (255, 192, 203))
redacted.save("redacted.png")- Actual redaction accuracy and false-positive rates for different image types and quality levels.
- Performance characteristics (processing time per image, memory usage) at scale.
- Whether DICOM metadata scrubbing is handled by this package or requires a separate tool as the docs suggest.
What it is and what it does
Presidio Image Redactor is a Python module that finds and masks PII text embedded in images and DICOM medical files. It combines OCR via pytesseract with named-entity recognition via spacy and presidio-analyzer to locate sensitive text, then covers those regions with a solid color or pattern of your choice.
The package handles both standard image formats and DICOM medical imaging files. For DICOM, it redacts only pixel-level burn-in text and does not modify metadata; the documentation explicitly recommends using a separate tool for metadata scrubbing. It supports batch processing from directories and can optionally return bounding boxes of redacted regions. The module runs as a Python library or as a Docker service with an HTTP API.
Use it for
- Anonymize photographs or scanned documents before sharing them in research or compliance workflows.
- Remove patient names and medical record numbers from DICOM images before publishing or archiving.
- Batch-redact burn-in text from medical imaging studies to prepare datasets for machine learning training.
- Build a privacy-preserving image pipeline in a web service using the HTTP API or Docker deployment.
- Detect and mask PII in user-uploaded images to enforce data minimization policies.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need to redact PII text from images or DICOM files.
The package is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and integrates with standard image libraries. The main gotcha is the Tesseract OCR system dependency—plan for that installation before deploying. Not suitable if you need metadata-level DICOM anonymization without a separate tool.
Install
presidio-image-redactor on PyPI
Before you install
Low friction install with a pure-Python wheel. Active maintenance (last commit 2026-08-11) and no known vulnerabilities. Requires Tesseract OCR as a system dependency; documentation specifies v5.2.0 for best results.
Tesseract OCR must be installed on your system (not via pip). Documentation recommends v5.2.0. The spacy model en_core_web_lg must be downloaded after package installation.
License in practice
MIT license permits commercial and private use with minimal restrictions—suitable for most production deployments.
Quickstart
pip install presidio-image-redactor
python -m spacy download en_core_web_lg
from presidio_image_redactor import ImageRedactorEngine
from pillow import Image
image = Image.open("image.png")
engine = ImageRedactorEngine()
redacted = engine.redact(image, (255, 192, 203))
redacted.save("redacted.png")
Verify before relying
- Actual redaction accuracy and false-positive rates for different image types and quality levels.
- Performance characteristics (processing time per image, memory usage) at scale.
- Whether DICOM metadata scrubbing is handled by this package or requires a separate tool as the docs suggest.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release <3.15,>=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 10 packagesazure-ai-formrecognizermatplotlibopencv-pythonpillowpresidio-analyzerpydicompypngpytesseractpython-gdcmspacy |
| Maintenance | Actively maintained 23 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 168,350 / month, #10,450 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14 |
Evidence: presidio_image_redactor-0.0.60-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “redact pii from images”
- presidio-image-redactorDetects and redacts personally identifiable information (PII) text in…
- scrubadubDetects and replaces personally identifiable information (names,…
- argus-redactDetects and redacts personally identifiable information (PII) in text…
Give your agent the search over MCP, or paste the wish link into any chat.
More Security packages
Provides Python bindings to the FreeDesktop.org Secret Service API for securely storing and retrieving passwords and secrets through GNOME Keyring, KWallet, or KeePassXC.
MSAL for Python handles OAuth2 and OpenID Connect authentication with Microsoft identity services, managing token acquisition, caching, and refresh for applications integrating with Microsoft Entra ID, Microsoft Accounts, and Azure AD B2C.
joserfc implements JOSE standards (JWS, JWE, JWK, JWT, and related RFCs) for signing, encrypting, and managing JSON-based cryptographic tokens in Python.
Authlib provides a complete implementation of OAuth 1.0, OAuth 2.0, and OpenID Connect 1.0 for building both authentication clients and servers, with built-in support for JWS, JWK, JWA, and JWT standards.
Provides low-level CFFI bindings to the official Argon2 password hashing algorithm for use by libraries and applications that need direct access to Argon2 without higher-level abstractions.
ADAL for Python authenticates applications with Azure Active Directory to obtain tokens for accessing Azure AD-protected resources.
Install only if maintaining existing code that already depends on it, and plan a migration.
See also argus-redact · presidio-analyzer · presidio-anonymizer · openmed · scrubadub · highdicom · pydicom · azure-ai-vision-imageanalysis · keras-ocr · dicom2nifti