kreuzberg
High-performance document intelligence library for Python. Extract text, metadata, and structured data from PDFs, Office documents, images, and 88+ formats. Powered by Rust core for 10-50x speed improvements.
Decision gist · record as of 2026-08-14
Yes, if you need multi-format document extraction with async support and don't mind the Python 3.10+ requirement. The library is actively maintained (v4 LTS through end of 2026), has no known vulnerabilities, and covers a broad format range. Install friction is moderate due to compiled bindings, but wheels are available for common platforms. Best suited for document processing pipelines where format heterogeneity is a pain point.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10+; optional Tesseract OCR or ONNX Runtime 1.22.x for full feature support.
- Medium install friction due to compiled Rust bindings; wheels available for modern Python (3.10+) on common platforms.
- Active maintenance with recent releases; v4 is long-term-support line receiving critical fixes through end of 2026.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; suitable for most projects without licensing concerns.
last release 2026-07-12 (33 days) · last repo commit 2026-08-13 · 9 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 215,433 downloads/mo, #9,401 on PyPI
Alternatives
Verify before relying
pip install kreuzberg
import asyncio
from kreuzberg import extract_file
async def main():
result = await extract_file("document.pdf")
print(result.content)
asyncio.run(main())- Performance improvement claims (10-50x speed) relative to other Python document extraction libraries.
- Actual number of supported programming languages for code intelligence (stated as 248).
- Completeness of OCR backend support (Tesseract, EasyOCR, PaddleOCR) and their integration maturity.
What it is and what it does
Kreuzberg is a document extraction library that reads text, tables, images, and metadata from a wide range of file formats—PDFs, Office documents, images, code files, and structured data formats. It wraps a Rust core with Python bindings and provides async/await support for non-blocking processing. The library handles 91+ formats across office documents, images with OCR, web markup, email, archives, and academic/scientific formats, plus code intelligence for 248 programming languages with structure extraction and docstring parsing.
You use it to automate document processing pipelines: feed it a file path, optionally configure extraction behavior (OCR backend, caching, quality processing), and receive structured output with content, tables, detected languages, and metadata. It's designed for batch processing, content indexing, and data extraction workflows where you need to handle heterogeneous document types without format-specific parsing code.
Use it for
- Extract text and tables from mixed document types (PDFs, Word, Excel) in a single pipeline.
- Process scanned documents with OCR using configurable backends (Tesseract, EasyOCR, PaddleOCR).
- Batch-extract code structure and docstrings from repositories for documentation or analysis.
- Build document indexing systems that preserve table structure and embedded metadata.
- Generate embeddings from document content using ONNX Runtime models for semantic search.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need multi-format document extraction with async support and don't mind the Python 3.10+ requirement.
The library is actively maintained (v4 LTS through end of 2026), has no known vulnerabilities, and covers a broad format range. Install friction is moderate due to compiled bindings, but wheels are available for common platforms. Best suited for document processing pipelines where format heterogeneity is a pain point.
Install
kreuzberg on PyPI
Before you install
Medium install friction due to compiled Rust bindings; wheels available for modern Python (3.10+) on common platforms. Active maintenance with recent releases; v4 is long-term-support line receiving critical fixes through end of 2026.
Requires Python 3.10+; optional Tesseract OCR or ONNX Runtime 1.22.x for full feature support.
License in practice
MIT license permits commercial and private use with minimal restrictions; suitable for most projects without licensing concerns.
Quickstart
pip install kreuzberg
import asyncio
from kreuzberg import extract_file
async def main():
result = await extract_file("document.pdf")
print(result.content)
asyncio.run(main())
Verify before relying
- Performance improvement claims (10-50x speed) relative to other Python document extraction libraries.
- Actual number of supported programming languages for code intelligence (stated as 248).
- Completeness of OCR backend support (Tesseract, EasyOCR, PaddleOCR) and their integration maturity.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 33 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 215,433 / month, #9,401 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: Implementation :: CPythonProgramming Language :: RustTopic :: Office/BusinessTopic :: Scientific/Engineering :: Information AnalysisTopic :: Software Development :: Libraries :: Python ModulesTopic :: Text ProcessingTopic :: Text Processing :: FiltersTopic :: Text Processing :: GeneralTyping :: Typed |
Evidence: kreuzberg-4.10.2-cp310-abi3-macosx_14_0_arm64.whl; kreuzberg-4.10.2-cp310-abi3-manylinux_2_28_aarch64.whl; kreuzberg-4.10.2-cp310-abi3-manylinux_2_28_x86_64.whl; kreuzberg-4.10.2-cp310-abi3-win_amd64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “office document extraction”
- kreuzbergExtracts text, tables, images, and metadata from 91+ file formats…
- pymupdfproPyMuPDF Pro extends PyMuPDF with commercial Office document…
- pymupdfPyMuPDF extracts, renders, converts, and manipulates PDF and other…
Give your agent the search over MCP, or paste the wish link into any chat.
More Python Modules packages
Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.
Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.
Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.
PyYAML parses and emits YAML 1.1 data format, enabling serialization and deserialization of configuration files and Python objects to and from human-readable YAML text.
Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.
Provides reusable metadata objects for use with PEP-593 `typing.Annotated` to express common constraints like bounds, collection sizes, and predicates on types.
Install it if you use or build libraries that need to express type constraints in a standardized, inspectable way—or if you want to annotate your own types with…
Provides runtime tools to inspect and introspect Python type annotations, enabling programmatic examination of type hints at execution time.
See also img2table · pymupdf · aurelio-sdk · pdf-oxide · paddleocr · textract · unPDF · marker-pdf · tika · factur-x