{"categories":[{"label":"Python Modules","url":"https://skillfed.io/packages/category/software-development-libraries-python-modules/15"},{"label":"Text Processing","url":"https://skillfed.io/packages/category/text-processing/3"},{"label":"Information Analysis","url":"https://skillfed.io/packages/category/scientific-engineering-information-analysis/2"},{"label":"Office/Business","url":"https://skillfed.io/packages/category/office-business"},{"label":"General","url":"https://skillfed.io/packages/category/text-processing-general"},{"label":"Filters","url":"https://skillfed.io/packages/category/text-processing-filters"}],"enrichment":{"capability":"Extracts text, tables, images, and metadata from 91+ file formats including PDFs, Office documents, and images, with native async/await support and multiple OCR backends.","skillfed_tags":["document-extraction","async-io","ocr"],"use_cases":["Extract text and tables from mixed document types (PDFs, Word, Excel) in a single pipeline.","Process scanned documents with OCR using configurable backends (Tesseract, EasyOCR, PaddleOCR).","Batch-extract code structure and docstrings from repositories for documentation or analysis.","Build document indexing systems that preserve table structure and embedded metadata.","Generate embeddings from document content using ONNX Runtime models for semantic search."],"what_it_does":"Kreuzberg is a document extraction library that reads text, tables, images, and metadata from a wide range of file formats\u2014PDFs, Office documents, images, code files, and structured data formats. It wraps a Rust core with Python bindings and provides async/await support for non-blocking processing. The library handles 91+ formats across office documents, images with OCR, web markup, email, archives, and academic/scientific formats, plus code intelligence for 248 programming languages with structure extraction and docstring parsing.\n\nYou use it to automate document processing pipelines: feed it a file path, optionally configure extraction behavior (OCR backend, caching, quality processing), and receive structured output with content, tables, detected languages, and metadata. It's designed for batch processing, content indexing, and data extraction workflows where you need to handle heterogeneous document types without format-specific parsing code.","worth_installing":"Yes, if you need multi-format document extraction with async support and don't mind the Python 3.10+ requirement. The library is actively maintained (v4 LTS through end of 2026), has no known vulnerabilities, and covers a broad format range. Install friction is moderate due to compiled bindings, but wheels are available for common platforms. Best suited for document processing pipelines where format heterogeneity is a pain point."},"id":"kreuzberg","links":{"html":"https://skillfed.io/packages/kreuzberg","md":"https://skillfed.io/packages/kreuzberg.md","pypi":"https://pypi.org/project/kreuzberg/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-07-12","license_spdx":null,"license_treatment":"permissive","name":"kreuzberg","python_support":"supports_current","summary":"High-performance document intelligence library for Python. Extract text, metadata, and structured data from PDFs, Office documents, images, and 88+ formats. Powered by Rust core for 10-50x speed improvements."},"popularity":{"monthly_downloads":215433,"position":9401,"tier":"top_15000"},"security":{"n_vulnerabilities":0},"version":"4.10.2"}
