liteparse
Python bindings for LiteParse - fast, lightweight PDF and document parsing
Decision gist · record as of 2026-08-14
Yes. LiteParse is actively maintained, has no known vulnerabilities, supports modern Python versions, and solves a real problem (document parsing for LLM/RAG workflows) with a clean API. The permissive Apache-2.0 license and zero runtime Python dependencies keep friction low. Install it if you need to parse PDFs or mixed document formats into structured text or Markdown; skip it only if you have a simpler use case or a strong preference for pure-Python solutions.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- LibreOffice is required to parse Microsoft Office and OpenDocument formats (.docx, .xlsx, .pptx, .odt, .ods, .odp).
- Medium install friction due to compiled wheels, but well-supported across Python 3.10–3.14 on Linux, macOS, and Windows.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; suitable for most production and proprietary projects.
last release 2026-08-13 (1 days) · last repo commit 2026-08-14 · 12,093 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 383,995 downloads/mo, #7,070 on PyPI
Alternatives
Verify before relying
pip install liteparse
from liteparse import LiteParse
parser = LiteParse()
result = parser.parse("document.pdf")
print(result.text)
print(f"Source document pages: {result.total_pages}")- Performance characteristics (parsing speed, memory usage) for large documents or batch operations.
- Accuracy of OCR and Markdown reconstruction quality on complex or scanned documents.
- Whether the `lit` CLI command is fully featured or a subset of the Python API.
- Compatibility or integration with specific RAG frameworks or LLM pipelines.
What it is and what it does
LiteParse is a Python wrapper around a Rust-based PDF and document parser that extracts text while preserving spatial layout information. It supports multiple input formats (PDF, Office documents, images) and can output structured data as JSON, plain text, or reconstructed Markdown. The package includes built-in OCR via Tesseract, configurable image extraction, annotation and form-field parsing, and a CLI tool (`lit`) for command-line workflows.
Typical use cases include feeding documents into RAG pipelines and LLMs by converting PDFs to clean Markdown, routing documents to different processing pipelines based on complexity detection, and extracting structured data (images, links, annotations) from mixed document types. The package is designed to be lightweight and fast, with no runtime Python dependencies, though some features (like OCR and Office format support) require optional system libraries.
Use it for
- Convert PDFs to Markdown for ingestion into LLM and RAG systems with preserved structure and links.
- Extract and route documents based on complexity signals (scanned vs. digital text) to optimize processing cost.
- Batch-parse large document collections via the CLI or Python API with configurable OCR and image extraction.
- Recover structured data from PDFs including images, annotations, form fields, and tagged logical structure.
- Generate PNG screenshots of specific document pages for preview or archival purposes.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
LiteParse is actively maintained, has no known vulnerabilities, supports modern Python versions, and solves a real problem (document parsing for LLM/RAG workflows) with a clean API. The permissive Apache-2.0 license and zero runtime Python dependencies keep friction low. Install it if you need to parse PDFs or mixed document formats into structured text or Markdown; skip it only if you have a simpler use case or a strong preference for pure-Python solutions.
Install
liteparse on PyPI
Before you install
Medium install friction due to compiled wheels, but well-supported across Python 3.10–3.14 on Linux, macOS, and Windows. Active maintenance with a recent release (1 day old) and 12093 repository stars suggest solid upkeep.
Requires Python 3.10 or later. LibreOffice is required to parse Microsoft Office and OpenDocument formats (.docx, .xlsx, .pptx, .odt, .ods, .odp).
License in practice
Apache-2.0 permissive license allows commercial and private use with minimal restrictions; suitable for most production and proprietary projects.
Quickstart
pip install liteparse
from liteparse import LiteParse
parser = LiteParse()
result = parser.parse("document.pdf")
print(result.text)
print(f"Source document pages: {result.total_pages}")
Verify before relying
- Performance characteristics (parsing speed, memory usage) for large documents or batch operations.
- Accuracy of OCR and Markdown reconstruction quality on complex or scanned documents.
- Whether the `lit` CLI command is fully featured or a subset of the Python API.
- Compatibility or integration with specific RAG frameworks or LLM pipelines.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 1 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 383,995 / month, #7,070 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: RustTopic :: Scientific/Engineering :: Information AnalysisTopic :: Text ProcessingTyping :: Typed |
Evidence: liteparse-2.12.0-cp310-cp310-manylinux_2_28_aarch64.whl; liteparse-2.12.0-cp310-cp310-manylinux_2_28_x86_64.whl; liteparse-2.12.0-cp310-cp310-win_amd64.whl; liteparse-2.12.0-cp311-cp311-macosx_10_12_x86_64.whl; liteparse-2.12.0-cp311-cp311-macosx_11_0_arm64.whl; liteparse-2.12.0-cp311-cp311-manylinux_2_28_aarch64.whl; liteparse-2.12.0-cp311-cp311-manylinux_2_28_x86_64.whl; liteparse-2.12.0-cp311-cp311-win_amd64.whl; liteparse-2.12.0-cp312-cp312-macosx_10_12_x86_64.whl; liteparse-2.12.0-cp312-cp312-macosx_11_0_arm64.whl; liteparse-2.12.0-cp312-cp312-manylinux_2_28_aarch64.whl; liteparse-2.12.0-cp312-cp312-manylinux_2_28_x86_64.whl; liteparse-2.12.0-cp312-cp312-musllinux_1_2_x86_64.whl; liteparse-2.12.0-cp312-cp312-win_amd64.whl; liteparse-2.12.0-cp312-cp312-win_arm64.whl; liteparse-2.12.0-cp313-cp313-macosx_10_12_x86_64.whl; liteparse-2.12.0-cp313-cp313-macosx_11_0_arm64.whl; liteparse-2.12.0-cp313-cp313-manylinux_2_28_aarch64.whl; liteparse-2.12.0-cp313-cp313-manylinux_2_28_x86_64.whl; liteparse-2.12.0-cp313-cp313-win_amd64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “ocr pdf extraction”
- liteparseLiteParse provides Python bindings for fast, lightweight PDF and…
- marker-pdfMarker converts PDFs, images, and other document formats (PPTX, DOCX,…
- pymupdf4llmConverts PDFs and documents into clean, structured Markdown, JSON, or…
Give your agent the search over MCP, or paste the wish link into any chat.
More Text Processing packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.
Install it if you need to parse text or define grammars programmatically.
fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.
Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.
Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.
tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.
Install it if your project requires CSS tokenization or syntax manipulation.
See also pymupdf4llm · opendataloader-pdf · mineru · llama-parse · unPDF · pymupdf · pdf-oxide · pdftotext · amazon-textract-caller · llama-cloud