$npx skillfedfor your agent

liteparse

Python bindings for LiteParse - fast, lightweight PDF and document parsing

Worth itPyPI Text ProcessingReleased Aug 2026384.0K downloads / moApache-2.0Platform wheel

Decision gist · record as of 2026-08-14

platform wheels — liteparse-2.12.0-cp310-cp310-manylinux_2_28_aarch64.whl · liteparse-2.12.0-cp310-cp310-manylinux_2_28_x86_64.whl · liteparse-2.12.0-cp310-cp310-win_amd64.whl
v2.12.0 · released 2026-08-13 · Python >=3.10

Yes. LiteParse is actively maintained, has no known vulnerabilities, supports modern Python versions, and solves a real problem (document parsing for LLM/RAG workflows) with a clean API. The permissive Apache-2.0 license and zero runtime Python dependencies keep friction low. Install it if you need to parse PDFs or mixed document formats into structured text or Markdown; skip it only if you have a simpler use case or a strong preference for pure-Python solutions.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later.
  • LibreOffice is required to parse Microsoft Office and OpenDocument formats (.docx, .xlsx, .pptx, .odt, .ods, .odp).
  • Medium install friction due to compiled wheels, but well-supported across Python 3.10–3.14 on Linux, macOS, and Windows.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; suitable for most production and proprietary projects.

last release 2026-08-13 (1 days) · last repo commit 2026-08-14 · 12,093 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 383,995 downloads/mo, #7,070 on PyPI

Verify before relying

pip install liteparse

from liteparse import LiteParse

parser = LiteParse()
result = parser.parse("document.pdf")
print(result.text)
print(f"Source document pages: {result.total_pages}")
  • Performance characteristics (parsing speed, memory usage) for large documents or batch operations.
  • Accuracy of OCR and Markdown reconstruction quality on complex or scanned documents.
  • Whether the `lit` CLI command is fully featured or a subset of the Python API.
  • Compatibility or integration with specific RAG frameworks or LLM pipelines.
Same gist for agents: .md · .json

What it is and what it does

LiteParse is a Python wrapper around a Rust-based PDF and document parser that extracts text while preserving spatial layout information. It supports multiple input formats (PDF, Office documents, images) and can output structured data as JSON, plain text, or reconstructed Markdown. The package includes built-in OCR via Tesseract, configurable image extraction, annotation and form-field parsing, and a CLI tool (`lit`) for command-line workflows.

Typical use cases include feeding documents into RAG pipelines and LLMs by converting PDFs to clean Markdown, routing documents to different processing pipelines based on complexity detection, and extracting structured data (images, links, annotations) from mixed document types. The package is designed to be lightweight and fast, with no runtime Python dependencies, though some features (like OCR and Office format support) require optional system libraries.

Use it for

  • Convert PDFs to Markdown for ingestion into LLM and RAG systems with preserved structure and links.
  • Extract and route documents based on complexity signals (scanned vs. digital text) to optimize processing cost.
  • Batch-parse large document collections via the CLI or Python API with configurable OCR and image extraction.
  • Recover structured data from PDFs including images, annotations, form fields, and tagged logical structure.
  • Generate PNG screenshots of specific document pages for preview or archival purposes.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

LiteParse is actively maintained, has no known vulnerabilities, supports modern Python versions, and solves a real problem (document parsing for LLM/RAG workflows) with a clean API. The permissive Apache-2.0 license and zero runtime Python dependencies keep friction low. Install it if you need to parse PDFs or mixed document formats into structured text or Markdown; skip it only if you have a simpler use case or a strong preference for pure-Python solutions.

Install

liteparse on PyPI

Before you install

Medium install friction due to compiled wheels, but well-supported across Python 3.10–3.14 on Linux, macOS, and Windows. Active maintenance with a recent release (1 day old) and 12093 repository stars suggest solid upkeep.

Requires Python 3.10 or later. LibreOffice is required to parse Microsoft Office and OpenDocument formats (.docx, .xlsx, .pptx, .odt, .ods, .odp).

License in practice

Apache-2.0 permissive license allows commercial and private use with minimal restrictions; suitable for most production and proprietary projects.

Quickstart

pip install liteparse

from liteparse import LiteParse

parser = LiteParse()
result = parser.parse("document.pdf")
print(result.text)
print(f"Source document pages: {result.total_pages}")

Verify before relying

  • Performance characteristics (parsing speed, memory usage) for large documents or batch operations.
  • Accuracy of OCR and Markdown reconstruction quality on complex or scanned documents.
  • Whether the `lit` CLI command is fully featured or a subset of the Python API.
  • Compatibility or integration with specific RAG frameworks or LLM pipelines.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionMedium. Platform-specific wheel
Runtime dependenciesNone
MaintenanceActively maintained 1 days since the last release
Last repo commit
First released
Downloads383,995 / month, #7,070 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: RustTopic :: Scientific/Engineering :: Information AnalysisTopic :: Text ProcessingTyping :: Typed

Evidence: liteparse-2.12.0-cp310-cp310-manylinux_2_28_aarch64.whl; liteparse-2.12.0-cp310-cp310-manylinux_2_28_x86_64.whl; liteparse-2.12.0-cp310-cp310-win_amd64.whl; liteparse-2.12.0-cp311-cp311-macosx_10_12_x86_64.whl; liteparse-2.12.0-cp311-cp311-macosx_11_0_arm64.whl; liteparse-2.12.0-cp311-cp311-manylinux_2_28_aarch64.whl; liteparse-2.12.0-cp311-cp311-manylinux_2_28_x86_64.whl; liteparse-2.12.0-cp311-cp311-win_amd64.whl; liteparse-2.12.0-cp312-cp312-macosx_10_12_x86_64.whl; liteparse-2.12.0-cp312-cp312-macosx_11_0_arm64.whl; liteparse-2.12.0-cp312-cp312-manylinux_2_28_aarch64.whl; liteparse-2.12.0-cp312-cp312-manylinux_2_28_x86_64.whl; liteparse-2.12.0-cp312-cp312-musllinux_1_2_x86_64.whl; liteparse-2.12.0-cp312-cp312-win_amd64.whl; liteparse-2.12.0-cp312-cp312-win_arm64.whl; liteparse-2.12.0-cp313-cp313-macosx_10_12_x86_64.whl; liteparse-2.12.0-cp313-cp313-macosx_11_0_arm64.whl; liteparse-2.12.0-cp313-cp313-manylinux_2_28_aarch64.whl; liteparse-2.12.0-cp313-cp313-manylinux_2_28_x86_64.whl; liteparse-2.12.0-cp313-cp313-win_amd64.whl

Tags

Capabilities
pdf parsing pythondocument text extractionocr pdf extractionmarkdown from pdfspatial text layoutlightweight pdf parserrag document processing
Topics
pdf-parsingocrrag-pipeline
PyPI keywords
pdfparsingocrdocumenttext-extraction

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “ocr pdf extraction”

  • liteparseLiteParse provides Python bindings for fast, lightweight PDF and…
  • marker-pdfMarker converts PDFs, images, and other document formats (PPTX, DOCX,…
  • pymupdf4llmConverts PDFs and documents into clean, structured Markdown, JSON, or…

Give your agent the search over MCP, or paste the wish link into any chat.

More Text Processing packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyparsing Worth it
PyPI · Text Processing · released Jan 2026

pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.

Install it if you need to parse text or define grammars programmatically.

MITpure Python · 3.9+
412.7Mdownloads / mo
fonttools Worth it
PyPI · Text Processing · released May 2026

fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.

Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.

permissive licensepure Python · 3.10+
235.9Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
RapidFuzz Worth it
PyPI · Text Processing · released Apr 2026

RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.

Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.

MITcompiled wheel · 3.10+
184.2Mdownloads / mo
tinycss2 Worth it
PyPI · Text Processing · released Nov 2025

tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.

Install it if your project requires CSS tokenization or syntax manipulation.

BSD-3-Clausepure Python · 3.10+
113.2Mdownloads / mo

See also pymupdf4llm · opendataloader-pdf · mineru · llama-parse · unPDF · pymupdf · pdf-oxide · pdftotext · amazon-textract-caller · llama-cloud