unstructured-ingest
Local ETL data pipeline to get data RAG ready
Decision gist · record as of 2026-08-14
Yes, if you need to prepare unstructured documents for AI/RAG workflows locally. The package is actively maintained, has low install friction, carries a permissive Apache-2.0 license, and depends on stable, lightweight libraries. No known vulnerabilities as of 2026-08-14. The main uncertainty is whether its supported document formats and transformation capabilities match your specific use case.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.11 or later (supports 3.11, 3.12, 3.13)
- Low install friction with a pure-Python wheel distribution.
- Active maintenance with a release on 2026-08-14.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 is a permissive open-source license. You can use, modify, and distribute this package freely in commercial and private projects, provided you include a copy of the license and state significant changes.
last release 2026-08-14 (0 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 620,054 downloads/mo, #5,727 on PyPI
Alternatives
Verify before relying
pip install unstructured-ingest
from unstructured_ingest import ...
# See documentation for specific ingestion and transformation workflows- What document formats (PDF, Word, HTML, etc.) does the ingestion pipeline actually support?
- Does the package require external services or APIs, or does it run entirely locally?
- What is the typical performance or throughput for document ingestion and transformation?
What it is and what it does
Unstructured Ingest is a Python package that runs as a local ETL pipeline designed to take raw, unstructured documents and prepare them for use in AI systems, particularly retrieval-augmented generation (RAG) applications. It handles the work of ingesting documents and transforming them into clean, structured formats that downstream AI models can consume.
The package depends on a lean set of runtime libraries: pydantic for data validation, click for CLI support, tqdm for progress indication, opentelemetry-sdk for observability, ijson for JSON streaming, certifi for SSL certificates, and python-dateutil for date handling. It supports Python 3.11, 3.12, and 3.13, and is currently in Beta status with active maintenance.
Use it for
- Ingest a folder of documents and transform them into text chunks for a vector database.
- Build a preprocessing step in a RAG pipeline that normalizes documents from multiple sources.
- Extract structured content from unstructured documents before feeding them to a language model.
- Automate local document processing workflows without relying on external APIs.
- Prepare datasets by cleaning and standardizing raw document collections.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need to prepare unstructured documents for AI/RAG workflows locally.
The package is actively maintained, has low install friction, carries a permissive Apache-2.0 license, and depends on stable, lightweight libraries. No known vulnerabilities as of 2026-08-14. The main uncertainty is whether its supported document formats and transformation capabilities match your specific use case.
Install
unstructured-ingest on PyPI
Before you install
Low install friction with a pure-Python wheel distribution. Active maintenance with a release on 2026-08-14. Runtime dependencies are all well-established libraries, suggesting a stable, dependency-light setup.
Requires Python 3.11 or later (supports 3.11, 3.12, 3.13)
License in practice
Apache-2.0 is a permissive open-source license. You can use, modify, and distribute this package freely in commercial and private projects, provided you include a copy of the license and state significant changes.
Quickstart
pip install unstructured-ingest
from unstructured_ingest import ...
# See documentation for specific ingestion and transformation workflows
Verify before relying
- What document formats (PDF, Word, HTML, etc.) does the ingestion pipeline actually support?
- Does the package require external services or APIs, or does it run entirely locally?
- What is the typical performance or throughput for document ingestion and transformation?
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release <3.14,>=3.11 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 7 packagescertificlickijsonopentelemetry-sdkpydanticpython-dateutiltqdm |
| Maintenance | Actively maintained 0 days since the last release |
| First released | |
| Downloads | 620,054 / month, #5,727 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Topic :: Scientific/Engineering :: Artificial Intelligence |
Evidence: unstructured_ingest-1.9.3-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “ETL pipeline for document ingestion”
- unstructured-ingestUnstructured Ingest is a local ETL pipeline that prepares…
- dagster-airbyteIntegrates Airbyte data connectors with Dagster's orchestration…
- dagster-fivetranIntegrates Fivetran data connectors with Dagster's orchestration…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also ingestr · openmetadata-ingestion · unstructured · langchain-unstructured · graphrag · nv-ingest-client · unstructured-client · embedchain · azure-ai-contentunderstanding · humiolib