nv-ingest-client
Python client for the nv-ingest service
Decision gist · record as of 2026-08-14
Yes, if you are building a document ingestion or data preparation pipeline that integrates with NVIDIA's nv-ingest microservice. The package has low install friction, active maintenance, permissive licensing, and no known vulnerabilities. Install only if you have a running nv-ingest microservice instance available and need Python-level control over job submission and task configuration; it is not a standalone tool.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires a running nv-ingest microservice instance accessible at the specified hostname and port (defaults to localhost:7670); Python 3.11 or later required.
- Low install friction with a pure Python wheel distribution.
- Active maintenance with recent releases; last commit 2026-08-14.
License · maintenance · safety
Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions beyond attribution and license inclusion.
last release 2026-03-16 (151 days) · last repo commit 2026-08-14 · 2,965 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 85,003 downloads/mo, #13,960 on PyPI
Alternatives
Verify before relying
pip install nv-ingest-client
from nv_ingest_client.client.client import NvIngestClient
from nv_ingest_client.primitives.jobs import JobSpec
from nv_ingest_client.primitives.tasks import ExtractTask
client = NvIngestClient(message_client_hostname="localhost", message_client_port=7670)
extract_task = ExtractTask(document_type="pdf", extract_text=True)
job_spec = JobSpec(payload={"data": "example"}, tasks=[extract_task])
response = client.submit_job(job_spec)- Whether the package works with Python versions newer than 3.11 (classifier only lists 3.11 explicitly)
- Performance characteristics and throughput limits for large-scale document ingestion
- Retry and error-handling behavior when the microservice is unavailable or slow to respond
What it is and what it does
NV-Ingest-Client is a Python library that acts as a client interface to NVIDIA's nv-ingest microservice, enabling programmatic submission and management of document processing jobs. It abstracts the complexity of communicating with the microservice by providing a high-level API for defining jobs, configuring extraction and splitting tasks, and submitting them for processing. The library includes task factories for common operations like text extraction from PDFs and document splitting by word, sentence, or passage boundaries, along with a command-line interface for direct terminal use.
The package is designed for workflows that require batch processing of large document collections—extracting text and images from PDFs, splitting documents into chunks for embedding or retrieval systems, and preparing data for downstream AI/ML pipelines. It depends on standard HTTP and data-handling libraries (httpx, requests, pydantic) to communicate with the microservice and manage job specifications, making it suitable for integration into data preparation and retrieval-augmented generation (RAG) systems.
Use it for
- Extract text and images from PDF documents in bulk and submit them to a processing pipeline via the nv-ingest service.
- Split large documents into smaller chunks with configurable overlap for use in vector databases or semantic search systems.
- Automate document ingestion workflows by defining job specifications with multiple extraction and splitting tasks.
- Build data preparation pipelines that transform raw documents into structured, embeddings-ready text chunks.
- Monitor and manage the status of long-running document processing jobs submitted to a remote nv-ingest microservice.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are building a document ingestion or data preparation pipeline that integrates with NVIDIA's nv-ingest microservice.
The package has low install friction, active maintenance, permissive licensing, and no known vulnerabilities. Install only if you have a running nv-ingest microservice instance available and need Python-level control over job submission and task configuration; it is not a standalone tool.
Install
nv-ingest-client on PyPI
Before you install
Low install friction with a pure Python wheel distribution. Active maintenance with recent releases; last commit 2026-08-14. Requires Python 3.11 or later. Depends on 12 runtime packages including pydantic, httpx, and lancedb, all widely available.
Requires a running nv-ingest microservice instance accessible at the specified hostname and port (defaults to localhost:7670); Python 3.11 or later required.
License in practice
Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions beyond attribution and license inclusion.
Quickstart
pip install nv-ingest-client
from nv_ingest_client.client.client import NvIngestClient
from nv_ingest_client.primitives.jobs import JobSpec
from nv_ingest_client.primitives.tasks import ExtractTask
client = NvIngestClient(message_client_hostname="localhost", message_client_port=7670)
extract_task = ExtractTask(document_type="pdf", extract_text=True)
job_spec = JobSpec(payload={"data": "example"}, tasks=[extract_task])
response = client.submit_job(job_spec)
Verify before relying
- Whether the package works with Python versions newer than 3.11 (classifier only lists 3.11 explicitly)
- Performance characteristics and throughput limits for large-scale document ingestion
- Retry and error-handling behavior when the microservice is unavailable or slow to respond
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.11 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 12 packagesbuildcharset-normalizerclickfsspechttpxpydanticpydantic-settingsrequestsurllib3setuptoolstqdmlancedb |
| Maintenance | Actively maintained 151 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 85,003 / month, #13,960 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Operating System :: OS IndependentProgramming Language :: Python :: 3.11 |
Evidence: nv_ingest_client-26.3.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “nvidia nv-ingest python”
- nv-ingest-clientPython client library for submitting and managing document ingestion…
- nvidia-cusparselt-cu12Provides NVIDIA's cuSPARSELt CUDA library for high-performance sparse…
- nvidia-cuda-ccclProvides NVIDIA CUDA C++ Core Compute Libraries (CCCL) for…
Give your agent the search over MCP, or paste the wish link into any chat.
More Text Processing packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.
Install it if you need to parse text or define grammars programmatically.
fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.
Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.
Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.
tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.
Install it if your project requires CSS tokenization or syntax manipulation.
See also nvidia-nat-eval · nvidia-nat-atif · unstructured-ingest · cuda-bindings · nvidia-nat-mcp · nvidia-nat-langchain · groundx · wmill · pydiscourse · llama-index-indices-managed-llama-cloud