$npx skillfedfor your agent

nv-ingest-client

Python client for the nv-ingest service

With conditionsPyPI Text ProcessingReleased Mar 202685.0K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — nv_ingest_client-26.3.0-py3-none-any.whl
v26.3.0 · released 2026-03-16 · Python >=3.11 · 12 runtime deps: build, charset-normalizer, click, fsspec, httpx, pydantic, pydantic-settings, requests

Yes, if you are building a document ingestion or data preparation pipeline that integrates with NVIDIA's nv-ingest microservice. The package has low install friction, active maintenance, permissive licensing, and no known vulnerabilities. Install only if you have a running nv-ingest microservice instance available and need Python-level control over job submission and task configuration; it is not a standalone tool.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires a running nv-ingest microservice instance accessible at the specified hostname and port (defaults to localhost:7670); Python 3.11 or later required.
  • Low install friction with a pure Python wheel distribution.
  • Active maintenance with recent releases; last commit 2026-08-14.

License · maintenance · safety

Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions beyond attribution and license inclusion.

last release 2026-03-16 (151 days) · last repo commit 2026-08-14 · 2,965 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 85,003 downloads/mo, #13,960 on PyPI

Verify before relying

pip install nv-ingest-client

from nv_ingest_client.client.client import NvIngestClient
from nv_ingest_client.primitives.jobs import JobSpec
from nv_ingest_client.primitives.tasks import ExtractTask

client = NvIngestClient(message_client_hostname="localhost", message_client_port=7670)
extract_task = ExtractTask(document_type="pdf", extract_text=True)
job_spec = JobSpec(payload={"data": "example"}, tasks=[extract_task])
response = client.submit_job(job_spec)
  • Whether the package works with Python versions newer than 3.11 (classifier only lists 3.11 explicitly)
  • Performance characteristics and throughput limits for large-scale document ingestion
  • Retry and error-handling behavior when the microservice is unavailable or slow to respond
Same gist for agents: .md · .json

What it is and what it does

NV-Ingest-Client is a Python library that acts as a client interface to NVIDIA's nv-ingest microservice, enabling programmatic submission and management of document processing jobs. It abstracts the complexity of communicating with the microservice by providing a high-level API for defining jobs, configuring extraction and splitting tasks, and submitting them for processing. The library includes task factories for common operations like text extraction from PDFs and document splitting by word, sentence, or passage boundaries, along with a command-line interface for direct terminal use.

The package is designed for workflows that require batch processing of large document collections—extracting text and images from PDFs, splitting documents into chunks for embedding or retrieval systems, and preparing data for downstream AI/ML pipelines. It depends on standard HTTP and data-handling libraries (httpx, requests, pydantic) to communicate with the microservice and manage job specifications, making it suitable for integration into data preparation and retrieval-augmented generation (RAG) systems.

Use it for

  • Extract text and images from PDF documents in bulk and submit them to a processing pipeline via the nv-ingest service.
  • Split large documents into smaller chunks with configurable overlap for use in vector databases or semantic search systems.
  • Automate document ingestion workflows by defining job specifications with multiple extraction and splitting tasks.
  • Build data preparation pipelines that transform raw documents into structured, embeddings-ready text chunks.
  • Monitor and manage the status of long-running document processing jobs submitted to a remote nv-ingest microservice.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are building a document ingestion or data preparation pipeline that integrates with NVIDIA's nv-ingest microservice.

The package has low install friction, active maintenance, permissive licensing, and no known vulnerabilities. Install only if you have a running nv-ingest microservice instance available and need Python-level control over job submission and task configuration; it is not a standalone tool.

Install

nv-ingest-client on PyPI

Before you install

Low install friction with a pure Python wheel distribution. Active maintenance with recent releases; last commit 2026-08-14. Requires Python 3.11 or later. Depends on 12 runtime packages including pydantic, httpx, and lancedb, all widely available.

Requires a running nv-ingest microservice instance accessible at the specified hostname and port (defaults to localhost:7670); Python 3.11 or later required.

License in practice

Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions beyond attribution and license inclusion.

Quickstart

pip install nv-ingest-client

from nv_ingest_client.client.client import NvIngestClient
from nv_ingest_client.primitives.jobs import JobSpec
from nv_ingest_client.primitives.tasks import ExtractTask

client = NvIngestClient(message_client_hostname="localhost", message_client_port=7670)
extract_task = ExtractTask(document_type="pdf", extract_text=True)
job_spec = JobSpec(payload={"data": "example"}, tasks=[extract_task])
response = client.submit_job(job_spec)

Verify before relying

  • Whether the package works with Python versions newer than 3.11 (classifier only lists 3.11 explicitly)
  • Performance characteristics and throughput limits for large-scale document ingestion
  • Retry and error-handling behavior when the microservice is unavailable or slow to respond

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.11
Install frictionLow. Pure-Python wheel
Runtime dependencies
12 packages
buildcharset-normalizerclickfsspechttpxpydanticpydantic-settingsrequestsurllib3setuptoolstqdmlancedb
MaintenanceActively maintained 151 days since the last release
Last repo commit
First released
Downloads85,003 / month, #13,960 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Operating System :: OS IndependentProgramming Language :: Python :: 3.11

Evidence: nv_ingest_client-26.3.0-py3-none-any.whl

Tags

Capabilities
document ingestion clientnvidia nv-ingest pythonbatch document processingpdf extraction and splittingdata preparation pipelinemicroservice job submissiontext extraction client
Topics
document-processingnvidia-ecosystemdata-preparation

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “nvidia nv-ingest python”

Give your agent the search over MCP, or paste the wish link into any chat.

More Text Processing packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyparsing Worth it
PyPI · Text Processing · released Jan 2026

pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.

Install it if you need to parse text or define grammars programmatically.

MITpure Python · 3.9+
412.7Mdownloads / mo
fonttools Worth it
PyPI · Text Processing · released May 2026

fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.

Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.

permissive licensepure Python · 3.10+
235.9Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
RapidFuzz Worth it
PyPI · Text Processing · released Apr 2026

RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.

Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.

MITcompiled wheel · 3.10+
184.2Mdownloads / mo
tinycss2 Worth it
PyPI · Text Processing · released Nov 2025

tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.

Install it if your project requires CSS tokenization or syntax manipulation.

BSD-3-Clausepure Python · 3.10+
113.2Mdownloads / mo

See also nvidia-nat-eval · nvidia-nat-atif · unstructured-ingest · cuda-bindings · nvidia-nat-mcp · nvidia-nat-langchain · groundx · wmill · pydiscourse · llama-index-indices-managed-llama-cloud