mineru-vl-utils
Utilities for MinerU Vision-Language models
Decision gist · record as of 2026-08-14
Yes. The package is actively maintained (released one day ago), has low install friction, carries a permissive Apache-2.0 license, and offers a clean abstraction over a capable multimodal model with flexible deployment options. Choose it if you need to extract structured content from documents and want to avoid managing model serving yourself; the http-client backend requires an external server, but other backends let you run inference locally.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- For http-client backend, you must have a separate vllm or LLM deployment tool serving the MinerU model as an HTTP server on the specified URL.
- Low friction: pure Python wheel with seven common runtime dependencies.
- Actively maintained with a release one day ago.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions, making it safe for most projects.
last release 2026-08-13 (1 days) · last repo commit 2026-08-13 · 136 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 198,484 downloads/mo, #9,729 on PyPI
Alternatives
Verify before relying
pip install mineru-vl-utils
from mineru_vl_utils import MinerUClient
client = MinerUClient(backend="http-client", server_url="http://127.0.0.1:8000")
extracted_blocks = client.two_step_extract(image)- Whether the package supports async operations across all seven backends or only specific ones.
- Performance characteristics and latency expectations for each backend type.
- Whether the http-client backend requires network connectivity or can work offline once the model is cached.
- Specific model size and memory requirements for each backend deployment mode.
What it is and what it does
mineru-vl-utils is a client library for the MinerU Vision-Language Model, a multimodal AI system that detects document layout and recognizes text, tables, equations, and images from visual input. It abstracts away the complexity of model serving and inference by providing a unified MinerUClient interface that works across seven different deployment modes: http-client (remote server), transformers (HuggingFace), mlx-engine (Apple Silicon), lmdeploy-engine, vllm-engine (synchronous), vllm-async-engine (asynchronous), and llama-cpp-engine (in-process). The model outputs structured ContentBlock objects containing block type, bounding box, rotation angle, and recognized content (text as strings, tables as HTML, equations as LaTeX).
You use it by instantiating a MinerUClient with your chosen backend, then calling two_step_extract() on images to get back a list of detected and recognized content blocks. The package handles all the marshalling between your code and the underlying model, whether that model runs locally via transformers, on Apple Silicon via mlx-engine, in a remote HTTP service, or in-process via llama-cpp-engine. Installation is modular: the base package includes http-client support, and optional extras pull in backend-specific dependencies.
Use it for
- Extract structured text, tables, and equations from scanned documents or PDFs for downstream processing.
- Detect and localize text regions in document images for document layout analysis.
- Convert document images to structured HTML or LaTeX for archival or republishing workflows.
- Run document understanding on Apple Silicon Macs without external servers using the mlx-engine backend.
- Build a document processing service using the http-client backend to call a centralized model server.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
The package is actively maintained (released one day ago), has low install friction, carries a permissive Apache-2.0 license, and offers a clean abstraction over a capable multimodal model with flexible deployment options. Choose it if you need to extract structured content from documents and want to avoid managing model serving yourself; the http-client backend requires an external server, but other backends let you run inference locally.
Install
mineru-vl-utils on PyPI
Before you install
Low friction: pure Python wheel with seven common runtime dependencies. Actively maintained with a release one day ago. Requires Python 3.10 or later.
For http-client backend, you must have a separate vllm or LLM deployment tool serving the MinerU model as an HTTP server on the specified URL.
License in practice
Apache-2.0 permissive license allows commercial and private use with minimal restrictions, making it safe for most projects.
Quickstart
pip install mineru-vl-utils
from mineru_vl_utils import MinerUClient
client = MinerUClient(backend="http-client", server_url="http://127.0.0.1:8000")
extracted_blocks = client.two_step_extract(image)
Verify before relying
- Whether the package supports async operations across all seven backends or only specific ones.
- Performance characteristics and latency expectations for each backend type.
- Whether the http-client backend requires network connectivity or can work offline once the model is cached.
- Specific model size and memory requirements for each backend deployment mode.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release <3.14,>=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 7 packageshttpxhttpx-retriesaiofilespillowpydanticlogurutqdm |
| Maintenance | Actively maintained 1 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 198,484 / month, #9,729 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Operating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: mineru_vl_utils-1.2.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “table and equation recognition”
- mineru-vl-utilsProvides a lightweight Python wrapper to interact with the MinerU…
- docling-ibm-modelsProvides AI models for table structure recognition and page layout…
- img2tableIdentifies and extracts tables from images and PDF files using…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also mineru · sgl-kernel · simple-lama-inpainting · vllm · llmcompressor · lm-eval · ell-ai · pytorch-tokenizers · unstructured.pytesseract · mlx-lm