colpali-engine
The code used to train and run inference with the ColPali architecture.
Decision gist · record as of 2026-08-14
Yes. Active maintenance, permissive MIT license, low install friction, no known vulnerabilities, and a focused scope for vision-based document retrieval make it a solid choice. Install if you need to search documents visually without OCR pipelines; skip if you only work with plain text or have existing OCR infrastructure.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >=3.10,<3.15 and PyTorch; Mac users with MPS and torch 2.6.0 should downgrade to torch 2.5.1 for ColQwen models.
- Low friction install with a pure Python wheel.
- Active maintenance with recent commits; last release 67 days ago.
License · maintenance · safety
permissive license (permissive) — MIT license (permissive) allows commercial and private use with minimal restrictions, making it suitable for both research and production deployment.
last release 2026-06-08 (67 days) · last repo commit 2026-08-03 · 2,736 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 154,293 downloads/mo, #10,856 on PyPI
Alternatives
Verify before relying
pip install colpali-engine
from colpali_engine.models import ColPali
from colpali_engine.utils import process_documents
model = ColPali.from_pretrained('vidore/colqwen2-v1.0')
embeddings = model.encode_documents(documents)- Whether the package includes full training utilities or primarily inference code
- Performance characteristics (throughput, latency) for typical document batch sizes
- GPU memory requirements for different model variants listed in the description
What it is and what it does
ColPali-engine is a PyTorch-based library for training and running inference with vision-language document retrieval models. It implements the ColPali architecture and variants (ColQwen, ColSmol, etc.) that convert document images into multi-vector embeddings using visual transformers, enabling efficient semantic search over documents without requiring separate OCR or layout recognition pipelines. The library depends on numpy, scipy, torch, torchvision, transformers, pillow, peft, and requests.
The package is designed for developers and researchers building document retrieval systems. It supports multiple pre-trained model variants with different performance-efficiency tradeoffs, from small models (256M parameters) to larger ones (4.5B+). The core approach follows ColBERT's late-interaction ranking method adapted to the visual domain, allowing both the textual and visual content (layout, charts, images) of documents to influence retrieval scoring.
Use it for
- Build a document search engine that retrieves pages from PDFs or scanned documents based on natural language queries without OCR
- Index and retrieve technical documentation, research papers, or forms where layout and visual structure matter for understanding
- Create a multilingual document retrieval system using models with support across multiple languages
- Fine-tune a pre-trained model on domain-specific documents using the training utilities and LoRA support
- Deploy efficient document ranking in production with optional fused MaxSim kernels for reduced memory usage
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Active maintenance, permissive MIT license, low install friction, no known vulnerabilities, and a focused scope for vision-based document retrieval make it a solid choice. Install if you need to search documents visually without OCR pipelines; skip if you only work with plain text or have existing OCR infrastructure.
Install
colpali-engine on PyPI
Before you install
Low friction install with a pure Python wheel. Active maintenance with recent commits; last release 67 days ago. Requires PyTorch and related deep learning dependencies, which are substantial but standard for this domain.
Requires Python >=3.10,<3.15 and PyTorch; Mac users with MPS and torch 2.6.0 should downgrade to torch 2.5.1 for ColQwen models.
License in practice
MIT license (permissive) allows commercial and private use with minimal restrictions, making it suitable for both research and production deployment.
Quickstart
pip install colpali-engine
from colpali_engine.models import ColPali
from colpali_engine.utils import process_documents
model = ColPali.from_pretrained('vidore/colqwen2-v1.0')
embeddings = model.encode_documents(documents)
Verify before relying
- Whether the package includes full training utilities or primarily inference code
- Performance characteristics (throughput, latency) for typical document batch sizes
- GPU memory requirements for different model variants listed in the description
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release <3.15,>=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 8 packagesnumpypeftpillowrequestsscipytorchtorchvisiontransformers |
| Maintenance | Actively maintained 67 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 154,293 / month, #10,856 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Topic :: Scientific/Engineering :: Artificial Intelligence |
Evidence: colpali_engine-0.3.17-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “document retrieval vision language models”
- colpali-engineColPali-engine provides training and inference code for…
- llama-index-multi-modal-llms-openaiIntegrates OpenAI's multi-modal language models with LlamaIndex,…
- paddleocrPaddleOCR extracts text, tables, and structured data from images and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also colbert-ai · FlagEmbedding · model2vec · fastembed · langchain-plaid · voyageai · sentence-transformers · FlashRank · gliner · sgl-kernel